← Reports

AI Radar

Monday, June 22, 2026 · auto-generated by tp3_ai_radar.py

Top 3 today

MosaicLeaks: Can your research agent keep a secret?

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face researchers introduced "MosaicLeaks," a novel red-teaming benchmark designed to evaluate the privacy leakage of research agents. This benchmark assesses how well agents can protect sensitive information when interacting with various tools and environments. The study highlights potential vulnerabilities in current agent designs regarding data privacy.

This is a critical signal for anyone building RAG or local LLM agents, especially for education or accessibility tools where data privacy is paramount. It's a must-read for the Gemma 4 contest.

research agents privacy leakage red-teaming rag local llms

Beyond LoRA: Can you beat the most popular fine-tuning technique?

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face explores fine-tuning methods beyond LoRA, evaluating techniques like QLoRA, DoRA, and ReFT. The article details the theoretical underpinnings and practical applications of these methods, comparing their efficiency and performance against traditional LoRA. It provides insights into when and why one might choose an alternative fine-tuning approach for large language models.

This deep dive into advanced fine-tuning techniques is crucial for anyone pushing the limits of local LLMs like Gemma, offering potential performance gains for the Gemma 4 hackathon.

fine-tuning lora qlora reft local llms

Is it agentic enough? Benchmarking open models on your own tooling

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face introduced a new benchmark for evaluating the agentic capabilities of open-source models, focusing on their ability to use external tools and APIs. The benchmark allows developers to test models against their own custom tooling, providing a more practical assessment of real-world performance. This initiative aims to bridge the gap between theoretical model capabilities and their effectiveness in complex, tool-augmented workflows.

This is a high-signal item for Mark, as it directly addresses the practical application of local LLMs like Gemma and Llama in tool-use scenarios, crucial for his Bidet and TP3 projects.

agentic ai open models tool use benchmarking hugging face

By topic

local llms (2)

MosaicLeaks: Can your research agent keep a secret?

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face researchers introduced "MosaicLeaks," a novel red-teaming benchmark designed to evaluate the privacy leakage of research agents. This benchmark assesses how well agents can protect sensitive information when interacting with various tools and environments. The study highlights potential vulnerabilities in current agent designs regarding data privacy.

This is a critical signal for anyone building RAG or local LLM agents, especially for education or accessibility tools where data privacy is paramount. It's a must-read for the Gemma 4 contest.

research agents privacy leakage red-teaming rag local llms

Beyond LoRA: Can you beat the most popular fine-tuning technique?

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face explores fine-tuning methods beyond LoRA, evaluating techniques like QLoRA, DoRA, and ReFT. The article details the theoretical underpinnings and practical applications of these methods, comparing their efficiency and performance against traditional LoRA. It provides insights into when and why one might choose an alternative fine-tuning approach for large language models.

This deep dive into advanced fine-tuning techniques is crucial for anyone pushing the limits of local LLMs like Gemma, offering potential performance gains for the Gemma 4 hackathon.

fine-tuning lora qlora reft local llms

other (3)

Temporary Cloudflare Accounts for AI agents

8/10

Simon Willison local LLM · Jun 21 · simonwillison.net

Cloudflare has introduced temporary accounts for AI agents, allowing users to deploy Cloudflare Workers projects for 60 minutes without needing to create a full account. This feature is accessible via `npx wrangler deploy --temporary` and creates an ephemeral project. While marketed for AI agents, it offers a quick deployment solution for any user.

This is a neat dev tool for quick testing, especially for AI agents or small utility functions, without the friction of account creation.

cloudflare workers ephemeral deployments developer tools ai agents

datasette-acl 0.6a0

8/10

Simon Willison local LLM · Jun 18 · simonwillison.net

Datasette-acl 0.6a0 has been released, expanding its capabilities from table-only permissions to a more general resource-sharing system. This update allows for finely-grained control over resource access within multi-user Datasette instances. Alex Garcia led the development for this release.

This Datasette plugin update offers improved access control, which is crucial for managing data access in multi-user environments, potentially relevant for collaborative AI development or data analysis projects.

datasette access control resource sharing multi-user systems

Is it agentic enough? Benchmarking open models on your own tooling

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face introduced a new benchmark for evaluating the agentic capabilities of open-source models, focusing on their ability to use external tools and APIs. The benchmark allows developers to test models against their own custom tooling, providing a more practical assessment of real-world performance. This initiative aims to bridge the gap between theoretical model capabilities and their effectiveness in complex, tool-augmented workflows.

This is a high-signal item for Mark, as it directly addresses the practical application of local LLMs like Gemma and Llama in tool-use scenarios, crucial for his Bidet and TP3 projects.

agentic ai open models tool use benchmarking hugging face

Full ranked list

MosaicLeaks: Can your research agent keep a secret?

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face researchers introduced "MosaicLeaks," a novel red-teaming benchmark designed to evaluate the privacy leakage of research agents. This benchmark assesses how well agents can protect sensitive information when interacting with various tools and environments. The study highlights potential vulnerabilities in current agent designs regarding data privacy.

This is a critical signal for anyone building RAG or local LLM agents, especially for education or accessibility tools where data privacy is paramount. It's a must-read for the Gemma 4 contest.

research agents privacy leakage red-teaming rag local llms

Beyond LoRA: Can you beat the most popular fine-tuning technique?

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face explores fine-tuning methods beyond LoRA, evaluating techniques like QLoRA, DoRA, and ReFT. The article details the theoretical underpinnings and practical applications of these methods, comparing their efficiency and performance against traditional LoRA. It provides insights into when and why one might choose an alternative fine-tuning approach for large language models.

This deep dive into advanced fine-tuning techniques is crucial for anyone pushing the limits of local LLMs like Gemma, offering potential performance gains for the Gemma 4 hackathon.

fine-tuning lora qlora reft local llms

Is it agentic enough? Benchmarking open models on your own tooling

9/10

Hugging Face blog local LLM · Jun 18 · huggingface.co

Hugging Face introduced a new benchmark for evaluating the agentic capabilities of open-source models, focusing on their ability to use external tools and APIs. The benchmark allows developers to test models against their own custom tooling, providing a more practical assessment of real-world performance. This initiative aims to bridge the gap between theoretical model capabilities and their effectiveness in complex, tool-augmented workflows.

This is a high-signal item for Mark, as it directly addresses the practical application of local LLMs like Gemma and Llama in tool-use scenarios, crucial for his Bidet and TP3 projects.

agentic ai open models tool use benchmarking hugging face

Temporary Cloudflare Accounts for AI agents

8/10

Simon Willison local LLM · Jun 21 · simonwillison.net

Cloudflare has introduced temporary accounts for AI agents, allowing users to deploy Cloudflare Workers projects for 60 minutes without needing to create a full account. This feature is accessible via `npx wrangler deploy --temporary` and creates an ephemeral project. While marketed for AI agents, it offers a quick deployment solution for any user.

This is a neat dev tool for quick testing, especially for AI agents or small utility functions, without the friction of account creation.

cloudflare workers ephemeral deployments developer tools ai agents

datasette-acl 0.6a0

8/10

Simon Willison local LLM · Jun 18 · simonwillison.net

Datasette-acl 0.6a0 has been released, expanding its capabilities from table-only permissions to a more general resource-sharing system. This update allows for finely-grained control over resource access within multi-user Datasette instances. Alex Garcia led the development for this release.

This Datasette plugin update offers improved access control, which is crucial for managing data access in multi-user environments, potentially relevant for collaborative AI development or data analysis projects.

datasette access control resource sharing multi-user systems

Improving health intelligence in ChatGPT

8/10

OpenAI news frontier · Jun 18 · openai.com

Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clearer communication, and physician-informed evaluations.

Using AI to help physicians diagnose rare genetic diseases affecting children

8/10

OpenAI news frontier · Jun 18 · openai.com

Researchers used an OpenAI reasoning model to help diagnose rare diseases, identifying 18 new diagnoses in previously unsolved cases.

The Professor of Outputmaxxing — Anjney Midha, AMP

8/10

Latent Space commentary · Jun 18 · www.latent.space

Last 4 days before regular tickets sell out at AI Engineer World’s Fair - this is the single biggest gathering of AI Engineers, Founders, Leaders, and Researchers in the world. Attendees get >$5000 worth of sponsor credits and talk tracks are looking FANTASTIC. Join us! The AI sc

The 5-Minute AI Weekly Recap: Realignment Week

8/10

AI Daily Brief YouTube commentary · Jun 21 · www.youtube.com

This week, the Fable fallout became a broader realignment across AI, pushing more attention toward open models, model routing, local control, and the risks of building around any single frontier system. GLM 5.2, OpenRouter’s Fusion, SpaceX’s Cursor acquisition, and Europe’s AI so

The Models Trying to Replace Fable

8/10

AI Daily Brief YouTube commentary · Jun 19 · www.youtube.com

G7 talks exposed geopolitical tension over access to US frontier models after the Anthropic Fable shutdown. Open-source and smaller efficient models: GLM 5.2, Kimi 2.7, Vibe Thinker, and Cursor Composer 2.5, are driving moves toward local hosting and lower-cost inference. Model p

Quoting Sean Lynch

7/10

Simon Willison local LLM · Jun 19 · simonwillison.net

The real valuable capability MCP offers over skills/CLI is isolating the auth flow outside of the agent’s context window, and potentially out of the harness completely. [...] Maybe the idealized form of MCP is just an auth gateway for the API and nothing else. That’d still be a w

datasette-apps 0.1a2

7/10

Simon Willison local LLM · Jun 15 · simonwillison.net

Release: datasette-apps 0.1a2 Custom network/CSP origins for apps are now guarded by a new apps-set-csp permission, with an optional allowed_csp_origins plugin allow-list for non-privileged users. The Datasette Agent app creation tool enforces the same rules. #24 Stored query pic

Samsung Electronics brings ChatGPT and Codex to employees

7/10

OpenAI news frontier · Jun 21 · openai.com

Samsung Electronics deploys ChatGPT Enterprise and Codex to employees worldwide, marking one of OpenAI’s largest enterprise AI rollouts.

[AINews] not much happened today

7/10

Latent Space commentary · Jun 20 · www.latent.space

GLM 5.2 is still trending very hard, but you knew that already . Regular Tickets for AIE WF 2026 will sell out by Monday. If you’re a Latent Space subscriber ($80 a year), a limited-time only $250 discount for select ticket classes is included below for the AIE-curious who have n

Datasette Apps: Host custom HTML applications inside Datasette

6/10

Simon Willison local LLM · Jun 18 · simonwillison.net

Today we launched a new plugin for Datasette, datasette-apps , with this launch announcement post on the Datasette project blog. That post has the what , but I'm going to expand on that a little bit here to provide the why . The TL;DR Datasette Apps are self-contained HTML+JavaSc

AI News: Fable Banned, New Open-Source Leader, Midjourney Shocker

6/10

Matt Wolfe YouTube commentary · Jun 19 · www.youtube.com

Here's The AI News You Probably Missed This Week. Learn more about how Box AI can unlock key insights for your business here: https://www.box.com/ai?utm_source=youtube&utm_medium=paidinfluencer&utm_theme=icm&utm_campaign=FY27_Q2_MattWolfe_June19 Discover More: 🛠️ Explore AI Tools

Your Company Doesn't Need an AI Strategy

6/10

AI Daily Brief YouTube commentary · Jun 20 · www.youtube.com

...it needs an AI learning system. This episode argues that the Fable 5 disruption exposed a deeper enterprise problem: companies can’t treat AI as a vendor strategy. The real advantage will come from building learning systems that capture institutional judgment, workflow traces,

Why Only AI Training Can Save the Economy

5/10

AI Daily Brief YouTube commentary · Jun 18 · www.youtube.com

Large-scale AI training offers the only practical path to align AI labs' token-driven revenue growth with enterprise spending constraints. A shift from seat-based subscriptions to agentic, usage-based consumption is driving surging token demand, massive infrastructure investment,

Pipeline status

Pipeline stats

Dead sources today

All sources green today.