I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 stable release that I felt truly comfortable about, since I try to keep to SemVer a
Release: llm-coding-agent 0.1a0 Another Fable 5 experiment. Now that my LLM library has evolved into more of an agent framework it's time to see what a simple coding agent would look like built on it. I started a new Python library using my python-lib-template-repository GitHub t
Simon Willison explored using DSPy to evaluate and refine the SQL system prompts within Datasette Agent. This research aimed to improve how Datasette Agent generates and executes read-only SQL queries to answer user questions. The process involved installing Datasette alpha, datasette-agent, and dspy, then leveraging dspy for prompt optimization.
→ This is a high-signal item for Mark, demonstrating a practical application of DSPy for RAG/retrieval improvement, directly relevant to his interest in tooling and prompt engineering for local LLMs.
Hugging Face and Cerebras have collaborated to optimize Gemma 4 for real-time voice AI applications. This partnership focuses on making Gemma 4 more efficient for speech-to-text tasks, leveraging Cerebras's hardware for faster inference. The goal is to enable more responsive and accessible voice AI experiences.
→ Gemma 4 + real-time voice AI is a high-signal combo, directly relevant to local LLM performance and STT improvements for accessibility and on-device use cases.
LiteRT-LM version 2.1.6 has been released, featuring refactored CC APIs that are now header-only and do not require linking Abseil. This update also includes the release of ARMv7 prebuilts and an expanded Accelerator Test Suite covering 43 operations across f16 and f32 data types. Additionally, built-in kernels have been improved to support high-dimensional tensors.
→ ARMv7 prebuilts for LiteRT-LM are a big win for on-device AI, making it easier to deploy local LLMs on a wider range of mobile and embedded hardware.
Hugging Face Kernels, a free Jupyter environment, has received significant updates including a new UI, improved performance, and enhanced collaboration features. Users can now access more powerful GPUs, increased storage, and better integration with Hugging Face Hub for datasets and models. These improvements aim to make it easier for researchers and developers to experiment with and share AI models directly within the platform.
→ This is a major upgrade for anyone working with open-source models, especially for Gemma 4 contest participants who need powerful, free compute and seamless integration with the Hugging Face ecosystem.
Hugging Face has integrated "Every Eval Ever" results directly onto model pages, providing comprehensive performance benchmarks for various models. This new feature allows users to easily compare models across a wide range of tasks and datasets, enhancing transparency and informed decision-making for model selection. The integration aims to standardize evaluation reporting and make it more accessible to the community.
→ This is a significant win for open-source model transparency, making it easier to quickly assess and compare local LLMs like Gemma, Llama, and Mistral for specific use cases, directly impacting your Kaggle Gemma 4 contest strategy.
Simon Willison explored using DSPy to evaluate and refine the SQL system prompts within Datasette Agent. This research aimed to improve how Datasette Agent generates and executes read-only SQL queries to answer user questions. The process involved installing Datasette alpha, datasette-agent, and dspy, then leveraging dspy for prompt optimization.
→ This is a high-signal item for Mark, demonstrating a practical application of DSPy for RAG/retrieval improvement, directly relevant to his interest in tooling and prompt engineering for local LLMs.
I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 stable release that I felt truly comfortable about, since I try to keep to SemVer a
Release: llm-coding-agent 0.1a0 Another Fable 5 experiment. Now that my LLM library has evolved into more of an agent framework it's time to see what a simple coding agent would look like built on it. I started a new Python library using my python-lib-template-repository GitHub t
Simon Willison explored using DSPy to evaluate and refine the SQL system prompts within Datasette Agent. This research aimed to improve how Datasette Agent generates and executes read-only SQL queries to answer user questions. The process involved installing Datasette alpha, datasette-agent, and dspy, then leveraging dspy for prompt optimization.
→ This is a high-signal item for Mark, demonstrating a practical application of DSPy for RAG/retrieval improvement, directly relevant to his interest in tooling and prompt engineering for local LLMs.
Nano Banana 2 Lite Also known as Gemini 3.1 Flash Lite Image ( gemini-3.1-flash-lite-image in their API ), this is the "fastest and cheapest Gemini image model, engineered for velocity and scale". I used AI studio to run this prompt: Do a where's Waldo style image but it's where
shot-scraper video is a new command introduced in today's shot-scraper 1.10 release which accepts a storyboard.yml file defining a routine to run against a web application and uses Playwright to record a video of that routine. I've written before about the importance of having co
Tiniest TIL, using AppleScript to count the number of open browser tabs in Safari: osascript -e 'tell application "Safari" to count tabs of every window' Tags: safari , til , applescript
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding This is an interesting new open weights (MIT licensed) model, the first model release from DeepReinforce. [...] with variants including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. Built on top of pretrained Gemma 4 and Qwen 3.5
Hugging Face Kernels, a free Jupyter environment, has received significant updates including a new UI, improved performance, and enhanced collaboration features. Users can now access more powerful GPUs, increased storage, and better integration with Hugging Face Hub for datasets and models. These improvements aim to make it easier for researchers and developers to experiment with and share AI models directly within the platform.
→ This is a major upgrade for anyone working with open-source models, especially for Gemma 4 contest participants who need powerful, free compute and seamless integration with the Hugging Face ecosystem.
Hugging Face and Cerebras have collaborated to optimize Gemma 4 for real-time voice AI applications. This partnership focuses on making Gemma 4 more efficient for speech-to-text tasks, leveraging Cerebras's hardware for faster inference. The goal is to enable more responsive and accessible voice AI experiences.
→ Gemma 4 + real-time voice AI is a high-signal combo, directly relevant to local LLM performance and STT improvements for accessibility and on-device use cases.
Hugging Face has integrated "Every Eval Ever" results directly onto model pages, providing comprehensive performance benchmarks for various models. This new feature allows users to easily compare models across a wide range of tasks and datasets, enhancing transparency and informed decision-making for model selection. The integration aims to standardize evaluation reporting and make it more accessible to the community.
→ This is a significant win for open-source model transparency, making it easier to quickly assess and compare local LLMs like Gemma, Llama, and Mistral for specific use cases, directly impacting your Kaggle Gemma 4 contest strategy.
LiteRT-LM version 2.1.6 has been released, featuring refactored CC APIs that are now header-only and do not require linking Abseil. This update also includes the release of ARMv7 prebuilts and an expanded Accelerator Test Suite covering 43 operations across f16 and f32 data types. Additionally, built-in kernels have been improved to support high-dimensional tensors.
→ ARMv7 prebuilts for LiteRT-LM are a big win for on-device AI, making it easier to deploy local LLMs on a wider range of mobile and embedded hardware.
Z.ai's GLM-5.2 is a 1 million token, MIT-licensed open weight model that costs a fraction of frontier AI prices, and I put it through real tests to show you exactly where it holds up and where it doesn't. Also join my newsletter: https://futuretools.io/newsletter I cover the thre
Building a World Map with only 500 bytes Iwo Kadziela (assisted by Codex) figured out a way to generate a credible ASCII world map using 445 bytes of data: The key trick is to use deflate compression, which is then wired together using this neat snippet of JavaScript. I didn't kn
One of the most interesting tips I got from the Fireside Chat I hosted with Cat Wu and Thariq Shihipar from the Claude Code team at AIE on Wednesday was to let Fable (and to a certain extent Opus) use their own judgement rather than dictating how they should work. The example the
I saw Geoffrey Litt speak at AIE yesterday, and one framing he used particularly resonated with me: Understand to participate Geoffrey was talking about the challenge of collaborating with coding agents as they construct increasingly large and sophisticated changes, and the need
We’ve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. We'll begin restoring access tomorrow, and will share an update soon. — Anthropic , on Twitter Tags: anthropic , claude , generative-ai , claude-mythos-fable , ai , ll
Vercel’s Andrew Qu on the AIEWF expo floor. Andrew Qu is Chief of Software at Vercel, where he works with the CTO across internal engineering, product experimentation and emerging technologies. He has built libraries for MCP, created skills.sh and led the development of eve, Verc
Impeccable’s Paul Bakaus at the AI Engineer World’s Fair. Paul Bakaus thinks the emerging discipline of “skill engineering” can make AI agents more capable — but he absolutely does not want to remove people from the creative process. He chats to Latent Space about his approach to
In separate announcements, Sonnet 5 was released today, and Fable/Mythos 5 were approved to be released again after some work with the government. The primary discussion around Sonnet 5’s efficiency was a damper on the excitement, driven by tokenizer changes and 3-6x more turn ta
Ahmad Osman at the AI Engineer World’s Fair today. Ahmad Osman has been advocating for local AI — running models on your own computer, workstation or dedicated hardware — long before it became a major theme at this year’s AI Engineer World’s Fair . He is also the founder of Osman
OpenAI and other labs are racing to slash inference costs, prompting debate about token-efficiency techniques and quality trade-offs. Anthropic regained global access to Fable 5 after export controls were lifted, triggering discussions about jailbreaking, tightened guardrails, an
The Commerce Department authorized limited, vetted access to Anthropic's Mythos and signaled a de facto licensing regime for frontier models. OpenAI previewed GPT‑5.6 (Sol, Terra, Luna) under a trusted‑partner rollout, igniting debate over benchmarks, safety measures, and governm
Release: sqlite-utils 4.0rc3 I hoped to release sqlite-utils 4.0 stable this weekend, but as I worked through the backlog of issues and PRs with a combination of Claude Fable 5 and GPT-5.5 the changelog since rc2 kept getting bigger . The biggest new feature is support for intros
Better Models: Worse Tools Armin reports on a weird problem he ran into while hacking on Pi: The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in the nested edits[] array. And not Haiku or some small model: Opus 4.8. The edit
Open Source AI Gap Map Current AI is "a global partnership building a public option for AI", founded as a non-profit at the AI Action Summit in Paris in February 2025 and backed by serious capital ($400m already committed). They launched their Gap Map a couple of days ago - an at
The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here . This month: Claude Fable 5, GPT-5.6, and US export restrictions GLM-5.2 is the new best open weights model Tokenmaxxing is so over Dat
What's new in Claude Sonnet 5 Claude Sonnet 5 came out this morning . I always head straight for the "what's new" developer docs because they tend to have more actionable information than the official announcement post. Anthropic say of Sonnet 5 that "its performance is close to
Release: shot-scraper 1.10 The big new feature is shot-scraper video storyboard.yml , described in detail in Have your agent record video demos of its work with shot-scraper video . Tags: shot-scraper
Tool: HTML table extractor Yet another in my growing collection of paste-conversion tools. This one accepts pasted rich text from browsers (with embedded HTML tables) and converts every detected table into HTML, Markdown, CSV, TSV, or JSON. Try it out by selecting everything on t
Google, the New York Jobs CEO Council and Urban Assembly hosted an AI summit for 150 education and industry leaders.
One of the highlights of the final day of the AI Engineer World’s Fair was a debate about loops. It nicely captured an argument running through the whole conference: are autonomous software factories viable now, or is the engineering discipline lagging behind the ambition? Allie
Adobe Principal Scientist Carlos Sanchez at AIEWF. For as long as I can remember (and I managed websites in the dot-com period), “personalization” has been a holy grail for websites. But up till now, that’s typically meant selecting from a predefined set of options. A retailer mi
Fable was relaunched on schedule, and AIE was on top of it with the first Field Guide to Fable talk , as well as the rest of the excellent coverage of AIEWF Day 3 across Autoresearch , Cursor FDE , and a followup to Zach Lloyd’s popular talk yesterday on Software Factories , as w
“You can’t one-shot design.” Paul Bakaus at AIEWF today. Wednesday was autoresearch day on the AI Engineer World’s Fair main stage. Autoresearch is — you guessed it — a kind of loop. Introspection co-founder Roland Gavrilescu explained it best in an interview with Latent Space th
Introspection’s Roland Gavrilescu at AIEWF. We’ve heard a lot about loops at the AI Engineer World’s Fair this week. Another buzzword is autoresearch , which involves building an “outer loop” where agents help maintain and improve the primary system, using feedback signals, evals
This episode has a fun personal twist: There’s a counterfactual world where I was employee #1 at Genesis Molecular AI , 1 the company behind today’s episode. A certain introduction happened a few weeks too late and I had already happily signed at Atomwise 2 , another ML-for-drug-
Warp founder Zach Lloyd in the AI Engineer World’s Fair expo hall. I’ve been covering Warp for a couple of years now, and its rapid evolution from a command-line interface tool to a software factory platform has been fascinating to watch. The company began in the pre-ChatGPT days
The market seems to think Anthropic just won the AI war against OpenAI. But what's the real story? Everyone is talking about Anthropic's massive $965B valuation filing compared to OpenAI's $852B. And if we look at the business fundamentals, Anthropic brought the receipts: 💰 $30B
Alleged credit-based, identity-verified access for Anthropic Fable sparks privacy and policy debate. Senate draft targets agent neutrality and a duty-of-loyalty while state and corporate contracts shift pricing and access. Exponential View estimates a $175 billion annual AI run r
I just launched my third course, Whimsical Animations, and so far, it’s on track to sell roughly ⅓ as many copies as a typical course launch. It’s a similar story with my two existing courses. Sales are down significantly from last year. There are likely a lot of reasons for this
The AI Compass This political compass style quiz by bambamramfan is pretty neat - answer 29 questions about AI and AI ethics to see which of the 30 archetypes you best fit. I'm impressed that my answers on my first time through the quiz categorized me as "The Garage Tinkerer", pa
Google UK shares its latest Economic Impact Report and how to enable more people to unlock the benefits of AI-powered technologies.
A Google expert explains what it means to take a full-stack approach to AI and why it’s been the foundation of our AI work for so long.
New OpenAI Signals data shows how ChatGPT adoption is growing globally, with users increasing usage, exploring more capabilities, and driving growth across regions and languages.
Pauline Brunet, VP of Forward Deployed Engineering at Cursor, at AIEWF. Forward deployed engineering has quickly become one of the most prominent roles in enterprise AI. Sitting somewhere between software engineering, product development and customer implementation, forward deplo
Agents are here to serve you in the software factory. Loops, loops and more loops. That word, loop, dominated conversations on day 2 of the AI Engineer World’s Fair — the first full day of keynotes and sessions. Perhaps knowing in advance what everyone would be talking about, AIE
BuseyBench is the best AI Benchmark Discover More: 🛠️ Explore AI Tools & News: https://futuretools.io/ 📰 Weekly Newsletter: https://futuretools.io/newsletter Socials: ❌ Twiter/X: https://x.com/mreflow 🖼️ Instagram: https://instagram.com/mr.eflow 🧵 Threads: https://www.threads.net
The future is here, and it’s ready to take your crap 😂 This is a self driving toilet called Xiaoban made by Chinese company Yueban that just debuted at a Shanghai aged-care expo. You just summon it to your bedside and it has a built-in bidet. When you’re done, it autonomously dri
This is the biggest mistake I see people making with AI. We’ve all fallen into the trap of trying to use AI for EVERYTHING. But the difference between an AI novice and a power user is knowing when to step back. The goal isn't to automate your entire life. It's to know which tasks
How to generate motion graphics using AI: Step 1: Choose your AI assistant These frameworks act as plugins or "skills" that work directly inside AI coding assistants. In the video, I used Claude Code, but they also work with VS Code, GitHub Copilot, Codex, Cowork, and more. Step
A new OpenAI report maps how AI could reshape jobs across the EU, highlighting which occupations may face automation, growth, or workflow changes.
Here's the AI News you probably missed this week. Try @GensparkProduct with free credits by registering at this link: https://www.genspark.ai/?utm_source=yt&utm_campaign=mattwolfe02 #Genspark #WorkWithGenspark Discover More: 🛠️ Explore AI Tools & News: https://futuretools.io/ 📰 W
tp3_memories_localgemma3:4b)