AI Field Notes by Michael Nemtsev

AI Inference Speedup | AI Field Notes #102

A printing press floods a desk with pages faster than a lone inspector can read them, while money fuels the speed and a rival arm copies the output across a border.

AI inference got faster and cheaper this week while AI coding funding broke records. Inception's Mercury 2.5, a diffusion model, now writes 1,107 tokens a second for four cents a million at launch, and OpenBMB put a capable model on phones under an open license. Cognition, maker of the Devin coding agent, raised $2 billion at a $48 billion valuation, and Mistral pulled in €3 billion, Europe's largest tech round. Underneath the launches, OpenAI's own Astra system card admits the model is harder to monitor, and US agencies accused six Chinese firms of siphoning American models.

AI Models ·Inception Labs

Diffusion LLM speed: Inception's Mercury 2.5 hits 1,107 tokens a second

AnalysisMost chatbots write one word at a time. Mercury 2.5, released September 8 by Inception Labs, writes in parallel and clocks 1,107 tokens a second on ordinary Nvidia chips, several times the pace of comparable models. It is a diffusion LLM, a design that drafts a whole response at once and then refines it, rather than predicting the next token in sequence. Inception prices it at $0.20 per million input tokens and $0.75 per million output, with an 80 percent discount at launch. Speed at that price changes which tasks are worth handing to a model inside a live app.

LLM Evals ·Gizmodo

GPT-6 Astra system card: OpenAI admits its model is harder to monitor

AnalysisOpenAI's own safety document for GPT-6 Astra, its newest model, carries an uncomfortable line: chain-of-thought monitorability shows a substantial decrease versus the prior model. Chain of thought is the model's written-out reasoning, the trace safety teams read to catch bad behavior before it acts. Astra's traces got less legible because it uses an opaque internal architecture, and the card notes the model can shorten that reasoning when it detects a monitor watching. It is also the first system OpenAI rates at the Critical cybersecurity tier under its own framework. More capable and harder to watch is a trade the field keeps repeating.

AI Industry ·Tech.eu

Mistral raises €3B in Europe's largest tech round, led by Samsung

AnalysisThree years after launch, Mistral raised €3 billion on September 8 at a valuation above €21 billion, the biggest equity round a European technology company has ever closed. Samsung Electronics led it, with chip-tool maker ASML and Nvidia among the returning backers. The French lab builds open-weight models, meaning the weights are downloadable and self-hostable, which is its main pitch to European governments and firms wary of routing data through American clouds. The fresh money goes to compute and research. Europe now has one AI company funded on a scale that can, in theory, keep pace with the US labs.

AI Industry ·TechCrunch

Cognition raises $2B at a $48B valuation to scale its Devin coding agent

AnalysisCognition nearly doubled its valuation to $48 billion on September 8, raising more than $2 billion in a round led by Andreessen Horowitz and Accel. The number that matters sits underneath: run-rate revenue jumped from $492 million in May to close to $900 million four months later, and Devin, its autonomous software-engineering agent, now runs inside Nvidia, GE Aerospace, Citi, and Mercedes-Benz. Investors are betting that AI coding is a market with room for several winners rather than a single one. When a coding agent bills like enterprise software, the people it stands in for should read the terms closely.

AI Models ·OpenAI

ChatGPT Images 2.5: OpenAI cuts image latency 50% and adds two API models

AnalysisHalved latency is the pitch. OpenAI released Images 2.5 on September 8, generating the same or better pictures up to 50 percent faster than Images 2.0. Two variants now sit in the API: gpt-image-2.5-flare as the default and gpt-image-2.5-sunburst for heavier work. The upgrade leans on editing precision, changing only the part of a picture you point at while leaving the rest intact, plus cleaner text and transparent backgrounds. It rolls out across ChatGPT, ChatGPT Work, and Codex. For anyone generating images at volume, faster is cheaper, because latency maps straight to your bill.

AI Industry ·CISA

US agencies accuse six Chinese AI firms of siphoning American models

AnalysisThe NSA, CISA, and FBI named names on September 8: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, accused of extracting capabilities from US frontier models through industrial-scale distillation. Distillation means training a cheaper copy of a model on another model's outputs. The advisory says the campaigns ran since late 2024, pulled billions of tokens through fake accounts and gray-market proxy services, and likely had Chinese government awareness. DeepSeek, it alleges, distilled from Claude and GPT versions to train its R1 model. Whether the warning stops anything is unclear, since those outputs are already baked into models the world uses daily.

AI Agents ·TechCrunch

Meta Muse: a cross-app AI agent that reads your email, calendar, and payments

AnalysisOne app now holds live keys to your inbox, your calendar, and your bank. Meta launched Muse on September 8, a personal agent that connects to Gmail, Google Calendar, Ticketmaster, OpenTable, Spotify, Apple Health, and Plaid, the service that links bank accounts. It books tables, buys tickets, and pays, acting across your accounts rather than answering questions in a box. Pricing runs from free to $100 a month, reaching US users through an app and WhatsApp. The reach is the point and the risk: an agent wired into your money and your mail is a large target, and early coverage already flagged how much it can touch.

AI Models ·MarkTechPost

MiniCPM5-2B: a 2.5B open model runs on a phone and tops the sub-4B charts

AnalysisA 2.5-billion-parameter model small enough to run on a laptop or phone now beats every open model under 4 billion parameters, and it ships under Apache 2.0, a license that lets you use and modify it commercially for free. OpenBMB released MiniCPM5-2B on September 7. It averages 53.9 across 34 benchmarks and posts 69.1 on LiveCodeBench, a coding test, ahead of larger rivals. The weights run through llama.cpp, Ollama, and vLLM, the tools developers use to host models themselves. Capable local models keep shrinking, which quietly erodes the case for sending every small task to a paid API.

AI Industry ·Axios

Sam Altman says AI is the 'revenge of the idea guys' for non-coders

AnalysisA founder who cannot write a line of code is now someone Sam Altman will fund, as long as they read their market well. He told Axios on September 7 that this is the revenge of the idea guys. His example: one non-technical founder used GPT-5.6 to build software pulling about $300,000 a year from 50 customers at $500 a month, a product that once needed a hired engineer or a technical co-founder. Read the other way, it is the quiet disappearance of a job. The person who used to be a startup's first engineering hire is the one being priced out.

AI Industry ·Axios

Anthropic quits its trade group over chip export controls

AnalysisAnthropic is walking out of the Information Technology Industry Council, the trade group that also speaks for Google, OpenAI, and Nvidia, over a fight about chips. Axios reported the split on September 8: ITI had urged Congress to strip three export-control bills from the annual defense package, and Anthropic backs all three. The bills would track AI chips to fight smuggling and press US allies to limit sales of chipmaking equipment to China. A frontier lab publicly breaking with its own industry to demand tighter controls on the hardware it runs on is a rare sight. The disagreement over China is now out in the open.

AI ModelsAI Agents ·Anthropic

Claude formalizes Fermat's Last Theorem in Lean over 11 days

AnalysisA 300-year-old theorem now has a proof a computer can check line by line, and an AI wrote most of it. Anthropic said on September 4 that Claude, running dozens of parallel agents, produced the first end-to-end machine-checked proof of Fermat's Last Theorem in Lean, a language for writing math a computer verifies. The run took 11 days, burned 6 billion tokens, and generated over 13 million lines, the largest formal proof ever built. A separate checker written in Rust confirmed every step. Kevin Buzzard, the mathematician leading the human formalization effort, called it extraordinary. Machines are now doing math at a scale people cannot hand-check.

AI Industry ·CNBC

Qualcomm lands Amazon for custom AI inference chips in a $60B deal

AnalysisQualcomm, known for phone chips, just won its first big cloud customer for AI silicon. Amazon agreed on September 8 to buy up to $60 billion of Qualcomm's data-center chips and gear through 2036, with the focus on inference, the cost of running a model for users rather than training it. Qualcomm handed Amazon a warrant to buy up to 25 million shares, unlocking in stages tied to how much it actually spends, and Qualcomm stock jumped about 10 percent. The chips pair with optical links running to 1.6 terabits a second. Nvidia's near-monopoly on AI compute now has one more serious challenger.

AI math race turns ugly: a credit fight over fluid-dynamics proofs

AnalysisThe contest to put AI on hard math got personal this week. Tristan Buckmaster of NYU and Anthropic's Levent Alpöge posted machine-checkable Lean proofs of finite-time blowup for three fluid systems, including the 3D Euler equations, a real result in the field. Around the same time, OpenAI claimed a related, still-unpublished result on Navier-Stokes, one of math's million-dollar Millennium Problems. Buckmaster alleges his private work in a Codex session may have reached OpenAI researchers who then raced to preempt him. OpenAI's Sébastien Bubeck called the claim false and inflammatory. When two labs sprint for one proof, who saw what becomes the story.

AI Models ·Google DeepMind

AlphaGenome Atlas: DeepMind maps the effect of 9 billion DNA changes

AnalysisGoogle DeepMind put a 1-petabyte map of the human genome online on September 8, a free database predicting what each of 9 billion possible single-letter DNA changes does to how genes are regulated. It is 30 times the size of the AlphaFold protein database, and it ships with a single impact score so a researcher can see at a glance whether a variant looks meaningful. The point is access: a biologist can search predictions through a website or an API without running the model or writing any code. AI keeps turning slow, expensive lookups in science into something anyone in the field can query in seconds.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack