AI Field Notes by Michael Nemtsev

Agents That Run Agents | AI Field Notes #104

A lone figure conducts a swarm of identical worker-drones assembling code below, suggesting coding agents now direct fleets of other agents as costs fall.

Coding agents crossed a line this week: several stopped writing code and started directing fleets of other agents, while inference costs kept collapsing. Cursor's new Projects runs a coordinator that hands work to a fleet of cloud subagents, DeepSeek's open V4.1 Flash runs near one seventeenth the cost of Claude Opus 5, and OpenAI opened a managed Agents API to every developer. Separately, OpenAI used about ten thousand of its own agents to dent a million-dollar math problem. The quieter development sits in Anthropic's threat report, which logs the first confirmed real-world attacks run through its models.

AI Agents ·OpenAI

OpenAI Agents API: the company will now host your agents, not just your model calls

AnalysisOpenAI wants to run your agents, not just answer your model calls. On September 10 it opened the Agents API in public beta, a managed service that runs autonomous agents in OpenAI's cloud and charges nothing beyond the model tokens they consume. Developers can plug in their own sandbox, an isolated space where an agent's code runs without touching the real system, or use an outside one like E2B. The pitch is that the plumbing every team was rebuilding by hand, memory, tool calls, safe execution, now comes from the vendor. The cost is plain: one more piece of your stack lives on OpenAI's servers.

AI Models ·AI Weekly

DeepSeek V4.1 Flash: open model runs near a seventeenth the cost of Claude Opus 5

AnalysisThe cheapest way to run a frontier-grade model keeps getting cheaper. DeepSeek released V4.1 Flash on September 9 under an MIT license, free to use or modify, built as a mixture-of-experts model, a design that wakes only part of itself per request to cut cost. It carries 552 billion parameters but fires 8 billion while reading a prompt and 16 billion while writing the answer, with a one-million-token context window. Testers put its running cost near one seventeenth of Claude Opus 5 for similar quality. The Chinese lab is doing to the cost of running models what it already did to training them.

AI Agents ·Cursor

Cursor Projects: a coordinator agent now directs a fleet of coding agents

AnalysisThe developer's job just moved up a rung. Cursor's Projects, released September 10, puts a coordinator agent in charge of a body of work and lets it direct other agents that do the actual coding, running on a cloud machine that keeps going after you shut your laptop. The coordinator watches Slack, runs on a schedule, and tracks pull requests, the proposed code changes a team reviews before shipping, then acts on its own. Cursor says new users merge 30 percent more pull requests and people who lean on Projects merge six times as many. The person who used to write the code now reviews what a team of agents produced overnight.

AI Industry ·CNBC

Oracle's AI cloud backlog hits $664B as demand outruns the capacity it can build

AnalysisOracle's order book now dwarfs its revenue, and that is the nervous part. In results reported September 10, the company said contracted future business, what it calls remaining performance obligations, reached 664 billion dollars, up 209 billion in a year, after booking more than 30 billion in new AI cloud deals in one quarter. Cloud infrastructure sales jumped 121 percent to 7.4 billion. Demand for AI training and inference is growing faster than Oracle can build capacity, it said. The debt funding that build-out sits near 125 billion, which is why the stock swung from up 5 percent to down 2 in a single day.

AI Agents ·Upstarts Media

Leaky sandboxes: security firm finds escapes in Claude Code, Codex, and Cursor

AnalysisThe agents that write your code can also break out of their box. Accomplish, a security firm, reported on September 10 that the sandboxes in Claude Code, OpenAI's Codex, and Cursor, meant to wall an agent's actions off from the rest of a machine, all leaked. Cursor flagged the hole in July and closed it in about a week, OpenAI fixed two in August, and Anthropic took roughly fifty days and thirty releases to patch. Accomplish's chief technology officer, Or Hiltch, asked the obvious question: if these models are so good, why do they not catch critical bugs in their own products? How fast a vendor patches is now part of the tool's safety.

LLM EvalsAI Industry ·Anthropic

Anthropic threat report: real espionage and data theft run through Claude

AnalysisThe abuse is no longer hypothetical. Anthropic's threat report, out September 10 and covering December 2025 through August, catalogs real operations run through Claude: a Russian espionage group that hit more than twenty government and defense targets, a data-theft crew that pulled over a terabyte of records including tens of millions of passenger files, and Chinese operators who surfaced more than a dozen zero-day flaws, previously unknown software holes, in a single month. One influence network posted 8,913 articles in about twenty languages across roughly seventy fake news sites. Five separate cases involved biological misuse. The tools built to help are being turned, at scale.

AI ModelsLLM Evals ·OpenAI (GitHub)

OpenAI agents crack a piece of the $1M Navier-Stokes prize, verified in Lean

AnalysisFor a century the Navier-Stokes equations, the math describing how fluids flow, resisted a full proof, and a correct answer carries a million-dollar Clay Institute prize. On September 10 OpenAI published results showing smooth starting conditions can blow up to infinity in finite time, a partial answer to the prize question, plus a matching result for the related Euler equations. The method is the real story: roughly ten thousand of its agents traded 2.7 million messages and 130 billion tokens over 88 hours, then software checked the proof in Lean, a language that verifies math step by step. Mathematicians are now grading a machine's homework.

AI Industry ·TechCrunch

Massachusetts orders data centers to bring their own clean power

AnalysisStates are done rolling out the welcome mat for data centers. Massachusetts governor Maura Healey signed an order this week requiring any data center drawing more than 25 megawatts, enough to power a small city, to supply its own electricity and meet 100 percent of that demand with clean energy, far above the state's standard target of 40 percent by 2030. She also told towns to stop signing the non-disclosure agreements that hide these deals from residents, and paused a tax break for the industry. Massachusetts is the third state in three months to clamp down, after Texas and New York. The cheap-power era for AI build-outs is closing.

AI Industry ·Suno

Suno v6 ships with Warner, BMG, and Believe on licensed data

AnalysisThe music industry spent two years suing AI song generators, and now it is shipping with one. Suno released v6 on September 9, its first model built alongside label partners Warner Music Group, BMG, and Believe, using licensed material rather than scraped catalogs. It ships in three tiers, from a polished flagship down to a free lightweight version, with a wilder experimental option in between. Suno also added filters that screen uploaded audio and lyrics for unauthorized use, the kind of guardrail the lawsuits demanded. The labels decided a cut of the machine beats a court fight they might lose.

AI Industry ·The Hollywood Reporter

AI now makes 95% of China's microdramas, and the crews are gone

AnalysisThe short vertical soap operas that fill phone screens are now made mostly by machines. In China, more than 95 percent of these microdramas were AI-generated by early this year, according to a Hollywood Reporter piece dated September 10. A traditional shoot ran 100,000 to 300,000 dollars, employed around fifty people, and took about twelve weeks. The AI version costs as little as 1,000 dollars and wraps in two. The quality gap is real, and for a genre watched in ninety-second bites on a commute, most viewers do not seem to mind. The crews lost their work to something cheaper that the audience accepts as good enough.

AI Industry ·Rest of World

China's white-collar workers train the AI for $15 to $74 a task

AnalysisChina's white-collar workers are training the machines that may replace them, for 15 to 74 dollars a task. As youth unemployment hit 17.9 percent in July, tech giants opened gig platforms, ByteDance's Xpert alone signed more than fifty thousand experts, that pay teachers, engineers, accountants, and composers to feed their specialized knowledge into AI models. The work fits people whose regular income has thinned in a slow economy. The data market behind it reached 7.8 billion yuan, about 1.1 billion dollars, this year, a quarter more than last. Each task done well makes the human who did it a little more optional.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack