AI Field Notes by Michael Nemtsev

The model got cheap. Control is where the bill moved.

A shrinking price tag sits beside a many-armed agent reaching through a broken fence toward a vault and calendar, showing cheap models but unguarded reach.

As capable models fall to a fraction of frontier prices, the real cost of AI moves from the model to what its agents are allowed to touch.

As capable models fall to a fraction of frontier prices, the real cost of AI moves from the model to what its agents are allowed to touch.

This week DeepSeek released V4.1 Flash under an MIT license, free to run and modify, and testers put its cost near one seventeenth of Claude Opus 5 for similar work (DeepSeek V4.1 Flash: open model runs near a seventeenth the cost of Claude Opus 5). The same week, Google's threat team reported that one attacker wired together an off-the-shelf coding chatbot and a set of playbooks and harvested 23,800 credentials in under six hours, with no human at the keyboard (Google: one attacker used AI agents to grab 23,800 credentials in six hours).

Read together, those two facts move the AI decision somewhere most budgets have not caught up to. The price of a capable model is falling toward zero. The price of letting one act is climbing.

Most planning still treats this as a build-versus-buy question: pick the cheapest model that clears your bar, wire it in, move on. That framing made sense while the model was the scarce, expensive part. This week it stopped being either.

The cheap part got cheaper, and it moved. DeepSeek's Flash is a mixture-of-experts model that fires a fraction of its parameters per request to hold cost down, and its maker says it beats the previous top tier (DeepSeek V4.1 Flash: a cheaper multimodal model it says beats its own Pro tier). Treat the beats-our-Pro-tier line as a vendor claim until your own traffic confirms it. The direction, though, is not in doubt. OpenBMB shipped a 2.5-billion-parameter model that runs on a phone and tops the sub-4B charts (MiniCPM5-2B), Inception's Mercury 2.5 clocked over a thousand tokens a second at sub-cent pricing (Diffusion LLM speed), and Mistral raised three billion euros, Europe's largest tech round, to keep an open-weight, self-hostable model near the frontier (Mistral raises €3B in Europe's largest tech round, led by Samsung). Capable inference is becoming something you own and run in your own building, not a metered call to someone else's.

The expensive part is control, and it broke in public. The plumbing that lets these models act, rather than just answer, failed all week. LiteLLM patched an 8.8-rated auth bypass in the MCP endpoint many teams use to plug models into their apps (LiteLLM patches an 8.8-rated auth bypass in its MCP endpoint). VulnCheck scored a flaw in DeepSeek Harness 9.4 that let an unauthenticated request switch an agent to full access and turn off its own sandbox (DeepSeek Harness flaw let AI agents switch off their own sandbox). A security firm found the sandboxes in Claude Code, Codex, and Cursor all leaked, and noted Anthropic took about fifty days and thirty releases to close its hole (Leaky sandboxes: security firm finds escapes in Claude Code, Codex, and Cursor). Anthropic's own threat report catalogued real espionage and data theft run through Claude, including a crew that pulled over a terabyte of records (Anthropic threat report: real espionage and data theft run through Claude).

The shift no one is pricing. When the model was expensive, the model was the decision. Now the model is a commodity you can swap by dropdown, and Adobe made that literal by putting five competing video generators inside the Premiere timeline (Adobe puts Firefly, Veo, Runway, Kling, and Luma inside the timeline). What you actually buy is the agent's reach: what it can read, write, approve, and execute, and how quickly the vendor closes the gap when that reach is exploited. Meta shipped Muse, a personal agent holding live keys to your inbox, calendar, and bank (Meta Muse: a cross-app AI agent that reads your email, calendar, and payments). The connectors are the product and the risk in one.

Watch who is trying to become the floor. OpenAI opened its Agents API in public beta to host your agents in its cloud, charging only for tokens, which moves one more layer of your stack onto its servers (OpenAI Agents API: the company will now host your agents). Cursor put a coordinator agent in charge of a fleet of coding agents and says heavy users merge six times as many pull requests (Cursor Projects: a coordinator agent now directs a fleet of coding agents); that multiple is the company's own number, worth testing before you reorganize a team around it. Underneath both, Oracle reported a $664 billion backlog it cannot build capacity fast enough to serve (Oracle's AI cloud backlog hits $664B), and Massachusetts told data centers to bring their own clean power (Massachusetts orders data centers to bring their own clean power). The cheap model still runs on someone's constrained, expensive floor. Cheap to call is not the same as cheap to depend on.

So the old assumption to drop is that the model is the thing you are choosing. The model is now the interchangeable part. The decision that carries real cost is the agency around it: the permissions you grant, the sandbox you trust, the vendor whose infrastructure and patch speed you inherit. Falling model prices are not a reason to move faster into agents. They are a reason to spend the money you just saved on the part that broke this week.

That points two audiences in the same direction, from opposite sides.

For buyers and operators, stop scoring vendors on benchmark and price alone and inventory reach instead. For each agent you run, write down what it can read, write, approve, and execute, then run a hundred-case drill that tries to make it act outside those bounds. Ask every agent vendor one question before you renew: median time to patch a critical sandbox or auth bug, with the DeepSeek Harness score and the fifty-day Claude Code fix as your reference points. Put a human approval step wherever an agent touches money, credentials, or customer data, and treat a cheaper model as a reason to fund that review, not skip it.

For sellers, consultants, and software teams, the week wrote your offer. Package a dependency-and-permission scorecard: map every model call and agent action a client runs, test the sandbox the way VulnCheck tested DeepSeek Harness, score each vendor's patch latency, and deliver a fallback plan for the day a leak is theirs rather than the news. The demand is concrete because the failures were. Sell the audit that turns "we use agents" into "we know exactly what ours can touch and how fast our vendor fixes it."

The labs spent this week racing to make the model free and to become the ground it runs on. The number that should move your budget is not the price per million tokens. It is how many things your cheapest agent is allowed to do while no one is watching.


The week in one line: Model cost is collapsing toward zero; the real AI bill is now the reach you grant your agents and the patch speed of the vendor whose sandbox you inherit.

Sources this week: DeepSeek V4.1 Flash cost, Google's autonomous credential heist, DeepSeek V4.1 Flash beta, MiniCPM5-2B, Mercury 2.5, Mistral's €3B round, LiteLLM auth bypass, DeepSeek Harness flaw, Leaky sandboxes, Anthropic threat report, Meta Muse, OpenAI Agents API, Cursor Projects, Oracle's AI cloud backlog, Massachusetts data-center power rules

Prefer email?

Get the daily brief and weekly deep dives delivered free.

Read on Substack

Want this in your inbox?

The week in AI, once a week.

A weekly long read on what actually shifted in AI and what it means for the work. Free, unsubscribe anytime.