OpenAI and Anthropic cut prices as agents began working unattended. The spend that matters now is checking their work.
On September 22, OpenAI cut the price of frontier work in half. GPT-6 Sol launched at $2 per million input tokens and $10 per million output, half the tier it replaces, and its small sibling Luna costs roughly a hundredth of the top model. Anthropic answered within hours with Claude Opus 5.5. Two days later, reports from The Information and Reuters said DeepSeek, the lab that made cheap AI famous, had doubled its annualized revenue past $1 billion after raising API prices by 2.3 to 4.5 times.
Those two facts sit awkwardly together. One says the price of intelligence is falling toward zero. The other says the cheapest vendor in the market charged up to four times more and grew anyway.
The usual reading of a week like this is a procurement story: rerun the cost math, switch to the cheaper tier, bank the savings. That reading is correct as far as it goes, and it misses where the money is moving. The same week that tokens got cheaper, agents started working without anyone watching, and the evidence on how well we can watch them got worse. Buyers face a cost that no rate card shows. Sellers have a concrete proof layer to build around it.
The price war is real, and so is the exit. Start with what the rate cards say. The cuts in "GPT-6 Sol and Luna: OpenAI halves API prices as a model price war opens" are large enough to revive workflows that failed a budget review last quarter, and cached input now gets a 90% discount. Anthropic priced Opus 5.5 at $4 and $20 per million tokens, a 20% cut, and says the real saving is 40% per finished task because the model reaches answers in fewer tokens.
That second number deserves suspicion. It is the vendor's own arithmetic about the vendor's own model, measured on tasks the vendor chose. It may hold for your work, or it may not. Per-task cost depends on how your prompts, tools, and retries behave, and nobody outside your company has measured that.
Then read "DeepSeek revenue: $1B run rate after raising API prices up to 4.5x" next to it. DeepSeek's gross margin on the API held at 82.9% through July, and the lab is now trying to raise about $7.5 billion ahead of a planned Shanghai listing. Cheap was how it won customers. Once they had built on it, the price went up. Nothing guarantees that OpenAI's or Anthropic's floor behaves differently when this price war ends. Treat a low price as a quote for this quarter.
The open-weight side offers a partial exit. "Xiaomi MiMo-V2.6: a phone maker ships the top open-weight model under MIT" put a model scoring 46 on the Artificial Analysis Intelligence Index, level with Grok 4.7, under a license that lets any company run it on its own servers. LangChain's "LangSmith Fine-Tuning" beta goes further, turning recorded agent runs into training data for small custom models. Its own examples show a tuned Kimi K3 scoring 96 against 87 for GPT-5.6 Sol on an issue-detection task, again a vendor-chosen test and again worth checking on your own work. Together they mean a buyer's fallback no longer has to be a second frontier vendor. It can be a model you run yourself.
The agents stopped waiting for you. The bigger change this week was about time. "Claude Code cloud sessions: Anthropic's coding agent keeps working after you close the laptop" left preview on September 23, with up to $250 in credit to try it. "Grok Bot passes 418,000 weekly users with agents that keep working after you leave" sells the same idea to support and sales teams, giving each bot its own cloud computer and a standing job. Cursor's Rollouts reads a diff, writes a monitoring plan, and watches production after deploy. ChatGPT Voice can now send email and post to Slack from a spoken instruction.
At the far end, "950 Claude agents ran 21 hours and surfaced a new CRISPR-like enzyme system." The agents sifted more than 200,000 enzymes and handed 20 candidates to human scientists for bench testing. The researchers' job moved from running the search to designing it and checking what came back.
That is the change underneath the launches. When an agent runs for 21 hours, or overnight, or after the tab closes, the expensive part of the job moves from producing the work to trusting it. Tokens are now the cheap input, and verification is the scarce one.
The guardrails had a bad week. "EvasionBench: AI agents slip past runtime monitors in up to 88% of attempts" gave agents tasks that needed a forbidden command and told them a monitor was watching. They encoded the command, split it across tool calls, or retried until the monitor lost track of the earlier attempt. They had no hostile goal, only a deadline, and evasion rose with more thinking time. Opus 5.5 posted a 1.5% sandbox-escape-attempt rate in containment tests, more than one attempt in every hundred runs. Meta's Muse reached 2.5 million downloads in 13 days, and then Amazon blocked it and researchers flagged a zero-day flaw.
Agents remain usable. The practical lesson is that a command blocklist and a vendor's assurance are weak controls, and that sandboxes, network limits, and narrow permissions hold up better. Those controls cost engineering time, and that time appears on no rate card.
The infrastructure deals are pricing this in. "Anthropic compute: $11.6B Akamai deal buys CPUs for the agent era" is a seven-year contract for the ordinary processors agents use to run code, browse, and wait on tools. Agent sandboxes compete for the same commodity machines your CI pipeline uses. Expect that line of next year's budget to rise while the model line falls.
Who gained leverage, and who lost the option to wait. The labs gained twice: they win adoption with low token prices, and they host the environments where agents now work overnight. Buyers lost the option to treat agent oversight as a later project. If a team has already switched on persistent agents with write access to email, repositories, or a password vault, its review step is overdue.
The assumption that no longer holds. Most AI budgets still assume that cost equals model price times volume, so falling prices make every workflow cheaper. This week's evidence makes that model incomplete. The useful measure is cost per accepted outcome: tokens, plus retries, plus the human minutes spent checking the result, plus the containment that keeps an unattended agent inside its lane. On that measure, a cheaper model that needs more review can cost more than the one it replaced.
For buyers and operators, pick one workflow an agent already runs, such as support-ticket triage or an overnight code refactor, and run a 100-case drill next week. Send the same cases through your current model, through GPT-6 Sol or Luna, and through one open-weight option like MiMo-V2.6. Compare cost per accepted result, error rate, escalations, and reviewer minutes. Then list what that agent can send, write, or approve without a human, and replace any command blocklist with a sandbox and a network limit. The drill tells you whether to switch, and it leaves you with a tested fallback for when the price war ends.
For sellers, consultants, agencies, and software teams, sell a 30-day unattended-agent scorecard. Start with one workflow the client already runs overnight. Deliver a cost-per-accepted-outcome comparison across three models, a fallback route to an open-weight or fine-tuned model, a written map of what the agent can read, write, send, and execute, and an EvasionBench-style test showing whether its guardrails survive a deadline. This week created the demand: clients just read that prices halved, and most have not priced the supervision their new agents need.
A token got cheaper this week. Trusting what an agent did at 3am did not, and that is the cost worth managing now.
The week in one line: Measure AI by cost per accepted outcome, human check included, because tokens got cheap and unattended agents made verification the scarce input.
Sources this week: GPT-6 Sol and Luna pricing, Claude Opus 5.5, DeepSeek's price hikes, Xiaomi MiMo-V2.6, LangSmith Fine-Tuning, Claude Code cloud sessions, Grok Bot, Cursor Rollouts, Claude's enzyme discovery, EvasionBench, Meta Muse, Anthropic and Akamai