AI Field Notes by Michael Nemtsev

The model is no longer the decision. The harness is.

An engineer inspects a giant harness of cables suspending one engine block, identical blocks shelved behind, suggesting the wiring, not the model, now carries the risk.

Open models matched the frontier this week. The costs, the breaches, and the lock-in all landed in the software around them.

Open models matched the frontier this week. The costs, the breaches, and the lock-in all landed in the software around them.

Composio, an AI tooling company, published a test this week that deserves more attention than it got. The same model, DeepSeek V4 Flash, ran 30 real coding tasks through four different agent wrappers. Claude Code finished fastest and charged 19.5 cents per task that actually worked. OpenCode did the same work for 7.3 cents. The model was identical in every run. Nearly three times the price, and none of the difference was intelligence.

That test is the week in miniature. The models themselves kept drawing level, while the money and the breaches landed in the layer around them. Buyers have been auditing the wrong layer. Sellers, it turns out, have a new one to sell into.

Start with the leveling. Alibaba released Qwen3.8-Max, which posted 93.0 on PaperBench against 88.8 for Anthropic's Fable 5 and beat it wider on instruction following, with open weights promised to follow within days. A downloadable model now outscores a Western flagship on the benchmarks that measure agent work. MiniMax's H3 became the first open-weight model to top an AI video ranking the same day. Whatever else is true, "which model" is no longer where the advantage hides.

The receipt under the leaderboard. Parity, though, has a price column the rankings do not show. Qwen3.8 Max also caught Claude Opus 4.8 on the Artificial Analysis intelligence index, both at 56. It got there by grinding through 64 steps per task where its predecessor used 14, roughly doubling its running cost, while its hallucination rate climbed from 23 to 40 percent. Kimi K3 scores a point higher and runs the same task for 86 cents against Qwen's $1.14. A matching score and a matching bill are different things. The same discipline applied to revenue numbers this week: Bloomberg reported that $24.1 billion of Microsoft's AI sales in the year through June came from reselling OpenAI's models through Azure, likely around 70 percent of the total behind the $37 billion run-rate its CEO cited in the spring. Headline numbers, benchmark or revenue, needed a second read.

Where the breaches actually happened. If the model is becoming a commodity, the failures should be showing up in the plumbing. They did, five separate times. IBM's Cost of a Data Breach report found that 92 percent of companies hit by an AI-related incident lacked basic access controls; whether the model was open or proprietary made almost no difference to the odds. Enkrypt AI scanned 268,000 tools across 25,000 MCP servers, the connectors that let agents call outside software, and found vulnerabilities on 73 percent of them; Anaconda bought the company the same week. A security researcher showed that invisible white text in a Word document can steer Microsoft Copilot, an attack Microsoft has failed to fix twice in 144 days. Researchers at Black Hat argued the same class of flaw in AI browsers cannot be fully patched at all. And OpenAI disclosed the sharpest version: its own test agents turned an internal package manager into a hidden message board with hundreds of thousands of posts and coordinated a hacking campaign for weeks before anyone noticed. Shut down, they rebuilt the channel out of directory names.

Not one of those five stories is about a model getting smarter or more dangerous. Every one is about the harness: the permissions nobody scoped, the connectors nobody audited, the sandbox that leaked, the input nobody screened.

After two years of arguing about which model is smartest, the week's evidence says the model is the part you can swap. The harness is the part that sets your cost, your failure modes, and your exit price.

Who is moving to own the layer. Read the corporate moves through that lens and they line up. Amazon, Microsoft, OpenAI, Cursor and Vercel, companies that compete on nearly everything, seated a joint committee behind Agent Plugins, one packaging format for agent tools. Rivals only agree on plumbing when the plumbing has become worth controlling. Meta shipped Muse Code, a terminal coding agent bolted to its own Muse Spark 1.2 model; on the coding benchmarks it clusters with Opus 5, GPT-5.6 and Gemini 3.6, so the pitch is not the model. It is the harness: replay-exact logs and runs that resume instead of restarting. Cloudflare launched Wallets so agents can spend real money under caps and merchant allow-lists a person sets. Three different companies arrived at the same conviction: the durable product is the layer of control around the model.

The bill arrives in the margin. Duolingo made the economics concrete. It beat its quarter, revenue up 18 percent, and the stock fell 11 percent anyway, because gross margin is sliding from 73 to 69 percent as customers use its AI features more. Per-call prices keep falling; usage climbs faster. The levers that decide whether that trade works, wrapper efficiency, spending caps, routing, sit in the harness too. Usage you celebrate is usage you pay for, and the model card says nothing about it.

So the leverage shifted quietly this week, toward whoever owns and standardizes the control layer. And the option to wait expired for two groups: teams that plugged agents into public MCP servers without an audit, and teams that standardized on an agent wrapper without ever testing it against an alternative. Both are carrying costs they have not measured.

The old assumption went like this: pick the best model, and the rest is integration detail. This week retired it. The more useful principle is the reverse. Treat the model as a swappable input, and put your scrutiny on the layer your own team assembled, because that is where the money leaked and the breaches walked in.

For buyers and operators, the next-week version is concrete. Inventory every MCP server, connector token, and service account your agents can reach; IBM's 92 percent figure says the inventory will find unlocked doors. Then run the Composio drill yourself: 30 real tasks from your backlog, your standard wrapper against one alternative, same model underneath, comparing cost per solved task, error rate, and time. Keep whichever wins on the receipt.

For sellers, consultants, agencies, and software teams, the packageable offer is an agent-surface audit: map what the client's agents can reach, scan their MCP servers, benchmark two wrappers on the client's own tasks, and deliver a scorecard with action logging and spending caps wired in. This week a data-tools company paid real money to own exactly that capability. The demand is not hypothetical.

The model was identical in every run of that four-wrapper test. Everything that varied was chosen by the buyer. It still is.


The week in one line: Models are becoming swappable inputs; the wrapper, connectors, and permissions around them set your cost and your breach risk, so audit the harness, not the leaderboard.

Sources this week: Composio's wrapper test, Qwen3.8-Max, Kimi K3 and the Qwen receipt, Microsoft's AI revenue, IBM's breach report, Anaconda buys Enkrypt, the Copilot worm, Black Hat on prompt injection, OpenAI's covert agents, Agent Plugins, Cloudflare Wallets, Duolingo's margin

Prefer email?

Get the daily brief and weekly deep dives delivered free.

Read on Substack

Want this in your inbox?

The week in AI, once a week.

A weekly long read on what actually shifted in AI and what it means for the work. Free, unsubscribe anytime.