AI Field Notes by Michael Nemtsev

AI Coding Cost Squeeze | AI Field Notes #106

A corporate hand turns a usage dial downward while its painted arrow still points up, as a lone coder eyes exits toward rival tools.

The cost of running AI coding agents shifted this week: Anthropic trimmed Claude Code's weekly limit about 17% while branding it a raise, and developers promptly rerouted the tool to cheaper rival models. New releases piled up alongside it, from Google's Gemini 3.8 Live voice models to Shanghai's open 744-billion-parameter Atria Dawn and TypeSafe's Jev, a model that claims it cannot hallucinate. Security got louder too, with an AI agent cracking a hosting firm's production code in 25 minutes and fresh research showing data-poisoning attacks were badly underrated. Off the tools, OpenAI, Anthropic, and Google admitted they have been coordinating on safety while Jensen Huang and Dario Amodei aired opposite views on regulation. A few items reach back across the past several days, where the fresh board ran thin.

AI Models ·Every

TypeSafe's Jev: a model that returns values, not words, and claims it can't hallucinate

AnalysisDiogo Almeida, a co-author of the training method behind ChatGPT, has released Jev through his startup TypeSafe with a large claim: a model that returns typed values instead of free text and, by design, cannot make things up. He argues that hallucination comes from the layer that tunes models to please people, so Jev drops it and forces every answer into a schema, a fixed structure that defines which outputs are allowed. TypeSafe prices it at $42 per billion tokens, against the per-million rates elsewhere, and clocks replies at 70 to 500 milliseconds. Treat the zero-hallucination line with suspicion until production proves it. The bet underneath is sound: for a job with one right shape of answer, a chatty generalist was always overkill.

Atria Dawn: Shanghai lab drops a 744B open agentic model with no press release

AnalysisA 744-billion-parameter agentic model appeared on Hugging Face on September 11 with an MIT license and no announcement, credited to the Shanghai Artificial Intelligence Laboratory, with the research paper following on September 14. Atria Dawn is built on the open GLM-5.2 base and trained with what the authors call a Verifiable Experience Pipeline, which grounds the model's tool use in environments where success can be checked by running the result. The paper reports top scores on five of sixteen benchmarks. MIT licensing (free to use and modify, including commercially) hands any team with enough GPUs a frontier-grade agent, with no contract to sign.

AI Models ·Google Blog

Gemini 3.8 Live: Google ships voice models that keep talking while they work

AnalysisTwo new voice models from Google, Gemini 3.8 Live and a heavier Extended Thinking variant, can switch among 97 languages mid-sentence and run tool calls in the background without pausing the conversation. Google put them out on September 15 through the Gemini API, so any developer building a phone-style agent can wire them in this week. The Extended Thinking version topped one independent speech-to-speech quality ranking at 82.6. The pitch is a voice assistant that reasons over several steps and acts while it listens, rather than the stop-start turn-taking that made older voice bots feel like hold music.

AI AgentsAI Industry ·Implicator AI

Claude Code limits: Anthropic trims weekly usage 17% and calls it a 25% raise

AnalysisThe weekly usage allowance for Claude Code, Anthropic's terminal coding agent, drops about 17% on September 14, though the company's own post frames it as a permanent 25% increase. Both numbers are honest, measured from different starting points: the cut lands against a temporary summer boost, the raise against the level before that boost existed. Anthropic said it could not secure enough compute to keep the boost running. Within two days, developers were already pointing Claude Code at rival models through OpenRouter, a service that reroutes the same command-line tool to GPT and other back ends. The lock-in loosened the moment the allowance shrank.

AI Industry ·404 Media

Project Lily: OpenAI paid contractors to read real ChatGPT prompts

AnalysisHundreds of contractors have been reading real ChatGPT conversations and scoring the replies one to seven, under an OpenAI program called Project Lily, per documents obtained by 404 Media. Many prompts carry sensitive personal detail. Reviewers do not see usernames, and OpenAI says it strips identifying data first, yet admits some gets through, and the dashboard often shows a summary of the user's past chats that can betray their job or rough location. ChatGPT has more than 900 million users, and training on their chats is on by default. The workers rating your words earn over $50 an hour to teach the model to sound less like a machine.

Data poisoning: which samples you pick swings backdoor success from 3% to 80%

AnalysisResearchers at Carnegie Mellon and Anthropic held everything fixed, the model, the clean data, the count of poisoned examples, and found attack success on a tampered language model swings from 3% to 80% based only on which poison samples get picked. Data poisoning means slipping corrupted examples into training data to plant a hidden trigger. Their method, SAILS, learns to rank candidate poison sets and beats earlier tricks by about 30 percentage points, carrying over to code generation and agent tasks. The takeaway is uncomfortable: the risk of poisoning has been measured against average attacks, not the sharp ones a real adversary would hunt for, so the published danger was set too low.

Baseten breach: an AI hacking agent found a 3-year-old token with admin access

AnalysisAn autonomous security agent named Strix, probing the AI hosting company Baseten, found an unguarded image registry and pulled a GitHub access token out of a Docker build log from March 2023. The token still worked, with push rights to Baseten's production code and a folder tree named after its customers. From first probe to admin access took about 25 minutes. Baseten rotated the credential within a day and handled the report well. The lesson is old and keeps not sticking: secrets pasted into build steps do not expire on their own, and a machine that never gets bored will read every layer of every image you ever shipped.

AI Industry ·NY Governor's Office

New York asks data centers to pay $1M per megawatt to the towns they land in

AnalysisNew York now recommends that data center developers pay about $1 million for every megawatt of electricity their project draws, sent to a local fund for housing, water, and childcare. Governor Kathy Hochul unveiled the framework on September 15, the follow-through on a July order that froze new data center approvals across the state. The reasoning is blunt: these sites pull enormous power and land while creating few permanent jobs per megawatt, so the host town should get paid for the strain. The benchmark is voluntary for now, which usually means it sets the floor for what communities will demand in the negotiations that follow. A 500-megawatt campus would carry a suggested $500 million community tab.

AI Industry ·SemiAnalysis

Vera Rubin: Nvidia's next chip shows 7x more tokens per megawatt in early tests

AnalysisEarly testing of Nvidia's next-generation Vera Rubin systems shows up to seven times more output per megawatt of power than today's Blackwell generation, per the analysis firm SemiAnalysis, and that is on pre-release software that usually improves with age. Power has become the real ceiling for AI companies, ahead of chip supply, so wringing more tokens (the chunks of text a model reads and writes) from each megawatt matters more than raw speed. If the gains hold, one substation serves several times the users. The catch is timing: these are lab figures on hardware still in early bring-up, and the distance from a benchmark to a full data center has swallowed rosy forecasts before.

AI Industry ·TechRadar

Dreamforce split: Huang calls AI doom a manufactured fear, Amodei wants a brake

AnalysisOn the same stage at this week's Dreamforce conference, Nvidia's Jensen Huang and Anthropic's Dario Amodei laid out opposite bets on how dangerous AI is. Huang dismissed warnings that AI could threaten humanity as manufactured fear with no scientific grounding, and pushed for a single federal rulebook rather than 50 state ones he says would choke the industry early. Amodei, fresh off an essay calling for a slower frontier, argued the risks are real enough to act on today. The split tracks the money more than the science. Huang sells the shovels and profits from speed; Amodei sells a model whose whole brand is caution. Each man's reading of the danger lines up with his revenue.

AI Industry ·TechCrunch

AI safety pact: OpenAI, Anthropic, and Google have been coordinating for weeks

AnalysisThe three largest US AI labs have been working together on safety for several weeks, OpenAI's policy chief Chris Lehane told reporters on September 15. The sketch so far includes a shared standards body and outside evaluators embedded to watch how each lab handles risk. It follows an essay days earlier by Anthropic's Dario Amodei urging the industry to slow the frontier race. The awkward part is antitrust: rivals coordinating can slide into illegal collusion. Amodei floated asking Washington for a narrow waiver; Lehane called it unnecessary. Three firms that fight over every engineer and every dollar now say they will police each other on safety, which is either maturity or a cartel in a lab coat.

AI Agents ·arXiv

Design agents: new research pushes AI past one good-enough interface

AnalysisAI tools that turn a written brief into a working interface tend to hand back one passable answer, and a research paper posted September 14 goes after that narrowness. The method runs a first pass proposing several structured design options, each scored for how typical or unusual it is, then a selector picks one before any code gets written. Tested across 168 prompts with more than a thousand paired comparisons, it widened the spread of layouts the system would try instead of collapsing toward the same safe template. The problem it names is familiar to anyone who has used these tools: they are fast and confident and keep returning the average of everything they saw in training.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack