AI Field Notes by Michael Nemtsev

AI Inference Cost Drop | AI Field Notes #78

A machine tightens its own gears while a price tag scratches lower and headset-wearing figures watch, showing AI cutting its own running costs and jobs.

AI inference costs are dropping fast, as retraining and self-optimizing model serving do the work that ever-bigger models used to. DeepSeek's re-post-trained V4-Flash beat its own flagship on nine agent and coding benchmarks at about a third of the price, while OpenAI says GPT-5.6 tuned its own serving stack to cut costs 20% and speed generation 15%. Cheaper voice and transcription models from xAI and OpenAI pushed speech work down the same week. In the background, Washington's August 1 frontier-model review deadline and a shift where chip supply, not demand, now gates AI round out a quiet late-July board.

AI ModelsAI Agents ·MarkTechPost

DeepSeek V4-Flash: a retrain beats the flagship at a third of the price

AnalysisA smaller model just outscored its own flagship, and DeepSeek pulled it off without touching the architecture. The V4-Flash build it shipped on July 31, a 13-billion-active mixture-of-experts model (a design that fires only part of itself per request to save compute), was re-post-trained rather than rebuilt. DeepSeek's own figures put it ahead of the larger V4-Pro preview on all nine agent and coding tests it published, including 82.7 on Terminal-Bench. Output runs about $0.28 per million tokens, near a third of Pro, and it now speaks OpenAI's Responses API, so Codex can call it directly. The scores are vendor-reported, but the lesson for anyone paying an inference bill is that post-training has become the cheapest place to buy capability.

xAI's Grok voice model answers in 0.7 seconds, live on Starlink support lines

AnalysisA synthetic voice that replies in seven-tenths of a second is fast enough to hold a real phone call, which is the whole point. xAI released Grok Voice Think Fast 2.0 on July 29, a speech-to-speech model (audio in, audio out, with no text step in the middle) that scores 82.9% on Artificial Analysis' Speech-to-Speech Quality Index with a 0.70-second time to first audio. xAI has already put it on Starlink's customer support lines at eight cents a minute. Frontline phone support has been the loudest promise of voice AI for years, and a sub-second agent running in production quietly turns that promise into a deployment.

AI AgentsAI Models ·OpenAI Platform docs

OpenAI's new transcription models cut word errors 41% below Whisper

AnalysisSpeech-to-text got sharper for everyone building on it. On July 28 OpenAI retired its older Whisper engines for two replacements, GPT-Transcribe for batch jobs and GPT-Live-Transcribe for streaming, reporting an 8.98% word error rate against Whisper-1's 15.21%, a 41% cut. The models accept a context prompt, so you can hand them the names, product terms, and jargon a recording is likely to contain and get back fewer mangled words. Transcription is the unglamorous layer under voice agents, meeting notes, and live captions, and a jump in accuracy there raises the ceiling on all of them at once.

OpenAI says GPT-5.6 cut its own serving cost 20% and dropped prices

AnalysisThe price cut matters less than who designed it. OpenAI said on July 30 that GPT-5.6 Sol ran hundreds of its own experiments on speculative decoding (a method that guesses several tokens ahead to speed generation) and on GPU kernels, and the payoff was 15% faster token generation with 20% lower serving costs. Those savings went straight into the API: the Luna tier fell 80% to $0.20 per million input tokens, Terra dropped 20%. A model tuning the machinery that runs it is a quiet milestone, and it arrives the same week Washington began asking labs for early access to exactly this kind of self-improving system.

AI Industry ·DIGITIMES

TSMC lifts 2026 capex to $64B as AI's bottleneck shifts from demand to supply

AnalysisFor three years the question was how many chips the big cloud buyers wanted; now the question is how many TSMC can physically make. The foundry raised its 2026 capital budget to between $60 billion and $64 billion, at least $4 billion above its earlier plan, and lifted its revenue growth forecast past 40%. Underneath the numbers is a change in what gates AI: advanced packaging and memory supply, not order books, with mature-node and packaging vendors raising prices as demand outruns capacity. Compute still expands, but the ceiling is set in a fab now, and that ceiling reaches the price and availability of every GPU downstream.

AI Industry ·Vorp Labs

US frontier-model review deadline hits August 1 with classified cyber tests

AnalysisThe government wants a look at frontier models before the public gets one. August 1 is the deadline under Executive Order 14409 for the Treasury, the NSA, and CISA to stand up a review framework that runs classified benchmarks on a model's offensive cyber ability and, for any that count as a 'covered frontier model,' asks the developer to hand over up to 30 days of pre-release access. The order calls this voluntary. The leverage is already visible: this summer Commerce suspended Anthropic's Claude before restoring it, and the administration shaped the timing of OpenAI's GPT-5.6 rollout. The threshold that defines a covered model stays classified, so a lab will not know it qualifies until the ask arrives.

AI ModelsAI Industry ·Google Flow (X)

Google's Lyria 3.5 makes full three-minute songs with key and tempo control

AnalysisThree-minute songs on demand, with the key and tempo dialed in to order: Google's Lyria 3.5 music model, out July 29, narrows much of the distance between a text prompt and a finished track. It also produces style-transfer covers and, paired with the Gemini Omni Flash model, lip-synced music videos, every output stamped with SynthID, Google's invisible watermark for AI-made media. The taste questions are real, and so is the arithmetic. A jingle, a stock-music bed, or a background loop that once meant a session player and a studio afternoon is now a prompt and a short wait. The cheap end of music production just got cheaper.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack