AI Field Notes by Michael Nemtsev

AI Model Price War | AI Field Notes #85

Identical mechanical brains with slashed price tags crowd a market stall while a support clerk is handed a headset, suggesting cheap AI flooding in.

An AI model price war broke out this week, with cheaper models and faster inference reshaping what developers can afford to run. Google cut Gemini 3.7 Flash to half price, Cerebras pushed OpenAI's GPT-5.6 Sol to 750 tokens a second, and DeepSeek moved V4-Pro to general availability days before a four-fold price hike, while Alibaba open-sourced Qwen3.8-Max at 2.4 trillion parameters. Cognition's Devin coding agent neared a $1 billion revenue run rate, and Twitch quietly switched on Amazon AI training for every streamer by default. Some coverage reaches back over the last 72 hours for a quiet week.

AI Models ·SiliconANGLE

Gemini 3.7 Flash: Google halves prices three weeks after 3.6 flopped

AnalysisGoogle put out Gemini 3.7 Flash on August 13, roughly three weeks after Gemini 3.6 Flash landed flat, and cut the price in half to pull developers back. Input runs $0.75 per million tokens and output $3.75, down from 3.6 Flash's list rate, though that discount expires December 31 and doubles on January 1. The model is a tuning pass on 3.6 rather than a fresh training run, and it takes text, images, audio, and video across a million-token context. Google aimed it straight at OpenAI's GPT-5.6 Terra and Meta's Muse Spark 1.2 on coding, where a cheaper competent model changes which one a team wires into its build pipeline.

AI Models ·RITS

Alibaba open-sources Qwen3.8-Max, the largest downloadable model yet

AnalysisFor the first time, anyone with the hardware can download a 2.4-trillion-parameter frontier model. On August 13 Alibaba published the open weights for Qwen3.8-2.4T-A95B, the Max-tier flagship it had sold only through an API ten days earlier, uploading it to Hugging Face and ModelScope under the Qwen organization. It is a sparse mixture-of-experts model, 2.4 trillion parameters with 95 billion firing per token, and it runs on standard inference stacks like vLLM. Alibaba has never released a Max-class model's weights before. The backdrop: Chinese open-weight models pulled 41% of Hugging Face downloads this spring, ahead of American ones.

AI Models ·Unite.AI

DeepSeek V4-Pro goes GA, then quadruples output pricing days later

AnalysisThe cheap era at DeepSeek is ending on a schedule. On August 12 the Chinese lab moved DeepSeek-V4-Pro-0813 to general availability across its app, web, and API, closing a preview that ran since late April, then said that on August 16 peak-hour output pricing rises to $3.96 per million tokens from a flat $0.87. The model is a mixture-of-experts design (1.6 trillion parameters total, 49 billion active per request to hold down cost) tuned for agent work: writing code, running tools, stitching multi-step tasks. DeepSeek's own numbers put its top tier above 80% on SWE-bench Verified, a real-bug-fixing test, though no outside lab has checked the claim.

AI ModelsAI Agents ·Cerebras

Cerebras runs OpenAI's GPT-5.6 Sol at 750 tokens a second

AnalysisA frontier model that answers 2,500 hard questions in eleven hours instead of seventy-eight changes what you can build on it. On August 13 Cerebras, a chipmaker that fits a whole model onto one wafer-sized processor, switched on Ultrafast mode for OpenAI's GPT-5.6 Sol, pushing up to 750 output tokens per second. That is about 14 times the standard tier's pace, with no quality drop on OpenAI's own GDP-Val benchmark of economically useful knowledge work. The mechanism is clearing the memory-bandwidth wall that slows the decode step on ordinary GPU clusters. For agent loops that chain dozens of model calls, latency has long been the tax, and this trims it hard.

AI Industry ·TechCrunch

Twitch turned on Amazon AI training for every streamer by default

AnalysisEvery Twitch creator woke up already enrolled in training Amazon's AI, with no email, no pop-up, and the off switch tucked under a Security and Privacy menu labeled Training for Generative AI. Twitch announced the opt-out on August 12, which is roughly when people learned the default had been on the whole time: their streams, clips, chat, and images were already fair game for Amazon's models. Asked why it was not opt-in, Twitch's Jason Minton said the quiet part plainly, that if it were opt-in nobody would opt in. He also could not say what Amazon had already trained on before the toggle existed.

LLM Evals ·The AI Insider

Design Arena raises $7.9M betting humans judge AI better than benchmarks

AnalysisAutomated benchmarks keep getting gamed, so a startup is selling human taste at scale instead. Intelligence, which runs the Design Arena evaluation platform, raised a $7.9 million seed led by Index Ventures. Its users, now 5.5 million across 190 countries, pit AI-generated websites and images against each other in blind head-to-head votes, and the aggregated rankings get sold to model labs as training signal. A team of ten scaled that enterprise business from $5 million to $60 million in annual recurring revenue in six months. Founder Grace Li started it in 2025 after noticing her AI game engine made games that worked but were not fun, a gap no automated score caught.

AI Industry ·TechCrunch

Cognition eyes $40B valuation as Devin revenue nears $1B run rate

AnalysisThree months after raising a billion dollars at $26 billion, Cognition is already back with investors, this time at a $40 billion price. The maker of Devin, the autonomous coding agent, is reportedly tying the new round to hitting a $1 billion annualized revenue run rate, up from $492 million in May. Enterprise use is growing 50% month over month, with Mercedes-Benz, NASA, and Goldman Sachs among the names. CEO Scott Wu frames Devin as a tool for long-tail grunt-work rather than a replacement for engineers, a careful line to walk while selling a product whose whole pitch is that it writes the code itself.

AI Industry ·CNBC

TSMC posts record July sales as 2nm chips reach commercial production

AnalysisThe choke point for every AI chip just moved a step, and it showed up in TSMC's July numbers. The Taiwanese foundry that makes the flagship processors for Nvidia, AMD, and Apple reported monthly revenue of NT$467.58 billion, about $14.5 billion, up 44.7% from a year earlier and a record. The driver TSMC singled out is its 2-nanometer process (its most advanced, denser and pricier per chip) entering commercial production, with high-performance computing already two-thirds of sales. It raised full-year capital spending to as much as $64 billion. When the sole maker of the leading node runs hotter, the supply that gates AI compute gets tighter rather than easier.

AI Agents ·GitHub Changelog

Agent Plugins 1.0 ships GA across VS Code and Copilot CLI

AnalysisThe plugin format that AWS, Cursor's maker Anysphere, Microsoft, OpenAI, and Vercel agreed to back reached 1.0 and went generally available across GitHub's tools on August 12, landing in VS Code, the Copilot CLI, and the Copilot app. Agent Plugins bundles agent skills and MCP servers (MCP is the tool-connection standard Anthropic introduced) into one installable package defined by a single plugin.json file, so a plugin built once runs on any client that adopts the spec. Install one through the Copilot CLI and VS Code discovers it automatically. Getting rival vendors to ship the same format in the same week is the part that makes a standard stick.

AI Industry ·Available Law

Colorado's chatbot safety law takes effect, teeth arrive in 2027

AnalysisColorado's Chatbot Safety Act (HB 26-1263) became effective on August 12, though the parts with teeth arrive later. The law makes any operator of a public conversational AI service tell users plainly that they are talking to software, keep a documented protocol for when someone mentions self-harm, and, for minors, block sexually explicit content and drop engagement-reward mechanics that hook young users. Operators also cannot claim a chatbot's output equals a licensed professional's advice. The substantive duties do not bind until January 1, 2027, with annual reporting to the attorney general after that, so the clock starts in August and the real obligations land sixteen months out.

AI AgentsAI Models ·GitHub Changelog

Microsoft's MAI-Code-1.1 lands in Copilot at 73% below its last model

AnalysisMicrosoft keeps pushing its own models deeper into the tools it owns, and undercutting on price as it goes. On August 11 MAI-Code-1.1-Flash, a coding model built by Microsoft's in-house MAI team rather than by OpenAI, landed in GitHub Copilot at $0.20 per million input tokens and $1.20 output, roughly 73% below the model it replaces. It adds native vision, so it can read a screenshot or UI mockup alongside code, and Microsoft claims 25% faster token streaming and 22% better command-line performance. The move fits a clear aim: cheaper in-house coding models that hold their own against the flood of low-cost Chinese options, instead of leaning on OpenAI's.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack