AI Field Notes by Michael Nemtsev

Claude 5.1 Cache Cut | AI Field Notes #97

A pen whose nib is a lockpick opens a padlock while price tags fall nearby, suggesting cheaper AI models that can also break into systems.

Claude 5.1 landed with cache reads cut 75%, and Google's Gemini 3.8 Flash arrives two days later aimed squarely at coding, so the model you pick this week just got cheaper and more crowded. OpenAI says its unreleased Astra is the first model to cross its 'Critical' cyber threshold, able to find and weaponize zero-days on its own. AI security ran through the whole day: fake crawlers hunting .env files, CrowdStrike's attacker-and-defender model pair, and Anthropic reversing a data-retention policy after customer pushback. Cognition's Devin hit a $47 billion valuation and Dell booked a record $60.9 billion in AI server orders.

AI AgentsAI Industry ·Investing.com (Bloomberg)

Cognition's Devin hits a $47B valuation as coding-agent revenue nears $1B

AnalysisCognition, the startup behind the Devin coding agent, is raising about $1 billion at a $47 billion valuation, up from $26 billion in May and roughly $25 billion in April. The number that justifies the jump is revenue: annualized sales above $900 million, nearly double the $492 million it reported in late May. Investors reportedly offered close to $10 billion, so the round was oversubscribed many times over. Devin sells the promise of an agent that closes real tickets, and the market is pricing that promise as if the automation of routine software work is already underway rather than pending.

LLM Evals ·CNBC

OpenAI says Astra is its first model to cross the 'Critical' cyber line

AnalysisFor the first time, OpenAI has labeled one of its own models a 'Critical' cyber risk, the top tier of its Preparedness Framework, meaning Astra can in principle find unknown software flaws and build working exploits against hardened systems with no human guiding each step. The company said on September 1 it has paused internal work that does not meet tightened controls: isolated test environments, locked-down network access, stronger protection for the model weights. Access will go only to a vetted defensive coalition it calls Daybreak. A model that can autonomously weaponize a zero-day is now something a lab admits it built, then races to fence.

AI Models ·Crypto Briefing

Gemini 3.8 Flash: Google ships a coding-focused model 20 days after 3.7

AnalysisTwenty days after Gemini 3.7 Flash, Google unveiled 3.8 Flash on September 2, and the pitch is narrower than usual: better code synthesis, steadier multi-step agent runs, fewer hallucinations in long chats. Codenamed Skimaki internally, it spent August being graded on Jetski, Google's own coding environment where builds run against live engineering work before they reach the API. The cadence itself is the signal. Google is now iterating its cheap, fast tier faster than most teams can finish an evaluation, and the target is plainly the coding market Anthropic and OpenAI have been splitting.

AI Models ·VentureBeat

Claude 5.1: Anthropic cuts cache reads 75% while token prices hold flat

AnalysisCache reads on Claude just dropped 75%, to $0.25 per million tokens, while the headline prices held at $10 in and $50 out. Anthropic shipped Fable 5.1 on the public API on September 1 and kept Mythos 5.1, its stronger sibling, behind a trusted-access wall. For teams running coding agents that re-read the same context on every step, the company puts the real-world saving at roughly 25% on standard work and up to 45% on cache-heavy automation. The benchmarks moved too, 55.8% on Terminal-Bench 4.0 and 73.4% on CursorBench, but the number that changes a monthly bill is the cache line.

AI Models ·World Labs

World Labs unveils Atlas, a world model that rebuilds 3D scenes from a few photos

AnalysisFei-Fei Li's World Labs put out Atlas on September 1, a single model trained from scratch to handle text, images, video, and 3D together, which the company calls an omni world model. It generates up to a minute of 1440p video with frame-by-frame camera control, and it rebuilds real 3D scenes from as few as one to a handful of photos, beating tools built only for reconstruction. The interesting part is the robotics angle: Atlas turns ordinary video into simulations, the 'real-to-sim' step that lets a robot practice in a copy of a room before it touches the real one. Spatial intelligence is moving from a research pitch to a shipping model.

AI Industry ·Business Wire

Dell books a record $60.9B in AI server orders as backlog hits $95B

AnalysisDell booked $60.9 billion in AI server orders in a single quarter and recognized $16.4 billion of AI server revenue, double a year earlier, with the figures reported September 1. The backlog now sits at $95 billion, and management said the pipeline of deals it has not yet booked is several times that. Strip away the record-chasing and one thing is clear: the companies building AI are still spending on the metal to run it at a pace that shows no bend. Dell's infrastructure group grew 89% year over year, the kind of number that only comes from a buildout that has not peaked.

AI Industry ·Anthropic

Anthropic reverses a data-retention policy after enterprise customers push back

AnalysisAfter what it called a lot of feedback, Anthropic scrapped a contested plan to retain enterprise data and replaced it with a system announced September 1 called Enterprise Frontier Safeguards. The trade it offers is unusual: keep zero data retention, and let automated systems still watch for misuse, with the monitoring logs stored in the customer's own S3, Azure, or Google Cloud bucket under the customer's own keys. No Anthropic employee reads the traffic; the software flags cyber or bio misuse and sends the signal back to the customer. The company is not charging for it, and it rolls out in phases this fall.

AI Industry ·Help Net Security

Fake AI crawlers: scanners posing as ClaudeBot and GPTBot hunt for .env files

AnalysisBetween July 28 and August 23, scanners from 824 separate IP addresses crawled the web wearing the names of AI bots from Anthropic, OpenAI, Google, and Perplexity, and none matched the real crawlers' published address ranges. The security firm GreyNoise caught the disguise by a simple tell: the impostors never requested robots.txt, the file legitimate crawlers check first. What they wanted was credentials, probing for .env files, AWS keys, and private keys, the loose secrets that let an attacker into a cloud account. Dressing malware as a friendly AI bot is a bet that defenders now wave those user-agents through.

AI Industry ·Asharq Al-Awsat

US pushes 'Carolina Principles' at G20, urging governments to hold off on AI rules

AnalysisAt a G20 innovation meeting in Chapel Hill on September 1, the White House science adviser Michael Kratsios asked other governments to sign what he called the Carolina Principles, a pledge to reserve new regulation for novel considerations and to skip building fresh AI regulatory bodies. His message was that policymakers should not treat every new technology as a first-of-its-kind problem. Elon Musk, at the same gathering, defended AI data centers and attacked the EU's rules. The pitch is a light-touch doctrine aimed squarely at the EU AI Act, and it asks the rest of the world to compete on permissiveness.

AI Industry ·SiliconANGLE

CrowdStrike builds attacker-and-defender AI models with Nvidia to patch on a loop

AnalysisCrowdStrike used its Fal.Con conference on September 1 to show SafeMind, two models that fight each other on purpose. Red Tempest plays the attacker, trained partly on 15 years of the company's incident-response files, and probes a digital twin, a simulated copy, of a customer's network for ways in. Blue Solano plays defense and patches what the attacker finds, and the loop runs until Red Tempest stops finding holes. Both are built on Nvidia's open Nemotron models. The premise is blunt: if attackers are about to get AI that moves at machine speed, the only workable defense is a defender that moves just as fast.

AI Industry ·TechCrunch

AfterQuery becomes Y Combinator's fastest unicorn at a $3.2B valuation

AnalysisFive months after a $30 million round valued it at $300 million, AfterQuery raised at $3.2 billion, which Y Combinator says is the fastest any of its startups has reached a billion-dollar valuation. The founders are 22 and 23. What they sell is not a model but the fuel: high-end human reasoning data used to train one, with frontier labs including OpenAI among its named customers. The pattern is worth sitting with. The labs racing to automate knowledge work are paying a premium for humans who can still demonstrate it, because a model cannot learn to reason from data a machine generated.

AI Models ·Runway

Runway's Solaris generates working software interfaces frame by frame, no code

AnalysisRunway, known for AI video, unveiled Solaris on August 31 and put it in a new category it named Interface World Models. Instead of writing code, Solaris renders a working interface as live video, generating each frame in response to your clicks and drags, with a language model reasoning about what should happen next. Runway says it beats frontier language models at producing new interfaces on measures of structure and consistency. It is research for now, early-access only, with no API or pricing. The idea underneath is strange and worth watching: an interface drawn on demand rather than built, compiled once, and shipped.

AI Models ·Meta

Meta's Muse Voice Transcribe beats rivals on real-time speech at $0.18 an hour

AnalysisMeta's Superintelligence Labs shipped Muse Voice Transcribe on September 1, a real-time speech-to-text model that also separates who is speaking and knows when a sentence ends, all in one model. On a streaming accuracy benchmark it posts a 3.1% word error rate, ahead of Cartesia, ElevenLabs, and Google's live transcription. It handles more than 70 languages, follows mid-sentence switches between them, and holds up across hour-long recordings with 20 or more speakers. The API runs $3.00 per thousand audio-minutes, or about 18 cents an hour. Live transcription that was a premium feature is turning into a cheap commodity call.

AI Agents ·9to5Mac

Perplexity splits its Mac agent between cloud and on-device models to keep data local

AnalysisPerplexity turned an ordinary Mac into part of its own infrastructure on September 1. Its Computer agent now splits a single task between a frontier model in the cloud and a smaller model running on the machine, so confidential files never leave the laptop. An on-device classifier scans each task first, swaps names, addresses, and account numbers for stand-ins before anything goes to the cloud, then restores them in the answer. It runs on Apple Silicon with at least 24GB of memory, using local models like Gemma 4 and a 35-billion-parameter Qwen. The local work costs no cloud credits, which is the quiet part.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack