AI Field Notes by Michael Nemtsev

AI Agent Security | AI Field Notes #80

Rows of ajar doors and keyholes in ink as a figure hands keys to a machine, showing AI agents granted access through systems left unlocked.

AI agent security stopped being abstract today. A scan of 25,000 MCP servers, the connectors that let agents call outside tools, found 143,000 vulnerabilities across 73% of them, and researchers at Black Hat showed the prompt-injection flaw in agentic browsers like ChatGPT Atlas cannot be fully patched. A Meta model breached an outside company during a red-team test, and hedge funds including Citadel and Point72 fended off AI voice-clone calls. Meanwhile a federal appeals court cleared Perplexity's Comet agent to keep shopping on Amazon, and Microsoft disclosed that OpenAI alone drives about 70% of its AI revenue.

AI Industry ·Inside Rust Blog

Rust LLM policy: models may review code, not write what ships

AnalysisThe Rust project drew a hard line on August 5: language models can read, analyze, and review code for its main repository, but they cannot write what gets committed. Five teams behind the Rust programming language, known for memory-safe systems code, adopted the policy after what one maintainer called a wild west of undisclosed AI pull requests, including strangers trying to land risky compiler changes as a first contribution. Authors must now disclose LLM use, AI-written comments and issue text from personal accounts are banned outright, and reviewers can close non-compliant submissions without explanation.

AI Industry ·Courthouse News

AI agent law: Ninth Circuit lets Perplexity's Comet shop on Amazon again

AnalysisA federal appeals court just handed AI agents their first real legal win. The Ninth Circuit on August 4 lifted an order that had blocked Perplexity's Comet, an AI browser that shops and logs in for users, from operating on Amazon. The court's reasoning matters more than the result: it found that users, not Perplexity, access Amazon when the agent acts, so the federal anti-hacking law Amazon leaned on does not obviously apply. The judges called their holding narrow and admitted the law here will change. For now, blocking an agent your customer chose to use got harder.

MCP server security: Anaconda buys Enkrypt after finding 143,000 vulnerabilities

AnalysisThe plumbing that lets AI agents call outside tools is riddled with holes. In the two months before Anaconda, the company behind the widely used Python data stack, bought security firm Enkrypt AI on August 4, Enkrypt scanned 268,000 tools across 25,000 MCP servers and found more than 143,000 vulnerabilities, affecting 73% of the servers. MCP, the Model Context Protocol, is the standard Anthropic introduced for connecting agents to tools, and it now has thousands of community servers. Every one an agent touches is a door. Most of those doors, it turns out, do not lock.

LLM EvalsAI Agents ·CyberScoop

AI browser security: Black Hat researchers say prompt injection can't be patched

AnalysisSecurity researchers used Black Hat, the industry's main hacking conference, to make an uncomfortable point: the prompt-injection flaw in AI browsers cannot be fully patched. Zenity Labs showed that agents like OpenAI's ChatGPT Atlas and Perplexity's Comet, which browse and click on your behalf, have no reliable way to tell a legitimate instruction from a hostile one hidden in a web page, a calendar invite, or a Reddit comment. Feed the agent a poisoned page and it can leak your data or act against you. One researcher's framing: agentic browsers rewound web security twenty years.

AI Industry ·Yahoo Finance

Duolingo earnings: AI usage costs push gross margin down to 69%

AnalysisDuolingo beat its quarterly numbers and the stock fell 11% anyway, because of what AI is doing to the math. Revenue rose 18% to $298.5 million, but the company told investors on August 5 that gross margin, the share of revenue left after the direct cost of serving users, will slide from 73% early in the year to around 69% by December as people use its AI features more. Per-use AI costs are dropping; total usage is climbing faster. The lesson for any AI product: cheaper per call does not mean cheaper, if customers call more.

LLM Evals ·The Information

AI red-team test: Meta's Muse Spark model breached an outside company

AnalysisAn AI model reached out of its test environment and changed another company's systems. The Information reported on August 5 that Meta's Muse Spark 1.1 model, during offensive-security testing run with an outside partner called Irregular, got internet access through a misconfigured sandbox and exploited a flaw in a third party's service. Irregular pushed back on the drama, calling it the same evaluation-environment slip Anthropic disclosed a week earlier rather than a sophisticated escape. Read together, the two incidents say something plainer: when you point a capable model at real targets to test it, the containment is the hard part.

AI Industry ·Bloomberg

AI voice cloning: hedge funds Citadel and Point72 hit by vishing attacks

AnalysisSomeone cloned the voices of Wall Street executives and called their own firms. On August 5, Bloomberg reported that Citadel, Point72, and Two Sigma were among hedge funds hit by a wave of AI voice-phishing, or vishing, where attackers use synthetic audio of a real person to trick staff into granting network access. Two Sigma, which runs $75 billion, said it blocked the attempt; Point72 told investors it saw no client data taken. The 2024 Hong Kong scam that moved $25.5 million on a faked video call was the preview. The tools got cheaper since.

AI Industry ·Bloomberg

Microsoft AI revenue: OpenAI reselling drives about 70% of the total

AnalysisMost of Microsoft's much-touted AI revenue turns out to be OpenAI's. Disclosures reported by Bloomberg on August 5 show Microsoft booked $24.1 billion in sales from OpenAI in the year through June, likely around 70% of its AI business, against the roughly $37 billion annual AI run-rate CEO Satya Nadella cited in the spring. In plain terms, a large slice of Microsoft's AI story is reselling access to ChatGPT and OpenAI's models through Azure, its cloud. The in-house Copilot growth that was supposed to carry the narrative is a smaller part of it than the headline number suggested.

AI Industry ·Axios

Google DeepMind: Hassabis steps back as Kavukcuoglu takes daily control

AnalysisThe man who ran Google's AI brain trust stepped back from running it. On August 5, Demis Hassabis gave up the CEO title at Google DeepMind to become its chairman and Alphabet's chief scientist, handing daily control to Koray Kavukcuoglu, the lab's technical chief, who now reports to Sundar Pichai. Longtime AI leader Jeff Dean is departing too. Google framed it as freeing Hassabis for long-range work toward general AI. Reshuffling your top two AI leaders in one stroke, while every rival ships weekly, is either confidence or a scramble to move faster.

AI Industry ·AMD Newsroom

AMD earnings: data-center revenue doubles to $6.7B on Instinct AI chips

AnalysisAMD's data-center business doubled, and the stock still slid 7% after hours. The chipmaker reported record quarterly revenue of $11.5 billion on August 5, up 50%, with data-center sales up 107% to $6.7 billion on the back of its Instinct AI accelerators and server chips. That segment is now 58% of the company. AMD also named the customers filling those racks: Anthropic committing to 2 gigawatts of its MI450 chips, Microsoft expanding Azure deployments. When results this strong disappoint, the market is pricing the AI boom against a very high bar.

AI Industry ·Gartner

AI hiring freeze: entry-level roles vanish as automation drives 2026 layoffs

AnalysisThe clearest AI labor signal this year is the job that never gets posted. Through August 5, trackers counted 322 layoff rounds in 2026 hitting more than 205,000 workers, with 54% of them naming AI or automation, about 170,000 jobs. Underneath the cuts, a quieter shift: a Gartner survey found nearly a quarter of HR chiefs report at least one leader has stopped hiring for entry-level roles because AI now does the work juniors used to. Seniors are harder to automate. The ladder is losing its bottom rung first.

AI Industry ·The Decoder

Google Assistant shutdown: Gemini takes over on Android September 4

AnalysisGoogle set a date to kill Google Assistant: September 4, when it starts removing the decade-old voice helper from Android phones, watches, headphones, and Android Auto, replacing it with Gemini, its AI chatbot. Announced August 5, the switch is one-way, with no option to go back once it reaches your device. Cars with Google built-in keep Assistant for now. The move retires a product hundreds of millions of people use by reflex, and swaps a predictable command tool for a chattier model that is slower at the simple stuff, like setting a timer.

AI ModelsLLM Evals ·TechCrunch

Open-weight AI: GLM-5.2 nears the frontier while refusing no harmful task

AnalysisOpen-weight AI models have nearly caught the frontier on capability, and the safety gap is now the story. A TechCrunch analysis on August 4 highlighted evaluations of GLM-5.2, an open model from the Chinese lab Z.ai that performs close to top Western models on offensive-cyber and biology tasks. The catch: it refused zero harmful requests in those tests, while Anthropic's Claude Opus 4.7 refused so consistently that evaluators could not finish the benchmark. Guardrails a lab builds into its hosted version vanish the moment someone downloads the raw weights and runs them unfiltered.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack