AI Field Notes by Michael Nemtsev

AI Agent Tooling | AI Field Notes #82

Hands fit rival tool-blocks into a single shared socket above a stamped chip and a padlocked purse, as agent tools converge over consolidating hardware.

AI agent tooling moved fast this week: two terminal coding agents shipped the same day, and Amazon, Microsoft, OpenAI, Cursor and Vercel agreed on one plugin format that bundles MCP servers and skills so you build a tool once instead of porting it per client. Prime Intellect's free Prime Agent edged the human baseline on a reasoning test while Meta's Muse Code couples its own model to your terminal, and Cloudflare gave agents real wallets with weekly caps. Underneath, the inference-chip race sharpened: AMD bought model-specific silicon startup Taalas, Anthropic began building chips in-house, and memory makers pitched an eight-times jump to clear the bottleneck. Jeff Dean left Google to automate the scientific method itself. This issue spans the last 72 hours.

AI AgentsAI Models ·The Register

Meta Muse Code: a terminal coding agent wired to its own Muse Spark 1.2 model

AnalysisMeta put a coding agent in the terminal and bolted it to a model it controls, Muse Spark 1.2, priced at $1.25 per million input tokens and $4.25 per million output with a one-million-token context window. Muse Code plans changes, writes code, and validates the result across a repository, keeping an append-only log of every model call and edit that Meta calls replay-exact, so a crashed run restarts where it stopped rather than from zero. Your working copy stays untouched until you approve. On Terminal-Bench and DeepSWE, two agentic coding tests, it clustered with Opus 5, GPT-5.6 and Gemini 3.6 with no clear winner. The bet is coupling: one team owning the model and the harness together.

AI Agents ·Agent Plugins

Agent Plugins: Amazon, Microsoft, OpenAI, Cursor and Vercel back one format

AnalysisBuild a tool for an AI agent once and it now runs everywhere the standard reaches, instead of getting rewritten for each client. Five companies, Amazon, Microsoft, OpenAI, Cursor and Vercel, seated a joint committee behind Agent Plugins, a vendor-neutral format that bundles agent skills and MCP servers (Model Context Protocol, the wiring that lets an agent call outside tools) into one portable package. It leaves MCP in place and wraps it, plus skills, in a single installable unit. The telling part is who signed: rivals agreeing on plumbing usually means the plumbing just became a moat none of them wants to lose.

AI Industry ·TechCrunch

Anthropic builds its own chip team, courts Samsung to make the silicon

AnalysisAnthropic is staffing an in-house silicon team to co-design chips with its own models, confirming a Business Insider report to TechCrunch, and is reportedly courting Samsung to manufacture them. The lab already buys compute from AWS, Google, Nvidia and AMD, and has decided renting is not enough as demand for Claude climbs. It joins a familiar pattern: OpenAI has a Broadcom-built chip in the works, Google has its TPUs (tensor processing units, its own AI accelerators), and Meta has MTIA. When the companies renting the most compute all start designing their own, the message to their landlords is plain.

AI Industry ·Benzinga

AMD buys Taalas: model-specific chips its maker clocks at 17,000 tokens a second

AnalysisAMD is buying Taalas, a Toronto startup that bakes a single model directly into silicon instead of running it on a general-purpose GPU, trading flexibility for raw inference speed (inference being the cost of running a model to answer a user, separate from training it). Taalas claims its HC1 chip pushes about 17,000 tokens per second per user against roughly 350 on Nvidia's Blackwell in its own testing, a gap wide enough to distrust and wide enough to matter. Terms were undisclosed. Fresh off doubling data-center revenue last quarter, AMD is buying an attack on the part of Nvidia's lead customers feel most: the cost to serve a model.

AI Agents ·Cloudflare Blog

Cloudflare Wallets: AI agents get spending money with weekly caps and allow-lists

AnalysisAI agents can now hold and spend real money under rules a person sets, after Cloudflare launched Wallets. A human funds an account wallet and delegates to a virtual wallet the agent uses through an API key, capped at, say, a hundred dollars a week, with per-transaction limits and a merchant allow-list. Payments ride the x402 protocol, which attaches a charge to an ordinary web request, and anomalous spending pauses for human review before more funds flow. Optional readable identifiers let a seller confirm which agent is buying. The awkward thing it says out loud: agents were already trying to pay for things.

AI Agents ·Open Source For You

Prime Agent: an open MIT coding harness that edged the human bar on ARC-AGI-3

AnalysisAn open, MIT-licensed coding harness scored 95.5 percent on ARC-AGI-3, a reasoning benchmark where the human expert baseline sits at 95.4, and wrote working emulators for the SEGA Genesis and Game Boy Color from scratch with no reference code. Prime Intellect built Prime Agent to bring your own model: a subscription account, a raw API, or a local model on your own machine. It runs tools as Python in a live kernel, spins sub-agents as asynchronous calls, and rewrites its own prompts and skills through a refine command. It also runs with your user permissions rather than sandboxed, which is power and a loaded gun in one.

AI Industry ·Implicator AI

Jeff Dean leaves Google to start Discovery Loop with three top researchers

AnalysisAlphabet shares fell more than five percent the day Jeff Dean, Google's 30th employee and the architect behind much of its AI stack, walked out after twenty-seven years to start a company. He took Sanjay Ghemawat, Oriol Vinyals and Quoc Le with him. Their venture, Discovery Loop, is a public benefit corporation built to automate the scientific method itself: have AI propose an experiment, build what it needs, run it, judge the result, and loop. Khosla, Radical Ventures and, oddly, Alphabet are backing it, with Alphabet supplying first-year compute. When your founders leave to disrupt research and you fund their exit, the org chart is telling on you.

AI Industry ·Allen Institute for AI

Ai2 triples its Hugging Face storage to 2 petabytes for fully open models

AnalysisFully open AI research needs somewhere to live, and Hugging Face just tripled the room, expanding the Allen Institute for AI's storage on its hub to nearly two petabytes and lifting the download rate limits that throttled big pulls. Ai2 ships more new artifacts a year than any other group Hugging Face tracks, over 900 models and 1,200 datasets, and it publishes the whole trail: training data, intermediate checkpoints, evaluations. Its releases have been downloaded more than fifty million times since spring 2024. Open weights make the headlines; the dull plumbing that lets anyone actually pull them decides whether open stays usable.

AI Industry ·Distill Intelligence

Memory becomes AI's bottleneck: Samsung's zHBM claims an 8x jump

AnalysisThe scarce resource in an AI server is shifting from the processor to the memory that feeds it, and the memory makers spent this week saying so out loud. Samsung unveiled zHBM, a next-generation high-bandwidth memory (HBM is the stacked memory glued next to an AI chip to move data fast enough to keep it busy) that it claims runs eight times faster than today's HBM5. SK Hynix and SanDisk floated a first standard for high-bandwidth flash, and Netlist signed a five-year memory pact with Samsung. A GPU computes only as fast as its memory can feed it, and that wall is where the next round of gains hides.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack