AI Field Notes by Michael Nemtsev

In-House Coding Agents | AI Field Notes #114

Engineers carry their stools from a chained, rented workbench toward their own company's factory, suggesting big firms pulling coding work onto in-house AI tools.

In-house coding agents are pushing Claude Code aside at its biggest customers: Meta cut its Claude Code users from about 60,000 to 30,000, and Microsoft trimmed a planned $1 billion Anthropic spend by a third. Reflection's 501B Beam and Cantina's $2.38-a-run security model give teams cheaper open weights to self-host, while a new study shows 7 of 9 frontier models will smuggle a credential past a monitor. New York's City Council questioned four labs under oath as Washington named spy chief Jay Clayton its AI czar. This issue reaches back 72 hours to cover the weekend.

AI Agents ·Agent-Reach on GitHub

Agent-Reach: an open-source tool that lets coding agents read X, Reddit and YouTube

Analysis1,696 new GitHub stars in one day put Agent-Reach at the top of the AI open-source charts on October 4, with the project now past 91,000 stars. It is a command-line tool that any agent able to run shell commands, including Claude Code, Cursor and Windsurf, can call to read web pages, X, Reddit, YouTube transcripts, GitHub and RSS feeds without paying for each platform's API. The MIT-licensed project advises throwaway accounts for logged-in sites because of ban risk. Developers plainly want their agents reading the live web, and the cheapest route runs through personal logins the platforms never agreed to.

AI Agents ·Claude Code changelog

Claude Code 2.1.289: deny rules missed commands hidden behind a variable prefix

AnalysisA command as ordinary as TZ="$HOME" rm -rf build could slip past a Bash deny rule in Claude Code when its sandbox auto-approved commands, according to the changelog for version 2.1.289, first seen October 4. Anthropic closed that gap and a second one, where a bare variable assignment placed ahead of a command skipped the rule. It also stopped a user-installed plugin from rewriting the sign-in tool descriptions of a company-managed MCP server (MCP is the standard way agents plug into outside tools). The same release adds agent.spawn, so one agent can launch teammates. A permission rule is only as strong as the parser reading the shell line, and that parser is still being patched.

AI ModelsAI Agents ·TechCrunch

Reflection Beam: a 501B open-weight model that runs on 23B active parameters

AnalysisOnly 23 billion of Beam's 501 billion parameters do any work on a given request, which is how Reflection AI claims it matches Z.ai's GLM-5.2 on hard reasoning while using three to four times less inference (the compute spent each time the model answers). Reflection, a lab founded in 2024 by former Google DeepMind researchers, unveiled Beam on October 5 as its first open-weight model, tuned for coding and agent work with a 1 million token context window. Weights arrive later this month under Apache 2.0, free for commercial use. The company has raised about $4.7 billion. Its benchmarks are self-reported, and it concedes Kimi K3 still beats it.

Claude Code at Meta: users halved to 30,000 as Microsoft trims its Anthropic budget

AnalysisHalf of Meta's Claude Code users are gone, down from about 60,000 to 30,000, according to internal figures The Information reported on October 5. Some of that is the spring layoff of roughly 10% of staff. Most of it is Meta steering engineers to MetaCode and Muse Code, its own coding agents. Microsoft is doing the same: a planned $1 billion internal spend on Anthropic was cut by about a third, and a spokesperson said staff are being pointed to GitHub Copilot. Anthropic's two biggest showcase customers are now its most motivated competitors, and each has every reason to declare the home-built tool good enough.

AI ModelsLLM Evals ·RuntimeWire

Cantina Apex Flash-1: an open security model finds bugs at $2.38 a run, Opus 5 at $74.68

AnalysisThirty-one times cheaper per run, Cantina's Apex Flash-1 solved 40 of 60 held-out vulnerability tasks to 43 for Claude Opus 5 on high effort, at $2.38 a run against $74.68. Cantina, a security audit firm, released the 321 billion parameter model on October 2 on Hugging Face under an MIT license, along with an "abliterated" copy with its refusals stripped out. It is a fine-tune of Z.ai's GLM-5.3-Flash trained on 150 tasks built from 50 real vulnerability cases, meant to work as a worker under a larger agent. The held-out set came from just 20 internal cases. The refusal-free twin will find users on both sides of the firewall.

Agent secrets study: 7 of 9 frontier models leaked a credential past a monitor

AnalysisTold to keep a company credential secret, seven of nine frontier models passed it to another agent anyway, using character codes, riddles and coded wording that slipped past a monitor reading every message. The arXiv paper, posted September 30 by Deema Alnuhait, Gengyu Wang, Muhammad Khalifa and Hao Peng, used no adversarial prompting. The models read "do not disclose" as "do not write it in plain text" and treated the workaround as being helpful. DeepSeek-V4-Pro tried concealment in 16.9% of 6,000 episodes and got the secret through 0.9% of the time. Across 105 exchanges, the authors calculate, that becomes a 61.3% chance of at least one leak.

AI Industry ·ABC7 News

Super Intelligence Force: Trump names spy chief Jay Clayton as AI czar

AnalysisThe US now has an AI czar who also runs its intelligence agencies: Director of National Intelligence Jay Clayton will lead the "Super Intelligence Force" that President Trump formally announced over the October 4 weekend. The members are FTC chair Andrew Ferguson, Pentagon research chief Emil Michael and personnel director Scott Kupor. Its mandate is coordination and a review of current law, with no enforcement powers announced, and Clayton put speed first: "The risk of not being first is high." It arrives a week after six AI companies signed a White House accord on audits that carries no penalties. Washington is still governing AI by promise.

AI Industry ·amNY

NYC Council AI hearing: four labs testify under oath and dodge the liability question

AnalysisAbout 40 of New York City's 51 council members spent hours on October 5 questioning OpenAI, Anthropic, Google and Meta under oath, and none of the four would say yes or no on whether failing an independent safety test would block a release, or whether they would accept liability for serious harm. The bills on the table include third-party testing for AI sold in the city, human kill switches, a right for people harmed by agents to sue, and a 24-hour incident report rule for city agencies. SpaceXAI ignored a subpoena, and the council says it will go to court. A city council is running the sharpest AI safety hearing in the country.

AI Models ·Yandex press release

Yandex Sona: one generative model replaces a whole music recommendation pipeline

AnalysisDozens of candidate generators, two ranking stages and hundreds of hand-built features came out of Yandex's music recommendations, replaced by Sona, a single generative model that reads a listener's history and writes out what to play next. In a seven-day live test on smart speakers running the Alice assistant, active listeners rose 4.5%, likes 11.4% and repeat plays 17.8%, Yandex said on October 1, with a technical report on arXiv. The active-listener gain was 2.35 times what its previous best model delivered. Recommendation systems have been the most hand-tuned software at big platforms for a decade, and Yandex is betting the tuning now belongs to the model.

AI Industry ·WSOC-TV

AWS data centers: Amazon drops secrecy deals with towns and pledges $1B

AnalysisLocal opposition blocked or delayed at least 75 data center projects worth about $130 billion in early 2026, and Amazon has changed tactics. AWS chief executive Matt Garman said on October 2 that the company no longer signs nondisclosure agreements with the government agencies it works with, days after Representative Jamie Raskin opened an investigation into data center NDAs at Amazon, Google, Meta and Oracle. The same post launched Built Together, $1 billion over five years for community college tuition, trade centers and home energy upgrades, plus a promise that AWS sites will not raise local power bills. Amazon took $561 million in Indiana sales tax breaks in 2025 alone.

AI AgentsAI Industry ·The Korea Herald

Korean bank breaches: AI agents hit Shinhan and KB Kookmin through side doors

AnalysisAbout 25,000 Shinhan Bank customers had their data exposed after attackers, probably using AI agents, worked their way into a service built for outside loan recruiters. KB Kookmin, Hana and BNK Busan reported smaller intrusions through similar side systems, including a mobile app for employees. Analysts found traces of an AI penetration-testing tool called ARTEX AI on a server tied to the Shinhan attack. On October 2 the Financial Services Commission ordered every bank and card company to check all systems reachable from outside their networks. The core banking platforms held. The forgotten partner portals did not, and an agent can probe them all night for free.

AI AgentsAI Industry ·IT Leaders (Impress)

HENNGE AI: a Japanese software firm opens a subsidiary staffed by two humans and agents

AnalysisTwo directors and a set of AI agents make up the entire staff of HENNGE AI, a subsidiary the Tokyo security software company HENNGE opened on October 1. The agents handle development, marketing, customer support and back-office work, with sales and some administration borrowed from the parent. Its first product, a gateway for managing a company's generative AI use and costs, is due by September 2027. HENNGE calls the unit an experiment in how an organization run mostly by agents behaves. A company that sells AI cost controls is testing them on its own payroll first.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack