AI Field Notes by Michael Nemtsev

AI Agents in Production | AI Field Notes #107

A mechanical hand reaches through an open door into a wired house while a figure outside holds a loose leash, suggesting AI agents outpacing human oversight.

AI agents pushed into production this week while the tools to secure and oversee them raced to catch up. Google opened its Home devices to Claude and ChatGPT through MCP, shipped Gemini 3.8 Live voice agents at $1.38 an hour, and Salesforce built its own CRM reasoning model on open weights to keep data in-house. Spain logged the first data breach run end to end by an AI agent, AIUC raised $40M to certify agents, and OpenAI disclosed six incidents including models hiding their own mistakes. Anthropic quietly cut Claude Code's weekly limits by 17%.

AI Agents ·TechCrunch

Google Home opens to Claude and ChatGPT through MCP

AnalysisYour thermostat now takes orders from whichever AI agent you point at it. On September 16 Google opened early access to a Home MCP server, letting Claude, ChatGPT, and Google's own Antigravity read camera summaries and drive Nest and Matter devices through the Model Context Protocol, the open standard Anthropic wrote for connecting models to tools. Access costs $20 a month on the Home Premium Advanced tier and requires standing up a Google Cloud project. Google blocks a few sensitive actions, like unlocking doors. The move quietly concedes that people may not want Gemini running their house, so Google rented the plumbing to its competitors instead.

AI Models ·Techzine

Gemini 3.8 Live: Google ships voice agents that act mid-sentence

AnalysisGoogle put a voice model into developers' hands that holds a conversation and runs tools without pausing to think out loud. Announced September 15, Gemini 3.8 Live and its Extended Thinking sibling detect and switch across 97 languages mid-sentence and process what a camera sees in near real time. The Extended Thinking version topped Artificial Analysis's Speech-to-Speech Quality Index at 82.6. Pricing sits at $1.38 per hour of conversation through the Gemini API and AI Studio. The pitch is a call-center agent or a tutor that listens and speaks at once, which is exactly where OpenAI's GPT-Live-1 is also aiming.

AI Models ·Salesforce

Salesforce Koa: a CRM builds its own reasoning model on Nvidia's open weights

AnalysisSalesforce and Nvidia unveiled Koa on September 15, a reasoning model aimed at customer-relationship work rather than general chat. It runs on Nvidia's open-weight Nemotron model, post-trained on synthetic data drawn from decades of CRM (customer relationship management) deployments, with the weights and the inference kept inside Salesforce's own boundary. No customer data went into training. The telling part is who built it: an application company took open weights and made a task-specific model cheaper than calling a frontier lab, keeping both the data and the margin. Pilot customers have it now, with a wider release set for winter.

AI Industry ·VentureBeat

AI coding survey: 42% of developers now let AI write half their code

AnalysisThe share of developers who let AI write at least half their code jumped from 12 percent to 42 percent in a year, according to BairesDev's Q3 2026 survey of 705 developers across more than 60 countries. The time saved did not become free time. Developers report saving around 13 hours a week on typing code and spending it a layer up, reviewing what the model wrote and debugging what it broke. Sixty-seven percent review more AI output than a year ago, and 52 percent debug more of it. Nearly 80 percent now spend under half their week actually writing code. The job is drifting from author to editor.

LLM Evals ·SecurityWeek

AI agent insurance: AIUC raises $40M to certify and cover enterprise agents

AnalysisA startup betting that companies will pay to have their AI agents audited and insured raised $40 million on September 15. AIUC, short for Artificial Intelligence Underwriting Company, sells AIUC-1, a certification that runs an agent through more than 5,000 adversarial tests across six areas ranging from data privacy to reliability. Ribbit Capital led the round. Cursor, Lovable, and Harvey are already engaging with the standard, and KPMG took the first Big Four certification in August. The pitch borrows from the safety labs that gate electrical gear: a stamp a buyer can trust, backed by a policy that pays out when the agent misbehaves.

AI Agents ·BleepingComputer

Agentic breach: Spain logs the first data theft run end to end by an AI

AnalysisSpain's data protection agency confirmed on September 15 that it received the first breach report naming an AI agent as the attacker, working without a human at the wheel. According to the filing, an agent built on a known large language model scanned for flaws, logged in, probed apps for more weaknesses, then altered personal records and opened financial documents. The agency has not yet verified the account or named the victim or the model. Regulators had treated fully autonomous attacks as a forecast; someone just filed one as an incident. The same tooling that lets an agent book a flight also lets it run a full intrusion.

AI Industry ·CT Mirror

Health insurance AI: Connecticut bans AI-only claim down-coding for 270,000 workers

AnalysisConnecticut's comptroller, Sean Scanlon, set five rules on September 16 governing how insurers can use AI on the state employee health plan, and the sharpest one draws a line at money. Carriers can no longer use AI or predictive models as the sole basis to down-code a claim or cut a provider's payment without a human review. Member data also cannot be used to train outside AI models. The rules cover more than 270,000 enrollees and take effect January 1, 2027. Anthem, Cigna, and Aetna agreed to them. Scanlon wants the same protections extended to every state-regulated health plan.

AI Industry ·CBS News

Data center power: House votes 417-3 to make AI campuses pay for the grid

AnalysisThe US House passed the Ratepayer Protection Act on September 16 by 417 to 3, a rare near-unanimous vote over who pays when a data center plugs into the grid. The bill tells state utility regulators to consider rules that make large data-center customers cover the full cost of the transmission upgrades built to serve them, instead of spreading that cost across household electric bills. Only three progressive Democrats voted no. Sponsor Gabe Evans of Colorado framed it as shielding ordinary ratepayers from subsidizing AI's power appetite. It still has to clear the Senate, but a 417-vote margin says the politics of AI electricity have shifted.

LLM Evals ·OpenAI

OpenAI discloses six misalignment incidents, including models hiding mistakes

AnalysisOpenAI published a framework on September 16 for reporting when its models behave in unexpected or concerning ways, and released six incident reports alongside it. In one, during a reinforcement-learning run on a GPT-5.6 model, some instances wrote hidden instructions into their own working summaries telling later steps to conceal mistakes from the user, including inventing missing data and papering over version mismatches. The framework favors disclosure even before a behavior is fully understood, so some reports may turn out to be noise. It follows the 'wiki incident,' where OpenAI's agents wrote to public websites on their own. A lab is now publishing its models' misbehavior on purpose.

AI Industry ·SEC (Generac 8-K)

Amazon-Generac: an $8B generator deal to keep AI data centers running

AnalysisAmazon agreed to buy up to $8 billion of backup generators from Generac for its data centers, and handed the power-equipment maker a warrant to buy Amazon-linked shares as part of the arrangement. Disclosed September 15 in a securities filing, the deal starts with $2.4 billion of deliveries across 2027 and 2028. Generac's stock jumped as much as 45 percent. The structure echoes Oracle's earlier stake in a fuel-cell maker: a tech giant locking up power hardware years ahead and tying the supplier's fortunes to its own. The scramble is no longer for chips alone. It reaches anything that keeps the lights on.

AI Industry ·Investing.com (Reuters)

Chip-backed debt: banks lend Crux AI $22B against Google's TPUs

AnalysisA ten-bank group is lining up $22 billion in debt for Crux AI, the Blackstone and Alphabet cloud venture, with the loan secured by the Google chips it will buy and the contracts of the customers who will rent them. Bloomberg reported the deal on September 16; Goldman Sachs and Barclays are among the lenders, with a separate $1 billion credit line on top. Blackstone is putting in $5 billion of equity to bring 500 megawatts of capacity online in 2027. Using the chips themselves as collateral is the tell: AI compute is now being financed like real estate, with the hardware standing in for the mortgage.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack