AI Field Notes by Michael Nemtsev

AI Agent Desktop Access | AI Field Notes #113

A mechanical hand reaches through a screen toward files as a locksmith padlocks the cabinet: AI agents gain desktop control while Apple tightens access.

AI agent desktop access moved in two directions at once: GitHub let Copilot click through Mac and Windows apps, while Apple said macOS Full Disk Access will soon need far more explicit permission because agents raise the risk. Google's Gemini 4 Argon scored 77.9% on a long coding test, then went to vetted cyber defenders before paying developers. Perplexity open-sourced a 27B decision model that answers yes or no for $0.04 per million tokens. StudentBench found AI GRE tutors matching human ones at 918 times lower cost.

AI ModelsLLM Evals ·Google Blog

Gemini 4 Argon: Google's new flagship goes to cyber defenders before developers

AnalysisA 77.9% score on DeepSWE v1.1, a test built from long real-world software engineering tasks, headlines Gemini 4 Argon, Google's first new frontier model since Gemini 3, announced September 30. It also leads CWE-bench at 68% for fixing security flaws, lifts the output ceiling from 64,000 tokens to 1 million, and costs $2 per million input tokens and $10 per million output at an introductory rate that later doubles to $4 and $20. Paying API customers have to wait. Argon is live only for Google's Fairwind Program of vetted security teams and agencies, and Bloomberg reports some Google staff doubt it codes as well on real work as the benchmarks suggest.

AI Agents ·GitHub Changelog

Copilot computer use: GitHub's agent can now click through your desktop apps

AnalysisOld software with no API just lost its best defence against automation. GitHub put computer use into public preview on October 1, letting Copilot CLI and the Copilot desktop app read screens, click, type, scroll and drag across macOS and Windows programs. Admins can block it through managed settings. The release notes hedge hard: GitHub tells developers to reach for APIs, MCP servers (a standard plug that connects agents to outside tools) or terminal commands first, and warns about misclicks, stalls and sensitive on-screen data leaking into the agent's context. A vendor that ships a feature and suggests trying something else first is telling you how ready it is.

AI Agents ·TechCrunch

macOS Full Disk Access: Apple tightens the gate as AI agents raise the risk

AnalysisOne checkbox in macOS settings hands an app every file, email, message and browser history on the machine, and Apple now says AI agents make that checkbox too easy to tick. On October 2 the company said it will add controls so users can grant Full Disk Access only through "very explicit user action." The trigger was a bad fortnight: an Inc. columnist said Meta's Muse agent read the columnist's private messages (Meta disputes it), and Wired documented a flaw in the ChatGPT Mac app that could have exposed sensitive data. Apple has not said which macOS version carries the change or what the new flow looks like.

AI Industry ·Anthropic

Claude Frontier Academy: Anthropic spends $100M to train 10,000 deployment engineers

AnalysisAnthropic's bottleneck is apparently people who can install its models, and it is paying $100 million to produce 10,000 of them by the end of 2027. Claude Frontier Academy, announced October 2, copies medical residency: an in-person bootcamp with simulated deployments, an assessment, then a 12-week residency running a real Claude project inside the engineer's own company. First cohorts are running in San Francisco, New York and London with staff from Accenture, Bain, Deloitte, McKinsey and Morgan Stanley, among others. A forward deployed engineer is a vendor's engineer embedded with a customer. Anthropic is training consultancies to play that role on its behalf.

AI IndustryAI Agents ·SiliconANGLE

Supabase buys Turso as agents now create 70% of its new databases

AnalysisRoughly 70% of the 4 million databases Supabase adds each month are now spun up by agents or AI tools. On October 2 the company raised $150 million led by Singapore's GIC, four months after a $500 million Series F, and agreed to buy Turso, which makes lightweight databases built on SQLite (the tiny single-file database inside most phones) for exactly that pattern. Terms were not disclosed. Turso founder Glauber Costa becomes head of agentic services, and Supabase also launched Supabase Compute, hosted sandboxes for long-running agents. The database customer it is chasing is a script that wants a fresh database for a ten-minute job and never logs in again.

AI Agents ·Earendil

Pi 1.0: the open-source terminal coding agent trims prompts by 40%

AnalysisAbout 2,000 prompt tokens disappear from a typical request in Pi 1.0, the first stable release of the open-source terminal coding agent that Earendil shipped October 1. Its default tools cut a GPT-5.6 request from roughly 5,300 prompt tokens to 3,300, partly by loading tools only when needed. The MIT-licensed harness (the loop that wraps a model with its tools and memory) now speaks MCP natively, can call non-language models such as Jev decision models, and pre-warms Anthropic's prompt cache. The GitHub repo shows 111,800 stars. A small agent that works with almost any provider makes it cheaper to walk away from any single lab's tool.

AI ModelsAI Agents ·Perplexity Docs

Perplexity Decisions API: an open 27B model that answers yes, no or a score

AnalysisOutput tokens are free on Perplexity's new Decisions API, because the model never writes any. Launched October 1, it runs pplx-decider-v1-27b, an open-weight model under the Apache license (free to use and modify, including commercially), fine-tuned from Alibaba's Qwen3.8-27B. Instead of prose it returns probabilities over fixed answers: yes or no, a pick from up to 255 options, or a rubric score. Input costs $0.04 per million tokens, and small requests return in under two seconds. Perplexity claims 85.71% on its own 11-benchmark panel against 84.51% for TypeSafe's Jev, the decision model that started this category. Self-graded tests deserve the usual salt.

StudentBench: AI tutors match human GRE coaching at 918 times lower cost

Analysis$4.81 for a human tutor against $0.0052 for an AI tutor, per percentage point of score gained, is the gap StudentBench measured. The study from Handshake's research team, revised on arXiv September 30, put 2,383 people through one-hour GRE sessions with AI tutors, expert human tutors or no tutor, and logged more than 175,000 student messages. AI tutoring came out statistically equivalent to human tutoring, the best AI tutor beat the human average in five of seven GRE areas, and the cheapest equivalent tutor was Google's open Gemma 4 31B. Humans kept the lead on verbal reasoning. Faster AI replies tracked with more practice and bigger gains.

LLM EvalsAI Industry ·arXiv Blog

arXiv submission cap: two papers a month after AI-made papers swamp moderators

Analysis40,363 papers hit arXiv in September 2026, nearly double the 20,569 of September 2024 and four times the 9,869 of 2016, along with almost 9,000 support tickets. Since October 1, the preprint server most AI research passes through caps every submitter at two papers per calendar month and three active submissions at once, and rejected papers still count. arXiv blames AI tools that make thin papers and "salami" papers (one study sliced into several) cheap to produce, and says AI submissions in computer science grew more than sixfold since 2024. Volunteer moderators turned out to be the cheapest part of science to overwhelm.

AI IndustryLLM Evals ·Microsoft Security Blog

Microsoft Digital Defense Report: attackers weaponise new flaws in under 24 hours

AnalysisLess than a day now separates a vulnerability turning up in the wild from attackers using it, according to Microsoft's 2026 Digital Defense Report, published October 1. Fixing a critical internet-facing flaw still takes a typical company 30 to 60 days. Phishing was the way in for 23% of the intrusions Microsoft's responders handled between July 2025 and June 2026, up from 7% a year earlier, and the report credits AI with making personalised lures cheap at scale. Spear-phishing used to need a skilled human per target. Microsoft's verdict is that, for now, AI has handed the advantage to attackers.

AI Industry ·StorageReview

DGX Spark 64GB: Nvidia's cheaper desktop AI box costs more than the original

AnalysisHalf the memory now costs more than the full machine did a year ago. Nvidia announced a 64GB DGX Spark on October 2 at $4,999, shipping October 23 through Acer, Asus, Dell, Gigabyte, HP and MSI with the same GB10 Grace Blackwell chip as the 128GB model. That bigger model launched at $3,999 and now sells for $6,950, a jump of nearly 75% that Nvidia's partners blame on soaring memory prices. The 64GB box runs open models up to about 100 billion parameters, and two can be clustered to reach 128GB. Local AI was supposed to be the cheap escape from cloud bills.

AI Models ·Tavus

Tavus Griffin: 48% of callers thought an AI video face was human

AnalysisNearly half the people on a one-minute video call with Griffin left believing they had talked to a person. Tavus, a startup that makes AI video avatars, unveiled the model October 1 and said 48% of 54 blind-test participants judged it human, against a 2% ceiling for its earlier systems. Griffin listens and watches while you talk, decides at sub-second intervals whether to nod, interrupt or stay quiet, and generates 720p video with about 0.43 seconds of audio-to-video lag on Nvidia H100 chips. It is a research preview for trusted testers only. Fifty-four people is a small sample, and a minute is a short call.

AI Industry ·Engadget

Google AI Overviews: judge tosses Penske and Chegg antitrust suits

Analysis"An expectation is not an agreement," Judge Amit Mehta wrote on October 1, dismissing antitrust suits from Penske Media, owner of Rolling Stone and Variety, and Chegg, the homework-help company, over Google's AI Overviews. The publishers argued Google forced them to choose between feeding its AI summaries and fading from search. Mehta, the same judge who ruled in 2024 that Google holds an illegal search monopoly, rejected all five claim types and said antitrust law is no substitute for lawmakers dealing with new technology. He called himself "not unsympathetic." The Penske dismissal came without prejudice, so a rewritten complaint remains possible.

AI-written web text: 31% of filtered tokens are synthetic, and they hurt training

AnalysisNearly a third of high-quality web text is now machine-written. Researchers from Pangram Labs and UMass Amherst ran August 2026 web data through FineWeb's quality filter and found 31.1% of tokens labelled AI-generated, up from 27.5% in June. Training on that material helps small, data-starved models at first, then flips: for well-fed models, AI tokens raise loss almost immediately while the same amount of fresh human text keeps lowering it. The paper, posted September 30, fits a scaling law (a formula predicting how models improve with more data) that forecasts the effect on models 3.6 times larger with 41% lower error. Human writing is turning into a scarce input.

AI Industry ·Air & Space Forces Magazine

Autonomous Warfare Command: Pentagon plans a four-star command for drones and AI

AnalysisA four-star command with its own budget, staff and career fields for drone operators is how the Pentagon plans to meet the battlefield it expects. Defense Secretary Pete Hegseth ordered the Autonomous Warfare Command on September 30 in a speech to 600 officers and senior enlisted at Quantico, with an October 1, 2027 target that still needs Congress. Until then Project Agincourt, led by Defense Innovation Unit director Owen West and Navy SEAL Senior Chief Max Strasiser, will draw on existing drone programmes and push buying decisions toward frontline units. Military AI is getting a permanent institution, and permanent institutions ask for a budget line every year.

AI ModelsAI Industry ·The Next Web

OpenAI says Moonshot-linked users sent 16,000 requests to copy its hidden reasoning

AnalysisSixteen thousand requests over July 24 and 25 mark the peak of what OpenAI calls an adversarial distillation campaign (using one model's outputs to train a rival) tied to people associated with Moonshot AI, the Chinese lab behind Kimi. Disclosed September 30, the method was crude: copy the encrypted reasoning a model produced in one chat, then ask another chat to decrypt and transcribe it. More than 4,000 users were active at the peak, and OpenAI says it shut the campaign down by July 28. No encryption was broken and no user data was touched. OpenAI banned accounts and tightened sign-up checks, which everyone else will now feel too.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack