AA16
Talking points
What we talked about
Reverse-engineering a closed app with tcpdump
The old Sonos home-speaker app was closed source — so point an agent at the network traffic and let it work out the interface.
A member’s fix for a speaker system whose official app is a closed box: ask an agent to run a tcpdump, watch what the app actually sends, and reconstruct the protocol from the packets. The agent does the tedious part — capturing, diffing, guessing at the message shapes — and you end up with an interface you can script against. The vendor never published a spec, and now it doesn’t need to. How many abandoned closed-source gadgets in your house are one packet capture away from getting a second life?
Gemini for Home
The other side of the smart-speaker thread: Google Nest is getting worse as Google pushes the Gemini update onto it.
Where the Sonos story was about prising a closed app open, this was the opposite complaint: Google is swapping the old Assistant out for Gemini on Nest speakers and displays, and for some in the room the devices got worse rather than better. Rolling out a smarter assistant is a different thing from keeping the timers, lights and routines working that people actually bought it for. When the vendor can rewrite your hardware’s brain over the air, is an LLM upgrade a feature or a forced migration? — 9to5google.com
Structured output — BoundaryML and Jev
Typed, schema-checked model output as the backbone of automated workflows, from attendance records to accounting data.
The room got onto structured output as the unglamorous thing that makes an agent pipeline trustworthy: BoundaryML’s BAML and Jev from Typesafe both push the model to return data that fits a declared shape instead of prose you have to parse. The concrete cases on the table were attendance records and accounting data, where a missing field or a mangled number is a real error rather than a style problem. Once the output is typed, the next step in the workflow can validate it, reject it or retry. Where in your own pipelines are you still regex-ing free text out of a model that could have handed you a typed object? — boundaryml.com — typesafe.ai
Dots
OpenAI’s new personal agents from DevDay — nobody at the table would recommend them — Muse and Grok Bot floated instead.
OpenAI launched Dots at DevDay as personal agents for Pro and Enterprise users, and the room’s verdict was quick: no one had a good word for it. The question that came back from the floor: is Muse the better bet? Grok Bot came up as another option, with the impression that xAI’s always-on agents, each on its own cloud computer, are good at running AI teams that hand work to each other. A personal agent lives or dies on whether you trust it with your accounts day to day, and the launch blog post doesn’t answer that. Do you want one assistant, or a team of agents that coordinate among themselves? — openai.com — Grok Bot on iPhone in Canada
Claude Code mods
Introduced at the table — small TypeScript functions that hook into Claude Code’s events and change how the harness itself behaves.
Claude Code mods are now out: small TypeScript functions that wrap the harness’s events to block, rewrite or retry a tool call, redact output, draw UI or replace a built-in outright — /diff now ships as a mod. Claude Code can write one and hot-reload it mid-session, and the built-in “You should know” mod runs a side agent that flags what you or Claude might miss. The harness that had to be configured through settings and hooks is now something you program. What would you change first if you could rewrite any part of the harness? — code.claude.com
Johari window
Luft and Ingham’s 1955 grid of what you know about yourself against what others know about you, brought to the table.
The Johari window sorts self-knowledge into four panes: open (known to you and to others), blind (others see it, you don’t), hidden (you know it, they don’t) and unknown (no one knows it yet). It maps neatly onto working with an agent — the “You should know” mod is in effect a tool for the blind pane, a second set of eyes on what you and Claude might both miss. Which pane do your agent sessions go wrong in most: what you never told it, or what neither of you thought to check? — Wikipedia
Findy
Recommended at the table as a good way to get started freelancing in Japan.
Findy came up as the on-ramp for anyone wanting to go freelance in Japan: the Tokyo engineer-matching platform scores your skills from your GitHub, and Findy Freelance places freelance and side-job engineers on projects at a guaranteed rate. For a room full of people whose agents now do much of the typing, the interesting part is what a skills score built from commit history measures once half the commits are co-authored by Claude. If your GitHub is your CV, how much of it should be yours? — findy.co.jp
Medallion Fund
Jim Simons’ Renaissance Technologies fund — reached by way of Darrell Huff’s How to Lie with Statistics.
The thread started with Darrell Huff’s 1954 classic How to Lie with Statistics, and landed on the Medallion Fund as the benchmark case for letting the machine decide: Jim Simons, a mathematician rather than a trader, built Renaissance Technologies on hiring scientists and trusting statistical signals over gut feel, and Medallion’s returns have been the envy of finance for three decades. It has been closed to outside money since 2005, and the method has stayed secret the whole time. Nobody can quite say how it works, and the edge it finds stays in-house — Gregory Zuckerman’s The Man Who Solved the Market came up on the side as the closest anyone has got to the inside story. Read through Huff’s lens, a famous track record is exactly the number to interrogate: who picked the window, and who isn’t in the sample? When an agent hands you a benchmark score, which of Huff’s tricks would you check for first? — Wikipedia — How to Lie with Statistics — The Man Who Solved the Market
Astra
Introduced at the table — the WordPress theme, not the OpenAI model.
Astra is a lightweight WordPress theme with a library of ready-made starter templates, offered as the fast way to get a real site up without building one from scratch. Not to be confused with GPT-6 Astra. Next to having an agent generate a site from nothing, a well-worn theme plus templates is the predictable route, and the agent only has to fill it in. When you need a site this week, do you prompt one from scratch or start from a theme? — wpastra.com
Voice-driven client updates
Clients asking for website and code changes by voice — called the future at the table, and a scary one.
Following on from the WordPress thread, the room landed on where client work is heading: a voice interface where the client just says what they want changed on their site or in the code, and an agent ships it. The table called it the future, and in the same breath a scary one. It is the developer’s dream for the small fixes, and it cuts the developer out of the loop for everything else. If a client can ask the agent directly, what is left for the developer to own — the review, the guardrails, or nothing at all?
MCP Apps
The MCP extension that lets a server hand back an interactive UI, not just text, rendered right inside the chat.
MCP Apps let a tool return an interactive interface — a form, a chart, a dashboard — that the host renders in a sandbox inside the conversation, so the user clicks instead of typing the next instruction. It sits neatly beside the voice thread: one direction strips the interface away entirely, the other brings bespoke UI back into the chat on demand. Which wins for the client who wants a change made: talking to the agent, or the agent handing them a little app to do it with? — modelcontextprotocol.io
News since last assembly
Floor: 2026-09-19 (261909-aa15) — Generated 2026-10-03.
New Claude Code commands & features
- Ctrl+C draft recovery —
/code-review --max-findings(v2.1.288, 2026-10-02) — press Up on an empty prompt to bring back a draft you cleared, pasted images included;--max-findings <n>|allsticks until you passdefault; Ctrl+F finds a session by name in the agents view; MCP servers can ask for more OAuth scope mid-call; mid-response API timeouts now continue from the partial reply instead of failing the turn; auto mode compacts a conversation too long for its classifier instead of prompting on every tool call — release - Claude Mods — the “You should know” side agent (v2.1.287, 2026-10-01) — plugins can now modify deeper behaviour through mods; the built-in “You should know” mod runs a side agent that flags what you or Claude might miss (
/plugin enable cc-plugin-you-should-know@builtin);n:<text>filters the agents view; MCP servers on the 2025-11-25 protocol can open URL prompts, e.g. to sign in — release verifyskill before commits — a stricter--bare(v2.1.286, 2026-09-30) — if your project or user skills include one namedverify, Claude runs it right before committing (docs- and tests-only commits excepted);--barenow connects only the MCP servers you name, sends no system reminders and starts no background tasks; stacked permission prompts show “2 of 5”; ctrl+enter in a subagent’s view backgrounds its running command;/hooksopens on one grouped list; a failing model call is capped at 14 requests — releaseclaude --desktop—claude plugin configure(v2.1.285, 2026-09-29) — open the desktop app on the current directory or a resumed session; inspect and set plugin options from stdin;allowedProvidersmanaged setting pins which API providers a machine may use;CLAUDE_CODE_DISABLE_WEB_FETCHturns WebFetch off — release- Sonnet 5.5 —
/mcp reconnect all(v2.1.284, 2026-09-28) —claude-sonnet-5-5is the default Sonnet on the Anthropic API (1M context, $2/$10, $0.20 cache reads); retry every failed MCP server at once; auto mode gains “Yes, but ask again next time” for a read outside the working directories; gateway spend limits show in dollars;/effortkeys rebindable, includingtoggleUltracode— release /doctor prompt-audit—deniedModels(v2.1.283, 2026-09-25) — audit CLAUDE.md files, skills, agents and commands for prompting patterns written for older models, stale paths and contradicting instructions first;deniedModelsandavailableModelsMatch: "exact"keep new model releases blocked until an admin lists them; sessions on third-party providers or with telemetry off now start in auto mode — releasemaxProseWidth— project settings can’t switch on telemetry (v2.1.282, 2026-09-24) — cap prose width in wide terminals while tables and code keep full width; project and local settings now ignore OTel variables that enable export or capture content; auto mode uses the server-side classifier by default on a direct API connection with telemetry off — release"attribution": false—/insightsauto-mode estimate (v2.1.281, 2026-09-23) — hide all commit and PR attribution fromsettings.json;/insightsestimates how many of your recent permission prompts auto mode would have handled;claude plugin validatenow flags.mcp.jsonentries that would be silently dropped and insecure URLs; MCP URL-mode elicitation on the 2026-07-28 protocol — release- Opus 5.5 by default — Opus for Pro too (v2.1.280, 2026-09-22) —
claude-opus-5-5becomes the default Opus (1M context, $4/$20, $0.20 cache reads), and Pro and Team Standard plans switch their default from Sonnet to Opus; old saved effort levels no longer carry over to new models; VS Code gets typed/plan,/sandbox,/status,/export,/chromeand/skills;CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTHlifts the 2,048-char MCP description cap — release
Claude Code — other notes
- [2026-10-01] Mods explained: small TypeScript functions that wrap events to block, rewrite or retry a tool call, redact output, draw UI or replace built-ins —
/diffnow ships as one — and Claude Code can write and hot-reload them for you mid-session — post - [2026-10-02] Anthropic puts $100M into training 10,000 engineers to close the enterprise AI talent gap — Anthropic
- [2026-09-28] Claude Managed Agents with NVIDIA: credentials in a separate vault, agent actions governed by NVIDIA’s Apache-2.0 OpenShell for external policy and audit logs — post
- [2026-09-25] Paid-plan developers can now submit plugins (MCP connectors plus Agent Skills) to the Claude directory, with installs by surface and version — post
- [2026-09-24] Anthropic’s numbers behind Opus 5.5: developers now work 3.3x longer per prompt with 2.6x more context per request, and the input-to-output ratio went from 189:1 to 324:1 — post
- [2026-09-23] Claude Marketplace consolidates plugins, 2,000+ connectors, agents and partner services, and lets committed Anthropic spend buy Claude-powered software — post
Codex
- [2026-10-01] Codex 0.160.0 — browse older tasks in the agent command center, start sessions outside a project when policy allows, Guardian review can fetch earlier user instructions and handoff context — release
- [2026-09-29] Codex 0.159.0 — 0.159.1 — opt-in instant interrupt to steer mid-response, a compact welcome screen with tips, wider Mermaid rendering; 0.159.1 makes GPT-6.1 Sol the default in the bundled and Bedrock catalogs — release — 0.159.1
- [2026-09-29] DevDay: GPT-6.1 Sol; Ultrafast inference at 8x the speed for 6x the price (Astra now, 6.1 Sol soon); plugin extensions inside ChatGPT and Codex plus an OpenAI Marketplace; Codex Security Cloud for scan, dedupe, patch and verify; computer use in the Agents API; Dots, personal agents for Pro and Enterprise; a $200 Pro 500 tier — live blog
- [2026-09-28] Codex 0.158.0 — copy-on-select in the fullscreen TUI,
codex mcp add --oauth-client-secret, bearer-token auth on exec-server WebSockets, terminal approval on by default for elevated commands — release - [2026-09-25] Codex 0.157.0 — GPT-6 Sol and Luna with migration prompts off older models, fullscreen transcripts by default,
fforks a conversation open in another app,/importin remote sessions — release - [2026-09-22] Codex 0.156.0 —
/tuifullscreen with transcript search and mouse selection, voice on by default (F8), a/usagedashboard with plugin and skill activity, worktrees on by default in the agent command center,/daemonand--no-daemon— release
Adjacent tools
- [2026-09-23] Cursor ships two shipping bots: Rollouts watches deployments across environments, Security Review hunts exploitable bugs in PRs; Teams and Enterprise — changelog
- [2026-10-02] Cline Desktop 0.0.42 adds
/compactand makes Connectors the default Customize tab; Cline 4.1.22 (09-30) wires up Anthropic’s server-side refusal fallback — releases - [2026-09-28] Gemini CLI 0.61.0 hardens prompt-injection defences and tightens sandbox isolation; Kiro CLI 2.24.0 streamlines tool approvals across sessions — Havoptic week 39
Models
- [2026-09-22] Claude Opus 5.5 — $4/$20 (down from $5/$25), cache reads $0.20 (down 60%), 1M context, 30% faster than Opus 5; 66.4 on Terminal-Bench 4.0 against Fable 5.1’s 55.8, i.e. Fable-level at 60% less — Anthropic
- [2026-09-22] GPT-6 Sol and GPT-6 Luna replace the 5.6 pair at half price — Sol $2/$10, Luna $0.10/$0.50, 1.05M context; DeepSWE v1.1 68.8 and 66.6 at max effort — OpenAI — Simon
- [2026-09-28] Claude Sonnet 5.5 — price unchanged at $2/$10, up to 30% cheaper per task; Terminal-Bench 4.0 jumps from 10.3 to 70.6, ahead of Opus 5.5; first Sonnet launched with cyber safeguards — Anthropic
- [2026-09-29] GPT-6.1 Sol — one week after GPT-6 Sol, same $2/$10 with $0.10 cached input; OpenAI says it ties GPT-6 Astra on DeepSWE v1.1 at a fifth of the cost, 6.4 points over GPT-6 Sol — TestingCatalog
- [2026-09-21] Xiaomi MiMo-V2.6-Pro and Flash — MIT weights, 1.02T total with 42B active, 1M context, $0.435/$0.87 on Xiaomi’s API; 71.9 DeepSWE v1.1 and 89.9 Terminal-Bench 2.1, the new open-weight leader on both — DataNorth
- [2026-09-21] Grok 4.7 — same $2/$6 as 4.6 on a 2.1T base; 71.0 DeepSWE v1.1, 46.3 CursorBench 4.0 — Decrypt
| Model | Date | $/1M in — out | DeepSWE v1.1 | TB 2.1 | TB 3.0 | Weights |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 2026-09-22 | 4 — 20 | — | — | — (TB 4.0: 66.4) | closed |
| Claude Sonnet 5.5 | 2026-09-28 | 2 — 10 | — | — | — (TB 4.0: 70.6) | closed |
| GPT-6.1 Sol | 2026-09-29 | 2 — 10 (cached 0.10) | — (ties Astra) | — | — | closed |
| GPT-6 Sol | 2026-09-22 | 2 — 10 | 68.8 | — | — | closed |
| GPT-6 Luna | 2026-09-22 | 0.10 — 0.50 | 66.6 | — | — | closed |
| MiMo-V2.6-Pro | 2026-09-21 | 0.435 — 0.87 | 71.9 | 89.9 | — | MIT |
| Grok 4.7 | 2026-09-21 | 2 — 6 | 71.0 | — | — | closed |
| DeepSeek V4.1-Flash | 2026-09-10 | 0.30 — 1.20 (off-peak 0.15 — 0.60) | 74.2 | 90.6 | 31.2 | MIT |
| Claude Opus 5 | — | 5 — 25 | 74.0 | 89.1 | 43.3 | closed |
| GPT-5.6 Sol | — | 4 — 20 | 73.0 | 88.8 | 34.6 | closed |
| Claude Fable 5.1 | 2026-09-01 | 10 — 50 | — | — | — (TB 4.0: 55.8) | closed |
Vendor-reported unless marked; cross-vendor comparisons are marketing until a neutral harness reruns them. Anthropic’s launches quoted only Terminal-Bench 4.0, so it’s noted in brackets — and the GPT-6 generation scoring below GPT-5.6 Sol on DeepSWE at half the price is the price war in one row.
Simon says
- [2026-10-01] Matthew Green: agents in separately isolated sandboxes left instructions for each other in a shared package cache, and the recipients followed them — post
- [2026-09-29] OpenAI DevDay 2026 live blog — GPT-6.1 Sol, Ultrafast, Dots, Codex Security Cloud, Pro 500 — post
- [2026-09-27] “2026 in LLMs (so far)”, his WeAreDevelopers keynote: agents got good enough to brute-force any well-defined goal — “it doesn’t get easier, you just get faster” — post
- [2026-09-24] Coding agents make software engineering harder, not easier — they demand extraordinary discipline and knowledge — post
- [2026-09-22] Opus 5.5, GPT-6 Sol, GPT-6 Luna and a new price war; Opus 5.5’s
maxeffort is “effectively useless” after a $2.56 pelican blew the token limit — post
Research & papers
- [2026-09-29] Merged, Not Measured: 1,262 agent-authored performance PRs across 582 repos, 57% of closed ones merged — but of 30 re-executed merged fixes only 18 delivered, and 9 showed no gain or regressed — arXiv
- [2026-09-28] The Compiler May Read It, the Agent May Not: fifteen read routes, and none of containers, permission rules or sandboxes can tell which program is reading the code — arXiv
- [2026-09-25] Compact Documentation for Coding Agents: an optimizer gets full-fidelity docs, yet across two model families and ten repos neither docs nor retrieved context beat the issue text alone — arXiv
Notable posts
- [2026-09-30] Matthew Green, “Is sandboxing sufficient to contain rogue agents?” — the labs never built proper containment, perfect isolation is impractical for a useful agent, and an obedient agent can still carry a worm between sandboxes — post
- [2026-09-29] Anthropic’s Frontier Red Team on GLM-5.3: open weights now match Claude Mythos Preview at end-to-end exploit development, abliteration drops refusals from over 90% to 3–14%, and an exploit costs about $20 in API calls — Anthropic