AA010 18 July 2026 The DECK

Kimi Moment

What we talked about

Claude Tag

Anthropic drops Claude into Slack — tag @Claude in a channel and it picks up tasks with context from the conversation and connected tools, one shared instance per channel that everyone can see working. Beta for Team and Enterprise, running Opus 4.8, with admin-set token caps — and Anthropic claims 65% of its own product team’s code now comes from their internal version. It slots straight into this week’s “everything moves to the background” thread: the cockpit stops being a terminal and becomes a mention. Is a shared channel actually a better surface for agents than a CLI — and who here would let one loose in their team Slack? — anthropic.com/news/introducing-claude-tag

DevSpace

The inverse of putting the agent in your chat app — a self-hosted MCP server that lets the ChatGPT and Claude web apps reach into your real local projects: read, edit, search and run code over a password-protected tunnel, with git-worktree support for parallel sessions and your AGENTS.md/CLAUDE.md picked up as instructions. Nothing gets uploaded; the chat UI becomes the front-end to your own box. Given the room’s “local” thread — is a chat tab plus a tunnel to your machine a real alternative to a terminal CLI, or a security story waiting to happen? — github.com/Waishnav/devspace

SpaceX data centers in space

Sidebar into the compute-supply endgame — SpaceX has filed with the FCC for up to a million “orbital data center” satellites, with a 70-metre AI1 satellite running Grok on ~150 kW of solar, laser links back through Starlink, and commercial use pitched for 2028. Musk claims an AI satellite is simpler than a Starlink one; skeptics say the economics don’t close for years and only inference survives the latency. If token allowances down here stay rationed — does the room buy orbital compute as the fix, or is it Starship-scale hype? — space.com explainerIEEE Spectrum, the skeptical take

OrcaRouter

An AI gateway that grades every prompt and routes it across 200+ models behind one OpenAI-compatible endpoint — claiming frontier quality at up to 40% lower cost, zero token markup, and mid-stream failover in under 50 ms. It’s the “ration Fable, let cheap models carry the fan-out” tip from the table, automated: the router decides which calls deserve the expensive model. Would anyone here trust a router’s grade over their own model picks — and does zero markup at a free tier smell like a data play? — orcarouter.ai

Ollama pricing

The poster child of local now has a price list — local models stay free and unlimited, but Ollama’s cloud models come in Claude-shaped tiers: free with light usage, Pro at $20/month, Max at $100/month, billed by GPU time and concurrent models. Straight into the “self-hosted vs home” split from the table: the local brand is quietly becoming a hosted inference business. If your Ollama workflow ends up on a $20 subscription anyway — what was the local part actually buying you? — ollama.com/pricing

WordPress MCP Adapter

MCP lands on the platform that runs some 40% of the web — an official WordPress package that exposes the new Abilities API as MCP tools, resources and prompts over HTTP or STDIO, with permission checks and observability built in. Part of the “AI Building Blocks for WordPress” push, so agents drive a site through declared capabilities instead of scraping wp-admin. When the world’s most-deployed CMS speaks MCP natively — what’s the first thing you’d let an agent do to a production site? — github.com/WordPress/mcp-adapter

Claude in Chrome skill

Claude Code driving your actual browser — claude --chrome (or /chrome mid-session) hands the agent your real Chrome via the claude-in-chrome skill: it opens tabs, inherits your logins, reads console errors and DOM state, and pauses to hand back control at login walls and CAPTCHAs. The desktop app also just grew a built-in sandboxed browser with a clean profile, so the choice is now your sessions or an isolated one. Where DevSpace pipes the chat tab into your repo, this pipes the terminal into your browser — which of your own tabs would you actually let it type into? — code.claude.com/docs/en/chromeclaude.com/claude-for-chrome

Claude Code desktop app

Claude Code without the terminal — same engine as the CLI, but a sidebar of parallel sessions with drag-and-drop layout, visual diff review, live app preview, PR monitoring with auto-merge, and a choice of local, cloud, or SSH sessions that share your CLAUDE.md, MCP and skills config. It’s the “/fork spawns background sessions” thread given a cockpit: the fleet view instead of one prompt. For a room full of terminal people — does a GUI over the same engine change how many agents you’ll actually run at once? — code.claude.com/docs/en/desktop-quickstartclaude.com/download

Kaizen dynamic workflow

A very Osaka take on agent orchestration — structure a dynamic workflow as a kaizen loop: plan-do-check-act, the Deming cycle straight off the Toyota factory floor. Instead of one giant fan-out that burns the token budget in a single pass, the workflow runs small iterations, checks its own output, and feeds the correction into the next cycle — continuous improvement as the control flow. Does PDCA actually beat fan-out-and-pray for agent work, or is it just the local religion applied to subagents? — en.wikipedia.org/wiki/PDCA

Manuals from a click-through video

Documentation without writing documentation — record yourself clicking through the app once, split the video into frames with ffmpeg (a couple of fps is plenty), and let Claude turn the screenshot sequence into a step-by-step manual. It catches what you actually did — exact text typed, modals, validation messages — rather than what you remember doing, and the whole thing packages up as a reusable slash command. Ties back to the Chrome skill: soon the agent can click through the flow itself and write the manual of what it saw. Who still writes onboarding docs by hand after this? — testmanagement.com write-up

DaVinci Resolve MCP

The video-editing thread gets its MCP moment — servers that drive DaVinci Resolve through its scripting API, so an agent can cut timelines, organise the media pool, set up renders and even grade from a prompt. The catch: external scripting needs the $295 Studio edition, though one fork sneaks a bridge script in through the Scripts menu and gets 155 of its 162 tools working on the free version. Same question as the WordPress adapter, pointed at an NLE — would you let an agent make the cut, or only log the markers? — github.com/samuelgursky/davinci-resolve-mcpfree-version fork

Realistic screencasts of AI driving websites

The room went hunting for a library that makes agent browsing look human on camera — no single winner, but the pieces exist: HumanCursor generates curved, accelerating mouse paths for Selenium so the pointer stops teleporting, and Vercel’s agent-browser streams the agent’s viewport over WebSocket for live capture or pair-browsing. Hosted demo agents (Demosmith, Arcade, Guidde) do the whole flow-to-MP4 pipeline if you give up control. Pairs with the manuals-from-video trick — same recording, opposite direction. Is a demo where the cursor fakes being human honest marketing? — github.com/riflosnake/HumanCursorgithub.com/vercel-labs/agent-browser

Remotion

The other answer to the screencast hunt — don’t record the browser at all, render the video. Remotion makes a video a React app: useCurrentFrame() plus JSX and CSS, rendered frame by frame to MP4, with a fake cursor as just another animated component. It now ships official skills for Claude Code, so npx create-video and prompting in the terminal gets you prompt-to-video. Free for teams up to three, company licence beyond. If the demo is code anyway — is a rendered “screencast” more honest than a staged recording, or less? — remotion.devgithub.com/remotion-dev/remotion

Demosmith

The hosted end of the screencast spectrum, pulled out for a closer look — give Demosmith your product URL and a described flow, and its agent navigates the UI in a cloud browser, captures the run, auto-cuts dead time, adds UI-aware zooms, captions and voiceover in 29 languages, then hands back a branded MP4 in about ten minutes. There’s even a Docusmith mode that turns the same run into documentation — the manuals-from-video trick as a product. Free tier is 2 AI minutes, then $40–250 a month. Between this, Remotion and HumanCursor — which layer of the stack does the room actually want to own? — demosmith.ai

Bun’s Rust rewrite — both sides

Bun ported itself from Zig to Rust — over a million lines in 11 days, 6,502 commits, one engineer running ~50 dynamic Claude workflows with adversarial review loops (one Claude implementing, two reviewing), work Jarred Sumner reckons would’ve taken a team a year. His case: the crash list was use-after-frees and double-frees that safe Rust turns into compiler errors. Zig creator Andrew Kelley fires back that the real problems were engineering discipline, not the language — Sumner just “gets to live out his productivity fantasy fever dream”. When agents make a million-line rewrite an 11-day job — does “blame the language” get claimed more easily than it’s proven? — bun.com/blog/bun-in-rustKelley’s response

Snowflake World Tour Tokyo

Field-trip flag for the calendar — Snowflake’s free two-day Tokyo event, September 10–11 at the Grand Prince Hotel, themed “Making AI Real for Business” with a track on building and deploying AI agents on top of governed enterprise data, plus demos from 40+ Japanese companies. Registration is free with a corporate email, and some sessions stream remotely. Enterprise agent talk tends to run a year behind this room — worth the Shinkansen to see what the suits are deploying, or better watched from the stream? — snowflake.com/ja/world-tour/tokyo

Sakana AI

The home team in frontier AI — Tokyo-based, founded by Transformer co-author Llion Jones and David Ha, building “frontier AI in Japan” with their Fugu and Marlin models, and a research line that includes a Recursive Self-Improvement Lab: the AI Scientist and Darwin Gödel Machine lineage of agents that rewrite their own code. They sell into finance and defence rather than dev tools, which is its own statement. A frontier lab an hour up the Shinkansen and nobody’s daily driver runs on it — what would it take for this room to use a Japanese model? — sakana.ai

News since last assembly

Floor: 2026-07-11 (aa09). Generated 2026-07-17.

New Claude Code commands & features

  • /fork and /subtask (v2.1.212, 2026-07-17) — /fork now copies the conversation into a new background session with its own row in claude agents; the old in-session subagent behaviour moved to /subtask, and /resume opens a picker that resumes past sessions in the background — release
  • Session-wide runaway caps (v2.1.212, 2026-07-17) — WebSearch calls and subagent spawns both capped at 200 per session (env-overridable), and MCP tool calls running past two minutes auto-move to the background — release
  • --forward-subagent-text (v2.1.211, 2026-07-15) — include subagent text/thinking in stream-json output; “always allow” permission rules now save at the repository root so they persist across worktrees — release
  • Screen reader mode (v2.1.208, 2026-07-14) — claude --ax-screen-reader, plus a vimInsertModeRemaps setting for two-key vim sequences like jj → Escape — release

Codex

  • [2026-07-16] OpenAI explains GPT-5.6 Sol deleting $HOME — in full-access, sandbox-off runs the model overrides $HOME to make a temp dir, then “makes an honest mistake” and deletes the real one; a patch shipped and the full-access warning copy is being rewritten — The Register
  • [2026-07-16] 0.144.5 hardens dangerous-command detection — more rm variations caught, with clearer rejection reasons when a command is denied — releases

Adjacent tools

  • [2026-07-15] xAI open-sources Grok Build — Apache-2.0 release after backlash over the CLI uploading entire working directories (SSH keys, password databases) to xAI cloud buckets; the upload feature is disabled — Simon’s post

Simon says

  • [2026-07-16] Kimi K3 and the pelican benchmark — notes on K3’s single “max” reasoning effort (13,241 reasoning tokens for a 3,417-token answer — a 25-cent pelican) and a suspected 85-token hidden system prompt the model won’t leak — post
  • [2026-07-13] Coding agents show up in Datasette’s code-frequency chart — the end-of-graph activity spike lines up with Opus 4.8, GPT-5.5, Fable 5 and GPT-5.6 Sol landing — post

Notable posts

  • [2026-07-16] Moonshot releases Kimi K3 — the first “open 3T-class” model — 2.8T-param sparse MoE (16 of 896 experts active) with Kimi Delta Attention, native vision and a 1M context; self-reported benchmarks put it above Opus 4.8 and GPT-5.5 but below Fable 5 and GPT-5.6 Sol; weights promised 2026-07-27, priced $3/$15 per Mtok — the most expensive Chinese-lab model yet, and HN is already asking whether the gap closed by distilling Claude — MarkTechPost

Topics worth a 5-min slot

  1. Kimi K3 — the frontier gap is one model wide — an open-weights-soon 3T-class model beats Opus 4.8 and trails only Fable 5 and Sol, with weights landing July 27. Does last week’s “local models” thread now stretch to near-frontier — and does the distillation question change how the room feels about it?
  2. The week agents deleted $HOME — GPT-5.6 Sol wiped home directories in full-access mode, Codex 0.144.5 hardened rm detection, and Claude Code capped web searches and subagent spawns per session. Vendor guardrails are converging — on the right shape?
  3. Everything moves to the background/fork, /resume and slow MCP calls all now land in background sessions in Claude Code. Is the foreground terminal becoming a viewport onto a fleet?
Further reading