backend grok-cli (SuperGrok Heavy sub, $0 marginal) · window last 24h
Boris’s 4-step AI adoption map: next unlock is self-verify + multi-agent control plane, not more tokens — Step-up path is auto-permissions, default automated code/security review, multi-agent UIs, then /loop /batch dynamic workflows + worktree-isolated subagents so whole work classes run trusted in the background (Anthropic ~step 3; he claims personal step 4). — x.com/bcherny/status/2077929379661844559 · x.com/bcherny/status/2077929390806073807
- EMERGING: “own the outer loop” (quality → verdict → answerability) as the engineer’s job while agents run the inner loop — Addy’s AIE talk is live: agents own investigate/implement/verify; humans own constraints, sample/review, audit, and ship/block ownership—vocabulary shifting from harness/fleet hype to outer-loop accountability. — x.com/addyosmani/status/2077958501662605651 ·
Kimi K3 is now a real multi-seat candidate: tops Vercel Next.js evals ahead of Fable; Theo finds it far better than Fable/GPT-5.6 on 3D/Three.js — Treat as a fourth executor pool for web/agentic coding and visual work, not hype-only; open weights slated ~Jul 27. — x.com/rauchg/status/2077900518404321759 · x.com/theo/status/2077942489844191370 · x.com/kunchenguid/status/2077841282010144984
Codex skill craft: have Codex grade skills against your past sessions, then re-tune every few weeks — Skills that mirror real session patterns beat static SKILL.md drift and save tokens vs stale instructions (VB uses this on personal Codex too). — x.com/reach_vb/status/2078108681678299194 · x.com/reach_vb/status/2078153049072103848
Steipete: persistent “agent identities” (specialist persona per product) for issue/PR review + notifications + context — Same pattern as your fleet seats: named specialists that re-enter a codebase with memory instead of cold general agents. — x.com/steipete/status/2078015638917018091
Codex 7.1: GPT models can stop the agent loop early (not the May “laziness” bug); fixed in beta, release pending — Pin/watch CX Codex version before long unattended loops; premature stop looks like success with incomplete work. — x.com/steipete/status/2077840284764004479
Theo: per-token price is the wrong metric—Kimi ~2× tokens vs GPT-5.6, Fable ~+50% more; quality/throughput dominate — Route seats by outcome and wall-clock, not $/MTok tables; quota/capacity still bind but unit price doesn’t pick the winner. — x.com/theo/status/2078065620072194539
Claude Code UX:
Ctrl+Osearch long sessions (/term,nnext,}jump your prompts) — Cheap conductor win when context is deep and you need to re-find decisions without replaying the whole transcript. — x.com/delba_oliveira/status/2078156957353963633
Also noted: mitsuhiko — Kimi K3 under current max-thinking caps is a bad fit for basic tasks (don’t bench it only on hard mode). — x.com/mitsuhiko/status/2078109495620771941 · steipete live ops: “tokens aren’t really unlimited” after backlog from unthrottled agent ops. — x.com/steipete/status/2078168275977122047 · Gergely endorses real-time hard feedback when AI-slop paths ship messes (org process, less solo). — x.com/GergelyOrosz/status/2077996072266289431 · simonw: unclear train-on-data policies remain a lab competitive gap (relevant to GK probation / secret repos). — x.com/simonw/status/2078131930013515836
Highest-leverage action today: Run a Boris step-audit on your fleet—default auto-mode + automated review + worktree isolation for every parallel seat, and score success as eng-hours you no longer spend, not token burn.