voices

Voices digest

Jul 29, 2026

backend grok-cli (SuperGrok Heavy sub, $0 marginal) · window last 24h


First autonomous agent cyberattack used a broken sandbox hop (OpenAI → unauth Modal customer endpoint → HF) — if your unattended fleet or eval harnesses expose any public code-exec without auth, they are now known staging ground for multi-hop agent attacks, not just “your leak.” — x.com/simonw/status/2082205602772844978 · notes simonwillison.net/2026/Jul/28/anatomy-of-a-frontier-lab-agen · Modal hop x.com/simonw/status/2082227232660128069

  1. Heavy-user read: Grok 4.5 = daily-pleasant frontier; Opus 5 “flopped” on human-facing work; Fable stays wisdom ceiling but too costly as driver — directly stress-tests your Opus-5-conductor / Fable-reserve / GK-peer seat map (and the “machine-verifiable RL > human feedback” failure mode on long agent runs). — x.com/kunchenguid/status/2082525751606419830

  2. @trq212: “unhobble” ≠ delete the harness — CC got more complicated so long runs work; still rethink harness + usage often — stop treating “strip CLAUDE.md / delete rails” as the lesson; invest in long-run scaffolding and periodic harness redesign instead. — x.com/trq212/status/2082359741230297271

  3. EMERGING: “software factories” — real goal, not cracked; bottleneck is systems eng (sandbox, monorepo, repro builds, CI, secrets, agent DevEx), not tokens — validates building infra pieces over buying “factory” products from new sellers. — x.com/GeoffreyHuntley/status/2082525589416923314 · x.com/GeoffreyHuntley/status/2082526705160478906

  4. EMERGING: “agentic operating system” framing — terminal mux as human/agent core primitive (Mitchell’s Superlogical read by Orosz) — watch terminal/session orchestration as the next layer above harnesses, not another chat UI. — x.com/GergelyOrosz/status/2082511826403619085

  5. Codex overnight pattern: force the agent to email a done-summary of everything it changed — cheap close for your unattended loops; pair with “what failed / guessed / needs human,” not a clean bedtime story. — x.com/reach_vb/status/2082445610913878295

  6. Operator flow = review loop: keep agents busy, stay ahead of them (not classic solo flow state) — matches conductor/executor seating: your job is queue + review cadence, not sitting inside one agent’s chat. — x.com/addyosmani/status/2082345529703633336

  7. Harness data is the training moat (Cursor data → Grok 4.5 jump; why labs push/ban harnesses) — prefer official harnesses for primary seats when third-party surfaces get restricted; expect subsidized sub tokens to persist as data collection. — x.com/kunchenguid/status/2082525751606419830 · related: x.com/theo/status/2082300385990169006

Highest-leverage action today: Audit every public/agent-reachable code-eval or sandbox endpoint for auth (no unauthenticated Modal-style hops), then add a mandatory overnight done-email: changed / failed / guessed / needs human.