voices

Voices digest

Jul 30, 2026

backend grok-cli (SuperGrok Heavy sub, $0 marginal) · window last 24h


GPT-5.6 Luna −80% ($0.20/$1.20), Terra −20%, Sol Fast ≤2.5×; Codex/ChatGPT Work usage counts now reflect the cut; auto-review moved to Luna (~10× cheaper) — reprice your CX fleet: bulk exec/eval/auto-review → Luna, keep Sol/Terra only where ceiling matters, overnight capacity just expanded a lot. — x.com/reach_vb/status/2082882388619641008 · x.com/gdb/status/2082885748337115632

  1. Software quality is migrating into the harness OS: tests/property/mutation gates as agent back-pressure when code outpaces human review; skills carry the nondeterministic half + judgment, only some gates earn hard enforcement — matches your automated-gate + hooks doctrine; audit which checks are hard-deny vs model judgment, not more CLAUDE.md prose. — x.com/addyosmani/status/2082723002836545641 · x.com/addyosmani/status/2082883008995168684

  2. codex exec for repeatable, isolated, scriptable runs (ephemeral session, pinned model/effort, no network/MCP, separate dir per run) — drop this into unattended loops/CI/eval harnesses as the CX equivalent of headless work orders, not interactive chat. — x.com/reach_vb/status/2082609245011198432

  3. EMERGING: “distil harnesses” (same move as model distillation—compress scaffolding, keep the load-bearing bits) — treat bloated skills/hooks/prompts as distillation targets; a thinner harness often beats more steering. — x.com/swyx/status/2082869478573183226

  4. EMERGING: “slop cannon” vs “slop mop”—and cannons are good at mopping (dead code, faster/stricter tests, flaky fixes) — schedule deliberate cleanup agents after generative bursts instead of only shipping new surface area. — x.com/jarredsumner/status/2082836181688221802

  5. Session portability push: labs quietly hiding session/data while stripping user control is bad for users and the multi-tool ecosystem — relevant under Grok trust probation + multi-seat (Claude/Codex/Grok): prefer exportable session/state, distrust opaque agent-to-agent payloads. — x.com/mitsuhiko/status/2082838283520786748 · earendil.com/posts/session-portability/

  6. Design CLI --help so it doubles as an agent skill (run tool --help is the onboarding doc) — any tool your fleet shells into should be agent-discoverable without a separate SKILL.md if possible. — x.com/simonw/status/2082828424243212708

  7. ARC-AGI-3 “victory” scores can be settings-tripped (OpenAI writeup: two settings tripled scores); don’t re-route models off lab leaderboard posts alone — keep routing on your own harness evals + day-to-day taste, not benchmark dunks. — x.com/steipete/status/2082617409408762124 · openai.com/index/how-two-settings-tripled-our-arc-agi-3-scor

Honorable (leverage but not top-8): Vercel E2E deploy −~7s + “custom software factory” via CLI/MCP/API for autonomous ship loops (https://x.com/rauchg/status/2082876367629381719); Huntley: VMs-per-agent still best practice, not just cgroups (https://x.com/GeoffreyHuntley/status/2082621236031750311); Theo: one full day can burn 100% of Fable weekly on $200 Max—keeps Fable as reserve, not daily driver (https://x.com/theo/status/2082692150207435154); T3 Code mobile remote for Claude+Codex (https://x.com/theo/status/2082613200441524514). Silent/no signal in-window: bcherny, indydevdan, karpathy, AndrewYNg, _catwu, delba_oliveira.

Action today: Point Codex bulk/CI/codex exec/auto-review at Luna (and treat the extra headroom as more overnight CX work), not more Claude Fable burn.