backend grok-cli (SuperGrok Heavy sub, $0 marginal) · window last 24h
Opus 5 already does 2h / 1M-token “stamina jobs” ($10) — the failure mode is audit, not generation — Unattended loops need stronger non-text verifiers (vision, playable runs, deterministic gates); screenshot-and-pray is how Karpathy’s 5.5k-LOC Three.js LoTR render went janky. — x.com/karpathy/status/2083749667410727319
Eval model tiers per workflow: Sol→Luna cut a daily structured-output job 96% ($5/$30 → $0.20/$1.20 per 1M in/out) with no quality drop — Route CX structured/batch seats off Sol/Opus to Luna-class after a quick A/B; keep expensive seats for hard coding only. — x.com/reach_vb/status/2083916246710710429
codex execis the practical rail for unsupervised cluster/triage over messy data — Fits peer-seat batch work (issue clustering, severity ranking) without burning interactive Sol sessions. — x.com/reach_vb/status/2083921970425741715EMERGING:
metaharness(and “every repo gets its own agent”) is the category people are baking off — Swyx ranked Omnigent over Flue/Eve as the meta-layer above coding agents; watch this name over “skills stack” / “orchestration doctrine.” — x.com/swyx/status/2083664187969196152 · x.com/swyx/status/2083695562004771063EMERGING:
slop-tolerantsystems > anti-slop purity — Design harnesses/tools to absorb noisy agent output (validate, recover, self-describe) instead of banning “slop” — higher leverage than more prompt hygiene. — x.com/swyx/status/2083753582160191988Don’t gut skills/AGENTS.md: steipete still calls skills “incredibly useful” against the vanilla-config wave — Keep load-bearing skills; only prune noise that isn’t fixing a repeated wrong/inefficient behavior. — x.com/steipete/status/2083721500490858974
Ignore viral “graph engineering” as Cherny doctrine — he says he never said “graph” and isn’t talking graphs — Don’t rename your conductor graph or adopt fake vocabulary from AI-slop recaps; watch primary sources. — x.com/bcherny/status/2083782540570279992
Practitioner signal: kunchenguid rates Opus 5 worst in the 4.6→5 line (personality/judgment), while long-running capability is up — Matches tight-leash conductor use: great for long agentic work, worse as a chatty trusted peer — don’t default-Opus conversational seats. — x.com/kunchenguid/status/2083788378664009738
Highest-leverage action today: A/B every recurring structured/batch job (and any Sol parent that only needs cheap subagents) onto Luna-class; keep Sol/Opus for hard coding — and add a real visual/playable gate to any multi-hour unattended loop.