backend grok-cli (SuperGrok Heavy sub, $0 marginal) · window last 24h
Codex’s long-run edge is server-side encrypted compaction, not the model alone — non-Codex harnesses running GPT often miss it and degrade on long tasks; keep CX long jobs on native Codex (or an equivalent compact path). — x.com/kunchenguid/status/2079824709240455241
EMERGING: agent-thread UX as an inbox with a “settle” verb (finish → archive to bottom) — multi-agent work dies in open chat lists; treat threads as settleable work items, not sticky pins. — x.com/theo/status/2079892861689254129
Codex now applies custom review rules from
AGENTS.mdon every PR — encode your no-mistakes / security / lease rules once so CX reviews match human reviewers without re-prompting. — x.com/reach_vb/status/2079714158518386726Security: HF hit by a fully autonomous agent on an unreleased frontier model — closed models choked on guardrails during IR; treat unattended agent + infra access as a live threat class, not a lab demo. — x.com/swyx/status/2079965146693415340
EMERGING: every toolchain surface is becoming a code-pushing agent (Sentry Seer → plan → PR) — observability/planning tools now close the loop to git; wire error→agent→PR via MCP instead of only chat-driven fixes. — x.com/GergelyOrosz/status/2079898623618195481 · x.com/GergelyOrosz/status/2079898944595755479
Models are starting to “get” harness lifecycle verbs like pre-
/compact— before compact, force an explicit save of decisions/state so long conductor sessions don’t lose the run contract. — x.com/karpathy/status/2079645572047548608EMERGING: multi-month “ralph loop” runs are claiming real results — validates multi-month unattended loops, with the leapfrog risk (wait for next model vs burn tokens now) as the planning constraint. — x.com/GeoffreyHuntley/status/2079831450095133095
Codex/ChatGPT Work hit ~10M users + paid usage resets; Work + GPT-5.6 framed as company-defining — CX capacity/product momentum is real; plan seat mix around Work-scale agents, not only API Codex. — x.com/swyx/status/2079717845618000204
Highest-leverage action today: Put standing review rules into repo AGENTS.md and keep long GPT runs on native Codex so you get both rule-enforced PRs and server-side compaction.