Context rot: which mechanism actually dominates, and whether hard rule #4 is calibrated to it
The question
Verbatim: "What is 'context rot' mechanistically (attention dilution vs positional bias vs retrieval degradation — which dominates), and how does the evidence map onto CLAUDE.md hard rule #4?"
Context: "context rot" appears 10+ times across the vault's harness-engineering cluster and is the stated justification for an AGENTS.md hard rule, but the vault's concept page ([[context-rot]]) describes the phenomenon without adjudicating the mechanism, and the rule's 5KB trigger has never been traced back to a measurement.
What we already know (from the vault)
- The rule's cited provenance is [[2026-04-15-thariq-claude-code-session-management-1m-context]]. Thariq's actual test is semantic, not size-based: "will I need the tool output again, or just the conclusion?" If just the conclusion, delegate. His post contains no byte threshold. The 5KB number in AGENTS.md is RDCO-invented, added for enforceability.
- [[context-rot]] states the mechanism as "attention spreads across more tokens and older, less-relevant content increasingly distracts." That is two mechanisms stated as one, and it is a vendor-and-blogger explanation, not a measured result.
- [[2026-07-07-claude-skill-count-degradation-skill-packs]] found the only hard numbers in the vault on this family of effects, and they are about tool-schema count (accuracy degrading past ~10 tools, hard failure past ~100), not conversational token count. The vault has been treating a selection-surface measurement as corroboration for a context-length claim.
- [[2026-07-24-thariq-context-engineering-claude-5-rules]] cuts the other way and is not reconciled anywhere: Anthropic deleted roughly 80% of Claude Code's system prompt for the Claude 5 generation with no measurable eval loss, and Thariq's advice became "spend scarce constraints only where they're truly yours." A mechanical byte rule is exactly the kind of over-constraint that pass was deleting.
- [[2026-09-21-indy-dev-dan-self-compact-pi-agent-zero-hype-devlog]] gives a field data point on where practitioners set thresholds: soft warning at 225k tokens, forced compaction at 270k. Nobody operating an agent in production sets a context trigger anywhere near 1,300 tokens.
What the web says
- Chroma's technical report (July 2025, 18 models) is the source of the term's popularity and it isolates length as real. Holding the needle-question pair fixed and varying only the volume of irrelevant content: "we isolate input size as the primary factor in performance decline." Degradation was consistent across all 18 models. (research.trychroma.com/context-rot)
- Chroma also found positional effects to be the weakest of the three, and task-dependent. Across 11 needle positions in needle-in-a-haystack (NIAH) there was "no notable variation in performance for this specific NIAH task." A primacy advantage appeared only in the repeated-words replication task. Same report: "we do not explain the mechanisms behind this performance degradation."
- Chroma's most awkward result is anti-intuitive and rules out a naive distractor story. "Models perform worse when the haystack preserves a logical flow of ideas. Shuffling the haystack and removing local coherence consistently improves performance" — across all 18 models and all needle-haystack configurations. Semantic coherence in the surrounding text is itself harmful, which no simple "irrelevant tokens dilute attention" account predicts.
- NoLiMa (Modarressi et al., ICML 2025, Adobe Research) gives the biggest effect sizes, and they are conditional on retrieval difficulty. Strip lexical overlap between question and needle so the model must infer a latent association: at 32K tokens, 11 of 13 models fall below 50% of their short-context baseline; GPT-4o drops 99.3% to 69.7%. GPT-4.1's effective context measures around 16K against a claimed 1M. Their attribution: "the increased difficulty the attention mechanism faces in longer contexts when literal matches are absent." (arxiv.org/abs/2502.05167)
- The decisive adjudication paper is "Context Length Alone Hurts LLM Performance Despite Perfect Retrieval" (arXiv 2510.05381). They remove the confounds one at a time: replace irrelevant tokens with whitespace, mask irrelevant tokens entirely so attention cannot reach them, and place all relevant evidence immediately before the question. Degradation persists at 13.9% to 85% across five models on math, question-answering and coding, "but remains well within the models' claimed lengths." Their working mitigation is to have the model recite the retrieved evidence before solving, converting a long-context task into a short one. (arxiv.org/abs/2510.05381)
- Anthropic's own first-party mechanism claim is an attention-budget story with a training-distribution rider. Transformers give "n² pairwise relationships for n tokens"; as length grows "a model's ability to capture these pairwise relationships gets stretched thin." Plus: "models develop their attention patterns from training data distributions where shorter sequences are typically more common than longer ones." This is an explanation offered, not an interpretability result demonstrated. (anthropic.com/engineering/effective-context-engineering-for-ai-agents)
- Positional bias is real, localized, and a modulator rather than a driver. The lost-in-the-middle U-curve persists into 128K+ windows and traces to an intrinsic position-dependent attention bias that calibration methods can partially correct. It tells you how to order a prompt; it does not explain why a long prompt is worse than a short one carrying the same information.
Convergences and contradictions
- Convergence: every primary source agrees the length effect is real and begins far below advertised context limits. NoLiMa puts the practical cliff at 32K with GPT-4.1's effective length near 16K; Chroma's stark contrast is 300 tokens versus 113K. The vault's directional belief is correct.
- Contradiction the vault has not absorbed: the vault's stated mechanism ("older irrelevant content distracts") is the one mechanism arXiv 2510.05381 explicitly falsifies as sufficient. Whitespace substitution and hard attention masking both remove the distraction and the degradation survives. Distraction amplifies; it is not the floor.
- Contradiction inside the literature: Chroma's coherent-haystack-is-worse result and the standard distractor-density account point in opposite directions. Neither has a mechanistic explanation. Chroma flags it as an open interpretability question and so should we.
- Contradiction between the rule and its own source: Thariq's test is "conclusion versus output." AGENTS.md substituted "bytes versus a constant." The vault records the substitution nowhere, and no measurement supports 5KB.
Synthesis for RDCO
Mechanistic verdict, stated as a ranking rather than a winner. There is no single dominant mechanism, and the entanglement is structured rather than mushy. Attention dilution in the strict, content-independent sense is the floor: an irreducible penalty for processing longer sequences that survives whitespace substitution, attention masking, and optimal positioning, and is best explained today by training-distribution mismatch plus representational crowding rather than by anything competing for attention. Retrieval degradation is the amplifier, and it owns most of the variance in realistic work: when the needed fact must be found by semantic rather than literal match, or competes with topically-adjacent distractors, effect sizes jump from "measurable" to "the answer is wrong." Positional bias is the modulator: genuine, persistent, and the smallest contributor to the failure mode hard rule #4 exists to prevent. If the founder wants one sentence: length sets the floor, retrieval difficulty sets the magnitude, position sets where it bites. What would change this ranking is an interpretability result explaining Chroma's coherent-haystack anomaly. If local semantic coherence turns out to be the driver, the control variable stops being length or distractor count and becomes something closer to "how many plausible answers does this context contain," which would reorder everything below.
The task dependence is nameable, and it is what the rule should key on. Across all three primary sources, the variable that predicts harm is not token count but the semantic distance between what the task needs and what the context contains, plus the number of plausible-but-wrong candidates the context supplies. Chroma's needle-question similarity sweep, NoLiMa's whole design, and the distractor conditions all measure the same axis. A 60KB artifact that is the task is a benign haystack. A 6KB artifact that is topically adjacent to the task but not the task is the malignant case, because it is a distractor with high similarity. Byte size does not distinguish these. That is the crux the 5KB threshold cannot see.
Verdict on hard rule #4: the policy is sound, the trigger is wrong, and the rule is right for a reason it does not state. Taking the three parts of the question in order.
Is 5KB the right control variable? No, on three independent counts. Wrong unit: bytes are a 1.6x-noisy proxy for tokens, since prose runs near 4 bytes per token and markup and JSON near 3. Token counts are trivially available and strictly better. Wrong magnitude: 5KB is roughly 1,300 tokens of prose or 1,700 of HTML. No result in the literature shows meaningful degradation from adding 1,300 tokens. The measured onset is 16K to 32K tokens. The threshold is set one to two orders of magnitude below the evidence, which means the rule fires constantly on artifacts that are harmless. Wrong scope, and this is the real defect: the rule is per-artifact and the damage is cumulative. Reading twenty separate 4KB files is fully compliant with the rule, totals 80KB, and is worse than the single 30KB read the rule forbids. The rule as written cannot catch the failure mode it is named for. Note also that the rule is token-negative at its own threshold: a dispatch prompt plus an agent boot plus a handback costs more tokens than the 1,300 it avoids spending, so every firing between 5KB and roughly 8-10K tokens makes the session strictly more expensive and no more reliable.
Does subagent routing avoid the failure mode or relocate it? Mostly avoids, and for a better reason than the rule claims. The subagent does eat the artifact, so the length penalty is paid, but it is paid once, on a fresh context, under the most favorable retrieval conditions available (the artifact is close to 100% of the haystack, the query is fixed, and there are few competing candidates). Because the parent's context stays short across N artifacts while each subagent's context stays short-plus-one, and because the penalty compounds with accumulated length, N short contexts genuinely beat one long one. Separately, the handback is structurally the mitigation arXiv 2510.05381 recommends: it re-emits the relevant evidence into a short, recent window before reasoning over it. The rule accidentally implements the best-measured fix and credits the wrong mechanism. The one case where delegation only relocates the problem is when the extraction itself requires reasoning across the whole artifact ("find the contradiction between section 2 and section 9"). Delegation converts long-context accumulation into long-context single-pass; it does not solve long-context reasoning, and the rule should stop implying it does.
Costs the rule does not price. Three. Lossy extraction with no error signal: the parent cannot distinguish a faithful extract from a confident fabrication, and the source is gone. This puts hard rule #4 in direct, undocumented tension with the workflow-agent-output-integrity memory, whose seven failure modes (pointer returns, false "verified" stamps, proposal-as-fact) are precisely what a mandatory-delegation rule mass-produces. A carve-out that fires too late: the rule exempts cases where you need to quote exactly, but that judgment must be made before reading the artifact, which is exactly when it cannot be made. The costs are asymmetric. Wrongly delegating is expensive to detect and sometimes impossible to repair if the source is not re-fetchable; wrongly reading costs some context. Judgment serialized ahead of information: the subagent filters using a spec written by a parent that has not seen the artifact, which makes the pattern actively worse for exploratory reading.
Proposed amendment, for founder greenlight (hard rules are immutable from below). Replace the single byte threshold with a three-part trigger plus a floor. (1) Primary, cumulative not per-artifact: delegate when a direct read would push live session context past a budget (propose 25-30% of the window as the soft mark) rather than when one artifact exceeds a size. (2) Secondary, restore Thariq's own test: delegate when the parent needs a conclusion; read directly when the parent needs the text (quotes, exact figures, diffs, or an open-ended "what is in here"). (3) Tertiary, distractor density: delegate aggressively when the artifact is topically adjacent to the live task but is not the task, which is the measured worst case; a large artifact that is the task is comparatively safe to read. (4) Floor: never delegate under roughly 8-10K tokens unless (3) fires, because the round trip costs more than it saves. Demote 5KB from mandate to a prompt to think. Keep the existing examples verbatim — newsletter HTML at 30-100KB is 10K-34K tokens and sits squarely in the measured danger zone, so the rule's illustrations were right even while its threshold was not.
Why this is in the vault
This is the evidence base for a specific pending amendment to AGENTS.md hard rule #4, and it resolves a live contradiction between that rule and the workflow-agent-output-integrity memory that governs how Ray's own subagent fan-outs are trusted. It also corrects [[context-rot]], which currently states as mechanism the one account the primary literature falsifies as sufficient, and which cites tool-count measurements as if they were context-length measurements.
Open follow-ups
- All the primary evidence measures single-pass input length on retrieval and reasoning benchmarks. Ray's actual failure mode is accumulated multi-turn agentic context — tool results, file reads and subagent handbacks piling up over dozens of turns, with structure and redundancy that no long-input benchmark reproduces. Is there any published measurement of a degradation curve for accumulated agentic context, and does the single-pass evidence transfer to it at all?
- Hard rule #4's unpriced cost is extraction fidelity. Is there published measurement of summarization or extraction faithfulness as a function of source-artifact length — specifically, does a subagent's handback get less faithful as the source grows, and at what size does the fidelity loss exceed the context-rot loss it was meant to avoid?
Related
- [[context-rot]] — the vault concept page this brief corrects and extends with mechanism
- [[2026-04-15-thariq-claude-code-session-management-1m-context]] — hard rule #4's cited provenance; contains the conclusion-versus-output test but no byte threshold
- [[2026-07-07-claude-skill-count-degradation-skill-packs]] — the vault's only hard degradation numbers, which are tool-count not context-length
- [[2026-07-24-thariq-context-engineering-claude-5-rules]] — the 80%-system-prompt-deletion result that argues against mechanical over-constraint
- [[2026-02-23-every-chatgpt-memory-context-rot]] — where the vault first picked up the term
- [[2026-09-21-indy-dev-dan-self-compact-pi-agent-zero-hype-devlog]] — practitioner threshold data (soft 225k, forced 270k tokens)
- [[2026-07-26-harness-seven-failure-mode-scorecard]] — the failure-mode frame that mandatory delegation feeds
- [[2026-07-21-technically-harness-engineering]] — the harness-engineering cluster this brief links into
Sources
Vault
06-reference/concepts/context-rot.md06-reference/2026-04-15-thariq-claude-code-session-management-1m-context.md06-reference/research/2026-07-07-claude-skill-count-degradation-skill-packs.md06-reference/2026-07-24-thariq-context-engineering-claude-5-rules.md06-reference/2026-09-21-indy-dev-dan-self-compact-pi-agent-zero-hype-devlog.md~/AGENTS.md(hard rule #4, current wording; file renamed from CLAUDE.md 2026-09-18)
Web (primary)
- Chroma Technical Report, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," July 2025 — https://www.trychroma.com/research/context-rot (code: https://github.com/chroma-core/context-rot)
- Modarressi et al., "NoLiMa: Long-Context Evaluation Beyond Literal Matching," ICML 2025 (Adobe Research) — https://arxiv.org/html/2502.05167v3
- "Context Length Alone Hurts LLM Performance Despite Perfect Retrieval," arXiv:2510.05381 — https://arxiv.org/pdf/2510.05381
- Anthropic Engineering, "Effective context engineering for AI agents" — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Hsieh et al., "Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization" — https://arxiv.org/pdf/2406.16008
- "Mitigate Position Bias in LLMs via Scaling a Single Hidden State," Findings of ACL 2025 — https://aclanthology.org/2025.findings-acl.316.pdf