06-reference/research

context rot mechanism

2026-09-27·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)

Context rot: which mechanism actually dominates, and whether hard rule #4 is calibrated to it

The question

Verbatim: "What is 'context rot' mechanistically (attention dilution vs positional bias vs retrieval degradation — which dominates), and how does the evidence map onto CLAUDE.md hard rule #4?"

Context: "context rot" appears 10+ times across the vault's harness-engineering cluster and is the stated justification for an AGENTS.md hard rule, but the vault's concept page ([[context-rot]]) describes the phenomenon without adjudicating the mechanism, and the rule's 5KB trigger has never been traced back to a measurement.

What we already know (from the vault)

What the web says

Convergences and contradictions

Synthesis for RDCO

Mechanistic verdict, stated as a ranking rather than a winner. There is no single dominant mechanism, and the entanglement is structured rather than mushy. Attention dilution in the strict, content-independent sense is the floor: an irreducible penalty for processing longer sequences that survives whitespace substitution, attention masking, and optimal positioning, and is best explained today by training-distribution mismatch plus representational crowding rather than by anything competing for attention. Retrieval degradation is the amplifier, and it owns most of the variance in realistic work: when the needed fact must be found by semantic rather than literal match, or competes with topically-adjacent distractors, effect sizes jump from "measurable" to "the answer is wrong." Positional bias is the modulator: genuine, persistent, and the smallest contributor to the failure mode hard rule #4 exists to prevent. If the founder wants one sentence: length sets the floor, retrieval difficulty sets the magnitude, position sets where it bites. What would change this ranking is an interpretability result explaining Chroma's coherent-haystack anomaly. If local semantic coherence turns out to be the driver, the control variable stops being length or distractor count and becomes something closer to "how many plausible answers does this context contain," which would reorder everything below.

The task dependence is nameable, and it is what the rule should key on. Across all three primary sources, the variable that predicts harm is not token count but the semantic distance between what the task needs and what the context contains, plus the number of plausible-but-wrong candidates the context supplies. Chroma's needle-question similarity sweep, NoLiMa's whole design, and the distractor conditions all measure the same axis. A 60KB artifact that is the task is a benign haystack. A 6KB artifact that is topically adjacent to the task but not the task is the malignant case, because it is a distractor with high similarity. Byte size does not distinguish these. That is the crux the 5KB threshold cannot see.

Verdict on hard rule #4: the policy is sound, the trigger is wrong, and the rule is right for a reason it does not state. Taking the three parts of the question in order.

Is 5KB the right control variable? No, on three independent counts. Wrong unit: bytes are a 1.6x-noisy proxy for tokens, since prose runs near 4 bytes per token and markup and JSON near 3. Token counts are trivially available and strictly better. Wrong magnitude: 5KB is roughly 1,300 tokens of prose or 1,700 of HTML. No result in the literature shows meaningful degradation from adding 1,300 tokens. The measured onset is 16K to 32K tokens. The threshold is set one to two orders of magnitude below the evidence, which means the rule fires constantly on artifacts that are harmless. Wrong scope, and this is the real defect: the rule is per-artifact and the damage is cumulative. Reading twenty separate 4KB files is fully compliant with the rule, totals 80KB, and is worse than the single 30KB read the rule forbids. The rule as written cannot catch the failure mode it is named for. Note also that the rule is token-negative at its own threshold: a dispatch prompt plus an agent boot plus a handback costs more tokens than the 1,300 it avoids spending, so every firing between 5KB and roughly 8-10K tokens makes the session strictly more expensive and no more reliable.

Does subagent routing avoid the failure mode or relocate it? Mostly avoids, and for a better reason than the rule claims. The subagent does eat the artifact, so the length penalty is paid, but it is paid once, on a fresh context, under the most favorable retrieval conditions available (the artifact is close to 100% of the haystack, the query is fixed, and there are few competing candidates). Because the parent's context stays short across N artifacts while each subagent's context stays short-plus-one, and because the penalty compounds with accumulated length, N short contexts genuinely beat one long one. Separately, the handback is structurally the mitigation arXiv 2510.05381 recommends: it re-emits the relevant evidence into a short, recent window before reasoning over it. The rule accidentally implements the best-measured fix and credits the wrong mechanism. The one case where delegation only relocates the problem is when the extraction itself requires reasoning across the whole artifact ("find the contradiction between section 2 and section 9"). Delegation converts long-context accumulation into long-context single-pass; it does not solve long-context reasoning, and the rule should stop implying it does.

Costs the rule does not price. Three. Lossy extraction with no error signal: the parent cannot distinguish a faithful extract from a confident fabrication, and the source is gone. This puts hard rule #4 in direct, undocumented tension with the workflow-agent-output-integrity memory, whose seven failure modes (pointer returns, false "verified" stamps, proposal-as-fact) are precisely what a mandatory-delegation rule mass-produces. A carve-out that fires too late: the rule exempts cases where you need to quote exactly, but that judgment must be made before reading the artifact, which is exactly when it cannot be made. The costs are asymmetric. Wrongly delegating is expensive to detect and sometimes impossible to repair if the source is not re-fetchable; wrongly reading costs some context. Judgment serialized ahead of information: the subagent filters using a spec written by a parent that has not seen the artifact, which makes the pattern actively worse for exploratory reading.

Proposed amendment, for founder greenlight (hard rules are immutable from below). Replace the single byte threshold with a three-part trigger plus a floor. (1) Primary, cumulative not per-artifact: delegate when a direct read would push live session context past a budget (propose 25-30% of the window as the soft mark) rather than when one artifact exceeds a size. (2) Secondary, restore Thariq's own test: delegate when the parent needs a conclusion; read directly when the parent needs the text (quotes, exact figures, diffs, or an open-ended "what is in here"). (3) Tertiary, distractor density: delegate aggressively when the artifact is topically adjacent to the live task but is not the task, which is the measured worst case; a large artifact that is the task is comparatively safe to read. (4) Floor: never delegate under roughly 8-10K tokens unless (3) fires, because the round trip costs more than it saves. Demote 5KB from mandate to a prompt to think. Keep the existing examples verbatim — newsletter HTML at 30-100KB is 10K-34K tokens and sits squarely in the measured danger zone, so the rule's illustrations were right even while its threshold was not.

Why this is in the vault

This is the evidence base for a specific pending amendment to AGENTS.md hard rule #4, and it resolves a live contradiction between that rule and the workflow-agent-output-integrity memory that governs how Ray's own subagent fan-outs are trusted. It also corrects [[context-rot]], which currently states as mechanism the one account the primary literature falsifies as sufficient, and which cites tool-count measurements as if they were context-length measurements.

Open follow-ups

  1. All the primary evidence measures single-pass input length on retrieval and reasoning benchmarks. Ray's actual failure mode is accumulated multi-turn agentic context — tool results, file reads and subagent handbacks piling up over dozens of turns, with structure and redundancy that no long-input benchmark reproduces. Is there any published measurement of a degradation curve for accumulated agentic context, and does the single-pass evidence transfer to it at all?
  2. Hard rule #4's unpriced cost is extraction fidelity. Is there published measurement of summarization or extraction faithfulness as a function of source-artifact length — specifically, does a subagent's handback get less faithful as the source grows, and at what size does the fidelity loss exceed the context-rot loss it was meant to avoid?

Related

Sources

Vault

Web (primary)