"Self-Compact Pi Agent: ZERO HYPE Agentic Coding Devlog" — IndyDevDan
Full transcript: [[2026-09-21-indy-dev-dan-self-compact-pi-agent-zero-hype-devlog-transcript]].
Why this is in the vault
IndyDevDan is a tracked tier-1 channel for RDCO's own agent-harness practice, and this devlog is a direct case study in agent-controlled context management — letting the agent, not a fixed harness default, decide when to compact — which is exactly the kind of context-discipline decision AGENTS.md hard rule #4 (route long artifacts through subagents) is already trying to solve by hand.
Episode summary
Dan builds a "self-compacting" Pi Coding Agent that owns its own compaction tool instead of relying on a fixed-percentage auto-compact default, giving it graduated soft/warning/forced thresholds plus agent-authored "note to self" hints carried across compaction. He plans the build once in a spec doc with an explicit rubric ("definition of ready/done" plus instant-failure conditions), then runs the identical plan through Claude Code (GLM 5.2), Codex (GPT-5.x "Astra"), and the Pi Coding Agent (Fable 5.1) side by side to compare which model/harness combination actually completes the work and how efficiently.
Key arguments / segments
- [00:00:00] Frames the "classic agent problem" as the context window itself — default fixed-percentage auto-compaction (Claude Code, Codex) is a decision engineers shouldn't cede to the tool.
- [00:04:00] Names the build "self-sealing": a standalone Pi Coding Agent with its own tool to autonomously compress its own context, run standalone, and expose three threshold levels (notification, warning, forced compression).
- [00:07:00] Plans the work through an F3 planning skill with an explicit "definition of ready" / "definition of done" rubric — Dan's framing: models are trained against graded rubrics, so writing your own grading rubric into the prompt lets the agent self-check against it.
- [00:11:00] Sets default thresholds tuned to real cost behavior — notes GPT-5.x pricing roughly doubles around the 270k-token mark, so soft warning at 225k, hard warning at 250k, forced compaction at 270k, giving the agent room to wrap up before the price jump.
- [00:13:00] Instant-failure conditions baked into the grading rubric: writing outputs outside the working directory, one agent reading another agent's isolated spec/programs folder — used deliberately to run clean side-by-side model comparisons without cross-contamination.
- [00:16:00]–[00:18:00] Runs the identical plan through three harness/model pairs simultaneously: GLM 5.2 (Claude Code) blows its context at 98% usage and fails to complete; Codex/GPT-5.x "Astra" finishes in ~21 minutes at ~136k tokens; Pi/Fable 5.1 finishes in ~50 minutes using ~500k of its 1M-token window.
- [00:21:00]–[00:22:00] Live demo of the mechanism working: Fable 5.1 digests a large file, crosses the soft threshold (accepts and continues), then crosses the harder threshold, writes a "note to self," and self-compresses mid-task without human intervention.
- [00:24:00] Notes harness-level control varies by tool: Pi lets you fully override the default compaction message, Codex partially, Claude Code not at all.
- [00:25:00] Qualitative model comparison on hint-authoring quality: judges Fable 5.1's self-written compaction notes as better-engineered than GPT-5.x "Astra"'s.
- [00:26:00] Thesis line: "if you master the four basic tools — the context, the tool, the model — you master the agent... you master knowledge work."
Notable claims
- GLM 5.2 hit 98% context usage and failed to complete the self-compaction build task, attributed to heavy internal "thinking" token consumption — used as the episode's central proof point for why self-sealing matters. [00:17:00]
- Codex/GPT-5.x "Astra" completed the same build in ~21 minutes using ~136k tokens; Pi/Fable 5.1 completed in ~50 minutes using ~500k of a 1M-token context — roughly 3.7x more tokens for over 2x the wall-clock time, on an identical plan. [00:18:00]
- Claimed GPT-5.x pricing roughly doubles once context crosses ~270k tokens — the stated basis for his 225k/250k/270k threshold defaults (unverified against OpenAI's published pricing tiers in this note; treat as Dan's own operating assumption, not confirmed vendor pricing). [00:11:00]
- Named technique: "self-sealing" / self-compacting agent pattern — agent owns a dedicated compaction tool with three graduated thresholds (soft notify, hard warning, forced compress) plus an agent-authored "note to self" hint carried across the compaction boundary.
- Harness override capability ranked: Pi Coding Agent (full override of default compact message) > Codex (partial) > Claude Code (no override). [00:24:00]
- Codebase referenced and shared:
github.com/disler/self-compact-pi-agent.
Guests
Solo — no guests.
Mapping against Ray Data Co
This maps directly onto RDCO's own operating discipline, not just abstractly. AGENTS.md hard rule #4 already mandates routing long artifacts (>5KB) to subagents rather than reading them raw into parent context — Dan's self-sealing pattern is the same problem attacked from the other side: instead of (or in addition to) pre-filtering what enters context, give the agent itself graduated awareness of its own context budget and a dedicated tool to act on it. Two concrete transferable ideas: (1) the "definition of ready / definition of done" rubric-as-prompt technique, which is close in spirit to the sub-agent dispatch acceptance-contract pattern already used in RDCO's own dispatch prompts (verify-dispatch skill); (2) the instant-failure isolation trick (each agent forbidden from touching another agent's spec/programs directory) as a clean way to run true side-by-side model bake-offs — relevant if RDCO ever wants to A/B different models against the same skill spec rather than trusting a single model's self-report. The GLM-5.2-context-blowout result is also a concrete data point for model selection on long-horizon autonomous work.
Related
- [[2026-04-15-thariq-claude-code-session-management-1m-context]] — the Thariq context-rot guidance cited directly in RDCO's own AGENTS.md hard rule #4; this video is a practitioner-side implementation of the same context-budget-awareness problem.
- [[2026-08-24-indy-dev-dan-intelligence-explosion-harness-engineering]] — same channel, same "harness engineering" framing, prior multi-model orchestration demo (opinion/debate/collaborate commands) that this video's side-by-side model bake-off extends.
- [[2026-08-31-indy-dev-dan-agentic-operating-level]] — Dan's "agentic operating level" framework, referenced in this video's own description as a companion piece ("WHERE to FOCUS your AGENTS").
- [[2026-06-01-indy-dev-dan-coding-agent-observability]] — related prior coverage of agent context/observability tooling from the same channel.