Transcript: "Self-Compact Pi Agent: ZERO HYPE Agentic Coding Devlog" — IndyDevDan
[00:00:00] Hello engineers, it's you and IndyDevDan. Welcome to the Agent Coding Developer's Journal, where we're going to tackle the classic agent problem — a stubborn problem with agents that won't go away no matter how great your agent engineering skills become. Every agent has it, no agent can avoid it, and every engineer must understand it: the context window. Most agent coding tools like Claude Code and Codex let you set a fixed percentage at which automatic compaction occurs — but your agents are now smart enough to be aware of their own context windows, so what if we let our agents decide when it's time to compress?
[00:01:00] Thanks to harness engineering, we can do much better than the default. If you're running agents every day with longer and longer running cycles, proper context window management pays off with every compaction. Dan's swarm system runs 10 to hundreds of agents working for hours, coordinating to achieve a common goal — this requires self-compaction awareness. Today: building a self-compacting Pi coding agent that is context-aware and has its own specialized tool for determining compaction time.
[00:02:00] The build starts with a rough plan — no software factories, no agents, no voice dictation, just Dan, his thoughts, and a notepad/keyboard. The ability to crystallize thoughts into a prompt your agents can use is the core engineer-agent advantage. In VS Code, Dan has prepared parts for the self-compacting agent: an empty spec catalog, empty prompts, Pi Coding Agent documentation, and a Fable-class "F3 plan" skill.
[00:03:00]
Writing the plan file final-draft-plan.md. Problem statement: "Self-sealing — long-running autonomous agents run out of context; a long-running context shows a decrease in productivity due to context destruction and cost wastage."
[00:04:00] Solution: build a standalone Pi Coding Agent with its own tool to manage this — a "self-shrink" capability with three threshold levels (notification, warning, forced compression), a UI that reflects the three levels, and compression tips/hints for every interaction.
[00:05:00] Key requirement: full control over compaction hints rather than relying on any tool's built-in default. Need soft/warning user prompts, a fallback "man-in-the-loop" manual-compression command, visible thresholds, and the ability to force compression manually — while the design target is long-term autonomous agents operating with no human in the loop.
[00:06:00] Five key pillars named: self-sealing, hint engineering (controlling all the compaction prompts), context control, the Pi interface, and man-in-the-loop commands. Moving into workflow: project root variables so multiple agents/tools can run the plan interchangeably.
[00:07:00] Workflow = three steps: planning, creation, verification, using the F3 planning skill. Introduces the "definition of ready" and "definition of done" sections — framed as a "hack" because models are trained against graded rubrics, so giving the agent a rubric lets it self-grade against it. Referenced back to a prior "Astro Swarms" video where OpenAI's agents lacked a clean stop mechanism.
[00:08:00] Defining "done": workflow complete, fasteners in the right place, self-compaction tool built. The tool's core job: let the agent (1) autonomously compress its own context, and (2) run standalone. Compression = a user prompt triggered "under the hood" at set thresholds or on manual invocation; a "note to self" gets added as an extra hint carried into the post-compaction summary.
[00:09:00] UI spec: a context panel showing usage and three threshold markers (message/notification, warning, strict/forced limit) — cache tokens, free context, and the three threshold tokens. Hint engineering = the actual user prompts triggered during compaction (soft self-compression note, file-location details).
[00:10:00] CLI flags: soft-compress-at (percentage or absolute token count), self-compression-warning-at (stricter, near a hard cutoff), and a user-supplied compression message that overrides the Pi Coding Agent's default compression request.
[00:11:00] Default thresholds tuned for a GPT-5.x "Astra" swarm scenario: Dan notes GPT-model pricing doubles around the 270k-token mark, so he sets soft warning at 225k, hard warning at 250k, and forced compaction at 270k — giving the agent room to wrap up before cost doubling kicks in.
[00:12:00]
Additional flags: compact-buffer, compact invitation message override. Man-in-the-loop commands: self-compress-info, self-compress-now, and the standard /compact. Focus stays on describing the end state, not the mechanism — "our agents are getting better and better at this."
[00:13:00] Grading rubric continues: continuous evaluation against each "definition of done" item and each workflow step. Instant-failure conditions: writing project outputs outside the working directory (except the plan directory), reading any file inside another agent's own spec folder, or one agent touching another agent's isolated "programs" subdirectory — used deliberately so Dan can compare models/tools side by side without cross-contamination.
[00:14:00] Extra grading notes: running tools outside designated directories and using temp files is allowed and not a grading failure; agents should stop immediately and report if they cause an accidental failure. Framed against a "top 5 agent engineering guidelines" video from the prior week — the throughline is building trust signals for which models can run long-horizon work out of the loop.
[00:15:00]
Creates a justfile to run the same plan across Codex, Claude Code, and the Pi Coding Agent against multiple models — testing GLM 5.2 among others. Dan: "if you think you have a winning formula, I can guarantee you that you don't — I don't have it... there are no tools that win right now, except maybe the Pi Coding Agent, because it's customizable, extensible, and you own it."
[00:16:00] Launches three terminal panes side by side: Claude Code, Codex, and Pi, running GLM 5.2, GPT-5.x "Astra," and Fable 5.1 respectively, each agent tasked with building the self-compacting tool per its own plan/build/test workflow. Expects completion in 20 minutes to an hour and steps away.
[00:17:00] Results check-in: GLM 5.2 blew its context (98% used) and could not complete the job — Dan notes it burns heavy "thinking" tokens, illustrating exactly the problem the video is about. Fable 5.1 completed in ~50 minutes.
[00:18:00] Codex/GPT-5.x "Astra" completed in ~21 minutes — roughly half the time and presumably fewer tokens. Fable operated in a 1M-token context window and used ~500k tokens; Astra's usage checked at ~136k tokens, i.e., meaningfully more token-efficient on this task. GLM's failure to produce a usable result is treated as the clearest illustration of why self-sealing matters.
[00:20:00] Live verification pass: sets distributed test thresholds (soft 10%, warning 20%, hard 30%) and re-runs Claude Code to watch the context panel visibly mark the wavy soft-threshold line and the hard cutoff line in the UI.
[00:21:00] Runs Fable 5.1 against a large file-digesting task (via an OpenRouter model) to trigger self-compaction live: watches the agent split a large plan file into 40KB chunks, read them, cross the soft threshold (message shown), and continue with an accepted warning.
[00:22:00] Fable crosses the harder threshold, writes a "note to self," and self-compresses — visibly logging the self-compression note, the compaction decision, and the next action before continuing to read the remaining file content. Confirms the self-sealing mechanism works end-to-end.
[00:23:00] Runs the Codex/GPT-5.x version with equivalent absolute thresholds (message at 100k, warning at 200k, forced compression at 300k) and separately spins up GLM 5.x flash for the same test — reiterating the design goal: a large gap between the notification and warning thresholds, a smaller gap between warning and forced compaction, and the harness refusing further tool calls once forced compaction triggers.
[00:24:00] Compares agent-authored compaction hint quality across tools: notes some harnesses (Pi) let you fully override the default compact message, others (Codex) partially, and Claude Code not at all — another point of harness-level control worth engineering around.
[00:25:00] Side-by-side qualitative comparison: Fable 5.1's self-authored compression notes/hints are judged better-engineered than Astra's (GPT-5.x) — "Fable is a better hint engineer than Astra."
[00:26:00] Wrap-up on mechanism: the agent passes itself a "note to self" message in addition to the standard compaction summary, carrying continuity into the next work cycle. Dan's thesis: "if you master the four basic tools — the context, the tool, the model — you master the agent. If you master the agent, [you master] knowledge work."
[00:27:00] Bigger-picture framing: self-compaction matters because the game is shifting from the "input loop" (sitting at a terminal switching between prompts) to the "output loop" — scaling autonomous agent cycles. Multiple graduated thresholds (soft/hard warnings) let the agent judge timing itself rather than hitting one blunt cutoff.
[00:28:00] The self-compacting Pi agent is now packaged and available for viewers to inspect — built end-to-end by the agents themselves from the plan, following the classic plan/build/test workflow.
[00:29:00] Closing: the level of detail you put into what you ask your agents for is the differentiator now — "anyone can write code... but your ability to ask for exactly what you want is a unique advantage." Teases "Phase 3" as the next step up the stack, reiterates the Pi Coding Agent as his preferred tool because it's fully customizable and un-gated, and signs off: "Focus and keep creating."