01-projects/phdata

Finance vertical — first production run of the brigade house (build + eval + stress)

·run-report·status: final
brigadeskill-agent-brigadefinanceexecution-evalreplication

Finance vertical — first production run of the consolidated brigade house

One line: the founder's post-refinement brigade house (2026-07-06 bundle) ran a full production vertical hands-off on the mini — 10 finance skills ordered, built, gated, and close-out-signed with zero station/infrastructure errors (quality defects were caught and are documented below — that's the gates working, not absence of problems); execution-eval then decided which four earned shipping (merged to ray-plugins as discipline-skills 0.2.0, PR #10), and the whole day doubled as the replication test the founder asked for: "make sure it works for others as well as it works for you."

1. Substrate verification (the replication test)

Ran the HANDOFF-README deploy runbook literally, as a stranger would, on the Discord zip (sha256-identical, now durably installed at ~/Projects/phdata-private/brigade-house/). All six mise gates green, all suites at-or-above README counts (factory 140 · assessment 327+6 · company-research 238 · sales-collateral 177+15 · frontend tsc/69/64/build clean).

Replication findings (the team will hit these):

  1. Runbook says re-point "five" mise.toml [roots]; there are six (ab-registrar). Layout table also says four ab- plugins; registrar makes five.
  2. The runbook's own mise command (uv run --with pyyaml --with jsonschema) FAILS the frontend's gate by construction — its pytest-importable check needs --with pytest (17 PASS / 0 FAIL once added).
  3. ab-sales-collateral suite: 177 passed/15 skipped here vs README's 164/28 — same 192 total; 13 skip conditions are environment-sensitive. Document to prevent count-drift panic.

2. The build (steward → rail → walk → close-out)

3. Execution-eval (334 agents, 0 errors, ~17 min)

Two-arm ablation per the execution-eval station contract: base vs base+skill, n=3/arm, oracle fixtures from the acceptance contracts, fixed per-fixture output schemas graded by code (no LLM name-matching), tier sweep (haiku/sonnet/opus) on two skills.

verdict skills evidence
ADVANCE (sonnet wins) annual-budget-build · close-management · treasury-liquidity-analysis +0.33 proration trap · +0.33/+0.67 accrual/correcting-entry sign conventions · +0.67 restricted-cash trap (base hit the documented trap verbatim: quick 1.30 vs 1.23)
ADVANCE at haiku (tier-floor asset) debt-schedule haiku base 0/3 on the annuity formula → 3/3 with skill; +0.33 on ACT/360. Sonnet/opus at ceiling. Value = minimum-viable-tier reduction — the third kind of lift, measured
Quality-lift, one named gap management-reporting-package base packs 5–7/10 rubric criteria → skill packs 9/10 ×3, all failing exactly criterion 8 (number traceability). Refire before ship
INCONCLUSIVE (base at ceiling) rolling-forecast-update · financial-statements · cash-flow-forecasting · reconciliation · capital-budgeting-analysis 2026 sonnet aces textbook fixtures — nothing could show lift. Not kills: fixtures need hardening (messier inputs, weaker-tier arms)

Verify-before-report caught two false negatives in the eval harness itself (two more grader-fragility catches, continuing the pattern first hit in the 2026-06-29 eval runs — grading defects manufacturing fake results until transcripts are read):

Critic-consistency finding: three of four shipped skills left the line with descriptions over the 1,024-char lint limit (≈1,350/2,000/1,810 by yaml-parsed description length) — the deterministic lint axis fired on close-management but evidently not on the others. The "deterministic" axis is only deterministic if the critic agent actually executes the code. Fix candidate: run skillLint() as a mechanical pre-critic step in the walk itself, not inside the critic agent's discretion. Trimmed at packaging; noted in PR #10.

4. Monarch out-of-sample stress arm (cash-flow-forecasting vs real personal data)

Verdict: degrades gracefully with a careful executor — every input class with no personal analog (AR aging, revolver, borrowing base) was named absent, nothing fabricated. Three hardening findings: (1) the scope fence gates on task shape, not entity type — personal forecasting sails through unflagged; (2) Step 8's corporate column vocabulary ("gross AR", "draw/repay") creates fabrication pressure on a literal executor; (3) "name what's missing" conflates missing-value with category-is-a-structural-non-fit. Full report: brigade-runs/cash-flow-forecasting/eval/monarch-stress-report.md (house cellar). Founder's data stayed local; nothing entered any shipped artifact.

5. Shipped + founder rulings today

6. Next-brigade brainstorm (the founder's "expand and compound" ask)

Ranked by leverage on today's evidence:

  1. Fixture-hardening capability (add-station or iterate-skill wave): today's #1 eval fact is 27/37 fixtures non-discriminating at sonnet — the bottleneck on proving skills is now fixture difficulty, not build quality. An oracle-hardening station (messier inputs, multi-step chains, noisy data, weaker-tier arms) unblocks the 5 held skills and every future vertical.
  2. Ops vertical (Six Sigma/CPIM anchors): strongest remaining cert anchor in Andrew's grid, computational → oracles self-generate exactly like Finance did.
  3. Data-engineering brigade (the founder's own domain): audit-model + generate-tests already exist as seeds; dbt/Snowflake skills (SQL optimization, model contracts, pipeline triage) are CAF P5-8 delivery assets. Founder-as-expert makes stewarding cheap.
  4. Marketing/Legal (generative) verticals: gated behind the mgmt-reporting criterion-8 refire — prove the generative eval loop closes first, then the exemplar-sourcing steward pattern (resolved 6/29) carries them.
  5. House-infrastructure code lane: the greenlit rail work (rename-claim, DB adapters) run as fire-lane tickets against the plugins repo — the ab-website pattern applied to the house's own canon.

Cross-refs