Finance vertical — first production run of the consolidated brigade house
One line: the founder's post-refinement brigade house (2026-07-06 bundle) ran a full production vertical hands-off on the mini — 10 finance skills ordered, built, gated, and close-out-signed with zero station/infrastructure errors (quality defects were caught and are documented below — that's the gates working, not absence of problems); execution-eval then decided which four earned shipping (merged to ray-plugins as discipline-skills 0.2.0, PR #10), and the whole day doubled as the replication test the founder asked for: "make sure it works for others as well as it works for you."
1. Substrate verification (the replication test)
Ran the HANDOFF-README deploy runbook literally, as a stranger would, on the Discord
zip (sha256-identical, now durably installed at ~/Projects/phdata-private/brigade-house/).
All six mise gates green, all suites at-or-above README counts (factory 140 ·
assessment 327+6 · company-research 238 · sales-collateral 177+15 · frontend
tsc/69/64/build clean).
Replication findings (the team will hit these):
- Runbook says re-point "five" mise.toml
[roots]; there are six (ab-registrar). Layout table also says four ab- plugins; registrar makes five. - The runbook's own mise command (
uv run --with pyyaml --with jsonschema) FAILS the frontend's gate by construction — its pytest-importable check needs--with pytest(17 PASS / 0 FAIL once added). - ab-sales-collateral suite: 177 passed/15 skipped here vs README's 164/28 — same 192 total; 13 skip conditions are environment-sensitive. Document to prevent count-drift panic.
2. The build (steward → rail → walk → close-out)
- Steward gathered 10 competency sets (30 cellar notes: core / worked-examples / interpretation per skill), CMA/CFA/CTP-anchored, all examples original. Every fixture's arithmetic verified before enqueue (authors script-verified; steward independently re-derived spot fixtures). Two authors self-caught errors pre-write.
- 10 contract-valid tickets enqueued (Gate A 8/8 at enqueue and at pull, all 10).
- Two service sweeps walked the rail via the Workflow-compatible rail-walk: 10/10 advance, 145 station agents, 0 errors, 6 tickets round-1, 4 round-2.
- Every refire was a legitimate catch: description over the 1,024-char lint limit (close-management) · fabricated fixture numbers caught against the cellar oracle (reconciliation) · missing flag-don't-recompute scenario the test station had demanded (management-reporting) · round-2 passes on debt-schedule + capital-budgeting.
- Service ended clean: journal + lock released, close-out sweep signed all 10, re-scan empty. Filed tickets carry complete build records.
3. Execution-eval (334 agents, 0 errors, ~17 min)
Two-arm ablation per the execution-eval station contract: base vs base+skill, n=3/arm, oracle fixtures from the acceptance contracts, fixed per-fixture output schemas graded by code (no LLM name-matching), tier sweep (haiku/sonnet/opus) on two skills.
| verdict | skills | evidence |
|---|---|---|
| ADVANCE (sonnet wins) | annual-budget-build · close-management · treasury-liquidity-analysis | +0.33 proration trap · +0.33/+0.67 accrual/correcting-entry sign conventions · +0.67 restricted-cash trap (base hit the documented trap verbatim: quick 1.30 vs 1.23) |
| ADVANCE at haiku (tier-floor asset) | debt-schedule | haiku base 0/3 on the annuity formula → 3/3 with skill; +0.33 on ACT/360. Sonnet/opus at ceiling. Value = minimum-viable-tier reduction — the third kind of lift, measured |
| Quality-lift, one named gap | management-reporting-package | base packs 5–7/10 rubric criteria → skill packs 9/10 ×3, all failing exactly criterion 8 (number traceability). Refire before ship |
| INCONCLUSIVE (base at ceiling) | rolling-forecast-update · financial-statements · cash-flow-forecasting · reconciliation · capital-budgeting-analysis | 2026 sonnet aces textbook fixtures — nothing could show lift. Not kills: fixtures need hardening (messier inputs, weaker-tier arms) |
Verify-before-report caught two false negatives in the eval harness itself (two more grader-fragility catches, continuing the pattern first hit in the 2026-06-29 eval runs — grading defects manufacturing fake results until transcripts are read):
- reconciliation "regression −0.33" was a key defect: the key demanded a magnitude (+5,600) where the competency's own table teaches the signed convention (−5,600, GL-minus-subledger). The skill arm followed its source. Keys must pin sign conventions.
- capital-budgeting "flat" was a tolerance defect: ±5 tighter than the oracle's own disclosed rounding-path drift (~$25). Tolerances must absorb legitimate rounding paths.
- Also: close-management D flat is REAL (both arms under-applied the stated materiality rule; the skill didn't fix it) — genuine open gap, documented.
Critic-consistency finding: three of four shipped skills left the line with
descriptions over the 1,024-char lint limit (≈1,350/2,000/1,810 by yaml-parsed
description length) — the deterministic
lint axis fired on close-management but evidently not on the others. The "deterministic"
axis is only deterministic if the critic agent actually executes the code. Fix
candidate: run skillLint() as a mechanical pre-critic step in the walk itself, not
inside the critic agent's discretion. Trimmed at packaging; noted in PR #10.
4. Monarch out-of-sample stress arm (cash-flow-forecasting vs real personal data)
Verdict: degrades gracefully with a careful executor — every input class with no
personal analog (AR aging, revolver, borrowing base) was named absent, nothing
fabricated. Three hardening findings: (1) the scope fence gates on task shape, not
entity type — personal forecasting sails through unflagged; (2) Step 8's corporate
column vocabulary ("gross AR", "draw/repay") creates fabrication pressure on a literal
executor; (3) "name what's missing" conflates missing-value with
category-is-a-structural-non-fit. Full report:
brigade-runs/cash-flow-forecasting/eval/monarch-stress-report.md (house cellar).
Founder's data stayed local; nothing entered any shipped artifact.
5. Shipped + founder rulings today
- ray-plugins PR #10 MERGED (
4ce5caf): discipline-skills 0.2.0 = the four proven skills + eval evidence doc. Five inconclusive + the generative capstone held back with reasons — the plugin's new description states the eval-evidence shipping rule. - Atomic rail claim greenlit (founder, 1:42pm): rename-based claim into
rail/.claimed/<walker-uuid>/composed with his session-UUID attribution idea; unlocks same-brigade multi-walker on the filesystem rail. Next factory ticket. - DB rail adapter targets set: SQLite → Postgres (FOR UPDATE SKIP LOCKED) → Snowflake native (conditional UPDATE + VARIANT). "Postgres on Snowflake" = covered by the PG adapter (Crunchy/Snowflake Postgres speaks PG wire).
6. Next-brigade brainstorm (the founder's "expand and compound" ask)
Ranked by leverage on today's evidence:
- Fixture-hardening capability (add-station or iterate-skill wave): today's #1 eval fact is 27/37 fixtures non-discriminating at sonnet — the bottleneck on proving skills is now fixture difficulty, not build quality. An oracle-hardening station (messier inputs, multi-step chains, noisy data, weaker-tier arms) unblocks the 5 held skills and every future vertical.
- Ops vertical (Six Sigma/CPIM anchors): strongest remaining cert anchor in Andrew's grid, computational → oracles self-generate exactly like Finance did.
- Data-engineering brigade (the founder's own domain): audit-model + generate-tests already exist as seeds; dbt/Snowflake skills (SQL optimization, model contracts, pipeline triage) are CAF P5-8 delivery assets. Founder-as-expert makes stewarding cheap.
- Marketing/Legal (generative) verticals: gated behind the mgmt-reporting criterion-8 refire — prove the generative eval loop closes first, then the exemplar-sourcing steward pattern (resolved 6/29) carries them.
- House-infrastructure code lane: the greenlit rail work (rename-claim, DB adapters) run as fire-lane tickets against the plugins repo — the ab-website pattern applied to the house's own canon.
Cross-refs
- Eval evidence (public repo): ray-plugins
plugins/discipline-skills/evals/2026-07-08-finance-vertical-eval.md - House cellar: filed tickets under
competencies/finance/*/tickets/· built skills underbrigade-runs/*/skill/· eval fixtures + Monarch stress underbrigade-runs/*/eval/ - Prior art: [[2026-07-07-brigade-house-handoff-review]] · the two-kinds-of-lift framing is in ab-skill-factory
examples/generate-tests-eval-report.md(DESIGN.md §5 covers the station mechanics) — today added the third kind: tier-floor