CAF Technical Architecture
Purpose. Give engineers enough shared picture to start work, and give the PM a backlog that sequences correctly. Sections 1–3 are the shared picture. Section 4 is the work. Section 5 is the ticket set.
Scope. CAF only.
Post-refactor structure (2026-07-27). CAF is now four phases, terminating in the C4 Build Manifest. Phases 5–8 have been cut.
| Phase | Name | Contract emitted |
|---|---|---|
| 1 | Engagement Framing | C1 Charter |
| 2 | Discovery Readiness | C2 Candidate Dossier |
| 3 | Classification | C3 Archetype Register |
| 4 | Prioritization | C4 Build Manifest |
Naming (settled 2026-07-27). Three repositories:
| Repo | Holds | Audience |
|---|---|---|
caf-website |
the sales-facing marketplace surface | sales reps, prospects |
caf-artifacts |
outputs generated by CAF runs | CAF engineers, the website build |
caf-engine |
the skills and orchestration that run CAF | CAF engineers |
Casing: the working list mixed
CAF-website/caf-artifacts/CAF-engine. Recommend all-lowercase kebab-case for all three — it is the Bitbucket/Git convention, avoids case-sensitivity bugs between macOS (case-insensitive) and Linux CI runners, and stops the mixed form propagating into clone paths and CI config. Cheap to fix now, annoying later.
Section 1 — Current structure
Where things live
| Component | Location | Status |
|---|---|---|
| CAF website | repo ___ |
TODO — need repo path |
| CAF artifacts | Google Drive | no version control |
| CAF engine | repo ___ |
TODO — need repo path |
Current flow
flowchart TD
A["Request submitted<br/>#CAF-requests Slack channel"] --> B["Glean agent<br/>parses request"]
B --> C["CAF-intake<br/>Jira board"]
C --> D["CAF-trained engineer<br/>self-assigns ticket"]
D --> E["Run CAF engine<br/>locally"]
E --> F["Artifacts produced<br/>on local disk"]
F -->|"MANUAL<br/>copy + paste"| G["Google Drive"]
G -->|"MANUAL<br/>copy + paste"| H["Claude design template"]
H --> I["CAF website<br/>generated"]
I --> J["Review:<br/>CAF PM + sales rep"]
J --> K["Sales rep presents<br/>marketplace"]
K --> L["Deal closes"]
style F fill:#ffe6e6,stroke:#cc0000
style G fill:#ffe6e6,stroke:#cc0000
style H fill:#ffe6e6,stroke:#cc0000
Red nodes are the manual handoffs. Every one of them is a place where state is lost, drifts, or has to be redone.
The rework loop
The failure case worth designing against explicitly:
sequenceDiagram
participant S as Sales rep
participant J as Jira
participant E1 as Engineer A
participant E2 as Engineer B
participant G as Google Drive
E1->>G: uploads SOME artifacts from original run
S->>J: requests updates to the demo
J->>E2: new ticket, different engineer
E2->>G: looks for run state
Note over E2,G: partial upload → context is gone
E2->>E2: re-runs CAF from scratch
Note over E2: tokens + hours burned reproducing<br/>work that already happened
This is the compounding cost. Everything else on the pain list is a papercut by comparison.
Section 2 — Pain points
Grouped by root cause rather than as a flat list, because the groups map to different fixes.
2.1 Non-determinism → inconsistent demos
- Claude design drifts run to run. It has guidelines, but output varies, so sales reps get an inconsistent demo experience.
- Artifacts drift run to run for the same reason. Two runs of the same input do not reliably produce the same shape.
Root cause: generation is doing work that should be done by a template. Anything that must be identical every time should not be regenerated every time.
2.2 Manual handoffs → lost engineering time
- Copy/paste from local run → Google Drive.
- Copy/paste from Google Drive → Claude design.
- Google Drive is poor for sharing and has no real version control between contributing engineers.
- Partial uploads cause the rework loop diagrammed above.
Root cause: no system of record. Drive is being used as one and is not one.
2.3 Token waste
- Claude design is an unnecessary token drag — regenerating a presentation layer that should be a static template hydrated with data.
- The engine was built without the website as a known target. It produces artifacts and makes decisions that never reach the final deliverable.
Root cause: the engine has no declared output contract. See §3.3 — this remains the highest-leverage fix in the document, and the refactor makes it harder rather than easier.
2.4 Operational robustness
- The engine is not turnkey and is not robust to re-runs.
- Parts cannot run headless — they need interactive execution to pass enterprise approval guardrails. This is a hard constraint on the "fully automated" target and is called out again in §3.2.
- Poor interplay between artifacts and engine, especially reusing research from previous runs (target-company and competitive research is the obvious case — that work is expensive and currently thrown away between runs).
Section 3 — Updated architecture
3.1 Target flow
flowchart TD
A["Request submitted<br/>#CAF-requests Slack"] --> B["Glean agent"]
B --> C["CAF-intake Jira board"]
C --> D["Run kicked off"]
D --> E["caf-engine<br/>remote sandboxed agent harness"]
E --> F["caf-artifacts repo<br/>version controlled"]
F --> G["CI/CD pipeline"]
G --> H["Object store<br/>S3 or equivalent"]
H --> I["caf-website<br/>hydrates dynamically"]
I --> J["Review: CAF PM + sales rep"]
J --> K["Sales rep presents"]
E -.->|"reads prior runs"| F
style E fill:#e6f7e6,stroke:#2d7a2d
style F fill:#e6f7e6,stroke:#2d7a2d
style H fill:#e6f7e6,stroke:#2d7a2d
No manual steps between kickoff and review. The dotted line is the reuse path — the engine reading prior artifacts is what kills the "re-run competitive research we already did" waste.
3.2 The headless constraint is real — design for it, don't wish it away
Some phases must run interactively to satisfy enterprise approval guardrails. A "remote sandboxed agent harness" that assumes full headless operation will hit this wall in implementation, not in planning, which is the expensive place to hit it.
Recommended shape: treat the run as resumable with explicit interactive hold points, rather than as one headless job. A run pauses at a guardrail, notifies the assigned engineer, and continues from the same state once cleared. Hold points should be placed by error-propagation cost rather than at uniform intervals — the pattern documented in [[2026-06-11-high-reliability-acceptance-gates-agent-contracts]].
This needs a spike before the harness is built. Ticket CAF-ENG-1.
3.3 After the refactor, trimming the engine has no clean seam — and that raises the stakes on the data model
This is the precise version of "the engine burns tokens on artifacts that don't reach the deliverable," and the phases 5–8 cut changes its shape significantly.
Before the refactor, the trim had a natural boundary. The marketplace is a pre-sales surface, phases 1–4 were the pre-sales half, and 5–8 were post-sale delivery. "Trim to what the website needs" meant stop at C4 — a clean cut at a contract line, low risk, easy to verify.
After the refactor, that seam is gone. CAF is phases 1–4. The website consumes C1–C4. There is no longer a block of phases to switch off, so every remaining token spent is spent inside a phase the website depends on. The trim is now field-level inside four phases rather than phase-level across eight.
Three consequences:
CAF-WEB-1(the data model) goes from important to load-bearing. It was previously one input to the trim; the phase boundary did most of the work. Now it is the only thing that can tell an engineer what is safe to drop. Trimming before it lands is guessing."Not rendered" does not mean "not needed" — this is the trap. An intra-phase artifact can be an input to a later contract without ever appearing on the site. Competitive research feeding C3 classification is the obvious case: the site may never show the research, but C3 is wrong without it. An engineer who reads "trim to what the website needs" literally will cut load-bearing intermediates and break C4 in a way that surfaces as subtly-degraded output rather than a failure. The trim must distinguish:
- terminal artifacts — rendered by the site → keep
- intermediate artifacts — feed a downstream contract → keep, but they need not be published to the object store
- orphan artifacts — neither rendered nor consumed downstream → these are the trim target
The artifact schema should mark which of the three each output is. That turns the trim from a judgment call into a query, and it stops the same analysis being redone every time someone touches the engine.
3.4 Design decisions that still need a call
These are not blockers for starting, but they should be decided before the relevant ticket is picked up rather than during it.
| # | Decision | Why it matters |
|---|---|---|
| D1 | caf-artifacts layout: one repo with per-client directories, or repo-per-client? |
Client-confidential material. Per-client repos give clean access control and clean deletion at engagement end; one repo is simpler to operate and to run CI over. Recommend one repo with per-client directories plus a documented retention/deletion policy, unless legal wants hard isolation. |
| D2 | Repo growth and binary artifacts. | Run outputs accumulate every engagement, forever. If artifacts include images/PDFs, Git LFS or "artifacts stay in object store, repo holds only manifests + text" needs deciding up front. Retrofitting is painful. |
| D3 | Object store access model. | The website hydrates from it. Public-read is off the table (client material). Signed URLs, or website-server-side fetch with credentials — decide before CAF-WEB-1 hardens. |
| D4 | Where does prior-run research live, and what is its staleness rule? | The reuse case in §2.4. Reusing 8-month-old competitive research on a fast-moving target is worse than regenerating it. Needs an explicit age/invalidations rule, not just a cache. |
Section 4 — What needs to be done
Sequencing note before the list: the data-model review is the critical path, and more so after the refactor (§3.3). It determines what the engine can safely trim and what shape artifacts are stored in. Starting the engine trim or the CI/CD pipeline before it lands means redoing both.
flowchart LR
DM["CAF-WEB-1<br/>data model review<br/>(C1-C4 mapping)"] --> ENG["engine trim"]
DM --> ART["artifact schema"]
ART --> CI["CI/CD to object store"]
CI --> HYD["website hydration"]
DM --> MOD["presentation module<br/>isolation"]
INFRA["DNS · Okta SSO ·<br/>repo creation"] -.->|"parallel, no dependency"| DM
style DM fill:#fff4cc,stroke:#cc9900
Infrastructure work (DNS, SSO, repo creation) has no dependency on the data model and should run in parallel from day one to avoid a serialized start.
caf-website
- Review the front end to establish the necessary supporting data model (map to C1–C4)
- Isolate the sales presentation as a module
- DNS configuration for
321.phdata.io - Generic marketplaces for industry demos
- Company-specific marketplaces behind password / invite code
- Okta SSO to prevent the site being fully public — sales-associate demos only for now
caf-artifacts
- Put under version control / create Bitbucket repo
- CI/CD process to upload artifacts to the object store that feeds the website
caf-engine
- PR-based workflow for all changes
- Capture feedback and improvement suggestions from engineers
- Trim skills to build only the artifacts the website needs (field-level — see §3.3)
Section 5 — Recommended Jira tickets
Four epics. Tickets are outcome-named so a reader recognises the result, not the operation.
[P1] = start now, [P2] = next, [P3] = after.
EPIC A — caf-website
CAF-WEB-1 · Establish the data model the marketplace actually needs [P1] ← critical path
Description. Audit the current front end and enumerate every field it renders. Map each to its source contract (C1 Charter, C2 Candidate Dossier, C3 Archetype Register, C4 Build Manifest). Produce a typed schema for what the site consumes.
Acceptance criteria.
- Every rendered element traced to a named C1–C4 field, or explicitly flagged as unsourced
- Unsourced elements listed as findings against the contracts (not silently added to a new schema)
- Typed schema published to
caf-artifactsas the contract between engine and website - Reviewed by CAF PM and at least one CAF-trained engineer
Blocks: CAF-ENG-3, CAF-ART-2, CAF-WEB-2.
CAF-WEB-2 · Sales presentation runs as a standalone module [P2]
Description. Extract the presentation surface so it can be versioned, tested, and rendered independently of a given client's data.
Acceptance criteria.
- Presentation renders from a fixture data file with no live backend
- No client-specific content in the module itself
- Documented props/inputs matching the CAF-WEB-1 schema
Depends on: CAF-WEB-1.
CAF-WEB-3 · Marketplace reachable at 321.phdata.io [P1]
Description. DNS configuration and TLS for the demo domain.
Acceptance criteria. Domain resolves, valid certificate, documented owner for renewal.
Note: parallelisable immediately — no dependency on the data model.
CAF-WEB-4 · Only authenticated phData staff can reach the marketplace [P1]
Description. Okta SSO in front of the site. Sales-associate demo access only for now.
Acceptance criteria.
- Unauthenticated request receives no marketplace content
- Okta group controls access; group membership documented
- Session behaviour verified for the live-demo case (a rep mid-presentation must not be logged out)
Note: the last criterion is the one that bites in production. Worth testing deliberately.
CAF-WEB-5 · Industry demo marketplaces available without a client engagement [P2]
Description. Generic per-industry marketplaces for top-of-funnel demos.
Acceptance criteria. At least one industry rendered end-to-end from fixture data; adding an industry requires no code change.
Depends on: CAF-WEB-2.
CAF-WEB-6 · Client-specific marketplaces gated by invite code [P3]
Description. Per-client marketplace behind a password / invite code, layered under the Okta boundary.
Acceptance criteria. Client A's code cannot reach Client B's marketplace; codes are revocable; access attempts logged.
Depends on: CAF-WEB-4. Decision needed: D3.
EPIC B — caf-artifacts
CAF-ART-1 · CAF run outputs live in version control [P1]
Description. Create the Bitbucket repo and migrate existing Drive artifacts.
Acceptance criteria.
- Repo created with branch protection
- Directory convention documented (see D1)
- Existing Drive artifacts migrated with provenance retained
- Drive marked read-only / deprecated so it stops accumulating new state
Decision needed: D1, D2. Note: the last criterion matters — leaving Drive writable means both systems stay half-true.
CAF-ART-2 · Artifacts conform to a validated schema [P2]
Description. Apply the CAF-WEB-1 schema as a validation gate so drift is caught at commit rather than in a demo. Schema must also classify each artifact as terminal, intermediate, or orphan per §3.3 — that classification is what makes the engine trim a query rather than a judgment call.
Acceptance criteria.
- Schema validation runs in CI
- A malformed artifact fails the build with a readable error
- Every artifact type carries a terminal/intermediate/orphan classification
- Validation covers the fields the website depends on
Depends on: CAF-WEB-1. Addresses: §2.1 artifact drift, and unblocks CAF-ENG-3.
CAF-ART-3 · Merged artifacts reach the website without human intervention [P2]
Description. CI/CD publishing artifacts to the object store on merge. Only terminal artifacts need publishing; intermediates stay in the repo.
Acceptance criteria.
- Merge to main publishes automatically
- Failed publish alerts a named owner
- Published objects carry run ID and timestamp
- Rollback path documented
Depends on: CAF-ART-1, CAF-ART-2. Decision needed: D3.
CAF-ART-4 · Client artifacts have a documented retention and deletion policy [P3]
Description. Define how long client material is retained, who can read it, and how it is deleted at engagement end.
Acceptance criteria. Written policy reviewed by whoever owns client-data obligations; deletion is executable and tested once.
Note: raised because artifacts are client-confidential and now accumulate permanently in a shared repo. Cheaper to answer now than during a client security review.
EPIC C — caf-engine
CAF-ENG-1 · Spike: which phases can run headless, and where must a run pause [P1]
Description. Enumerate every point requiring interactive execution for enterprise approval guardrails. Produce the hold-point map and a recommended resumable-run design.
Acceptance criteria.
- Every interactive-only step listed with the guardrail it satisfies
- Recommendation on resumable-run state (what is persisted, how a run resumes)
- Explicit statement of what cannot be automated and why
Note: do this before building the remote harness. Discovering interactive guardrails during harness implementation is the expensive path. See §3.2.
CAF-ENG-2 · Engine changes ship through peer-reviewed PRs [P1]
Description. PR-based workflow with review requirements and CI.
Acceptance criteria. Direct pushes to main blocked; ≥1 review required; CI runs on PR; contribution guide written.
Note: parallelisable immediately.
CAF-ENG-3 · Engine stops producing artifacts nothing consumes [P2]
Description. Using the CAF-ART-2 classification, remove generation of orphan artifacts — outputs that are neither rendered by the site nor consumed by a downstream contract. Terminal and intermediate artifacts are both retained.
Per §3.3, this is a field-level trim inside phases 1–4, not a phase-level cut. The phases 5–8 refactor removed the clean seam, so the classification from CAF-ART-2 is the safety mechanism — without it this ticket is guesswork with a silent failure mode.
Acceptance criteria.
- Only artifacts classified
orphanare removed - C1–C4 contracts still validate after the trim
- A full run's output is diffed pre/post to prove no terminal or intermediate artifact was lost
- Token usage per run measured before and after, with the delta recorded
Depends on: CAF-WEB-1, CAF-ART-2. Addresses: §2.3 token waste.
Note: the before/after measurement is the point. Without it this is a refactor with an unverified benefit claim. The pre/post diff is the guard against the §3.3 trap.
CAF-ENG-4 · Re-running CAF reuses prior research instead of regenerating it [P2]
Description. Let a run read prior artifacts for the same target company — competitive and target-company research especially — and reuse rather than regenerate.
Acceptance criteria.
- A second run against the same target reuses prior research
- Reuse is visible in the run log (what was reused, from which run)
- Staleness rule defined and enforced — how old is too old to reuse
Depends on: CAF-ART-1. Decision needed: D4. Note: the staleness rule is not optional. Silently reusing stale competitive research is a worse failure than regenerating it, because it degrades the demo invisibly.
CAF-ENG-5 · Engineer feedback is captured where it can act on the backlog [P2]
Description. A lightweight route for CAF-trained engineers to log friction and improvement suggestions during a run.
Acceptance criteria. Documented capture path; submissions land somewhere the PM triages on a stated cadence.
Note: keep this genuinely lightweight. A feedback process that costs an engineer five minutes mid-run will not be used.
CAF-ENG-6 · CAF runs execute on the remote sandboxed harness [P3]
Description. Move execution from engineer laptops to the remote sandboxed agent harness, incorporating CAF-ENG-1's hold-point design.
Acceptance criteria.
- A run completes end-to-end on the harness, pausing and resuming at each hold point
- Artifacts land in
caf-artifactswith no manual copy - Run is reproducible from its run ID
Depends on: CAF-ENG-1, CAF-ART-1, CAF-ART-3. This is the ticket that closes the loop — after it, §1's red nodes are gone.
EPIC D — Presentation consistency
CAF-PRES-1 · The demo looks the same every time it is generated [P2]
Description. Replace per-run generation of the presentation layer with a static template hydrated by data. Addresses §2.1 drift and the §2.3 "Claude design is a token drag" cost in a single change.
Acceptance criteria.
- Same input data produces byte-identical presentation output across runs
- Design guidelines encoded in the template rather than in a prompt
- Token cost of presentation generation measured before and after
Depends on: CAF-WEB-2.
Note: this is the highest-leverage ticket in the document that is not on the critical path. Determinism here is what makes the sales experience consistent, and it removes the Claude-design token cost as a side effect rather than as a separate effort.
Open items for the PM
- Repo paths for
caf-websiteandcaf-engine— blanks in §1. Needed before CAF-WEB-1 and CAF-ENG-2 can be assigned. - Decisions D1–D4 (§3.4) — each blocks a specific ticket, noted inline.
- Confirm C1–C4 survived the refactor unchanged. This document assumes the four remaining phases still emit the same four contracts. If the refactor also changed what phases 1–4 produce, CAF-WEB-1's mapping target moves and several tickets shift with it.