01-projects/phdata

CAF Technical Architecture — current state, target state, and delivery backlog

2026-07-27·architecture + planning·status: draft — for review
phdatacafarchitecturebacklogjira321gomarketplace

CAF Technical Architecture

Purpose. Give engineers enough shared picture to start work, and give the PM a backlog that sequences correctly. Sections 1–3 are the shared picture. Section 4 is the work. Section 5 is the ticket set.

Scope. CAF only.

Post-refactor structure (2026-07-27). CAF is now four phases, terminating in the C4 Build Manifest. Phases 5–8 have been cut.

Phase Name Contract emitted
1 Engagement Framing C1 Charter
2 Discovery Readiness C2 Candidate Dossier
3 Classification C3 Archetype Register
4 Prioritization C4 Build Manifest

Naming (settled 2026-07-27). Three repositories:

Repo Holds Audience
caf-website the sales-facing marketplace surface sales reps, prospects
caf-artifacts outputs generated by CAF runs CAF engineers, the website build
caf-engine the skills and orchestration that run CAF CAF engineers

Casing: the working list mixed CAF-website / caf-artifacts / CAF-engine. Recommend all-lowercase kebab-case for all three — it is the Bitbucket/Git convention, avoids case-sensitivity bugs between macOS (case-insensitive) and Linux CI runners, and stops the mixed form propagating into clone paths and CI config. Cheap to fix now, annoying later.


Section 1 — Current structure

Where things live

Component Location Status
CAF website repo ___ TODO — need repo path
CAF artifacts Google Drive no version control
CAF engine repo ___ TODO — need repo path

Current flow

flowchart TD
    A["Request submitted<br/>#CAF-requests Slack channel"] --> B["Glean agent<br/>parses request"]
    B --> C["CAF-intake<br/>Jira board"]
    C --> D["CAF-trained engineer<br/>self-assigns ticket"]
    D --> E["Run CAF engine<br/>locally"]
    E --> F["Artifacts produced<br/>on local disk"]
    F -->|"MANUAL<br/>copy + paste"| G["Google Drive"]
    G -->|"MANUAL<br/>copy + paste"| H["Claude design template"]
    H --> I["CAF website<br/>generated"]
    I --> J["Review:<br/>CAF PM + sales rep"]
    J --> K["Sales rep presents<br/>marketplace"]
    K --> L["Deal closes"]

    style F fill:#ffe6e6,stroke:#cc0000
    style G fill:#ffe6e6,stroke:#cc0000
    style H fill:#ffe6e6,stroke:#cc0000

Red nodes are the manual handoffs. Every one of them is a place where state is lost, drifts, or has to be redone.

The rework loop

The failure case worth designing against explicitly:

sequenceDiagram
    participant S as Sales rep
    participant J as Jira
    participant E1 as Engineer A
    participant E2 as Engineer B
    participant G as Google Drive

    E1->>G: uploads SOME artifacts from original run
    S->>J: requests updates to the demo
    J->>E2: new ticket, different engineer
    E2->>G: looks for run state
    Note over E2,G: partial upload → context is gone
    E2->>E2: re-runs CAF from scratch
    Note over E2: tokens + hours burned reproducing<br/>work that already happened

This is the compounding cost. Everything else on the pain list is a papercut by comparison.


Section 2 — Pain points

Grouped by root cause rather than as a flat list, because the groups map to different fixes.

2.1 Non-determinism → inconsistent demos

Root cause: generation is doing work that should be done by a template. Anything that must be identical every time should not be regenerated every time.

2.2 Manual handoffs → lost engineering time

Root cause: no system of record. Drive is being used as one and is not one.

2.3 Token waste

Root cause: the engine has no declared output contract. See §3.3 — this remains the highest-leverage fix in the document, and the refactor makes it harder rather than easier.

2.4 Operational robustness


Section 3 — Updated architecture

3.1 Target flow

flowchart TD
    A["Request submitted<br/>#CAF-requests Slack"] --> B["Glean agent"]
    B --> C["CAF-intake Jira board"]
    C --> D["Run kicked off"]
    D --> E["caf-engine<br/>remote sandboxed agent harness"]
    E --> F["caf-artifacts repo<br/>version controlled"]
    F --> G["CI/CD pipeline"]
    G --> H["Object store<br/>S3 or equivalent"]
    H --> I["caf-website<br/>hydrates dynamically"]
    I --> J["Review: CAF PM + sales rep"]
    J --> K["Sales rep presents"]

    E -.->|"reads prior runs"| F

    style E fill:#e6f7e6,stroke:#2d7a2d
    style F fill:#e6f7e6,stroke:#2d7a2d
    style H fill:#e6f7e6,stroke:#2d7a2d

No manual steps between kickoff and review. The dotted line is the reuse path — the engine reading prior artifacts is what kills the "re-run competitive research we already did" waste.

3.2 The headless constraint is real — design for it, don't wish it away

Some phases must run interactively to satisfy enterprise approval guardrails. A "remote sandboxed agent harness" that assumes full headless operation will hit this wall in implementation, not in planning, which is the expensive place to hit it.

Recommended shape: treat the run as resumable with explicit interactive hold points, rather than as one headless job. A run pauses at a guardrail, notifies the assigned engineer, and continues from the same state once cleared. Hold points should be placed by error-propagation cost rather than at uniform intervals — the pattern documented in [[2026-06-11-high-reliability-acceptance-gates-agent-contracts]].

This needs a spike before the harness is built. Ticket CAF-ENG-1.

3.3 After the refactor, trimming the engine has no clean seam — and that raises the stakes on the data model

This is the precise version of "the engine burns tokens on artifacts that don't reach the deliverable," and the phases 5–8 cut changes its shape significantly.

Before the refactor, the trim had a natural boundary. The marketplace is a pre-sales surface, phases 1–4 were the pre-sales half, and 5–8 were post-sale delivery. "Trim to what the website needs" meant stop at C4 — a clean cut at a contract line, low risk, easy to verify.

After the refactor, that seam is gone. CAF is phases 1–4. The website consumes C1–C4. There is no longer a block of phases to switch off, so every remaining token spent is spent inside a phase the website depends on. The trim is now field-level inside four phases rather than phase-level across eight.

Three consequences:

  1. CAF-WEB-1 (the data model) goes from important to load-bearing. It was previously one input to the trim; the phase boundary did most of the work. Now it is the only thing that can tell an engineer what is safe to drop. Trimming before it lands is guessing.

  2. "Not rendered" does not mean "not needed" — this is the trap. An intra-phase artifact can be an input to a later contract without ever appearing on the site. Competitive research feeding C3 classification is the obvious case: the site may never show the research, but C3 is wrong without it. An engineer who reads "trim to what the website needs" literally will cut load-bearing intermediates and break C4 in a way that surfaces as subtly-degraded output rather than a failure. The trim must distinguish:

    • terminal artifacts — rendered by the site → keep
    • intermediate artifacts — feed a downstream contract → keep, but they need not be published to the object store
    • orphan artifacts — neither rendered nor consumed downstream → these are the trim target
  3. The artifact schema should mark which of the three each output is. That turns the trim from a judgment call into a query, and it stops the same analysis being redone every time someone touches the engine.

3.4 Design decisions that still need a call

These are not blockers for starting, but they should be decided before the relevant ticket is picked up rather than during it.

# Decision Why it matters
D1 caf-artifacts layout: one repo with per-client directories, or repo-per-client? Client-confidential material. Per-client repos give clean access control and clean deletion at engagement end; one repo is simpler to operate and to run CI over. Recommend one repo with per-client directories plus a documented retention/deletion policy, unless legal wants hard isolation.
D2 Repo growth and binary artifacts. Run outputs accumulate every engagement, forever. If artifacts include images/PDFs, Git LFS or "artifacts stay in object store, repo holds only manifests + text" needs deciding up front. Retrofitting is painful.
D3 Object store access model. The website hydrates from it. Public-read is off the table (client material). Signed URLs, or website-server-side fetch with credentials — decide before CAF-WEB-1 hardens.
D4 Where does prior-run research live, and what is its staleness rule? The reuse case in §2.4. Reusing 8-month-old competitive research on a fast-moving target is worse than regenerating it. Needs an explicit age/invalidations rule, not just a cache.

Section 4 — What needs to be done

Sequencing note before the list: the data-model review is the critical path, and more so after the refactor (§3.3). It determines what the engine can safely trim and what shape artifacts are stored in. Starting the engine trim or the CI/CD pipeline before it lands means redoing both.

flowchart LR
    DM["CAF-WEB-1<br/>data model review<br/>(C1-C4 mapping)"] --> ENG["engine trim"]
    DM --> ART["artifact schema"]
    ART --> CI["CI/CD to object store"]
    CI --> HYD["website hydration"]
    DM --> MOD["presentation module<br/>isolation"]

    INFRA["DNS · Okta SSO ·<br/>repo creation"] -.->|"parallel, no dependency"| DM

    style DM fill:#fff4cc,stroke:#cc9900

Infrastructure work (DNS, SSO, repo creation) has no dependency on the data model and should run in parallel from day one to avoid a serialized start.

caf-website

caf-artifacts

caf-engine


Section 5 — Recommended Jira tickets

Four epics. Tickets are outcome-named so a reader recognises the result, not the operation. [P1] = start now, [P2] = next, [P3] = after.

EPIC A — caf-website


CAF-WEB-1 · Establish the data model the marketplace actually needs [P1] ← critical path

Description. Audit the current front end and enumerate every field it renders. Map each to its source contract (C1 Charter, C2 Candidate Dossier, C3 Archetype Register, C4 Build Manifest). Produce a typed schema for what the site consumes.

Acceptance criteria.

Blocks: CAF-ENG-3, CAF-ART-2, CAF-WEB-2.


CAF-WEB-2 · Sales presentation runs as a standalone module [P2]

Description. Extract the presentation surface so it can be versioned, tested, and rendered independently of a given client's data.

Acceptance criteria.

Depends on: CAF-WEB-1.


CAF-WEB-3 · Marketplace reachable at 321.phdata.io [P1]

Description. DNS configuration and TLS for the demo domain.

Acceptance criteria. Domain resolves, valid certificate, documented owner for renewal.

Note: parallelisable immediately — no dependency on the data model.


CAF-WEB-4 · Only authenticated phData staff can reach the marketplace [P1]

Description. Okta SSO in front of the site. Sales-associate demo access only for now.

Acceptance criteria.

Note: the last criterion is the one that bites in production. Worth testing deliberately.


CAF-WEB-5 · Industry demo marketplaces available without a client engagement [P2]

Description. Generic per-industry marketplaces for top-of-funnel demos.

Acceptance criteria. At least one industry rendered end-to-end from fixture data; adding an industry requires no code change.

Depends on: CAF-WEB-2.


CAF-WEB-6 · Client-specific marketplaces gated by invite code [P3]

Description. Per-client marketplace behind a password / invite code, layered under the Okta boundary.

Acceptance criteria. Client A's code cannot reach Client B's marketplace; codes are revocable; access attempts logged.

Depends on: CAF-WEB-4. Decision needed: D3.


EPIC B — caf-artifacts


CAF-ART-1 · CAF run outputs live in version control [P1]

Description. Create the Bitbucket repo and migrate existing Drive artifacts.

Acceptance criteria.

Decision needed: D1, D2. Note: the last criterion matters — leaving Drive writable means both systems stay half-true.


CAF-ART-2 · Artifacts conform to a validated schema [P2]

Description. Apply the CAF-WEB-1 schema as a validation gate so drift is caught at commit rather than in a demo. Schema must also classify each artifact as terminal, intermediate, or orphan per §3.3 — that classification is what makes the engine trim a query rather than a judgment call.

Acceptance criteria.

Depends on: CAF-WEB-1. Addresses: §2.1 artifact drift, and unblocks CAF-ENG-3.


CAF-ART-3 · Merged artifacts reach the website without human intervention [P2]

Description. CI/CD publishing artifacts to the object store on merge. Only terminal artifacts need publishing; intermediates stay in the repo.

Acceptance criteria.

Depends on: CAF-ART-1, CAF-ART-2. Decision needed: D3.


CAF-ART-4 · Client artifacts have a documented retention and deletion policy [P3]

Description. Define how long client material is retained, who can read it, and how it is deleted at engagement end.

Acceptance criteria. Written policy reviewed by whoever owns client-data obligations; deletion is executable and tested once.

Note: raised because artifacts are client-confidential and now accumulate permanently in a shared repo. Cheaper to answer now than during a client security review.


EPIC C — caf-engine


CAF-ENG-1 · Spike: which phases can run headless, and where must a run pause [P1]

Description. Enumerate every point requiring interactive execution for enterprise approval guardrails. Produce the hold-point map and a recommended resumable-run design.

Acceptance criteria.

Note: do this before building the remote harness. Discovering interactive guardrails during harness implementation is the expensive path. See §3.2.


CAF-ENG-2 · Engine changes ship through peer-reviewed PRs [P1]

Description. PR-based workflow with review requirements and CI.

Acceptance criteria. Direct pushes to main blocked; ≥1 review required; CI runs on PR; contribution guide written.

Note: parallelisable immediately.


CAF-ENG-3 · Engine stops producing artifacts nothing consumes [P2]

Description. Using the CAF-ART-2 classification, remove generation of orphan artifacts — outputs that are neither rendered by the site nor consumed by a downstream contract. Terminal and intermediate artifacts are both retained.

Per §3.3, this is a field-level trim inside phases 1–4, not a phase-level cut. The phases 5–8 refactor removed the clean seam, so the classification from CAF-ART-2 is the safety mechanism — without it this ticket is guesswork with a silent failure mode.

Acceptance criteria.

Depends on: CAF-WEB-1, CAF-ART-2. Addresses: §2.3 token waste.

Note: the before/after measurement is the point. Without it this is a refactor with an unverified benefit claim. The pre/post diff is the guard against the §3.3 trap.


CAF-ENG-4 · Re-running CAF reuses prior research instead of regenerating it [P2]

Description. Let a run read prior artifacts for the same target company — competitive and target-company research especially — and reuse rather than regenerate.

Acceptance criteria.

Depends on: CAF-ART-1. Decision needed: D4. Note: the staleness rule is not optional. Silently reusing stale competitive research is a worse failure than regenerating it, because it degrades the demo invisibly.


CAF-ENG-5 · Engineer feedback is captured where it can act on the backlog [P2]

Description. A lightweight route for CAF-trained engineers to log friction and improvement suggestions during a run.

Acceptance criteria. Documented capture path; submissions land somewhere the PM triages on a stated cadence.

Note: keep this genuinely lightweight. A feedback process that costs an engineer five minutes mid-run will not be used.


CAF-ENG-6 · CAF runs execute on the remote sandboxed harness [P3]

Description. Move execution from engineer laptops to the remote sandboxed agent harness, incorporating CAF-ENG-1's hold-point design.

Acceptance criteria.

Depends on: CAF-ENG-1, CAF-ART-1, CAF-ART-3. This is the ticket that closes the loop — after it, §1's red nodes are gone.


EPIC D — Presentation consistency


CAF-PRES-1 · The demo looks the same every time it is generated [P2]

Description. Replace per-run generation of the presentation layer with a static template hydrated by data. Addresses §2.1 drift and the §2.3 "Claude design is a token drag" cost in a single change.

Acceptance criteria.

Depends on: CAF-WEB-2.

Note: this is the highest-leverage ticket in the document that is not on the critical path. Determinism here is what makes the sales experience consistent, and it removes the Claude-design token cost as a side effect rather than as a separate effort.


Open items for the PM

  1. Repo paths for caf-website and caf-engine — blanks in §1. Needed before CAF-WEB-1 and CAF-ENG-2 can be assigned.
  2. Decisions D1–D4 (§3.4) — each blocks a specific ticket, noted inline.
  3. Confirm C1–C4 survived the refactor unchanged. This document assumes the four remaining phases still emit the same four contracts. If the refactor also changed what phases 1–4 produce, CAF-WEB-1's mapping target moves and several tickets shift with it.