06-reference

Vercel — how our agents build on-brand pages with design.md

2026-09-01·reference·status: assessed·source: https://vercel.com/blog/how-our-agents-build-on-brand-pages-with-design-md
design-contractsagent-harnessscribble-worksevals

Vercel: agents build on-brand pages with design.md

Why this is in the vault

Founder-shared 2026-09-01 ~02:47 ET; verdict READ. Independent convergence on RDCO's design-contract architecture (per-brand DESIGN-*.md + critic gates), with four mechanisms we lack. Assessment based on a subagent summary of the page (fetched 2026-09-01; no embedded instructions found).

Their system (as claimed)

One public design.md (page structure, copywriting standards, visual composition, publishing specs, named anti-patterns) loaded into agent context; brand tokens live in a runtime stylesheet (vercel-brand.css) so agents pick class names — tokens never ride in the prompt. Enforcement: deterministic pre-ship code checks mirroring prose rules; an eval harness of 7 frozen real scenarios for blind A/B of guidance versions; versioned run records (prompt, inputs, model config, design.md version, screenshots, feedback); weekly complaint clustering with each accepted fix routed to exactly one layer (prose / stylesheet / code check / skill) at the narrowest enforcement point. Claimed result: 57% fewer deterministic-check failures with design.md in context (39 vs 91 across 6 pages — sample honestly caveated by them; every page still had ≥1 ship-blocking failure).

Deltas vs RDCO's system

They have, we lack: (1) deterministic pre-ship checks ahead of any critic; (2) an eval harness for CONTRACT changes (we iterate artifacts, never A/B edits to the contract itself); (3) versioned run records tying output → contract version; (4) stylesheet-as-constraint (prompt-free tokens); (5) formalized named-anti-patterns section.

We have, they don't mention: post-deploy Playwright capture at 2 widths, a scoring critic (5 brand tells + checklist) with a bounded ≤3-round iterate loop, per-brand contract separation, and the /improve complaint-clustering analog already running weekly.

Their gates are pre-and-mechanical; ours post-and-perceptual. Complementary — the complete system is both.

Queued actions (studio queue items 24-25)

  1. Deterministic pre-check layer for Scribble Works pages (mechanical rules from [[DESIGN-scribble-works]] non-negotiables that code can check: token-only colors, min tap targets, type-scale steps, no naked previews) running in prebuild before design-critic.
  2. Frozen-scenario eval harness for DESIGN-*.md contract edits — A/B a contract change on fixed scenarios before adopting (extends the skill-retune discipline to design contracts).
  3. Consider: named anti-patterns section in [[DESIGN-scribble-works]]; move remaining prompt-borne tokens into the site stylesheet contract.

Cf. also the prior scoring critic run at [[2026-08-31-design-critic-findings-merged-site]], which supplies the "we have" baseline above.