Sivulka: "You just hired a million bad employees"
Source: X long-form article by George Sivulka (Hebbia CEO, ~10.8k followers), posted 2026-07-14. Traction at fetch time (~18h): 1.1M impressions, 2,359 bookmarks, 201 RTs. Full text pulled via xmcp article.plain_text.
Bias flag: Sivulka runs Hebbia (enterprise AI for knowledge work — finance/legal document analysis). "Encode each firm's nuances into agents" is his product pitch wearing an essay. The historical framing (railroads → modern management) and the workforce analogy are still genuinely sharp. He credits "@ClaudeAI Fable 5, running on far too many loops" for drafting help.
Thesis in one line
AI is a workforce, not software — and the missing trillion-dollar layer is management: evals, context curation, and token discipline, sold as ongoing transformation to incumbents.
The 7 parallels (his structure)
- Tokenmaxxing = throwing bodies at the problem. People overspend on tokens because ~1 in 100 employees can articulate a process clearly enough to give AI context.
- Loops = meetings about meetings. Self-correcting agent loops are brute force compensating for a task never articulated cleanly. "You are spending tokens on spending tokens."
- Wasted tokens = headcount bloat. "Just like 80% of employees do nothing, 80% of tokens today do nothing. Looping is the new empire building."
- 100X tokens = the new 10X engineers. For any job, some token context cuts AI effort by orders of magnitude. "Humans are cheaper than tokens on average, but good tokens are cheaper at scale. Management converts one into the other."
- Context hoarding = job security. "Nobody trains their replacement for free." Tribal knowledge as guild secret; firms are emotionally/structurally/politically wired to reject the tech that needs their context most.
- Evals = the new OKRs. Coding won because it has built-in evals. "A firm's eval suite will become its most valuable resource… No two firms will have the same eval set. Evals will be key to competitive advantage. An organization running generic evals or generic agents has no edge."
- The transformation company = the next trillion-dollar opportunity. Bigger than "neofirms" eating services spend: sell ongoing AI transformation to incumbents (Jevons: every adopted use case surfaces ten more). Palantir comp: it was never selling software, it was selling transformation.
Mapping to RDCO / phData positioning
- §6 is the RDCO moat argument, independently derived, at 1.1M impressions. Compare [[2026-07-09-anthropic-plugin-ecosystem-vs-rdco-brigade-plugins]] — "evals prove it works at all, MISE proves it works HERE"; demo-grade vs delivery-grade. Sivulka adds the firm-specificity claim: generic evals = no edge, which is exactly the per-brigade eval-suite + cellar-exemplar architecture (EVAL-SPEC, tasting sets, freshness watch).
- §7 names phData's DIE/CAF play as the category winner. "AI transformation companies will be 10X larger than any neofirm" = the founder's main bet, argued from the Palantir comp. Useful sales vocabulary for the CAF roadmap conversations (incl. 2026-07-16 CAF roadmap & planning meeting).
- §5 context hoarding rhymes with the 2026-07-14 Sam Hall / Chick-fil-A knowledge-graph thread and the cellar fill problem: the hard part of filling a cellar isn't storage, it's extraction politics. See [[2026-07-14-cellar-ports-adapters-kb-standards]].
- §4 "100X tokens" = skills/exemplars in our vocabulary. The claim that lift comes from encoded context, not model strength, matches the execution-eval findings (judgment lift vs convention lift, 2026-06-29).
Honest counterpoint — the "loops" jab (§2)
His target is brute-force self-fixing loops (agents calling themselves because the task was never specified), not scheduled/managed operations. The eval-gated brigade architecture is the managed alternative he's arguing FOR. But expect "loops are meetings about meetings" to circulate as a quote against agent harnesses generally — worth having the distinction ready: unmanaged retry loops vs contract-gated stations with deterministic exits.
Disposition
- Positioning evidence, not a Sanity Check topic — an SC piece restating this would be derivative (per the no-derivative-pieces rule). Use as a citation when the harness/evals thesis needs external validation.
- Vocabulary worth adopting in sales contexts: "100X tokens," "evals are the new OKRs," "transformation company."