06-reference

Sivulka: 'You just hired a million bad employees' — evals-as-moat goes mainstream

2026-07-15·source-assessment·source: https://x.com/gsivulka/status/2077070925154161101 (X long-form article, published 2026-07-14)·by George Sivulka (founder/CEO, Hebbia)
evalsagent-managementpositioning-evidencecafbrigade-housetransformation

Sivulka: "You just hired a million bad employees"

Source: X long-form article by George Sivulka (Hebbia CEO, ~10.8k followers), posted 2026-07-14. Traction at fetch time (~18h): 1.1M impressions, 2,359 bookmarks, 201 RTs. Full text pulled via xmcp article.plain_text.

Bias flag: Sivulka runs Hebbia (enterprise AI for knowledge work — finance/legal document analysis). "Encode each firm's nuances into agents" is his product pitch wearing an essay. The historical framing (railroads → modern management) and the workforce analogy are still genuinely sharp. He credits "@ClaudeAI Fable 5, running on far too many loops" for drafting help.

Thesis in one line

AI is a workforce, not software — and the missing trillion-dollar layer is management: evals, context curation, and token discipline, sold as ongoing transformation to incumbents.

The 7 parallels (his structure)

  1. Tokenmaxxing = throwing bodies at the problem. People overspend on tokens because ~1 in 100 employees can articulate a process clearly enough to give AI context.
  2. Loops = meetings about meetings. Self-correcting agent loops are brute force compensating for a task never articulated cleanly. "You are spending tokens on spending tokens."
  3. Wasted tokens = headcount bloat. "Just like 80% of employees do nothing, 80% of tokens today do nothing. Looping is the new empire building."
  4. 100X tokens = the new 10X engineers. For any job, some token context cuts AI effort by orders of magnitude. "Humans are cheaper than tokens on average, but good tokens are cheaper at scale. Management converts one into the other."
  5. Context hoarding = job security. "Nobody trains their replacement for free." Tribal knowledge as guild secret; firms are emotionally/structurally/politically wired to reject the tech that needs their context most.
  6. Evals = the new OKRs. Coding won because it has built-in evals. "A firm's eval suite will become its most valuable resource… No two firms will have the same eval set. Evals will be key to competitive advantage. An organization running generic evals or generic agents has no edge."
  7. The transformation company = the next trillion-dollar opportunity. Bigger than "neofirms" eating services spend: sell ongoing AI transformation to incumbents (Jevons: every adopted use case surfaces ten more). Palantir comp: it was never selling software, it was selling transformation.

Mapping to RDCO / phData positioning

Honest counterpoint — the "loops" jab (§2)

His target is brute-force self-fixing loops (agents calling themselves because the task was never specified), not scheduled/managed operations. The eval-gated brigade architecture is the managed alternative he's arguing FOR. But expect "loops are meetings about meetings" to circulate as a quote against agent harnesses generally — worth having the distinction ready: unmanaged retry loops vs contract-gated stations with deterministic exits.

Disposition