06-reference/research

oi skill description token cost audit

2026-09-14·research-brief·source: deep-research·by Ray Data Co (deep-research synthesis)
skillsprogressive-disclosuremcpcontext-managementorganizational-intelligence

The OI skill descriptions cost about 3-5K tokens always-on, none of them bundle an MCP (Model Context Protocol) tool surface, and the real exposure is that the listing overflows the harness budget on a 200K window

The question

Verbatim: "What is the actual token cost of the CAF skill descriptions specifically (they may exceed Anthropic's ~80-token median), and does any single CAF skill bundle an MCP/tool surface that reintroduces the expensive tool-schema always-on cost?"

Context: this is follow-up 3 from [[2026-07-07-claude-skill-count-degradation-skill-packs]]. That brief concluded that skills are cheap because of progressive disclosure, and that the real ceiling is (a) any bundled tool surface and (b) trigger collision. This brief tests the first assumption empirically. Naming note: the Catalyst Assessment Framework ("CAF") name was retired on 2026-08-10. The work now runs under the Organizational Intelligence (OI) umbrella, so this brief says OI from here on.

What we already know (from the vault)

What the web says

The audit

Method. I read every SKILL.md frontmatter, agent file, and command file read-only, using a small YAML-aware parser that handles folded and literal block scalars. No tokenizer was installed, and I did not call the count_tokens API. Token counts are therefore estimates, and each is given as a range:

OI descriptions are dense with jargon: the average word is 7.1-7.5 characters, against about 5 for plain English. The real figure is probably nearer the chars/4 end, so the per-skill figures below use chars/4.

MCP check. For each tree I searched for:

Aggregate, by skill set:

Skill set Skills Median desc tokens Max Over 80 tok Always-on desc total (tokens) Bundled tool surface
OI caf plugin skills (checkout 2026-07-03) 106 ~24 ~102 10 ~2.7K-3.8K (~4.7K with namespaced names) None. Every skill has allowed-tools: Read Grep (built-in, pre-approval only). No .mcp.json, no mcpServers
OI caf plugin subagents 9 ~90 ~118 6 ~0.8K (Agent-tool listing) None. Built-in tools only (Read, Grep, Glob, Edit, Write, Skill)
OI caf plugin command 1 ~12 - 0 ~12 None
OI project .claude/skills (phase-level tree) 12 ~141 ~178 12 ~1.3K-1.7K None. allowed-tools lists built-ins and Skill(...) only
ab-assessment (brigade-house v2 OI brigade) 11 ~148 ~194 11 ~1.6K None
All brigade-house plugins (the wider house, for comparison) 56 ~175 ~255 48 ~7.0K-9.5K None across all 13 plugins
OI spec repo .claude/skills (orgmap-*) 2 ~83 ~84 2 ~0.17K None
OI Snowflake frontend .claude/skills 2 ~129 ~129 2 ~0.26K None. settings.json holds only enabledPlugins

Per-skill outliers in the 106-skill plugin (all descriptions above about 80 tokens; the remaining 96 sit at or below that):

Skill Desc tokens (chars/4) Bundled tool surface Always-on cost
m-3-expectation-steward ~102 No (Read, Grep pre-approval) ~105
4g-3-deliverable-expert-reviewer ~102 No ~105
2-10-orchestrator ~100 No ~103
s4-7-archetype-propagation-wiring ~95 No ~98
3-9a-solution-archetype-classifier ~95 No ~98
4g-4-real-world-value-validator ~91 No ~94
2-9-exec-translator ~91 No ~94
2-4-process-input-discovery-detector ~89 No ~92
6d-8-adversarial-correctness-validator (in the cut 36) ~84 No ~87
s4-6-platform-native-decomposition-engine ~83 No ~86

Per-skill table for the current v2 set (ab-assessment): 01-engagement-framing ~148 · 02-discovery-readiness ~140 · 03-classification ~178 · 04-prioritization ~165 · 05-skill-contracts ~164 · 06-execution ~150 · 07-quality-gates ~133 · 08-productize ~134 · catalyst-assessment ~109 · expo ~194 · prime-catalyst-assessment ~115. No skill bundles a tool surface. The always-on cost of each equals its description tokens plus about 3-5 tokens for the name.

No description in any tree comes near the 1,536-character per-entry cap. The largest, standards-regulatory-sourcing in ab-domain-research, is 1,021 characters.

Convergences and contradictions

Synthesis for RDCO

The token question is settled: OI's always-on skill cost is small, and no OI skill carries a tool surface. The worst case is the full 106-skill plugin plus its 9 subagents: about 5.5K tokens including names. The live ab-assessment brigade is about 1.6K. Both are rounding errors next to one mid-size MCP server: Slack alone is about 21K, and Anthropic's five-server example is about 55K. So nothing in the OI build breaks the progressive-disclosure assumption the parent brief rested on. The skills-are-cheap conclusion survives the empirical test. The ~80-token median framing also needs a correction. The original catalog comes in under it. The newer phase-level skills come in about 1.8x over it, and that is a deliberate house style that buys routing precision with tokens.

The binding constraint moves from token cost to the listing budget, which is the parent brief's trigger-collision finding in harness form. Claude Code does not let skill descriptions grow context without limit. It reportedly caps the listing at about 1% of the window and drops descriptions from the least-used skills first. So a client installing the 106-skill OI plugin on a 200K-window model would pay a bounded token cost, with no drop in answer quality. The cost is that many of the micro-skills would become name-only entries Claude cannot route to. This matters for any OI deliverable that ships the full catalog into a client harness, and it favors the v2 shape: 11 phase skills that load the ~106 dimensions as reference files (Level 3). That shape was presumably chosen for other reasons, but it is also the budget-safe one. Keep new OI skills on the phase-skill-plus-reference pattern, not on one skill per dimension.

The real tool-schema risk is prospective, not present. It sits in the planned remote OI data MCP server and in the Glean MCP endpoint that the restructure brief proposed to consume. If either ships inside the OI plugin through .mcp.json, it becomes an always-on schema cost for every user of the plugin, and that is exactly the failure mode this question worried about. Two mitigations exist:

Keep that separation for when the server is built. It is a design rule to carry forward, not a problem to fix now.

Why this is in the vault

This closes open follow-up 3 of [[2026-07-07-claude-skill-count-degradation-skill-packs]]. It also gives the OI plugin-packaging decision in [[2026-06-14-caf-restructure-organizing-brief]] (skills plugin versus separate remote MCP server) a measured number: keep the MCP server out of the skills plugin, and prefer phase skills with reference files over 106 top-level skills for any client-harness install.

Open follow-ups

Related

Sources