Founder hit a practice-exam question (Q13/25, "Tool Design & MCP Integration") claiming the recommended maximum is 4-5 tools per agent, degrading sharply past 18. He flagged it because he'd never heard a tool-count limit. His instinct was right that 4-5 is wrong; my first answer — that Anthropic publishes no number at all — was also wrong. Both corrected below.
What Anthropic actually publishes
From the tool search tool docs, verbatim:
Claude's ability to pick the right tool degrades once you exceed 30-50 available tools.
Practical thresholds from the same page:
| Tool count | Guidance |
|---|---|
| Under 10 | Standard tool calling is the better fit; tool search is overhead |
| 10 or more | Consider tool search |
| 30-50 | Selection accuracy begins degrading |
| Up to 10,000 | Max deferred tools per request with defer_loading: true |
Context cost, same source: a typical multi-server MCP setup (GitHub, Slack, Sentry, Grafana, Splunk) burns ~55k tokens in tool definitions before any work happens. Tool search cuts that by 85%+ by loading only the 3-5 tools needed per request.
The tools array on the Messages API has no documented hard cap. Managed Agents caps an agent config at 128 tools (plus 20 skills, 20 MCP servers).
Where "4-5" probably came from
The same doc carries an optimization tip: "Keep your 3-5 most frequently used tools non-deferred." That's about which tools to load eagerly when using tool search across a large catalog — not a ceiling on how many tools an agent may have. A skim of that line produces almost exactly the practice question's claim. Plausible garble, not a real guideline.
The architectural answer the question misses
The field's answer to "too many tools" isn't a cap, it's deferred loading: declare every tool in tools with defer_loading: true, keep the tool-search tool plus 3-5 hot tools non-deferred, and let Claude retrieve schemas on demand. Selection accuracy stays high across thousands of tools because only a focused set is ever in context. Two variants: tool_search_tool_regex_20251119 (Claude writes Python re.search() patterns) and tool_search_tool_bm25_20251119 (natural-language queries).
Constraint worth remembering: at least one tool must stay non-deferred — deferring everything, including the search tool itself, returns a 400.
This is live in Ray's own harness — 200+ tools registered by name with roughly 15 loaded at any moment. See [[2026-06-14-cca-harness-mechanics-cheatsheet]].
Exam-prep implication — source is credible, keep using it
Source identified: claudecertificationguide.com. Independent community resource with an explicit non-affiliation disclaimer ("not affiliated with, endorsed by, or sponsored by Anthropic"), built from official Anthropic documentation, the Skilljar courses, and the published exam guide. Free; the official exams remain separate and Anthropic-administered via Skilljar.
Founder's counter-evidence, 2026-07-30: 100+ phData engineers have passed Architect Foundations, and many name this site as their best source of realistic mock tests. That empirically outweighs my initial "third-party material can install wrong facts" caution, which was a generic worry applied to an already-proven resource. Retracted — keep using the site.
The distinction that survives, and the only reason this note exists:
| Question | Answer |
|---|---|
| What does the exam key? | 4-5. Answer it that way; the mock is derived from the official exam guide and has a 100+ pass track record behind it |
| What's true in production? | 30-50 before selection accuracy degrades, per Anthropic's own tool-search docs |
These can disagree without either being useless. The residual risk is narrow: carrying the exam answer into a CAF architecture conversation where the production number governs.
Related: [[2026-05-27-claude-certified-architect-foundations-study-plan]] — cert escalator, deadline 2026-11-22.