October 2, 2026 · 8 min read
Engineering

When an MCP server has too many tools: 171,665 tokens a turn, and how on-demand discovery cut it to 4,800

TL;DR

An MCP server with too many tools quietly costs every client the whole tool list on every turn: a fully connected key on ours was 171,665 input tokens before the model did anything. Listing tools on demand — search, describe, call — cut that to about 4,800 with the same tool choices on three models.

Where the tokens go

The Model Context Protocol's tools/list returns every tool's name, description and input schema, and an MCP host puts all of it into the model's context. That is fine at twenty tools. Our server exposes the platform tools of every connected ad, analytics, commerce and CRM platform — about nine hundred at the time of the test, roughly 359 KB of JSON — and a fully connected key paid 171,665 input tokens on every single request before the first word of the task.

Trimming descriptions only bends the curve. Every new platform adds dozens of tools, so a fixed list grows with the product while the useful part of it for any one task stays at three or four tools.

The design: three tools instead of nine hundred

A platform tool stays callable by its own name, and a guessed name that does not exist is answered with the closest real tools rather than an error. A key lists on demand automatically once it can reach more than 100 tools or 64 KB of definitions; it can be pinned to the full list or to discovery.

  1. 1search_tools — a ranked search over the tools this key can reach: BM25 over names, families, descriptions and argument names, with verb synonyms, and delete / remove / archive tools pushed down unless the query asks for them. Free to call.
  2. 2describe_tool — the full documentation and input schema of one tool.
  3. 3call_tool — runs it, through exactly the same path as a direct call: the same gates, approval holds, scopes and credit checks.

The A/B

Twenty real tasks against the real catalogue, with tool execution mocked so only the choice was measured. With on-demand discovery, Claude Sonnet 5 picked the right tool 20 times out of 20, GPT-6 Luna 20/20 and Claude Haiku 4.5 19/20 — as accurate as with the full list — at about 4,800 input tokens a turn.

Two ranking changes came out of the misses rather than the plan: a separate class for report queries (spend, impressions, CTR) next to plain reads, which moved "Snapchat spend last 7 days" from outside the top six to first; and indexing the platform's own tool description as search-only text.

The same problem inside an agent loop

Our in-app autonomous loop has a hard cap of 120 registered tools. A six-platform run registers about 319, and the original fix — trimming to the cap — made 7 of 12 test targets unreachable on every model. Keeping the trimmed tools reachable through the same search (find a tool, then call it) reached 11 or 12 of 12 on every model tested, and used fewer tokens.

If you build an MCP server: measure your tools/list in tokens, not tools. Every host pays it on every turn.

Frequently asked questions

Why does my MCP server use so many tokens?
MCP hosts load the full tools/list — every name, description and input schema — into the model's context on every turn. A server with hundreds of tools can cost over 100,000 input tokens per turn; ours was 171,665 for a fully connected key.
How do you handle hundreds of tools in an MCP server?
List them on demand: expose a search tool, a describe tool and a call tool, and keep the real tools callable by name. In our A/B this cut a turn from 171,665 to about 4,800 input tokens with the same tool choices on Claude Sonnet 5, GPT-6 Luna and Claude Haiku 4.5.
Does tool search make an AI agent pick the wrong tool?
Not in our measurement — 20/20, 20/20 and 19/20 on three models over twenty tasks, matching the full list — provided the search ranks well: synonyms, a report class apart from plain reads, and destructive tools ranked down unless asked for.
Run your ad ops with an agent you can trust

Start free — 2,500 credits a month, no credit card. Every write waits for your approval.

Keep reading