The MCP server lists its tools on demand: from 171,665 input tokens a turn to about 4,800, with the same choices
The MCP server now lists platform tools on demand once a key can reach more than 100. A fully connected key's whole tool list was 171,665 input tokens that every MCP host loaded on every turn.
What changed
- Three tools stand in for the platform tools:
search_tools(free, ranked, sees exactly what the key can reach),describe_tool(the full documentation) andcall_tool(every gate, hold, scope and credit check the direct call has). - A platform tool is still callable by name in either mode, and a guessed name is answered with the closest real tools.
- Per key: auto (on demand past 100 tools or 64 KB), always full, or always on demand — switchable on the API & MCP page.
The A/B behind it
Twenty real tasks against the real catalogue, with tool execution mocked: on demand, Claude Sonnet 5 chose the right tool 20/20, GPT-6 Luna 20/20 and Claude Haiku 4.5 19/20 — as accurate as with the full list — at about 4,800 input tokens a turn instead of 171,665.
The in-app autonomous loop had the mirror-image problem: a six-platform run registers about 319 tools against a cap of 120, and the trim made 7 of 12 test targets unreachable. Trimmed tools are now reachable through the same search, and 11–12 of 12 targets were reached on every model tested.
Frequently asked questions
- Why is my MCP server using so many tokens?
- Every MCP host loads the whole tools/list into the model's context on each turn. A server exposing hundreds of tools can cost over 100,000 input tokens a turn before the model does anything; listing tools on demand (search, describe, call) avoids it.
- Does on-demand tool discovery hurt tool choice?
- Not in our A/B: on twenty real tasks, Claude Sonnet 5 and GPT-6 Luna chose the right tool 20/20 with on-demand discovery, the same as with the full list, and Claude Haiku 4.5 19/20.