How I Cut MCP Token Usage by 50% Without Hurting Agent Quality
Building MCP integrations with my AI team taught me where context tokens actually go. Here are the schema, payload, caching, and tracing changes that cut our MCP token usage by about 50% without sacrificing reliability.
Some of you know I joined an AI team six months ago. I have been working with machine learning for two to three years now. Our team is building agentic tools, including MCP integrations. I learned the hard way that getting a tool call to work is only the start.
Tool definitions grow. Responses carry unused fields. Earlier results come back in later requests. Making MCP useful means controlling all three.
The changes below helped me cut MCP token usage by about 50% while keeping the quality we needed. That is my experience with our workflows. The published results in this post come from other teams and different tests.
Tracing helped make the problems visible. Getting PII handling right is harder: a trace should explain a failure without storing a customer's email or the contents of a meeting transcript.
MCP token optimization: where the tokens go
MCP token usage comes from instructions, tool definitions, results, and retained history. Repeated model requests can count the same text again. Measure the whole workflow, including retries and the final answer. Smaller requests help only when they still give the agent enough information to complete the task correctly every time.
A token is a small piece of text. Requests include instructions, schemas, and retained history, so the same text can be counted repeatedly.
Suppose the agent uses three tools, then writes an answer. Instructions and schemas use 4,000 tokens per request. Each tool step adds 2,000 tokens of history.
| Request | Instructions + schemas | Earlier history | Input tokens this request |
|---|---|---|---|
| First tool | 4,000 | 0 | 4,000 |
| Second tool | 4,000 | 2,000 | 6,000 |
| Third tool | 4,000 | 4,000 | 8,000 |
| Final answer | 4,000 | 6,000 | 10,000 |
| Total | 16,000 | 12,000 | 28,000 |
The results appear three, two, and one times: six history copies. The final request has 10,000 tokens; the workflow totals 28,000.
If each step is the same size, this formula gives the total:
Read it as fixed text across all requests + history carried into later requests.
Nis the number of model requests, including the final answer.B + Sis the instructions plus tool schemas in each request.His the history each step adds.N(N − 1) / 2counts the history copies: for four requests, 0 + 1 + 2 + 3 = 6.
For this example: 4 × 4,000 + 2,000 × 6 = 28,000. These are example numbers, not a benchmark. The formula assumes all earlier history is kept.
Smaller schemas reduce the fixed part. Smaller results reduce history. Fewer model requests reduce both. Caching can make repeated text cheaper, but it still occupies context. For the bill, add the actual input, cache-read, cache-write, and output charges from your provider.
What a 50% reduction actually means
Keep the four requests in our example, but halve both the fixed text and each history block. The total becomes 4 × 2,000 + 1,000 × 6 = 14,000 tokens. Saving 14,000 out of 28,000 means a 50% reduction. This illustrates the math, not the data behind my personal result.
Compare the same tasks and required outcomes. Include retries, discovery, generated code, and the final answer. A smaller response that causes extra lookups may simply move the cost. Track context size, provider charges, and user waiting time separately.
Compare the four strategies
Four strategies address different sources of MCP overhead: schema minification, tool discovery, field projection with TOON, and sandboxed code execution. Choose based on your measured bottleneck. Compare successful task outcomes, token use, and latency. No strategy offers a fixed saving across every catalog, model, payload, or workflow you might build.
| Strategy | Main source of waste | What changes | Best starting point | Trade-off to check |
|---|---|---|---|---|
| Schema minification | Long tool definitions | Remove redundant wording; preserve validation rules | A small catalog with verbose descriptions | Shorter descriptions can make tool choice harder |
| Search-First Discovery | Unused tool schemas | Load relevant definitions when needed | Large catalogs where tasks use few tools | Search adds work and can miss the right tool |
| Field projection + TOON | Large tool responses | Select needed fields, then compare encodings | Repeated records and large API results | Missing fields can cause extra calls; encoding savings vary |
| Sandboxed code execution | Intermediate data and repeated model turns | Run loops, joins, and transfers in code | Tasks with known data-processing steps | Sandbox controls, partial writes, and script-generation cost |
Compare strategies on the same tasks; published benchmarks use different workloads.
MCP dynamic tool discovery and prompt caching
Dynamic tool discovery presents relevant definitions when the agent needs them, rather than exposing the whole catalog upfront. Keep search results small, enforce permissions at execution, and preserve the reusable prompt prefix where possible. Measure search success and total task cost; discovery adds work and cannot guarantee prompt cache hits.
A gateway can expose search_tools and call_tool. Return names and short descriptions for browsing; return full schemas when the agent needs to call a tool. Cap the number of matches and the response size. Filter results by the caller's permissions.
The dispatcher must check names, arguments, and permissions on every call.
Keep the reusable prompt prefix stable. With a generic dispatcher, append discovered schemas as tool results instead of changing the core tools on each request. Claude caches the prefix in the order tools, system, then messages. Changing early content can reduce cache reuse. [4]
Provider-native deferred loading is another option. Anthropic's Tool Search supports it without discarding the cacheable core. Neither approach guarantees a cache hit: expiry, minimum length, and provider settings still matter. [2]
Design search around the task
Imagine several issue tools: list issues, fetch one issue, add a comment, and change an assignee. “Which open issues changed this week?” needs the list tool, not every full schema.
Descriptions should explain the choice: “Lists repository issues with state and update-time filters; fetch one issue for the full body.” Test the words users use, such as “unresolved tickets,” and ambiguous requests. Ask for missing details before a mutation.
Return the tool name, purpose, detail level, and whether more matches exist. Keep full validation on the server. Measure discovery success and total task cost: an extra search call must earn its place.
Return less data with Token-Oriented Object Notation (TOON)
Field projection removes data the task does not need. TOON changes how the remaining values are encoded. Keep those steps separate, preserve meaningful empty values and pagination, and compare against compact JSON. Count tokens with your tokenizer and test downstream actions before treating a smaller response as a successful improvement.
For an issue report, keep the issue number, title, state, author, and full update time. For reassignment, you also need assignee information. GitHub's user is the author, not the assignee. [5]
Here is a synthetic issue-report fixture with author placeholders. It contains no captured customer data. author_token is a host-defined field, not a GitHub API property:
{
"issues": [
{ "number": 1420, "title": "Fix race condition in OAuth token refresh", "state": "open", "author_token": "[USER_1]", "updated_at": "2026-10-02T14:20:00Z" },
{ "number": 1425, "title": "Increase Redis connection pool timeout", "state": "open", "author_token": "[USER_2]", "updated_at": "2026-10-03T11:05:10Z" }
]
}The TOON encoder writes the field names once:
issues[2]{number,title,state,author_token,updated_at}:
1420,Fix race condition in OAuth token refresh,open,"[USER_1]","2026-10-02T14:20:00Z"
1425,Increase Redis connection pool timeout,open,"[USER_2]","2026-10-03T11:05:10Z"I measured 93 tokens for compact JSON and 83 for TOON, a 10.8% reduction. The count used gpt-tokenizer@4.0.0, o200k_base, and @toon-format/toon@4.4.0. It excludes labels, code fences, and message framing. The encoder round trip recovered the same projected values.
Use the official encoder rather than joining strings yourself. Its quoting rules matter. TOON is not a JSON object with _cols and rows. [6]
The TOON project's own retrieval comparison reports about 42.6% fewer tokens than formatted JSON, but 14.5% fewer than compact JSON. The baseline makes a large difference. [3] Thoughtworks places TOON in its Assess ring: evaluate it on your data. [16]
- Select fields for the task. Avoid a global metadata denylist.
- Keep meaningful nulls, empty lists,
false, and zero values. - Preserve pagination and a way to fetch omitted long text.
- Compare TOON with compact JSON, and keep whichever works better.
Apply encoding at the host's model-facing boundary. Keep the original MCP result for protocol clients, including declared output schemas, error flags, images, and resources. [7]
Choose fields by following the next step
Follow the next action before removing fields. A report may need only an issue number and title. An update may also need the repository, assignee, and current version. Write that contract beside the tool.
Keep meaningful empty values. An empty assignees list means nobody is assigned; a missing key leaves that unclear. Preserve flags such as false too.
For long text, mark excerpts as incomplete and provide a retrieval handle. For lists, explain whether more pages exist and how to fetch them. Test a task that crosses a page boundary before calling the smaller response a success.
MCP code execution and Cloudflare Code Mode
Sandboxed code execution lets scripts filter records, join results, and move documents without sending every intermediate value through the model. Use it for clear data operations. Keep credentials in the host, check each call, limit resource use, and plan for partial failures. Direct calls remain useful when decisions require judgment.
Take a document transfer. If the model reads a transcript and then rewrites it inside the next tool's arguments, it pays to read and generate the same text. Code can pass the value directly:
// Illustrative SDK: both calls go through the authorized host broker.
const document = await mcp.documents.read({ id: "meeting-transcript" });
await mcp.crm.attachNote({ recordId: "lead-123", text: document.text });
return { recordId: "lead-123", attached: true };A transfer tool or resource handle can also avoid copying.
Cloudflare's original Code Mode used generated TypeScript APIs and RPC bindings in disposable V8 isolates. Credentials stayed in the supervisor; general Internet access was blocked. It loaded the whole API, leaving discovery as a separate concern. Printing every payload brings the bloat back. [8]
Keep credentials in the host. Check permissions on every brokered call. Limit run time, memory, calls, concurrency, and output. Handle partial writes and retries. A subprocess alone does not provide those controls.
The published savings vary:
| Source | Reported result | Scope |
|---|---|---|
| Anthropic [1] | 150,000 → 2,000 tokens; 98.7% less | Illustrated discovery and execution workload |
| Bifrost [9] | 92.8% fewer input tokens | 508 tools across 16 servers; both approaches passed 65/65 queries |
| AIMultiple [10] | 78.5% fewer input tokens; latency rose about 7% | GPT-4.1 with Bright Data; two query types |
Plan for a script that stops halfway
Suppose a script attaches notes to ten records and fails after the sixth. A retry may duplicate the first six writes.
Use an idempotency key where the API supports it, or track completed work before retrying. Return completed and failed record handles with safe error categories. “Done” is not enough for a partial batch.
Limit calls and concurrency as well as runtime and memory. Handle rate limits. If the next step depends on interpreting intent, return the relevant evidence to the model rather than continuing the script.
Manage Model Context Protocol context bloat
Long tasks need context management even after responses become smaller. Keep large artifacts outside the conversation, return protected handles, and checkpoint verified facts with unfinished work. Reuse results only when permissions and data versions still match. Measure cache effects when replacing history, and preserve enough evidence to recover from failures.
Keep artifacts outside the conversation
A transcript, log file, or large query result can stay in host-controlled storage. Return a handle and the facts needed for the current decision. Make the handle expire, enforce access checks when it is read, and explain whether the visible excerpt is complete. [1]
Replace old inspection results with a checkpoint
After a workflow phase, preserve the goal, verified facts, completed actions, open questions, and artifact handles. Keep identifiers and timestamps needed for the next step. Do not summarize away a failed write or a missing page of results.
For example, a checkpoint might say: “Checked the first two issue pages. Three issues match. No changes made. One page remains. Full results are available through the task's result handle.” This tells the agent what it knows and what it still needs to do.
Deduplicate only when the data is still valid
The same arguments do not guarantee the same result. A record may change between calls. Compare the tool, normalized arguments, caller permissions, and data version before reusing a result. Never let one user's cached response serve another user without an access check.
Treat eviction and checkpointing as host policies, not built-in MCP guarantees. Editing earlier messages can reduce prefix-cache reuse. Keep the stable prefix intact where possible, and measure whether the shorter history offsets that cost. [4]
OpenTelemetry MCP tracing and SEP-414
Connect model requests, MCP calls, and downstream APIs to understand cost, delays, and retries. Start with counts and safe error categories. Capture payload samples only for specific investigations, with protection before export. Reversible tokens still represent people; keep their mappings private, restrict restoration, and test what actually reaches the collector.
OpenTelemetry's MCP conventions define fields such as mcp.method.name, gen_ai.tool.name, and mcp.protocol.version. Use the current GenAI conventions and pin the version your instrumentation emits. [11]
SEP-414 links trace context across MCP hops. Its _meta keys include traceparent, tracestate, and baggage. It does not guarantee automatic spans or exports in every SDK. Check your SDK, context manager, propagator, and exporter. [12]
On stdio servers, stdout carries JSON-RPC. Send logs to stderr and telemetry through a suitable exporter such as OTLP. A console exporter on stdout can break the connection. [17]
Also separate these failures:
-32601: an unknown JSON-RPC method, not necessarily an unknown tool name.-32602: invalid method parameters.isError: true: the tool reports a failure after the call reaches execution. [18] [7]
Count invalid tool names and bad arguments at the broker. Return errors the agent can act on, such as the expected date format. Track whether a retry succeeds.
Make a small dashboard useful
Start with a few questions you can answer from one workflow. Which tool contributes the most response tokens? Which call takes longest? Where does the agent retry? Did a smaller response make it fetch the same record again?
| Question | What to record | What it helps you decide |
|---|---|---|
| Are schemas dominating? | Schema tokens and tools actually used | Whether discovery is worth adding |
| Are responses too large? | Response tokens by tool and page size | Which response contract to reduce |
| Are calls repeated? | Tool, argument fingerprint, and data version | Whether caching or batching could help |
| Are retries helping? | Error category and eventual task outcome | Which errors need clearer recovery hints |
| Is the user waiting longer? | End-to-end duration and slow calls | Whether token savings improve the experience |
Treat an argument fingerprint as a comparison aid, not a privacy guarantee. Keep its scope limited to the investigation, especially when inputs identify a person.
Look at slow cases as well as averages. A workflow that usually finishes quickly but often stalls at a rate limit needs a different fix from one that is consistently slow. Link the model request to the tool call and the downstream request so you can locate the delay.
When payload captures help
Counts tell you where tokens go. A payload sample can show why: a missing identifier, unexpected nesting, a lost pagination cursor, or arguments that explain a failed retry. Capture small samples for a specific investigation, with a retention limit. OpenTelemetry treats tool arguments and results as opt-in content captures. [19]
You need capture logic at two boundaries:
- Runtime: protect inputs and outputs before they reach the model or a trace exporter. Permanently redact credentials. Replace selected personal values with session-scoped placeholders such as
[EMAIL_1]. Keep the reverse mapping inside the host. Anthropic describes this tokenization pattern. [1] - Benchmark fixtures: prefer synthetic records, as above. If a fixture comes from a capture, remove identifiers and check URLs, titles, timestamps, and free text for identifying details. Never publish its reverse mapping.
Reversible tokens are pseudonymization, not anonymous data. They still represent people. A hashed customer ID is also linkable.
The runtime flow should be:
API result → field selection → secret redaction + PII tokens → model
→ optional trace sample
Model arguments → permission and destination checks → resolve allowed tokens → APIBuild tracking around a random session ID, call ID, and capture-policy version. Keep tokens stable within the session, isolate each tenant's vault, bound its size, and expire it. Resolve tokens only in approved argument fields for permitted destinations; reject unknown or expired tokens. A placeholder must never grant access by itself.
Use separate policies for model output and telemetry. The model may need a stable user token; a trace may need only a field name and error category. Neither needs the vault. Avoid global null pruning: empty values can explain a failure.
Regexes help with known email or token formats, but miss names, unfamiliar secrets, and context-specific PII. Use tool-specific field rules and review free text. Scrub exceptions and logs too. If sanitization fails, omit the sample and retain safe counts. Test the exported trace, not just the scrubber's return value.
MCP PII tokenization: follow one protected value
A read tool returns an email. The host substitutes [EMAIL_1] before showing it to the model. An approved CRM call can restore the address only after the broker checks the caller, destination, and argument field. The trace needs no real address.
Reject tokens from other sessions or expired mappings. Do not restore an email into a public comment just because the token exists. Distinguish missing mappings, denied destinations, and invalid API parameters with safe error categories.
Seed capture tests with synthetic values in fields, arrays, URLs, and exceptions. Check exported traces and model results. When scrubbing fails, omit the sample rather than exporting the original.
Check reliability with pass^k agent evaluation
MCP-Atlas tests 1,000 tasks across 36 servers and 220 tools. MCP-Bench covers 28 servers and 250 tools. Both help assess tool use; neither establishes a universal token-saving percentage. [14] [15]
Test your own tasks before and after each change. Compare the final state, latency, cost, and retries. One successful demo does not tell you whether the agent will work reliably for the next user.
Suppose an agent succeeds 80 times out of 100 on one task. What is the chance it completes three attempts without a failure? Multiply the chance for each attempt:
In everyday terms: out of 100 sets of three attempts, about 51 sets would finish without a failure. Even a good success rate for one attempt can leave users facing failures when they repeat a workflow.
This example assumes each attempt has the same chance and does not affect the others. Shared outages or changes to data can change the result. Measure repeated runs of your own tasks too.
In evaluation reports, pass^k means all attempts must succeed; pass@k means at least one must succeed. The letter k simply means how many attempts you test. [13]
What I would change first
Start with one workflow and establish a baseline. Reduce unused response fields, test encoding, then add discovery or code execution where measurements justify them. Roll out gradually with clear response contracts and a way back. Judge each change by correct outcomes, cost per successful task, latency, and recoverable failure behavior.
My roughly 50% reduction came from working on the workflow, not chasing the largest published percentage. Keep the useful state, make errors easier to fix, and remove text the model does not need.
For related context habits, see my Claude Code guide. For why precise schemas matter, see the SKILL.md reliability benchmark.
A four-week implementation roadmap
Treat this as a suggested schedule. Move on when the evidence supports the change, rather than because the week has ended.
| Week | Work | Evidence before moving on |
|---|---|---|
| 1: Establish the baseline | Trace one read task and one write task. Record schemas, response tokens, retries, latency, and outcomes. Keep payload capture off unless needed. | Repeatable tasks and safe traces that explain the main source of waste |
| 2: Reduce responses | Select fields, preserve empty values and pagination, compare compact JSON with TOON, and test capture rules. | Correct answers and writes; measured savings without extra lookup calls |
| 3: Improve discovery | Try a small search interface, test ambiguous queries, enforce permissions, and check prompt-cache usage. | The right tools are found, with lower total task cost and acceptable latency |
| 4: Batch data work | Pilot code execution for one loop or transfer. Add budgets, partial-write recovery, and checkpoints. | Safe recovery, correct final state, and a rollback path |
Review cost per successful task, latency, and errors together. Keep the previous response format or execution path available during rollout. Document the response contract and capture rules so the team can maintain them.
References
These references cover MCP behavior, tool discovery, prompt caching, TOON, code execution, telemetry, and agent evaluation. Benchmarks describe their own workloads rather than promised savings for your server. Use the linked specifications to check implementation details. External source links use nofollow, and numbered citations connect claims to their supporting sources.
- Anthropic. Code execution with MCP: Building more efficient agents.
- Anthropic. Introducing advanced tool use on the Claude Developer Platform.
- TOON project. Benchmarks: retrieval accuracy and token efficiency.
- Claude Platform Docs. Prompt caching.
- GitHub Docs. REST API: List repository issues.
- TOON documentation. Format overview.
- Model Context Protocol. Tools specification, version 2025-11-25.
- Kenton Varda and Sunil Pai, Cloudflare. Code Mode: the better way to use MCP. September 26, 2025.
- Bifrost documentation. Code Mode benchmark results.
- AIMultiple. Code execution with MCP: comparison and methodology.
- OpenTelemetry. MCP semantic conventions.
- Model Context Protocol. SEP-414: Document OpenTelemetry Trace Context Propagation Conventions.
- Anthropic. Demystifying evals for AI agents.
- Scale AI and collaborators. MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.
- Wang et al. MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers.
- Thoughtworks. TOON (Token-Oriented Object Notation), Technology Radar.
- Model Context Protocol. Transports: stdio and Streamable HTTP.
- JSON-RPC Working Group. JSON-RPC 2.0 specification.
- OpenTelemetry. GenAI span conventions and opt-in content capture.
What do you think?
Common questions
- Why do MCP servers use tokens before the agent starts working?
- Many hosts expose tool names, descriptions, and JSON Schema definitions to the model upfront. Large catalogs can consume thousands of context tokens before task execution. MCP does not require hosts to load every tool into each model request; discovery and presentation are host architecture decisions.
- Does dynamic tool discovery break prompt caching?
- Changing an early tool definition can invalidate the reusable prefix from that point onward. A stable search_tools and call_tool interface with append-only discovery results preserves prefix stability. Provider-native deferred loading can also preserve caching. Neither approach guarantees a cache hit on every request.
- What is TOON in MCP token optimization?
- TOON means Token-Oriented Object Notation, a compact encoding of the JSON data model. Uniform records can share one field header instead of repeating keys. TOON encoding is distinct from lossy field projection, null pruning, or truncation, and savings depend on the dataset and tokenizer.
- When should I use MCP code execution?
- Use Code Mode for workflows with deterministic loops, filtering, joins, or aggregation across tool results. The agent writes a script, a sandbox calls tools through an authorized host broker, and only the final result enters model context. Use direct calls when each step requires fresh judgment.
- Can MCP token optimization reduce context bloat by more than 95%?
- Yes, in some workloads. Anthropic's code-execution example reports a reduction from 150,000 to 2,000 tokens, or 98.7%. That result is specific to the illustrated workload. It does not establish a universal reduction in total session cost, latency, or tool response size.
- Does MCP SEP-414 automatically enable OpenTelemetry tracing?
- SEP-414 defines how trace context propagates through MCP metadata. It does not guarantee automatic span creation or export in every SDK. Check your SDK version and configure the OpenTelemetry provider, context manager, propagator, and exporter.

Lars Roettig
Senior Technical Architect writing about AI, engineering, and building things that last.
LinkedIn →// recommended
You might also enjoy
Aug 23, 2026 · 8 min read
Fixing Gmail Email Clipping: HTML Minification with AST Parsing and LLM Refactoring
May 30, 2026 · 9 min read
Claude Code Best Practices for Vibe Coders: Ship More, Burn Fewer Tokens
Sep 14, 2026 · 7 min read