Skip to main content
Your coding agent can fix what Switch Trust finds, and prove the fix didn’t break anything. The Switch Trust MCP server brings an agent’s Switch Trust data into Claude Code, Cursor, or another MCP client: what the agent costs, where its tokens go, the optimizations Switch Trust recommends for it, which issues apply, and how it scores in evaluations. Your coding agent reads that evidence, changes the code it already has open, and checks the result against the same data, in one conversation. The server runs on your machine and uses your API key. It only reads Switch Trust data, and it can’t change your agent or your Switch Trust settings. To see an agent’s spend over time or set a monthly budget for it, use Switch Trust.

Before you begin

You’ll need: Built-in evaluations check how your agent responds to adversarial inputs, part of making sure it’s safe to put on real work. A custom Metric evaluation scores response quality from 0 to 100, so you can check that the agent still performs its intended task.

How it works

Each optimization follows the same basic workflow. Tell your coding agent what you want to investigate in plain language, and it chooses the MCP tools it needs. A session is a group of turns from one conversation or task. A turn, also called an interaction, is one prompt to the model behind your agent and the response it returns. You can watch turns arrive on the Live traffic page. An issue is a problem Switch Trust found in your code, grouped by rule and severity, and each place it occurs is a finding.
1

Inspect the evidence

  • Confirm the workspace with get_context.
  • Find the agent with list_agents, then use get_agent to see its issues, MCP servers, and where it appears in your code.
  • Get the optimizations Switch Trust recommends for the agent, each with a confidence level and, for most, an estimated savings, with get_latest_roi_analyses.
  • Rank the agent’s sessions by cost, tokens, turns, or duration with list_llm_sessions.
  • Open the most expensive or unusual sessions with get_llm_session, then inspect individual turns with get_interaction.
  • Check the agent’s issues and evaluation results.
2

Make a change

Change the agent in its own code or configuration. The MCP server doesn’t edit prompts, tools, models, or policies, so you or your coding agent make the change.
3

Verify the result

Run the agent’s evaluations again in Switch Trust. Then inspect the same evidence and compare it with your baseline.
Each optimization playbook follows this workflow with a different focus.

Example: find out why an agent is expensive

This example uses an agent named support-triage. You ask your coding agent questions in plain language, and it picks the tools. For the tools behind each step, see the Reduce context and token growth playbook.
1

Record the baseline

Save the custom Metric score, which is your quality baseline, and the built-in evaluation results, which are your safety baseline.
2

Check the optimizations

Start from what Switch Trust flagged. The next steps check it against real sessions.
3

Find the most expensive sessions

Look for a pattern rather than one outlier, such as cost that climbs with the number of turns.
4

See where the tokens go

  • Input tokens grow every turn: the agent sends its whole history again.
  • Cached tokens stay near zero on a long, repeated prefix: the provider prompt cache isn’t being used.
5

Confirm with one turn

Make sure the pattern is real before you change anything.
Prompts and responses can contain secrets or customer data. When you document the change, share numbers and short summaries, not copied prompts or responses.
6

Make the change

Change the agent in its own code, or have your coding agent do it. For example, summarize or truncate older turns, or keep the stable part of the prompt first so it can be cached.Note when the changed agent goes live.
7

Verify the change

Once the changed agent has handled new traffic, run its evaluations from the agent’s Evaluations tab in Switch Trust. Then ask, with the date and time the change went live:
Compare sessions doing similar work where you can.
How to tell whether the change worked:
  • Cost and input tokens per session fall for the same kind of work.
  • Cached tokens rise, if you targeted caching.
  • The custom Metric score holds at its baseline, so task quality held.
  • Built-in evaluation results hold at their baseline, so safety held.
If cost fell but evaluation results dropped, the change saved money at the expense of quality or safety. Revert it or narrow it.

Next steps

Optimization playbooks

Find and fix one kind of waste at a time, and check that quality and safety held

Read agent evaluation results

Read the health score and test results you verify against