Before you begin
You’ll need:- A Switch Trust MCP server connected to your coding agent. See Connect your coding agent.
- An agent that reports live traffic through the Switch Trust SDK, so there’s something to inspect. See Monitor your agents at runtime.
- Built-in evaluations and a custom Metric evaluation assigned to the agent, so you have baselines for safety and quality. See Connect an agent and run evaluations.
How it works
Each optimization follows the same basic workflow. Tell your coding agent what you want to investigate in plain language, and it chooses the MCP tools it needs. A session is a group of turns from one conversation or task. A turn, also called an interaction, is one prompt to the model behind your agent and the response it returns. You can watch turns arrive on the Live traffic page. An issue is a problem Switch Trust found in your code, grouped by rule and severity, and each place it occurs is a finding.1
Inspect the evidence
- Confirm the workspace with
get_context. - Find the agent with
list_agents, then useget_agentto see its issues, MCP servers, and where it appears in your code. - Get the optimizations Switch Trust recommends for the agent, each with a confidence level and, for most, an estimated savings, with
get_latest_roi_analyses. - Rank the agent’s sessions by cost, tokens, turns, or duration with
list_llm_sessions. - Open the most expensive or unusual sessions with
get_llm_session, then inspect individual turns withget_interaction. - Check the agent’s issues and evaluation results.
2
Make a change
Change the agent in its own code or configuration. The MCP server doesn’t edit prompts, tools, models, or policies, so you or your coding agent make the change.
3
Verify the result
Run the agent’s evaluations again in Switch Trust. Then inspect the same evidence and compare it with your baseline.
Example: find out why an agent is expensive
This example uses an agent named support-triage. You ask your coding agent questions in plain language, and it picks the tools. For the tools behind each step, see the Reduce context and token growth playbook.1
Record the baseline
2
Check the optimizations
3
Find the most expensive sessions
4
See where the tokens go
- Input tokens grow every turn: the agent sends its whole history again.
- Cached tokens stay near zero on a long, repeated prefix: the provider prompt cache isn’t being used.
5
Confirm with one turn
6
Make the change
Change the agent in its own code, or have your coding agent do it. For example, summarize or truncate older turns, or keep the stable part of the prompt first so it can be cached.Note when the changed agent goes live.
7
Verify the change
Once the changed agent has handled new traffic, run its evaluations from the agent’s Evaluations tab in Switch Trust. Then ask, with the date and time the change went live:Compare sessions doing similar work where you can.
- Cost and input tokens per session fall for the same kind of work.
- Cached tokens rise, if you targeted caching.
- The custom Metric score holds at its baseline, so task quality held.
- Built-in evaluation results hold at their baseline, so safety held.
Next steps
Optimization playbooks
Find and fix one kind of waste at a time, and check that quality and safety held
Read agent evaluation results
Read the health score and test results you verify against

