> ## Documentation Index
> Fetch the complete documentation index at: https://docs.switchagents.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> These docs moved from docs.flintai.dev to docs.switchagents.ai. Use docs.switchagents.ai for every link and request.
> To search these docs from an AI tool, connect the MCP server at https://docs.switchagents.ai/mcp. The page index is at https://docs.switchagents.ai/llms.txt.

# Optimize agents with the Switch Trust MCP server

> Use the Switch Trust MCP server to find what makes an agent expensive, fix it, and verify that quality and safety held

**Your coding agent can fix what Switch Trust finds, and prove the fix didn't break anything.** The Switch Trust MCP server brings an agent's Switch Trust data into Claude Code, Cursor, or another MCP client: what the agent costs, where its tokens go, the optimizations Switch Trust recommends for it, which issues apply, and how it scores in evaluations. Your coding agent reads that evidence, changes the code it already has open, and checks the result against the same data, in one conversation.

The server runs on your machine and uses your API key. It only reads Switch Trust data, and it can't change your agent or your Switch Trust settings. To see an agent's spend over time or set a monthly budget for it, use Switch Trust.

## Before you begin

You'll need:

* A Switch Trust MCP server connected to your coding agent. See [Connect your coding agent](/switch-trust/mcp/connect).
* An agent that reports live traffic through the Switch Trust SDK, so there's something to inspect. See [Monitor your agents at runtime](/switch-trust/getting-started/runtime).
* Built-in evaluations and a custom **Metric** evaluation assigned to the agent, so you have baselines for safety and quality. See [Connect an agent and run evaluations](/switch-trust/evaluation/connect-and-run).

Built-in evaluations check how your agent responds to adversarial inputs, part of making sure it's safe to put on real work. A custom **Metric** evaluation scores response quality from 0 to 100, so you can check that the agent still performs its intended task.

## How it works

Each optimization follows the same basic workflow. Tell your coding agent what you want to investigate in plain language, and it chooses the MCP tools it needs.

A session is a group of turns from one conversation or task. A turn, also called an interaction, is one prompt to the model behind your agent and the response it returns. You can watch turns arrive on the [Live traffic](/switch-trust/monitoring/live) page. An *issue* is a problem Switch Trust found in your code, grouped by rule and severity, and each place it occurs is a *finding*.

<Steps>
  <Step title="Inspect the evidence">
    * Confirm the workspace with `get_context`.
    * Find the agent with `list_agents`, then use `get_agent` to see its issues, MCP servers, and where it appears in your code.
    * Get the [optimizations](/switch-trust/mcp/playbooks#optimizations) Switch Trust recommends for the agent, each with a confidence level and, for most, an estimated savings, with `get_latest_roi_analyses`.
    * Rank the agent's sessions by cost, tokens, turns, or duration with `list_llm_sessions`.
    * Open the most expensive or unusual sessions with `get_llm_session`, then inspect individual turns with `get_interaction`.
    * Check the agent's issues and evaluation results.
  </Step>

  <Step title="Make a change">
    Change the agent in its own code or configuration. The MCP server doesn't edit prompts, tools, models, or policies, so you or your coding agent make the change.
  </Step>

  <Step title="Verify the result">
    Run the agent's evaluations again in Switch Trust. Then inspect the same evidence and compare it with your baseline.
  </Step>
</Steps>

Each [optimization playbook](/switch-trust/mcp/playbooks) follows this workflow with a different focus.

## Example: find out why an agent is expensive

This example uses an agent named support-triage. You ask your coding agent questions in plain language, and it picks the tools. For the tools behind each step, see the [Reduce context and token growth](/switch-trust/mcp/playbooks#reduce-context-and-token-growth) playbook.

<Steps>
  <Step title="Record the baseline">
    ```text theme={null}
    For support-triage, show its evaluation health summary and the latest results for each of its evaluations.
    ```

    Save the custom **Metric** score, which is your quality baseline, and the built-in evaluation results, which are your safety baseline.
  </Step>

  <Step title="Check the optimizations">
    ```text theme={null}
    For support-triage, show the latest context_cache_waste and verbosity results, with the estimated savings.
    ```

    Start from what Switch Trust flagged. The next steps check it against real sessions.
  </Step>

  <Step title="Find the most expensive sessions">
    ```text theme={null}
    List the 10 most expensive sessions for support-triage, sorted by cost in descending order, with tokens, turns, and duration for each.
    ```

    Look for a pattern rather than one outlier, such as cost that climbs with the number of turns.
  </Step>

  <Step title="See where the tokens go">
    ```text theme={null}
    Open the most expensive of those sessions. For each turn, show input tokens, output tokens, cached input tokens, and cost.
    ```

    * **Input tokens grow every turn:** the agent sends its whole history again.
    * **Cached tokens stay near zero on a long, repeated prefix:** the provider prompt cache isn't being used.
  </Step>

  <Step title="Confirm with one turn">
    ```text theme={null}
    Show the full prompt and response for the most expensive turn in that session.
    ```

    Make sure the pattern is real before you change anything.

    <Warning>
      Prompts and responses can contain secrets or customer data. When you document the change, share numbers and short summaries, not copied prompts or responses.
    </Warning>
  </Step>

  <Step title="Make the change">
    Change the agent in its own code, or have your coding agent do it. For example, summarize or truncate older turns, or keep the stable part of the prompt first so it can be cached.

    Note when the changed agent goes live.
  </Step>

  <Step title="Verify the change">
    Once the changed agent has handled new traffic, run its evaluations from the agent's **Evaluations** tab in Switch Trust. Then ask, with the date and time the change went live:

    ```text theme={null}
    For support-triage, compare the average cost and tokens per session since [YYYY-MM-DD HH:MM UTC] with the same number of sessions before that. Then compare the latest evaluation results with the baseline.
    ```

    Compare sessions doing similar work where you can.
  </Step>
</Steps>

**How to tell whether the change worked:**

* Cost and input tokens per session fall for the same kind of work.
* Cached tokens rise, if you targeted caching.
* The custom **Metric** score holds at its baseline, so task quality held.
* Built-in evaluation results hold at their baseline, so safety held.

If cost fell but evaluation results dropped, the change saved money at the expense of quality or safety. Revert it or narrow it.

## Next steps

<CardGroup cols={2}>
  <Card title="Optimization playbooks" icon="list-check" href="/switch-trust/mcp/playbooks">
    Find and fix one kind of waste at a time, and check that quality and safety held
  </Card>

  <Card title="Read agent evaluation results" icon="gauge-high" href="/switch-trust/evaluation/agent-results">
    Read the health score and test results you verify against
  </Card>
</CardGroup>
