> ## Documentation Index
> Fetch the complete documentation index at: https://docs.switchagents.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> These docs moved from docs.flintai.dev to docs.switchagents.ai. Use docs.switchagents.ai for every link and request.
> To search these docs from an AI tool, connect the MCP server at https://docs.switchagents.ai/mcp. The page index is at https://docs.switchagents.ai/llms.txt.

# Optimization playbooks

> Cut an agent's cost with the Switch Trust MCP server, starting from the optimizations Switch Trust recommends, and check that quality and safety held

Each cost playbook takes one kind of waste from a Switch Trust optimization to a verified fix. The last playbook fixes issues in your code without degrading evaluation results. Your coding agent reads the evidence through the Switch Trust MCP server, and you make the change in the agent's own code or configuration. To set up the server, see [Connect your coding agent](/switch-trust/mcp/connect). For a worked example, see [Optimize agents](/switch-trust/mcp/index).

**Record a baseline before running any playbook.** Ask your coding agent for the agent's latest evaluation results, and save the custom **Metric** score for task quality and the built-in evaluation results for safety. *Evaluation results hold* means both stay at that baseline after your change. See [Connect an agent and run evaluations](/switch-trust/evaluation/connect-and-run) to set up both.

<Warning>
  **Prompts, responses, and input previews can contain secrets, customer data, or other sensitive content.** Share numbers and short summaries in pull requests and tickets, not copied prompts or responses.
</Warning>

This page covers:

* [Optimizations](#optimizations)
* [Reduce context and token growth](#reduce-context-and-token-growth)
* [Reduce model cost](#reduce-model-cost)
* [Reduce unnecessary tool calls](#reduce-unnecessary-tool-calls)
* [Reduce retries and repeated calls](#reduce-retries-and-repeated-calls)
* [Fix findings without degrading evaluation results](#fix-findings-without-degrading-evaluation-results)

*A session groups the turns from one conversation or task. A turn, also called an interaction, is one prompt to the model behind your agent and the response to it. An issue is a problem Switch Trust found in your code, grouped by rule and severity, and each place it occurs is a finding.*

## Optimizations

Switch ROI, a Switch Trust capability, analyzes each agent's recent traffic for cost waste and recommends optimizations for it. Each optimization is a short write-up for that agent, with a confidence level and, usually, an estimated savings in dollars over 30 days. On the **Improvements** tab, Verbosity shows Yes or No for whether the agent is a candidate, in place of a dollar figure. Weigh the confidence level before acting on a dollar figure. For context and retry optimizations, it's lower when few sessions back the finding.

Switch ROI re-runs each analysis about once a day, over the agent's last 90 days of traffic. Traffic from before your change stays in that window for a while, so a later result reflects the fix gradually. When the waste is gone, the optimization's write-up says so and it shows no savings.

In Switch Trust, an agent's optimizations appear on its **Improvements** tab. Switch ROI recommends each optimization, and you decide whether to apply it. Costs and savings are estimates, based on each model's published price or the rate your organization has set for it.

| Shown in Switch Trust as | Analysis type | What it looks for |
| - | - | - |
| **Context cache waste** | `context_cache_waste` | Context that keeps growing, and how often the prompt is served from cache |
| **Verbosity** | `verbosity` | Responses longer than the task needs |
| **Model optimizations** | `model_optimizer` | Whether a cheaper model could do the same job |
| **Tool bloat** | `tool_bloat` | Tools the agent is offered but doesn't use |
| **Retry and Loop waste** | `retry_loop_waste` | Repeated tool calls and duplicate prompts within a session |

*Tools: `get_latest_roi_analyses` for an agent's current optimizations, `list_roi_analyses` for past results, `get_roi_analysis` for one in full, and `list_roi_interactions` for the traffic behind them.*

## Reduce context and token growth

**Start from the agent's context and verbosity optimizations.**

```text theme={null}
For support-triage, show the latest context_cache_waste and verbosity results, with the estimated savings.
```

| The optimization flags | What to do |
| - | - |
| Context that keeps growing and is never trimmed | Summarize or truncate history |
| A low cache hit rate | Check whether the prompt structure is defeating the cache |
| Frequent cache invalidation | Keep the stable part of the prompt first, and don't insert changing content before it |
| Tool definitions that change between turns | Keep tool schemas stable, so the cached prefix survives |
| Large outputs | Review responses for verbosity |

To see the pattern in a real session, ask for the input, output, and cached tokens of each turn in the agent's most expensive session. Input tokens that grow every turn, or cached tokens near zero, are the same patterns at turn level.

**Verify:** tokens and cost per session fall for comparable work, and evaluation results hold.

*Tools: `get_latest_roi_analyses`, `list_llm_sessions`, `get_llm_session`, `get_interaction`.*

## Reduce model cost

**Start from the agent's model optimization, then check it against real traffic.**

```text theme={null}
For support-triage, show the latest model_optimizer result, then compare it with the model and cost of each recent interaction.
```

**Fix:** use the cheaper model the optimization names for the work it identifies.

**Verify:** cost per session falls, and evaluation results hold.

*Tools: `get_latest_roi_analyses`, `list_roi_interactions`.*

## Reduce unnecessary tool calls

**Start from the agent's tool optimization, then compare the tools offered with the tools called.**

```text theme={null}
For support-triage, show the latest tool_bloat result, then compare the tool definitions with the tool calls in recent interactions.
```

**Fix:** remove the unused tool definitions the optimization lists. The analysis covers a window of recent traffic, so check that a tool isn't needed for rarer tasks before you remove it.

**Verify:** tokens per session fall, because fewer tool definitions go into every prompt, and evaluation results hold.

*Tools: `get_latest_roi_analyses`, `list_roi_interactions`.*

## Reduce retries and repeated calls

**Start from the agent's retry optimization, then find the sessions it describes.**

```text theme={null}
For support-triage, show the latest retry_loop_waste result. Then list the 10 sessions with the most turns, sorted by turns in descending order, and show the input preview for each turn in the one with the most turns.
```

| The optimization flags | What to do |
| - | - |
| The same tool called again and again | Cache the result, or check state before calling the tool again |
| Duplicate or near-duplicate prompts | Find out why the first attempt wasn't enough |

**Verify:** turns and duration per session fall, and evaluation results hold.

*Tools: `get_latest_roi_analyses`, `list_llm_sessions` sorted by `turns` or `duration_ms`, `get_llm_session`, `get_interaction`.*

## Fix findings without degrading evaluation results

**Ask for the issues on a file before you change it, with the fix for each.**

```text theme={null}
Show the Switch Trust issues for src/agent/tools.py, with the evidence and remediation guidance for each.
```

Apply the remediation in your code. Your coding agent can do this directly from the guidance.

**Verify:** the finding clears the next time Switch Trust scans your code, and evaluation results hold. Switch Trust scans your code from a GitHub Action or GitLab component in your CI pipeline. See [How discovery works](/switch-trust/discovery/how-discovery-works).

*If the file check returns `complete: false`, it stopped before checking every issue, so treat the result as unknown, not clean. Ask your coding agent to raise `max_issues`, up to 500, or narrow the search. Tools: `find_issues_for_file`, `list_issues`, `get_issue`, `get_finding_detail`, `get_remediation_guidance`.*

## Next steps

<CardGroup cols={2}>
  <Card title="Optimize agents" icon="gauge-high" href="/switch-trust/mcp/index">
    Walk through the context and token playbook end to end
  </Card>

  <Card title="MCP tools" icon="list" href="/switch-trust/mcp/tools">
    Every tool the server offers and what it returns
  </Card>
</CardGroup>
