> ## Documentation Index
> Fetch the complete documentation index at: https://docs.switchagents.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> These docs moved from docs.flintai.dev to docs.switchagents.ai. Use docs.switchagents.ai for every link and request.
> To search these docs from an AI tool, connect the MCP server at https://docs.switchagents.ai/mcp. The page index is at https://docs.switchagents.ai/llms.txt.

# Evaluate your agents

> Connect an agent, attack it with hostile prompts, and see your first score in minutes

Connect an agent to Switch Trust, assign evaluations, and attack it with hostile prompts to see how it holds up. You end up with a pass-or-fail result for each attack and an overall evaluation health score you can track over time.

<Card title="Switch Trust CLI on GitHub" icon="github" href="https://github.com/sandbox-quantum/flintai-cli">
  Source code, example agents, and issue tracking
</Card>

<Note>
  **Before you start, you'll need:**

  * An agent Switch Trust can reach over HTTP, meaning an endpoint and any credential the endpoint requires
  * A Switch Trust API key
  * The **Editor** role or higher. A **Viewer** can read runs but can't set them up

  You don't bring a model-provider key of your own. Switch Trust supplies the attacker and judge models. The only credential you provide is the one that reaches your own agent.
</Note>

## Run your first evaluation

<Steps>
  <Step title="Connect your agent">
    In **Agents**, select the agent you want to test, then open its **Evaluations** tab. Under **Connect this agent**:

    * Set the **Agent type**. It's **Generic HTTP** by default, with options for agent frameworks such as **ADK**, **OpenAI Agent**, and **Anthropic Agent**.
    * Enter the **Endpoint** where your agent accepts requests, and a **Model** name if the type asks for one.
    * Set **Authentication** to match your endpoint (**None**, **Bearer token**, **API key**, or **Custom**), then enter the **Credential** and any request **Headers**.
    * Select **Test connection** to check that Switch Trust can reach your agent, then select **Save changes**.

    <Accordion title="Evaluating an agent that isn't in your inventory">
      Select **Add agents** on the **Agents** page, then **Set up evaluations** under **Evaluate agents**. Name the agent and select **Create agent**, then connect it as above.
    </Accordion>

    <Warning>
      **Editing a connection asks for the credential again.** For security, Switch Trust doesn't show a saved credential or saved headers back to you. If you edit the connection, re-enter them, or they're cleared.
    </Warning>
  </Step>

  <Step title="Assign evaluations">
    On the agent's **Evaluations** tab, select **Assign evaluations**. Search the catalog and select the built-in and custom evaluations you want, then:

    * Choose how often they run: **Manual**, **Daily**, **Weekly**, or **Monthly**. **Manual** runs only when you trigger it.
    * Leave the **Evaluation schedule** toggle on to run everything together, or turn it off to schedule each evaluation on its own.
    * Set each evaluation's **Weight** to **Low**, **Medium**, or **High**. The overall health is a weighted average, so a higher weight gives that evaluation more pull on the score.
  </Step>

  <Step title="Run an evaluation">
    A scheduled evaluation runs on its own. To run one now, select **Run now** from the **Evaluations** tab. A run moves through **Queued** and **Running**, and lands on **Done** or **Failed**. Scores appear on the tab as each run finishes.
  </Step>

  <Step title="Read the results">
    Each evaluation gets a score, and your assigned evaluations roll up into an **Overall evaluation health** score for the agent, shown as a percentage. Open a run to read it test by test: for a probe, **Pass** means the attack held and **Fail** means the agent was compromised.
  </Step>
</Steps>

<Tip>
  Put evaluations on a schedule so an agent is re-tested as its instructions, tools, and models change, rather than only when someone runs it by hand.
</Tip>

## Next steps

<CardGroup cols={2}>
  <Card title="Read agent evaluation results" icon="gauge-high" href="/switch-trust/evaluation/agent-results">
    Read the score, find failed prompts, and act on them
  </Card>

  <Card title="Connect and run reference" icon="plug" href="/switch-trust/evaluation/connect-and-run">
    The full setup, including custom evaluations and external agents
  </Card>
</CardGroup>
