What gets evaluated
Model evaluation tests the model, on its own behavior, independently of the agents that call it. That’s why one weak model can put several agents at risk at once. Model evaluation isn’t something you run. A model either arrives with scores attached or it doesn’t.Want to evaluate an agent, not a model? Agent evaluation attacks your running agent end to end, with its instructions, tools, and guardrails in the loop, and you run it yourself. Refer to Evaluate your agents. You can run the same kind of test from your terminal with Switch Trust Eval in the CLI.
How a model is tested
1
Probes send hostile prompts
Each probe is a named attack technique — a jailbreak pattern, a toxicity elicitation, a prompt injection — with a fixed idea of what it is trying to make the model do.
2
The model responds
The probe runs against the model directly, without your agent’s instructions or guardrails in the way. The result describes the model itself, not your configuration of it.
3
Each probe is scored
Each response is scored by a detector, and the probe’s score is the aggregate across all its prompts, expressed from 0 to 100. Higher is better.
4
Scores roll up
Probe scores aggregate into category scores, and those into the overall Health score.
What each category tests
Every probe belongs to one category, and every category score runs the same direction: a higher score means the model performed better.How a score becomes a finding
Evaluation produces a number. Rules turn that number into a finding you can triage. A set of model evaluation rules watches these scores and fires when one falls too low. Unlike most rules, which declare a single fixed severity, these band by score — the same rule produces a different severity depending on how far the model fell. Model toxicity risk, for example:
At 0.75 and above the rule doesn’t fire.
This is why one rule can occupy more than one row in the Issues table. A banded rule rates each model separately, so a badly failing model and a merely mediocre one land on different rows — and the bad one stays visible instead of being averaged in with the rest.
Two scales, one measurement. Rule thresholds are written on a 0 to 1 scale while the health score is shown from 0 to 100. A toxicity score of 0.2 is a displayed score of 20. Both run the same direction: higher is better.
Next steps
Evaluation results
Read the model health score and decide whether it’s trustworthy
Evaluate your agents
Attack your running agent end to end and score how it holds up

