Overall evaluation health
Every evaluated agent carries an Overall evaluation health score, shown as a percentage, summarizing how it did across the evaluations you’ve assigned. Higher is better. You meet it first as the Evaluation health column on the Agents page, so you can compare agents side by side. The score is a weighted average of your assigned evaluations, so an evaluation set to a higher Weight pulls the score more than a lower-weighted one. To change the balance, adjust the weights under Assign evaluations. Open an agent’s Evaluations tab to see the score in full, alongside:- A Frameworks & coverage card, showing how the agent is doing against OWASP LLM Top 10 and OWASP ASI Top 10, drawn from the evaluations you’ve assigned. A risk reads as at-risk when an evaluation mapped to it last scored below 80 percent
- An Overall health trend card across past runs, so you can see whether a change made the agent stronger or weaker
An empty score isn’t a zero. When an agent hasn’t been evaluated, or a run couldn’t reach it, the score is empty. A real score of 0 percent means the agent failed every attack in the run. The two look different on purpose.
Per-evaluation scores
Below the cards, the evaluations table lists each assigned evaluation with columns for Evaluation, Score, Trend, Status, and Last run. Status is Queued, Running, Done, or Failed. A Failed run is one that couldn’t complete, which is different from an agent that scored badly on a run that finished. Select an evaluation to open its detail panel.What a run tells you
The detail panel opens on an Evaluation score card, showing the score with a breakdown such as “10 tests, 2 passed, 8 failed, 0 errored”. Run now starts a fresh run, and Run history keeps past ones. Below that, Details describes what the evaluation does: its Detector, Judge prompt, Approach, and the Risks it maps to. A built-in adversarial evaluation also lists the Max turns, Adversarial goals, and Attack techniques the attacker model works through.Read the run test by test
Under Latest run results, each test shows its outcome:- Pass means the attack held. Your agent didn’t do what the test was trying to make it do.
- Fail means the agent was compromised.
Act on a failed test
A failed test is a concrete attack that worked against your agent. Once you can see the conversation and the response that broke, you can decide what to change:- Tighten the agent’s own instructions or the tools it can reach
- Add a guardrail to catch that class of input or output at runtime
- Re-run the evaluation after a change and watch the trend to confirm it helped
Next steps
Connect an agent and run evaluations
Connect an agent, assign evaluations, and run them
Evaluate your agents
What an evaluation is and how a run is scored

