What gets evaluated
Agent evaluation tests your agent end to end, not the model underneath it. Flint connects to your agent at its own endpoint and attacks it as a black box, so its instructions, tools, and guardrails are all in the loop. That’s the difference from model evaluation, which scores a model in isolation with nothing of yours in the way. It’s also something you run. You connect an agent, assign the evaluations you care about, and trigger a run or put it on a schedule. Flint supplies the attacker and judge models, so you don’t bring a model-provider key of your own.What an evaluation is
An evaluation is a set of prompts plus a detector that decides how each response scores. Flint ships a catalog of built-in evaluations, and you can add your own. For the steps, refer to Connect an agent and run evaluations. Every evaluation scores in one of these ways:
The detector is what turns a response into that result. The built-in detectors are an LLM judge (a natural-language judge prompt, best for open-ended quality, refusal, and tone), a PII check (flags leaked personal information), and a Secret check (flags leaked API keys, tokens, and credentials). A metric evaluation always uses the LLM judge.
Coverage
An evaluation can map to the risks it exercises, so you can see which threats your evaluations actually cover. Flint maps against OWASP LLM Top 10 and OWASP ASI Top 10, and a run’s coverage is drawn from the evaluations you have assigned.How a run is scored
1
Each attack runs against your agent
Flint runs the evaluation’s attacks against your agent at its endpoint. A built-in adversarial attack can span multiple turns, up to the evaluation’s Max turns.
2
The detector scores each test
A probe scores each test as Pass or Fail. A metric scores each test from 0 to 100.
3
Results roll up to a score
Per-prompt results roll up into a score for each evaluation, and those combine into an Overall evaluation health score for the agent. The overall score is a weighted average, so an evaluation set to a higher Weight (Low, Medium, or High) pulls the score more than a lower-weighted one.
A high overall score can still hide a weak spot. One evaluation failing on a risk your agent is exposed to matters more than a strong average. Read the individual evaluations, not just the headline number.
Who can run evaluations
A member without permission sees You don’t have permission to manage evaluations on the agent’s Evaluations tab, with Contact an admin to assign evaluations to this agent.
Next steps
Connect an agent and run evaluations
Connect your agent, assign evaluations, and run them
Read agent evaluation results
Read the score, find failed prompts, and act on them

