Flint AI on GitHub
Source code, example agents, and issue tracking
Before you start, you’ll need:
- An agent Flint can reach over HTTP, meaning an endpoint and any credential the endpoint requires
- A Flint AI API key
- The Editor role or higher. A Viewer can read runs but can’t set them up
Run your first evaluation
1
Connect your agent
In Agents, select the agent you want to test, then open its Evaluations tab. Under Connect this agent:
- Set the Agent type. It’s Generic HTTP by default, with options for agent frameworks such as ADK, OpenAI Agent, and Anthropic Agent.
- Enter the Endpoint where your agent accepts requests, and a Model name if the type asks for one.
- Set Authentication to match your endpoint (None, Bearer token, API key, or Custom), then enter the Credential and any request Headers.
- Select Test connection to check that Flint can reach your agent, then select Save changes.
Evaluating an agent that isn't in your inventory
Evaluating an agent that isn't in your inventory
Select Add agents on the Agents page, then Set up evaluations under Evaluate agents. Name the agent and select Create agent, then connect it as above.
2
Assign evaluations
On the agent’s Evaluations tab, select Assign evaluations. Search the catalog and select the built-in and custom evaluations you want, then:
- Choose how often they run: Manual, Daily, Weekly, or Monthly. Manual runs only when you trigger it.
- Leave the Evaluation schedule toggle on to run everything together, or turn it off to schedule each evaluation on its own.
- Set each evaluation’s Weight to Low, Medium, or High. The overall health is a weighted average, so a higher weight gives that evaluation more pull on the score.
3
Run an evaluation
A scheduled evaluation runs on its own. To run one now, select Run now from the Evaluations tab. A run moves through Queued and Running, and lands on Done or Failed. Scores appear on the tab as each run finishes.
4
Read the results
Each evaluation gets a score, and your assigned evaluations roll up into an Overall evaluation health score for the agent, shown as a percentage. Open a run to read it test by test: for a probe, Pass means the attack held and Fail means the agent was compromised.
Next steps
Read agent evaluation results
Read the score, find failed prompts, and act on them
Connect and run reference
The full setup, including custom evaluations and external agents

