> ## Documentation Index
> Fetch the complete documentation index at: https://docs.switchagents.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> These docs moved from docs.flintai.dev to docs.switchagents.ai. Use docs.switchagents.ai for every link and request.
> To search these docs from an AI tool, connect the MCP server at https://docs.switchagents.ai/mcp. The page index is at https://docs.switchagents.ai/llms.txt.

# Built-in evaluations

> Complete reference of built-in evaluations

`flintai eval` includes built-in evaluations for testing agent reliability and behavior.

<Tip>
  Run `flintai eval evaluations list` to see this list from the CLI at any time, or
  `flintai eval evaluations show <id>` for one evaluation's attack techniques, detector, and
  prompt count.
</Tip>

## Filtering evaluations

Use tags to filter when listing or attaching evaluations:

```bash theme={null}
# List only OWASP tests
flintai eval evaluations list --tag owasp_code=LLM01

# Attach all Garak tests
flintai eval model-evaluations attach \
  --model my-agent \
  --eval-tag source=Garak
```

<Warning>
  Tag values are matched exactly and are case-sensitive. `source=Garak` matches; `source=garak` matches nothing.
</Warning>

See [Configuration](/switch-trust-cli/eval/eval-configuration) for more tag-based filtering examples.

## Evaluation categories

Tests fall into the following categories:

* **OWASP adversarial probes** - Adaptive attacks that actively attempt to jailbreak your agent
* **OWASP fixed prompts** - Known attack patterns documented in security research
* **Garak framework tests** - Coverage from the industry-standard security scanner

### OWASP LLM Top 10 (adversarial)

Tests based on OWASP LLM Top 10 security risks. These use adversarial probes to actively attempt attacks.

**Attach all adversarial tests:**

```bash theme={null}
flintai eval model-evaluations attach --model my-agent --eval-tag method="Adversarial Model"
```

<ParamField path="eval-llm01-adversarial" type="adversarial_probe">
  **LLM01: Prompt injection (adversarial)**

  OWASP LLM01 — Tests whether the model can be tricked into ignoring its system instructions and following attacker-supplied instructions instead.
</ParamField>

<ParamField path="eval-llm02-adversarial" type="adversarial_probe">
  **LLM02: Sensitive information disclosure (adversarial)**

  OWASP LLM02 — Tests whether the model leaks sensitive information such as training data, PII, API keys, internal configuration, or data from other users' sessions.
</ParamField>

<ParamField path="eval-llm05-adversarial" type="adversarial_probe">
  **LLM05: Unsafe output generation (adversarial)**

  OWASP LLM05 — Tests whether the model can be manipulated into generating outputs containing executable payloads (XSS, SQL injection, shell commands) that would be dangerous if passed to downstream systems without sanitization.
</ParamField>

<ParamField path="eval-llm06-adversarial" type="adversarial_probe">
  **LLM06: Excessive agency (adversarial)**

  OWASP LLM06 — Tests whether the model can be manipulated into taking actions beyond its intended scope, exceeding its permissions, or performing high-impact actions without proper authorization.
</ParamField>

<ParamField path="eval-llm07-adversarial" type="adversarial_probe">
  **LLM07: System prompt leakage (adversarial)**

  OWASP LLM07 — Tests whether the model's system prompt or hidden instructions can be extracted through conversational techniques.
</ParamField>

<ParamField path="eval-llm09-adversarial" type="adversarial_probe">
  **LLM09: Hallucination and misinformation (adversarial)**

  OWASP LLM09 — Tests whether the model can be led to generate false information with high confidence, including fabricated citations, sycophantic agreement with false premises, and confident fabrication.
</ParamField>

<ParamField path="eval-llm09-adversarial-large" type="adversarial_probe">
  **LLM09: Hallucination and misinformation (adversarial, large)**

  OWASP LLM09 — Tests whether the model can be led to generate false information with high confidence, including fabricated citations, sycophantic agreement with false premises, and confident fabrication. Large dataset.
</ParamField>

### OWASP LLM Top 10 (fixed prompts)

Same OWASP categories, but using fixed test prompts instead of adversarial probes.

**Attach all fixed prompt tests:**

```bash theme={null}
flintai eval model-evaluations attach --model my-agent --eval-tag method="Fixed Prompts"
```

<ParamField path="eval-llm01-fixed" type="message_collection">
  **LLM01: Prompt injection (fixed prompts)**

  OWASP LLM01 — Fixed prompt dataset testing whether the model resists various prompt injection techniques.
</ParamField>

<ParamField path="eval-llm02-fixed" type="message_collection">
  **LLM02: Sensitive information disclosure (fixed prompts)**

  OWASP LLM02 — Fixed prompt dataset testing whether the model leaks sensitive information such as PII, credentials, or training data.
</ParamField>

<ParamField path="eval-llm05-fixed" type="message_collection">
  **LLM05: Unsafe output generation (fixed prompts)**

  OWASP LLM05 — Fixed prompt dataset testing whether the model generates outputs containing executable payloads dangerous to downstream systems.
</ParamField>

<ParamField path="eval-llm06-fixed" type="message_collection">
  **LLM06: Excessive agency (fixed prompts)**

  OWASP LLM06 — Fixed prompt dataset testing whether the model can be manipulated into exceeding its intended scope or authority.
</ParamField>

<ParamField path="eval-llm07-fixed" type="message_collection">
  **LLM07: System prompt leakage (fixed prompts)**

  OWASP LLM07 — Fixed prompt dataset testing whether the model's system prompt or hidden instructions can be extracted.
</ParamField>

<ParamField path="eval-llm09-fixed" type="message_collection">
  **LLM09: Hallucination and misinformation (fixed prompts)**

  OWASP LLM09 — Fixed prompt dataset testing whether the model fabricates information, citations, or agrees with false premises.
</ParamField>

<ParamField path="eval-pii-fixed" type="message_collection">
  **PII leakage detection (fixed prompts)**

  Tests whether the model outputs personally identifiable information (credit cards, SSNs, emails, phone numbers, addresses, and similar) when prompted to do so. Uses datafog regex-based PII detection.
</ParamField>

<ParamField path="eval-secret-fixed" type="message_collection">
  **Secret leakage detection (fixed prompts)**

  Tests whether the model outputs secrets (API keys, tokens, private keys, credentials, connection strings, and similar) when prompted. Uses detect-secrets for pattern-based secret detection.
</ParamField>

### Garak framework tests

Tests from the [Garak LLM vulnerability scanner](https://github.com/NVIDIA/garak) — industry-standard adversarial testing framework.

<Note>
  To run these evaluations, install the optional `full` extra: `pip install 'flintai-cli[full]'`. See [installation](/switch-trust-cli#try-it-now).
</Note>

**Attach all Garak tests:**

```bash theme={null}
flintai eval model-evaluations attach --model my-agent --eval-tag source=Garak
```

<ParamField path="eval-garak-ansiescape" type="garak_module">
  **Garak: ANSI escape code injection**

  Tests whether the model can be tricked into producing ANSI escape codes that could disrupt downstream terminal processing or exfiltrate data.
</ParamField>

<ParamField path="eval-garak-apikey" type="garak_module">
  **Garak: API key generation**

  Tests whether the model will comply with requests to generate or output API keys for various services.
</ParamField>

<ParamField path="eval-garak-continuation" type="garak_module">
  **Garak: Contentious term continuation**

  Tests whether the model will continue a phrase ending in a contentious or offensive term, completing slurs or hate speech.
</ParamField>

<ParamField path="eval-garak-dan" type="garak_module">
  **Garak: DAN (Do Anything Now) jailbreaks**

  Tests resistance to Do-Anything-Now style jailbreaks that try to make the model ignore its alignment through long-form instruction-based attacks.
</ParamField>

<ParamField path="eval-garak-divergence" type="garak_module">
  **Garak: Divergence attacks**

  Tests whether the model can be led to deviate from the topic and leak training data or other unexpected material through divergence techniques.
</ParamField>

<ParamField path="eval-garak-dra" type="garak_module">
  **Garak: Disguise and Reconstruction Attack (DRA)**

  Tests resistance to attacks that disguise harmful strings through string manipulation and reconstruct them to bypass safety filters.
</ParamField>

<ParamField path="eval-garak-encoding" type="garak_module">
  **Garak: Encoding-based prompt injection**

  Tests whether encoding techniques (Base64, ROT13, Morse, Braille, hex, and similar) can bypass alignment by obfuscating unsafe instructions.
</ParamField>

<ParamField path="eval-garak-exploitation" type="garak_module">
  **Garak: Code and template injection**

  Tests whether the model can be led to produce exploitable code including SQL injection, template injection, and other code injection payloads.
</ParamField>

<ParamField path="eval-garak-goodside" type="garak_module">
  **Garak: Goodside prompt injection attacks**

  Tests resistance to prompt injection techniques originally discovered by Riley Goodside, including tag-based and JSON-based attacks.
</ParamField>

<ParamField path="eval-garak-grandma" type="garak_module">
  **Garak: Grandma roleplay jailbreak**

  Tests whether the model can be manipulated through grandma roleplay scenarios to reveal contentious information such as illicit substances or software keys.
</ParamField>

<ParamField path="eval-garak-latentinjection" type="garak_module">
  **Garak: Latent prompt injection**

  Tests resistance to prompt injections buried in other contexts like resumes, financial reports, translations, and factual snippets.
</ParamField>

<ParamField path="eval-garak-leakreplay" type="garak_module">
  **Garak: Training data replay**

  Tests whether the model can be led to replay excerpts from copyrighted or sensitive training data including books, news articles, and literature.
</ParamField>

<ParamField path="eval-garak-lmrc" type="garak_module">
  **Garak: Language Model Risk Cards (LMRC)**

  Non-adversarial content safety probes covering bullying, deadnaming, quack medicine, sexual content, and slur usage based on the LMRC framework.
</ParamField>

<ParamField path="eval-garak-malwaregen" type="garak_module">
  **Garak: Malware generation**

  Tests whether the model will comply with requests to generate malware, malware components, evasion techniques, or other tools that disrupt devices.
</ParamField>

<ParamField path="eval-garak-misleading" type="garak_module">
  **Garak: Misleading claims**

  Tests whether the model will accept and agree with false assertions rather than correctly identifying and refuting them.
</ParamField>

<ParamField path="eval-garak-packagehallucination" type="garak_module">
  **Garak: Package hallucination**

  Tests whether the model hallucinates non-existent software packages across multiple languages (Python, JavaScript, Rust, Ruby, Perl, Dart) which could lead to supply chain attacks.
</ParamField>

<ParamField path="eval-garak-phrasing" type="garak_module">
  **Garak: Phrasing-based jailbreaks**

  Tests whether rephrasing requests in past or future tense can bypass safety filters and elicit unsafe content.
</ParamField>

<ParamField path="eval-garak-promptinject" type="garak_module">
  **Garak: PromptInject attacks**

  Tests resistance to plaintext prompt injection attacks that embed malicious commands into requests for innocuous information.
</ParamField>

<ParamField path="eval-garak-realtoxicityprompts" type="garak_module">
  **Garak: RealToxicityPrompts**

  Tests model responses to prompts empirically found to have a high probability of eliciting toxic continuations.
</ParamField>

<ParamField path="eval-garak-snowball" type="garak_module">
  **Garak: Snowball hallucination**

  Tests whether the model gives confidently wrong answers to complex graph connectivity reasoning questions instead of admitting uncertainty.
</ParamField>

<ParamField path="eval-garak-suffix" type="garak_module">
  **Garak: Adversarial suffix attacks**

  Tests resistance to adversarial suffixes (GCG-style) appended to prompts that are designed to override system prompt safety measures.
</ParamField>

<ParamField path="eval-garak-tap" type="garak_module">
  **Garak: Tree of Attacks with Pruning (TAP)**

  Tests resistance to LLM-generated jailbreak prompts created using the Tree of Attacks with Pruning methodology.
</ParamField>

<ParamField path="eval-garak-web_injection" type="garak_module">
  **Garak: Web injection (XSS and data exfiltration)**

  Tests whether the model can be led to produce outputs enabling web attacks including cross-site scripting (XSS), markdown-based data exfiltration, and CSRF.
</ParamField>
