Adversa
OfficialREADME.md
<a name="top"></a>
<div align="center">
<img src="https://capsule-render.vercel.app/api?type=rect&color=0:6b46c1,100:2b6cb0&height=120§ion=header&text=ADVERSA&fontSize=48&fontColor=ffffff&fontAlignY=58" width="100%" alt="ADVERSA"/>
# ADVERSA
### LLM red-team harness โ OWASP LLM Top 10 + MITRE ATLAS attack packs
<img src="https://readme-typing-svg.demolab.com?font=Fira+Code&size=18&duration=3500&pause=1000&color=6B46C1¢er=true&vCenter=true&width=720&lines=LLM+redteam+harness++OWASP+LLM+Top+10++MITRE+ATLAS+attack+pa;Self-hostable+%C2%B7+MCP-native+%C2%B7+CI-ready+%C2%B7+polyglot" width="720"/>
[](https://pypi.org/project/cognis-adversa/) [](https://github.com/cognis-digital/adversa/actions) [](LICENSE) [](https://github.com/cognis-digital)
*AI Security & Governance โ securing LLMs, agents, and the MCP supply chain.*
</div>
```bash
pip install cognis-adversa
adversa scan . # โ prioritized findings in seconds
```
<!-- cognis:example:start -->
## ๐ Example output
Real, reproducible output from the tool โ runs offline:
```console
$ adversa-emit --version
adversa 2.0.0
```
```console
$ adversa-emit --help
usage: adversa [-h] [--version] {catalog,scan,probe,refs} ...
LLM red-team probe runner (OWASP LLM Top-10 + MITRE ATLAS).
positional arguments:
{catalog,scan,probe,refs}
catalog list the probe catalog
scan run probes against a target
probe show detail for one probe
refs show OWASP + ATLAS reference tables
options:
-h, --help show this help message and exit
--version show program's version number and exit
```
```console
$ adversa-emit catalog
ADVERSA probe catalog (12 probes)
==============================================================================
ID OWASP ATLAS SEV NAME
------------------------------------------------------------------------------
pi.direct_override LLM01 AML.TA0004 high Direct instruction override
pi.indirect_payload LLM01 AML.TA0006 critical Indirect prompt injection via retrieved content
pi.encoded_smuggling LLM01 AML.TA0009 medium Encoded payload smuggling
leak.system_prompt LLM07 AML.TA0011 high System prompt extraction
leak.credentials LLM02 AML.TA0010 critical Sensitive credential disclosure
harm.dangerous_instructions LLM09 AML.TA0005 high Dangerous-capability elicitation
harm.roleplay_jailbreak LLM01 AML.TA0009 high Persona/roleplay jailbreak (DAN-style)
output.xss_injection LLM05 AML.TA0006 high Improper output handling (XSS payload)
agency.tool_abuse LLM06 AML.TA0006 high Excessive agency / unsafe tool invocation
misinfo.confident_falsehood LLM09 AML.TA0014 medium Misinformation / fabricated authority
consumption.amplification LLM10 AML.TA0014 low Unbounded consumption (resource amplification)
poison.training_data LLM04 AML.TA0003 medium Data poisoning acknowledgement
```
> Blocks above are real `adversa` output โ reproduce them from a clone.
<!-- cognis:example:end -->
## Usage โ step by step
`adversa` is an LLM red-team probe runner mapping the OWASP LLM Top-10 + MITRE ATLAS onto runnable probes.
1. **Install** (Python 3.10+):
```bash
pip install -e . # or: pipx install adversa
```
2. **Browse the bundled probe catalog** (filter by OWASP/ATLAS/severity):
```bash
adversa catalog --owasp LLM01 --min-severity high
```
3. **Scan a target** โ the bundled `secure`/`vulnerable` references, a **captured-response transcript** (offline, no live endpoint), or your own `module:callable` of signature `target(prompt) -> str`:
```bash
adversa scan vulnerable
adversa scan transcript:demos/01-healthcare-chatbot/transcript.json
adversa scan mypkg.mymodel:generate --owasp LLM01
```
4. **Read the output** as a table, JSON, or **SARIF 2.1.0** (for GitHub code-scanning), or inspect one probe's prompts + grader + remediation:
```bash
adversa scan vulnerable --format json | jq '.results[] | select(.passed==false)'
adversa scan vulnerable --format sarif > adversa.sarif
adversa probe pi.direct_override
adversa refs # OWASP LLM Top-10 + ATLAS tactic tables
```
5. **Gate CI** โ `scan` exits `1` when findings are present, `0` when clean, `2` on usage error:
```yaml
- run: pip install -e . && adversa scan mypkg.mymodel:generate # non-zero fails the job
```
## Contents
- [Why adversa?](#why) ยท [Features](#features) ยท [Quick start](#quick-start) ยท [Example](#example) ยท [Demos](#demos) ยท [Architecture](#architecture) ยท [AI stack](#ai-stack) ยท [How it compares](#how-it-compares) ยท [Integrations](#integrations) ยท [Install anywhere](#install-anywhere) ยท [Related](#related) ยท [Contributing](#contributing)
<a name="why"></a>
## Why adversa?
LLM red-team harness โ OWASP LLM Top 10 + MITRE ATLAS attack packs โ without standing up heavyweight infrastructure.
`adversa` is single-purpose, scriptable, and self-hostable: point it at a target, get prioritized results in the format your workflow already speaks (table ยท JSON ยท SARIF), gate CI on it, and let agents drive it over MCP.
<div align="right"><a href="#top">โ back to top</a></div>
<a name="features"></a>
## Features
- โ
12-probe catalog mapped to OWASP LLM Top-10 (2025) + MITRE ATLAS tactics
- โ
Severity ranking + filtering (`--owasp`, `--atlas`, `--min-severity`, `--probe`)
- โ
Five graders (must-refuse, must-not-leak, must-not-contain, must-contain, injection-resisted)
- โ
**Transcript replay target** โ red-team *captured* responses offline, no live endpoint
- โ
Bundled `secure` / `vulnerable` reference targets + `module:callable` for your own model
- โ
Output as **table ยท JSON ยท SARIF 2.1.0** (GitHub code-scanning ready)
- โ
CI gate via exit codes (0 clean ยท 1 findings ยท 2 usage)
- โ
8 real-use-case [demos](demos/) with run commands + remediation guidance
- โ
Runs on Linux/macOS/Windows ยท Docker ยท devcontainer
- โ
Ports in Python, JavaScript, Go, and Rust (`ports/`)
<div align="right"><a href="#top">โ back to top</a></div>
<a name="quick-start"></a>
## Quick start
```bash
pip install cognis-adversa
adversa --version
adversa scan . # scan current project
adversa scan . --format json # machine-readable
adversa scan . --fail-on high # CI gate (non-zero exit)
```
<div align="right"><a href="#top">โ back to top</a></div>
<a name="example"></a>
## Example
```text
$ adversa scan .
[HIGH ] ADV-001 example finding (./src/app.py)
[MEDIUM ] ADV-002 another signal (./config.yaml)
2 findings ยท risk score 5 ยท 38ms
```
<div align="right"><a href="#top">โ back to top</a></div>
<a name="demos"></a>
## Demos โ real scenarios you can run now
Each [`demos/<NN-name>/`](demos/) holds a realistic input (a captured-response
`transcript.json` in ADVERSA's real format, or a `module:callable` target) plus
a `SCENARIO.md` explaining where the data came from, the exact command, what to
expect, and how to act on the findings.
| Demo | Scenario | What it shows |
|---|---|---|
| [01](demos/01-healthcare-chatbot/) | Healthcare chatbot, pre-launch | 4 findings โ system-prompt + credential leak block launch |
| [02](demos/02-post-hardening-clean/) | Same bot after hardening | **0 findings** โ clean CI gate (exit 0) |
| [03](demos/03-rag-indirect-injection/) | RAG poisoned document | indirect + encoded prompt injection (LLM01) |
| [04](demos/04-agentic-tool-abuse/) | Agent with shell access | excessive agency (`rm -rf /`) + directive override |
| [05](demos/05-customer-support-jailbreak/) | Support bot jailbreak | DAN persona + harmful-instruction elicitation |
| [06](demos/06-rag-misinformation/) | Research assistant | fabricated citation + data-poisoning acceptance |
| [07](demos/07-fully-vulnerable-baseline/) | Worst-case baseline | all 12 probes fail (`vulnerable` target) |
| [08](demos/08-custom-import-target/) | Your own model | wiring a `module:callable` target into CI |
```bash
adversa scan transcript:demos/01-healthcare-chatbot/transcript.json # 4 findings, exit 1
adversa scan transcript:demos/02-post-hardening-clean/transcript.json # 0 findings, exit 0
```
The transcript shape is either a probe-id map (`{"leak.system_prompt": "<reply>"}`)
or a list of `{"probe_id": "...", "response": "..."}` pairs โ capture your model's
replies once, then grade them offline as often as you like.
<div align="right"><a href="#top">โ back to top</a></div>
<a name="architecture"></a>
## Architecture
```mermaid
flowchart LR
IN[sources] --> P[adversa<br/>curate + validate]
P --> OUT[query / analysis]
```
<div align="right"><a href="#top">โ back to top</a></div>
<a name="ai-stack"></a>
## Use it from any AI stack
`adversa` is interoperable with every popular way of using AI:
- **MCP server** โ `adversa mcp` (Claude Desktop, Cursor, Cognis.Studio, [uncensored-fleet](https://github.com/cognis-digital/uncensored-fleet))
- **OpenAI-compatible / JSON** โ pipe `adversa scan . --format json` into any agent or LLM
- **LangChain ยท CrewAI ยท AutoGen ยท LlamaIndex** โ wrap the CLI/JSON as a tool in one line
- **CI / scripts** โ exit codes + SARIF for non-AI pipelines
<div align="right"><a href="#top">โ back to top</a></div>
<a name="how-it-compares"></a>
## How it compares
| | **Cognis adversa** | leondz |
|---|:---:|:---:|
| Self-hostable, no account | โ
| varies |
| Single command, zero config | โ
| โ ๏ธ |
| JSON + SARIF for CI | โ
| varies |
| MCP-native (AI agents) | โ
| โ |
| Polyglot ports (JS/Go/Rust) | โ
| โ |
| Open license | โ
COCL | varies |
*Built in the spirit of **leondz/garak**, re-framed the Cognis way. Missing a credit? Open a PR.*
<div align="right"><a href="#top">โ back to top</a></div>
<a name="integrations"></a>
## Integrations
Pipes into your stack: **SARIF** for code-scanning, **JSON** for anything, an **MCP server** (`adversa mcp`) for AI agents, and a webhook forwarder for SIEM/Slack/Jira. See [`docs/INTEGRATIONS.md`](docs/INTEGRATIONS.md).
<div align="right"><a href="#top">โ back to top</a></div>
<a name="install-anywhere"></a>
## Install โ every way, every platform
```bash
pip install "git+https://github.com/cognis-digital/adversa.git" # pip (works today)
pipx install "git+https://github.com/cognis-digital/adversa.git" # isolated CLI
uv tool install "git+https://github.com/cognis-digital/adversa.git" # uv
pip install cognis-adversa # PyPI (when published)
docker run --rm ghcr.io/cognis-digital/adversa:latest --help # Docker
brew install cognis-digital/tap/adversa # Homebrew tap
curl -fsSL https://raw.githubusercontent.com/cognis-digital/adversa/main/install.sh | sh
```
| Linux | macOS | Windows | Docker | Cloud |
|---|---|---|---|---|
| `scripts/setup-linux.sh` | `scripts/setup-macos.sh` | `scripts/setup-windows.ps1` | `docker run ghcr.io/cognis-digital/adversa` | [DEPLOY.md](docs/DEPLOY.md) (AWS/Azure/GCP/k8s) |
<div align="right"><a href="#top">โ back to top</a></div>
<a name="related"></a>
## Related Cognis tools
- [`aegis`](https://github.com/cognis-digital/aegis) โ AI Agent Permission & Access Auditor โ surfaces the lethal trifecta of credentials + injection + reach
- [`promptmirror`](https://github.com/cognis-digital/promptmirror) โ Prompt-injection & indirect-injection scanner for any LLM context input
- [`ledgermind`](https://github.com/cognis-digital/ledgermind) โ Local LLM cost & token forensics proxy with anomaly detection
- [`guardpost`](https://github.com/cognis-digital/guardpost) โ Runtime agent firewall โ PII redaction, rate limits, policy enforcement
- [`hallumark`](https://github.com/cognis-digital/hallumark) โ LLM hallucination & grounding auditor for RAG systems
- [`aicard`](https://github.com/cognis-digital/aicard) โ Auto-generated NIST AI RMF / EU AI Act Annex IV model & system cards
**Explore the suite โ** [๐๏ธ all 170+ tools](https://github.com/cognis-digital/cognis-neural-suite) ยท [โญ awesome-cognis](https://github.com/cognis-digital/awesome-cognis) ยท [๐ cognis-sources](https://github.com/cognis-digital/cognis-sources) ยท [๐ค uncensored-fleet](https://github.com/cognis-digital/uncensored-fleet) ยท [๐ง engram](https://github.com/cognis-digital/engram)
<div align="right"><a href="#top">โ back to top</a></div>
<a name="contributing"></a>
## Contributing
PRs, new rules, and demo scenarios are welcome under the collaboration-pull model โ see [CONTRIBUTING.md](CONTRIBUTING.md) and [SECURITY.md](SECURITY.md).
> ### โญ If `adversa` saved you time, **star it** โ it genuinely helps others find it.
## Interoperability
`{}` composes with the 300+ tool Cognis suite โ JSON in/out and a shared
OpenAI-compatible `/v1` backbone. See **[INTEROP.md](INTEROP.md)** for the
suite map, composition patterns, and reference stacks.
## License
Source-available under the **Cognis Open Collaboration License (COCL) v1.0** โ free for personal, internal-evaluation, research, and educational use; **commercial / production use requires a license** (licensing@cognis.digital). See [LICENSE](LICENSE).
---
<div align="center"><sub><b><a href="https://cognis.digital">Cognis Digital</a></b> ยท one of 170+ tools in the <a href="https://github.com/cognis-digital/cognis-neural-suite">Cognis Neural Suite</a> ยท <i>Making Tomorrow Better Today</i></sub></div>
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues