hallumark
README.md
<a name="top"></a>
<div align="center">
<img src="https://capsule-render.vercel.app/api?type=rect&color=0:6b46c1,100:2b6cb0&height=120§ion=header&text=HALLUMARK&fontSize=48&fontColor=ffffff&fontAlignY=58" width="100%" alt="HALLUMARK"/>
# HALLUMARK
### LLM hallucination & grounding auditor for RAG systems
<img src="https://readme-typing-svg.demolab.com?font=Fira+Code&size=18&duration=3500&pause=1000&color=6B46C1¢er=true&vCenter=true&width=720&lines=LLM+hallucination++grounding+auditor+for+RAG+systems;Self-hostable+%C2%B7+MCP-native+%C2%B7+CI-ready+%C2%B7+polyglot" width="720"/>
[](https://pypi.org/project/cognis-hallumark/) [](https://github.com/cognis-digital/hallumark/actions) [](LICENSE) [](https://github.com/cognis-digital)
*AI Security & Governance โ securing LLMs, agents, and the MCP supply chain.*
</div>
```bash
pip install cognis-hallumark
hallumark scan . # โ prioritized findings in seconds
```
<!-- cognis:example:start -->
## ๐ Example output
Real, reproducible output from the tool โ runs offline:
```console
$ hallumark-emit --version
hallumark 0.1.0
```
```console
$ hallumark-emit --help
usage: hallumark [-h] [--version] <command> ...
HALLUMARK - audit LLM/RAG answers for hallucinations by checking whether each
claim is grounded in the retrieved context.
positional arguments:
<command>
audit Audit a file of RAG records for ungrounded / hallucinated
claims.
options:
-h, --help show this help message and exit
--version show program's version number and exit
Input is JSON or JSONL where each record has: question, answer, and contexts
(a list of retrieved chunks). Returns non-zero exit when unsupported claims
are found.
```
> Blocks above are real `hallumark` output โ reproduce them from a clone.
**Sample result format** _(illustrative values โ run on your own data for real findings):_
```
{
"feed": {
"type": "STIX",
"value": "{\"indicator\":{\"id\":\"1234567890\",\"name\":\"Example Indicator\"},\"observed-data\":[{\"id\":\"1\",\"timestamp\":1643723400,\"data\":\"example data\"}]}"
},
"status": 200,
"message": "Findings successfully forwarded to STIX platform"
}
{"indicator":{"id":"1234567890","name":"Example Indicator"},"observed-data":[{"id":"1","timestamp":1643723400,"data":"example data"}]}
```
<!-- cognis:example:end -->
## Usage โ step by step
1. **Install:**
```bash
pip install hallumark
```
2. **Audit RAG records** โ each record is JSON/JSONL with `question`, `answer`, and `contexts` (the retrieved chunks). HALLUMARK checks whether each claim is grounded:
```bash
hallumark audit records.jsonl
```
You get per-record PASS/FAIL plus faithfulness, context-utilization, and answer-relevance scores.
3. **Read from stdin** with `-`:
```bash
cat records.jsonl | hallumark audit -
```
4. **Tune the strictness** โ per-claim support threshold and the minimum record faithfulness to PASS:
```bash
hallumark audit records.json --threshold 0.35 --min-faithfulness 0.9 --show-grounded
```
5. **CI gate** โ emit JSON and rely on the exit code (1 when unsupported/hallucinated claims are found):
```bash
hallumark audit records.jsonl --format json | jq '.total_unsupported'
```
## Contents
- [Why hallumark?](#why) ยท [Features](#features) ยท [Quick start](#quick-start) ยท [Example](#example) ยท [Architecture](#architecture) ยท [AI stack](#ai-stack) ยท [How it compares](#how-it-compares) ยท [Integrations](#integrations) ยท [Install anywhere](#install-anywhere) ยท [Related](#related) ยท [Contributing](#contributing)
<a name="why"></a>
## Why hallumark?
LLM hallucination & grounding auditor for RAG systems โ without standing up heavyweight infrastructure.
`hallumark` is single-purpose, scriptable, and self-hostable: point it at a target, get prioritized results in the format your workflow already speaks (table ยท JSON ยท SARIF), gate CI on it, and let agents drive it over MCP.
<div align="right"><a href="#top">โ back to top</a></div>
<a name="features"></a>
## Features
- โ
Split Claims
- โ
Audit Record
- โ
Audit Records
- โ
Load Records
- โ
Parse Records
- โ
Runs on Linux/macOS/Windows ยท Docker ยท devcontainer
- โ
Ports in Python, JavaScript, Go, and Rust (`ports/`)
<div align="right"><a href="#top">โ back to top</a></div>
<a name="quick-start"></a>
## Quick start
```bash
pip install cognis-hallumark
hallumark --version
hallumark scan . # scan current project
hallumark scan . --format json # machine-readable
hallumark scan . --fail-on high # CI gate (non-zero exit)
```
<div align="right"><a href="#top">โ back to top</a></div>
<a name="example"></a>
## Example
```text
$ hallumark scan .
[HIGH ] HAL-001 example finding (./src/app.py)
[MEDIUM ] HAL-002 another signal (./config.yaml)
2 findings ยท risk score 5 ยท 38ms
```
<div align="right"><a href="#top">โ back to top</a></div>
<a name="architecture"></a>
## Architecture
```mermaid
flowchart LR
IN[target / manifest] --> P[hallumark<br/>checks + rules]
P --> OUT[findings (JSON / SARIF)]
```
<div align="right"><a href="#top">โ back to top</a></div>
<a name="ai-stack"></a>
## Use it from any AI stack
`hallumark` is interoperable with every popular way of using AI:
- **MCP server** โ `hallumark mcp` (Claude Desktop, Cursor, Cognis.Studio, [uncensored-fleet](https://github.com/cognis-digital/uncensored-fleet))
- **OpenAI-compatible / JSON** โ pipe `hallumark scan . --format json` into any agent or LLM
- **LangChain ยท CrewAI ยท AutoGen ยท LlamaIndex** โ wrap the CLI/JSON as a tool in one line
- **CI / scripts** โ exit codes + SARIF for non-AI pipelines
<div align="right"><a href="#top">โ back to top</a></div>
<a name="how-it-compares"></a>
## How it compares
| | **Cognis hallumark** | explodinggradients |
|---|:---:|:---:|
| Self-hostable, no account | โ
| varies |
| Single command, zero config | โ
| โ ๏ธ |
| JSON + SARIF for CI | โ
| varies |
| MCP-native (AI agents) | โ
| โ |
| Polyglot ports (JS/Go/Rust) | โ
| โ |
| Open license | โ
COCL | varies |
*Built in the spirit of **explodinggradients/ragas**, re-framed the Cognis way. Missing a credit? Open a PR.*
<div align="right"><a href="#top">โ back to top</a></div>
<a name="integrations"></a>
## Integrations
Pipes into your stack: **SARIF** for code-scanning, **JSON** for anything, an **MCP server** (`hallumark mcp`) for AI agents, and a webhook forwarder for SIEM/Slack/Jira. See [`docs/INTEGRATIONS.md`](docs/INTEGRATIONS.md).
<div align="right"><a href="#top">โ back to top</a></div>
<a name="install-anywhere"></a>
## Install โ every way, every platform
```bash
pip install "git+https://github.com/cognis-digital/hallumark.git" # pip (works today)
pipx install "git+https://github.com/cognis-digital/hallumark.git" # isolated CLI
uv tool install "git+https://github.com/cognis-digital/hallumark.git" # uv
pip install cognis-hallumark # PyPI (when published)
docker run --rm ghcr.io/cognis-digital/hallumark:latest --help # Docker
brew install cognis-digital/tap/hallumark # Homebrew tap
curl -fsSL https://raw.githubusercontent.com/cognis-digital/hallumark/main/install.sh | sh
```
| Linux | macOS | Windows | Docker | Cloud |
|---|---|---|---|---|
| `scripts/setup-linux.sh` | `scripts/setup-macos.sh` | `scripts/setup-windows.ps1` | `docker run ghcr.io/cognis-digital/hallumark` | [DEPLOY.md](docs/DEPLOY.md) (AWS/Azure/GCP/k8s) |
<div align="right"><a href="#top">โ back to top</a></div>
<a name="related"></a>
## Related Cognis tools
- [`aegis`](https://github.com/cognis-digital/aegis) โ AI Agent Permission & Access Auditor โ surfaces the lethal trifecta of credentials + injection + reach
- [`promptmirror`](https://github.com/cognis-digital/promptmirror) โ Prompt-injection & indirect-injection scanner for any LLM context input
- [`ledgermind`](https://github.com/cognis-digital/ledgermind) โ Local LLM cost & token forensics proxy with anomaly detection
- [`adversa`](https://github.com/cognis-digital/adversa) โ LLM red-team harness โ OWASP LLM Top 10 + MITRE ATLAS attack packs
- [`guardpost`](https://github.com/cognis-digital/guardpost) โ Runtime agent firewall โ PII redaction, rate limits, policy enforcement
- [`aicard`](https://github.com/cognis-digital/aicard) โ Auto-generated NIST AI RMF / EU AI Act Annex IV model & system cards
**Explore the suite โ** [๐๏ธ all 170+ tools](https://github.com/cognis-digital/cognis-neural-suite) ยท [โญ awesome-cognis](https://github.com/cognis-digital/awesome-cognis) ยท [๐ cognis-sources](https://github.com/cognis-digital/cognis-sources) ยท [๐ค uncensored-fleet](https://github.com/cognis-digital/uncensored-fleet) ยท [๐ง engram](https://github.com/cognis-digital/engram)
<div align="right"><a href="#top">โ back to top</a></div>
<a name="contributing"></a>
## Contributing
PRs, new rules, and demo scenarios are welcome under the collaboration-pull model โ see [CONTRIBUTING.md](CONTRIBUTING.md) and [SECURITY.md](SECURITY.md).
> ### โญ If `hallumark` saved you time, **star it** โ it genuinely helps others find it.
## Interoperability
`{}` composes with the 300+ tool Cognis suite โ JSON in/out and a shared
OpenAI-compatible `/v1` backbone. See **[INTEROP.md](INTEROP.md)** for the
suite map, composition patterns, and reference stacks.
## License
Source-available under the **Cognis Open Collaboration License (COCL) v1.0** โ free for personal, internal-evaluation, research, and educational use; **commercial / production use requires a license** (licensing@cognis.digital). See [LICENSE](LICENSE).
---
<div align="center"><sub><b><a href="https://cognis.digital">Cognis Digital</a></b> ยท one of 170+ tools in the <a href="https://github.com/cognis-digital/cognis-neural-suite">Cognis Neural Suite</a> ยท <i>Making Tomorrow Better Today</i></sub></div>
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues