BlackBook MCP
README.md
<div align="center">
<img src="assets/blackbook-mcp-final-refined.png" alt="BlackBook MCP Logo" width="220" style="margin-bottom: 20px;"/>
# BlackBook MCP v0.9.0
### Source-Grounded Cybersecurity and bugbounty Knowledge & Research MCP
[](#)
[](https://www.python.org/)
[](https://modelcontextprotocol.io/)
[](#available-mcp-tools)
[](#retrieval-architecture)
[](#what-it-is)
[](#testing)
[](#roadmap)
[](LICENSE)
[](#security-model)
[](https://github.com/Daniel-wambua/BlackBook/actions/workflows/ci.yml)
**The read-only cybersecurity knowledge & research teammate: source-grounded search, exact citations, a knowledge graph, and investigation context, running alongside an execution MCP.**
[π What It Is](#what-it-is) β’ [ποΈ Architecture](#architecture-overview) β’ [π Installation](#installation) β’ [π οΈ MCP Tools](#available-mcp-tools) β’ [πΈοΈ Knowledge Graph](#knowledge-graph) β’ [π Security](#security-model)
</div>
***
**BlackBook MCP** is a source-grounded cybersecurity **knowledge & research** server
that speaks the Model Context Protocol (MCP). It is the research teammate that runs
*alongside* an execution MCP such as HexStrike inside Claude Code, Cursor, VS Code,
or any MCP-compatible client.
```
CLAUDE / AI AGENT
|
+------------+------------+
| |
v v
HEXSTRIKE BlackBook MCP
EXECUTION KNOWLEDGE
| |
+------+------+ +--------+--------+
| | | | | |
Nmap ffuf nuclei HackTricks 0xdf ATT&CK
| | | | | |
+------+------+ +--------+--------+
| |
+------------+------------+
v
AI REASONING LOOP
```
* **HexStrike** answers: *"what can I execute or test?"*
* **BlackBook** answers: *"what is documented about this situation, which similar
cases exist, which techniques are relevant, and what source material supports
that conclusion?"*
Claude is the orchestrator.
> **Read-only by design.** BlackBook never runs commands, scans hosts, or exploits
> targets. It indexes a controlled corpus and retrieves source-grounded knowledge
> with exact, verifiable citations. Execution belongs to a separate MCP.
***
## Architecture Overview
BlackBook MCP v0.9.0 is a source-grounded knowledge system: every query flows through
a hybrid retrieval facade, is enriched (never gated) by a knowledge graph, and returns
results that resolve to exact, verifiable citations. Nothing is executed.
```mermaid
%%{init: {"themeVariables": {
"primaryColor": "#7f1d1d",
"secondaryColor": "#dc2626",
"tertiaryColor": "#ef4444",
"background": "#1a0505",
"edgeLabelBackground":"#7f1d1d",
"fontFamily": "monospace",
"fontSize": "15px",
"fontColor": "#fee2e2",
"nodeTextColor": "#fee2e2"
}}}%%
graph TD
A[AI Agent - Claude / Cursor / VS Code] -->|MCP Protocol over stdio| B[BlackBook MCP Server v0.9.0]
B --> C[Hybrid Retrieval Facade]
B --> D[12 Knowledge Tools]
B --> E[Knowledge Graph]
C --> F[FTS5 BM25 - always on]
C --> G[Local Semantic - optional]
C --> H[Reranker + Source Diversity]
D --> I[knowledge_search]
D --> J[knowledge_source]
D --> K[knowledge_technique]
D --> L[knowledge_case_search]
D --> M[knowledge_research]
D --> M2[knowledge_graph]
D --> N[knowledge_context]
E --> O[Technique / Tool / Service / OS]
E --> P[Writeup / Source entities]
E --> Q[Evidence-linked edges]
B --> R[Corpus - SQLite FTS5 + JSON1]
R --> S[HackTricks]
R --> T[0xdf Writeups]
R --> U[Local PDFs]
R --> X["MITRE ATT&CK"]
R --> Y[GTFOBins + LOLBAS + LOOBins]
R --> Z[Payloads + Recipes + WADComs]
R --> AA[InternalAllTheThings + HTB Writeups]
B --> V[Exact Citations and Provenance]
V --> W[chunk_id resolves to verifiable excerpt]
style A fill:#7f1d1d,stroke:#ef4444,stroke-width:3px,color:#fee2e2
style B fill:#dc2626,stroke:#7f1d1d,stroke-width:4px,color:#ffffff
style C fill:#ef4444,stroke:#7f1d1d,stroke-width:2px,color:#1a0505
style D fill:#ef4444,stroke:#7f1d1d,stroke-width:2px,color:#1a0505
style E fill:#ef4444,stroke:#7f1d1d,stroke-width:2px,color:#1a0505
style R fill:#ef4444,stroke:#7f1d1d,stroke-width:2px,color:#1a0505
style V fill:#ef4444,stroke:#7f1d1d,stroke-width:2px,color:#1a0505
```
### How It Works
1. **AI Agent Connection**: Claude, Cursor, VS Code, or any MCP-compatible client
connects over stdio. The server owns stdout for the JSON-RPC protocol; every byte
of banner/log chrome goes to stderr, so the stream is never corrupted.
2. **Source-Grounded Retrieval**: a query flows through metadata filters β FTS5 BM25
(always available) β optional **local** semantic search β reranking β a
per-document cap that enforces source diversity.
3. **Graph Enrichment**: the knowledge graph annotates technique dossiers and
similar-case results with evidence-linked edges. It *enhances* retrieval and never
gates it: everything works with an empty graph.
4. **Verifiable Provenance**: every result resolves through its `chunk_id` to the
exact indexed excerpt. BlackBook never fabricates a citation.
5. **Read-Only by Design**: no command execution, host scanning, or arbitrary URL
fetching through tool parameters. Execution belongs to a separate MCP such as
HexStrike.
***
## What it is
A **hybrid knowledge system**, not a queryβembeddingβdump pipeline:
```
query β metadata filter β FTS5 (BM25) β [optional semantic] β rerank β
source diversity β provenance β exact citations
```
Lexical retrieval (SQLite FTS5) is the always-available backbone. Semantic search is
optional and local. Nothing is presented as fact unless it traces to an indexed
source chunk.
## Features (current phase)
* **Source-grounded search** across 26 configured sources (26 enabled by
default): HackTricks, 0xdf
writeups, MITRE ATT&CK, GTFOBins, LOLBAS, LOOBins, WADComs,
PayloadsAllTheThings, The Hacker Recipes, Internal All The Things,
Moamen Basel's HTB writeups, local PDFs, OWASP WSTG, OWASP ASVS,
OWASP API Security, the Bug Bounty Cheatsheet, PortSwigger Academy,
Google Bug Hunters, Bugcrowd VRT, GitHub Security Lab research, Hacker101,
and four public HackerOne report/index repositories, plus WebHackList. Report
archives are labeled with unknown authority and should be treated as local
research data, not official HackerOne guidance.
* **Exact, verifiable citations**: every reference resolves to real indexed text
* **Structure-preserving chunking**: heading breadcrumbs and code blocks intact
* **Hybrid retrieval facade** with reranking + source diversity: lexical (FTS5
BM25) always on, **local** semantic embeddings merged in when enabled
* **Local semantic search** (`sentence-transformers`, offline): paraphrased
queries with no keyword overlap still find the right chunk; degrades gracefully
to lexical when the extra isn't installed
* **Source filtering & platform/category filters**
* **Knowledge graph** (Technique/Tool/Service/OS/Writeup/Source) built from the
index; evidence-linked edges enrich technique dossiers and case search without
ever gating retrieval
* **Modular ingestion** via a `SourceAdapter` interface (add sources without a rewrite)
* **CLI** for ingestion, search, graph, stats, sources, diagnostics
* **MCP server over stdio** (default) for Claude Code / Cursor / VS Code, with an
optional **streamable-http** transport (`serve --http`) for a shared network server
## Installation
Requires Python β₯ 3.10.
```bash
# with uv (recommended)
uv pip install -e .
# or with pip
pip install -e .
# optional: semantic/embedding search (Phase 3)
uv pip install -e ".[semantic]"
# development / tests
uv pip install -e ".[dev]"
```
This installs two CLI entry points: `blackbook` and `cyber-knowledge` (alias).
## Configuration
BlackBook reads, in increasing priority: built-in defaults β a YAML config file β
`BLACKBOOK_*` environment variables.
```bash
cp config.example.yaml ~/.blackbook/config.yaml
# edit paths/sources; see config.example.yaml for every option
```
Key settings:
```yaml
home: ~/.blackbook # data dir (db, caches, raw checkouts)
sources:
- id: hacktricks
enabled: true
- id: "0xdf" # quote hex-like ids (YAML parses 0xdf as 223)
enabled: true
- id: local_pdfs
enabled: true
directory: ~/knowledge/pdfs
authority: user # NOT assumed authoritative
- id: attack # MITRE ATT&CK STIX bundle (authority: official)
enabled: true
embeddings:
enabled: false # set true + install [semantic] for local semantic search
model: sentence-transformers/all-MiniLM-L6-v2
device: cpu
retrieval:
default_limit: 8
per_document_cap: 2 # source diversity: max chunks per document
per_source_cap: 4 # source diversity: max results per source
query_log:
enabled: true # record every search locally (see Query log)
max_entries: 5000 # newest entries kept; older ones pruned
```
Twenty-six sources are configured and enabled by default (run `blackbook sources`
to list them). Website sources are bounded to their configured origin and optional
`path_prefix`; they skip non-HTML assets and respect `max_files`,
`max_document_bytes`, and `request_delay`.
GitHub-backed sources accept a few extra keys: `ref` (branch), `include_glob`
(which files to index), `exclude_glob` (skip repo plumbing / link indexes),
`content_root` (restrict to a subtree), and `site_url` (map citations to the
published site instead of the GitHub blob URL; a Jekyll `permalink` in a
page's front matter wins over the path-derived URL). See
`config.example.yaml`.
### Source freshness
Every ingest run stamps the source it pulled: when it was last fetched
successfully, and at which revision. `blackbook sources` shows both, and the
`knowledge_sources` tool returns them as `last_fetched` and `version`.
The timestamp means *last successful pull*, not *last attempt*: a fetch that
fails leaves the previous stamp in place, because a run that did not complete
should not look like one that did. The revision is the commit the extracted
tree actually came from, read back from the fetch marker rather than
remembered in memory, so a run that skipped the download still reports what it
is working from. Sources with no revision to speak of (a website crawl, a
local directory) report a timestamp and no version, and a source that has
never been fetched reports neither rather than a misleading zero.
## Initial ingestion
```bash
blackbook ingest --source hacktricks # markdown book (tarball over HTTPS)
blackbook ingest --source 0xdf # HTB/CTF writeups
blackbook ingest --source attack # MITRE ATT&CK STIX bundle (~54 MB download)
blackbook ingest --source gtfobins # Unix binary abuse (YAML corpus)
blackbook ingest --source lolbas # Windows living-off-the-land binaries
blackbook ingest --source loobins # macOS living-off-the-land binaries
blackbook ingest --source wadcoms # offensive Windows/AD command cheat sheets
blackbook ingest --source payloads # PayloadsAllTheThings
blackbook ingest --source hacker_recipes # The Hacker Recipes
blackbook ingest --source internal_all_the_things # AD / internal network cheat sheets
blackbook ingest --source htb_writeups # Moamen Basel's HTB writeups + cheatsheets
blackbook ingest --source owasp_wstg # OWASP web testing methodology
blackbook ingest --source owasp_asvs # OWASP verification requirements
blackbook ingest --source owasp_api_security # OWASP API security guidance
blackbook ingest --source bugbounty_cheatsheet # practical bug bounty workflow
blackbook ingest --source portswigger # Web Security Academy guidance
blackbook ingest --source google_bug_hunters # Google bug bounty guidance
blackbook ingest --source bugcrowd_vrt # vulnerability severity taxonomy
blackbook ingest --source github_security_lab # GitHub Security Lab research
blackbook ingest --source hacker101 # HackerOne educational material
blackbook ingest --source hackerone_reports_index # report index
blackbook ingest --source hackerone_disclosed_reports # report bodies
blackbook ingest --source hackerone_reports_metadata # report metadata
blackbook ingest --source hackerone_bug_bounty_reports # report index
blackbook ingest --source webhacklist # web hacking technique archive
blackbook ingest # all enabled sources
# bound the size during a first run:
# set `max_files: 25` on a source in config.yaml
```
`ingest` is incremental by default. For the normal daily workflow, use
`blackbook update`: it checks every enabled source, downloads only changed
GitHub revisions or missing cached pages, skips unchanged documents by content
hash, prints `status=up-to-date` for sources with no changes, and continues if
another source has an error. New sources are ingested normally on the same run.
Use `blackbook ingest --force` only when you explicitly need a full refresh.
Citation metadata (URLs, titles) is refreshed in place when it drifts, so
a URL-mapping fix or a source re-publishing under new permalinks self-heals
without re-chunking or new chunk ids.
WebHackList ingests its complete Markdown repository, including yearly lists,
archived references, and evaluation notes. It enables exact normalized
cross-document deduplication, so repeated chunks keep the first citation while
unique material remains searchable.
## PDF ingestion
```bash
# point the local_pdfs source at your directory in config.yaml, then:
blackbook ingest --source local_pdfs
```
PDFs are chunked per page with page-number citations. They default to
`authority: user` and are **not** treated as authoritative.
## Searching
```bash
blackbook search "kerberoasting"
blackbook search "windows service privilege escalation" --source hacktricks
blackbook search "NTLM relay" --platform windows --limit 5
blackbook search "crack service account passwords" --mode semantic # paraphrase-friendly
blackbook stats [--json] # corpus counts (machine-readable with --json)
blackbook sources [--json] # configured sources, index counts, freshness
blackbook queries [--empty] [--clear] # what was asked, and what came back empty
blackbook graph build [--full] # (re)build the knowledge graph, reusing cached terms
blackbook graph show [--json] # graph counts + writeup coverage by source
blackbook graph neighbors kerberoasting -d 2 # walk the graph from an entity
blackbook doctor # diagnostics: db, index, sources, embeddings
blackbook rebuild-index # rebuild the FTS5 index
blackbook case export MY-CASE # export an investigation case as Markdown
blackbook backup # snapshot the knowledge base (VACUUM INTO)
```
`platform` and `categories` are **hard filters**: results only come from
documents carrying the tag (e.g. `windows`/`linux`, `htb`, `Easy`/`Insane`).
The MCP tools' `techniques` parameter resolves through the controlled
vocabulary and biases results toward chunks whose heading names the technique;
unknown terms are searched as plain keywords and flagged in the response note.
Search modes: `hybrid` (default, lexical + semantic), `keyword` (FTS5 only),
`semantic` (embeddings only), plus two intent-biased modes: `technique` (nudges
canonical technique/reference material up) and `case_similarity` (favours hands-on
writeups). The intent modes *nudge* ranking, they never filter results out. Semantic
and hybrid use vectors only when `embeddings.enabled` and the `[semantic]` extra is
installed; otherwise they fall back to lexical automatically.
### Query log
Every search is recorded locally: the phrasing, the mode, which sources were
searched, how many results came back, the best score, and the latency.
```bash
blackbook queries # recent entries, newest first
blackbook queries --empty # only the ones that returned nothing
blackbook queries --json # stats + entries, machine-readable
blackbook queries --clear # delete the log
```
The useful rows are the empty ones. A query that returns nothing is the only
evidence the index gives about what it *cannot* answer, and that evidence is
invisible from the corpus side: no amount of reading the indexed documents
tells you which phrasing a user tried and missed. `--empty` is that view.
The log is local to the SQLite file and never transmitted anywhere. It is
bounded (the newest `query_log.max_entries` entries, default 5000, pruned on
insert) so it stays recent history rather than an ever-growing table, and it
can be turned off completely:
```yaml
query_log:
enabled: true # set false to record nothing at all
max_entries: 5000
```
A log write never affects the search it describes: failures are swallowed and
reported at debug level, so a locked database or a full disk degrades the log
and never the answer. `blackbook eval` does not write to the log at all, since
it measures retrieval rather than reporting usage.
### Semantic embeddings
With `embeddings.enabled: true` and the `[semantic]` extra, ingestion embeds new
chunks inline. To (re)build the semantic index without re-ingesting:
```bash
blackbook embed # embed chunks missing a current-model vector
blackbook embed --source local_pdfs # only one source
blackbook embed --reembed # drop existing vectors first, then re-embed
```
Embeddings are computed **locally** and never leave the machine. `blackbook doctor`
reports coverage (`N/M embedded`).
## Transport modes
BlackBook speaks MCP over two transports. **stdio is the default** and is what
you almost always want.
| Transport | How it starts | Who launches it | Use it for |
|-----------|---------------|-----------------|------------|
| **stdio** (default) | `blackbook serve` | the MCP client spawns it automatically | Claude Code / Cursor / VS Code |
| **streamable-http** | `blackbook serve --http` | you start it, it stays running | a shared always-on server, remote access, or just to see the banner |
* **stdio**: the client (Claude Code, Cursor, VS Code) owns the process. It
spawns `blackbook serve` on session start, talks over stdin/stdout, and stops
it on exit. Nothing to launch by hand. This is the mode all the setup snippets
below use.
* **streamable-http**: a long-lived network server you run yourself, reachable
at `http://<host>:<port>/mcp` with a `GET /health` check. Handy for a shared
instance or when you want the banner in front of you. It is **not** something a
client auto-launches, so don't register an HTTP endpoint that isn't already
running or the client will just fail to connect each session.
```bash
blackbook serve # stdio (default)
blackbook serve --http # streamable-http on 127.0.0.1:8890/mcp
blackbook serve --http --port 9000 # override port
blackbook serve --http --host 0.0.0.0 # bind all interfaces (see auth note below)
```
Equivalently, set `BLACKBOOK_TRANSPORT=streamable-http` (or `sse`) in the
environment. HTTP host/port/path come from the `server:` block in
`config.yaml` (defaults `127.0.0.1` / `8890` / `/mcp`); `--host` and `--port`
override them for a single run.
> **Bind safety.** BlackBook refuses to start on a non-loopback address (e.g.
> `0.0.0.0`) unless a bearer token is set, guarding against an accidentally
> exposed, unauthenticated server. Set `server.auth_token` in `config.yaml` (or
> `BLACKBOOK_SERVER__AUTH_TOKEN`) and clients must then send
> `Authorization: Bearer <token>`. Only flip `require_auth_off_loopback: false`
> if you understand the exposure.
### Health endpoint
`GET /health` is for monitoring, and it reports the whole surface rather than
just liveness: version, transport, corpus counts, and the names of every tool,
resource and prompt the server exposes. Opening the base URL instead gives a
landing page with the same information.
Because a monitor may poll `/health` on a timer, the corpus counts there (and
on the landing page) come from a short-lived cache rather than a fresh count.
Counting the chunk table is the one query that scales with the whole corpus,
around 4 ms at half a million chunks, which is pointless to repeat per poll for
a number nothing acts on. Any write **this** process commits clears the cache
outright, so a server that has just ingested reports the new totals on its next
request; a write from another process (a CLI ingest while the server is up) is
picked up within five seconds. Everything an answer depends on, including the
`blackbook://corpus` resource and the CLI, reads the exact counts.
## Claude Code setup
```bash
claude mcp add blackbook -- blackbook serve # this project (local scope)
claude mcp add blackbook -s user -- blackbook serve # every directory (user scope)
```
Use `-s user` to register it once for all your projects; Claude Code then
auto-launches it everywhere over stdio. Or put it in your MCP config
(`.mcp.json` / `~/.config/claude/...`):
```json
{
"mcpServers": {
"blackbook": { "command": "blackbook", "args": ["serve"] }
}
}
```
Using a virtualenv? Point `command` at it: `"/home/you/venv/bin/blackbook"`.
## Cursor
`Settings β MCP β Add server`:
```json
{ "mcpServers": { "blackbook": { "command": "blackbook", "args": ["serve"] } } }
```
## VS Code
`.vscode/mcp.json` (with an MCP-capable extension):
```json
{ "servers": { "blackbook": { "command": "blackbook", "args": ["serve"] } } }
```
## Startup banner
Launching the server prints a banner and then streams status logs. Every byte of
this chrome goes to **stderr**; stdout is reserved for the JSON-RPC protocol, so
the banner and logs never corrupt an MCP client's stream.
```
βββββββ βββ ββββββ ββββββββββ ββββββββββ βββββββ βββββββ βββ βββ
βββββββββββ βββββββββββββββββββ βββββββββββββββββββββββββββββββββ ββββ
βββββββββββ βββββββββββ βββββββ βββββββββββ ββββββ ββββββββββ
βββββββββββ βββββββββββ βββββββ βββββββββββ ββββββ ββββββββββ
βββββββββββββββββββ ββββββββββββββ ββββββββββββββββββββββββββββββββ βββ
βββββββ βββββββββββ βββ ββββββββββ ββββββββββ βββββββ βββββββ βββ βββ
Source-grounded cybersecurity knowledge & research MCP
v0.9.0 Β· stdio Β· read-only Β· no execution Β· every claim cited
corpus <live database count> sources Β· <live count> docs Β· <live count> chunks Β· <live count> embeddings
graph <live count> entities Β· <live count> relationships Β· <live count> cases
```
In a real terminal the wordmark is gradient-lit (cyanβindigo, intentionally
distinct from an execution MCP's red). The corpus and graph lines always reflect
the live database, so the counts above are placeholders rather than a fixed
snapshot. The transport line reflects how you started it (`stdio`, or
`streamable-http Β· http://127.0.0.1:8890/mcp` under `--http`, where the MCP
endpoint and `/health` URL are also printed). Suppress the banner with
`blackbook serve --no-banner`.
Status and log lines use a compact, level-styled prefix, showing the successes and
failures at a glance:
```
[+] Embedded 18630 chunks. Total vectors: 18630 success (green)
[*] server ready info (cyan)
[!] Graph rebuild skipped: no chunks changed warning (yellow)
[-] hacktricks: fetch failed (offline) error (red)
```
Rich strips the colour automatically when output is piped or redirected, so log
files stay clean.
## Available MCP tools
| Tool | Status | Purpose |
|------|--------|---------|
| `knowledge_search` | β
| Source-grounded search with provenance-tagged results |
| `knowledge_source` | β
| Resolve a reference to the exact supporting excerpt |
| `knowledge_technique` | β
| Structured technique dossier (graph-enriched, always cited; official ATT&CK tactics/platforms/link when mapped) |
| `knowledge_graph` | β
| Bounded graph walk outward from any entity, evidence-linked edges |
| `knowledge_case_search` | β
| Similar-case (writeup) retrieval, techniques annotated |
| `knowledge_research` | β
| Observation-driven, source-grounded research packets |
| `knowledge_context` | β
| Local investigation state (cases + observations) |
| `knowledge_hunt_plan` | β
| Cited, non-executing bug bounty validation plans |
| `knowledge_finding_review` | β
| Evidence-gap and severity-guidance review |
| `knowledge_report_draft` | β
| Cautious report drafts from local case evidence |
| `knowledge_sources` | β
| Configured sources, index counts, and freshness (last fetch, revision) |
| `knowledge_compare` | β
| Independent multi-source evidence comparison |
Only implemented tools are registered; nothing is stubbed or faked.
### MCP resources
Resources are read-only views of local state. A tool answers a question; a
resource is there for an agent to see what the corpus holds without spending a
tool call on it. Each one is computed on read, so it can never describe a
database other than the one it was just read from, and none of them write or
reach the network.
| Resource | Returns |
|----------|---------|
| `blackbook://sources` | Configured sources with index counts and freshness. The same payload as `knowledge_sources` |
| `blackbook://corpus` | Counts: sources, documents, chunks, embeddings, graph entities and relationships, cases, query-log totals |
| `blackbook://vocabulary` | The controlled vocabulary: service, technique and tool terms, the ATT&CK id each technique maps to, and the aliases the filters resolve |
| `blackbook://cases` | Local investigation cases: name, target, platform, observation count |
| `blackbook://case/{name}` | One case rendered as portable Markdown. A missing case returns a short note naming `blackbook://cases`, not an error |
`blackbook://vocabulary` is the one worth reading first if you are driving the
server yourself. It lists the exact spellings the `techniques` filter resolves
against, so you can ask for `kerberoasting` rather than guess at a phrasing the
index will not match.
### MCP prompts
Prompts are framings for the questions this corpus can answer. The failure mode
with a knowledge server is not "no answer", it is a fluent answer the corpus
never supported, so every one of these routes the agent through the tools and
asks it to cite what comes back and name the gap where nothing did.
| Prompt | Arguments | Use |
|--------|-----------|-----|
| `triage_observation` | `observation`, `target` (optional) | Turn a raw observation into a source-grounded triage starting point |
| `explain_technique` | `technique`, `platform` (optional) | Explain a technique strictly from what the indexed sources document |
| `review_finding` | `finding` | Review a suspected finding for evidence gaps before it is reported |
| `draft_report` | `case` | Draft a cautious report from a local case, keeping the warnings where evidence is missing |
### Example Claude Code interaction
```
You: What does HackTricks document about Kerberoasting, and has 0xdf
covered a similar HTB machine?
Claude: (calls knowledge_search {query: "kerberoasting", sources: ["hacktricks","0xdf"]})
HackTricks documents Kerberoasting under Active Directory β Kerberos β¦
Similar 0xdf case: HTB: Forest β¦
[cites chunk refs]
Claude: (calls knowledge_source {chunk_id: β¦} to read the exact section)
Here's the exact HackTricks enumeration procedure β¦
```
## Knowledge graph
A lightweight graph of `Technique / Tool / Service / OS / Writeup / Source`
entities and their relationships (`documented_by`, `demonstrated_in`, `uses`,
`targets`, `runs_on`, β¦), derived **from the already-indexed corpus**, with no fetching
or execution. Every non-structural edge carries the document it was extracted from
(`evidence_doc_id`), a `confidence`, and an `inferred` flag; nothing is fabricated,
and a citation always resolves to real indexed text.
The graph **enhances** retrieval, it never gates it: search and both new tools work
with an empty graph and simply gain neighbours/annotations once it is built.
```bash
blackbook graph build # (re)build the graph from the index, idempotent
blackbook graph build --full # re-extract every document, ignoring the term cache
blackbook graph show # current entity/relationship counts, no rebuild
blackbook graph neighbors kerberoasting --depth 2 --direction out
```
`blackbook graph show` also reports **writeup coverage**: how many documents count as
hands-on writeups and which sources contribute none. The graph's own writeup total is a
single number, so a corpus where every writeup comes from one source while the largest
source contributes none looks fully populated; the coverage table names the gaps
instead. `blackbook doctor` reports the same figures as a single check.
Edges come in two kinds, and they are stored differently on purpose. A **documentary**
edge (`documented_by`, `demonstrated_in`, `used_in`, `present_in`, `runs_on`) is a fact
about one document, so it is stored once per document with `support = 1` and that
document as its citation. A **co-occurrence** edge (`uses`, `targets`) is one claim that
many documents can witness: two terms appearing in the same page says the same thing
however many pages say it. Those collapse to a single row per pair, with `support`
counting how many distinct documents backed the claim and the lowest such document kept
as the citation. Without that, a real corpus stores ~43 rows per pair and
`knowledge_technique` repeats the same tool or service hundreds of times; with it, each
neighbour is listed once and `support` is what separates a claim made everywhere from
one made once.
Ingesting also refreshes the graph automatically (skip with `ingest --no-graph`).
**A rebuild is incremental.** Term extraction is a regex pass over every document's
full text and accounts for essentially all of a build's cost, so the vocabulary terms
each document contributed are cached, keyed by a fingerprint of everything the
extraction reads (content hash, title, source, categories, metadata). A rebuild
re-extracts only the documents whose fingerprint moved, then assembles the graph from
all documents' terms as before, which is what makes the result identical to a
from-scratch build: only the input to assembly is cached, never its output. When
nothing changed at all, the rebuild is skipped outright and the command says so. On a
25k-document corpus this takes a rebuild from about two minutes to under a second when
nothing moved, and to a couple of seconds after a small ingest. Use `--full` to
distrust the cache and re-extract everything. The fingerprint carries a cache version
constant that is bumped whenever the extraction itself changes (a new vocabulary list,
a changed alias table, a different scan cap), so a cache written by an older version
cannot keep serving terms the current code would no longer produce.
Three tools consume the graph:
* `knowledge_technique`: returns which sources document a technique, which
tools/services/writeups the graph associates with it (each edge with confidence,
support, and its backing document), plus real cited excerpts. Works before the graph
exists; it always returns indexed references.
* `knowledge_graph`: walks the graph outward from any entity (technique, tool,
service, os, writeup or source), up to 3 hops, following any predicate in either
direction, and returns the reached entities plus the edges between them. Where
`knowledge_technique` answers a fixed question about one technique, this answers
the open one: what is around this entity at all. Resolution is exact-match, so a
near miss returns candidate names rather than traversing the wrong entity and
producing a plausible, entirely wrong neighbourhood. Both bounds on the walk (a
per-node `limit` and an overall `max_nodes`) are reported through `truncated` and
`note` instead of being applied silently, so a partial result is never mistaken
for a complete one. Available from the terminal as `blackbook graph neighbors`.
* `knowledge_case_search`: finds hands-on writeups similar to a situation and,
when the graph is built, annotates each with the techniques it demonstrates.
## Retrieval architecture
See `docs/retrieval.md`. FTS5 BM25 is always available; semantic search is an
optional local backend merged into the same facade. Reranking combines lexical
score, source authority, platform/category match, and keyword overlap, then a
per-document cap enforces source diversity.
## Source diversity
Two caps shape a result set, and they differ in kind.
The **per-document cap** (`retrieval.per_document_cap`, default 2) is hard.
Past it a chunk is discarded, because several chunks of one page are near
certainly the same evidence restated.
The **per-source cap** (`retrieval.per_source_cap`, default 4) is a
preference. HackTricks alone is over a thousand documents here, so a query it
covers heavily can otherwise return eight HackTricks chunks and nothing else
while the writeup repos, the cheatsheets and the local pdfs hold comparable
evidence. Past the cap a hit is set aside rather than dropped, and any slot the
diverse hits did not fill is spent on exactly those set-aside hits.
The consequence worth stating plainly is that **the cap never shortens a
result list**. Ask for eight and you get eight; on a corpus where one source
holds everything relevant the output is identical, and the cap only reorders
when there was genuinely another source to prefer. Set it to `0` to disable.
On the current corpus this changes four of the twenty four answerable gold
queries, and changes none of their metrics: `hit_rate`, `mrr`,
`mean_recall_at_k` and `citation_integrity` are identical with the cap on and
off, so nothing relevant was displaced. What it does change is the spread. One
query that returned six chunks from a single source now returns four sources.
## Source provenance
Every claim carries provenance. `knowledge_search` returns a `ref` (chunk_id,
doc_id, source, url, page, section_path); `knowledge_source` resolves it to the
exact indexed text. BlackBook never fabricates a citation.
## Troubleshooting
```bash
blackbook doctor --verbose # full diagnostics
```
* **`Unknown or disabled source`**: check `blackbook sources`; quote `"0xdf"` in YAML.
* **Empty results**: run `blackbook ingest` first; check `blackbook stats`.
* **PDF dir missing**: set `sources[].directory` for `local_pdfs`.
* **Logs**: add `--verbose` to any command for structured debug output.
## Security model
BlackBook is read-only with respect to external systems, confines filesystem reads
to configured knowledge directories, validates all tool inputs, and never fetches
arbitrary URLs through tool parameters. See `docs/security.md`.
## Development
```bash
uv pip install -e ".[dev]"
python -m pytest tests # run the suite
```
Layout:
```
src/blackbook/
config.py layered settings
server.py FastMCP wiring (stdio)
mcp/ tool schemas + implementations
ingestion/ SourceAdapter + per-source adapters + pipeline
retrieval/ lexical / hybrid / reranker / chunking
knowledge/ source resolution (citation -> excerpt)
storage/ SQLite (FTS5 + JSON1), models, migrations
cli/ Typer CLI
utils/ path-safety helpers
tests/ unit + fixtures (+ integration)
docs/ architecture, ingestion, retrieval, mcp, security
```
## Testing
```bash
python -m pytest tests -q
```
Tests cover chunking, storage/FTS5 sync, every source adapter (offline
fixtures: HackTricks, 0xdf, GitHub markdown, GTFOBins, LOLBAS, ATT&CK STIX),
retrieval & reranking, semantic embeddings & hybrid merge, MCP tools,
provenance round-trips, and path safety. Semantic tests use a deterministic
model-free embedder so they run offline with no model download; one real-model
test skips cleanly when the `[semantic]` extra isn't installed. PDF tests use a
generated PDF and skip if `reportlab` isn't installed.
## Roadmap
- [x] **Phase 1**: MCP server, SQLite+FTS5, HackTricks + 0xdf ingestion, search, citations
- [x] **Phase 2**: font-aware PDF adapter (heading/code detection, page-level
citations), cross-document near-duplicate detection, structural chunking, CLI
- [x] **Phase 3**: local embeddings (`all-MiniLM-L6-v2`), hybrid retrieval, reranking
- [x] **Phase 4**: knowledge graph, technique relationships, case similarity
- [x] **Phase 5**: `knowledge_research` (observation β source-grounded packet),
`knowledge_context` (local investigation state)
- [x] **Phase 6**: offline evaluation suite (`blackbook eval`), citation-integrity
gate, FTS5 optimize on ingest, adversarial/hardening tests
- [x] **Phase 7**: generic GitHub source adapter (tarball over HTTPS, config-driven)
with PayloadsAllTheThings, The Hacker Recipes, GTFOBins, LOLBAS, LOOBins,
WADComs; MITRE ATT&CK STIX source with technique-dossier enrichment
- [x] **Phase 8**: MCP prompts and resources (corpus, sources, vocabulary, cases as
readable views); an ATT&CK ID lookup derived from the indexed corpus rather than a
hand-kept table; a bounded local query log (`blackbook queries`) covering what was
asked, what came back empty, and when a semantic request fell back to lexical;
per-source freshness reporting; a per-document source-diversity cap in retrieval;
search diagnostics in `blackbook doctor`; and an incremental knowledge-graph rebuild
that caches per-document term extraction
## License
MIT. See `LICENSE`.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues