bd-graph MCP Server
by halaprix
README.md
# bd-graph
**A temporal knowledge graph over your [beads](https://github.com/gastownhall/beads)
issue tracker — built by coding agents, queried over MCP, no LLM API keys required.**
Your bd store accumulates decisions, invariants, supersessions, and provenance —
but keyword search can't answer *shape* questions: *which rules currently govern
this component, and since when? Was that memory superseded? Which lasting
decisions came out of that epic?* bd-graph derives a Neo4j graph from your bd
corpus (issues, memories, ADRs) that answers exactly those, and exposes it to
coding agents through a read-only MCP server.
The temporal layer is the point: edges carry `valid_from` / `invalid_at`, and
`SUPERSEDES` / `INVALIDATES` edges record when guidance flipped — so agents
retrieve the *current* rule instead of a stale one.
**The graph is a derived, disposable index. bd stays the source of truth.**
Rebuild it any time with one command; never write project state to it.
## How it works
```
bd store ──export──▶ corpus/ ──┬─▶ load_p1.py (deterministic: structure + regex)
docs/adr ──copy───▶ └─▶ load_p2.py (semantic: agent-extracted JSON)
│
Neo4j ◀───────────┘
▲
mcp-neo4j-cypher (read-only MCP) ◀── your coding agents
```
**Phase 1 — deterministic (no LLM).** Issue dependency trees (`CHILD_OF`,
`DEPENDS_ON`, `DISCOVERED_FROM`), ADR→issue tracking links, ADR supersession
status lines, and regex cross-mentions between issues/ADRs/memories/PRs. This
alone is ~90% of the edges.
**Phase 2 — semantic (agents, not APIs).** Your coding agents (Claude Code
subagents, Codex, whatever you run) extract `Component` and `Rule` entities and
`GOVERNS` / `ESTABLISHES` / `SUPERSEDES` / `INVALIDATES` edges with dates, using
the prompt template in `prompts/extraction.md`. Because the extractor is your
agent harness, this costs subscription tokens — no `OPENAI_API_KEY`, no
embeddings provider, nothing.
**Precision pass.** A second wave of agents adversarially verifies every
extracted edge against its source (`prompts/verification.md`); `apply_verdicts.py`
deletes refuted edges and tags unverifiable ones `confidence: "unsupported"`.
In practice verifiers kill 5–15% of edges — mostly fabricated dates and
misattributed components.
## Data model
| Node | Key | From |
|---|---|---|
| `Issue` | `id` | every bd issue incl. closed (title, status, dates, close_reason, summary) |
| `Memory` | `key` | `bd memories --json` |
| `ADR` | `id` | `NNNN-slug.md` files (title, status, date) |
| `PR` | `number` | regex over all text |
| `Component`, `Rule` | `name` | Phase-2 extraction |
| `Meta` | — | freshness stamp: `generated_at`, corpus counts |
Phase-1 edges: `CHILD_OF`, `DEPENDS_ON`, `DISCOVERED_FROM`, `TRACKED_IN`,
`SUPERSEDES` (ADR status lines), `MENTIONS`, `REFERENCES_PR`.
Phase-2 edges: `ABOUT`, `ESTABLISHES`, `GOVERNS`, `CONSTRAINS`, `SUPERSEDES`,
`INVALIDATES`, `PART_OF` — each with `fact`, `valid_from`, `invalid_at`,
`confidence`, `source`.
## Quickstart
Requirements: Python ≥3.11, Docker, a beads store (tested against bd 1.1.0);
`pipx install mcp-neo4j-cypher` for the MCP server (no pipx? `python3 -m pip
install --user pipx` first).
```bash
git clone <this repo> && cd bd-graph
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
docker run -d --name bd-graph-neo4j --restart unless-stopped \
-p 127.0.0.1:7474:7474 -p 127.0.0.1:7687:7687 \
-v bd-graph-data:/data -e NEO4J_AUTH=neo4j/<pick-a-password> neo4j:5-community
cp config.example.toml config.toml # edit: id_prefix, password, repo_root, vocab
# repo_root IS REQUIRED for the first run —
# it's what lets regenerate.sh export your corpus
mkdir -p corpus extracted
./regenerate.sh # exports corpus, builds Phase 1
.venv/bin/python test_graph.py # structural + recall assertions
```
Then wire the MCP server into your agent tool (Claude Code shown):
```bash
claude mcp add bd-graph --scope local -- mcp-neo4j-cypher \
--db-url bolt://localhost:7687 --username neo4j --password <password> \
--database neo4j --transport stdio --read-only
```
For Phase 2, dispatch extraction agents with `prompts/extraction.md`, verify
with `prompts/verification.md`, then:
```bash
.venv/bin/python apply_verdicts.py && ./regenerate.sh
```
### Try it without a bd store
`demo/` contains a synthetic corpus (5 issues, 3 memories, 2 ADRs, a sample
extraction) — setup commands in `demo/config.demo.toml`. The demo exercises
every edge type including a supersession chain.
## Example queries
```cypher
// pre-flight: what binds this component, and since when
MATCH (r:Rule)-[g:GOVERNS|CONSTRAINS]->(c:Component {name:'Api'})
OPTIONAL MATCH (src)-[e:ESTABLISHES]->(r)
RETURN r.name, g.fact, e.valid_from, coalesce(src.id, src.key);
// staleness chain: what superseded what, in date order
MATCH (a)-[s:SUPERSEDES|INVALIDATES]->(b)
RETURN coalesce(a.id,a.key,a.name), type(s), s.valid_from,
coalesce(b.id,b.key,b.name) ORDER BY s.valid_from;
// epic rollup: lasting decisions that came out of an epic
MATCH (adr:ADR)-[:TRACKED_IN]->(:Issue)-[:CHILD_OF*1..2]->(:Issue {id:'demo-ep1'})
MATCH (adr)-[:ESTABLISHES]->(r:Rule) RETURN adr.id, r.name;
// issue provenance: where it came from, what it left behind
MATCH (i:Issue {id:'demo-bug7'})
OPTIONAL MATCH (i)-[:DISCOVERED_FROM]->(p)
OPTIONAL MATCH (i)-[:REFERENCES_PR]->(pr)
RETURN p.id, collect(pr.number);
```
## Keeping it honest
- **Freshness is queryable in-band**: `MATCH (m:Meta) RETURN m.generated_at,
m.issues`. Tell your agents to check it at first use per session.
- **`./doctor.sh`** names the failing layer (container / bolt / data /
staleness) with the exact fix command; exit 0/2/1 = healthy/stale/broken.
`--full` chains both test suites.
- **`test_graph.py`** asserts corpus↔graph parity, dependency-edge parity,
no dupes/self-loops, ISO dates, all corpus supersession lines present — plus
any project-specific `[[probes]]` you define in config.
- **`test_mcp.py`** drives the MCP server over raw stdio JSON-RPC and proves
the write path is closed.
- Suggested agent-instruction snippet (CLAUDE.md / AGENTS.md):
> Use the `bd-graph` MCP for shape questions over the tracker (what governs a
> component, supersession chains, epic rollups, provenance). It is a derived
> index — on any mismatch trust bd. Check `Meta.generated_at` at first use;
> on any problem run `doctor.sh` and say explicitly that you're falling back
> to the bd CLI. Never fail silently.
## Design notes & limitations
- **Recall is deliberately partial** (~8 edges/doc): the graph is a map, not a
substitute for reading the document it points to.
- ADR parsing assumes the common `# ADR-NNNN: title` + `- **Status:** / - **Date:**
/ - **Tracked in:**` header block; adjust `ADR_META` in `load_p1.py` for
other formats. Duplicate ADR numbers merge into one node — treat that as a
lint error in your repo.
- Extraction quality depends on your `[vocab]` component list; keep it short
and canonical.
- Inspired by the Graphiti/Neo4j temporal-graph approach
([getzep/graphiti](https://github.com/getzep/graphiti)); this project trades
Graphiti's embedding-based hybrid search for zero API keys and
agent-native extraction. If you have API keys and want semantic search,
Graphiti is the heavier-duty option.
## Security
Everything is local-only by design: Neo4j binds to 127.0.0.1, the MCP server
is registered `--read-only`, and `config.toml` (credentials, project vocab) is
gitignored. Your corpus (`corpus/`, `extracted/`) contains your project's
internal knowledge — both are gitignored; never commit or publish them.
Note that `mcp-neo4j-cypher` takes the password as a CLI argument (visible in
`ps` on the local machine) — one more reason the password must be a
local-only throwaway, never a reused credential.
## License
MIT
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues