ecfr-cartography-mcp
by apraba05
README.md
# ecfr-cartography-mcp
An MCP server that turns the US Code of Federal Regulations into an agent-navigable
citation graph — so agents can **traverse the regulatory graph instead of skimming
PDFs**. Point Claude at 45 CFR part 170 and ask "which section is the hub? which
citations are stale?"; you get structured, edge-by-edge answers grounded in the
live [eCFR API](https://www.ecfr.gov/developers/documentation/api/v1) instead of
best-effort text search.
Under the hood: a small pure-Python citation extractor (regex + reference graph),
a FastMCP server that wires six tools over stdio, and a D3 visualizer for the
README hero. All three share the same graph code so the picture can't drift from
what the model sees.

*45 CFR part 170 in the Cartography investigation workspace. Node size scales
with in-degree; the federal-blue cluster centers on the twin hubs § 170.315
and § 170.299; orphans (definitions, one-off subparts) sit ochre on the
periphery; and the red dashed ghost is § 170.503 — a citation target
§ 170.599 still points at, even though § 170.503 was removed in the 2020
ONC final rule.*
---
## Tools
| Tool | What it does | When to reach for it |
| --- | --- | --- |
| `search_regulations(query, per_page)` | Full-text search across all 50 CFR titles. | Discover which title/part governs a topic. |
| `get_part_structure(title, part)` | Ordered list of every section + its heading. | Get a table of contents cheaply. |
| `get_section_text(title, part, section, date?)` | Full section text + every citation extracted from it. | Read one section and see where it points. |
| `map_part_references(title, part)` | **Flagship.** Builds the intra-part citation graph — nodes, edges, top-10 hubs, orphans, external refs, statutory authorities. | Ask "which section anchors this part?" or "what other parts does this part depend on?" |
| `find_stale_references(title, part)` | Same-part citations whose target section doesn't exist in the fetched part. | Surface likely-broken cross-references for a compliance-review workflow. |
| `get_section_history(title, part, section)` | Every amendment version recorded for a section. | Point-in-time questions ("what did this look like before 2020?"). |
Every tool returns a JSON string. Errors come back as `{"error": "..."}` rather
than exceptions, so an LLM caller never sees a raw traceback.
---
## Install
Requires Python 3.11+.
```bash
git clone https://github.com/apraba05/ecfr-cartography-mcp
cd ecfr-cartography-mcp
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
```
Run the offline test suite:
```bash
pytest -q # 18 tests, all pure — no network
pytest -m integration -q # 5 tests, hits the live eCFR API
```
Confirm the server registers all six tools:
```bash
python server.py --list
# search_regulations
# get_section_text
# get_part_structure
# map_part_references
# find_stale_references
# get_section_history
```
---
## Wire it up to Claude Desktop / Claude Code
Add this to your MCP client's config
(`~/Library/Application Support/Claude/claude_desktop_config.json` on macOS, or
`.claude/mcp.json` for Claude Code):
```json
{
"mcpServers": {
"ecfr-cartography": {
"command": "python",
"args": ["/absolute/path/to/ecfr-cartography-mcp/server.py"]
}
}
}
```
Restart the client. The six tools show up under the `ecfr-cartography` server.
---
## Three-minute demo (45 CFR part 170 — ONC health-IT certification)
Copy-paste these into a Claude conversation with the server connected. The
narrative shows an agent walking the graph rather than word-searching PDFs.
**1. Discover the part with search.**
> Tool: `search_regulations`
> Args: `query="electronic health records certification API"`, `per_page=5`
Every result surfaces its title + part. Expect 45 CFR 170 (ONC Health IT
Certification) at the top for this query.
**2. Map the entire part in one call.**
> Tool: `map_part_references`
> Args: `title="45"`, `part="170"`
Returns ~51 sections and ~63 intra-part edges. The `top_referenced` list
identifies **§ 170.315 (Certification criteria)** and **§ 170.299 (Incorporation
by reference)** as the twin hubs — each with in-degree 9. Roughly a third of the
sections are orphans (definitions, purpose statements, one-off subparts).
**3. Ask the model to spot broken cross-references.**
> Tool: `find_stale_references`
> Args: `title="45"`, `part="170"`
Returns one candidate: **§ 170.599 → § 170.503**. Section 170.599
("Incorporation by reference") says "IBR approved for § 170.503" — but 170.503
isn't in the current part.
**4. Prove it with the versioner (the point-in-time story).**
> Tool: `get_section_history`
> Args: `title="45"`, `part="170"`, `section="170.503"`
Response: three versions, the last one on **2020-06-30** with
`"removed": true`. The ISO/IEC standard incorporated by reference in § 170.599
is anchored to a section that was retired six years ago. That's a real
regulatory-graph defect surfaced without a human reading the whole part.
**5. Optional flourish — recover the removed text.**
> Tool: `get_section_text`
> Args: `title="45"`, `part="170"`, `section="170.503"`, `date="2020-05-01"`
Fetches the part as it stood before the removal so the model can quote the old
language back to you.
---
## Honest limitations
- **Recall isn't perfect.** Citation extraction is regex-based. It handles the
common forms exhaustively — `§ 170.404`, `§§ 170.401 through 170.404`,
`45 CFR 164.512`, `42 CFR part 2`, `part 171 of this title`, `42 U.S.C.
300jj-11`, `42 U.S.C. §§ 1301 et seq.` — but obscure phrasings ("the
regulation at Title 45, Section 170.315") will slip through. A statistical
extractor would push recall higher at the cost of the "reproducible, offline
unit-testable" property that made regex the right call for a portfolio piece.
- **`find_stale_references` returns candidates, not conclusions.** A dangling
target may be a reserved section, an intentional cross-title cite, or a
paragraph-level reference the extractor misread. The tool response says so
explicitly. Always verify against the source before acting.
- **Large parts pull a lot of XML.** 40 CFR part 60 is ~13 MB of XML and yields
1,800+ sections and 5,600+ edges. `map_part_references` handles it, but the
D3 visualizer starts to strain — the tool is happier on chunks the size of
45 CFR 170.
- **In-memory caching only.** The httpx client caches within one server
process. There's no on-disk cache; a restart re-fetches.
- **Titles-list resolution is authoritative for dates.** The versioner is
point-in-time. If you don't pass a `date`, we use the title's
`up_to_date_as_of` — which may lag reality by a business day or two.
---
## Next steps
- **Statute → regulation authorization graph.** The `statutory_refs` field
already collects USC citations. Cross-link them against
[uscode.house.gov](https://uscode.house.gov) to build the "which statute
authorizes this rule?" graph — the reverse of what agencies publish.
- **State administrative codes.** Same shape, different sources — Cornell's LII
and per-state SoS APIs. A `search_state_regulations` tool + a
`map_state_part_references` tool would drop straight in beside these six.
- **Diff two point-in-time snapshots.** `get_section_history` already exposes
version dates; a `diff_section(section, date_a, date_b)` tool would surface
what changed between rulemakings.
- **Federal Register cross-linking.** Each amendment date maps back to a
Federal Register notice; adding a `get_amendment_notice` tool would connect
the graph to the rulemaking record.
---
## Files
```
server.py FastMCP server — six tools over stdio
citations.py Pure citation extractor + ReferenceGraph (no HTTP)
visualize.py CLI: python visualize.py 45 170 → graph.html
tests/
test_citations.py 18 offline tests
test_integration.py 5 live-API smoke tests
```
MIT licensed. Contributions welcome.
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues