Skip to main content
Glama
cstebs

clinical-guidance-mcp

by cstebs
README.md
# clinical-guidance-mcp

An MCP server that lets any MCP client (Claude Desktop, Claude Code, an agent you
wrote yourself) search and read 127 FDA guidance documents on clinical research.

The retrieval is not new. It is the hybrid BM25 + dense retriever from
[clinical-rag-evals](https://github.com/cstebs/clinical-rag-evals), configured
exactly as run `008-retrieval-hybrid`, which scored **recall@10 0.918, MRR 0.821**
on that repo's 85-question hand-built gold set. The point of this repo is the
interface. A retriever with measured numbers behind it becomes a tool any model can
call, without anyone copying the index or the corpus into their own app.

## Tools

| Tool | Input | Returns |
|---|---|---|
| `search_guidance` | `query`, `top_k` (1-20, default 10), optional `center` (`CDER` / `CBER`) | Ranked passages with title, center, issue date, section, pages, and the ids `fetch_guidance` accepts |
| `fetch_guidance` | `id`, optional `max_chars` | A document's metadata and section outline, a whole section, or one passage, depending on the id |
| `lookup_guidance_metadata` | `doc_id` | Title, issuing center, issue date, status, docket, PDF and landing-page URLs |

Plus one resource, `guidance://catalog`: every document in the corpus with its
metadata, and the retrieval config the server is running.

All three tools are marked read-only and idempotent, and all return typed
structured output (each has an output schema), so a client gets JSON it can rely
on rather than text it has to parse.

### Ids

One id scheme covers the whole corpus, and every search hit carries all three
levels:

```
fda-70685           document   outline of sections, metadata
fda-70685::s003     section    that section's full text
fda-70685::0012     passage    one chunk, exactly as the retriever returned it
```

Sections and passages include `previous_id` and `next_id`, so a model that finds a
useful passage can widen to its section or read on to the next one without a
second search.

## Example session

From the stdio smoke test, a real subprocess launched the way Claude Desktop
launches it:

```
handshake 1.0 s
['search_guidance', 'fetch_guidance', 'lookup_guidance_metadata']
first search 8.2 s
  1 fda-172258 fda-172258::s003 p 24 | Considerations for the Conduct of Clinical Trials of Medical ...
  2 fda-115172 fda-115172::s006 p 16 | Expansion Cohorts: Use in First-In-Human Clinical Trials to ...
  3 fda-169090 fda-169090::s006 p 24 | E6(R3) Good Clinical Practice (GCP)
meta: {'center': 'CDER', 'issue_date': '09/21/2023', 'status': 'Final', 'docket': 'FDA-2023-D-3550'}
CBER search 0.02 s
```

The first search waits for the index to finish building. After that a search takes
about 20 ms.

## Design decisions

**Reuse the measured retriever, do not rebuild it.** The chunker, index and
manifest are imported from `rag-evals` with run 008's settings (section chunking,
512 tokens, no overlap, hybrid index, no reranking, no query rewriting). At this
chunk size, adding a reranker or query rewriting lowered MRR in that repo's runs,
and rewriting would also add an LLM call to every search, so both stay off. A server that quietly used a different
config would be serving numbers it cannot back up. `tests/test_live_corpus.py`
pins the passage count (4,224) to the eval run so drift shows up as a failing test.

**Outline first, text on request.** Fetching a document id returns its outline, not
its full text. Some of these guidances run past 100 pages, and dumping that into a
model's context to answer one question is the wrong default. The model reads the
outline and asks for the section it needs.

**Sections are contiguous runs, not heading text.** FDA guidances reuse headings
like "A. Background" under different parts. Grouping by heading text would stitch
unrelated passages together, so a section is a run of consecutive chunks under one
heading, and a repeated heading gets its own id.

**The center filter does not change the ranking.** `center` is a post-filter over a
deeper candidate list (8x `top_k`), so results inside the filter are still in run
008 order rather than coming from a new, unmeasured search.

**Errors are for the model to read.** An unknown id returns a tool error that says
what the valid id formats look like and which tool to call next, so the model can
correct itself. A corpus that failed to load returns a tool error pointing at the
setup steps instead of crashing the server.

**Nothing blocks the handshake.** The index takes about 10 seconds to build with a
warm embedding cache. It builds on a background thread started when the server
starts, so clients connect in about a second and only the first tool call waits.
Everything `rag-evals` and `fastembed` print during the build goes to stderr,
because on the stdio transport stdout is the protocol stream.

## Limitations

- **Retrieval quality is the eval's, good and bad.** Recall@10 of 0.918 means about
  1 question in 12 does not get the right passage in the top 10. The server returns
  passages, not answers, so whatever the calling model does with them (and the eval
  found multi-hop synthesis was where answers went wrong, not retrieval) is outside
  what this repo measures.
- **Section detection is heuristic.** Headings come from PDF typography and outline
  patterns. Some documents split cleanly; others come out as a handful of large
  sections. `max_chars` caps what a section fetch returns.
- **A snapshot, not live FDA data.** The corpus is final CDER and CBER guidance on
  clinical research as of the manifest date (2026-08-28). Guidance can be revised or
  withdrawn after that, and the server's instructions tell the model so.
- **3 of 127 documents are metadata only.** Their PDFs yielded no usable text at
  ingest. `lookup_guidance_metadata` still works for them and reports
  `indexed: false`.

## Setup

This depends on `clinical-rag-evals`, which is not on PyPI. Clone both side by side:

```bash
git clone https://github.com/cstebs/clinical-rag-evals.git
git clone https://github.com/cstebs/clinical-guidance-mcp.git

cd clinical-rag-evals
python -m venv .venv && .venv/bin/pip install -e .      # Windows: .venv\Scripts\pip
.venv/bin/rag-evals ingest   # fetches the corpus; slow on purpose, see that repo's README

cd ../clinical-guidance-mcp
python -m venv .venv
.venv/bin/pip install -e ../clinical-rag-evals -e ".[dev]"
.venv/bin/pytest
```

The corpus PDFs are not redistributed. `ingest` fetches them from fda.gov and
honours the crawl delay in FDA's robots.txt, so a full ingest takes over an hour.
The first server start after that embeds the corpus (a few minutes on CPU, no GPU
needed) and caches the vectors, after which startup is about 10 seconds.

If `rag-evals` is installed somewhere other than an editable clone, set
`RAG_EVALS_ROOT` to the directory that holds its `data/` folder.

## Connecting a client

**Claude Code:**

```bash
claude mcp add clinical-guidance -- /path/to/clinical-guidance-mcp/.venv/bin/clinical-guidance-mcp
```

**Claude Desktop:** add the server to `claude_desktop_config.json`. See
`examples/claude_desktop_config.json`, and use the absolute path to the
executable inside this repo's `.venv`.

**Over HTTP**, for clients that connect to a URL instead of launching a process:

```bash
clinical-guidance-mcp --http --port 8000     # serves http://127.0.0.1:8000/mcp
```

It binds to localhost by default and has no authentication, so do not expose it
publicly as is.

## Tests

```bash
pytest
```

- `test_corpus.py`: id scheme, section grouping, truncation, center filter, error
  messages, against a three-document fake corpus.
- `test_server.py`: the same behaviour through a real MCP client session
  (in-process): tool listing, schemas, read-only annotations, structured results,
  argument validation, tool errors, the catalog resource.
- `test_live_corpus.py`: against the real ingested corpus. Pins the corpus to run
  008 and checks that every search hit round-trips through `fetch_guidance`. Skipped
  if the corpus is not ingested.

## Built with

The official MCP Python SDK (`mcp` 2.x, whose `MCPServer` class was called
`FastMCP` in 1.x), pydantic for the output schemas, and `rag-evals` for everything
retrieval-related.

## License

MIT