FAERS MCP Server
# FAERS MCP Server
Wraps the OpenFDA Drug Adverse Event API (`https://api.fda.gov/drug/event.json`) as an MCP
server with 15 pharmacovigilance tools: case search, disproportionality (ROR/PRR), empirical
Bayes signal scores (MGPS/EBGM), and bulk screening. `faers_disproportionality` and
`faers_ebgm` (and `faers_warm_cache`) accept `stratify_by` for confounder adjustment; the
other tools do not.
## Layout
```
faers/
server.py tool definitions and the faers-mcp entry point
client.py shared openFDA client: Basic-auth key, retry, ceilings
query.py Lucene construction, escaping, the case-sensitivity rules
projection.py compact case cards, per-record suspect verification
stats.py ROR / PRR / chi-square, Mantel-Haenszel, Breslow-Day
screen.py bulk signal screen with the cached global marginal table
ebgm.py DuMouchel gamma-Poisson shrinker (MGPS) against scipy
background.py the drug x event background table the prior is fitted on
strata.py sex / age / year stratifiers and their resolution
fields.py field catalogue served by faers_describe_fields
errors.py structured {code, reason, recovery} failures
tests/ offline tests on recorded fixtures, plus live openFDA contract tests
```
## Installation
From a clone:
```bash
pip install -e .
pip install -e ".[test]" # if you want to run tests
```
Or skip the clone:
```bash
uvx --from git+https://github.com/drvvek/faers_mcp faers-mcp --version
```
The `mcp` dependency is pinned to `<2`: mcp 2.x renamed `FastMCP` to `MCPServer`, and an
unpinned `>=1.0` resolves to 2.x on a fresh install.
## API key
Without a key: 1,000 requests/day and a count `limit` ceiling of 999.
With a key: 120,000 requests/day and `limit` up to 1,000.
Free key at https://open.fda.gov/apis/authentication/
Set `OPENFDA_API_KEY` in your MCP config or the process environment. It is sent as HTTP
Basic auth, never in the query string. Do not commit it; `.gitignore` excludes `.env`.
## Running
Run with the same interpreter you installed into:
```bash
python -m faers # stdio
python -m faers --transport http --port 8010 # Streamable HTTP on 127.0.0.1:8010/mcp
```
`faers-mcp` is the same entry point if that interpreter's Scripts directory is on your PATH.
`FAERS_MCP_TRANSPORT`, `FAERS_MCP_HOST` and `FAERS_MCP_PORT` set the defaults. Binding HTTP
off loopback prints a warning: the transport has no authentication.
## Interface
- **Arguments are flat.** Every tool takes `drug_name`, `events`, ... at the top level of
the arguments object; nothing is wrapped in a `params` key. Enumerated arguments
(`role_basis`, `stratify_by`, `mode`, `sort_by`, `category`) are declared as enums in the
schema, and `date_from`/`date_to` carry a `^\d{8}$` pattern, so invalid values are
rejected before any API call.
- **Results are structured.** Tools return objects; the server publishes an `outputSchema`
and sends the result as `structuredContent` with a JSON text fallback.
- **Failures are errors.** A failed call is an MCP error result (`isError: true`) whose
text is a `{code, reason, recovery}` object. It is never a success payload with
`ok: false`.
- **Progress.** `faers_signal_screen`, `faers_ebgm` and `faers_warm_cache` report progress
to clients that support it.
### Timeouts
The first `faers_ebgm` call for a FAERS release builds a background table: 3 + 100 API calls
and a prior fit, about 100 s in total, which exceeds many MCP clients' per-call timeout.
Call **`faers_warm_cache`** first — it does that work on a turn that expects it and reports
whether each cache was already warm — after which `faers_ebgm` takes a few seconds. Year
stratification adds ~47 calls on its first use, likewise cached.
### Prompt
`faers_signal_workup(drug_name, event=None)` returns a fixed sequence of tool calls for a
defensible work-up — counts, screen, crude and stratified ROR/PRR, EBGM with year
stratification, trend, suspect-verified cases — ending with the disclaimer. Clients that
support MCP prompts can offer it directly.
## Deployment options
Choose whether you spend your own openFDA quota or share one key.
### 1. Git URL + `uvx`
Add this to your MCP config with your own key:
```json
{
"mcpServers": {
"faers": {
"command": "uvx",
"args": ["--from", "git+https://github.com/drvvek/faers_mcp", "faers-mcp"],
"env": { "OPENFDA_API_KEY": "<key>" }
}
}
}
```
No clone and no virtualenv; updates apply on the next launch. `uvx` must be on the PATH
your MCP host uses when it starts the server — otherwise use Local install with an
absolute interpreter. Caches land in `~/.faers_mcp_cache`.
### 2. Shared HTTP endpoint
One process serves Streamable HTTP. Put `OPENFDA_API_KEY` on that process; the server does
not read a `.env` file.
Git Bash:
```bash
OPENFDA_API_KEY=<key> python -m faers --transport http --host 127.0.0.1 --port 8010
```
cmd:
```bat
set OPENFDA_API_KEY=<key>
python -m faers --transport http --host 127.0.0.1 --port 8010
```
PowerShell:
```powershell
$env:OPENFDA_API_KEY="<key>"
python -m faers --transport http --host 127.0.0.1 --port 8010
```
Use the interpreter that has the package. Bind to loopback or a private interface. Your
MCP config then only needs the URL:
```json
{ "mcpServers": { "faers": { "url": "http://127.0.0.1:8010/mcp" } } }
```
The EBGM background, priors and stratified marginals are built once. Everyone on that
endpoint shares one key's 120,000/day quota. The transport has no authentication — keep it
off the public internet, or put it behind a proxy.
| | own quota | install | shared caches | auth |
|---|---|---|---|---|
| git + uvx | yes | none | no | n/a |
| shared HTTP | no — one key | none | yes | none built in |
### Local install
```bash
pip install -e .
```
Put an **absolute path** to that interpreter in your MCP config. On Claude Desktop that
file is `%APPDATA%\Claude\claude_desktop_config.json`. Fully quit and reopen the app after
editing.
```json
{
"mcpServers": {
"faers": {
"command": "C:\\Users\\<you>\\AppData\\Local\\Programs\\Python\\Python314\\python.exe",
"args": ["-m", "faers"],
"env": { "OPENFDA_API_KEY": "<key>" }
}
}
}
```
`Python314` is an example: use the folder for the interpreter you ran `pip install -e .`
with (3.10+). On Hermes Desktop, Settings → Connectors → Local command runs in a remote
sandbox and will not see this install; use the config file, or Shared HTTP if you need a
URL. Claude Desktop and Cursor only need the JSON above.
## Tools
| Tool | Description |
|------|-------------|
| `faers_search_cases` | Case reports as compact triage cards (`full=true` for raw ICSRs) |
| `faers_case_counts` | Total / serious / fatal counts |
| `faers_disproportionality` | 2x2 table, ROR, PRR, chi-square, named criteria; optional Mantel-Haenszel adjustment |
| `faers_count_by_field` | Aggregate by any FAERS field (GROUP BY equivalent) |
| `faers_top_events` | Top MedDRA PTs for a drug |
| `faers_get_report` | Full ICSR by Safety Report ID |
| `faers_demographic_profile` | Sex, age (coded group and onset-age bands, each with coverage), reporter, country |
| `faers_outcome_breakdown` | Reaction outcomes + seriousness criteria |
| `faers_time_trend` | Yearly reporting trend |
| `faers_coreported_drugs` | Substances co-reported with an index drug |
| `faers_signal_screen` | Bulk ROR/PRR screen: a drug against all its events, or vice versa |
| `faers_ebgm` | Empirical Bayes signal scores (EBGM, EB05/EB95) via MGPS; optional stratification |
| `faers_raw_search` | Arbitrary Lucene query passthrough, validated |
| `faers_describe_fields` | Catalogue of searchable field paths, stratifiers and their traps |
| `faers_warm_cache` | Build the EBGM background and fit the prior ahead of time; report cache state |
## Date windows and filters
Every query-shaped tool accepts `date_from` / `date_to` (YYYYMMDD, both or neither) and a
`raw_filter` Lucene clause. On `faers_disproportionality` and `faers_signal_screen` these
are applied to **all four marginals including the grand total N**, so the table stays
internally consistent — which is what makes `raw_filter="patient.patientsex:2"` a genuine
restricted analysis rather than a broken one.
## Bulk screening
`faers_signal_screen` computes ROR, PRR and chi-square for every event reported with a
drug (or every drug reported with an event). The database-wide marginals come from a
single cached count call rather than one call per term:
```
screening 200 events, cold cache 13 API calls, 13.3 s
same, warm cache, another drug 10 API calls, 3.9 s
one-call-per-term equivalent ~202 calls, ~60 s floor
```
The cache is keyed on openFDA's own `meta.last_updated`, so a FAERS refresh invalidates it.
Set `FAERS_CACHE_DIR` to relocate it; it defaults to `~/.faers_mcp_cache`.
## Stratification
Crude ROR/PRR/EBGM compare a drug against the whole database, so anything predicting both
exposure and reporting confounds them. Pass `stratify_by` to adjust:
| stratifier | strata | field | coverage |
|---|---|---|---|
| `sex` | male, female | `patient.patientsex` | 87.9% |
| `age` | 0-17, 18-64, 65+ | `patient.patientonsetage` (years) | 55.3% |
| `year` | calendar year, measured from the data | `receivedate` | ~100% |
| `age_sex` | age band x sex | both | lower |
`patientagegroup` is not used for age: it is populated on only 18.4% of reports. Age
bands are derived from `patientonsetage` instead.
`faers_disproportionality` pools stratum-specific tables by **Mantel-Haenszel** (Robins-
Breslow-Greenland variance for ROR, Greenland-Robins for PRR) and reports crude and
adjusted side by side with a **Breslow-Day** homogeneity test:
```
ATORVASTATIN x RHABDOMYOLYSIS, by age coverage 55.3%
crude ROR 6.273 adjusted ROR 4.888 (-22.1%, material)
Breslow-Day p = 0.0 -> heterogeneous
0-17 a=11 ROR 16.95
18-64 a=963 ROR 5.58
65+ a=1538 ROR 4.51
```
Statins are prescribed predominantly to older patients, who also report more
rhabdomyolysis; adjustment removes that confounding. A small Breslow-Day p indicates that
the stratum-specific estimates differ, in which case the per-stratum rows are more
informative than the pooled value.
Cost is four calls per outer band, because openFDA can count *on* the stratifying field and
the bins are formed locally. Reports missing the field cannot enter any stratum, so every
stratified result states its coverage.
`faers_ebgm` accepts all four. Stratification enters MGPS only through the expected count —
`E = sum_k (n_drug,k * n_event,k) / n_k` — and the prior is refitted on those expectations.
Year strata are not knowable in advance, so they are measured from the data and the
negligible tail is pruned: FAERS spans 38 calendar years, but 1986–2003 hold **385 reports
between them (0.002%)** while costing two calls each to stratify. Years holding less than
0.1% of the window are dropped internally, and more than 30 remaining years is refused.
That cap is not a tool argument — pass a narrower `date_from`/`date_to` window instead.
Year is the only stratifier with full coverage (`receivedate` is present on every report).
On empagliflozin it produces the largest correction:
| event | obs | EBGM crude | EBGM by sex | EBGM by year |
|---|---|---|---|---|
| FOURNIER^S GANGRENE | 1,102 | 140.0 | 124.3 | **87.0** |
| DIABETIC KETOACIDOSIS | 3,902 | 50.7 | 52.9 | 52.6 |
| URINARY TRACT INFECTION | 2,024 | 3.70 | 4.36 | 3.76 |
| NAUSEA | 3,595 | 1.42 | 1.66 | 1.55 |
Empagliflozin was approved in 2014 and Fournier's gangrene reporting spiked after FDA's
2018 safety communication, so both are concentrated in the same recent years. Adjusting for
report year removes that stimulated-reporting effect and drops EBGM by 38%.
## EBGM / MGPS
`faers_ebgm` fits DuMouchel's Gamma-Poisson Shrinker and reports EBGM with an EB05/EB95
credibility interval. EB05 is the conventional screening statistic; EB05 > 2 is the usual
threshold.
Shrinkage is what distinguishes EBGM from ROR and PRR, which treat two cases against an
expectation of 0.02 as a strong signal; EBGM discounts a ratio in proportion to how little
evidence supports it. Teplizumab (29 reports in total) illustrates the difference:
| event | n | expected | RRR | ROR lo95 | EBGM | EB05 |
|---|---|---|---|---|---|---|
| HEPATIC CYTOLYSIS | 2 | 0.017 | 118.3 | **30.2** | 11.0 | **1.5** |
| BLOOD POTASSIUM INCREASED | 2 | 0.022 | 89.1 | 22.7 | 10.1 | 1.5 |
| DEEP VEIN THROMBOSIS | 2 | 0.090 | 22.1 | 5.6 | 5.0 | 1.1 |
An ROR lower confidence bound of 30 on two cases is a false alarm; the EB05 of 1.5 sits
below the screening threshold. On a heavily-reported drug the two agree closely — for empagliflozin's
top 200 events (all n >= 201) EBGM retains 98-100% of the raw ratio and the rankings match.
The prior is fitted across a drug x event background table, not the single pair, so the
first call for a FAERS release builds and caches it (3 + `background_drugs` calls; ~100 s
for the default 100 drugs / 74,000 cells). Later calls take about 6.
**Caveats, repeated in every payload:** expected counts use any-role marginals, the prior
comes from a truncated background, and the likelihood is zero-truncated because openFDA
reports only co-occurring pairs. Unstratified by default — pass `stratify_by` for sex, age
or report year. These values will not reproduce FDA's published EBGMs.
## Reading the output
**Role basis.** openFDA cannot scope a count to one drug's role within a report: the
clause `drugcharacterization:1` filters the *report*, not the matched drug. Counts are
therefore any-role (suspect, concomitant or interacting), and are labelled as such in
`role_basis_note`.
`faers_search_cases` accepts `role_basis="suspect_verified"`, which checks each returned
record individually — the one place the distinction is computable.
**Screening criteria.** There is no single "signal detected" verdict. `criteria_met`
reports each convention separately:
| Criterion | Rule |
|---|---|
| `ema_ror` | ROR lower 95% CI > 1 and a >= 3 |
| `evans_prr` | PRR >= 2 and chi-square >= 4 and a >= 3 |
**Counts are not de-duplicated.** They do not match FAERS Public Dashboard case counts.
These are reporting-rate comparisons, not incidence, and cannot support causal inference.
Every count payload carries this disclaimer.
**Approximate terms.** FAERS stores apostrophes as a caret (`CROHN^S DISEASE`), and
openFDA rejects that character in a search however it is escaped. Such terms fall back to a
tokenised phrase match, which is slightly over-inclusive (measured +0.08% to +1.36%). Rows
affected carry the flag `approximate_marginal`.
**Response size.** ICSRs average ~72 KB; ten raw records measured 724 KB. Case tools
return compact cards by default and declare what was dropped in `fields_omitted`.
## Tests
```bash
pip install -e ".[test]" # pytest is not a runtime dependency
python -m pytest
```
CI runs the offline suite on Python 3.10–3.12 for every push and pull request. The live
contract suite runs on pushes to `master`/`main` only (not on PRs), needs `OPENFDA_API_KEY`
as a repository secret, and is `continue-on-error` so an openFDA behaviour change shows up
as a failed step without turning the workflow red.
Offline tests run against recorded fixtures. Contract tests that hit the live API — they
pin undocumented openFDA behaviour such as `.exact` case sensitivity and the `time` key on
date histograms — are deselected by default:
```bash
python -m pytest -m live
```
## Example queries
- *"What are the top adverse events for empagliflozin in FAERS?"*
- *"Calculate ROR and PRR for metformin hydrochloride and lactic acidosis, adjusted for age"*
- *"Give me a demographic profile of levetiracetam rhabdomyolysis cases"*
- *"Show the yearly trend of pancreatitis reports with sitagliptin"*
- *"Find atorvastatin myopathy cases where atorvastatin is the suspect drug"*
- *"EBGM for empagliflozin's top events, stratified by report year"*
TDQS
Scored across 15 tools
Most tools have clearly distinct purposes, but there is notable overlap among counting/aggregation tools (faers_case_counts, faers_count_by_field, faers_top_events) and signal detection tools (faers_signal_screen vs faers_disproportionality). The detailed descriptions clarify scope (e.g., top_events is a convenience form, signal_screen screens all terms), so an agent can usually tell them apart, but a few could still be confused.
All tools use a consistent faers_ prefix and snake_case, which is predictable. However, the naming pattern is not strictly verb_noun throughout; it mixes verb phrases (search_cases, get_report) with noun phrases (case_counts, demographic_profile) and acronyms (ebgm). This is a minor deviation from the ideal verb_noun consistency.
15 tools is appropriate for a comprehensive FAERS analysis server, covering search, retrieval, aggregation, signal detection, and specialized metrics. Each tool appears to earn its place, with only minor redundancy (e.g., top_events as a convenience wrapper). The count is at the upper end of the ideal 3-15 range but well-scoped.
The tool set covers the full lifecycle of FAERS analysis: search (structured and raw), retrieval, counts, disproportionality, EBGM, demographics, outcomes, trends, and co-reported drugs. Field description and cache warming are also included. No obvious gaps for typical signal detection workflows.