Skip to main content
Glama
musharna

data-aggregator-mcp

by musharna
README.md
# πŸ”Ž data-aggregator-mcp

**One MCP server to find and fetch research data across archives, omics
registries, and literature β€” behind a single normalized model.**

[![PyPI](https://img.shields.io/pypi/v/data-aggregator-mcp.svg)](https://pypi.org/project/data-aggregator-mcp/)
[![Python](https://img.shields.io/pypi/pyversions/data-aggregator-mcp.svg)](https://pypi.org/project/data-aggregator-mcp/)
[![Downloads](https://img.shields.io/pypi/dm/data-aggregator-mcp.svg)](https://pypi.org/project/data-aggregator-mcp/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://github.com/musharna/data-aggregator-mcp/blob/main/LICENSE)
[![CI](https://github.com/musharna/data-aggregator-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/musharna/data-aggregator-mcp/actions/workflows/ci.yml)
[![Glama](https://glama.ai/mcp/servers/musharna/data-aggregator-mcp/badges/score.svg)](https://glama.ai/mcp/servers/musharna/data-aggregator-mcp)
[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.21636332.svg)](https://doi.org/10.5281/zenodo.21636332)

`search` one query across **17 sources** β€” **Zenodo, DataCite** (Dryad /
Figshare / Dataverse / OSF / OpenNeuro / Mendeley), **NCBI omics**
(GEO / SRA / BioProject), **BioStudies** (EBI, incl. ArrayExpress),
**literature** (PubMed / OpenAIRE), **HuggingFace** datasets, **DataONE**
(eco / environmental), **OmicsDI** (proteomics / metabolomics), **DANDI**
(neurophysiology), **CZ CELLxGENE** (single-cell), **OpenML** (ML datasets),
**RCSB PDB** (structures), **UniProtKB** (proteins), the **GWAS Catalog**,
**GBIF** (biodiversity), **data.gov** (US federal open data), and **NASA CMR**
(Earth science) β€” deduplicated, normalized, and cross-linked. `resolve` any hit to its file
manifest, citation, trust signals, and the data it points at. `fetch` it to
disk, checksum-verified where the source publishes a checksum.

mcp-name: io.github.musharna/data-aggregator-mcp

<p align="center">
  <img src="https://raw.githubusercontent.com/musharna/data-aggregator-mcp/main/examples/assets/demo.svg"
       alt="data-aggregator-mcp stdio demo β€” initialize, tools/list (search, resolve, fetch, operate, relate, list_sources), and a live list_sources call showing the wired sources across archives, omics, and literature"
       width="820">
</p>

## ✨ Why this

Many research-data MCP servers wrap one source each. This one **unifies** many
behind six tools and one `DataResource` model, so an agent searches once and gets
back comparable records:

- **Multi-domain, one model** β€” generalist archives + raw omics + literature,
  deduplicated by DOI (the fetchable record wins over bare metadata).
- **Taxonomy synonym expansion** β€” `organism="Orobanche aegyptiaca"` also matches
  `Phelipanche aegyptiaca` (NCBI Taxonomy), so a species rename doesn't cost you
  results.
- **Paper β†’ data bridge** β€” resolve a paper and get links to the GEO / SRA /
  BioProject / DataCite records it produced.
- **Checked fetch** β€” streams to disk with md5 / sha-256 verification where the
  source publishes a checksum (a mismatch raises), and optional archive
  unpacking. Many sources publish no checksum (see the Checksum column below);
  those downloads are not verified, and the only content check is an HTML sniff
  on files declared as PDF or XML, which rejects a paywall page served as a
  "PDF".
- **Citations, access & full text** β€” render a citation in any CSL style, get
  normalized access/license, and pull open-access full text β€” all in one
  `resolve`.
- **Trust signals** β€” usage `metrics` (citations / views / downloads / likes),
  version status (`is_latest` / `superseded_by`), and `last_updated` freshness,
  surfaced wherever the source exposes them.
- **Interop exports** β€” `resolve(format="croissant")` or `"ro-crate"` hands a
  dataset to an ML or research-packaging pipeline as standard JSON-LD.
- **Operate on data in place** β€” `operate` reads the schema, previews rows, or
  runs a read-only SQL `SELECT` against a remote Parquet/CSV/TSV **without
  downloading it** (Parquet footer + DuckDB httpfs range reads). Optional
  `[operate]` extra; base install is unchanged.
- **Relate across records** β€” `relate` takes a handful of resolved ids and
  reports how they connect β€” shared accession, shared cross-identifier, an
  explicit link, or version lineage β€” naming the literal shared value as
  evidence. Metadata hints only: it never reads files or executes a join.

β†’ Full rationale and a comparison vs. single-source servers, breadth gateways, and
ML-dataset tools: **[docs/POSITIONING.md](https://github.com/musharna/data-aggregator-mcp/blob/main/docs/POSITIONING.md)**.

<p align="center">
  <img src="https://raw.githubusercontent.com/musharna/data-aggregator-mcp/main/docs/assets/architecture.svg"
       alt="Architecture: an MCP client speaks stdio to data-aggregator-mcp's six tools, which fan out through one router (DOI dedup, ontology expansion, ranking) to archives (Zenodo, DataCite, HuggingFace, DataONE, OpenML, RCSB PDB, UniProtKB, GBIF, data.gov, NASA CMR), omics (GEO, SRA, BioProject, OmicsDI, BioStudies, DANDI, CELLxGENE, GWAS Catalog), and literature (PubMed, OpenAIRE, EuropePMC, Unpaywall)"
       width="760">
</p>

## ⚑ Quickstart

Run with no install:

```bash
uvx data-aggregator-mcp
```

Register with Claude Code:

```bash
claude mcp add data-aggregator -- uvx data-aggregator-mcp
```

A typical agent flow:

```text
search("drought stress RNA-seq", organism="Sorghum bicolor")
  β†’ [ geo:GSE..., sra:SRX..., zenodo:..., pubmed:... ]   # deduped, taxa-normalized

resolve("sra:SRX079566")
  β†’ DataResource{ files: [ENA FASTQ urls…], access: "open", taxa: [...] }

fetch("sra:SRX079566", dest="./data")
  β†’ ["./data/SRX079566_1.fastq.gz", …]                   # md5-verified
```

<details>
<summary>Other ways to run (pip, python -m, raw client config)</summary>

```bash
pip install data-aggregator-mcp
data-aggregator-mcp        # or: python -m data_aggregator_mcp
```

To use the `operate` tool (query remote tabular files in place), install the
optional extra:

```bash
pip install "data-aggregator-mcp[operate]"
```

Add to a client's MCP config (e.g. Claude Desktop `claude_desktop_config.json`):

```json
{
  "mcpServers": {
    "data-aggregator": {
      "command": "uvx",
      "args": ["data-aggregator-mcp"],
      "env": { "NCBI_API_KEY": "your-optional-key" }
    }
  }
}
```

</details>

## 🌐 Transports

**stdio (default)** β€” the server runs as a child of the client, so `fetch()`
writes to your own disk. Nothing to configure; every command above uses it.

**Streamable HTTP** β€” the same six tools, prompts, and resources over HTTP:

```bash
data-aggregator-mcp --transport http     # β†’ http://127.0.0.1:8000/mcp/
```

| flag                       | default          | notes                                                             |
| -------------------------- | ---------------- | ----------------------------------------------------------------- |
| `--transport {stdio,http}` | `stdio`          |                                                                   |
| `--host`                   | `127.0.0.1`      | this machine only; any non-loopback value requires `--allow-host` |
| `--port`                   | `8000`           |                                                                   |
| `--allow-host HOST:PORT`   | auto on loopback | permitted `Host` header, repeatable β€” **required off loopback**   |
| `--allow-origin ORIGIN`    | derived          | permitted browser `Origin` header, repeatable                     |
| `--stateless`              | off              | fresh transport per request, no session affinity                  |
| `--json-response`          | off              | plain JSON responses instead of SSE streams                       |

The endpoint is served at **`/mcp/`** β€” with the trailing slash. `/mcp` answers
`307` redirecting there, which is fine for any client that follows redirects (a
`307` preserves the POST body); point one that doesn't straight at `/mcp/`. In
stateful mode, sessions idle for 30 minutes are reaped.

**DNS-rebinding protection is always on.** A loopback bind derives its own
host/origin allowlist, so the default needs no configuration. A non-loopback bind
(`--host 0.0.0.0`, a LAN address, a container interface) **refuses to start**
without at least one explicit `--allow-host` β€” guessing an allowlist there is
precisely the hole the protection exists to close, so it fails loud instead of
open:

```bash
data-aggregator-mcp --transport http --host 0.0.0.0 \
  --allow-host data.example.org:8000
```

Once running, a request whose `Host` header is outside the allowlist is refused
with `421 Invalid Host header`.

> ⚠️ **`fetch(dest=…)` writes to the _server's_ filesystem, not the client's.**
> Over stdio those are the same disk; over HTTP they may be different machines,
> and the caller gets back paths it cannot read. Treat `dest` on an HTTP
> deployment as server-side staging, or use stdio when you need the bytes
> locally. `search`, `resolve`, `operate`, `relate`, and `list_sources` are
> unaffected β€” they return data, not paths.

## πŸ—‚οΈ Sources

| Source                       | Discover |       Fetch       |     Checksum     |
| ---------------------------- | :------: | :---------------: | :--------------: |
| Zenodo                       |    βœ…    |        βœ…         |       md5        |
| DataCite β†’ Figshare          |    βœ…    |        βœ…         |       md5        |
| DataCite β†’ Dataverse         |    βœ…    |        βœ…         |       md5        |
| DataCite β†’ OSF               |    βœ…    |        βœ…         |       md5        |
| DataCite β†’ Dryad             |    βœ…    |  manifest onlyΒΉ   | sha-256 (listed) |
| DataCite β†’ Mendeley & others |    βœ…    |         β€”         |        β€”         |
| NCBI SRA                     |    βœ…    |  βœ… (ENA FASTQ)   |       md5        |
| NCBI GEO                     |    βœ…    |   βœ… (`suppl/`)   |      noneΒ²       |
| NCBI BioProject              |    βœ…    |    β†’ SRA links    |        β€”         |
| PubMed / OpenAIRE            |    βœ…    | βœ… (OA full text) |      noneΒ³       |
| HuggingFace datasets         |    βœ…    | βœ… (resolve URL)  |      noneΒ²       |
| DataONE (eco/env)            |    βœ…    | βœ… (Member Node)  |  md5 / sha-256   |
| OmicsDI β†’ PRIDE              |    βœ…    |  βœ… (HTTPS FTP)   |      noneΒ²       |
| OmicsDI β†’ MetaboLights       |    βœ…    |  βœ… (HTTPS FTP)   |     sha-256      |
| OmicsDI β†’ other MS repos     |    βœ…    |         β€”         |        β€”         |
| DataCite β†’ OpenNeuro         |    βœ…    |   βœ… (snapshot)   |      noneΒ²       |
| DANDI (neurophysiology)      |    βœ…    |    βœ… (302β†’S3)    |     sha-256      |
| CZ CELLxGENE (single-cell)   |    βœ…    |   βœ… (H5AD/RDS)   |      noneΒ²       |
| OpenML (ML datasets)         |    βœ…    |     βœ… (ARFF)     |       md5        |
| RCSB PDB (structures)        |    βœ…    |  βœ… (.cif/.pdb)   |      noneΒ²       |
| UniProtKB (proteins)         |    βœ…    |    βœ… (FASTA)     |      noneΒ²       |
| BioStudies (EBI)             |    βœ…    | βœ… (study files)  |      noneΒ²       |
| GBIF (biodiversity)          |    βœ…    | βœ… (Darwin Core)⁴ |      noneΒ²       |
| data.gov (DCAT-US)           |    βœ…    |  βœ… (file URL)⁴   |      noneΒ³       |
| NASA CMR (Earth science)     |    βœ…    |        —⁡         |        β€”         |
| GWAS Catalog                 |    βœ…    |   β†’ PMID bridge   |        β€”         |

ΒΉ Dryad downloads are token / bot-challenge gated, so `fetch` fails loud;
`resolve` still lists the files.
Β² No upstream checksum, so `fetch` does not verify these bytes. It still fails
loud on an HTTP error or when the download exceeds `max_bytes`.
Β³ No upstream checksum. Files declared as PDF or XML (literature full text, and
data.gov distributions with that mediaType) get an HTML sniff: an HTML login or
paywall page served in their place fails loud. Other files are not checked.
⁴ Only records that carry a downloadable file (a GBIF Darwin Core Archive, a
data.gov distribution URL); metadata-only records are discovery-only.
⁡ Discovery-only: granule downloads need an Earthdata login, which is not wired.
`resolve` returns the DOI and a data-access portal link.

## πŸ› οΈ Tools

### `search(query?, size?, sources?, organism?, disease?, tissue?, chemical?, assay?, kind?, published_after?, published_before?, rank?, cursor?, collapse_mirrors?, understand?, multi_query?, provenance?)`

Fan out across all wired sources in parallel and return compact `DataResource`
records, deduped by DOI. Per-source failures land in `errors{}` β€” never silently
dropped.

- `organism` β€” expand the query with NCBI-Taxonomy synonyms; the expansion is
  echoed in `taxon_expansion`, and results carry normalized `taxa[]`
  (`{taxid, name}`) plus a `described_in` link to plant-genomics-mcp for plant
  taxa.
- `sources` β€” restrict the fan-out, e.g. `["omics"]`.
- `size` β€” max results (1–50).
- `kind` β€” keep only `dataset` / `sequencing_run` / `study` / `publication` /
  `software`. A record whose upstream type none of these covers (a Zenodo image, a
  DataCite `Audiovisual`, an untyped record) is kind `other` and matches no filter.
- `published_after` / `published_before` β€” filter by publication year.
- `rank` β€” `relevance` (default) or `semantic` (re-rank the fetched page by
  embedding similarity to the query; needs `EMBEDDING_API_BASE`, degrades to
  relevance order otherwise).
- `understand` β€” opt into LLM query understanding (default false). A free-text
  query is **normalized** into a focused keyword query: conversational fluff
  (`"I'm looking for…"`, `"where can I find…"`) is stripped while the scientific
  and entity terms are kept so they still match by text. The LLM also detects
  structured entities (organism/disease/tissue/chemical/assay, kind) β€” these are
  **echoed in `query_understanding.extracted` for transparency but not
  auto-applied**, because ANDing LLM-_inferred_ facets across free-text keyword
  upstreams over-constrains and hurts recall. Only the cleaned `keyword_core` and
  explicit `year` scopes are applied; the ontology resolvers still run on the
  facets **you** pass (the LLM proposes, you dispose). Needs an LLM endpoint
  (`LLM_API_BASE`); with none configured the search runs unchanged and notes it in
  `errors['understand']`. **Effectiveness is query- and model-dependent β€” opt-in /
  default-off; validate the recall lift on your own corpus and LLM (see the eval
  harness below).** `understand=` was measured once (v0.38.0, 2026-06-11) on a
  5-query verified gold set: mean recall@20 lift βˆ’0.10 against the plain query,
  with 4 of the 5 queries neutral or better. `multi_query=` has not been measured.
- `multi_query` β€” opt into diverse multi-query recall expansion (default false).
  An LLM generates up to a few deliberately-diverse reformulations of your query
  (different facets/synonyms/framings, not paraphrases), each is fanned out across
  every source, and the deduped union is re-ranked against your **original** query,
  aiming to reach records a single keyword query would miss. Bounded at
  `MAX_QUERY_VARIANTS` (4, incl. the original), so it costs at most NΓ— the upstream
  calls. The original query's results are always among the candidates, but only the
  top `size` of the re-ranked union are returned, so a result the plain query would
  have returned can be displaced by one from a variant. Composes with
  `understand=` (which structures variant 0). The variants used are echoed in
  `query_expansion`. Needs an LLM endpoint (`LLM_API_BASE`); with none configured
  the search runs as a normal single query and notes it in `errors['multi_query']`.
- `cursor` β€” opaque token from a prior result's `next_cursor`; pages forward
  across every source. In `cursor` mode the other params are read from the
  token, so `query` is optional.

### `resolve(id, cite?, format?, trust?, fair?, use?)`

Full record + files manifest. Routes by id shape β€” `zenodo:7654321`, a bare DOI,
`datacite:10.5061/dryad.x`, an omics id (`sra:SRX079566`, `geo:GSE332789`,
`bioproject:PRJNA1468572`), a literature id (`pubmed:34320281`, `openaire:<id>`),
a HuggingFace id (`hf:owner/name`), a DataONE id (`dataone:doi:10.5063/F1HT2M7Q`),
or an OmicsDI id (`omicsdi:pride:PXD000001`). Attaches, where available:

- **`files[]`** β€” ENA FASTQ manifest (SRA), GEO `suppl/`, or the host repo's
  native manifest (Figshare / Dataverse / OSF / Dryad).
- **`links[]`** β€” paper β†’ data: `pubmed:` β†’ `sra:` / `geo:` / `bioproject:` (NCBI
  elink); `openaire:` β†’ `datacite:` (ScholeXplorer Scholix).
- **`access` / `license`** β€” normalized status
  (`open` / `embargoed` / `restricted` / `closed` / `unknown`) and license where
  the source exposes it.
- **`identifiers`** β€” normalized `{pmid, pmcid, doi}`, plus an open-access
  full-text `FileEntry` (EuropePMC XML, or an Unpaywall PDF fallback) for papers.
- **`citation`** β€” pass `cite=<format>`: `bibtex`, `ris`, `csl-json`, or any CSL
  style name (`apa`, `mla`, `vancouver`, …). DOI records use content
  negotiation; others render CSL-JSON from metadata. Off by default; failures
  degrade quietly.
- **trust signals** β€” `metrics` (citations / views / downloads / likes),
  `is_latest` / `superseded_by` (derived from version links), and `last_updated`
  freshness, where the source provides them.
- **`errors`** β€” `{step: message}` when an enrichment step failed on this record
  (e.g. `taxonomy` during an NCBI rate limit); the rest of the record stands. Such a
  record is not cached, so the next resolve retries the step.
- **`truncated`** β€” `{field: note}` when a list on this record is deliberately partial,
  e.g. a BioProject's `links` past 100 SRA runs: `first 100 of 891 SRA runs; …`. Empty
  when every list is complete.
- **`trust=true`** β€” attach retraction status (via Crossref) under `trust{}`.
  One extra Crossref call; meaningful for DOI-bearing records only.
- **`fair=true`** β€” attach an RDA-grounded FAIRness score (0–100 + F/A/I/R
  sub-scores + actionable gaps) computed from the record metadata under `fair{}`.
  Pure/local β€” no extra network call.
- **`use=<intent>`** β€” attach a licence-compatibility advisory under
  `license_compat{}` for the intended use (`commercial` / `redistribute` /
  `modify` / `ml-training`). Returns ALLOW/REVIEW/DENY with the governing clause.
  Metadata-derived advisory, **not legal advice**; an absent/unrecognized licence
  yields REVIEW.
- **`format`** β€” pass `format="croissant"` (file-level Croissant JSON-LD),
  `"ro-crate"` (minimal RO-Crate 1.1), or `"provenance"` (one-call RO-Crate 1.1
  data-availability dossier bundling version-currency, licence+SPDX, FAIR score,
  and retraction status) to attach a standard manifest under the matching field.

### `fetch(id, dest?, files?, max_bytes?, force?, extract?)`

Download files to disk and return their paths. Streams under a `max_bytes` guard
(`force` to override) with md5 / sha-256 verification wherever the source
publishes a checksum.

- `files` β€” restrict to a subset of the resolved manifest.
- `extract` β€” unpack downloaded zip / tar archives in place, guarded against
  path traversal and runaway extracted size. Off by default.
- Sources without a checksum are downloaded unverified. The one content check
  there is an HTML sniff on files declared as PDF or XML (literature full text,
  some data.gov distributions): it fails loud if the body is actually an HTML page.
- Checksum-verified: **Zenodo**, **SRA** (ENA FASTQ), **DataONE** (Member-Node
  objects), DataCite-hosted **Figshare** / **Dataverse** / **OSF**, **OpenML**
  (ARFF), **MetaboLights** (via OmicsDI; sha-256 from the study's `HASHES/`) and
  **DANDI** (sha-256; an asset whose hash DANDI has not computed yet is unverified).
- Fetchable but unverified: **GEO** `suppl/`, **HuggingFace** datasets,
  **PRIDE** (via OmicsDI), DataCite-hosted **OpenNeuro**,
  **CZ CELLxGENE**, **RCSB PDB**, **UniProtKB**, **BioStudies**,
  **GBIF** (Darwin Core Archives), **data.gov** distributions, and **literature**
  open-access full text.
- **Dryad**, other DataCite repos, other OmicsDI repos (MassIVE / GNPS / ...),
  **BioProject**, **NASA CMR**, and the **GWAS Catalog** are discovery-only and
  raise `FetchNotSupportedError`.

### `list_sources()`

Wired sources with their capabilities β€” layer, kinds, supported filters,
fetchability, `operable` flag, id examples, auth, and rate limits.

### `operate(op, id, file?, query?, n?, columns?)`

Inspect or query a remote tabular file (Parquet / CSV / TSV) **without
downloading it**. Addresses a file by catalog `id` + `file` name (defaults to the
first tabular file on the resolved record). Ops:

- `schema` β€” column names + types (reads the Parquet footer / sniffs the CSV
  header; no full load).
- `preview` β€” a small sample of rows.
- `head` β€” the first `n` rows (default 20), optionally restricted to `columns`.
- `sql` β€” a read-only `SELECT` (the file is the view `data`), e.g.
  `SELECT col, count(*) FROM data GROUP BY 1`.
- `peek` β€” per-column profile via DuckDB `SUMMARIZE` (type, null-rate,
  approximate distinct count, min/max, numeric quartiles) **without
  downloading** the file. Like `head`/`sql`, reads the whole file and honors
  the source-size ceiling.

Backed by the Parquet footer reader + DuckDB `httpfs` range reads. `sql` runs in
a locked-down DuckDB (read-only, local filesystem disabled, single-SELECT
validation, row / wall-clock caps). Requires the optional `[operate]` extra
(`pip install data-aggregator-mcp[operate]`); without it, `operate` returns a
clear install-the-extra message and the other five tools are unaffected.

Any HuggingFace dataset with a datasets-server converted view is operable
(`schema` / `preview` / `head` / `sql`): `resolve` surfaces the auto-converted
Parquet files (`source="hf-datasets-server"`) even for datasets stored as
JSON/JSONL/arrow, so pass `file=<config>/<split>/...parquet` to pick a split when
there are several. A split HF converted only in part (its first 5 GB) is named
`<config>/partial-<split>/...`, so a query on it covers that part, not the whole split.

### `relate(ids)`

Cross-resource join/harmonization **hints**. Given 2–10 resource ids, `relate` resolves
each (TTL-cached) and reports how they relate and on what key they could be joined:

- **`shared_accession`** β€” same BioProject/SRA/GEO accession on β‰₯2 records β†’ joinable key.
- **`shared_identifier`** β€” same doi/pmid/pmcid across records β†’ same work / paper↔data link.
- **`explicit_link`** β€” one record's `links[]` points at another input record.
- **`version_lineage`** β€” one record supersedes another (dedupe, don't join, those).

**Hints only.** `relate` never reads file columns, fetches files, or executes a
join/merge/conversion β€” every hint names the shared value as evidence. Per-id resolve
failures are reported in `errors`, not fatal; an empty result carries an explanatory
`note`.

### Prompts

Three workflow prompts surface in clients (e.g. `/mcp__data_aggregator__*` in
Claude Code):

- **`find_data`** β€” find datasets for a topic, optionally scoped to an organism.
- **`data_behind_paper`** β€” find the datasets / accessions behind a paper.
- **`search_resolve_fetch`** β€” walk the end-to-end search β†’ resolve β†’ fetch flow.

## βš™οΈ Configuration

All optional, set via environment variables:

- `NCBI_API_KEY` β€” raises the NCBI E-utilities rate limit (3 β†’ 10 req/s) used by
  the omics, literature, and taxonomy lookups.
- `DATA_GOV_API_KEY` β€” optional; data.gov works without it through the keyless
  catalog API (`catalog.data.gov`). With a free
  [api.data.gov](https://api.data.gov/signup/) key set, data.gov requests go
  through the api.data.gov gateway instead (1,000 requests/hour per key).
- `UNPAYWALL_EMAIL` β€” enables the Unpaywall fallback leg of literature full-text
  retrieval (the EuropePMC leg works without it).
- `NCBI_EMAIL` β€” contact address sent to NCBI's ID converter; falls back to
  `UNPAYWALL_EMAIL` when unset.
- `DATAVERSE_BASE_URL` β€” resolve Dataverse DOIs against a different installation
  (default `https://dataverse.harvard.edu`).
- `CACHE_TTL_SECONDS` β€” resolve-cache lifetime in seconds (default `3600`; an
  unparseable value falls back to that default).
- `EMBEDDING_API_BASE` / `EMBEDDING_API_KEY` / `EMBEDDING_MODEL` β€” an
  OpenAI-compatible embeddings endpoint enabling `rank=semantic`. Absent β‡’
  semantic re-rank degrades to relevance order. Key is optional (keyless local
  servers supported); model defaults to `text-embedding-3-small`.
- `LLM_API_BASE` / `LLM_API_KEY` / `LLM_MODEL` β€” an OpenAI-compatible
  `/chat/completions` endpoint enabling `search(understand=true)` (NL→structured
  query rewriting) **and** `search(multi_query=true)` (diverse multi-query recall
  expansion). Absent β‡’ both run the raw query unchanged and note it in
  `errors['understand']` / `errors['multi_query']`. Key is optional (keyless local
  servers supported); model defaults to `gpt-4o-mini` (a passthrough string β€” set
  it to whatever your endpoint serves). `multi_query` fans out at most
  `MAX_QUERY_VARIANTS` (4, incl. the original) variants, bounding the NΓ— cost.

To measure the recall lift of `understand=true` / `multi_query=true` on a small
labeled set, run the gated eval harnesses (need a live LLM endpoint):

```bash
DATA_AGGREGATOR_MCP_LIVE=1 LLM_API_BASE=... python scripts/eval_understand.py
DATA_AGGREGATOR_MCP_LIVE=1 LLM_API_BASE=... python scripts/eval_multi_query.py
```

They print per-query and mean recall@20 (understand / multi-query off vs. on). See
the fixtures at `scripts/eval_understand_fixture.json` and
`scripts/eval_multi_query_fixture.json`.

## πŸ§ͺ Develop

```bash
uv venv && uv pip install -e ".[dev]"
uv run pytest -q
uv run ruff check src tests
DATA_AGGREGATOR_MCP_LIVE=1 uv run pytest -k live -q   # real-API probes
```

The README demo (`examples/assets/demo.svg`) is recorded network-free from
`examples/_demo_stdio.py` β€” see the header of that file to re-record.

## License

MIT β€” see [LICENSE](https://github.com/musharna/data-aggregator-mcp/blob/main/LICENSE).

TDQS

A4.3/5.0

Scored across 6 tools

Disambiguation5/5

Each tool occupies a distinct phase of the data lifecycle: list_sources (source metadata), search (fan-out discovery), resolve (full record), operate (remote tabular inspection), fetch (download), and relate (join hints). The descriptions explicitly state boundaries (e.g. resolve returns the full record, fetch returns paths, relate is HINTS ONLY), so an agent can pick the right tool without confusion.

Naming Consistency4/5

All names are lowercase snake_case and verb-oriented (search, resolve, fetch, operate, relate), giving a predictable pattern. The only minor deviation is list_sources, a two-word verb_noun form among single verbs, but it remains readable and consistent in style.

Tool Count5/5

Six tools is well-scoped for a multi-repo data aggregation server; each tool maps to a necessary capability (source listing, search, resolve, fetch, inspect, relate) with no redundant or filler entries. Nothing feels thin or bloated.

Completeness5/5

The surface covers the full workflow from source awareness (list_sources) through discovery (search), record retrieval (resolve), file download (fetch), remote inspection (operate), and cross-dataset linkage (relate). No obvious dead ends for the stated aggregation purpose; relate being hints-only is an explicit design choice rather than a gap.

Maintenance

ActivityActive
ResponsivenessSlow