biodatafinder
# biodatafinder
An MCP server and CLI that searches several public life-science data repositories with one
query and returns a single ranked list of datasets.
Finding reusable data usually means repeating the same search in GEO, ENA, CELLxGENE, PRIDE
and a few generalist repositories, each with its own query syntax, organism filter and
metadata fields. biodatafinder sends one query to all of them concurrently, expands disease,
tissue and cell-type terms through EBI's Ontology Lookup Service (OLS4), maps every result
onto one record shape (accession, title, organism, assay, sample count, access, license, DOI,
link), ranks the merged list, and reports any source that failed instead of hiding it.
| Source | What it covers | API |
| --- | --- | --- |
| `geo` | NCBI GEO series (GSE): expression and other functional genomics | E-utilities (`db=gds`) |
| `ena` | European Nucleotide Archive studies (PRJ*/ERP/SRP) | ENA Portal API |
| `cellxgene` | CZ CELLxGENE Discover single-cell collections | Curation API v1 |
| `pride` | PRIDE Archive mass-spectrometry proteomics (PXD) | PRIDE Archive API v3 |
| `datacite` | Dataset DOIs from Zenodo, Figshare, Dryad, Mendeley Data, MassIVE and others | DataCite REST API |
No API keys are needed. Metadata is taken as each API returns it: when a repository does not
report a sample count, license or access level, the field is left empty rather than guessed.
## Install
Requires Python 3.11+ and [uv](https://docs.astral.sh/uv/).
Run without installing:
```sh
uvx --from git+https://github.com/fmdelgado/biodatafinder biodatafinder search "lung adenocarcinoma"
```
Or from a checkout:
```sh
git clone https://github.com/fmdelgado/biodatafinder
cd biodatafinder
uv sync
uv run biodatafinder search "lung adenocarcinoma"
```
Optional environment variables:
- `NCBI_API_KEY`: raises the NCBI E-utilities limit from 3 to 10 requests/s (GEO).
- `NCBI_EMAIL`: sent to NCBI with each request, as NCBI asks.
## Use with Claude
### Claude Code
```sh
claude mcp add biodatafinder -- uvx --from git+https://github.com/fmdelgado/biodatafinder biodatafinder-mcp
```
### Claude Desktop
Add to `claude_desktop_config.json` (Settings > Developer > Edit Config):
```json
{
"mcpServers": {
"biodatafinder": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/fmdelgado/biodatafinder",
"biodatafinder-mcp"
],
"env": { "NCBI_API_KEY": "optional" }
}
}
}
```
### Skill
[`skills/find-datasets/SKILL.md`](skills/find-datasets/SKILL.md) is a Claude skill that tells
Claude how to use these tools with a scientist: clarify requirements, search broad then
narrow, check top hits with `get_dataset`, and report only datasets the tools returned.
Copy the `skills/find-datasets` folder into `~/.claude/skills/` (Claude Code) or upload it
as a skill in Claude.
## CLI
```text
biodatafinder search TEXT [--organism NAME] [--data-type TYPE ...] [--source NAME ...]
[--min-samples N] [--open-only] [--limit N] [--no-ontology] [--json]
biodatafinder get SOURCE ACCESSION
```
`--limit` is per source. `--json` prints the full response, including the ontology terms the
query was expanded with and any per-source errors. `n=?` means the source did not report a
sample count.
Output observed on 2026-09-24 (results change as repositories grow):
```text
$ biodatafinder search "lung adenocarcinoma single-cell RNA-seq" --organism "Homo sapiens" --limit 3
3.60 [ena] PRJDB5904 n=? Combinatory use of the distinct single cell RNA-seq analytical platforms reveals heterogen
https://www.ebi.ac.uk/ena/browser/view/PRJDB5904
3.60 [ena] PRJNA647741 n=? Long and short-read single cell RNA-seq profiling of human lung adenocarcinoma cell lines
https://www.ebi.ac.uk/ena/browser/view/PRJNA647741
3.43 [ena] PRJNA1471638 n=? Single-Cell RNA Sequencing Identifies Stem-Like Subpopulations and Their Gene Signature in
https://www.ebi.ac.uk/ena/browser/view/PRJNA1471638
2.93 [datacite] 10.7910/dvn/iieaj5 n=? Neuroendocrine-related prognostic risk model and tumor immune environment modulation in lu
https://doi.org/10.7910/dvn/iieaj5
2.85 [cellxgene] 0bebef1a-4607-4584-9070-dacf89a0d635 n=? Defining the cellular and molecular identities of histologic subtypes in lung adenocarcino
https://cellxgene.cziscience.com/collections/0bebef1a-4607-4584-9070-dacf89a0d635
2.60 [geo] GSE335637 n=6 Tumor-derived exosomal METTL3 reprograms macrophages through the m6A–IGF2BP2–ATG2A axis to
https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE335637
2.43 [cellxgene] 6f6d381a-7701-4781-935c-db10d30de293 n=? The integrated Human Lung Cell Atlas
https://cellxgene.cziscience.com/collections/6f6d381a-7701-4781-935c-db10d30de293
...
```
```text
$ biodatafinder search "Alzheimer's disease" --data-type proteomics --limit 5
3.10 [pride] PXD079144 n=? Serine-65 phosphorylated ubiquitin associates with soluble Tau and seeding activity in Alz
https://www.ebi.ac.uk/pride/archive/projects/PXD079144
3.10 [pride] PXD075747 n=? Small Molecule Guided Small Molecule Guided Photocatalytic Proteomics Profiling of Amyloid
https://www.ebi.ac.uk/pride/archive/projects/PXD075747
3.10 [pride] PXD075438 n=? Glycoproteomics of Human and Mouse Brains in Alzheimer’s Disease
https://www.ebi.ac.uk/pride/archive/projects/PXD075438
3.00 [datacite] 10.25934/pr00012874.0 n=? Available datapackage for study 'Donanemab Follow-On Study: Safety, Tolerability, And Effi
https://doi.org/10.25934/pr00012874.0
...
2.00 [datacite] 10.25345/c5bn9xh7v n=? MassIVE MSV000103338 - Alzheimer's Disease Neuroimaging Initiative (ADNI) cerebrospinal fl
https://doi.org/10.25345/c5bn9xh7v
...
```
```text
$ biodatafinder search "gut microbiome metabolomics" --organism "Mus musculus" --limit 3
3.38 [ena] PRJNA1347090 n=? Short-Term Exposure to Environmentally Relevant doses of Chlorpyrifos Impact both the Gut
https://www.ebi.ac.uk/ena/browser/view/PRJNA1347090
...
2.88 [datacite] 10.5281/zenodo.22865761 n=? Multi-omics dataset of gut microbiome, serum metabolome, and whole-blood transcriptome in
https://doi.org/10.5281/zenodo.22865761
2.50 [geo] GSE316703 n=50 Microbiota-Derived Isovalerate Ameliorates Sex-Specific Gut Barrier Dysfunction in Malnutr
https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE316703
...
```
```text
$ biodatafinder get geo GSE2034
{
"source": "geo",
"accession": "GSE2034",
"title": "Breast cancer relapse free survival",
"url": "https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE2034",
"description": "This series represents 180 lymph-node negative relapse free patients and 106 lymph-node negate patients that developed a distant metastasis. ...",
"organisms": ["Homo sapiens"],
"data_type": "transcriptomics",
"assay": "Expression profiling by array",
"tissues": [],
"diseases": [],
"sample_count": 286,
"access": "open",
"license": null,
"publication_date": "2005-02-23",
"doi": null,
"pubmed_ids": ["15721472"]
}
```
Tips: put the organism in `--organism` rather than in the text (each source applies it in
its own way), and keep the text to the topic. The ontology expansion only fires when the
whole text matches an ontology label or synonym, so `"NSCLC"` expands to
*non-small cell lung carcinoma* (MONDO:0005233) while `"NSCLC scRNA-seq"` is searched as
written.
## MCP tools
### `search_datasets`
Searches all registered sources at once and returns a ranked `SearchResponse`.
| Parameter | Type | Default | Meaning |
| --- | --- | --- | --- |
| `text` | string | required | What the data should contain, e.g. `"lung adenocarcinoma scRNA-seq"` |
| `organism` | string | none | Scientific or common name, e.g. `"Homo sapiens"`, `"mouse"` |
| `data_types` | list | all | Any of `transcriptomics`, `single_cell`, `proteomics`, `metabolomics`, `genomics`, `imaging`, `general`; skips sources that cover none of them |
| `sources` | list | all | Restrict to source names from `list_sources` |
| `min_samples` | int | none | Drop records with a known sample count below this; records with unknown count are kept |
| `open_access_only` | bool | false | Drop records marked `controlled` |
| `limit_per_source` | int | 10 | 1 to 100 |
| `expand_ontology` | bool | true | Expand the text and organism through OLS4 |
The response has `results` (records sorted by `score`), `errors` (one entry per source that
failed or timed out after 20 s), `sources_queried`, and `query` (including `expanded_terms`).
Each record has: `source`, `accession`, `title`, `url`, `description`, `organisms`,
`data_type`, `assay`, `tissues`, `diseases`, `sample_count`, `access` (`open`, `controlled`,
`unknown`), `license`, `publication_date`, `doi`, `pubmed_ids`, `score`.
### `get_dataset`
`get_dataset(source, accession)` returns the full record for one accession, or `null` if the
source does not have it. Accepted accessions: GEO `GSE...`; ENA `PRJ...`, `ERP...`, `SRP...`;
CELLxGENE collection UUID; PRIDE `PXD...`; DataCite DOI (bare, `doi:` or `https://doi.org/`).
### `list_sources`
Lists the sources with their description, data types and homepage.
## What each source reports
| Field | geo | ena | cellxgene | pride | datacite |
| --- | --- | --- | --- | --- | --- |
| sample count | yes (samples) | no | no (cell count in description) | no | no |
| organism | yes | yes | yes | yes | no |
| assay | GEO series type | no | yes | experiment type + instrument | no |
| tissues / diseases | no | no | yes | yes | no |
| access | open | open | open | open | from rights statements, else unknown |
| license | no | no | no | detail only (`get_dataset`) | from rights list |
| DOI | no | no | paper DOI | PRIDE dataset DOI | the dataset DOI |
Source-specific behaviour:
- **geo**: esearch then esummary; the API has no relevance sort, so the per-source list is the
newest matching series and ordering comes from biodatafinder's ranker. Requests are spaced
to stay under NCBI's rate limit, with one retry on HTTP 429.
- **ena**: every word (3+ characters) of a search term must appear in the study title or
description; terms are ORed. Common organism names (human, mouse, rat, zebrafish) are mapped
to scientific names; a numeric organism is treated as an NCBI taxon ID (includes subtaxa).
- **cellxgene**: the API has no search endpoint, so the full collection listing (~3 MB) is
downloaded once per process (cached 6 h) and matched locally. The first search takes ~7 s.
- **pride**: keyword search has no OR, so expanded terms are extra requests used only to top
up the results. Words such as "proteomics" are dropped from the keyword because every PRIDE
project is proteomics and the search requires every word. The organism filter is applied
locally on the returned labels.
- **datacite**: DataCite has no organism field, so `organism` does not affect DataCite results. Zenodo/Figshare
version DOIs with the same title and publisher are collapsed to one record.
## Architecture
```text
src/biodatafinder/
models.py SearchQuery, DatasetRecord, SearchResponse (pydantic)
ontology.py OlsExpander: query text and organism -> OLS4 terms + synonyms (best effort)
sources/ one adapter per repository, all subclasses of sources/base.py:Source
registry.py source name -> adapter class
federation.py expand, fan out to sources concurrently (20 s timeout each), collect errors
ranking.py filter (min_samples, open_access_only), dedupe by DOI, score, sort
server.py MCP server (mcp 2.x MCPServer): search_datasets, get_dataset, list_sources
cli.py `biodatafinder search` / `biodatafinder get`
```
Ranking is a transparent heuristic, not a learned model. For each phrasing of the query (the
original text, then ontology labels and synonyms at weight 0.8) it takes the share of query
words found in the title (weight 2) and in the other metadata (weight 1); assay words such as
"single", "cell", "RNA", "seq" count a quarter of a topic word. Small bonuses are added for a
matching organism (+0.5), sample count (up to +0.3) and open access (+0.1). The best phrasing
wins.
## Adding a source
1. Create `src/biodatafinder/sources/<name>.py` with a `Source` subclass:
```python
"""<Repository> via <API name>. Docs: <link to API docs>"""
from biodatafinder.models import Access, DatasetRecord, DataType, SearchQuery
from biodatafinder.sources.base import Source
API = "https://example.org/api"
class ExampleSource(Source):
name = "example"
description = "One line on what the repository holds."
data_types = (DataType.METABOLOMICS,)
homepage = "https://example.org"
async def search(self, query: SearchQuery) -> list[DatasetRecord]:
r = await self.client.get(
f"{API}/search",
params={"q": query.search_terms()[0], "size": query.limit_per_source},
)
r.raise_for_status()
return [self._to_record(hit) for hit in r.json()["hits"]]
async def get(self, accession: str) -> DatasetRecord | None:
r = await self.client.get(f"{API}/studies/{accession}")
if r.status_code == 404:
return None
r.raise_for_status()
return self._to_record(r.json())
def _to_record(self, hit: dict) -> DatasetRecord:
return DatasetRecord(
source=self.name,
accession=hit["id"],
title=hit["title"],
url=f"https://example.org/studies/{hit['id']}",
)
```
Rules: use only `self.client` for HTTP; return at most `query.limit_per_source` records;
apply `query.organism` in the API query when the API supports it; `get()` returns `None`
for not-found and raises on other HTTP errors; leave fields empty when the API does not
provide them (never guess sample counts, access or license).
2. Add one line to `SOURCE_CLASSES` in `registry.py`:
```python
"example": "biodatafinder.sources.example:ExampleSource",
```
3. Add tests in `tests/sources/test_<name>.py`:
- save a few real, trimmed API responses under `tests/fixtures/<name>/` (keep each under 50 KB);
- offline tests with [respx](https://lundberg.github.io/respx/) covering field mapping
(check concrete values), the organism and limit parameters, `get()` found,
`get()` not found returning `None`, and a 5xx raising;
- one `@pytest.mark.live` test with a realistic query against the real API.
Also add the name to the parametrized contract test in `tests/test_core.py`.
## Development
```sh
uv sync
uv run pytest # offline tests (live tests are deselected)
uv run pytest -m live -o addopts="" # live tests against the real APIs
uv run ruff check src tests && uv run ruff format --check src tests
```
## Roadmap
v0.2, more sources:
- ArrayExpress / BioStudies
- SRA
- Human Cell Atlas Data Portal
- Single Cell Portal (Broad)
- MetaboLights
- Metabolomics Workbench
- NCI Genomic Data Commons (GDC)
- GTEx
- gnomAD
- 1000 Genomes
- dbGaP (study metadata only)
- Image Data Resource (IDR)
- EMPIAR
- PDB
Later: an optional locally harvested index for GEO and ENA, so that results can be ranked by
relevance rather than by the order the APIs return them.
Known limitations in v0.1:
- Ontology expansion only matches the whole query text; multi-concept queries are not split.
- GEO, ENA and PRIDE return matches in their own order, not by relevance, so with a small
`limit_per_source` the most relevant older datasets may not be fetched.
- Sample counts are only available from GEO.
## License
MIT, see [LICENSE](LICENSE).
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: search_datasets finds datasets, get_dataset retrieves metadata for a specific one, and list_sources enumerates available sources. No overlap or ambiguity between them.
All tool names follow the same verb_noun snake_case convention (search_datasets, get_dataset, list_sources). The naming pattern is consistent and predictable.
With 3 tools, the server is lean but well-scoped for its purpose of searching and retrieving biodata metadata. Each tool earns its place and the count falls within the ideal 3-15 range.
The tool set covers the core lifecycle of dataset discovery: listing sources, searching across them, and fetching full metadata. There are no obvious gaps or dead ends for the stated domain.