Skip to main content
Glama
musharna

data-aggregator-mcp

by musharna

data-aggregator-mcp

Search 17 research-data sources at once (data archives, omics repositories and papers) and get one list back. Organism, disease and tissue names are expanded with their synonyms in the 11 sources that accept them, records that share a DOI are collapsed to one, and downloads are checked against the source's checksum where it publishes one.

PyPI Python License: MIT CI Glama DOI

Install in VS Code Install in Cursor

mcp-name: io.github.musharna/data-aggregator-mcp

Install

Claude Code:

claude mcp add data-aggregator -- uvx data-aggregator-mcp

VS Code and Cursor: the buttons above. Any other MCP client: run uvx data-aggregator-mcp as a stdio server.

Claude Desktop (claude_desktop_config.json) and most other clients:

{
  "mcpServers": {
    "data-aggregator": {
      "command": "uvx",
      "args": ["data-aggregator-mcp"],
      "env": { "NCBI_API_KEY": "your-optional-key" }
    }
  }
}

With pip:

pip install data-aggregator-mcp
data-aggregator-mcp        # or: python -m data_aggregator_mcp

operate (SQL and previews over remote files) needs an optional extra: pip install "data-aggregator-mcp[operate]".

Related MCP server: Academic MCP

What you can ask

"Find transcriptomes for Orobanche aegyptiaca." The species has been renamed. NCBI Taxonomy maps it to Phelipanche aegyptiaca, and the sources that accept synonyms match both names (the demo above).

"What data came out of this paper?"

resolve("pubmed:40098680")
  links  geo:GSE284240, bioproject:PRJNA1198054, and 18 sra: experiments
  files  PMC11910882.xml   (open-access full text)

"Can I train a model on this dataset?"

resolve("hf:scikit-learn/iris", use="ml-training")
  license_compat  ALLOW: "CC0-1.0 grants the permission(s) required for
                  ml-training: commercial-use, modifications"

The verdict is read from the record's licence metadata. It is advice, not a legal opinion.

resolve also renders citations in any CSL style, checks Crossref for retractions, scores FAIRness, and exports Croissant or RO-Crate. relate reports how a set of records connect (a shared accession, DOI or explicit link).

How a search runs

A source that fails is named in the result's errors, never dropped silently.

Compared with other servers

Server

Searches

One search

Checksum

Remote SQL

Licence check

data-aggregator-mcp

17 sources: data archives, omics, papers

yes, DOI dedup

where published

yes

yes

Mobus

20 general and ML data platforms

yes

—

row preview

yes

paper-search-mcp

papers: arXiv, PubMed, OpenAlex, Crossref and more

yes, deduplicated

—

—

—

ToolUniverse

1000+ tools: models, datasets, APIs, packages

papers

—

—

—

BioMCP

genes, variants, trials, drugs, proteins, papers

yes (search all)

—

—

—

— means the project's README doesn't describe it (READMEs read 2026-10-04). Where they are ahead: BioMCP reaches clinical trials, ChEMBL and AlphaFold, which this server doesn't; ToolUniverse covers far more ground overall; paper-search-mcp covers far more literature sources; Mobus also compares datasets and checks schema compatibility. More detail: docs/POSITIONING.md.

Sources

Source

Discover

Fetch

Checksum

Zenodo

✅

✅

md5

DataCite → Figshare

✅

✅

md5

DataCite → Dataverse

✅

✅

md5

DataCite → OSF

✅

✅

md5

DataCite → Dryad

✅

manifest only¹

sha-256 (listed)

DataCite → Mendeley & others

✅

—

—

NCBI SRA

✅

✅ (ENA FASTQ)

md5

NCBI GEO

✅

✅ (suppl/)

none²

NCBI BioProject

✅

→ SRA links

—

PubMed / OpenAIRE

✅

✅ (OA full text)

none³

Hugging Face datasets

✅

✅ (resolve URL)

none²

DataONE (eco/env)

✅

✅ (Member Node)

md5 / sha-256

OmicsDI → PRIDE

✅

✅ (HTTPS FTP)

none²

OmicsDI → MetaboLights

✅

✅ (HTTPS FTP)

sha-256

OmicsDI → other MS repos

✅

—

—

DataCite → OpenNeuro

✅

✅ (snapshot)

none²

DANDI (neurophysiology)

✅

✅ (302→S3)

sha-256

CZ CELLxGENE (single-cell)

✅

✅ (H5AD/RDS)

none²

OpenML (ML datasets)

✅

✅ (ARFF)

md5

RCSB PDB (structures)

✅

✅ (.cif/.pdb)

none²

UniProtKB (proteins)

✅

✅ (FASTA)

none²

BioStudies (EBI)

✅

✅ (study files)

none²

GBIF (biodiversity)

✅

✅ (Darwin Core)⁴

none²

data.gov (DCAT-US)

✅

✅ (file URL)⁴

none³

NASA CMR (Earth science)

✅

—⁵

—

GWAS Catalog

✅

→ PMID bridge

—

¹ Dryad downloads are token / bot-challenge gated, so fetch returns an error; resolve still lists the files.

² No upstream checksum, so fetch does not verify these bytes. It still returns an error on an HTTP error or when the download exceeds max_bytes.

³ No upstream checksum. Files declared as PDF or XML (literature full text, and data.gov distributions with that mediaType) get an HTML sniff: an HTML login or paywall page served in their place returns an error. Other files are not checked.

⁴ Only records that carry a downloadable file (a GBIF Darwin Core Archive, a data.gov distribution URL); metadata-only records are discovery-only.

⁵ Discovery-only: granule downloads need an Earthdata login, which is not wired. resolve returns the DOI and a data-access portal link.

Reference

Every tool and parameter, the HTTP transport (--transport http), and the environment variables: docs/reference.md.

Develop

uv venv && uv pip install -e ".[dev]"
uv run pytest -q
uv run ruff check src tests
DATA_AGGREGATOR_MCP_LIVE=1 uv run pytest -k live -q   # real-API probes

The README demo (examples/assets/demo.svg) is recorded from live calls by examples/_demo_search.py; its header has the commands to re-record it.

License

MIT — see LICENSE.

Available Tools

6 tools
fetchA

Download a resource's files to local disk and return the PATHS (never the file contents). Fetchable backends: Zenodo (md5-verified); SRA via ENA FASTQ (md5-verified); GEO supplementary files (unverified); DataCite sub-repos — Figshare/Dataverse/OSF (md5-verified), OpenNeuro (snapshot manifest, unverified), Dryad is manifest-only (resolve lists files, fetch fails loud), Mendeley + other DataCite repos fail loud; PubMed/OpenAIRE open-access full text (EuropePMC XML / Unpaywall PDF, unverified); HuggingFace Hub (unverified); DataONE Member-Node objects (md5/SHA-256-verified); OmicsDI — PRIDE (unverified) + MetaboLights (sha-256-verified) only, MassIVE/jPOST/iProX/PeptideAtlas/Panorama Public/GNPS/Metabolomics Workbench fail loud; DANDI dandisets (302→S3, sha-256-verified); CZ CELLxGENE H5AD/RDS assets (unverified); OpenML ARFF (md5-verified); RCSB PDB .cif/.pdb structure files (unverified); UniProtKB FASTA (unverified); BioStudies study files (unverified); GBIF Darwin Core Archives (unverified); data.gov dataset distributions (unverified). Fails loud if selected files exceed max_bytes unless force=true. Verifies md5/SHA-256 where the source publishes one; files marked unverified get no integrity check. Writes a .dataresource.json sidecar.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesSource-prefixed id or bare Zenodo id
destNoDestination dir (default managed cache)
filesNoGlob over file names (default all)
forceNoOverride max_bytes
extractNoUnpack downloaded zip/tar archives into the destination (default false). Path-traversal-guarded; counts against max_bytes.
max_bytesNoByte ceiling before failing loud

Output Schema

ParametersJSON Schema
NameRequiredDescription
bytesNo
pathsNo
resumedNo
skippedNo
unverifiedNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: per-source md5/SHA-256 verification vs unverified status, the fail-loud behavior on max_bytes (unless force), the path-traversal-guarded extract behavior, and the .dataresource.json sidecar write. With readOnlyHint=false and idempotentHint=false declared, this explains the mutation profile in depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and mutation semantics are front-loaded in the first two sentences, and the backend enumeration is information-dense rather than filler. The long run-on source list is harder to scan than a structured list would be, costing a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers every remaining agent-relevant concern: supported sources, integrity verification, failure modes, sidecar output, and archive extraction. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: force is framed as overriding max_bytes specifically, extract 'counts against max_bytes', and the return contract (paths, not contents) contextualizes id/files/dest. It stops short of clarifying dest-vs-files interaction.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+output contract: 'Download a resource's files to local disk and return the PATHS (never the file contents).' That single clause tells an agent exactly what it gets back and distinguishes it from a contents-returning read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives detailed context on which backends are fetchable and which 'fail loud', and explicitly routes listing to the sibling ('Dryad is manifest-only (resolve lists files, fetch fails loud)'). It does not, however, state a general when-to-use-fetch vs search/operate condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sourcesA
Read-only

List wired data sources and their capabilities (layer, kinds, supported filters, auth requirement, rate limit, status). filters_supported names the search parameters a source applies itself: cursor = it pages past page 1; published_after/published_before/kind = filtered upstream, so its total is filtered (any other source is filtered after fetch); organism/disease/tissue/chemical/assay = the synonym expansion reaches its query.

ParametersJSON Schema
NameRequiredDescriptionDefault
check_healthNoWhen true, probe 5 sources (zenodo, datacite, omics, literature, huggingface) and attach a 'health' field ({status: up|down, latency_ms, detail}) to those entries; every other source gets health: null. Default false: returns the static catalog with no network.

Output Schema

ParametersJSON Schema
NameRequiredDescription
sourcesYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, and the description adds real behavioral context beyond that: which filter values are applied upstream by a source versus 'after fetch', and that some sources have auth and rate-limit constraints. It stops short of describing the tool's own failure modes or cost of the network probe, which the schema covers instead.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in a short first sentence; the second sentence is dense but every clause (cursor, date/kind, synonym-expansion filters) earns its place by defining a returned field's values. The semicolon-chained list is slightly run-on but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and the single parameter fully documented in the schema, the description supplies what structured fields cannot: the semantics of filters_supported values. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter and schema description coverage is 100%, so the schema already fully documents check_health, including which five sources are probed and the health field shape. The description adds nothing about it, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List wired data sources') and enumerates exactly what each entry carries (layer, kinds, filters, auth requirement, rate limit, status). No sibling tool (search/operate/resolve/fetch/relate) overlaps with enumerating sources, so the scope is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the discovery step before querying, but the description never says when to call it versus running a search directly, nor when to re-list. The filters_supported explanation helps interpret results but is not a when-to-use instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

operateA
Read-only

Inspect or query a remote tabular file (Parquet/CSV/TSV) WITHOUT downloading it. op='schema' returns columns+types; 'preview' a small sample; 'head' the first n rows; 'sql' a read-only SELECT against the file (exposed as the view 'data', e.g. "SELECT * FROM data WHERE x > 1"). op='peek' profiles every column WITHOUT downloading — type, null-rate, approximate distinct count, min/max, and numeric quartiles (a DuckDB SUMMARIZE; like head/sql it reads the whole file, so it honors the source-size ceiling). Addresses a file by catalog id + file name (resolve the id first to see files[] and access_modes). Requires the [operate] extra; fails loud if the file is not an operable tabular file.

ParametersJSON Schema
NameRequiredDescriptionDefault
nNoRow count for head/preview
idYesDataResource id (e.g. 'zenodo:7654321')
opYes
fileNoFile name within the record; optional when exactly one operable file is present.
queryNoRead-only SELECT for op='sql'.
columnsNoOptional column projection for head.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, and the description adds substantial behavior beyond that: SQL is restricted to read-only SELECT against the view 'data', peek/head/sql read the whole file and therefore honor a source-size ceiling, the tool requires the [operate] extra, and it fails loud on non-operable files. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but every clause carries information; the no-download constraint and op list are front-loaded before the addressing/permission details. Slightly long as a single paragraph, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of describing what each op returns, and covers the addressing model, extra requirement, and failure mode. A couple of minor gaps remain (e.g. default n behavior and pagination/limits on sql results) but the tool is callable correctly as written.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 83%, but the description adds real meaning the schema lacks: it defines each op value's semantics and output (columns+types, small sample, first n rows, read-only SELECT, full profiling with null-rate/distinct/min-max/quartiles). It also explains the SQL view alias 'data' and its example, which the schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Inspect or query a remote tabular file') plus the key differentiator ('WITHOUT downloading it') and enumerates the five operations with their distinct return shapes. An agent can distinguish it from fetch (which downloads) and from search/resolve without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational routing: resolve the id first to see files[] and access_modes, notes the [operate] extra requirement, and notes file is optional when exactly one operable file exists. It stops short of explicitly naming the sibling alternative (e.g. fetch) that would be used when download is desired, so it is clear context but not full when/when-not coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

relateA
Read-only

Given 2-10 resource ids, return metadata-level join/harmonization HINTS: how the datasets relate and on what key they could be joined. Detects shared accessions (BioProject/SRA/GEO), shared cross-identifiers (doi/pmid/pmcid), explicit links between the inputs, and version lineage. HINTS ONLY — it does not read file columns, fetch files, or execute any join/merge/conversion; each hint names the shared value as evidence. Resolve ids first if you only have a search result. Per-id resolve failures are reported, not fatal.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYes2-10 source-prefixed resource ids to relate.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
hintsNo
errorsNo
resolvedYes
input_idsYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond readOnlyHint=true by disclosing that output is hints only, that no file columns are read or files fetched, that each hint cites the shared value as evidence, and that per-id resolve failures are reported rather than fatal. These are exactly the behavioral traits an agent needs to set expectations for a non-authoritative, non-mutating analysis call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and return type, then constraints and failure behavior in descending priority. Dense but almost every clause carries distinct information; the enumeration of detection categories is slightly listy but justified by the tool's discovery purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-format explanation is unnecessary, and the description still covers scope boundaries, prerequisites, evidence semantics, and partial-failure behavior. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the description restates the same facts (2-10 ids, source-prefixed) without adding syntax or format guidance beyond the schema. With a single fully documented parameter, this is the expected baseline rather than added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (relate) and resource (2-10 resource ids) and states precisely what is returned: metadata-level join/harmonization hints, with enumerated detection categories (shared accessions, cross-identifiers, explicit links, version lineage). It is clearly distinguishable from resolve/search/fetch because it explicitly says it does not read columns, fetch files, or execute joins.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit prerequisite ('Resolve ids first if you only have a search result') and a clear when-not-to-use boundary (it does not execute joins/merges, so it is not a substitute for an operate-style action). It does not, however, name the sibling tool an agent should reach for when it actually wants to execute a join, leaving that routing implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolveA
Read-only

Fetch the full DataResource for a known id (e.g. 'zenodo:7654321', 'datacite:10.5061/dryad.x', 'hf:owner/name', a bare Zenodo record id, or a DOI), including the complete files[] manifest. Publication resolve also attaches normalized identifiers (pmid/pmcid/doi) and, when open access, a full-text file. Pass cite= to render a citation onto the result (citation field); omitted means no citation. Pass trust=true to attach retraction status (via Crossref) under trust{}. Pass fair=true to attach an RDA-grounded FAIRness score (0–100 + F/A/I/R sub-scores + actionable gaps) computed from the record under fair{}. Pass use= (commercial/redistribute/modify/ml-training) to attach a licence-compatibility advisory (ALLOW/REVIEW/DENY, not legal advice) under license_compat{}. Pass format=provenance for a one-call RO-Crate 1.1 data-availability dossier (under provenance{}) composing version-currency, licence+SPDX, FAIR score, retraction status, and the source/DOI/ID chain — it auto-attaches fair + trust.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesSource-prefixed id, bare Zenodo id, or DOI
useNoWhen set, attach a licence-compatibility advisory under license_compat{} for an intended use of the record. Supported intents: 'commercial', 'redistribute', 'modify', 'ml-training' (training = a derivative+commercial use, our stated interpretation). The verdict is ALLOW/REVIEW/DENY computed from a bundled choosealicense.com licence matrix keyed on the normalized SPDX id, naming the governing clause — a metadata-derived advisory, NOT legal advice. An unrecognized or absent licence yields REVIEW (never a fabricated ALLOW/DENY); an unknown intent is an error.
citeNoOptional citation format to render onto the result: 'bibtex', 'ris', 'csl-json', or any CSL style name ('apa', 'mla', 'vancouver', ...). DOI-bearing records render via DOI content negotiation; non-DOI records support 'csl-json' only. Omitted = no citation. Failures degrade quietly (citation stays null).
fairNoWhen true, attach an RDA-grounded FAIRness assessment under fair{}: a 0–100 overall score plus findable/accessible/interoperable/reusable sub-scores, the count of indicators evaluated, and actionable gaps each naming its RDA FAIR Data Maturity Model indicator id. Pure/local — no network call. Only the machine-evaluable subset is scored (never fabricates what the metadata cannot show).
trustNoWhen true, attach trust signals (retraction status via Crossref) to the result under trust{}. One extra Crossref call; only meaningful for DOI-bearing records (a DataCite data DOI Crossref does not register leaves retracted=null = unknown, never a false clean claim).
formatNoOptional export to render onto the result. 'croissant' attaches a file-level Croissant JSON-LD manifest (croissant field); 'ro-crate' attaches a minimal RO-Crate 1.1 manifest (ro_crate field); 'provenance' attaches a one-call RO-Crate 1.1 data-availability dossier (provenance field) bundling version-currency, licence+SPDX, FAIR score, retraction status, and the source/DOI/ID chain — it auto-attaches fair{} and trust{} so the dossier is complete in one call (unknown signals are reported as unknown, never as a clean claim).

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
doiNo
fairNo
kindYes
taxaNo
yearNo
filesNo
linksNo
titleYes
trustNo
accessNo
errorsNo
sourceYes
fundingNo
licenseNo
metricsNo
mirrorsNo
citationNo
creatorsNo
organismNo
ro_crateNo
subjectsNo
croissantNo
is_latestNo
truncatedNo
accessionsNo
provenanceNo
descriptionNo
identifiersNo
access_modesNo
last_updatedNo
superseded_byNo
license_compatNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond readOnlyHint=true: it discloses cost/network behavior ('pure/local — no network call' for fair, 'one extra Crossref call' for trust), graceful degradation (citation failures stay null, retraction is null=unknown rather than a false clean claim), and legal/scope caveats (advisory 'NOT legal advice', unrecognized licence yields REVIEW). Unknowns are explicitly reported as unknown, never fabricated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose (fetch + files[] manifest) before enumerating optional flags, one sentence per flag, no filler. It is on the long side and duplicates schema parameter docs, which costs a point but not clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with an output schema and only readOnlyHint annotations, the description carries the behavioral burden well: auto-attach relationships, error/degradation behavior, and advisory framing are all covered. The remaining gap is sibling routing (when to prefer search or fetch instead), which no sentence addresses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mostly restates what the schema already documents (citation formats, fair/trust/use semantics, format enum, provenance auto-attaching fair+trust); its only additive value is the concrete id-format examples that the schema's brief 'Source-prefixed id' note lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — fetch the full DataResource for a known id — and names the distinguishing payload ('the complete files[] manifest'). The id-format examples show exactly what a queryable 'known id' looks like, so an agent can tell this apart from search/fetch without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The precondition is clear: this is for a KNOWN id, and format=provenance is framed as a 'one-call' alternative to assembling the pieces separately. However, it never names the sibling tools (search, fetch, relate) or states the when-not condition (e.g. 'use search when you do not have an id'), leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.61.0
    • Changedsearch10 fields changed
      • changedInput schema / properties / chemical / description
        Previous value: -"Optional chemical/compound name. Resolved via ChEBI (EBI OLS); the query is expanded with the canonical name + exact synonyms (e.g. 'caffeine' also matches '1,3,7-trimethylxanthine'), capped to a bounded number of synonyms. An unknown term yields no expansion; an OLS failure surfaces in errors. The expansion is echoed in chemical_expansion."New value: +"Optional chemical/compound name. Resolved via ChEBI (EBI OLS); the query is expanded with the canonical name + exact synonyms (e.g. 'caffeine' also matches '1,3,7-trimethylxanthine'), capped to a bounded number of synonyms. An unknown term yields no expansion; an OLS failure surfaces in errors. The expansion is echoed in chemical_expansion. A name several terms carry resolves to the one it is the label of, else the one OLS ranks first; the others are listed in chemical_expansion.alternatives."
      • changedInput schema / properties / collapse_mirrors / description
        Previous value: -"Opt into conservative cross-repo content dedup (default false). On top of the always-on exact-DOI dedup, folds records that are the SAME dataset deposited under different (or no) DOIs — e.g. a Zenodo mirror of a figshare deposit, GEO<->ArrayExpress — into one record, annotating the survivor with the folded copies under mirrors[]. Conservative: a merge needs a shared file checksum OR identical (normalized-title, first-author name, year); title-only or partial matches never merge. Intra-page / best-effort only (a mirror on a different page is not collapsed), so a page may return fewer than size items; pagination is unaffected."New value: +"Opt into conservative cross-repo content dedup (default false). On top of the always-on exact-DOI dedup, folds records that are the SAME dataset deposited under different (or no) DOIs — e.g. a Zenodo mirror of a figshare deposit, GEO<->ArrayExpress — into one record, annotating the survivor with the folded copies under mirrors[]. Conservative: a merge needs a shared file checksum OR identical (normalized-title, first-author name, year); title-only or partial matches never merge. Records of one repository never fold into each other: versions and a concept DOI stay separate hits, and a copy elsewhere folds with the latest one. Intra-page / best-effort only (a mirror on a different page is not collapsed), so a page may return fewer than size items; pagination is unaffected."
      • changedInput schema / properties / organism / description
        Previous value: -"Optional organism name. Resolved via NCBI Taxonomy; the query is expanded with the canonical name + synonyms (e.g. 'Orobanche aegyptiaca' also matches 'Phelipanche aegyptiaca'). The expansion is echoed in taxon_expansion."New value: +"Optional organism name. Resolved via NCBI Taxonomy; the query is expanded with the canonical name + synonyms (e.g. 'Orobanche aegyptiaca' also matches 'Phelipanche aegyptiaca'). The expansion is echoed in taxon_expansion. A name matching several taxa resolves to the one it is the scientific name of, else the one with the most nucleotide records ('Drosophila' → the fly genus, not the fungus); the others are listed in taxon_expansion.alternatives."
      • changedInput schema / properties / rank / description
        Previous value: -"Result ordering. 'relevance' (default) = upstream/merged order. 'semantic' re-ranks the fetched page by embedding similarity to the query (needs EMBEDDING_API_BASE; degrades to relevance order with an errors['semantic'] note if unconfigured). In semantic mode pagination is window-based (each page consumes its full fetched window)."New value: +"Result ordering. 'relevance' (default) = hits that name more of the search first (each facet, then each query word, in the title, description, subjects or organism), each source's own order kept among equal hits. 'semantic' re-ranks the fetched page by embedding similarity to the query (needs EMBEDDING_API_BASE; degrades to relevance order with an errors['semantic'] note if unconfigured). In semantic mode pagination is window-based (each page consumes its full fetched window)."
      • changedInput schema / properties / tissue / description
        Previous value: -"Optional tissue/anatomy name. Resolved via UBERON (EBI OLS); the query is expanded with the canonical term + exact synonyms (e.g. 'liver' also matches 'iecur'/'jecur'). The expansion is echoed in tissue_expansion."New value: +"Optional tissue/anatomy name. Resolved via UBERON (EBI OLS); the query is expanded with the canonical term + exact synonyms (e.g. 'liver' also matches 'iecur'/'jecur'). The expansion is echoed in tissue_expansion. A name several terms carry resolves to the one it is the label of, else the one OLS ranks first ('skin' → 'zone of skin', whose exact synonym it is); the others are listed in tissue_expansion.alternatives."
      • addedOutput schema / $defs / ChemicalExpansion / properties / alternatives
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/TermAlternative"
        +  },
        +  "title": "Alternatives",
        +  "type": "array"
        +}
      • addedOutput schema / $defs / TaxonAlternative
        Added value: +{
        +  "description": "Another taxon the organism name matched, not chosen.",
        +  "properties": {
        +    "name": {
        +      "title": "Name",
        +      "type": "string"
        +    },
        +    "taxid": {
        +      "title": "Taxid",
        +      "type": "integer"
        +    }
        +  },
        +  "required": [
        +    "taxid",
        +    "name"
        +  ],
        +  "title": "TaxonAlternative",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / TaxonExpansion / properties / alternatives
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/TaxonAlternative"
        +  },
        +  "title": "Alternatives",
        +  "type": "array"
        +}
      • addedOutput schema / $defs / TermAlternative
        Added value: +{
        +  "description": "Another ontology term the param matched, not chosen.",
        +  "properties": {
        +    "id": {
        +      "title": "Id",
        +      "type": "string"
        +    },
        +    "label": {
        +      "title": "Label",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "id",
        +    "label"
        +  ],
        +  "title": "TermAlternative",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / TissueExpansion / properties / alternatives
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/TermAlternative"
        +  },
        +  "title": "Alternatives",
        +  "type": "array"
        +}
  2. 1 tool updatev0.54.14
    • Changedsearch1 field changed
      • changedInput schema / properties / collapse_mirrors / description
        Previous value: -"Opt into conservative cross-repo content dedup (default false). On top of the always-on exact-DOI dedup, folds records that are the SAME dataset deposited under different (or no) DOIs — e.g. a Zenodo mirror of a figshare deposit, GEO<->ArrayExpress — into one record, annotating the survivor with the folded copies under mirrors[]. Conservative: a merge needs a shared file checksum OR identical (normalized-title, first-author-surname, year); title-only or partial matches never merge. Intra-page / best-effort only (a mirror on a different page is not collapsed), so a page may return fewer than size items; pagination is unaffected."New value: +"Opt into conservative cross-repo content dedup (default false). On top of the always-on exact-DOI dedup, folds records that are the SAME dataset deposited under different (or no) DOIs — e.g. a Zenodo mirror of a figshare deposit, GEO<->ArrayExpress — into one record, annotating the survivor with the folded copies under mirrors[]. Conservative: a merge needs a shared file checksum OR identical (normalized-title, first-author name, year); title-only or partial matches never merge. Intra-page / best-effort only (a mirror on a different page is not collapsed), so a page may return fewer than size items; pagination is unaffected."
  3. 6 tool updatesv0.54.1
    • Changedfetch1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedlist_sources1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedoperate1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedrelate1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedresolve3 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedOutput schema / properties / errors
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "title": "Errors",
        +  "type": "object"
        +}
      • addedOutput schema / properties / truncated
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "title": "Truncated",
        +  "type": "object"
        +}
    • Changedsearch3 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedOutput schema / $defs / DataResource / properties / errors
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "title": "Errors",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / DataResource / properties / truncated
        Added value: +{
        +  "additionalProperties": {
        +    "type": "string"
        +  },
        +  "title": "Truncated",
        +  "type": "object"
        +}
  4. 1 tool updatev0.46.0
    • Changedoperate1 field changed
      • addedInput schema / properties / n / minimum
        Added value: +1
  5. 1 tool updatev0.45.3
    • Changedfetch1 field changed
      • addedOutput schema / properties / unverified
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "title": "Unverified",
        +  "type": "array"
        +}
  6. 2 tool updatesv0.45.1
    • Addedresolve
    • Addedsearch
  7. 3 tool updatesv0.43.0
    • Addedrelate
    • Removedresolve
    • Removedsearch
  8. 3 tool updatesv0.42.0
    • Changedlist_sources1 field changed
      • changedInput schema / properties / check_health / description
        Previous value: -"When true, probe 5 sources (zenodo, datacite, omics, literature, huggingface) and attach a 'health' field ({status: up|down, latency_ms, detail}) to those entries; the remaining 7 sources get health: null. Default false: returns the static catalog with no network."New value: +"When true, probe 5 sources (zenodo, datacite, omics, literature, huggingface) and attach a 'health' field ({status: up|down, latency_ms, detail}) to those entries; every other source gets health: null. Default false: returns the static catalog with no network."
    • Removedrelate
    • Changedsearch4 fields changed
      • changedInput schema / properties / sources / description
        Previous value: -"Restrict fan-out to these sources (default: all). Available: zenodo, dataone, cellxgene, datacite, dandi, omics, literature, huggingface, omicsdi, openml, pdb, uniprot, gwas"New value: +"Restrict fan-out to these sources (default: all). Available: zenodo, dataone, gbif, cellxgene, datacite, dandi, omics, literature, huggingface, datagov, nasacmr, omicsdi, openml, pdb, uniprot, gwas, biostudies"
      • addedOutput schema / $defs / QueryUnderstanding / properties / confidence
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "number"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Confidence"
        +}
      • addedOutput schema / $defs / UnresolvedEntity
        Added value: +{
        +  "description": "Echo of an ontology-typed search param that was supplied but matched NO term,\nso the query ran WITHOUT that expansion.\n\nDistinct from ``SearchResult.errors``: an entry there means the ontology LOOKUP\nFAILED (HTTP/parse) and the two are mutually exclusive per field. An entry here\nmeans the lookup SUCCEEDED and legitimately returned no match — the common case\nfor a common name the registry does not index (NCBI Taxonomy has no ``yeast``,\n``oak`` or ``cedar``; UBERON has no bare ``root``).\n\nWithout this echo the response for a silently-dropped param is byte-identical to\none where the param was never passed, so the caller cannot tell that the filter\nthey asked for was not applied.",
        +  "properties": {
        +    "field": {
        +      "title": "Field",
        +      "type": "string"
        +    },
        +    "input": {
        +      "title": "Input",
        +      "type": "string"
        +    },
        +    "note": {
        +      "title": "Note",
        +      "type": "string"
        +    },
        +    "ontology": {
        +      "title": "Ontology",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "field",
        +    "input",
        +    "ontology",
        +    "note"
        +  ],
        +  "title": "UnresolvedEntity",
        +  "type": "object"
        +}
      • addedOutput schema / properties / unresolved
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/UnresolvedEntity"
        +  },
        +  "title": "Unresolved",
        +  "type": "array"
        +}
  9. 1 tool updatev0.41.1
    • Changedsearch1 field changed
      • changedInput schema / properties / sources / description
        Previous value: -"Restrict fan-out to these sources (default: all). Available: zenodo, dataone, cellxgene, datacite, dandi, omics, literature, huggingface, omicsdi, openml, pdb, gwas"New value: +"Restrict fan-out to these sources (default: all). Available: zenodo, dataone, cellxgene, datacite, dandi, omics, literature, huggingface, omicsdi, openml, pdb, uniprot, gwas"
  10. 5 tool updatesv0.40.0
    • Changedlist_sources1 field changed
      • changedInput schema / properties / check_health / description
        Previous value: -"When true, probe each source's base endpoint and attach a 'health' field ({status: up|down, latency_ms, detail}) to each source. Default false: returns the static catalog with no network."New value: +"When true, probe 5 sources (zenodo, datacite, omics, literature, huggingface) and attach a 'health' field ({status: up|down, latency_ms, detail}) to those entries; the remaining 7 sources get health: null. Default false: returns the static catalog with no network."
    • Changedoperate1 field changed
      • changedInput schema / properties / op / enum
        Previous value: -[
        -  "schema",
        -  "preview",
        -  "head",
        -  "sql"
        -]New value: +[
        +  "schema",
        +  "preview",
        +  "head",
        +  "sql",
        +  "peek"
        +]
    • Addedrelate
    • Changedresolve14 fields changed
      • addedInput schema / properties / fair
        Added value: +{
        +  "description": "When true, attach an RDA-grounded FAIRness assessment under fair{}: a 0–100 overall score plus findable/accessible/interoperable/reusable sub-scores, the count of indicators evaluated, and actionable gaps each naming its RDA FAIR Data Maturity Model indicator id. Pure/local — no network call. Only the machine-evaluable subset is scored (never fabricates what the metadata cannot show).",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / format / description
        Previous value: -"Optional export to render onto the result. 'croissant' attaches a file-level Croissant JSON-LD manifest (croissant field); 'ro-crate' attaches a minimal RO-Crate 1.1 manifest (ro_crate field)."New value: +"Optional export to render onto the result. 'croissant' attaches a file-level Croissant JSON-LD manifest (croissant field); 'ro-crate' attaches a minimal RO-Crate 1.1 manifest (ro_crate field); 'provenance' attaches a one-call RO-Crate 1.1 data-availability dossier (provenance field) bundling version-currency, licence+SPDX, FAIR score, retraction status, and the source/DOI/ID chain — it auto-attaches fair{} and trust{} so the dossier is complete in one call (unknown signals are reported as unknown, never as a clean claim)."
      • changedInput schema / properties / format / enum
        Previous value: -[
        -  "croissant",
        -  "ro-crate"
        -]New value: +[
        +  "croissant",
        +  "ro-crate",
        +  "provenance"
        +]
      • addedInput schema / properties / trust
        Added value: +{
        +  "description": "When true, attach trust signals (retraction status via Crossref) to the result under trust{}. One extra Crossref call; only meaningful for DOI-bearing records (a DataCite data DOI Crossref does not register leaves retracted=null = unknown, never a false clean claim).",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / use
        Added value: +{
        +  "description": "When set, attach a licence-compatibility advisory under license_compat{} for an intended use of the record. Supported intents: 'commercial', 'redistribute', 'modify', 'ml-training' (training = a derivative+commercial use, our stated interpretation). The verdict is ALLOW/REVIEW/DENY computed from a bundled choosealicense.com licence matrix keyed on the normalized SPDX id, naming the governing clause — a metadata-derived advisory, NOT legal advice. An unrecognized or absent licence yields REVIEW (never a fabricated ALLOW/DENY); an unknown intent is an error.",
        +  "type": "string"
        +}
      • addedOutput schema / $defs / FairAssessment
        Added value: +{
        +  "description": "FAIRness assessment attached on resolve(fair=True). PURE-function output:\na 0–100 overall score plus 0–100 per-dimension sub-scores, grounded in the\nmachine-evaluable subset of the RDA FAIR Data Maturity Model. ``assessed`` is\nthe count of indicators actually evaluated (transparency — we never score what\nthe metadata can't show). ``gaps`` are failed-indicator reasons, each naming its\nRDA indicator id and framed as a metadata-exposure gap, not a value judgement.",
        +  "properties": {
        +    "accessible": {
        +      "title": "Accessible",
        +      "type": "integer"
        +    },
        +    "assessed": {
        +      "title": "Assessed",
        +      "type": "integer"
        +    },
        +    "findable": {
        +      "title": "Findable",
        +      "type": "integer"
        +    },
        +    "gaps": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Gaps",
        +      "type": "array"
        +    },
        +    "interoperable": {
        +      "title": "Interoperable",
        +      "type": "integer"
        +    },
        +    "reusable": {
        +      "title": "Reusable",
        +      "type": "integer"
        +    },
        +    "score": {
        +      "title": "Score",
        +      "type": "integer"
        +    }
        +  },
        +  "required": [
        +    "score",
        +    "findable",
        +    "accessible",
        +    "interoperable",
        +    "reusable",
        +    "assessed"
        +  ],
        +  "title": "FairAssessment",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / LicenseVerdict
        Added value: +{
        +  "description": "Licence-compatibility advisory attached on resolve(use=<intent>). PURE-function\noutput: an ALLOW / REVIEW / DENY verdict for an intended use of the resolved record,\ncomputed from a bundled licence matrix (choosealicense.com flag vocabulary) keyed on\nthe normalized SPDX id. ``spdx_id`` is None exactly when the licence was unrecognized\nor absent (→ REVIEW, never a fabricated ALLOW/DENY). ``reason`` names the governing\nclause; ``disclaimer`` states this is a metadata-derived advisory, not legal advice.",
        +  "properties": {
        +    "disclaimer": {
        +      "title": "Disclaimer",
        +      "type": "string"
        +    },
        +    "license_raw": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "title": "License Raw"
        +    },
        +    "reason": {
        +      "title": "Reason",
        +      "type": "string"
        +    },
        +    "spdx_id": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "title": "Spdx Id"
        +    },
        +    "use": {
        +      "title": "Use",
        +      "type": "string"
        +    },
        +    "verdict": {
        +      "enum": [
        +        "ALLOW",
        +        "REVIEW",
        +        "DENY"
        +      ],
        +      "title": "Verdict",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "use",
        +    "verdict",
        +    "spdx_id",
        +    "license_raw",
        +    "reason",
        +    "disclaimer"
        +  ],
        +  "title": "LicenseVerdict",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / Mirror
        Added value: +{
        +  "description": "A same-dataset copy folded into this record by content dedup (resolve the\nmirror's id to reach the original deposit). Only populated when a search ran\nwith the opt-in ``collapse_mirrors`` flag.",
        +  "properties": {
        +    "doi": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Doi"
        +    },
        +    "id": {
        +      "title": "Id",
        +      "type": "string"
        +    },
        +    "source": {
        +      "title": "Source",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "source",
        +    "id"
        +  ],
        +  "title": "Mirror",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / TrustSignals
        Added value: +{
        +  "description": "Integrity/provenance signals attached on resolve(trust=True). All nullable:\nNone = not checked or not determinable (e.g. a DOI Crossref doesn't register) —\nNEVER a negative claim. A *found* Crossref work yields definitive booleans.",
        +  "properties": {
        +    "concern": {
        +      "anyOf": [
        +        {
        +          "type": "boolean"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Concern"
        +    },
        +    "retracted": {
        +      "anyOf": [
        +        {
        +          "type": "boolean"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Retracted"
        +    },
        +    "retraction_doi": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Retraction Doi"
        +    }
        +  },
        +  "title": "TrustSignals",
        +  "type": "object"
        +}
      • addedOutput schema / properties / fair
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/FairAssessment"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / license_compat
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/LicenseVerdict"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / mirrors
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/Mirror"
        +  },
        +  "title": "Mirrors",
        +  "type": "array"
        +}
      • addedOutput schema / properties / provenance
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Provenance"
        +}
      • addedOutput schema / properties / trust
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/TrustSignals"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
    • Changedsearch31 fields changed
      • addedInput schema / properties / assay
        Added value: +{
        +  "description": "Optional assay/method name. Resolved via EDAM topics (EBI OLS); the query is expanded with the canonical name + exact synonyms (e.g. 'ChIP-seq' also matches 'ChIP-sequencing'/'ChIP-exo'). An unknown term yields no expansion; an OLS failure surfaces in errors. The expansion is echoed in assay_expansion.",
        +  "type": "string"
        +}
      • addedInput schema / properties / chemical
        Added value: +{
        +  "description": "Optional chemical/compound name. Resolved via ChEBI (EBI OLS); the query is expanded with the canonical name + exact synonyms (e.g. 'caffeine' also matches '1,3,7-trimethylxanthine'), capped to a bounded number of synonyms. An unknown term yields no expansion; an OLS failure surfaces in errors. The expansion is echoed in chemical_expansion.",
        +  "type": "string"
        +}
      • addedInput schema / properties / collapse_mirrors
        Added value: +{
        +  "default": false,
        +  "description": "Opt into conservative cross-repo content dedup (default false). On top of the always-on exact-DOI dedup, folds records that are the SAME dataset deposited under different (or no) DOIs — e.g. a Zenodo mirror of a figshare deposit, GEO<->ArrayExpress — into one record, annotating the survivor with the folded copies under mirrors[]. Conservative: a merge needs a shared file checksum OR identical (normalized-title, first-author-surname, year); title-only or partial matches never merge. Intra-page / best-effort only (a mirror on a different page is not collapsed), so a page may return fewer than size items; pagination is unaffected.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / disease
        Added value: +{
        +  "description": "Optional disease/phenotype name. Resolved via MeSH (NCBI E-utilities); the query is expanded with the canonical descriptor + entry-term synonyms (e.g. 'breast cancer' also matches 'Breast Neoplasms'). The expansion is echoed in mesh_expansion.",
        +  "type": "string"
        +}
      • addedInput schema / properties / multi_query
        Added value: +{
        +  "default": false,
        +  "description": "Opt into diverse multi-query recall expansion: an LLM generates up to a few deliberately-diverse reformulations of your query, each is fanned out across all sources, and the deduped union is re-ranked against your original query — surfacing relevant records a single keyword query would miss. Costs N× the upstream calls (bounded). Requires an LLM endpoint (LLM_API_BASE); with none configured the search runs as a normal single query and notes it in errors['multi_query']. The variants used are echoed in query_expansion. Composes with understand=. NOTE: multi_query=true ALWAYS applies semantic re-ranking of the window internally regardless of rank=; the rank= param has no effect in this mode.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / provenance
        Added value: +{
        +  "default": false,
        +  "description": "Opt into a whole-search RO-Crate 1.1 Run Crate (default false). Attaches provenance_crate{} — a machine-readable manifest documenting this search: the query, the sources queried, the ontology expansions that fired, the per-source errors (a partial search is disclosed), and per-hit provenance for every result (version-currency, licence + normalized SPDX, FAIR score). Per-hit RETRACTION is omitted — it would need one Crossref call per hit; use per-record resolve(format=provenance) for that. Covers THIS search page only (intra-page; each page of a paginated search gets its own crate).",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / sources / description
        Previous value: -"Restrict fan-out to these sources (default: all). Available: zenodo, datacite, omics, literature, huggingface, dataone, omicsdi"New value: +"Restrict fan-out to these sources (default: all). Available: zenodo, dataone, cellxgene, datacite, dandi, omics, literature, huggingface, omicsdi, openml, pdb, gwas"
      • addedInput schema / properties / tissue
        Added value: +{
        +  "description": "Optional tissue/anatomy name. Resolved via UBERON (EBI OLS); the query is expanded with the canonical term + exact synonyms (e.g. 'liver' also matches 'iecur'/'jecur'). The expansion is echoed in tissue_expansion.",
        +  "type": "string"
        +}
      • addedInput schema / properties / understand
        Added value: +{
        +  "default": false,
        +  "description": "Opt into LLM query understanding: a free-text query is rewritten into a keyword core + structured params (organism/disease/tissue/chemical/assay, kind, year) before fan-out; extracted entities are validated by the same ontology resolvers (a hallucinated entity that doesn't resolve is simply dropped), explicit params you pass always win, and the interpretation is echoed in query_understanding. Requires an LLM endpoint (LLM_API_BASE); with none configured the search runs unchanged and notes it in errors['understand'].",
        +  "type": "boolean"
        +}
      • addedOutput schema / $defs / AssayExpansion
        Added value: +{
        +  "description": "Echo of EDAM assay-synonym expansion that fired for a search (transparency).",
        +  "properties": {
        +    "canonical_name": {
        +      "title": "Canonical Name",
        +      "type": "string"
        +    },
        +    "edam_id": {
        +      "title": "Edam Id",
        +      "type": "string"
        +    },
        +    "input": {
        +      "title": "Input",
        +      "type": "string"
        +    },
        +    "synonyms": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Synonyms",
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "input",
        +    "edam_id",
        +    "canonical_name",
        +    "synonyms"
        +  ],
        +  "title": "AssayExpansion",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / ChemicalExpansion
        Added value: +{
        +  "description": "Echo of ChEBI chemical-synonym expansion that fired for a search (transparency).",
        +  "properties": {
        +    "canonical_name": {
        +      "title": "Canonical Name",
        +      "type": "string"
        +    },
        +    "chebi_id": {
        +      "title": "Chebi Id",
        +      "type": "string"
        +    },
        +    "input": {
        +      "title": "Input",
        +      "type": "string"
        +    },
        +    "synonyms": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Synonyms",
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "input",
        +    "chebi_id",
        +    "canonical_name",
        +    "synonyms"
        +  ],
        +  "title": "ChemicalExpansion",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / DataResource / properties / fair
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/FairAssessment"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / $defs / DataResource / properties / license_compat
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/LicenseVerdict"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / $defs / DataResource / properties / mirrors
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/Mirror"
        +  },
        +  "title": "Mirrors",
        +  "type": "array"
        +}
      • addedOutput schema / $defs / DataResource / properties / provenance
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Provenance"
        +}
      • addedOutput schema / $defs / DataResource / properties / trust
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/TrustSignals"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / $defs / FairAssessment
        Added value: +{
        +  "description": "FAIRness assessment attached on resolve(fair=True). PURE-function output:\na 0–100 overall score plus 0–100 per-dimension sub-scores, grounded in the\nmachine-evaluable subset of the RDA FAIR Data Maturity Model. ``assessed`` is\nthe count of indicators actually evaluated (transparency — we never score what\nthe metadata can't show). ``gaps`` are failed-indicator reasons, each naming its\nRDA indicator id and framed as a metadata-exposure gap, not a value judgement.",
        +  "properties": {
        +    "accessible": {
        +      "title": "Accessible",
        +      "type": "integer"
        +    },
        +    "assessed": {
        +      "title": "Assessed",
        +      "type": "integer"
        +    },
        +    "findable": {
        +      "title": "Findable",
        +      "type": "integer"
        +    },
        +    "gaps": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Gaps",
        +      "type": "array"
        +    },
        +    "interoperable": {
        +      "title": "Interoperable",
        +      "type": "integer"
        +    },
        +    "reusable": {
        +      "title": "Reusable",
        +      "type": "integer"
        +    },
        +    "score": {
        +      "title": "Score",
        +      "type": "integer"
        +    }
        +  },
        +  "required": [
        +    "score",
        +    "findable",
        +    "accessible",
        +    "interoperable",
        +    "reusable",
        +    "assessed"
        +  ],
        +  "title": "FairAssessment",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / LicenseVerdict
        Added value: +{
        +  "description": "Licence-compatibility advisory attached on resolve(use=<intent>). PURE-function\noutput: an ALLOW / REVIEW / DENY verdict for an intended use of the resolved record,\ncomputed from a bundled licence matrix (choosealicense.com flag vocabulary) keyed on\nthe normalized SPDX id. ``spdx_id`` is None exactly when the licence was unrecognized\nor absent (→ REVIEW, never a fabricated ALLOW/DENY). ``reason`` names the governing\nclause; ``disclaimer`` states this is a metadata-derived advisory, not legal advice.",
        +  "properties": {
        +    "disclaimer": {
        +      "title": "Disclaimer",
        +      "type": "string"
        +    },
        +    "license_raw": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "title": "License Raw"
        +    },
        +    "reason": {
        +      "title": "Reason",
        +      "type": "string"
        +    },
        +    "spdx_id": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "title": "Spdx Id"
        +    },
        +    "use": {
        +      "title": "Use",
        +      "type": "string"
        +    },
        +    "verdict": {
        +      "enum": [
        +        "ALLOW",
        +        "REVIEW",
        +        "DENY"
        +      ],
        +      "title": "Verdict",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "use",
        +    "verdict",
        +    "spdx_id",
        +    "license_raw",
        +    "reason",
        +    "disclaimer"
        +  ],
        +  "title": "LicenseVerdict",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / MeshExpansion
        Added value: +{
        +  "description": "Echo of MeSH-synonym expansion that fired for a search (transparency).",
        +  "properties": {
        +    "canonical_name": {
        +      "title": "Canonical Name",
        +      "type": "string"
        +    },
        +    "input": {
        +      "title": "Input",
        +      "type": "string"
        +    },
        +    "mesh_ui": {
        +      "title": "Mesh Ui",
        +      "type": "string"
        +    },
        +    "synonyms": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Synonyms",
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "input",
        +    "mesh_ui",
        +    "canonical_name",
        +    "synonyms"
        +  ],
        +  "title": "MeshExpansion",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / Mirror
        Added value: +{
        +  "description": "A same-dataset copy folded into this record by content dedup (resolve the\nmirror's id to reach the original deposit). Only populated when a search ran\nwith the opt-in ``collapse_mirrors`` flag.",
        +  "properties": {
        +    "doi": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Doi"
        +    },
        +    "id": {
        +      "title": "Id",
        +      "type": "string"
        +    },
        +    "source": {
        +      "title": "Source",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "source",
        +    "id"
        +  ],
        +  "title": "Mirror",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / QueryExpansion
        Added value: +{
        +  "description": "Transparency echo of A2.P2 multi-query recall expansion (search multi_query=true).\n\nWhen enabled and an LLM endpoint is configured, the LLM generates deliberately-diverse\nreformulations of the query; each variant is fanned out across all sources, and the\ndeduped union is re-ranked against the ORIGINAL query. ``variants`` lists the RAW variants\nactually fanned out, the original query first. Each variant received the same ontology\nexpansion (shown by the ``*_expansion`` echoes); results are the deduped union re-ranked\nagainst ``input``.",
        +  "properties": {
        +    "input": {
        +      "title": "Input",
        +      "type": "string"
        +    },
        +    "variants": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Variants",
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "input",
        +    "variants"
        +  ],
        +  "title": "QueryExpansion",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / QueryUnderstanding
        Added value: +{
        +  "description": "Echo of the LLM query-understanding rewrite that fired (transparency, A2.P1).\n\nThe LLM proposes; explicit caller params win; the ontology resolvers then VALIDATE the\nproposed entities. ``applied`` lists the fields the caller left None that were FED into\nthis search as parameters — for ontology entities (organism/disease/tissue/chemical/\nassay) this means \"passed to the resolver\", NOT \"resolved\": whether it actually expanded\nis shown by the corresponding ``*_expansion`` echo (None there ⇒ the entity did not\nresolve, and was never silently treated as a match). ``overridden`` lists fields the LLM\nproposed but the caller had set explicitly (so the LLM's value was ignored).",
        +  "properties": {
        +    "applied": {
        +      "additionalProperties": true,
        +      "title": "Applied",
        +      "type": "object"
        +    },
        +    "extracted": {
        +      "additionalProperties": true,
        +      "title": "Extracted",
        +      "type": "object"
        +    },
        +    "input": {
        +      "title": "Input",
        +      "type": "string"
        +    },
        +    "keyword_core": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "title": "Keyword Core"
        +    },
        +    "overridden": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Overridden",
        +      "type": "array"
        +    }
        +  },
        +  "required": [
        +    "input",
        +    "keyword_core"
        +  ],
        +  "title": "QueryUnderstanding",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / TissueExpansion
        Added value: +{
        +  "description": "Echo of UBERON tissue-synonym expansion that fired for a search (transparency).",
        +  "properties": {
        +    "canonical_name": {
        +      "title": "Canonical Name",
        +      "type": "string"
        +    },
        +    "input": {
        +      "title": "Input",
        +      "type": "string"
        +    },
        +    "synonyms": {
        +      "items": {
        +        "type": "string"
        +      },
        +      "title": "Synonyms",
        +      "type": "array"
        +    },
        +    "uberon_id": {
        +      "title": "Uberon Id",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "input",
        +    "uberon_id",
        +    "canonical_name",
        +    "synonyms"
        +  ],
        +  "title": "TissueExpansion",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / TrustSignals
        Added value: +{
        +  "description": "Integrity/provenance signals attached on resolve(trust=True). All nullable:\nNone = not checked or not determinable (e.g. a DOI Crossref doesn't register) —\nNEVER a negative claim. A *found* Crossref work yields definitive booleans.",
        +  "properties": {
        +    "concern": {
        +      "anyOf": [
        +        {
        +          "type": "boolean"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Concern"
        +    },
        +    "retracted": {
        +      "anyOf": [
        +        {
        +          "type": "boolean"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Retracted"
        +    },
        +    "retraction_doi": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Retraction Doi"
        +    }
        +  },
        +  "title": "TrustSignals",
        +  "type": "object"
        +}
      • addedOutput schema / properties / assay_expansion
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/AssayExpansion"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / chemical_expansion
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/ChemicalExpansion"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / mesh_expansion
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/MeshExpansion"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / provenance_crate
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Provenance Crate"
        +}
      • addedOutput schema / properties / query_expansion
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/QueryExpansion"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / query_understanding
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/QueryUnderstanding"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / tissue_expansion
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/TissueExpansion"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
  11. 3 tool updatesv0.20.0
    • Addedoperate
    • Changedresolve9 fields changed
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "Optional export to render onto the result. 'croissant' attaches a file-level Croissant JSON-LD manifest (croissant field); 'ro-crate' attaches a minimal RO-Crate 1.1 manifest (ro_crate field).",
        +  "enum": [
        +    "croissant",
        +    "ro-crate"
        +  ],
        +  "type": "string"
        +}
      • addedOutput schema / $defs / Metrics
        Added value: +{
        +  "description": "Usage/impact signals, each a separate axis — NO blended score. All\nnullable: a source that does not expose an axis leaves it None.",
        +  "properties": {
        +    "citations": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Citations"
        +    },
        +    "downloads": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Downloads"
        +    },
        +    "likes": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Likes"
        +    },
        +    "views": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Views"
        +    }
        +  },
        +  "title": "Metrics",
        +  "type": "object"
        +}
      • addedOutput schema / properties / access_modes
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "title": "Access Modes",
        +  "type": "array"
        +}
      • addedOutput schema / properties / croissant
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Croissant"
        +}
      • addedOutput schema / properties / is_latest
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "boolean"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Is Latest"
        +}
      • addedOutput schema / properties / last_updated
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Last Updated"
        +}
      • addedOutput schema / properties / metrics
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/Metrics"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / properties / ro_crate
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Ro Crate"
        +}
      • addedOutput schema / properties / superseded_by
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Superseded By"
        +}
    • Changedsearch9 fields changed
      • changedInput schema / properties / sources / description
        Previous value: -"Restrict fan-out to these sources (default: all). Available: zenodo, datacite, omics, literature, huggingface"New value: +"Restrict fan-out to these sources (default: all). Available: zenodo, datacite, omics, literature, huggingface, dataone, omicsdi"
      • addedOutput schema / $defs / DataResource / properties / access_modes
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "title": "Access Modes",
        +  "type": "array"
        +}
      • addedOutput schema / $defs / DataResource / properties / croissant
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Croissant"
        +}
      • addedOutput schema / $defs / DataResource / properties / is_latest
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "boolean"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Is Latest"
        +}
      • addedOutput schema / $defs / DataResource / properties / last_updated
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Last Updated"
        +}
      • addedOutput schema / $defs / DataResource / properties / metrics
        Added value: +{
        +  "anyOf": [
        +    {
        +      "$ref": "#/$defs/Metrics"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}
      • addedOutput schema / $defs / DataResource / properties / ro_crate
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Ro Crate"
        +}
      • addedOutput schema / $defs / DataResource / properties / superseded_by
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Superseded By"
        +}
      • addedOutput schema / $defs / Metrics
        Added value: +{
        +  "description": "Usage/impact signals, each a separate axis — NO blended score. All\nnullable: a source that does not expose an axis leaves it None.",
        +  "properties": {
        +    "citations": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Citations"
        +    },
        +    "downloads": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Downloads"
        +    },
        +    "likes": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Likes"
        +    },
        +    "views": {
        +      "anyOf": [
        +        {
        +          "type": "integer"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Views"
        +    }
        +  },
        +  "title": "Metrics",
        +  "type": "object"
        +}
  12. 4 tool updatesv0.16.0
    • Changedfetch1 field changed
      • addedOutput schema / properties / resumed
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "title": "Resumed",
        +  "type": "array"
        +}
    • Changedlist_sources1 field changed
      • addedInput schema / properties / check_health
        Added value: +{
        +  "default": false,
        +  "description": "When true, probe each source's base endpoint and attach a 'health' field ({status: up|down, latency_ms, detail}) to each source. Default false: returns the static catalog with no network.",
        +  "type": "boolean"
        +}
    • Changedresolve5 fields changed
      • addedOutput schema / $defs / Creator
        Added value: +{
        +  "properties": {
        +    "name": {
        +      "title": "Name",
        +      "type": "string"
        +    },
        +    "orcid": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Orcid"
        +    }
        +  },
        +  "required": [
        +    "name"
        +  ],
        +  "title": "Creator",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / FundingRef
        Added value: +{
        +  "properties": {
        +    "award": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Award"
        +    },
        +    "funder": {
        +      "title": "Funder",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "funder"
        +  ],
        +  "title": "FundingRef",
        +  "type": "object"
        +}
      • addedOutput schema / properties / creators / items / $ref
        Added value: +"#/$defs/Creator"
      • removedOutput schema / properties / creators / items / type
        Removed value: -"string"
      • addedOutput schema / properties / funding
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/FundingRef"
        +  },
        +  "title": "Funding",
        +  "type": "array"
        +}
    • Changedsearch12 fields changed
      • addedInput schema / properties / cursor
        Added value: +{
        +  "description": "Opaque pagination token from a prior search's next_cursor. When set, all other search params are read from the cursor.",
        +  "type": "string"
        +}
      • addedInput schema / properties / kind
        Added value: +{
        +  "description": "Keep only results of this kind.",
        +  "enum": [
        +    "dataset",
        +    "sequencing_run",
        +    "study",
        +    "publication",
        +    "software"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / published_after
        Added value: +{
        +  "description": "Keep results with year >= this.",
        +  "type": "integer"
        +}
      • addedInput schema / properties / published_before
        Added value: +{
        +  "description": "Keep results with year <= this.",
        +  "type": "integer"
        +}
      • addedInput schema / properties / rank
        Added value: +{
        +  "default": "relevance",
        +  "description": "Result ordering. 'relevance' (default) = upstream/merged order. 'semantic' re-ranks the fetched page by embedding similarity to the query (needs EMBEDDING_API_BASE; degrades to relevance order with an errors['semantic'] note if unconfigured). In semantic mode pagination is window-based (each page consumes its full fetched window).",
        +  "enum": [
        +    "relevance",
        +    "semantic"
        +  ],
        +  "type": "string"
        +}
      • changedInput schema / properties / sources / description
        Previous value: -"Restrict fan-out to these sources (default: all). Available: zenodo, datacite, omics, literature"New value: +"Restrict fan-out to these sources (default: all). Available: zenodo, datacite, omics, literature, huggingface"
      • removedInput schema / required
        Removed value: -[
        -  "query"
        -]
      • addedOutput schema / $defs / Creator
        Added value: +{
        +  "properties": {
        +    "name": {
        +      "title": "Name",
        +      "type": "string"
        +    },
        +    "orcid": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Orcid"
        +    }
        +  },
        +  "required": [
        +    "name"
        +  ],
        +  "title": "Creator",
        +  "type": "object"
        +}
      • addedOutput schema / $defs / DataResource / properties / creators / items / $ref
        Added value: +"#/$defs/Creator"
      • removedOutput schema / $defs / DataResource / properties / creators / items / type
        Removed value: -"string"
      • addedOutput schema / $defs / DataResource / properties / funding
        Added value: +{
        +  "items": {
        +    "$ref": "#/$defs/FundingRef"
        +  },
        +  "title": "Funding",
        +  "type": "array"
        +}
      • addedOutput schema / $defs / FundingRef
        Added value: +{
        +  "properties": {
        +    "award": {
        +      "anyOf": [
        +        {
        +          "type": "string"
        +        },
        +        {
        +          "type": "null"
        +        }
        +      ],
        +      "default": null,
        +      "title": "Award"
        +    },
        +    "funder": {
        +      "title": "Funder",
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "funder"
        +  ],
        +  "title": "FundingRef",
        +  "type": "object"
        +}
  13. 4 tool updatesv0.11.0
    • First observedfetch
    • First observedlist_sources
    • First observedresolve
    • First observedsearch

TDQS

A4.3/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct role: search discovers resources, resolve fetches full metadata for known ids, fetch downloads files, operate inspects remote tabular data, list_sources reports source capabilities, and relate returns join hints. The descriptions explicitly cross-reference each other (e.g., 'Use resolve for the full record... then fetch to download files'), making boundaries unambiguous.

Naming Consistency4/5

All names are lowercase snake_case and start with an imperative verb, which is consistent overall. The only minor deviation is list_sources, which is verb_noun, while the other five are bare verbs (search, operate, resolve, fetch, relate).

Tool Count5/5

Six tools is well-scoped for a federated data aggregator; each tool covers a distinct stage of the discovery-to-retrieval workflow and generalizes across dozens of backends. No tool feels redundant or missing at the count level.

Completeness4/5

The surface covers discovery, full-record resolution, file download, remote tabular inspection, source capability listing, and join hints, which is strong lifecycle coverage for a discovery/retrieval server. However, it stops short of executing any actual cross-resource aggregation, join, or merge, and lacks batch resolve/fetch operations, which are minor gaps for a server named data-aggregator.

Maintenance

ActivityActive
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    Enables searching and downloading academic papers from 14 platforms including arXiv, PubMed, Google Scholar, Web of Science, Springer, and Sci-Hub with unified data format and intelligent rate limiting.
    21
    363 npm
    187
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    ▎ Provides 32 tools for plant-genomics locus lookup across 11 free public backends (Ensembl Plants, Phytozome, UniProtKB, Europe PMC, QuickGO, NCBI BLAST, Gramene, KEGG, STRING-DB, ATTED-II, BAR). Takes a TAIR-style locus plus optional organism and returns gene metadata, functional/pathway annotation, interactions, co-expression, and literature — in single-locus, batch, and cross-source synthesis.
    56
    795 PyPI
    7
    MIT