Skip to main content
Glama
musharna

data-aggregator-mcp

by musharna

fetch

Download research dataset files to local disk, return file paths (not contents), and verify md5/SHA-256 when available.

Instructions

Download a resource's files to local disk and return the PATHS (never the file contents). Fetchable backends: Zenodo (md5-verified); SRA via ENA FASTQ (md5-verified); GEO supplementary files (unverified); DataCite sub-repos — Figshare/Dataverse/OSF (md5-verified), OpenNeuro (snapshot manifest, unverified), Dryad is manifest-only (resolve lists files, fetch fails loud), Mendeley + other DataCite repos fail loud; PubMed/OpenAIRE open-access full text (EuropePMC XML / Unpaywall PDF, unverified); HuggingFace Hub (unverified); DataONE Member-Node objects (md5/SHA-256-verified); OmicsDI — PRIDE (unverified) + MetaboLights (sha-256-verified) only, MassIVE/GNPS/PeptideAtlas/Metabolomics Workbench fail loud; DANDI dandisets (302→S3, sha-256-verified); CZ CELLxGENE H5AD/RDS assets (unverified); OpenML ARFF (md5-verified); RCSB PDB .cif/.pdb structure files (unverified); UniProtKB FASTA (unverified); BioStudies study files (unverified); GBIF Darwin Core Archives (unverified); data.gov dataset distributions (unverified). Fails loud if selected files exceed max_bytes unless force=true. Verifies md5/SHA-256 where the source publishes one; files marked unverified get no integrity check. Writes a .dataresource.json sidecar.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idYesSource-prefixed id or bare Zenodo id
destNoDestination dir (default managed cache)
filesNoGlob over file names (default all)
forceNoOverride max_bytes
extractNoUnpack downloaded zip/tar archives into the destination (default false). Path-traversal-guarded; counts against max_bytes.
max_bytesNoByte ceiling before failing loud

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
bytesNo
pathsNo
resumedNo
skippedNo
unverifiedNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.54.1
    • addedInput schema / additionalProperties
      Added value: +false
  2. Changed1 schema field changedv0.45.3
    • addedOutput schema / properties / unverified
      Added value: +{
      +  "items": {
      +    "type": "string"
      +  },
      +  "title": "Unverified",
      +  "type": "array"
      +}
  3. Changed1 schema field changedv0.16.0
    • addedOutput schema / properties / resumed
      Added value: +{
      +  "items": {
      +    "type": "string"
      +  },
      +  "title": "Resumed",
      +  "type": "array"
      +}
  4. First observedv0.11.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the provided annotations: it discloses per-backend integrity verification (md5/SHA-256 vs unverified), failure conditions ('fails loud if selected files exceed max_bytes unless force=true'), the writing of a .dataresource.json sidecar, and the fact that it returns paths rather than contents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the tool's core action and return behavior, but the long run-on sentence enumerating every backend is dense and could be more structured (e.g., a table or bullet list). It is information-rich but not concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex multi-backend nature of the tool, existing output schema, and safety annotations, the description is complete enough: it covers supported sources, verification behavior, failure modes, side effects (sidecar), and return semantics without needing to re-explain outputs or safety hints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters. The description restates the interaction between max_bytes and force but adds no new semantic detail about the other parameters (id, dest, files, extract), so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Download a resource's files to local disk and return the PATHS (never the file contents).' It also implicitly distinguishes itself from the sibling 'resolve' by noting that Dryad is manifest-only and that resolve lists files while fetch fails loud.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when the tool is applicable by listing all fetchable backends and their verification/failure behavior. It also provides a specific alternative for Dryad ('resolve lists files, fetch fails loud'), but it lacks general routing guidance like 'use search to find an id first' or 'use this instead of resolve when you need the actual files.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools