Skip to main content
Glama

Get Preprint by DOI

biorxiv_get_preprint
Read-onlyIdempotent

Fetch full metadata, abstract, all revision history, JATS XML full-text links, and published-journal DOI for one or more preprints by DOI. Each DOI returns all revisions in one response. When server="both" (default), each DOI is checked against both bioRxiv and medRxiv; the response includes which server the preprint was found on. Failed lookups are reported per-DOI in failed[] rather than aborting the batch, each carrying a reason (not_found, invalid_doi_format, upstream_unavailable, rate_limited) and a retryable flag; a rate_limited entry also carries the wait in seconds the origin asked for. DOIs must match the pattern 10.NNNN/…; a doi.org or article URL, a doi: label, and a trailing vN / .full suffix are stripped first, and results report the bare DOI.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
doisYesOne or more preprint DOIs to look up (max 10).
serverNoServer to query. "both" checks bioRxiv and medRxiv in parallel for each DOI.both

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoPresent when the call failed. Absent on success.
failedNoDOIs that could not be resolved, with per-DOI error details.
preprintsNoSuccessfully resolved preprints with their full revision history.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • addedOutput schema / properties / preprints / items / properties / revisions / items / properties / awards
      Added value: +{
      +  "description": "Grant award numbers from the funding statement, verbatim and deduplicated — one value can hold several grants run together without a separator. Absent when none. Funder names are not included: api.biorxiv.org attributes them to unrelated organizations.",
      +  "items": {
      +    "description": "One award value as upstream records it.",
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • removedOutput schema / properties / preprints / items / properties / revisions / items / properties / funder
      Removed value: -{
      -  "description": "Funder information.",
      -  "type": "string"
      -}
  2. Changed4 schema fields changed
    • changedInput schema / properties / dois / items / description
      Previous value: -"Preprint DOI (e.g. 10.1101/2024.01.15.575123 or 10.64898/2026.05.07.723463)."New value: +"Preprint DOI (e.g. 10.1101/2024.01.15.575123 or 10.64898/2026.05.07.723463). A doi.org or biorxiv.org/medrxiv.org article URL, a doi: prefix, or a vN / .full suffix is accepted and stripped to the bare DOI; every revision is returned either way."
    • changedOutput schema / properties / error / properties / data / properties / reason / description
      Previous value: -"Machine-readable failure mode. Declared by this tool: `doi_not_found`: ALL requested DOIs resolve to empty collections on all requested servers, with every server answering. `invalid_doi_format`: One or more input DOIs do not match the 10.NNNN/ pattern. `upstream_unavailable`: No DOI resolved and at least one lookup failed against api.biorxiv.org, so absence could not be established for any requested DOI. `rate_limited`: No DOI resolved and at least one lookup was rejected with HTTP 429 by api.biorxiv.org. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `doi_not_found`: ALL requested DOIs resolve to empty collections on all requested servers, with every server answering. `invalid_doi_format`: Every input DOI fails to match the 10.NNNN/ pattern, even after URL, doi:, and suffix stripping. `upstream_unavailable`: No DOI resolved and at least one lookup failed against api.biorxiv.org, so absence could not be established for any requested DOI. `rate_limited`: No DOI resolved and at least one lookup was rejected with HTTP 429 by api.biorxiv.org. Other values are possible when a failure originates below the handler."
    • changedOutput schema / properties / failed / items / properties / doi / description
      Previous value: -"DOI that failed to resolve."New value: +"DOI that failed to resolve — in bare form, or exactly as sent for invalid_doi_format."
    • changedOutput schema / properties / preprints / items / properties / doi / description
      Previous value: -"The requested DOI."New value: +"The requested DOI in bare form (any URL, doi: prefix, or suffix removed)."
  3. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds substantial behavioral context beyond annotations: it discloses that each DOI returns all revisions in one response, that server='both' checks both bioRxiv and medRxiv, that failed lookups are reported per-DOI in failed[] with specific reason codes and a retryable flag, and that rate_limited entries carry the wait time. It also explains input normalization (stripping doi.org URLs, doi: labels, vN/.full suffixes). This is rich, non-obvious behavior that an agent needs to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place. It front-loads the core purpose, then covers the multi-revision behavior, the server default, error handling, and input normalization in a logical order. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (batch lookup, multi-revision responses, error handling, input normalization) and the presence of an output schema, the description covers all the behavioral context an agent needs: what is returned, how failures are reported, how inputs are normalized, and what the server parameter does. The output schema presumably documents the response shape, so the description doesn't need to repeat return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters well. The description adds value by explaining the normalization behavior for the dois parameter (stripping URL prefixes, doi: labels, version suffixes) and the meaning of server='both' (parallel check of both repositories). It doesn't add much beyond the schema for the server enum, but the DOI normalization detail is genuinely useful and not fully captured in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Fetch') and a precise resource ('full metadata, abstract, all revision history, JATS XML full-text links, and published-journal DOI for one or more preprints by DOI'). It clearly distinguishes this tool from siblings like biorxiv_get_fulltext (which presumably fetches full text) and biorxiv_get_published_version (which focuses on the published-journal DOI). The scope and output are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: to fetch metadata and revision history by DOI, and it contrasts with the sibling biorxiv_get_fulltext by mentioning JATS XML full-text links are included but the tool is not the full-text fetcher. It also explains the server='both' default behavior and how failed lookups are handled, giving an agent clear context for choosing this tool over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.