Skip to main content
Glama
algonacci

mcp-crossref

by algonacci

search_crossref

Search 170 million scholarly works across publishers by keyword, author, title, journal, or citation string to find papers and retrieve DOI metadata.

Instructions

Search ~170 million scholarly works registered in Crossref (all publishers: Elsevier, IEEE,
Springer, ACM, MDPI, Indonesian journals, ...). Every result carries a DOI you can pass to the
other crossref tools.

When to use:
    - Broad literature discovery across publishers ("papers on X since 2022").
    - Finding a specific paper from a messy citation string (use `bibliographic`).
    - Listing an author's or a journal's works (use `author` / `orcid` / `issn`).

How matching works (important):
    Crossref has NO exact-phrase search. `query="retrieval augmented generation"` matches any work
    containing ANY of those words, so total_results is inflated and the tail is noise. Rely on the
    top relevance-sorted results, add filters, or use analyze_crossref_topic for phrase-accurate trends.

Args:
    query: Free-text keywords over all metadata, e.g. "graph neural network traffic forecasting".
    author: Author name, e.g. "Geoffrey Hinton". Fuzzy; combine with `orcid` for precision.
    title: Words that should appear in the title.
    bibliographic: A full or partial citation string, e.g.
        "LeCun Bengio Hinton 2015 Deep learning Nature". Best tool for "find this exact paper".
    journal: Journal / proceedings name, e.g. "Expert Systems with Applications".
    publisher: Publisher name, e.g. "IEEE".
    affiliation: Author affiliation text, e.g. "Universitas Indonesia" (only works where deposited).
    funder: Funder name, e.g. "LPDP" or "National Science Foundation".
    from_date: Earliest publication date, "YYYY", "YYYY-MM" or "YYYY-MM-DD".
    until_date: Latest publication date, same formats.
    work_type: Crossref type id, e.g. "journal-article", "proceedings-article", "book-chapter",
        "posted-content" (preprints), "dissertation", "dataset".
    issn: Restrict to one journal by ISSN, e.g. "0957-4174".
    orcid: Restrict to works carrying this ORCID iD, e.g. "0000-0002-1825-0097".
    has_abstract: Only works with a deposited abstract (many publishers do not deposit them).
    has_full_text: Only works with full-text links deposited (does NOT mean open access).
    sort: "relevance" (default), "published", "is-referenced-by-count" (most cited),
        "references-count", "updated", "created".
    order: "desc" (default) or "asc".
    rows: Results to return, 1-100 (default 10).
    offset: Skip this many results for paging (Crossref caps offset at 10,000).
    include_abstract: Add abstracts to results (longer output). Default False to save tokens;
        fetch a single paper's abstract with get_crossref_work instead.

Returns:
    {"total_results": int, "returned_results": int, "items": [work, ...]} where each work has
    doi, title, authors, year, journal, publisher, type, cited_by_count, reference_count, url, ...

Examples:
    search_crossref(query="large language model education", from_date="2023", work_type="journal-article")
    search_crossref(author="Yoshua Bengio", sort="is-referenced-by-count", rows=5)
    search_crossref(bibliographic="Vaswani 2017 Attention is all you need")
    search_crossref(query="deep learning", issn="2169-3536", sort="published")

Note: cited_by_count counts only citations registered in Crossref; it is not a quality measure.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
issnNo
rowsNo
sortNorelevance
orcidNo
orderNodesc
queryNo
titleNo
authorNo
funderNo
offsetNo
journalNo
from_dateNo
publisherNo
work_typeNo
until_dateNo
affiliationNo
has_abstractNo
bibliographicNo
has_full_textNo
include_abstractNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: Crossref has no exact-phrase search, query matches ANY word so total_results is inflated, offset is capped at 10,000, has_full_text does not imply open access, and cited_by_count only counts Crossref-registered citations. These are exactly the non-obvious behaviors an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then usage, then the critical matching caveat, then args/returns/examples — a logical, skimmable structure. The Args section is lengthy, but with 20 undocumented parameters that length is largely earned; only minor trimming would be possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 20 params, no annotations, and a defined output schema, the description still supplies the return shape, paging limits, abstract-token tradeoff, and interpretation caveats. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 20 parameters, so the description must compensate — and it does, documenting every argument with format rules (date formats, offset cap), accepted values (work_type ids, sort keys, order), and worked examples. This is well above the baseline expected when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scale (~170M scholarly works) and explicitly names the sibling it is not ('use analyze_crossref_topic for phrase-accurate trends'). An agent can distinguish it from get_crossref_work, analyze_crossref_topic, and the author/journal tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

A dedicated 'When to use' block gives three concrete scenarios and routes each alternative (bibliographic, author, orcid, issn) to the right tool. It also names the exact alternative for phrase-accurate trends, so when/when-not is fully covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.