| search_crossrefA | Search ~170 million scholarly works registered in Crossref (all publishers: Elsevier, IEEE,
Springer, ACM, MDPI, Indonesian journals, ...). Every result carries a DOI you can pass to the
other crossref tools.
When to use:
- Broad literature discovery across publishers ("papers on X since 2022").
- Finding a specific paper from a messy citation string (use `bibliographic`).
- Listing an author's or a journal's works (use `author` / `orcid` / `issn`).
How matching works (important):
Crossref has NO exact-phrase search. `query="retrieval augmented generation"` matches any work
containing ANY of those words, so total_results is inflated and the tail is noise. Rely on the
top relevance-sorted results, add filters, or use analyze_crossref_topic for phrase-accurate trends.
Args:
query: Free-text keywords over all metadata, e.g. "graph neural network traffic forecasting".
author: Author name, e.g. "Geoffrey Hinton". Fuzzy; combine with `orcid` for precision.
title: Words that should appear in the title.
bibliographic: A full or partial citation string, e.g.
"LeCun Bengio Hinton 2015 Deep learning Nature". Best tool for "find this exact paper".
journal: Journal / proceedings name, e.g. "Expert Systems with Applications".
publisher: Publisher name, e.g. "IEEE".
affiliation: Author affiliation text, e.g. "Universitas Indonesia" (only works where deposited).
funder: Funder name, e.g. "LPDP" or "National Science Foundation".
from_date: Earliest publication date, "YYYY", "YYYY-MM" or "YYYY-MM-DD".
until_date: Latest publication date, same formats.
work_type: Crossref type id, e.g. "journal-article", "proceedings-article", "book-chapter",
"posted-content" (preprints), "dissertation", "dataset".
issn: Restrict to one journal by ISSN, e.g. "0957-4174".
orcid: Restrict to works carrying this ORCID iD, e.g. "0000-0002-1825-0097".
has_abstract: Only works with a deposited abstract (many publishers do not deposit them).
has_full_text: Only works with full-text links deposited (does NOT mean open access).
sort: "relevance" (default), "published", "is-referenced-by-count" (most cited),
"references-count", "updated", "created".
order: "desc" (default) or "asc".
rows: Results to return, 1-100 (default 10).
offset: Skip this many results for paging (Crossref caps offset at 10,000).
include_abstract: Add abstracts to results (longer output). Default False to save tokens;
fetch a single paper's abstract with get_crossref_work instead.
Returns:
{"total_results": int, "returned_results": int, "items": [work, ...]} where each work has
doi, title, authors, year, journal, publisher, type, cited_by_count, reference_count, url, ...
Examples:
search_crossref(query="large language model education", from_date="2023", work_type="journal-article")
search_crossref(author="Yoshua Bengio", sort="is-referenced-by-count", rows=5)
search_crossref(bibliographic="Vaswani 2017 Attention is all you need")
search_crossref(query="deep learning", issn="2169-3536", sort="published")
Note: cited_by_count counts only citations registered in Crossref; it is not a quality measure.
|
| get_crossref_workA | Get the complete, normalized metadata record for one DOI.
When to use:
- You have a DOI (from any search tool, a PDF, a reference list) and need authors with ORCID
and affiliations, journal, volume/issue/pages, abstract, license, funders, citation counts.
- Checking whether a DOI is real before citing it (a 404 means Crossref does not know it).
Args:
doi: Any DOI spelling: "10.1145/3065386", "https://doi.org/10.1145/3065386",
"doi:10.1145/3065386". Case does not matter.
Returns:
The work record plus "metadata_quality": {"metadata_completeness": 0-1, "missing": [...]},
which tells you which fields the publisher did not deposit (e.g. abstract). It describes the
metadata only, never the scientific quality of the paper.
Tips:
- abstract is often null because many publishers do not deposit abstracts to Crossref;
snowball_doi returns OpenAlex's abstract when it has one.
- arXiv DOIs (10.48550/arXiv.*) are registered with DataCite, not Crossref, so this returns
"not found"; cite_dois still works for them.
- Use get_crossref_references for the reference list.
|
| cite_doisA | Format citations for one or more DOIs in any citation style or export format.
Uses DOI content negotiation, so it works for Crossref AND DataCite DOIs (arXiv, Zenodo, ...).
When to use:
- Building a bibliography / reference list for a thesis or paper.
- Exporting references to Zotero, Mendeley or EndNote (bibtex / ris).
Args:
dois: List of DOIs (max 50), e.g. ["10.1038/nature14539", "10.1145/3065386"].
style: Citation style or export format:
- "apa" (default), "ieee", "vancouver", "chicago", "mla", "harvard", "nature"
- any other CSL style id from https://github.com/citation-style-language/styles,
e.g. "american-medical-association"
- "bibtex", "ris" or "csl-json" for reference-manager exports
Returns:
{"style": str, "citations": [{"doi": str, "citation": str}], "failed": [{"doi", "error"}]}
Example:
cite_dois(["10.1038/nature14539"], style="ieee")
-> "Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, ..."
|
| resolve_doisA | Find every DOI mentioned in free text and resolve each one to clean metadata.
When to use:
- The user pastes a messy reference list, a PDF's text, notes or a URL list and wants to know
what the papers are, check that the DOIs exist, or turn them into a clean table.
- Verifying DOIs produced by another tool or model before citing them.
Args:
text: Any text containing DOIs in any form ("doi:10.x/y", "https://doi.org/10.x/y", bare).
max_dois: Resolve at most this many unique DOIs (default 25, max 100).
Returns:
{"found": int, "works": [compact work], "failed": [{"doi", "error"}]}. A DOI in "failed"
with "Not found" is unknown to Crossref (typo, fabricated, or registered with DataCite).
Follow-up: pass the resolved DOIs to cite_dois to format a bibliography.
|
| get_crossref_referencesA | List the references (bibliography) of a paper, i.e. the older works it cites.
When to use:
- Backward snowballing in a literature review: find the foundational papers a key paper builds on.
- Checking which sources a paper relies on.
Args:
doi: DOI of the citing paper.
resolve: Also fetch full metadata for the first N references that have a DOI (max 20),
sorted by citation count, to spot the most influential ones. Default 0 (no extra calls).
Returns:
{"doi", "title", "reference_count", "references": [{doi, title, author, year, journal,
unstructured}], "resolved": [compact work]}
Notes:
- Only references the publisher deposited are available; some publishers deposit none.
- Crossref does not expose the reverse direction (papers that cite this one); use snowball_doi.
|
| find_related_worksA | Find works bibliographically similar to a given paper (same title vocabulary and subjects).
When to use:
- "More like this" from one good paper, without needing its references or citations.
Args:
doi: DOI of the seed paper.
rows: Number of similar works to return (1-50, default 10).
Returns:
{"seed": compact work, "related": [compact work + relevance_score]}
Tip: snowball_doi gives citation-based neighbours (references, citing works, OpenAlex related),
which are usually more meaningful than text similarity.
|
| analyze_crossref_topicA | Describe a research topic: publications per year, top venues, publishers and funders, and the
most cited works, counting only titles that contain the exact phrase.
When to use:
- Trend questions: "is X growing?", "when did X take off?", thesis/proposal background.
- "Where is X published?", "who funds X?" (venue and funder landscape).
How it works:
Crossref has no phrase search, so a plain query for "retrieval augmented generation" matches
about a million works. This tool scans the top `scan` relevance-ranked hits and keeps only
works whose title contains the exact phrase, then aggregates them. It is a sample of the most
relevant works, not a complete count; older years may be under-represented.
Args:
phrase: The topic phrase, e.g. "retrieval augmented generation" (hyphens/case ignored).
scan: Relevance hits to scan, 100-2000 (default 500). Larger = slower but more complete.
from_year: Optional earliest publication year.
until_year: Optional latest publication year.
work_type: Optional Crossref type, e.g. "journal-article".
Returns:
{"phrase", "fuzzy_total" (all keyword matches, for context), "scanned", "matched",
"per_year": {year: count}, "top_venues", "top_publishers", "top_funders": [[name, count]],
"most_cited": [compact work]}
Counts describe metadata only; they do not rank venue or funder quality.
|
| get_crossref_authorA | Build an author profile: publications, years active, frequent co-authors, venues, affiliations
and ORCID iDs seen in Crossref metadata.
When to use:
- "Who is this researcher / what do they work on / who do they collaborate with?"
- Finding collaborators or research groups around a person.
Args:
name: Author name, e.g. "Geoffrey Hinton". Matching is fuzzy, so namesakes can be mixed in.
orcid: ORCID iD (e.g. "0000-0002-1825-0097"). When given, only works carrying this iD are used,
which removes namesakes. If the result lists several ORCID iDs, rerun with one of them.
max_works: Works to scan, 20-500 (default 100).
Returns:
{"name", "orcid_filter", "scanned", "matched", "total_citations", "years": {year: count},
"orcids_seen", "affiliations", "coauthors", "venues": [[name, count]],
"most_cited": [compact work]}
Counts describe Crossref metadata, not research impact; centrality is not quality.
|
| get_crossref_journalA | Look up a journal: publisher, ISSNs, subjects, DOI counts, metadata coverage and latest works.
When to use:
- "Tell me about journal X", "what does X publish lately?", "does X deposit abstracts/ORCIDs?"
- Finding a journal's ISSN from its name (then filter search_crossref by issn).
Args:
issn_or_name: An ISSN like "0957-4174" for the full profile, or a name like
"expert systems with applications" to list matching journals with their ISSNs.
latest: Number of most recent works to include for an ISSN lookup (0-20, default 5).
Returns:
For an ISSN: {"title", "publisher", "issn", "subjects", "total_dois", "coverage" (share of
current works with abstracts, ORCIDs, references, licenses, ...), "dois_by_year", "latest"}.
For a name: {"matches": [{"title", "publisher", "issn", "total_dois"}]}.
Coverage is metadata completeness, not a journal quality ranking.
|
| get_crossref_funderA | Look up a research funder and the works it funded (as declared by publishers in Crossref).
When to use:
- "What has LPDP / NSF / Horizon Europe funded?", grant output tracking, funder landscape.
Args:
name_or_id: Funder name ("LPDP", "National Science Foundation") or Crossref Funder ID
("501100014538").
latest: Number of most recent funded works to include (0-20, default 5).
Returns:
{"funder": {"id", "name", "location", "alt_names"}, "other_matches", "funded_works_total",
"per_year": {year: count}, "top_venues": [[name, count]], "latest": [compact work]}
|
| snowball_doiA | Citation snowballing around one paper: its references (backward), the works that cite it
(forward), and related works, plus open-access status and an abstract when available.
Forward citations are not available from Crossref, so this tool uses OpenAlex (free; set
OPENALEX_API_KEY for a 10x daily budget, otherwise it runs keyless).
When to use:
- Systematic literature review snowballing from one or two seed papers.
- "Who built on this paper?", "what newer work cites it?", "is there a free PDF?"
Args:
doi: DOI of the seed paper.
rows: Works per direction, sorted by citation count (1-50, default 10).
Returns:
{"seed": work with open_access + abstract, "backward_total", "backward": [...],
"forward_total", "forward": [...], "related": [...],
"strong_candidates": works found by more than one direction}
Notes:
- OpenAlex matching is automatic and occasionally wrong; sanity-check odd entries.
- open_access.url is a legal free copy when OpenAlex knows one; DOI != free PDF.
- Costs about 4 OpenAlex list calls; single lookups are free.
|