Skip to main content
Glama
Rishab-Ghosh

Reviewer Zero

by Rishab-Ghosh

check_citations

Resolve every reference in a paper PDF and flag unresolvable citations, wrong years or authors, unpublished arXiv preprints, and duplicates.

Instructions

Check every reference of a paper PDF: resolve it (the Reviewer Zero index first, then OpenAlex, Crossref and arXiv) and flag references that cannot be found, have the wrong year or wrong authors, are cited as arXiv preprints although now published, or are listed twice. No LLM is used and no API key other than the index key is needed.

Sent off your machine: the titles, DOIs and arXiv ids of the works the paper cites (to the index and to OpenAlex, Crossref and arXiv) — never the paper's own text. The reference list is read by a local GROBID container if one is running, otherwise from the PDF's text layer; the result says which (parser). Report each flag with the reference as printed and what it resolved to; a reference that did not resolve is not proof it does not exist.

Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pdf_pathYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
flagsYes
notesYes
parserYesgrobid: a local GROBID container parsed the reference list; fallback: the reference list was split from the PDF's text layer (see text_engine).
n_resolvedYes
referencesYes
text_engineNo
n_referencesYes
index_availableYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it discloses the resolution chain and its order (index first, then OpenAlex/Crossref/arXiv), the exact data sent off-machine, that a local GROBID container is used when available with a reported `parser` field, and the important caveat that an unresolved reference is not proof of non-existence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is correctly front-loaded in the first sentence, but the 'Sent off your machine' paragraph and the following 'Privacy' paragraph substantially duplicate each other, and the privacy block is partly boilerplate shared with review_paper. The definition is longer than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the description still covers behavior, resolution order, flag semantics, parsing fallback, and data-egress boundaries. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (pdf_path) with 0% schema description coverage, so the description is the sole source of meaning. It implies a paper PDF is supplied but adds nothing about path format, single-vs-multiple files, or local/remote path expectations; the parameter name is self-explanatory enough that this is an acceptable but unremarkable 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check every reference of a paper PDF') and enumerates exactly what it resolves and flags (unresolvable refs, wrong year/authors, preprint-now-published, duplicates). This is clearly distinguishable from siblings like verify_quotes or find_prior_work.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the applicable context clear (validating a paper's bibliography) and even notes the prerequisite that no LLM or extra API key is needed. It does not, however, name alternatives such as verify_quotes or check_format or state when NOT to use it, so the routing guidance stops short of explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.