Reviewer Zero
Integrates with arXiv to resolve cited arXiv identifiers and flag arXiv-now-published references during citation checks.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Reviewer ZeroReview my paper at ~/papers/draft.pdf for ICLR 2027"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Reviewer Zero
An open-source MCP server that reviews your ML paper before submission the way a strict PI would: format and anonymization, every reference, prior work for each claim, a methodology checklist and a writing check, with every quote checked against your PDF. It runs inside Claude Code or Claude Desktop, on your own Claude plan. It never gives a score, a rating or an accept/reject verdict, and it never rewrites your text.
MIT licensed. Package reviewer-zero (import name reviewer).
Install
You need uv and a free index key. Get a free key at https://tryreviewerzero.com. Questions: hello@tryreviewerzero.com. The key meters use of the hosted search index; it is not a payment method.
Claude Code
claude mcp add reviewer-zero -e REVIEWER_ZERO_INDEX_KEY=<your key> -- uvx reviewer-zero
# the review workflow (the skill), so Claude knows the order and the rules:
git clone https://github.com/rishab-ghosh/reviewer-zero
mkdir -p ~/.claude/skills && cp -r reviewer-zero/skills/reviewer-zero ~/.claude/skills/Then ask: "Review my paper at ~/papers/draft.pdf for ICLR 2027."
Claude Desktop: add to claude_desktop_config.json (Settings → Developer → Edit Config):
{
"mcpServers": {
"reviewer-zero": {
"command": "uvx",
"args": ["reviewer-zero"],
"env": { "REVIEWER_ZERO_INDEX_KEY": "<your key>" }
}
}
}and add skills/reviewer-zero as a skill (zip the folder and upload it in Claude's skills
settings), or paste skills/reviewer-zero/SKILL.md into a project's instructions.
Optional, recommended: GROBID for the best reference parsing (and required by
review_paper). It runs locally; your PDF never leaves your machine for it:
docker run --rm -p 8070:8070 grobid/grobid:0.9.1-crfWithout it, check_citations reads the reference list from the PDF's text layer and resolves
about 72% as many references (see "Measured" below).
Related MCP server: Paper Scout MCP Server
Two ways to review
The default: the skill, on your Claude plan, no API key. Claude reads your paper and
writes the review following skills/reviewer-zero/SKILL.md; the tools do retrieval and the
checks that need code. Every quote goes through verify_quotes before Claude may use it.
The measured pipeline: review_paper, on your own ANTHROPIC_API_KEY. Our prompts and
models, the pipeline we measured on the dev set. It needs a local GROBID and the key in the
server's environment (-e ANTHROPIC_API_KEY=...). It always shows an estimate and a hard spending limit first
(confirm=false spends nothing) and runs only after you agree. Two novelty settings:
| Candidates per claim sent to the reranker | Dev recall@10 | Demo paper, end to end | Demo paper, cost |
| 150 | 43.2% (32 / 74) | 172 s | $1.61 |
| 60 | 36.5% (27 / 74) | ≈ 2 min | ≈ $0.88 |
Recall is on the same 50 dev cases (74 prior works named by real reviewers); times and costs
are one run each on our 4-page demo paper (docs/demo), so a long paper takes longer.
Timeouts and the cheap resume. Claude Code stops waiting for a tool after its MCP tool
timeout; start it with a longer one for review_paper:
MCP_TOOL_TIMEOUT=900000 claudeIf a call still times out, call review_paper again on the same PDF: finished steps are
served from the local cache (~/.cache/reviewer-zero) and only the unfinished ones are paid
for.
Tools
Tool | What it does | Needs | Sends off your machine |
| desk-reject risks: names, emails, self-citations, identifying URLs, PDF metadata, page limits, required sections | nothing | nothing |
| resolves every reference; flags unresolved, wrong year, wrong authors, arXiv-now-published, duplicates | index key | titles, DOIs and arXiv ids of the works you cite |
| 40 candidates for one claim from 6–10 component queries, with BibTeX | index key | the queries Claude writes |
| one paper's details; one hop of its citation graph | index key | the paper key |
| checks each quote appears verbatim in your PDF, with its page | nothing | nothing |
| the measured pipeline (above) | index key, | your paper's text to Anthropic under your key; queries and cited titles as above |
Quota. A key allows 600 units an hour and 5,000 a day. One find_prior_work call costs
about 17 units (6–10 searches plus two expansion calls), so about 35 calls an hour.
Privacy
Reviewer Zero's tools never send your PDF or its text anywhere, except to Anthropic under your own key in
review_paper. (In the skill path, Claude itself reads the paper through your Claude client, as with any file you give it.)Only search queries, paper keys and the titles, DOIs and arXiv ids of works you cite go to our index; cited works' identifiers may also go to OpenAlex, Crossref and arXiv to resolve them.
GROBID is used only at a local address (
localhost); any otherGROBID_URLis ignored.No telemetry. The index counts requests per key and stores no query text.
tests/test_mcp_privacy.pyruns every tool on a test paper and fails if any request goes to another host or carries a sentence of the paper's body.
Measured
Dev-set numbers, measured on 2026-10-02 unless noted; details, caveats and suggested wording
in docs/MEASURED_CLAIMS.md.
Prior-work recall@10, measured pipeline (
review_paper): 43.2% full, 36.5% lite, on 50 dev cases with 74 reviewer-named prior works.Prior-work recall, MCP path (Claude writes the queries from the tool description, then ranks the 40 candidates): 25.7% with Sonnet 5, 27.0% with Opus 5.5 (recall@10, same 50 cases; today's index). The gap to the pipeline is mostly retrieval: the pipeline reranks 150 candidates per claim, the tool returns 40. For the most thorough prior-work search, use
review_paper.Skill path, end to end (Opus 5.5, 6 dev papers): about 4 minutes per paper; 184 of 188 quotes Claude submitted were verified (it must fix or drop the rest); no score or verdict sentence in any review. Too few papers for a recall number.
find_prior_worklatency: median 5.2 s for 6 queries against the hosted index (16 calls from Austin, TX, 4.7–9.3 s): about 2.4 s of searches, 2.3 s of citation-graph expansion and 0.3 s of metadata, with at most 4 index requests in flight (the per-key limit). On a local copy of the index the same call takes a median 1.6 s.check_citationswithout GROBID: 72.3% as many references resolved as with GROBID (30 dev papers).Index: about 1.09 million papers (577k arXiv papers in ML and adjacent fields since 2017, plus works they cite and venue papers not on arXiv), 28.3 million citation edges, updated nightly.
Licence
MIT (LICENSE). Runtime dependencies are permissively licensed: mcp, anthropic,
pydantic, pydantic-settings, typer, pyyaml, pdfplumber (MIT); httpx, lxml,
pypdf, numpy (BSD); the optional docling extra (MIT). GROBID, run separately as a
container, is Apache-2.0.
Development
make setup # uv sync, pre-commit install
make lint test # ruff and pytest; tests never touch the networkThe pipeline behind review_paper is in src/reviewer/: parse/ (GROBID), claims/,
novelty/, methods/, writing/, citations/, format/, meta/ (including the leak
filter that keeps scores and verdicts out of every output), review/ (orchestration, budget)
and mcp/ (the server and its tools). Prompts are versioned files next to each step.
Available Tools
7 toolscheck_citationsA
Check every reference of a paper PDF: resolve it (the Reviewer Zero index first, then OpenAlex, Crossref and arXiv) and flag references that cannot be found, have the wrong year or wrong authors, are cited as arXiv preprints although now published, or are listed twice. No LLM is used and no API key other than the index key is needed.
Sent off your machine: the titles, DOIs and arXiv ids of the works the paper cites (to the index and to OpenAlex, Crossref and arXiv) — never the paper's own text. The reference list is read by a local GROBID container if one is running, otherwise from the PDF's text layer; the result says which (parser). Report each flag with the reference as printed and what it resolved to; a reference that did not resolve is not proof it does not exist.
Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| pdf_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| flags | Yes | |
| notes | Yes | |
| parser | Yes | grobid: a local GROBID container parsed the reference list; fallback: the reference list was split from the PDF's text layer (see text_engine). |
| n_resolved | Yes | |
| references | Yes | |
| text_engine | No | |
| n_references | Yes | |
| index_available | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses the resolution chain and its order (index first, then OpenAlex/Crossref/arXiv), the exact data sent off-machine, that a local GROBID container is used when available with a reported `parser` field, and the important caveat that an unresolved reference is not proof of non-existence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is correctly front-loaded in the first sentence, but the 'Sent off your machine' paragraph and the following 'Privacy' paragraph substantially duplicate each other, and the privacy block is partly boilerplate shared with review_paper. The definition is longer than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and the description still covers behavior, resolution order, flag semantics, parsing fallback, and data-egress boundaries. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (pdf_path) with 0% schema description coverage, so the description is the sole source of meaning. It implies a paper PDF is supplied but adds nothing about path format, single-vs-multiple files, or local/remote path expectations; the parameter name is self-explanatory enough that this is an acceptable but unremarkable 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Check every reference of a paper PDF') and enumerates exactly what it resolves and flags (unresolvable refs, wrong year/authors, preprint-now-published, duplicates). This is clearly distinguishable from siblings like verify_quotes or find_prior_work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the applicable context clear (validating a paper's bibliography) and even notes the prerequisite that no LLM or extra API key is needed. It does not, however, name alternatives such as verify_quotes or check_format or state when NOT to use it, so the routing guidance stops short of explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_formatA
Check a paper PDF against a venue's format and anonymization rules, locally and for free (no LLM, nothing sent anywhere). venue: iclr2027, neurips2026, icml2026, cvpr2027, acl or other; stage: submission, camera_ready or preprint. At submission, anonymization problems (author names or affiliations on page 1, emails, first-person self-citations, identifying URLs, PDF metadata) are violations. Page limits and required sections are checked only for venues whose rules are built (ICLR 2027, NeurIPS 2026, ICML 2026); for the others the result says so. Report every violation with its page and evidence; never rewrite the paper.
Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| stage | Yes | ||
| venue | Yes | ||
| pdf_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| notes | Yes | |
| rules | Yes | The venue rules applied ('ICLR 2027'), or why none were (D18). |
| stage | Yes | |
| venue | Yes | |
| n_pages | Yes | |
| findings | Yes | |
| page_limit | Yes | |
| text_engine | Yes | pdftotext (poppler installed) or pypdf. |
| detected_header | Yes | What page 1 itself declares; used only to warn about a mismatch. |
| main_text_pages | Yes | For the chosen venue; None when its page rules are not built. |
| author_names_from | Yes | 'none': GROBID was not reachable, so names on page 1 were not checked (affiliation, email, self-citation, URL and metadata checks still ran). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: local-only execution, no LLM, nothing sent anywhere, exact data egress points for the index and OpenAlex/Crossref/arXiv, and the guarantee that it 'never rewrite[s] the paper'. It also discloses a real limitation ('for the others the result says so') and the output contract (every violation with page and evidence). This is unusually complete behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a well-formed, front-loaded purpose statement. The privacy block, however, is long and largely concerns sibling tools – it explains what review_paper and check_citations send – which is off-target for this tool and dilutes the description rather than tightening it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the description still covers the key scoping caveat (only ICLR 2027, NeurIPS 2026 and ICML 2026 have page-limit/section rules) plus the reporting format. The gap is scope confusion: it documents data flows for two other tools inside this tool's description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it lists the venue values (including 'other') and the stage values, and explains the semantic difference of 'submission' (anonymization checks apply) versus camera_ready/preprint. pdf_path is never described, but its meaning is self-evident from the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Check a paper PDF against a venue's format and anonymization rules') and immediately narrows scope ('locally and for free'). That is clearly distinct from review_paper, check_citations and verify_quotes among the siblings, so an agent can route without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives strong operational context – 'At submission, anonymization problems ... are violations' and which venues have page-limit/section rules built – which implies when the tool is useful. However, it never explicitly says when to prefer this over review_paper or the other checker siblings, and review_paper only appears incidentally in the privacy note, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citationsA
List the papers one paper cites (direction "cites") or the papers in the Reviewer Zero index that cite it (direction "cited_by"), with metadata and BibTeX for each. Use it to follow a candidate from find_prior_work one hop. limit is 1 to 200 (default 50).
Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| limit | No | ||
| direction | No | cites |
Output Schema
| Name | Required | Description |
|---|---|---|
| key | Yes | |
| papers | Yes | |
| direction | Yes | 'cites': works this paper cites; 'cited_by': works in the index that cite it. |
| truncated | Yes | True when `limit` cut the list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose data egress (paper keys and search queries to the Reviewer Zero index, per-key request counting, nothing stored). It does not describe output/pagination behavior beyond the limit range, and a large share of the privacy text concerns other tools (review_paper, check_citations) rather than this one.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The functional sentences are tight and front-loaded, but the privacy block is long and much of it is off-topic for this tool (it explains what happens when calling review_paper and check_citations). That dilutes the signal for an agent trying to select this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. Function, direction semantics, and limit bounds are all covered. The only real gap is the meaning of the required `key` parameter, which is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it defines both directions and gives the limit range (1–200, default 50). The `key` parameter is not explicitly defined, though the surrounding text implies it is a paper key.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (list) and resource (papers a paper cites / papers that cite it), and it disambiguates the two directions using the exact enum values. An agent can tell this apart from find_prior_work or paper without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a concrete workflow: 'follow a candidate from find_prior_work one hop,' naming the sibling that leads into it. There is no explicit when-not guidance or statement of when the other direction is preferred, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_prior_workA
Find prior work that may already contain a claim of an ML paper.
Before calling: decompose ONE claim into 6 to 10 component queries, one per technique, objective, data-construction step or problem framing. Phrase each the way a prior paper's abstract would describe that component: a declarative sentence of 10 to 40 words, not a question, not naming this paper or its method name. Set before to the paper's submission date so later work is excluded. Call once per claim.
Returns up to k (default 40) candidates in a cheap first-stage order: title, year, venue, authors, abstract, link, seed_count, the queries that found it and a BibTeX entry. The order is not a judgement. You must read the abstracts and label each candidate yourself: same (it already contains the claim), close, builds on, or different, with a one-line reason; show same and close first. Never declare the paper novel or not novel overall.
Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| before | No | ||
| queries | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| before | No | |
| queries | Yes | |
| candidates | Yes | |
| n_considered | Yes | Distinct candidates scored (search hits plus one-hop citation neighbours). |
| index_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and mostly succeeds: it discloses that results come in a 'cheap first-stage order' that 'is not a judgement,' that the agent must read and label candidates itself, and that it must 'Never declare the paper novel or not novel overall.' It also spells out the privacy/data-flow model. It omits edge cases like error behavior or index failure modes, keeping it short of exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and usage guidance are front-loaded and well organized under a 'Before calling' block. However, the closing 'Privacy' paragraph is long and tangential to invocation, and the return-value enumeration (title, year, venue, authors, abstract, link...) is largely redundant given the tool already has an output schema. Together these inflate the description beyond what an agent needs to select and call it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search/retrieval tool with an output schema and no annotations, the description covers purpose, procedural guidance, ordering caveats, and privacy. The main gap is that it does not help the agent distinguish this tool from neighboring citation/paper tools, but otherwise an agent has what it needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does well for the two important parameters: `queries` gets detailed semantics (6-10 per claim, declarative 10-40 word sentences, not questions, not naming the method) and `before` is tied to the paper's submission date to exclude later work. `k` receives only 'up to `k` (default 40),' leaving its effect under-specified, which is why this is not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Find prior work that may already contain a claim of an ML paper,' and the scoping ('one claim,' candidate retrieval) is clear. However, it never explicitly distinguishes this from siblings like citations or check_citations, so an agent must infer the boundary from context rather than being told.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Before calling' block gives unusually concrete guidance: decompose ONE claim into 6-10 component queries, phrase each as a 10-40 word declarative sentence, set `before` to the submission date, and call once per claim. This is clear context for how to use the tool, but it names no alternative tools or conditions under which another sibling would be preferred, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paperA
Look up one paper in the Reviewer Zero index by key: a bare arXiv id (2106.09685) or a key returned by find_prior_work (openalex:W..., s2:...). Returns title, authors, year, venue, DOI, arXiv id, abstract, a link and a BibTeX entry. Fails with a clear message if the key is unknown.
Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| doi | No | |
| key | Yes | |
| kind | No | |
| link | No | |
| year | No | |
| title | No | |
| venue | No | |
| bibtex | Yes | |
| authors | No | |
| abstract | No | |
| arxiv_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose error behavior ('Fails with a clear message if the key is unknown') plus which data leaves the machine and to whom. It omits rate limits and any explicit read-only statement, but 'look up' plus the privacy section make the safety profile clear enough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose, key format, return shape and failure mode are front-loaded into one dense sentence, which is exactly the right ordering. The multi-sentence privacy block is relevant but long and largely generic across the toolset, diluting focus on selection-relevant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described (they are, harmlessly), and the only parameter is fully explained in prose. Error behavior and data-handling are covered, so nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter is undocumented in the schema, but the description compensates by giving the accepted key syntaxes with concrete examples (bare arXiv id, 'openalex:W...', 's2:...'). Minor gap: DOI-style keys are mentioned as inputs to check_citations but never as valid keys here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb and resource ('Look up one paper in the Reviewer Zero index') and pins the scope to a single record, which separates it from the search-style sibling find_prior_work. It also enumerates the returned fields, so an agent knows exactly what it gets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the required input and where valid keys originate ('a key returned by find_prior_work'), which implies the workflow: find_prior_work to discover, paper to fetch one. It does not, however, explicitly say when to prefer this over the citations or verify_quotes siblings that also operate on works.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_paperA
Run the full Reviewer Zero review of a paper PDF (claims, prior work, citations with support checks, methodology, writing, format) with our prompts, using the user's own ANTHROPIC_API_KEY. It costs real money (usually about $1.50 to $2.00) and takes 7 to 12 minutes.
ALWAYS call it first with confirm=false: that spends nothing and returns the estimate and the hard limit. Show both to the user and call again with confirm=true only after they agree. If a call times out, call again with the same PDF: finished steps are served from the cache and only what is left is paid for.
novelty="full" (default) sends the 150 best candidates per claim to the reranker; novelty="lite" sends 60 (quicker and cheaper, lower recall; see the README for the measured recall and time of each).
Report findings evidence-first, each with its quote. Never give a score, rating or accept/reject verdict, and never rewrite the user's text. Without ANTHROPIC_API_KEY or a local GROBID container it explains what is missing; check_format, check_citations, find_prior_work, paper and citations work without a key.
Sent off your machine by this tool: the paper's text to Anthropic under your own key; search queries and cited works' titles/DOIs/arXiv ids to the index, OpenAlex, Crossref and arXiv.
Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| venue | Yes | ||
| confirm | No | ||
| novelty | No | full | |
| pdf_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| notes | No | |
| phase | Yes | 'plan' (confirm=false): nothing spent; show the estimate and ask the user. 'review': the full run. |
| steps | No | |
| format | Yes | |
| findings | No | |
| limit_usd | Yes | Hard limit: the run stops calling the model before spending more than this. |
| spent_usd | No | |
| meta_review | No | The allowlisted meta-review (no number anywhere; every sentence passed the D12 leak filter). None when the meta step did not run or failed. |
| estimate_usd | Yes | About this much (sum of per-step medians for the steps not yet done; D5). |
| novelty_mode | No | full: 150 candidates per claim go to the LLM reranker; lite: 60 (quicker, cheaper). |
| steps_to_run | Yes | |
| steps_already_done | Yes | Finished in an earlier call on the same PDF; rerunning them is served from the cache. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations at all, the description carries the full burden and does so richly: concrete cost ($1.50-$2.00), runtime (7-12 min), that confirm=false spends nothing, cache-served idempotent retries, prerequisite credentials (ANTHROPIC_API_KEY, local GROBID container), and exactly what data leaves the machine and to whom. It also constrains output behavior (evidence-first, no scores or accept/reject verdicts, never rewrite user text).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, cost and the confirm workflow, which is the right order. However the 'Sent off your machine by this tool' paragraph and the following 'Privacy' paragraph restate the same egress facts (search queries and cited works' titles/DOIs/arXiv ids to the index, OpenAlex, Crossref and arXiv), which is redundant bulk in an already long description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description still covers cost/limit transparency, the two-step confirmation flow, credential and container prerequisites, caching, privacy/egress, and the reporting style. For a long-running, paid, multi-step tool this is complete enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains confirm (dry-run that returns estimate and hard limit), novelty semantics with concrete candidate counts and the recall/time tradeoff, and implicitly the pdf_path reuse rule on retries. Only venue is left unexplained, but its enum of venue names is self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run the full Reviewer Zero review of a paper PDF') and enumerates the covered dimensions (claims, prior work, citations, methodology, writing, format), so it is clearly the comprehensive sibling of check_format/check_citations/find_prior_work. It never explicitly says 'use the other tools for a cheaper subset', so the differentiation is implied through the word 'full' rather than spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit mandatory workflow: 'ALWAYS call it first with confirm=false', show the estimate and hard limit to the user, then call again with confirm=true only after agreement. Also states the timeout retry policy and the tradeoff between novelty='full' (150 candidates, default) and novelty='lite' (60, quicker/cheaper, lower recall), which is exactly the when-to-use guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_quotesA
Check that quotes you intend to cite from a paper appear in its PDF word for word, and on which page. Use it on every quote before putting it in a review: a quote that is not verified must not be presented as the paper's text. Whitespace, line-break hyphenation, ligatures and punctuation are forgiven; added, dropped or changed words are not. Join separate fragments with "..." only when they appear in order, close together. Quotes need at least 5 words. Up to 50 quotes per call. Local only: nothing is sent anywhere.
Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| quotes | Yes | ||
| pdf_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| quotes | Yes | |
| n_verified | Yes | |
| text_engine | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses matching tolerance (whitespace, hyphenation, ligatures, punctuation forgiven; word changes not), fragment-joining rules, a 5-word floor, a 50-quote cap, and that the operation is local-only. This is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and matching rules are front-loaded and every sentence in the first block earns its place. The privacy paragraph is long and spends several sentences on other tools' data flows (review_paper, check_citations), which is tangential to invoking this tool, though the local-only claim itself is valuable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. For a two-parameter verification tool, the description supplies matching semantics, limits, workflow placement, and privacy posture — everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds substantial meaning for `quotes` (minimum 5 words, max 50 per call, join fragments with '...' in order), but says nothing about `pdf_path` (format, absolute vs relative). Half the parameters remain undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: 'Check that quotes you intend to cite from a paper appear in its PDF word for word, and on which page.' An agent immediately knows this verifies quote fidelity against a PDF. However, it never explicitly contrasts itself with siblings like check_citations or check_format, so routing relies on the reader inferring the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage trigger: 'Use it on every quote before putting it in a review: a quote that is not verified must not be presented as the paper's text.' This tells the agent when to reach for the tool. It stops short of naming an alternative tool or an explicit when-not case, so it is strong context without full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
check_citations - First observed
check_format - First observed
citations - First observed
find_prior_work - First observed
paper - First observed
review_paper - First observed
verify_quotes
TDQS
Scored across 7 tools
Each tool targets a distinct operation: searching prior work, single-paper lookup, citation-graph traversal, format checking, reference resolution, quote verification, and full review. The main overlap is that review_paper orchestrates the same steps the atomic tools perform individually, but the descriptions frame the atomic tools as composable primitives, so an agent can tell which to pick.
check_format, check_citations, verify_quotes, find_prior_work, and review_paper follow a clear verb_noun pattern. The bare nouns 'paper' and 'citations' are the only deviations, and they remain readable and unambiguous.
Seven tools is well-scoped for a paper-review assistant, with each tool earning its place: atomic checks, lookup, graph traversal, and one orchestrating full review. No redundant or filler tools.
The surface covers the core review lifecycle: prior-work search, paper lookup, citation following, format/anonymization checks, reference validation, quote verification, and a full pipeline. Minor gaps exist (e.g. no standalone claim-extraction or title/author search), but agents can work around them.
Related MCP Connectors
Research-backed linting + generation for agent context files (CLAUDE.md, AGENTS.md, Cursor rules).
Reviews of arXiv papers for AI agents: verdicts, claims, flaws, compiled records, citation graphs.
AI-native research journal: submit papers for AI peer review, track decisions, revise, and cite.
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Related MCP Servers
- AlicenseAqualityDmaintenanceConnects Claude to OpenReview to facilitate conference reviewer and area chair workflows through the Model Context Protocol. It enables users to manage paper assignments, read submissions and reviews, and perform write actions like submitting reviews or comments using a preview-and-confirm system.13MIT
- AlicenseNot gradedqualityDmaintenanceEnables reviewing and managing AI research paper candidates from arXiv, Semantic Scholar, and Hugging Face Daily Papers, scored against personal interests, via MCP tools in Claude Desktop or Cowork.9MIT
- AlicenseAqualityCmaintenanceSearch and read arXiv papers directly from Claude. Supports keyword, author, category, and date filtering plus full PDF text extraction so Claude can read, summarise, and reason over entire papers, not just abstracts.525MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for agentic-first PDF review of LaTeX papers, enabling commenting, compiling, and visual inspection via Claude Code.8MIT