Skip to main content
Glama
Rishab-Ghosh

Reviewer Zero

by Rishab-Ghosh

find_prior_work

Search for prior work that may already contain a claim of an ML paper. Decompose the claim into component queries and search before your submission date.

Instructions

Find prior work that may already contain a claim of an ML paper.

Before calling: decompose ONE claim into 6 to 10 component queries, one per technique, objective, data-construction step or problem framing. Phrase each the way a prior paper's abstract would describe that component: a declarative sentence of 10 to 40 words, not a question, not naming this paper or its method name. Set before to the paper's submission date so later work is excluded. Call once per claim.

Returns up to k (default 40) candidates in a cheap first-stage order: title, year, venue, authors, abstract, link, seed_count, the queries that found it and a BibTeX entry. The order is not a judgement. You must read the abstracts and label each candidate yourself: same (it already contains the claim), close, builds on, or different, with a one-line reason; show same and close first. Never declare the paper novel or not novel overall.

Privacy: runs on your machine. Your PDF and its text never leave it, except to Anthropic under your own API key when you call review_paper. What is sent: search queries and paper keys to the Reviewer Zero index (it counts requests per API key and stores nothing else), and, for check_citations, the titles, DOIs and arXiv ids of the works the paper cites, to the index and to OpenAlex, Crossref and arXiv. No telemetry.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
kNo
beforeNo
queriesYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
beforeNo
queriesYes
candidatesYes
n_consideredYesDistinct candidates scored (search hits plus one-hop citation neighbours).
index_versionYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and mostly succeeds: it discloses that results come in a 'cheap first-stage order' that 'is not a judgement,' that the agent must read and label candidates itself, and that it must 'Never declare the paper novel or not novel overall.' It also spells out the privacy/data-flow model. It omits edge cases like error behavior or index failure modes, keeping it short of exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and usage guidance are front-loaded and well organized under a 'Before calling' block. However, the closing 'Privacy' paragraph is long and tangential to invocation, and the return-value enumeration (title, year, venue, authors, abstract, link...) is largely redundant given the tool already has an output schema. Together these inflate the description beyond what an agent needs to select and call it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search/retrieval tool with an output schema and no annotations, the description covers purpose, procedural guidance, ordering caveats, and privacy. The main gap is that it does not help the agent distinguish this tool from neighboring citation/paper tools, but otherwise an agent has what it needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does well for the two important parameters: `queries` gets detailed semantics (6-10 per claim, declarative 10-40 word sentences, not questions, not naming the method) and `before` is tied to the paper's submission date to exclude later work. `k` receives only 'up to `k` (default 40),' leaving its effect under-specified, which is why this is not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: 'Find prior work that may already contain a claim of an ML paper,' and the scoping ('one claim,' candidate retrieval) is clear. However, it never explicitly distinguishes this from siblings like citations or check_citations, so an agent must infer the boundary from context rather than being told.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Before calling' block gives unusually concrete guidance: decompose ONE claim into 6-10 component queries, phrase each as a 10-40 word declarative sentence, set `before` to the submission date, and call once per claim. This is clear context for how to use the tool, but it names no alternative tools or conditions under which another sibling would be preferred, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.