Skip to main content
Glama
amos689

paper-preflight

Check a paper's references

preflight_check
Read-onlyIdempotent

Verify every cited reference in a LaTeX project or .bib file against Crossref, dblp, arXiv, and more. Checks citation keys, bibliography, and retractions to ensure accurate scholarly records.

Instructions

Verify every cited reference of a LaTeX project (or a .bib file) against real scholarly records, and check citation keys and the bibliography.

Returns a summary (error/warning counts, one verdict per reference, whether the run was complete) and the findings, most severe first, max_findings at a time; call again with offset=next_offset for more. complete: false means a source was unavailable, so the paper cannot be declared clean yet. offline: true answers from the local cache only. The first run of a paper may take a minute or two (progress is reported while the references are searched); later runs are served from the cache.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
langNoen
pathNo.
offsetNo
offlineNo
include_infoNo
max_findingsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly/openWorld/idempotent), the description discloses pagination via offset/next_offset, the meaning of complete:false (a source was unavailable), offline cache-only behavior, and first-run latency of a minute or two with progress reporting. This is exactly the kind of behavioral context structured fields cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then the return/pagination behavior, then latency caveats. Every sentence earns its place, though the paragraph is dense and could be broken up slightly for scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description helpfully still summarizes the return shape and the critical complete:false caveat. It is nearly complete for a read-only verification tool, with the only gap being the undocumented lang and include_info parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It meaningfully explains max_findings, offset/next_offset, and offline, but says nothing about lang or include_info, leaving two of six parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: verify every cited reference of a LaTeX project or .bib file against real scholarly records, plus check citation keys and the bibliography. This is clearly distinguishable from purely fixing or lookup tools, though it never names the sibling tools (preflight_bib_fix, preflight_bib_lookup) it should be contrasted with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the usage context (verifying references before declaring a paper clean) and explains the pagination loop and offline mode, but never states when to choose this over preflight_bib_lookup or preflight_bib_fix, nor any explicit exclusions. Usage is inferable rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.