Skip to main content
Glama
amos689

paper-preflight

paper-preflight

Check every reference of a LaTeX paper against real scholarly records before you submit. No LLM guessing, no false accusations.

CI PyPI version MIT license paper-preflight MCP server on Glama

Python 3.11 to 3.14 Input: LaTeX, BibTeX and PDF Checked against six scholarly databases No LLM in the verdicts Read-only MCP tools

Tested on Windows, Linux and macOS English and Simplified Chinese

English · 简体中文 · Quick start · Releases · Feedback

paper-preflight checking the demo paper: errors for an undefined citation key, a DOI that belongs to another paper, a reference no source knows, a retracted paper and a duplicate entry key; warnings for a published preprint, two entries for the same work, a wrong year and a LaTeX-escaped DOI

Language models invent references, and copy-pasted BibTeX carries wrong years, wrong authors and dead DOIs. paper-preflight reads your .tex and .bib files and asks Crossref, dblp, arXiv, DataCite, PubMed and OpenAlex (and Semantic Scholar, if you have a key) about every cited work:

  • Does it exist?

  • Does it match what you wrote?

  • Has it been retracted?

  • Has the preprint you cite been published since?

When it cannot tell, it says so instead of guessing.

Status: v0.4, an early release. False positives are the bugs we most want to hear about: please open an issue.

The repository's demo paper cites eleven works, several of them wrong on purpose. A real run, against the live sources:

$ paper-preflight check examples/demo-paper
paper-preflight 0.4.0 · main.tex · 12 entries, 12 cited keys

error   CIT001 main.tex:31
    Citation key 'nonexistent2023' is not defined in any bibliography file (1 use(s)).
error   REF001 refs.bib:43
    The doi of 'devlin2019bert' (10.1109/cvpr.2016.90) resolves to a different work in Crossref: "Deep Residual Learning for Image Recognition" (He et al., 2016).
error   REF003 refs.bib:66
    'lindqvist2024quantum' was not found in Crossref, dblp and Semantic Scholar, and every source responded. Check that the work exists and that its title is correct.
error   REF004 refs.bib:73
    'wakefield1998ileal' has been retracted (reported by Crossref, OpenAlex). Cite it only if the text discusses the retraction.
error   CIT002 refs.bib:127
    Entry key 'kingma2015adam' is already defined at line 47; BibTeX ignores this one.
warning REF015 refs.bib:31
    'he2015residual' cites a preprint that has been published in CVPR (2016), DOI 10.1109/cvpr.2016.90. Cite the published version and keep the eprint field.
warning CIT004 refs.bib:37
    Entries 'devlin2019bert' and 'he2016deep' look like the same work (same DOI).
warning REF013 refs.bib:51
    'kingma2015adam' gives the year 2016, but dblp records 2014, 2015.
warning REF017 refs.bib:111
    The doi of 'tacl2019example' contains LaTeX escapes: '10.1162/tacl\_a\_00276'. Write it as: 10.1162/tacl_a_00276
info    REF005 refs.bib:73
    'wakefield1998ileal' has a published correction (reported by Crossref).
info    REF090 refs.bib:86
    'goodfellow2016deep' could not be verified: grey literature without an identifier (book, report, software, web page).
info    REF090 refs.bib:94
    'zhou2016ml' could not be verified: non-Latin titles are not supported yet; grey literature without an identifier (book, report, software, web page).
info    CIT003 refs.bib:115
    Entry 'lecun1998gradient' is never cited.

References: 6 verified · 1 metadata mismatch · 1 identifier conflict · 1 not found · 2 cannot determine
5 error(s) · 4 warning(s) · 4 info

Each finding is backed by a record (or by every source answering "no"). The correct NeurIPS paper is verified through dblp even though Crossref only holds fake copies of it, and the two books without identifiers are reported as "cannot determine" instead of "not found".

What it catches

Rule

Finding

REF001

The DOI or arXiv ID points to a different paper

REF002

The DOI or arXiv ID does not exist

REF003

The work was not found in any source, and every source answered

REF004 · REF005

The work was retracted, or has an expression of concern or a correction

REF010–REF014

Authors, title, year or venue differ from the real record

REF015

A cited preprint has been formally published

REF016

The registry has a DOI the entry lacks (offered as a safe fix)

REF017

An identifier is written so that links break (10.1162/tacl\_a\_00276, …v1)

CIT001–CIT008

Undefined, duplicate, unused or near-duplicate citation keys; broken .bib syntax

REF090

Cannot determine, always with the reason (source unavailable, grey literature, …)

paper-preflight explain REF003 describes any rule.

Related MCP server: Citation Guard

How accurate is it?

Four measurements, all against the live sources: hallucinations found in published papers, the bibliographies of real papers, a head-to-head with published tools, and a public benchmark.

On hallucinations that got past peer review

GPTZero published 151 hallucinated references it found in NeurIPS 2025 papers and ICLR 2026 submissions, each confirmed by its staff. Pasted as plain text, as the papers printed them:

References

Flagged

Cannot determine

Verified

151

135 (89%)

16

0

  • None of them is verified. The 16 left undecided are web pages and blog posts, titles too short to search with confidence, real titles given with invented authors where several works share the title, and two references the plain-text reader could not take apart. Each is listed with its reason in evals/results/gptzero.md.

  • GPTZero's own tool found these, so they are the hallucinations a search can find; recall on every kind of hallucination is lower (see HALLMARK below).

On real papers

The bibliographies of 20 arXiv papers first submitted in mid-August 2026 (cs, stat, q-bio, quant-ph and astro-ph), chosen mechanically and collected only after every fix in this release, with every warning and error reviewed by hand:

References

Flags

Real problems

False positives

Unclear

False positives per 100 references

753

86

77

9

0

1.2

  • One false alarm every two papers (38 references on average), against 77 real problems: 36 errors in the entries (invented authors and titles, wrong given names, years and titles, identifiers written so that links break) and 41 cited preprints that have since been published.

  • The false alarms are mostly registry records with errors of their own (two misspelt titles, an affiliation mark inside a name, a workshop filed under a joint volume) and real works no source describes as cited (a Substack post, a technical report, a database cited by its access year, an article's early-access year).

  • Six earlier batches of 20 papers were used to find false positives, each first measured as it came out (0.1.0: 4.5 per 100 references; 0.1.1: 2.3; 0.1.2 before its last fixes: 3.0; 0.1.2: 1.7; 0.2.1: 1.9; 0.3.0: 1.9). Details in evals/README.md.

Next to other tools

Badalova & Mayr (2026) checked 104 references by hand and published what five tools flagged. On the same references, with their labels:

Tool

Precision [95% CI]

Recall

False flags per 100 correct references

CheckIfExist

47.7% [36.0%, 59.6%]

93.9%

47.9

HalluCiteChecker

47.4% [32.5%, 62.7%]

54.5%

28.2

Hallucinator

50.9% [38.3%, 63.4%]

87.9%

39.4

HalRef

31.2% [21.9%, 42.2%]

72.7%

74.6

RefChecker

47.1% [35.7%, 58.8%]

97.0%

50.7

paper-preflight

72.5% [57.2%, 83.9%]

87.9%

15.5

The sample is small, so the intervals are wide. Some flags count as false here because the study labels a reference correct when the work exists: five of paper-preflight's flags on such references point at real errors (a wrong author, a broken DOI). Two causes of false flags found in this data were fixed, and four names the study's CSV garbled were restored, before the run above; the first run measured 62.8%. See evals/results/badalova-mayr.md.

On a benchmark: HALLMARK

HALLMARK is a public benchmark of real and hallucinated BibTeX entries.

Split

Mode

Precision

Recall

False-positive rate

Coverage

test_public: 831 entries, never used during development

Any issue

98.1%

88.9%

2.2%

97.0%

Fabrication

99.0%

49.0%

0.6%

97.0%

dev_public: 1,119 entries, used during development

Any issue

97.6%

90.7%

2.1%

98.4%

Fabrication

98.1%

52.7%

1.0%

98.4%

HALLMARK v1.2.3, every entry of both public splits, run on 2026-10-04. Fabrication counts a wrong identifier, a work not found and no author in common; any issue also counts wrong authors, title, year or venue.

  • The held-out split confirms the development numbers: the same precision and two points less recall on entries no rule was ever tuned on.

  • Every flag on a dev_public entry labelled VALID was checked by hand. The 11 that remain are not correct citations: DOIs that belong to other papers, author lists naming people who did not write the paper, a shifted year and a truncated title.

  • Without them, both modes reach 100% precision and 0% false positives. The list, each item with a reason one lookup confirms, is in evals/hallmark_disputed.toml.

  • What is still missed: invented venues on papers known only as preprints (an arXiv record cannot contradict a venue) and author lists that merely leave people out. See evals/results/ for every hallucination type.

Precision comes first: a reference is called fabricated only on positive evidence, and an unanswered or ambiguous lookup is reported as "cannot determine", never as "not found". The evaluation harness and every run's summary are in evals/.

Quick start

With uv nothing needs installing (or pip install paper-preflight):

uvx paper-preflight check path/to/paper

path/to/paper is the project directory, its main .tex file, or a single .bib file. A project that ships no .bib, as many arXiv sources do, is read from its compiled .bbl (checked, but never edited).

No LaTeX at all? A reference list as plain text works too, in the common styles (APA, IEEE, ACM, Nature, Vancouver, Springer, Elsevier, Chicago, MLA), one reference per line, per paragraph or numbered:

uvx paper-preflight check references.txt
pbpaste | uvx paper-preflight check -      # or from stdin

Any arXiv paper, by its ID: the source is downloaded to a temporary folder, checked, and deleted.

uvx paper-preflight check arxiv:2607.06922

Only the PDF? Its reference list is read too, with the pdf extra:

uvx --from 'paper-preflight[pdf]' paper-preflight check paper.pdf

Option

Effect

--format json / --format sarif

Machine-readable output (SARIF works with GitHub code scanning)

--offline

Never touch the network; use only answers already in the local cache

--refresh

Ask every source again instead of using cached answers (after a correction, say)

--fail-on warning

Make warnings fail the run too (the default is errors)

--lang zh

Chinese messages (also chosen automatically from your locale)

Exit codes:

Code

Meaning

0

Nothing at or above --fail-on was found

1

Blocking findings

2

No blocking findings, but a source was unavailable, so the paper cannot be called clean yet

3

Usage error

Fetch verified BibTeX

Instead of writing an entry from memory, ask for it by DOI, arXiv ID or title. Every field comes from the registry record, which a comment above the entry names:

paper-preflight bib fetch 1810.04805
% Verified with paper-preflight against dblp (conf/naacl/DevlinCLT19), 2026-10-03
@inproceedings{devlin2019bert,
  title         = {{BERT:} Pre-training of Deep Bidirectional Transformers for Language Understanding},
  author        = {Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
  booktitle     = {NAACL-HLT (1)},
  year          = {2019},
  doi           = {10.18653/v1/n19-1423},
  eprint        = {1810.04805},
  archivePrefix = {arXiv},
}
  • Preprints: an arXiv preprint that has been published comes back as the published version, with its eprint kept (--prefer preprint for the preprint itself).

  • Titles: --title (with --author/--year if needed) lists the candidates instead of choosing when several works match.

  • Retractions: a retracted work comes with a warning.

  • Agents: --format json is for scripts and agents.

Fix the bibliography

bib fix turns findings into edits of your .bib files, taken from the verified records. It prints a diff and changes nothing until you add --apply:

paper-preflight bib fix path/to/paper --level unsafe
--- a/refs.bib
+++ b/refs.bib
@@ -48,7 +47,7 @@
   title     = {Adam: A Method for Stochastic Optimization},
   author    = {Kingma, Diederik P. and Ba, Jimmy},
   booktitle = {International Conference on Learning Representations (ICLR)},
-  year      = {2016},
+  year      = {2015},
 }
  • --level safe (the default) only fixes what cannot change which work is cited: identifiers written so that links break, and DOIs the registry has but the entry lacks.

  • --level unsafe also rewrites authors, title, year and venue from the record, and removes identifiers that point to another work. Review the diff first.

  • Only the affected fields change; comments, formatting, line endings and encoding are kept. A reference nobody could find is never "fixed": only you can say what was meant.

Silence a finding you have checked

A comment directly above an entry silences rules for that entry, with an optional reason:

% preflight: ignore[REF003] reason="internal technical report, not indexed anywhere"
@techreport{lab2024internal,
  ...
}

The verdict stays in the JSON report; only the finding is dropped. A suppression that silenced nothing is reported as CFG001 (info), so stale comments do not pile up. Reference rules are only judged after a complete online run, since offline answers and outages may leave them unrun.

Experimental: find the passage behind each citation

support looks in each cited work for a passage that says what the citing sentence claims. It reads the work's text: the arXiv source, an open-access full text or PDF, or else the abstract. A small local model (HHEM-2.1-open, 0.4 GB) then scores the passages ranked best for the claim.

pip install "paper-preflight[support]"
paper-preflight support path/to/paper --download-model --all

--download-model fetches the model's weights once. --all also lists the confirmed citations with their quotes, and arxiv:<id> works as a target, as it does for check.

  • What it says: "confirmed", with the passage quoted word for word, or "could not confirm". A citation it could not confirm comes with the reason: no text, only the abstract, or no passage close enough. A citation that only names what it cites ("Adam \cite{...}") is confirmed when the cited work's title carries the name.

  • It never calls a citation wrong. On a gold set of 298 citations, 93% of its confirmations were right [95% CI 82%, 98%]. But it confirms only about one real citation in six, and a low score pointed at a mis-citation less than half the time. The gold set's labels were made by AI models, not experts; see evals/results/support.md.

  • Or let your agent judge. The MCP tool preflight_cited_passages returns each claim with the cited work's best passages, for Claude, Codex or another agent to judge by the same rules. It needs no model and no support extra. On 100 gold-set citations, a Claude agent judging from it confirmed 39% of the real ones, against HHEM's 9%, and every confirmation was at least partially supported (the labels come from the same model family).

  • What leaves your machine: the claims are scored locally. Only the cited works' identifiers go out, to fetch their text, which is then kept in the local cache.

Use it from your coding agent

Claude Code — install the plugin. It bundles an MCP server and a skill that makes Claude check the references before calling a paper finished, fix only what is proven wrong, and never invent a reference.

claude plugin marketplace add amos689/paper-preflight
claude plugin install paper-preflight@paper-preflight

Codex, Gemini CLI, Copilot, Cursor and other agents — install the same skill. It runs the CLI when no MCP server is configured:

npx skills add amos689/paper-preflight

gh skill install amos689/paper-preflight paper-preflight installs it too.

MCP clients (Codex, Cursor, VS Code, …) — run paper-preflight mcp. The tools are read-only and confined to your workspace; see docs/mcp.md.

pre-commit — check citation keys and cached verdicts on every commit in seconds; see docs/pre-commit.md.

GitHub Actions — uses: amos689/paper-preflight@main checks the paper on every push, with the report in the job summary and optional code-scanning alerts; see docs/github-action.md.

Better results with free credentials

paper-preflight works without any account. These optional environment variables make it faster and more complete; their values are never printed or logged.

Variable

Effect

PAPER_PREFLIGHT_EMAIL

Crossref's polite pool: faster, more reliable lookups

OPENALEX_API_KEY

A larger OpenAlex budget for retraction checks

S2_API_KEY

Semantic Scholar as a rescue source for references nobody else found

paper-preflight doctor shows which are set and whether each source answers right now.

How it works

  1. Source-first. It reads the LaTeX project as LaTeX sees it: comments, \iffalse blocks and \includeonly are respected, .aux files are used when they are fresh, and the first definition of a duplicated key wins, as in BibTeX.

  2. Identifier-first routing. DOIs go to their registration agency (doi.org tells which: Crossref, DataCite, …). arXiv IDs go to arXiv, with DataCite as a fallback, and PMIDs and PMCIDs to PubMed (which also marks retracted articles). Entries without identifiers are searched by title in dblp and Crossref. dblp is read through its SPARQL endpoint, which still answers scripts now that dblp's search API is behind a bot challenge.

  3. Field-by-field matching with guards. It compares titles (including earlier arXiv version titles), authors (tolerating transcriptions such as Reiß/Reis), year and venue. A search result is used only when enough of these agree and no other work fits as well; known fake DOI copies are skipped.

  4. One verdict per reference: verified, metadata mismatch, identifier conflict, not found, or cannot determine with a reason. "Not found" needs every required source to answer "no".

  5. No LLM anywhere in the verdict. Answers are cached locally (SQLite), so re-runs are fast and --offline works.

Design principles

  • Positive confirmation or abstain. Rate limits, outages and unindexed works lead to "cannot determine", never to "not found".

  • Neutral wording. Findings state observations ("not found in Crossref, dblp and Semantic Scholar, and every source responded"), never accusations.

  • Local-first, no telemetry. Only the metadata of the cited works (DOIs, titles, authors) is sent to the public scholarly APIs above. Your manuscript never leaves your machine.

What it will never do

Help evade plagiarism or AI-text detection, scrape paywalled or bot-protected sites, recommend or "complete" references from memory, or name and shame authors.

Roadmap

  • Done: releases on PyPI (v0.1); references from a .bbl, plain text, a PDF or an arXiv ID (v0.2); an experimental evidence finder for citations, support (v0.3); installs into more agents, an agent-judged support, and recall measured on hallucinations found in published papers (v0.4)

  • Next: catch more of what is still missed (real titles with invented authors, references without titles, invented venues), each round measured on a new week of real papers

  • Later: Chinese-language references

Progress is tracked in docs/PROGRESS.md (in Chinese) and the changelog.

Contributing

Bug reports with a reproducible .bib entry are the most valuable contribution, especially false positives. See CONTRIBUTING.md.

License

MIT. See THIRD_PARTY_NOTICES.md for adapted code.

Available Tools

5 tools
preflight_bib_fixPropose fixes for a bibliographyA
Read-onlyIdempotent

Propose edits to the .bib files, taken from the verified records, as a unified diff. Nothing is written: apply the diff with your own editing tools.

level="safe" only fixes identifiers written so that links break and adds DOIs the registry has; level="unsafe" also rewrites authors, title, year and venue from the record and removes identifiers that point to another work. Show unsafe diffs to the user before applying them. References nobody could find are never "fixed".

ParametersJSON Schema
NameRequiredDescriptionDefault
keysNo
pathNo.
levelNosafe
offlineNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and idempotentHint, and the description reinforces this with the load-bearing detail that nothing is written and the agent must apply the diff itself. It further discloses what safe/unsafe actually mutate and the invariant that unfindable references are never fixed — context well beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core behavior and output format, then layers level semantics compactly. Sentences earn their place, though the level paragraph is dense enough to require careful reading.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers the write-safety semantics and level behavior well. The remaining gap is the un-annotated keys/path/offline parameters, which an agent must infer from names alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden. It fully explains the most consequential parameter (level, including the enum values' behavioral differences), but keys, path, and offline are left undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (propose edits) and resource (.bib files), plus the concrete output form (unified diff) and the source of the edits (verified records). This is clearly distinguishable from preflight_bib_lookup, preflight_check, and preflight_explain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit guidance on level="safe" vs level="unsafe" and instructs showing unsafe diffs to the user before applying. It does not, however, route the agent against any named sibling (e.g. implying bib_lookup must run first), so alternative-selection guidance is strong but incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preflight_bib_lookupGet verified BibTeXA
Read-onlyIdempotent

Get a BibTeX entry for a DOI, an arXiv ID or a title, built from the registry record instead of written from memory. Give identifier, or title (with author and year when you know them).

status is "found" (use bibtex as is), "ambiguous" (several works match: pick from candidates with the user, or retry with author/year), "not_found" (ask the user for the source; do not invent one) or "unavailable" (a source did not answer; retry later). A published preprint comes back as its published version with the eprint kept; status_flags lists notices such as "retracted".

ParametersJSON Schema
NameRequiredDescriptionDefault
yearNo
titleNo
authorNo
preferNopublished
offlineNo
identifierNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld, but the description goes well beyond them: it enumerates the four status values with their meanings, warns that a published preprint resolves to the published version with the eprint retained, and notes that status_flags can carry notices such as 'retracted'. These are exactly the edge behaviors an agent needs and cannot get from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and how to invoke it, then groups the outcome semantics into one tight paragraph. Dense but every clause carries information; no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, status/candidates/status_flags need not be re-explained, yet the description does so helpfully for the retraction and preprint-versioning cases. The only real omission is guidance on `prefer` and `offline`, which leaves an agent guessing about preprint preference and offline behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the full burden. It explains identifier, title, author, and year, but leaves `prefer` (published/preprint enum) and `offline` entirely unaddressed, so two of six parameters remain opaque. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb (get a BibTeX entry) and scope (for a DOI, arXiv ID, or title) and emphasizes it is built from the registry record rather than from memory. This is distinctive against siblings like preflight_bib_fix, which implies fixing an existing entry, so an agent can pick this one for retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear invocation guidance ('Give `identifier`, or `title` (with `author` and `year` when you know them)') and prescribes the action for each outcome: use bibtex when found, disambiguate with the user or retry with author/year, and explicitly not to invent a source on not_found. It does not name sibling tools as alternatives, so it falls short of the full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preflight_checkCheck a paper's referencesA
Read-onlyIdempotent

Verify every cited reference of a LaTeX project (or a .bib file) against real scholarly records, and check citation keys and the bibliography.

Returns a summary (error/warning counts, one verdict per reference, whether the run was complete) and the findings, most severe first, max_findings at a time; call again with offset=next_offset for more. complete: false means a source was unavailable, so the paper cannot be declared clean yet. offline: true answers from the local cache only. The first run of a paper may take a minute or two (progress is reported while the references are searched); later runs are served from the cache.

ParametersJSON Schema
NameRequiredDescriptionDefault
langNoen
pathNo.
offsetNo
offlineNo
include_infoNo
max_findingsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly/openWorld/idempotent), the description discloses pagination via offset/next_offset, the meaning of complete:false (a source was unavailable), offline cache-only behavior, and first-run latency of a minute or two with progress reporting. This is exactly the kind of behavioral context structured fields cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then the return/pagination behavior, then latency caveats. Every sentence earns its place, though the paragraph is dense and could be broken up slightly for scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description helpfully still summarizes the return shape and the critical complete:false caveat. It is nearly complete for a read-only verification tool, with the only gap being the undocumented lang and include_info parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It meaningfully explains max_findings, offset/next_offset, and offline, but says nothing about lang or include_info, leaving two of six parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: verify every cited reference of a LaTeX project or .bib file against real scholarly records, plus check citation keys and the bibliography. This is clearly distinguishable from purely fixing or lookup tools, though it never names the sibling tools (preflight_bib_fix, preflight_bib_lookup) it should be contrasted with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the usage context (verifying references before declaring a paper clean) and explains the pagination loop and offline mode, but never states when to choose this over preflight_bib_lookup or preflight_bib_fix, nor any explicit exclusions. Usage is inferable rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preflight_cited_passagesFind the passages a citation rests onA
Read-onlyIdempotent

For each sentence of the paper that cites key: the claim it makes and the passages of the cited work ranked best for that claim, for you to judge whether the work supports it. The first call downloads the work's text (its arXiv source, an open-access full text or PDF, or else its abstract) into the local cache.

Judge each claim conservatively. Say "confirmed" only when a passage states it (quote the passage word for word) or, for a citation set right after a name such as "Adam \cite{...}", when name_in_title is true. Otherwise say "could not confirm", with the reason: no text, only the abstract, or no passage says it. Never call a citation wrong or invented on this evidence: the text may say it in other words, or be only an abstract. evidence.level is "full_text", "abstract" or "none".

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
pathNo.
offlineNo
max_passagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, idempotent, openWorld), but the description adds genuinely new behavior: the first call downloads and caches the work's text from arXiv/OA/PDF or falls back to the abstract, and it enumerates evidence.level values ('full_text', 'abstract', 'none'). It stops short of describing cache staleness or rate/failure behavior on repeated calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what the tool returns, followed by concrete judging instructions that are actionable for an agent rather than filler. It is somewhat long, but nearly every sentence constrains agent behavior (quote verbatim, prefer 'could not confirm', never declare a citation wrong).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return fields need not be re-explained, and the description adds the workflow and interpretive guidance needed for a non-trivial fetch-and-judge tool. The main completeness gap is the undocumented optional parameters, which the description does not address.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameter burden, yet only `key` is meaningfully explained (the citation key being traced). `path`, `offline`, and `max_passages` are never mentioned, leaving their semantics to be guessed despite the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: for each citing sentence it returns the claim plus the best-ranked passages of the cited work. It is clearly distinguishable from siblings like preflight_bib_lookup or preflight_bib_fix, which operate on bibliography data rather than tracing a citation to source text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (deciding whether a cited work supports a claim) and gives a judgment protocol, but it never states when to choose this tool over preflight_check or preflight_explain, nor any exclusions. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preflight_explainExplain a ruleA
Read-onlyIdempotent

Explain a paper-preflight rule (for example REF003 or CIT001): what it detects, its default severity, the message template and whether a fix can be applied safely.

ParametersJSON Schema
NameRequiredDescriptionDefault
rule_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: the categories of information the caller will receive (what a rule detects, default severity, message template, fix applicability), which helps an agent decide whether this call answers its question.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tightly written sentence with the verb and resource front-loaded and the return payload enumerated after the colon. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only explainer with an output schema present, the description supplies everything needed to invoke it correctly and does not need to describe return values. Minor gap: no guidance on what happens for an unknown or malformed rule_id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single rule_id parameter, so the description must compensate, and it does by giving concrete example identifiers (REF003, CIT001) that convey the expected format. It stops short of stating whether unknown IDs are rejected or aliased, hence not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Explain') and resource ('paper-preflight rule'), gives example IDs, and enumerates the content returned (detection, severity, message template, fix safety). This clearly separates it from sibling tools like preflight_check and preflight_bib_fix.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the read-only nature and the name, but the description never states when to call this instead of siblings (e.g., before running preflight_check, or to interpret a rule ID emitted by a check). No explicit when/when-not guidance or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedpreflight_bib_fix
    • First observedpreflight_bib_lookup
    • First observedpreflight_check
    • First observedpreflight_cited_passages
    • First observedpreflight_explain

TDQS

A4.2/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct action: verify references, explain a rule, look up a BibTeX entry, propose .bib diffs, and gather cited passages. The check/fix and check/cited_passages pairs are separated by verb (verify vs edit vs ground claims) and the descriptions reinforce the boundaries.

Naming Consistency5/5

All names share the preflight_ prefix and snake_case form with a predictable verb/noun tail (check, explain, bib_lookup, bib_fix, cited_passages). Minor sub-namespacing (bib_*) does not break the pattern.

Tool Count5/5

Five tools map cleanly onto the whole preflight workflow (check, explain, lookup, fix, evidence-gathering) with no redundancy. The count is well-scoped for the domain.

Completeness4/5

The surface covers verification, rule explanation, record lookup, repair diffs, and claim-to-passage grounding, which is most of a preflight lifecycle. There is no obvious way to enumerate all rules or configure/batch settings, a minor gap agents can work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Searches academic references from arXiv, DBLP, Semantic Scholar, and OpenAlex concurrently and generates BibTeX citations.
    4
    11
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Checks whether a citation has been retracted, corrected, or flagged with an expression of concern by querying Crossref — including retractions that Crossref backfills from the Retraction Watch database, which publishers often never record in their own metadata. Lets an AI agent verify a DOI, or every DOI in a reference list, before using it in research or writing.
    2
    1
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables verifying citations in reference lists and BibTeX by detecting fabricated or mismatched references and retractions, and returning corrected BibTeX.
    3
    MIT