CiteCheck MCP
Locates and fetches PDFs from arXiv so cited sources can be attached and checked during manuscript auditing.
Uses the Semantic Scholar API to search for PDFs and to detect preprints that have a published version, improving source discovery and citation verification.
Reads the local Zotero attachment storage folder to locate PDFs for cited sources and attach them for verification.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CiteCheck MCPaudit citations in chapter 2 of ~/tese/main.tex and show only problems"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CiteCheck MCP
Servidor MCP que confere cada citação de um manuscrito (LaTeX ou Markdown, em português ou inglês) contra o
trecho correspondente na fonte citada. Ele também sugere reescritas e as aplica no .tex com validação e backup.
O servidor não chama nenhuma API de LLM. O trabalho determinístico fica com ele:
parsing do LaTeX e do BibTeX;
busca e extração dos PDFs;
busca lexical;
verificação literal das citações;
checagem da bibliografia;
edição segura dos arquivos.
O julgamento (se o trecho sustenta a afirmação, como reescrever a frase) fica com o modelo do cliente que você já usa (Claude Code ou Claude Desktop), dentro da sua assinatura.
manuscrito.tex ─► frases + \cite ─► .bib ─► PDFs (pasta local, Zotero, arXiv, OA) ─► texto com página
│
relatório HTML ◄── veredito (com citação literal verificada) ◄── modelo do host ◄── busca BM25Instalação
Requer o uv (brew install uv), que já instala o Python 3.12.
git clone https://github.com/lucasgris/citecheck.git
cd citecheck
uv sync
uv run pytest # 19 testes, offlineClaude Code
claude mcp add citecheck -s user -- uv run --directory /caminho/para/citecheck citecheck-mcpClaude Desktop
Em ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"citecheck": {
"command": "/opt/homebrew/bin/uv",
"args": ["run", "--directory", "/caminho/para/citecheck", "citecheck-mcp"],
"env": { "CITECHECK_EMAIL": "seu@email" }
}
}
}Use em command o caminho que which uv mostra na sua máquina.
Variáveis de ambiente (todas opcionais)
Variável | Para quê |
| Entra no "polite pool" do Crossref e habilita o Unpaywall (mais PDFs de acesso aberto). |
| Sem a chave, o Semantic Scholar costuma responder 429. A chave é gratuita e melhora a detecção de "preprint com versão publicada" e a busca de PDFs. |
| Pasta de anexos do Zotero (padrão |
| Abre um manuscrito automaticamente ao iniciar. |
Related MCP server: CiteStamp MCP server
Uso
No Claude Code, peça em linguagem natural ou use os prompts prontos:
/mcp__citecheck__auditcom o caminho domain.tex: faz a auditoria completa e gera o relatório./mcp__citecheck__rewrite: sugere reescritas para as afirmações com problema, emptouen, mostra o diff e só aplica depois que você aprovar./mcp__citecheck__write_with_evidence: redige um parágrafo usando apenas trechos verificados das fontes escolhidas e o insere no.tex.
Exemplos de pedidos:
Audita as citações do capítulo 2 de ~/tese/main.tex e me mostra só os problemas.
Reescreve em inglês as frases marcadas como "partial", mantendo os \cite.
Todo o estado fica em .citecheck/, ao lado do manuscrito: SQLite, cache dos PDFs, backups e report.html.
Essa pasta tem um .gitignore próprio, porque PDFs de terceiros não devem ir para o repositório.
Ferramentas
Ferramenta | O que faz |
| Lê o manuscrito (segue |
| Lista as frases citadas com o status; mostra o detalhe com o LaTeX bruto e o contexto. |
| Localiza os PDFs (campo |
| Busca BM25 com números como âncora; leitura por página. |
| Confere se o trecho está literalmente no PDF (tolerante a ligaduras, hifenização e aspas). |
| Grava o veredito; rejeita citações que não estão no PDF. |
| Encontra frases que parecem precisar de citação (heurística EN/PT). |
| Chaves ausentes, entradas não usadas, duplicatas, DOI ausente (com sugestão), metadados divergentes do Crossref, retratações e preprint com versão publicada. |
| Valida a reescrita ( |
| Escrita atômica que preserva a codificação (UTF-8 ou latin-1) e o CRLF, com backup. |
| Relatório HTML (pt/en) com vereditos, trechos e páginas. |
| Progresso da auditoria. |
Suporte a LaTeX
Funciona bem via MCP, porque o servidor lê e escreve os arquivos .tex direto no disco:
\cite,\citep,\citet,\parencite,\textcite,\autocite,\cites{a}{b}e variantes com*e argumentos opcionais;\nocite;thebibliography/\bibitem;\bibliographye\addbibresource.Ignora comentários, o preâmbulo, equações, tabelas,
verbatime afins; trata legendas e\itemcomo unidades separadas.Converte acentos LaTeX (
\c{c},\~a,{\'e}) para comparar o texto, mas escreve de volta exatamente o que foi aprovado. Avisa quando a reescrita usa acentos Unicode num arquivo que usa macros.Não compila o documento. Para conferir a compilação depois de aplicar reescritas, rode
latexmk.
Limitações conhecidas
PDFs escaneados não têm camada de texto. Rode
ocrmypdfantes.Tabelas e figuras são extraídas como texto corrido; afirmações sustentadas só por uma figura precisam de verificação manual.
A busca é lexical. Para uma afirmação em português e uma fonte em inglês, o modelo passa os termos em inglês em
query(os prompts já instruem isso). Os números servem de âncora entre idiomas.Ainda não lê
.docx(é o próximo formato a suportar).
Licença
MIT. Veja LICENSE.
Available Tools
19 toolsapply_rewriteA
Write a proposed rewrite into the manuscript file (atomic write, original encoding and line endings kept, backup in .citecheck/backups). Only call after the user approved the diff.
| Name | Required | Description | Default |
|---|---|---|---|
| rewrite_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full disclosure burden and does well: atomic write, preservation of original encoding and line endings, and a backup location (.citecheck/backups) are all concrete behavioral traits. It doesn't state whether the rewrite is consumed/marked applied, idempotency, or failure behavior, which are meaningful gaps for a file mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, action and side effects front-loaded, with the gating condition as the closing clause. No filler and nothing redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema, the description covers the essentials an agent needs: what it writes, the safety guarantees, and the approval prerequisite. Missing details about the applied rewrite's subsequent state and error cases are minor against what is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single rewrite_id parameter, so the schema adds no meaning. The description never explains where rewrite_id comes from (presumably list_rewrites/propose_rewrite) or its expected form, though the name and surrounding workflow make it largely inferable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Write a proposed rewrite into the manuscript file'), which clearly implies committing an existing proposal as opposed to propose_rewrite's drafting step. It does not explicitly name the sibling tools (revert_rewrite, discard_rewrite, list_rewrites), so the differentiation is inferential rather than spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Only call after the user approved the diff' is an explicit, actionable precondition that routes the agent correctly in the propose/approve/apply workflow. It stops short of naming the alternatives (e.g. use propose_rewrite first, discard_rewrite to abandon), which would make it a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_pdfA
Attach a local PDF to a reference (for paywalled papers or to replace a wrong match).
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| pdf_path | Yes | Local path to the PDF for this reference. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies overwriting via 'replace a wrong match' but does not state whether an existing attachment is destroyed, what permissions are required, or whether the action is reversible. For a mutation tool with zero annotation coverage, this is a significant disclosure gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core action front-loaded and the usage context trailing in parentheses. Every clause earns its place; nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers purpose and use cases. However, with no annotations it leaves behavioral edges (overwrite behavior, error cases, auth) unspecified, so it is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: pdf_path is documented in the schema, but key is not. The description's phrase 'to a reference' implicitly maps key to a reference identifier, partially compensating, but it adds no format or constraint detail. A 3 reflects the partial bridging of the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (attach), resource (local PDF), and target (a reference). This is a distinct action from any sibling tool like fetch_sources or record_verdict, so an agent can identify it. No explicit sibling differentiation is given, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical names two concrete situations for use: paywalled papers and replacing a wrong match. That is real when-to-use guidance. It stops short of naming an alternative tool or stating exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_statusB
Progress overview: claims per status, sources per status, rewrites per status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it discloses only the aggregation dimensions. It never states that the call is read-only, that it is a computed summary rather than stored data, or what shape/grouping the counts take — though the three listed categories do convey the returned content at a high level.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single compact sentence with the core concept (progress overview) front-loaded and the three dimensions enumerated, so every word earns its place. The sentence-fragment style and lack of a verb are minor demerits.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must explain the return content, which it partly does by naming the three per-status breakdowns. It omits the status vocabulary, the shape of the counts, and any read-only assurance — adequate but with clear gaps for a status/summary tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero properties, so there are no parameters to document; the baseline for a no-param tool is 4. The description does not need to add param syntax and nothing is missing or misleading.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Progress overview: claims per status, sources per status, rewrites per status' states a specific aggregation resource and its three dimensions, so an agent can tell it apart from atomic verbs like get_claim or list_claims. It lacks a full verb phrase and gives no explicit sibling differentiation, but the resource is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when a caller should prefer this over list_claims/get_claim or any other sibling, and no prerequisites. The closest thing to guidance is the implicit assumption that the caller wants a summary rather than raw records.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_bibliographyB
Check the bibliography: cited keys missing from the .bib, unused entries, duplicates, missing DOIs, metadata that differs from Crossref, retractions/corrections, preprints that have a published version.
| Name | Required | Description | Default |
|---|---|---|---|
| keys | No | Only these keys; default: every used key. | |
| online | No | Query Crossref/Semantic Scholar (DOI, metadata, retractions, published versions). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It reveals that the tool may perform network calls via the online parameter but does not state whether the check is read-only, whether it mutates the .bib file, what permissions are needed, or how errors from Crossref are surfaced. For a tool combining network calls and file analysis with zero annotation coverage, this is a significant transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-front-loaded sentence that packs the full scope without filler. The enumerative style is efficient, though the absence of punctuation beyond the initial colon makes it slightly dense; it remains readable and non-redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The described checks are comprehensive for a bibliography audit, matching the complexity of the task. However, with no output schema and no annotations, the description omits how results are returned, whether the tool mutates files, and how it handles network failures, leaving gaps an agent would need to probe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (keys, online) are already documented in the schema, including defaults and the Crossref scope. The description adds only the high-level framing of what online queries cover. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource (the bibliography), then enumerates the exact classes of problems it detects: missing keys, unused entries, duplicates, missing DOIs, Crossref mismatches, retractions, and preprint-to-published cases. This is far more specific than any sibling and tells the agent precisely what it audits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumerated checks, but there is no explicit when-to-use or when-not-to-use guidance, and no routing to siblings like scan_uncited or fetch_sources which could overlap with certain checks (unused entries, DOI fetch). The agent must infer that this is the holistic bibliography audit tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discard_rewriteC
Discard a proposed (not applied) rewrite.
| Name | Required | Description | Default |
|---|---|---|---|
| rewrite_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it only says a proposal is discarded. It does not state whether the discard is reversible, whether confirmation is needed, what happens if the rewrite was already applied or doesn't exist, or what permissions are required for this destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the scoping qualifier up front and no filler. Nothing is wasted; it is simply brief because the description is thin, not because it is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and an undocumented parameter, the description is too sparse. It omits reversibility, error cases (already applied / not found), and any tie-in to the proposal lifecycle (propose_rewrite, list_rewrites, apply_rewrite) that an agent needs to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one required parameter, rewrite_id, with 0% schema description coverage, so the description must compensate. It never mentions the parameter or what value is expected (an id from list_rewrites / propose_rewrite), leaving the identifier semantics implied at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (discard) and resource (a proposed rewrite), and the parenthetical '(not applied)' scopes it away from applied rewrites, implicitly separating it from revert_rewrite. It does not name the sibling explicitly, but the scope qualifier is enough for an agent to distinguish it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '(not applied)' implies usage: call this only for proposals that have not been applied, and use revert_rewrite for applied ones. That condition is implied rather than stated, and no prerequisites or exclusions are given explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_reportC
Write a self-contained HTML report with every claim, verdict, verified quote (with page), rewrite and bibliography problem.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Output .html path. Default: .citecheck/report.html next to the manuscript. | |
| language | No | Report language. Default: the manuscript's language. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the output is 'self-contained HTML' and that a file is written, but says nothing about overwrite behavior of an existing path, permissions, or idempotency of repeated exports – significant gaps for a file-writing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the artifact first, with the report contents enumerated compactly. It is somewhat run-on at the end but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description adequately conveys the report's contents and format. However, for a write tool with no annotations it omits when to invoke it and how it interacts with existing files, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with 0 required parameters, so the schema already documents the path default and language default fully. The description adds no parameter detail beyond what the schema provides, making baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Write') and resource ('a self-contained HTML report') and enumerates what the report aggregates (claims, verdicts, verified quotes with page, rewrites, bibliography problems). This distinguishes it from the check-specific siblings like verify_quote or check_bibliography, though it does not name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites (e.g. checks must have been run first), and no alternatives named. The aggregating role is only implied by the surrounding sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_sourcesA
Locate the PDF of each cited reference (bib file field, PDF folders, Zotero storage, then open-access
repositories), extract its text with page numbers and index it. Reports missing/paywalled sources and
PDFs whose title does not match the reference.
| Name | Required | Description | Default |
|---|---|---|---|
| keys | No | Keys to fetch; default: every cited key. | |
| download | No | Download open-access PDFs (arXiv, Semantic Scholar, OpenAlex, Unpaywall). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: the ordered fallback chain, that text is extracted with page numbers and indexed, and that the run reports missing, paywalled, and title-mismatched PDFs. It stops short of stating the side effects explicitly (files written to disk, index mutation, possible network calls when download is enabled) or permissions/rate-limit considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no wasted words; the resolution order and the reporting outcome are both front-loaded. Every clause earns its place by conveying either behavior or output content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-param, no-output-schema tool with no annotations, the description adequately covers scope, approach, and what gets reported back. It could add a note on side effects (persisted index/PDF files) to be fully complete for a state-mutating batch operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters (keys, download) are already documented in the schema, including the open-access repositories relevant to download. The description adds no syntax or format detail for `keys` beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs and resources ('Locate the PDF of each cited reference', 'extract its text ... and index it'), and spells out the resolution order (bib `file` field → PDF folders → Zotero storage → open-access repositories). This clearly distinguishes it from siblings like attach_pdf (manual attach) and read_source (read a single source), without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The multi-source fallback chain implies this is the batch-acquisition tool, and reporting on missing/paywalled sources hints at a verification workflow context. However, it never explicitly states when to use this versus attach_pdf or read_source, nor any exclusions or prerequisites (e.g., when to prefer manual attachment).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_evidenceA
Rank the source's sentences by relevance to the claim (BM25 + number matching). Each hit has the exact text, page and surrounding context; copy quotes verbatim from here.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| key | Yes | Citation key of the source to search. | |
| query | No | Extra search terms, ideally in the SOURCE's language (needed when the claim is in another language). | |
| claim_id | No | Claim to search evidence for (its text is used as query). | |
| include_references | No | Also search the source's own bibliography section. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the scoring method and the shape of each hit (exact text, page, surrounding context), but never states that it is a read-only operation with no side effects, nor whether it reads from disk/network or has any cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the purpose is front-loaded and the actionable instruction is last. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the return fields (text, page, surrounding context) and how to use them. Nothing critical is missing for a 1-required-parameter evidence search, though a note on the claim_id/query relationship and read-only nature would make it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema already documents key, query, claim_id and include_references. The description adds no parameter-level meaning beyond that (it mentions 'the claim' generically), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Rank the source's sentences') and the relevance criterion ('by relevance to the claim'), plus the retrieval method (BM25 + number matching). An agent knows exactly what this returns, though it does not explicitly contrast itself with siblings like verify_quote or read_source.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the closing instruction 'copy quotes verbatim from here' tells the agent this is the quote-sourcing step, but there is no explicit when-to-use, when-not-to-use, or named alternative (e.g. read_source for full text, verify_quote to check a quote).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_claimB
Full detail for one claim: plain text, raw LaTeX, neighbouring sentences, cited references (with source availability) and existing verdicts.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | ||
| claim_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full disclosure burden, and it does enumerate the returned fields (including source availability flags), which is genuinely useful. However it says nothing about error behavior for invalid claim_ids, permissions, or whether existing verdicts are mutable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a dense, comma-separated inventory of returns. No filler, though the list format loses a little readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read getter with no annotations and no output schema, the return-field inventory compensates reasonably well. But the 'context' parameter's meaning and range are unexplained, leaving a gap an agent must resolve from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 2 parameters. The phrase 'neighbouring sentences' hints at the role of the 'context' parameter but never connects it, and the 0-5 range/default of 1 is left entirely to the schema. claim_id is only self-explanatory by name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Full detail for one claim') and enumerates the exact contents returned: plain text, raw LaTeX, neighbouring sentences, cited references, and verdicts. The singular 'one claim' implicitly distinguishes it from list_claims, though it never names the sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, prerequisites, or alternatives are stated. Nothing tells the agent to prefer this over list_claims when it needs a single claim's detail, nor that a claim_id must already be known.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_claimsC
List cited sentences ('claims') with their keys, location and verification status (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | Only this file (relative path as listed in open_project). | |
| limit | No | ||
| offset | No | ||
| status | No | Filter by status. | all |
| section | No | Only sections whose title contains this text. | |
| include_uncited | No | Also list sentences without citations. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden but discloses little beyond 'paginated'. It does not state what pagination returns, the default filtering to cited-only sentences, auth/permission needs, or cost. The listing behavior is only implicitly conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb and resource leading and no wasted words. It is appropriately sized, though it could sacrifice a word or two to add usage context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter list tool with no output schema and no annotations, the description is minimal but conveys the shape of results. It omits usage guidance, pagination mechanics, and the cited-vs-uncited default, leaving clear gaps for an agent acting without structured support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema already documents file, status, section and include_uncited, while limit/offset are undescribed there. The description adds only 'paginated', which loosely maps to limit/offset but gives no defaults or bounds. Baseline 3 is appropriate given partial schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List), a clearly defined resource ('cited sentences (\'claims\')'), and the returned fields (keys, location, verification status). It is far clearer than a tautology, but it never distinguishes itself from siblings like get_claim or scan_uncited.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no prerequisites, and no mention of alternatives. The only operational hint is the '(paginated)' clause. An agent gets no help deciding this tool over get_claim or scan_uncited.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_rewritesC
List rewrites with their diffs.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It never confirms this is a read-only operation or describes pagination, ordering, or result limits; the one behavioral hint ('with their diffs') largely restates what the existing output schema already covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short, front-loaded sentence with no filler or redundancy. It is efficient, though its brevity borders on under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, and the tool is simple with one optional parameter. However, the description omits the status-filter semantics and any usage context, leaving meaningful gaps for an agent deciding between this and its many sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'status' parameter has 0% schema description coverage, so the description must compensate and does not — it never mentions that results can be filtered by status or what those enum values mean. The enum values (proposed/applied/reverted/discarded) are only self-explanatory because they mirror sibling tool names, not because the definition explains them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List rewrites') and even indicates the content of the listing ('with their diffs'). It is clear what the tool does, but it does nothing to distinguish it from sibling listing/inspection tools like list_claims or audit_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the many sibling rewrite tools (propose_rewrite, apply_rewrite, revert_rewrite, discard_rewrite) or the other listing tools. Nothing is said about filtering, ordering, or prerequisites, so the agent must infer all usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_projectA
Open (or re-open after edits) a manuscript. Parses sentences, citations, the bibliography and restores previous verdicts. Returns an overview: language, counts, keys missing from the .bib.
| Name | Required | Description | Default |
|---|---|---|---|
| bib | No | Extra .bib/CSL-JSON files, if not declared in the manuscript. | |
| pdf_dirs | No | Folders with the cited PDFs (searched recursively). | |
| manuscript | Yes | Path to the main .tex file (or .md). \input/\include files are followed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does substantive work: it discloses side effects (parsing sentences/citations/bibliography, restoring prior verdicts) and specifies the returned overview (language, counts, missing .bib keys). It does not state performance cost, failure modes, or whether restored verdicts can be reset, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences, front-loaded with the core verb and the re-open condition. Every clause carries information (behavior plus return shape). Minor density in the middle sentence keeps it just below maximal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly describes the return value (language, counts, missing .bib keys), and it explains the side effects of loading. Combined with full schema coverage, an agent has enough to invoke it, though it lacks any note on dependencies with sibling tools that consume the opened project.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (bib, pdf_dirs, manuscript) are already documented in the schema, including that \input/\include files are followed. The description adds no parameter-level detail beyond this, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open ... a manuscript') and enumerates the concrete work done: parsing sentences, citations, bibliography, and restoring previous verdicts. This clearly positions it as the session/entry-point loader, distinct from read_source, verify_quote, or check_bibliography. It could be sharper by explicitly naming a sibling it is not, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(or re-open after edits)' implies a when-to-use condition, but there is no explicit statement of alternatives, prerequisites, or when-not to call it (e.g., whether it must precede verify_quote/record_verdict). Guidance is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_insertionB
Propose new text after an existing sentence (e.g. a paragraph drafted from verified evidence). Applied with apply_rewrite like any rewrite.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| language | No | ||
| new_text | Yes | New sentence(s) in the manuscript's markup, with citations. | |
| after_claim_id | Yes | Sentence (or heading) after which the new text goes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral load. It does disclose the key trait that the proposal is staged and later 'applied with apply_rewrite,' implying reversibility via discard/revert siblings. It says nothing about permissions, validation of the claim/sentence target, or what happens if the anchor sentence doesn't exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the action front-loaded and zero filler. 'like any rewrite' is slightly vague but functional, and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core action and its follow-up step, which is the most important missing piece for a staged-mutation tool. It omits what the call returns (a rewrite handle?) and how the proposal is inspected or discarded, which the sibling set suggests matters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: new_text and after_claim_id are documented, while note and language are undocumented in both places, and the description does not compensate. It also loosely calls the anchor an 'existing sentence' when the parameter is a claim_id, adding only marginal meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and effect: 'Propose new text after an existing sentence,' which is clearly distinct from the sibling propose_rewrite (which replaces existing text rather than inserting). It does not name propose_rewrite explicitly, but the insertion semantics are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence ('Applied with apply_rewrite like any rewrite') gives useful workflow context — that this stages a change rather than committing it. However, it never states when to choose this over propose_rewrite, nor any prerequisites, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_rewriteA
Propose a replacement for one sentence. Validates LaTeX (braces, $, special chars), citation keys (none lost, all new ones in the .bib) and language. Nothing is written until apply_rewrite.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Why (e.g. 'match the hedging of the source'). | |
| claim_id | Yes | ||
| language | No | Target language. Default: manuscript language. | |
| new_text | Yes | Replacement for the whole sentence, in the manuscript's markup (LaTeX: keep \cite commands, ~, macros). | |
| allow_citation_changes | No | Allow removing/adding citation keys (e.g. to cite the original source). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: validation of LaTeX constructs, citation-key preservation, and .bib membership, plus the fact that nothing is persisted until apply_rewrite. It omits what happens on validation failure (rejected vs. returned with warnings) and whether the proposal is stored and retrievable, which are the remaining behavioral gaps for a mutation-adjacent tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the scope constraint is front-loaded before the validation contract and the deferral note. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must cover returns and it partially does by framing this as a non-persisting proposal gated behind apply_rewrite. It is nearly complete for a 5-parameter tool, with only failure semantics and proposal lifecycle left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80% and the schema already documents note, language, new_text and allow_citation_changes, so the baseline is 3. The description's citation rule ("none lost, all new ones in the .bib") gives useful context for why allow_citation_changes exists, but it adds no syntax, format or default details beyond what the schema supplies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with bounded scope ("Propose a replacement for one sentence") and explicitly names the sibling apply_rewrite as the separate commit step. An agent can distinguish this from apply_rewrite, revert_rewrite and propose_insertion without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase "Nothing is written until apply_rewrite" clearly establishes the propose-then-apply workflow and implicitly tells the agent when this tool (not apply_rewrite) is the right call. It does not, however, contrast with the adjacent sibling propose_insertion or state when a rewrite is appropriate versus an insertion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_sourceB
Read a source page by page (or from a sentence id), e.g. the abstract, or around a search hit.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| page | No | Page to read (1-based). Default: page 1. | |
| from_sentence | No | Or read sequentially from this sentence id. | |
| max_sentences | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It discloses the paginated/sentence-based reading mode, but does not explicitly confirm read-only safety, describe permissions or rate limits, explain the return format, or mention error behavior for a page/sentence that doesn't exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. The core purpose and reading modes are stated immediately, and the examples earn their place by clarifying intent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, no annotations, no output schema, and only 50% schema description coverage, the definition leaves key and max_sentences unexplained. It is not complete enough for an agent to confidently call the tool without inferring the missing parameter meanings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%. The description references the page and from_sentence modes, which loosely maps to the page and from_sentence parameters. It adds no semantics for the key parameter or max_sentences, so it only partially compensates for the undocumented half of the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Read) and resource (a source), and describes the reading modes (page by page or from a sentence id). It does not explicitly differentiate from siblings like fetch_sources or find_evidence, so sibling routing is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The examples ('the abstract, or around a search hit') imply common use cases, giving some usage context. However, there is no explicit guidance on when to prefer this tool over alternatives such as fetch_sources or find_evidence, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_verdictA
Record the verdict for one (claim, key) pair. Every quote is verified against the PDF first; the verdict is rejected if any quote is not verbatim.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| issues | No | Tags such as overclaim, wrong_number, wrong_population, wrong_condition, misattribution, needs_page, outdated. | |
| quotes | No | Verbatim passages from the source (required for supported/partial/contradicted/secondary). | |
| verdict | Yes | supported: The source states the claim as written.; partial: The source supports part of it, or with caveats (overclaim, different number/population/condition).; contradicted: The source says the opposite, or the numbers/facts conflict.; secondary: The source only repeats a claim from another work; the original should be cited.; not_found: No relevant passage found in the available text. NOT the same as unsupported.; unavailable: The source PDF could not be obtained or has no text layer. | |
| claim_id | Yes | ||
| rationale | Yes | One or two sentences explaining the verdict, in the manuscript's language. | |
| confidence | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral load. It does disclose a real trait – every quote is checked against the PDF and the verdict is rejected if a quote is not verbatim – which is more than the schema states. But it omits what 'rejected' looks like (error vs. silent drop), whether re-recording overwrites a prior verdict, and any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, the core action front-loaded and the enforcement rule immediately after. Nothing is padding and nothing needs to be re-read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no annotations, no output schema, and 57% schema coverage, the description is adequate but thin. It never says what is returned on success, whether recording is idempotent, or how claim_id/key relate to the sibling lookup tools, leaving real gaps for an agent wiring this into a workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 57%, with the verdict enum, quotes, and issues well documented in the schema itself. The description's verbatim-quote rule adds meaning to the quotes parameter, but claim_id, key, and confidence remain undescribed in both places, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource+scope: 'Record the verdict for one (claim, key) pair.' That is precise enough to distinguish it from read-oriented siblings like get_claim or list_claims, though it never names a sibling explicitly to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the quote-verification precondition tells the agent when the call will succeed, but there is no explicit 'use this after X' or comparison to alternatives such as verify_quote or audit_status. An agent can infer the workflow but must supply the context itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revert_rewriteC
Undo an applied rewrite.
| Name | Required | Description | Default |
|---|---|---|---|
| rewrite_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It says 'undo' but does not disclose whether the revert requires authorization, whether it is reversible, whether it restores prior text, or what happens to the rewrite record afterwards. For a mutation tool with zero annotation coverage, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, front-loaded with the action. It is efficient, though arguably too terse for a mutation tool, so it falls short of a full 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a mutation action, no annotations, no output schema, and an undocumented parameter, the description omits most of what an agent needs to invoke this safely. It should at minimum state the precondition and the effect on the underlying claim text.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter with 0% schema description coverage, so the description must compensate and does not. It never explains what rewrite_id is, whether it refers to an applied rewrite specifically, or where the agent obtains it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('undo') and resource ('an applied rewrite'), so the agent knows this reverses an already-applied rewrite. It does not, however, distinguish itself from close siblings like discard_rewrite or propose_rewrite, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'applied rewrite' hints at the precondition, but there is no explicit when-to-use guidance and no mention of alternatives such as discard_rewrite or apply_rewrite. Given four sibling rewrite tools, this omission is significant.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_uncitedB
Find sentences without citations that look like factual claims needing one (numbers, comparisons, 'studies show'-style wording in EN/PT). Heuristic: the host must judge each candidate.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | ||
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one valuable trait: results are heuristics and the host must judge each candidate (i.e., expect false positives, not verdicts). It does not address read-only safety, scope (whole project vs. a file), or the shape of what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action and examples, with the heuristic caveat reserved for the end. No filler, though the phrasing is slightly terse for the amount of schema burden it must carry.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-annotation, no-output-schema tool, the description covers purpose and the candidate/heuristic caveat but omits any explanation of the file vs. project scope and the limit behavior. Adequate but with clear gaps an agent would want filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two parameters, so the description must compensate and it does not: neither 'file' (optional, default null) nor 'limit' (default 30, max 100) is mentioned at all. The agent gets no hint that 'file' scopes the scan or that 'limit' caps candidates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Find) and resource (sentences without citations that look like factual claims), with concrete examples of the patterns it matches (numbers, comparisons, 'studies show' wording, EN/PT). This is clearly distinguishable from verification-oriented siblings like verify_quote or check_bibliography, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: it surfaces candidates for later judgment, which fits as a pre-screening step before tools like verify_quote or find_evidence. However, there is no explicit statement of when to use it versus those siblings, and no exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_quoteB
Check that a quote is really in the source (fuzzy-tolerant to PDF artifacts, threshold 90/100). Returns the page and the exact source passage.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| quote | Yes | Text that should appear verbatim in the source. Use '...' to join separate passages. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the fuzzy matching tolerance for PDF artifacts, the 90/100 threshold, and what is returned (page and exact passage), but says nothing about failure behavior, permissions, or what happens when no match is found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, the core purpose front-loaded, and every clause (tolerance, threshold, return content) adds information without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter verification tool with no output schema and no annotations, the description covers mechanics and return content but omits the meaning of 'key' and any failure/no-match behavior. Adequate but with a clear gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: the 'quote' parameter is documented (including the '...' joining syntax) but 'key' has no schema description and the description never explains it. The description does not compensate for the undocumented parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Check that a quote is really in the source') and adds scope detail (fuzzy tolerance, threshold). It distinguishes the operation from sibling read/scan tools by naming quote verification, but does not explicitly contrast with close siblings like find_evidence or check_bibliography.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for confirming quotes exist, but gives no when-to-use guidance, no alternatives, and no prerequisites. With siblings such as find_evidence and check_bibliography, an agent gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
19 tool updates
v0.1.0- First observed
apply_rewrite - First observed
attach_pdf - First observed
audit_status - First observed
check_bibliography - First observed
discard_rewrite - First observed
export_report - First observed
fetch_sources - First observed
find_evidence - First observed
get_claim - First observed
list_claims - First observed
list_rewrites - First observed
open_project - First observed
propose_insertion - First observed
propose_rewrite - First observed
read_source - First observed
record_verdict - First observed
revert_rewrite - First observed
scan_uncited - First observed
verify_quote
TDQS
Scored across 19 tools
Each tool has a clearly distinct role: verify_quote checks, find_evidence ranks, and record_verdict persists, while read_source, open_project, fetch_sources, and attach_pdf all target different resource stages. The rewrite lifecycle (propose_rewrite, propose_insertion, apply_rewrite, revert_rewrite, discard_rewrite, list_rewrites) is also cleanly separated by action.
Every tool follows a consistent verb_noun snake_case pattern (verify_quote, open_project, list_claims, apply_rewrite, etc.). There are no mixed conventions or vague verbs; even the rewrite variants follow a predictable propose/apply/revert/discard/list scheme.
19 tools is on the heavy side of the ideal 3-15 range, but the domain (citation checking of a manuscript) is genuinely multi-stage and each tool earns its place. No redundant tools inflate the count.
The surface covers the full workflow: opening a project, listing/getting claims, verifying and recording verdicts, bibliography checks, source fetching, rewrite proposal through rollback, and report export. Minor gaps like closing/archiving a project or a search-within-bibliography helper exist but are workable.
Maintenance
Related MCP Connectors
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Catch AI-fabricated citations (real DOI + fake title). Retraction, open-access, 10,000+ CSL styles.
Real-time fact-check, citation verification, and source-freshness for AI agents.
Verifies legal citations vs primary sources: existence, quote match, proposition support.
Related MCP Servers
- AlicenseAqualityDmaintenancePrevents citation hallucination by verifying academic citations against CrossRef's database of 150+ million publications before they can be mentioned, ensuring every citation includes a valid DOI.315MIT

CiteStamp MCP serverofficial
AlicenseNot gradedqualityBmaintenanceGround citations before your agent emits them by checking references against public scholarly registries and flagging hallucinated or retracted ones.MIT- AlicenseAqualityBmaintenanceProvides local academic literature search and writing support by querying OpenAlex, arXiv, and Crossref, downloading open-access PDFs, extracting IMRaD sections, and appending BibTeX references.42MIT
- AlicenseAqualityBmaintenanceEnables AI writing assistants to retrieve citable evidence from local PDFs and verify draft citations against their sources locally, providing verifiable support for claims.21MIT