Skip to main content
Glama

litcheck-mcp

CI

A prior-work checker for scientists and their AI agents. litcheck checks quoted passages against pinned open-access PubMed Central texts and checks retraction status; every quote check and every search it runs goes into an append-only, hash-chained log.

It never tells you a hypothesis is novel. It gives you a dated record of what was searched and what was found, and, where open-access text exists, checks whether each quoted passage appears in it.

Install

Clone it, then add it to your MCP client (uv must be on your PATH):

git clone https://github.com/javrodriguez/litcheck-mcp

Claude Code:

claude mcp add litcheck -- uv --directory /ABS/PATH/TO/litcheck-mcp run litcheck-mcp

Codex:

codex mcp add litcheck -- uv --directory /ABS/PATH/TO/litcheck-mcp run litcheck-mcp

Replace /ABS/PATH/TO/litcheck-mcp with the clone's absolute path. The first launch installs the server's dependencies (mcp, pydantic, and the small development group: pytest, pytest-asyncio, ruff, mypy) into the clone's own .venv. The server needs Python 3.12 or 3.13; uv fetches one if none is installed.

Related MCP server: traceable-research-mcp

Worked example

docs/EXAMPLE.md takes one CC BY paper through the whole loop: a true quote comes back FOUND with its offsets, the same quote with one word changed comes back NOT_FOUND, the paper's retraction status is reported with its sources, and the log's chain is verified. Every output on that page is re-derived from recorded responses by scripts/check_readme.py, so it shows what the code does.

What it checks, and what it does not

  • Quotes, against the paper's open-access text from the PMC Cloud Service: the latest version, recorded by version number and sha256 (a later check may meet a newer version), and accepted only if its bytes match the md5 in PMC's metadata. Matching is an exact substring search after a pinned normalisation (q1: Unicode NFKC, straight quotes, one dash, collapsed whitespace; case kept); it does not respect word boundaries, so a quote that starts or ends mid-word can still be FOUND. The text searched is PMC's whole plain-text file: a metadata header (journal, identifiers, affiliations), the article, and its reference list, so a cited work's title is FOUND too. Paragraph breaks become single spaces, so a match can also run across two blocks; paragraph names the block it starts in. Check it, or read the passage, before saying the paper itself states it. Verdicts: FOUND, NOT_FOUND, TOO_SHORT (under 20 characters), NOT_CHECKABLE (no open-access text to check: none in PMC, PMC's copy missing, or the identifier not found at all; the resolved status and the reason say which; a short quote with no text to check is NOT_CHECKABLE, not TOO_SHORT), UNVERIFIABLE (a source could not be read or verified, or litcheck itself failed; the reason says which).

  • Only open-access full text can be quote-checked. Paywalled papers, and author manuscripts PMC does not mark as open access, come back NOT_CHECKABLE.

  • Retraction status, from PMC's is_retracted flag and Crossref's notices about the DOI (the notice records filter=updates:<doi> returns: retraction, withdrawal and removal count; corrections do not; an expression of concern is reported separately; a retraction recorded only in the paper's own updated-by field is not read). NOT_RETRACTED_AS_OF <date> means no retraction notice was found in the sources consulted, as of that date; it needs the identifier to have resolved and every consulted source to have answered, otherwise the status is UNVERIFIABLE. PMC's flag is read for papers PMC lists (if their metadata cannot be read, PMC's copy is missing, or the PMC Cloud Service holds no copy of a paper PMC lists, that counts as a source that did not answer); a paper PMC does not hold is checked on Crossref alone, and a paper with no known DOI on PMC's flag alone; sources shows which answered. A PMID or PMCID outside PMC has no DOI to ask Crossref about: pass the DOI instead. Retraction checks are returned, not logged, except inside a check_quote evidence line.

  • No support verdict is computed. Whether a passage supports, contradicts or is absent from a claim is recorded only when a person or a named judge gives it (record_support).

  • Searches are recorded, and a search can miss papers. Europe PMC, LitSense 2.0, PubMed and OpenAlex citations are queried as asked; each query, its parameters, the sha256 of the search response (for PubMed the esearch response, not the esummary that adds titles) and one id per returned item go into the log. Searches that were sent and failed are recorded too; a request rejected before sending (an empty query, say) is not. An empty result is a record of one query, not evidence that nothing exists.

Tools

Tool

What it does

check_quote

resolve the paper, fetch its pinned open-access text, look for the quote; licence and retraction status alongside

resolve_identifier

DOI, PMID or PMCID to the other two (PMC ID converter, Europe PMC fallback for DOIs only: a PMID or PMCID outside PMC comes back NOT_FOUND, which does not mean it does not exist)

retraction_status

PMC's flag and Crossref's notices, with the date and response hash of each check (returned, not logged)

search_literature

one recorded search on europe_pmc, litsense or pubmed

citing_papers

works citing a paper, from OpenAlex, recorded

verify_log

re-walk the log's hash chain

record_support

store a person's or a named judge's verdict (supports, contradicts, absent) about an evidence line

Command line

The core is standard-library Python that runs on Python 3.6 or newer, with no install (./litcheck is a POSIX shell wrapper; on Windows run python -m litcheck with core on PYTHONPATH, and note that appends to the log are not locked there, so two writers, including concurrent calls to the MCP server on Windows, can break the chain):

./litcheck check --id PMC10496602 --quote "Arteriosclerosis consists of functional depletion of large-artery elasticity."
./litcheck search litsense "large-artery elasticity" --limit 5
./litcheck verify-log
./litcheck annotate --seq 1 --support supports --by "your name"

check exits 0 only for FOUND; verify-log exits 0 only for an intact chain.

The log

One JSON object per line: seq, a UTC timestamp, the kind (evidence, search, annotate), the sha256 of the previous line, its payload, and its own sha256. Nothing in litcheck rewrites a line. verify-log (or verify_log) reports the first line that does not fit, which catches accidental and naive edits. The hashes are not keyed: someone who rewrites a line and recomputes every hash after it, or cuts off the tail, leaves a chain that checks out. To detect that, keep the head hash verify-log prints somewhere else; later, the sha256 of line N's bytes (without its newline) must still equal the head you saved when the log had N lines. litcheck has no command for that comparison yet. The chain shows order and integrity, not authorship: anyone who can write the file can append well-formed lines, and each line's timestamp is the writing machine's clock. A log that does not exist yet is reported as not intact (exit 1, chain_ok: false), and a damaged last line stops new checks and searches from being written until it is dealt with.

System

Default location

Windows

%LOCALAPPDATA%\litcheck\log.jsonl

macOS

~/Library/Application Support/litcheck/log.jsonl

Linux

$XDG_DATA_HOME/litcheck/log.jsonl, else ~/.local/share/litcheck/log.jsonl

Set LITCHECK_LOG to put it elsewhere (for one project, say); --log does the same per call (verify-log takes the path as its argument). Only check_quote, the searches and record_support write to the log; resolve_identifier and retraction_status do not.

Etiquette

Set LITCHECK_CONTACT to an address the services can reach you at. litcheck sends it only with each request: in the User-Agent, to every host it calls (the PMC bucket on Amazon S3 included), and as the email/mailto parameter NCBI, Europe PMC, Crossref and OpenAlex ask for. It is not part of any URL litcheck stores, and it is never written to a recorded fixture; a value that is not printable ASCII is ignored. With Claude Code, add -e LITCHECK_CONTACT=you@example.org before --; with Codex, --env LITCHECK_CONTACT=you@example.org.

Rate limits are honoured per host within one process, concurrent tool calls included (separate CLI runs or servers do not share the spacing): NCBI E-utilities and the ID converter 3 requests a second (no API key), LitSense one a second, Europe PMC 5 a second, OpenAlex and Crossref 10 a second. Answers 429, 500, 502, 503 and 504, and network errors, are retried at most twice, with backoff (honouring a Retry-After in seconds, up to 10); other answers are not retried. A call that still fails makes the result UNVERIFIABLE, with one exception: when PubMed's esearch answers but the esummary that adds titles fails, the search is SEARCHED with ids only and a reason saying the titles are missing.

Tests

uv sync --dev
uv run pytest -q                                              # the MCP server
python3 -m unittest discover -s core/tests -t core            # the core, offline
uv run python scripts/check_readme.py                         # README and worked example

The core's tests replay responses recorded live by core/tools/record_fixture.py; a replay refuses any URL it was not given and any body whose sha256 differs from its record. LITCHECK_NETWORK_TESTS=1 runs the live smoke test against the real services.

Credits

litcheck reads public services and is grateful to the people who run them: Europe PMC (EMBL-EBI); NCBI, for E-utilities, LitSense 2.0, the PMC Cloud Service and the PMC ID Converter; Crossref, including the Retraction Watch data it distributes; and OpenAlex.

The pinned-PMC gate and quote matcher are ported from peerpanel.

Built by Javier Rodríguez Hernáez · LinkedIn · GitHub

Built for GARS, the genomics agentic research system.

Licence

MIT, for the source code. The recorded responses under core/tests/fixtures/ belong to the services that served them; the one article text among them is CC BY and attributed in docs/EXAMPLE.md.

Available Tools

7 tools
check_quoteA

Check whether a quoted passage appears in the paper's pinned open-access PMC text.

Also reports the text's version, sha256 and licence, and the paper's retraction status, and appends an evidence line to the log. The text searched is PMC's whole plain-text file, header and reference list included, so check paragraph before attributing a FOUND passage to the paper. Exact substring match (word boundaries are not checked) after normalisation q1 (Unicode NFKC, straight quotes, one dash, collapsed whitespace; case kept). It does not judge whether the passage supports the claim.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimNoOptional: the claim you are citing this passage for, kept in the log
quoteYesThe passage exactly as you would quote it (20 characters or more; shorter is TOO_SHORT when there is text to check)
identifierYesOne paper: a DOI ("10.1590/1516-3180.2016.1344090516" or a doi.org URL), a PMCID ("PMC10496602") or a PMID written with its cue ("PMID 27355798")

Output Schema

ParametersJSON Schema
NameRequiredDescription
claimYes
quoteYes
reasonYes
log_seqYes
verdictYesFOUND: the quote occurs, as a substring, in the pinned open-access text (which includes its metadata header and reference list; see paragraph). NOT_FOUND: it does not. TOO_SHORT: under 20 characters, not searched. NOT_CHECKABLE: no open-access text in PMC, PMC's copy missing, or the identifier was not found (see resolved.status and reason). UNVERIFIABLE: a source could not be read or verified, or litcheck failed (see reason); nothing was concluded.
log_pathYes
resolvedYes
paragraphYes0-based index of the blank-line-separated block of the text
identifierYes
retractionYes
pmc_outcomeYes
pmc_versionYes
text_md5_okYes
text_sha256Yes
license_codeYes
normalisationYes
quote_offsetsYes[start, end) of the match in the normalised text
casefold_foundYesThe quote occurs when case is ignored

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses a side effect ('appends an evidence line to the log'), the exact matching semantics (substring, no word boundaries, after normalisation q1), the text scope searched, and the auxiliary facts returned (version, sha256, licence, retraction status). This is substantially more than a bare 'check' verb would convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then layered with scope, semantics, and caveats in separate sentences. Slightly verbose with parenthetical detail about normalisation, but every sentence carries usable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, and the description still covers the write side effect, matching semantics, and the scope caveat needed to interpret a FOUND result correctly. Nothing an agent needs to call this safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters, giving a baseline of 3. The description adds only matching/normalisation context rather than new per-parameter meaning, so it does not rise above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check') and resource ('whether a quoted passage appears in the paper's pinned open-access PMC text'), which an agent can distinguish from siblings like retraction_status or record_support. The scope (PMC full plain-text file including header and references) is unusually precise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear exclusion ('It does not judge whether the passage supports the claim') and warns to check `paragraph` before attributing a FOUND passage, which is implied usage guidance. However, it never states when to prefer this tool over siblings such as record_support or verify_log, so routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

citing_papersB

List works that cite a paper, from OpenAlex, and record the search in the log.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many items to return (1-100)
identifierYesThe cited work: a DOI or an OpenAlex work id (W123...)

Output Schema

ParametersJSON Schema
NameRequiredDescription
itemsYes
queryYes
engineYes
reasonYes
statusYesUNVERIFIABLE means the search did not complete: an empty item list then says nothing about what exists
log_seqYes
log_pathYes
hit_countYes
response_sha256Yes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, but it does disclose a real side effect: the search is recorded in the log. It says nothing about pagination, rate limits, or whether the log write can fail, so the behavioral picture is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence naming the action, resource, source, and side effect with zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and both parameters are documented in the schema. The description covers purpose, source, and the logging side effect, leaving only alternative-tool routing unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (identifier, limit) are already documented in the schema. The description only alludes to the identifier ('a paper') and never mentions the limit, so it adds little beyond structured data; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('works that cite a paper'), plus the data source (OpenAlex), which an agent can act on immediately. It does not name or differentiate itself from the sibling search_literature, which is the nearest overlapping capability, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of when a different sibling (e.g. search_literature, resolve_identifier) is the better choice, and no prerequisites stated. The only usage signal is implied by the description's subject matter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_supportB

Record a person's or a named judge's support verdict about an evidence line.

litcheck computes no support verdict itself; this stores one that was given, with who gave it, as a new line in the log.

ParametersJSON Schema
NameRequiredDescriptionDefault
byYesWho judged: a person's name or a named judge; litcheck never does
seqYesThe log line (an evidence line) being judged
supportYesThe verdict: does the quoted passage support the claim?

Output Schema

ParametersJSON Schema
NameRequiredDescription
byYes
seqYesThe annotation's own line number in the log
supportYes
log_pathYes
target_seqYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does usefully disclose that the record is appended 'as a new line in the log', implying append-only, non-overwriting behavior. However it says nothing about permissions, idempotency, duplicate verdicts, or what happens if the referenced seq is invalid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by one clarifying sentence about what the tool does not do. No filler, though the line-broken prose is slightly looser than necessary for a three-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the append-to-log behavior is stated. But for a mutating tool with zero annotations, the description leaves permission requirements, error handling, and duplicate/conflicting verdict behavior unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents seq, support, and by (including the enum values). The description only echoes the 'who gave it' idea, adding no format, constraint, or interpretation beyond the structured fields, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource: it records a support verdict about an evidence line, and it clarifies the semantics of 'support' as a verdict given by someone rather than computed. It implicitly separates itself from verification-style siblings (check_quote, verify_log) by stressing that litcheck computes no verdict, though it never names a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: 'litcheck computes no support verdict itself; this stores one that was given' tells the agent this is the write path for an externally-supplied verdict, but there is no explicit when-to-use, when-not-to-use, or named alternative among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_identifierA

Map one DOI, PMID or PMCID to the other two, through the PMC ID converter.

The converter knows only articles in PubMed Central; a DOI it cannot map is also looked up in Europe PMC. NOT_FOUND means not in PMC (and, for a DOI, not in Europe PMC either): a PMID or PMCID is not looked up elsewhere, so NOT_FOUND does not mean the paper does not exist. UNVERIFIABLE means a service could not be read.

ParametersJSON Schema
NameRequiredDescriptionDefault
identifierYesOne paper: a DOI ("10.1590/1516-3180.2016.1344090516" or a doi.org URL), a PMCID ("PMC10496602") or a PMID written with its cue ("PMID 27355798")

Output Schema

ParametersJSON Schema
NameRequiredDescription
doiYes
pmidYes
pmcidYes
reasonYesWhy it is not RESOLVED, or why there is no PMCID
sourceYesWhich service answered: pmc-idconv or europe-pmc
statusYesNOT_FOUND: not in PubMed Central (and, for a DOI, not in Europe PMC); a PMID or PMCID is not looked up elsewhere, so NOT_FOUND does not mean it does not exist
identifierYesThe identifier as it was given

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the converter's PMC-only coverage, the Europe PMC fallback that applies to DOIs but not PMIDs/PMCIDs, and the exact meaning of NOT_FOUND and UNVERIFIABLE, including the important caveat that NOT_FOUND does not mean the paper is nonexistent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, and the remaining sentences are dense but entirely devoted to failure-mode semantics that an agent cannot get elsewhere. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with a full output schema and no annotations, the description covers everything an agent needs: what it converts, coverage limits, fallback behavior, and error-code interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents the accepted formats (DOI, doi.org URL, PMCID, cued PMID). The description adds no syntax or format detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: mapping one of DOI/PMID/PMCID to the other two via the PMC ID converter. The scope is narrow and unambiguous, and clearly distinct from siblings like search_literature, citing_papers, and retraction_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: use it when you hold one identifier and need the others. It never names an alternative tool or an explicit when-not-to-use condition, so the agent must infer the routing from the sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retraction_statusA

Is the paper retracted? Reads PMC's is_retracted flag and Crossref's notices.

Retraction, withdrawal and removal notices count; corrections do not; an expression of concern is reported separately. NOT_RETRACTED_AS_OF means no retraction notice was found in the sources consulted, as of its date; it needs the identifier to have resolved and every consulted source to have answered, otherwise the status is UNVERIFIABLE. A paper PMC does not hold is checked on Crossref alone, and one with no known DOI on PMC's flag alone; sources shows which answered. For a paper outside PMC, pass its DOI. Not logged.

ParametersJSON Schema
NameRequiredDescriptionDefault
identifierYesOne paper: a DOI ("10.1590/1516-3180.2016.1344090516" or a doi.org URL), a PMCID ("PMC10496602") or a PMID written with its cue ("PMID 27355798")

Output Schema

ParametersJSON Schema
NameRequiredDescription
as_ofYesFor NOT_RETRACTED_AS_OF: the date of the earliest check; no retraction notice was found in the sources consulted as of then
reasonYes
statusYes
concernYesAn expression of concern is on record
sourcesYes
identifierNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it defines NOT_RETRACTED_AS_OF, explains that it is only returned when the identifier resolved and all sources answered, describes the UNVERIFIABLE fallback, documents source-fallback behavior for papers not in PMC, exposes the `sources` field, and notes 'Not logged.' These are genuine behavioral traits beyond any structured field.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core question and behavior, then organized into status definitions and edge cases; every sentence carries information. It is dense and slightly long for a one-parameter tool, but the length is mostly justified by the status semantics that cannot live in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be explained, yet the description still clarifies the `sources` field and the meaning of each status, fully covering the edge cases (non-PMC papers, missing DOI, unresolved identifiers). Nothing an agent needs to invoke and interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: it clarifies the identifier addresses a single paper and gives routing advice ('For a paper outside PMC, pass its DOI'), plus the requirement that the identifier resolve for a definitive status. That is meaningful meaning beyond the format examples in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific question-verb framing ('Is the paper retracted?') and names the exact mechanism ('Reads PMC's is_retracted flag and Crossref's notices'). An agent can distinguish this from siblings like check_quote or resolve_identifier, which concern quotations and identifier resolution respectively.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear routing context: 'For a paper outside PMC, pass its DOI,' and explains that corrections are excluded while expressions of concern are reported separately. It stops short of explicitly naming a sibling alternative or a when-not-to-use condition, so it is strong context without full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_literatureA

Run one literature search and record it, with its response hash, in the log.

hit_count is what the engine reported (for litsense, the number of results it returned, at most 100); items are the first limit of them. A search can miss papers: status SEARCHED with few items is a record of this query, not of the field.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many items to return (1-100)
queryYesThe query, in the chosen engine's syntax
engineYeseurope_pmc: Europe PMC query syntax (TITLE_ABS:, FIRST_PDATE:[a TO b], ...); litsense: a sentence-level semantic search of PubMed and PMC (no date filter); pubmed: PubMed query syntax through E-utilities

Output Schema

ParametersJSON Schema
NameRequiredDescription
itemsYes
queryYes
engineYes
reasonYes
statusYesUNVERIFIABLE means the search did not complete: an empty item list then says nothing about what exists
log_seqYes
log_pathYes
hit_countYes
response_sha256Yes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the call has a persistent side effect (logging the query with a response hash), explains that hit_count reflects what the engine reported while items are only the first `limit`, and warns that a SEARCHED status with few items reflects the query rather than the field. It stops short of covering failure modes, auth, or engine-specific limits beyond litsense's cap of 100.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action and its logging side effect are front-loaded in the first sentence, and every subsequent sentence adds interpretive value. The parenthetical about litsense and the 100 cap is slightly dense, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not required, yet the description still clarifies hit_count vs items semantics, which is the main interpretive risk. For a three-parameter tool with a full schema this is close to complete, with only error/edge behavior left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema by explaining the relationship between `limit` and the returned items and the distinct meaning of hit_count. The engine-specific syntax nuance is left to the enum descriptions rather than repeated here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Run one literature search') and adds the side effect ('record it, with its response hash, in the log'). No sibling tool performs a search, so it is unambiguously distinguishable from check_quote, resolve_identifier, citing_papers, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the fact that this is the only search tool among the siblings, and the note about missing papers hints at how to interpret results. However, there is no explicit 'when to use this vs alternatives' guidance, no prerequisites, and no statement of when not to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_logA

Re-walk the log's hash chain and report the first line that does not fit, if any.

This catches accidental and naive edits. The hashes are not keyed, so a rewrite that recomputes every later hash, or a cut-off tail, is caught only by comparing head_sha256 with a copy kept elsewhere. A log that does not exist yet reports chain_ok false.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
pathYes
linesYes
reasonYes
chain_okYes
head_sha256Yessha256 of the last line's bytes (without its newline; not its own sha256 field). The chain catches accidental and naive edits; a rewrite that recomputes every hash, or a cut-off tail, is caught only by comparing against a head hash kept elsewhere
first_bad_seqYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the detection model, its key limitation (unkeyed hashes; recomputed chains and truncated tails evade it), and an edge case (nonexistent log returns chain_ok false). It omits side-effect/read-only and cost/performance characteristics, which keeps it just short of 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, followed by the value proposition and the adversarial-limitation caveat. Every sentence earns its place, though the limitation sentence is dense and could be split or tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not required, yet the description still usefully explains the chain_ok=false edge case. Combined with a zero-parameter schema, the definition is essentially complete; only security/side-effect framing is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is no parameter syntax for the description to clarify, and it correctly adds no redundant parameter discussion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Re-walk the log's hash chain and report the first line that does not fit.' It also states the exact output of interest (first failing line, if any), so an agent knows both the action and the result shape without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implicitly defines the use case ('catches accidental and naive edits') and, importantly, states the condition under which the tool is NOT sufficient (rewrites that recompute later hashes or cut tails are only caught by comparing head_sha256 with an external copy). No sibling among check_quote/search_literature/etc. competes for this job, so explicit alternative routing isn't needed, but it stops short of saying when in a workflow to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.0
    • First observedcheck_quote
    • First observedciting_papers
    • First observedrecord_support
    • First observedresolve_identifier
    • First observedretraction_status
    • First observedsearch_literature
    • First observedverify_log

TDQS

A4/5.0

Scored across 7 tools

Disambiguation5/5

Each tool targets a clearly distinct task: quote checking, identifier mapping, log integrity, retraction status, literature search, citation listing, and support recording. Although retraction_status and check_quote both read PMC, their purposes (retraction notices vs. textual quote matching) are unambiguous from the descriptions.

Naming Consistency4/5

Most tools follow a verb_noun pattern (check_quote, resolve_identifier, verify_log, search_literature, record_support). Two deviate into noun phrases (retraction_status, citing_papers), which is a minor inconsistency but still readable and predictable.

Tool Count5/5

Seven tools is well-scoped for a literature-verification server, and each tool earns its place with a distinct role. No redundancy or filler tools.

Completeness4/5

The surface covers the core verification lifecycle: quote checking, ID resolution, retraction checking, searching, citations, and support recording plus log auditing. Minor gaps exist (e.g., no tool to fetch full paper metadata or export/report the log), but agents can work around them.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables scientific due diligence by grading claims against public literature, clinical trials, and filings, with explicit citations and optional attestation.
    4
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables traceable scholarly literature reviews using free APIs, generating reports where every claim links to evidence IDs.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Ground citations before your agent emits them by checking references against public scholarly registries and flagging hallucinated or retracted ones.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to conduct academic research workflows such as paper discovery, literature mapping, citation chasing, author pivots, citation repair, and regulatory or species document retrieval.
    MIT