litcheck-mcp
Fetches pinned open-access full texts from PMC's Amazon S3 bucket for quote checking, verifying the retrieved bytes against PMC metadata.
Provides recorded literature searches against PubMed via NCBI E-utilities, returning matching IDs and optional titles, and logs each query, parameters, and response hash.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@litcheck-mcpCheck if this quote is in PMC1234567: 'CRISPR enables gene editing'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
litcheck-mcp
A prior-work checker for scientists and their AI agents. litcheck checks quoted passages against pinned open-access PubMed Central texts and checks retraction status; every quote check and every search it runs goes into an append-only, hash-chained log.
It never tells you a hypothesis is novel. It gives you a dated record of what was searched and what was found, and, where open-access text exists, checks whether each quoted passage appears in it.
Install
Clone it, then add it to your MCP client (uv must be on your PATH):
git clone https://github.com/javrodriguez/litcheck-mcpClaude Code:
claude mcp add litcheck -- uv --directory /ABS/PATH/TO/litcheck-mcp run litcheck-mcpCodex:
codex mcp add litcheck -- uv --directory /ABS/PATH/TO/litcheck-mcp run litcheck-mcpReplace /ABS/PATH/TO/litcheck-mcp with the clone's absolute path. The first launch installs
the server's dependencies (mcp, pydantic, and the small development group: pytest,
pytest-asyncio, ruff, mypy) into the clone's own .venv. The server needs Python 3.12 or 3.13; uv fetches one
if none is installed.
Related MCP server: traceable-research-mcp
Worked example
docs/EXAMPLE.md takes one CC BY paper through the whole loop: a true quote
comes back FOUND with its offsets, the same quote with one word changed comes back
NOT_FOUND, the paper's retraction status is reported with its sources, and the log's chain
is verified. Every output on that page is re-derived from recorded responses by
scripts/check_readme.py, so it shows what the code does.
What it checks, and what it does not
Quotes, against the paper's open-access text from the PMC Cloud Service: the latest version, recorded by version number and sha256 (a later check may meet a newer version), and accepted only if its bytes match the md5 in PMC's metadata. Matching is an exact substring search after a pinned normalisation (
q1: Unicode NFKC, straight quotes, one dash, collapsed whitespace; case kept); it does not respect word boundaries, so a quote that starts or ends mid-word can still beFOUND. The text searched is PMC's whole plain-text file: a metadata header (journal, identifiers, affiliations), the article, and its reference list, so a cited work's title isFOUNDtoo. Paragraph breaks become single spaces, so a match can also run across two blocks;paragraphnames the block it starts in. Check it, or read the passage, before saying the paper itself states it. Verdicts:FOUND,NOT_FOUND,TOO_SHORT(under 20 characters),NOT_CHECKABLE(no open-access text to check: none in PMC, PMC's copy missing, or the identifier not found at all; theresolvedstatus and thereasonsay which; a short quote with no text to check isNOT_CHECKABLE, notTOO_SHORT),UNVERIFIABLE(a source could not be read or verified, or litcheck itself failed; thereasonsays which).Only open-access full text can be quote-checked. Paywalled papers, and author manuscripts PMC does not mark as open access, come back
NOT_CHECKABLE.Retraction status, from PMC's
is_retractedflag and Crossref's notices about the DOI (the notice recordsfilter=updates:<doi>returns: retraction, withdrawal and removal count; corrections do not; an expression of concern is reported separately; a retraction recorded only in the paper's ownupdated-byfield is not read).NOT_RETRACTED_AS_OF <date>means no retraction notice was found in the sources consulted, as of that date; it needs the identifier to have resolved and every consulted source to have answered, otherwise the status isUNVERIFIABLE. PMC's flag is read for papers PMC lists (if their metadata cannot be read, PMC's copy is missing, or the PMC Cloud Service holds no copy of a paper PMC lists, that counts as a source that did not answer); a paper PMC does not hold is checked on Crossref alone, and a paper with no known DOI on PMC's flag alone;sourcesshows which answered. A PMID or PMCID outside PMC has no DOI to ask Crossref about: pass the DOI instead. Retraction checks are returned, not logged, except inside acheck_quoteevidence line.No support verdict is computed. Whether a passage supports, contradicts or is absent from a claim is recorded only when a person or a named judge gives it (
record_support).Searches are recorded, and a search can miss papers. Europe PMC, LitSense 2.0, PubMed and OpenAlex citations are queried as asked; each query, its parameters, the sha256 of the search response (for PubMed the esearch response, not the esummary that adds titles) and one id per returned item go into the log. Searches that were sent and failed are recorded too; a request rejected before sending (an empty query, say) is not. An empty result is a record of one query, not evidence that nothing exists.
Tools
Tool | What it does |
| resolve the paper, fetch its pinned open-access text, look for the quote; licence and retraction status alongside |
| DOI, PMID or PMCID to the other two (PMC ID converter, Europe PMC fallback for DOIs only: a PMID or PMCID outside PMC comes back |
| PMC's flag and Crossref's notices, with the date and response hash of each check (returned, not logged) |
| one recorded search on |
| works citing a paper, from OpenAlex, recorded |
| re-walk the log's hash chain |
| store a person's or a named judge's verdict ( |
Command line
The core is standard-library Python that runs on Python 3.6 or newer, with no install
(./litcheck is a POSIX shell wrapper; on Windows run python -m litcheck with core on
PYTHONPATH, and note that appends to the log are not locked there, so two writers, including
concurrent calls to the MCP server on Windows, can break the chain):
./litcheck check --id PMC10496602 --quote "Arteriosclerosis consists of functional depletion of large-artery elasticity."
./litcheck search litsense "large-artery elasticity" --limit 5
./litcheck verify-log
./litcheck annotate --seq 1 --support supports --by "your name"check exits 0 only for FOUND; verify-log exits 0 only for an intact chain.
The log
One JSON object per line: seq, a UTC timestamp, the kind (evidence, search, annotate),
the sha256 of the previous line, its payload, and its own sha256. Nothing in litcheck rewrites
a line. verify-log (or verify_log) reports the first line that does not fit, which catches
accidental and naive edits. The hashes are not keyed: someone who rewrites a line and
recomputes every hash after it, or cuts off the tail, leaves a chain that checks out. To
detect that, keep the head hash verify-log prints somewhere else; later, the sha256 of
line N's bytes (without its newline) must still equal the head you saved when the log had N
lines. litcheck has no command for that comparison yet. The chain shows order and
integrity, not authorship: anyone who can write the file can append well-formed lines, and
each line's timestamp is the writing machine's clock.
A log that does not exist yet is reported as not intact (exit 1, chain_ok: false), and a
damaged last line stops new checks and searches from being written until it is dealt with.
System | Default location |
Windows |
|
macOS |
|
Linux |
|
Set LITCHECK_LOG to put it elsewhere (for one project, say); --log does the same per call
(verify-log takes the path as its argument). Only check_quote, the searches and
record_support write to the log; resolve_identifier and retraction_status do not.
Etiquette
Set LITCHECK_CONTACT to an address the services can reach you at. litcheck sends it only
with each request: in the User-Agent, to every host it calls (the PMC bucket on Amazon S3
included), and as the email/mailto parameter NCBI, Europe PMC, Crossref and OpenAlex ask
for. It is not part of any URL litcheck stores, and it is never written to a recorded
fixture; a value that is not printable ASCII is ignored. With Claude Code, add
-e LITCHECK_CONTACT=you@example.org before --; with Codex,
--env LITCHECK_CONTACT=you@example.org.
Rate limits are honoured per host within one process, concurrent tool calls included
(separate CLI runs or servers do not share the spacing): NCBI E-utilities and the ID
converter 3 requests a second (no API key), LitSense one a second, Europe PMC 5 a second,
OpenAlex and Crossref 10 a second. Answers 429, 500, 502, 503 and 504, and network errors,
are retried at most twice, with backoff (honouring a Retry-After in seconds, up to 10);
other answers are not retried. A call that still fails makes the result UNVERIFIABLE, with
one exception: when PubMed's esearch answers but the esummary that adds titles fails, the
search is SEARCHED with ids only and a reason saying the titles are missing.
Tests
uv sync --dev
uv run pytest -q # the MCP server
python3 -m unittest discover -s core/tests -t core # the core, offline
uv run python scripts/check_readme.py # README and worked exampleThe core's tests replay responses recorded live by core/tools/record_fixture.py; a replay
refuses any URL it was not given and any body whose sha256 differs from its record.
LITCHECK_NETWORK_TESTS=1 runs the live smoke test against the real services.
Credits
litcheck reads public services and is grateful to the people who run them: Europe PMC (EMBL-EBI); NCBI, for E-utilities, LitSense 2.0, the PMC Cloud Service and the PMC ID Converter; Crossref, including the Retraction Watch data it distributes; and OpenAlex.
The pinned-PMC gate and quote matcher are ported from peerpanel.
Built by Javier Rodríguez Hernáez · LinkedIn · GitHub
Built for GARS, the genomics agentic research system.
Licence
MIT, for the source code. The recorded responses under core/tests/fixtures/ belong to the
services that served them; the one article text among them is CC BY and attributed in
docs/EXAMPLE.md.
Available Tools
7 toolscheck_quoteA
Check whether a quoted passage appears in the paper's pinned open-access PMC text.
Also reports the text's version, sha256 and licence, and the paper's retraction status,
and appends an evidence line to the log. The text searched is PMC's whole plain-text file,
header and reference list included, so check paragraph before attributing a FOUND
passage to the paper. Exact substring match (word boundaries are not
checked) after normalisation q1 (Unicode
NFKC, straight quotes, one dash, collapsed whitespace; case kept). It does not judge
whether the passage supports the claim.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | No | Optional: the claim you are citing this passage for, kept in the log | |
| quote | Yes | The passage exactly as you would quote it (20 characters or more; shorter is TOO_SHORT when there is text to check) | |
| identifier | Yes | One paper: a DOI ("10.1590/1516-3180.2016.1344090516" or a doi.org URL), a PMCID ("PMC10496602") or a PMID written with its cue ("PMID 27355798") |
Output Schema
| Name | Required | Description |
|---|---|---|
| claim | Yes | |
| quote | Yes | |
| reason | Yes | |
| log_seq | Yes | |
| verdict | Yes | FOUND: the quote occurs, as a substring, in the pinned open-access text (which includes its metadata header and reference list; see paragraph). NOT_FOUND: it does not. TOO_SHORT: under 20 characters, not searched. NOT_CHECKABLE: no open-access text in PMC, PMC's copy missing, or the identifier was not found (see resolved.status and reason). UNVERIFIABLE: a source could not be read or verified, or litcheck failed (see reason); nothing was concluded. |
| log_path | Yes | |
| resolved | Yes | |
| paragraph | Yes | 0-based index of the blank-line-separated block of the text |
| identifier | Yes | |
| retraction | Yes | |
| pmc_outcome | Yes | |
| pmc_version | Yes | |
| text_md5_ok | Yes | |
| text_sha256 | Yes | |
| license_code | Yes | |
| normalisation | Yes | |
| quote_offsets | Yes | [start, end) of the match in the normalised text |
| casefold_found | Yes | The quote occurs when case is ignored |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses a side effect ('appends an evidence line to the log'), the exact matching semantics (substring, no word boundaries, after normalisation q1), the text scope searched, and the auxiliary facts returned (version, sha256, licence, retraction status). This is substantially more than a bare 'check' verb would convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then layered with scope, semantics, and caveats in separate sentences. Slightly verbose with parenthetical detail about normalisation, but every sentence carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the description still covers the write side effect, matching semantics, and the scope caveat needed to interpret a FOUND result correctly. Nothing an agent needs to call this safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters, giving a baseline of 3. The description adds only matching/normalisation context rather than new per-parameter meaning, so it does not rise above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Check') and resource ('whether a quoted passage appears in the paper's pinned open-access PMC text'), which an agent can distinguish from siblings like retraction_status or record_support. The scope (PMC full plain-text file including header and references) is unusually precise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear exclusion ('It does not judge whether the passage supports the claim') and warns to check `paragraph` before attributing a FOUND passage, which is implied usage guidance. However, it never states when to prefer this tool over siblings such as record_support or verify_log, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citing_papersB
List works that cite a paper, from OpenAlex, and record the search in the log.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many items to return (1-100) | |
| identifier | Yes | The cited work: a DOI or an OpenAlex work id (W123...) |
Output Schema
| Name | Required | Description |
|---|---|---|
| items | Yes | |
| query | Yes | |
| engine | Yes | |
| reason | Yes | |
| status | Yes | UNVERIFIABLE means the search did not complete: an empty item list then says nothing about what exists |
| log_seq | Yes | |
| log_path | Yes | |
| hit_count | Yes | |
| response_sha256 | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it does disclose a real side effect: the search is recorded in the log. It says nothing about pagination, rate limits, or whether the log write can fail, so the behavioral picture is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the action, resource, source, and side effect with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and both parameters are documented in the schema. The description covers purpose, source, and the logging side effect, leaving only alternative-tool routing unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (identifier, limit) are already documented in the schema. The description only alludes to the identifier ('a paper') and never mentions the limit, so it adds little beyond structured data; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('works that cite a paper'), plus the data source (OpenAlex), which an agent can act on immediately. It does not name or differentiate itself from the sibling search_literature, which is the nearest overlapping capability, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when a different sibling (e.g. search_literature, resolve_identifier) is the better choice, and no prerequisites stated. The only usage signal is implied by the description's subject matter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_supportB
Record a person's or a named judge's support verdict about an evidence line.
litcheck computes no support verdict itself; this stores one that was given, with who gave it, as a new line in the log.
| Name | Required | Description | Default |
|---|---|---|---|
| by | Yes | Who judged: a person's name or a named judge; litcheck never does | |
| seq | Yes | The log line (an evidence line) being judged | |
| support | Yes | The verdict: does the quoted passage support the claim? |
Output Schema
| Name | Required | Description |
|---|---|---|
| by | Yes | |
| seq | Yes | The annotation's own line number in the log |
| support | Yes | |
| log_path | Yes | |
| target_seq | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does usefully disclose that the record is appended 'as a new line in the log', implying append-only, non-overwriting behavior. However it says nothing about permissions, idempotency, duplicate verdicts, or what happens if the referenced seq is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by one clarifying sentence about what the tool does not do. No filler, though the line-broken prose is slightly looser than necessary for a three-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the append-to-log behavior is stated. But for a mutating tool with zero annotations, the description leaves permission requirements, error handling, and duplicate/conflicting verdict behavior unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents seq, support, and by (including the enum values). The description only echoes the 'who gave it' idea, adding no format, constraint, or interpretation beyond the structured fields, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource: it records a support verdict about an evidence line, and it clarifies the semantics of 'support' as a verdict given by someone rather than computed. It implicitly separates itself from verification-style siblings (check_quote, verify_log) by stressing that litcheck computes no verdict, though it never names a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: 'litcheck computes no support verdict itself; this stores one that was given' tells the agent this is the write path for an externally-supplied verdict, but there is no explicit when-to-use, when-not-to-use, or named alternative among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_identifierA
Map one DOI, PMID or PMCID to the other two, through the PMC ID converter.
The converter knows only articles in PubMed Central; a DOI it cannot map is also looked up in Europe PMC. NOT_FOUND means not in PMC (and, for a DOI, not in Europe PMC either): a PMID or PMCID is not looked up elsewhere, so NOT_FOUND does not mean the paper does not exist. UNVERIFIABLE means a service could not be read.
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | One paper: a DOI ("10.1590/1516-3180.2016.1344090516" or a doi.org URL), a PMCID ("PMC10496602") or a PMID written with its cue ("PMID 27355798") |
Output Schema
| Name | Required | Description |
|---|---|---|
| doi | Yes | |
| pmid | Yes | |
| pmcid | Yes | |
| reason | Yes | Why it is not RESOLVED, or why there is no PMCID |
| source | Yes | Which service answered: pmc-idconv or europe-pmc |
| status | Yes | NOT_FOUND: not in PubMed Central (and, for a DOI, not in Europe PMC); a PMID or PMCID is not looked up elsewhere, so NOT_FOUND does not mean it does not exist |
| identifier | Yes | The identifier as it was given |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the converter's PMC-only coverage, the Europe PMC fallback that applies to DOIs but not PMIDs/PMCIDs, and the exact meaning of NOT_FOUND and UNVERIFIABLE, including the important caveat that NOT_FOUND does not mean the paper is nonexistent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence, and the remaining sentences are dense but entirely devoted to failure-mode semantics that an agent cannot get elsewhere. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with a full output schema and no annotations, the description covers everything an agent needs: what it converts, coverage limits, fallback behavior, and error-code interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents the accepted formats (DOI, doi.org URL, PMCID, cued PMID). The description adds no syntax or format detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource: mapping one of DOI/PMID/PMCID to the other two via the PMC ID converter. The scope is narrow and unambiguous, and clearly distinct from siblings like search_literature, citing_papers, and retraction_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: use it when you hold one identifier and need the others. It never names an alternative tool or an explicit when-not-to-use condition, so the agent must infer the routing from the sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retraction_statusA
Is the paper retracted? Reads PMC's is_retracted flag and Crossref's notices.
Retraction, withdrawal and removal notices count; corrections do not; an expression of
concern is reported separately. NOT_RETRACTED_AS_OF means no retraction notice was found
in the sources consulted, as of its date; it needs the identifier to have resolved and
every consulted source to have answered, otherwise the status is UNVERIFIABLE. A paper PMC
does not hold is checked on Crossref alone, and one with no known DOI on PMC's flag alone;
sources shows which answered. For a paper outside PMC, pass its DOI. Not logged.
| Name | Required | Description | Default |
|---|---|---|---|
| identifier | Yes | One paper: a DOI ("10.1590/1516-3180.2016.1344090516" or a doi.org URL), a PMCID ("PMC10496602") or a PMID written with its cue ("PMID 27355798") |
Output Schema
| Name | Required | Description |
|---|---|---|
| as_of | Yes | For NOT_RETRACTED_AS_OF: the date of the earliest check; no retraction notice was found in the sources consulted as of then |
| reason | Yes | |
| status | Yes | |
| concern | Yes | An expression of concern is on record |
| sources | Yes | |
| identifier | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it defines NOT_RETRACTED_AS_OF, explains that it is only returned when the identifier resolved and all sources answered, describes the UNVERIFIABLE fallback, documents source-fallback behavior for papers not in PMC, exposes the `sources` field, and notes 'Not logged.' These are genuine behavioral traits beyond any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core question and behavior, then organized into status definitions and edge cases; every sentence carries information. It is dense and slightly long for a one-parameter tool, but the length is mostly justified by the status semantics that cannot live in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be explained, yet the description still clarifies the `sources` field and the meaning of each status, fully covering the edge cases (non-PMC papers, missing DOI, unresolved identifiers). Nothing an agent needs to invoke and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics: it clarifies the identifier addresses a single paper and gives routing advice ('For a paper outside PMC, pass its DOI'), plus the requirement that the identifier resolve for a definitive status. That is meaningful meaning beyond the format examples in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific question-verb framing ('Is the paper retracted?') and names the exact mechanism ('Reads PMC's is_retracted flag and Crossref's notices'). An agent can distinguish this from siblings like check_quote or resolve_identifier, which concern quotations and identifier resolution respectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear routing context: 'For a paper outside PMC, pass its DOI,' and explains that corrections are excluded while expressions of concern are reported separately. It stops short of explicitly naming a sibling alternative or a when-not-to-use condition, so it is strong context without full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_literatureA
Run one literature search and record it, with its response hash, in the log.
hit_count is what the engine reported (for litsense, the number of results it returned,
at most 100); items are the first limit of them. A search can
miss papers: status SEARCHED with few items is a record of this query, not of the field.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many items to return (1-100) | |
| query | Yes | The query, in the chosen engine's syntax | |
| engine | Yes | europe_pmc: Europe PMC query syntax (TITLE_ABS:, FIRST_PDATE:[a TO b], ...); litsense: a sentence-level semantic search of PubMed and PMC (no date filter); pubmed: PubMed query syntax through E-utilities |
Output Schema
| Name | Required | Description |
|---|---|---|
| items | Yes | |
| query | Yes | |
| engine | Yes | |
| reason | Yes | |
| status | Yes | UNVERIFIABLE means the search did not complete: an empty item list then says nothing about what exists |
| log_seq | Yes | |
| log_path | Yes | |
| hit_count | Yes | |
| response_sha256 | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the call has a persistent side effect (logging the query with a response hash), explains that hit_count reflects what the engine reported while items are only the first `limit`, and warns that a SEARCHED status with few items reflects the query rather than the field. It stops short of covering failure modes, auth, or engine-specific limits beyond litsense's cap of 100.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and its logging side effect are front-loaded in the first sentence, and every subsequent sentence adds interpretive value. The parenthetical about litsense and the 100 cap is slightly dense, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required, yet the description still clarifies hit_count vs items semantics, which is the main interpretive risk. For a three-parameter tool with a full schema this is close to complete, with only error/edge behavior left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema by explaining the relationship between `limit` and the returned items and the distinct meaning of hit_count. The engine-specific syntax nuance is left to the enum descriptions rather than repeated here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Run one literature search') and adds the side effect ('record it, with its response hash, in the log'). No sibling tool performs a search, so it is unambiguously distinguishable from check_quote, resolve_identifier, citing_papers, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the fact that this is the only search tool among the siblings, and the note about missing papers hints at how to interpret results. However, there is no explicit 'when to use this vs alternatives' guidance, no prerequisites, and no statement of when not to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_logA
Re-walk the log's hash chain and report the first line that does not fit, if any.
This catches accidental and naive edits. The hashes are not keyed, so a rewrite that recomputes every later hash, or a cut-off tail, is caught only by comparing head_sha256 with a copy kept elsewhere. A log that does not exist yet reports chain_ok false.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| path | Yes | |
| lines | Yes | |
| reason | Yes | |
| chain_ok | Yes | |
| head_sha256 | Yes | sha256 of the last line's bytes (without its newline; not its own sha256 field). The chain catches accidental and naive edits; a rewrite that recomputes every hash, or a cut-off tail, is caught only by comparing against a head hash kept elsewhere |
| first_bad_seq | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the detection model, its key limitation (unkeyed hashes; recomputed chains and truncated tails evade it), and an edge case (nonexistent log returns chain_ok false). It omits side-effect/read-only and cost/performance characteristics, which keeps it just short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, followed by the value proposition and the adversarial-limitation caveat. Every sentence earns its place, though the limitation sentence is dense and could be split or tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required, yet the description still usefully explains the chain_ok=false edge case. Combined with a zero-parameter schema, the definition is essentially complete; only security/side-effect framing is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is no parameter syntax for the description to clarify, and it correctly adds no redundant parameter discussion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Re-walk the log's hash chain and report the first line that does not fit.' It also states the exact output of interest (first failing line, if any), so an agent knows both the action and the result shape without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implicitly defines the use case ('catches accidental and naive edits') and, importantly, states the condition under which the tool is NOT sufficient (rewrites that recompute later hashes or cut tails are only caught by comparing head_sha256 with an external copy). No sibling among check_quote/search_literature/etc. competes for this job, so explicit alternative routing isn't needed, but it stops short of saying when in a workflow to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
check_quote - First observed
citing_papers - First observed
record_support - First observed
resolve_identifier - First observed
retraction_status - First observed
search_literature - First observed
verify_log
TDQS
Scored across 7 tools
Each tool targets a clearly distinct task: quote checking, identifier mapping, log integrity, retraction status, literature search, citation listing, and support recording. Although retraction_status and check_quote both read PMC, their purposes (retraction notices vs. textual quote matching) are unambiguous from the descriptions.
Most tools follow a verb_noun pattern (check_quote, resolve_identifier, verify_log, search_literature, record_support). Two deviate into noun phrases (retraction_status, citing_papers), which is a minor inconsistency but still readable and predictable.
Seven tools is well-scoped for a literature-verification server, and each tool earns its place with a distinct role. No redundancy or filler tools.
The surface covers the core verification lifecycle: quote checking, ID resolution, retraction checking, searching, citations, and support recording plus log auditing. Minor gaps exist (e.g., no tool to fetch full paper metadata or export/report the log), but agents can work around them.
Maintenance
Related MCP Connectors
Checks AI-written references against Crossref, PubMed and OpenAlex. Formats citations, PRISMA.
Real-time fact-check, citation verification, and source-freshness for AI agents.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Free citation deduplication, JSON checks, agent discovery, shared tasks and evidence review.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables scientific due diligence by grading claims against public literature, clinical trials, and filings, with explicit citations and optional attestation.4MIT
- AlicenseNot gradedqualityAmaintenanceEnables traceable scholarly literature reviews using free APIs, generating reports where every claim links to evidence IDs.MIT

CiteStamp MCP serverofficial
AlicenseNot gradedqualityBmaintenanceGround citations before your agent emits them by checking references against public scholarly registries and flagging hallucinated or retracted ones.MIT- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to conduct academic research workflows such as paper discovery, literature mapping, citation chasing, author pivots, citation repair, and regulatory or species document retrieval.MIT