Skip to main content
Glama

verify_reference_list

Read-onlyIdempotent

Verify a plain-text reference list against PubMed evidence to confirm citation accuracy and flag unresolved entries.

Instructions

Verify a plain-text reference list against PubMed evidence.

First version scope: - Reference-list verification only - Client supplies the extracted reference list text - Backend parses entries and resolves them via PMID / DOI / ECitMatch

Second version scope: - Adds unresolved review workflow for partial_match and unresolved rows - Returns a manual-review queue with retry queries and review checklist - Supports human-in-the-loop acceptance/rejection in client-side workflows

Args: reference_text: Plain-text references, ideally one per line or a numbered reference list extracted from a file. Limited to 200,000 characters / 400,000 UTF-8 bytes; each entry is limited to 4,000 characters / 8,000 UTF-8 bytes. source_name: Optional single-line file label for reporting (up to 255 characters / 512 UTF-8 bytes). max_references: Hard input-entry limit from 1 through 200. Inputs above the selected limit are rejected instead of truncated.

Returns: JSON verification report with parsed fields, matched PubMed evidence, per-reference verification status, and explicit source_unavailable / not_checked rows when evidence could not be assessed.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
source_nameNo
max_referencesNo
reference_textYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changedv0.7.7
    • addedInput schema / properties / max_references / anyOf
      Added value: +[
      +  {
      +    "default": 100,
      +    "maximum": 200,
      +    "minimum": 1,
      +    "title": "Max References",
      +    "type": "integer"
      +  },
      +  {
      +    "description": "ASCII decimal integer; the integer branch's bounds apply after conversion.",
      +    "maxLength": 32,
      +    "pattern": "^[ \\t\\r\\n]*-?(?:0|[1-9][0-9]*)[ \\t\\r\\n]*$",
      +    "type": "string"
      +  }
      +]
    • removedInput schema / properties / max_references / maximum
      Removed value: -200
    • removedInput schema / properties / max_references / minimum
      Removed value: -1
    • removedInput schema / properties / max_references / type
      Removed value: -"integer"
  2. Changed7 schema fields changedv0.7.2
    • addedInput schema / additionalProperties
      Added value: +false
    • addedInput schema / properties / max_references / maximum
      Added value: +200
    • addedInput schema / properties / max_references / minimum
      Added value: +1
    • addedInput schema / properties / reference_text / maxLength
      Added value: +200000
    • addedInput schema / properties / reference_text / minLength
      Added value: +1
    • addedInput schema / properties / source_name / maxLength
      Added value: +255
    • changedOutput schema / (root)
      Previous value: -{
      -  "properties": {
      -    "result": {
      -      "title": "Result",
      -      "type": "string"
      -    }
      -  },
      -  "required": [
      -    "result"
      -  ],
      -  "title": "verify_reference_listOutput",
      -  "type": "object"
      -}New value: +null
  3. First observedv0.5.16

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a read-only, idempotent, open-world, non-destructive operation, so the safety bar is low. The description still adds real behavioral content: hard size/entry limits, that over-limit input is rejected rather than truncated, and that unresolved evidence surfaces as explicit source_unavailable / not_checked rows. It stops short of describing performance or retry behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The Args/Returns structure is clear and front-loaded, but the "Second version scope" block describes future behavior that is not invocable today, consuming roughly a third of the text without helping an agent call the tool now. That section is the main argument for a mid score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on return-value explanation (parsed fields, matched evidence, per-reference status, explicit unavailable rows), and it covers all limits and required inputs. Complete enough to invoke correctly; only the ambiguity about whether v2 behavior is currently active leaves a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load and it does: reference_text limits (200,000 chars / 400,000 bytes, 4,000 per entry), source_name as an optional single-line reporting label with a stated length cap, and max_references bounds (1-200) plus the rejection-instead-of-truncation rule. All three parameters gain meaning the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: "Verify a plain-text reference list against PubMed evidence." No sibling performs reference-list verification, so an agent can route to it without ambiguity. The resolution mechanisms (PMID / DOI / ECitMatch) further pin down what the tool actually does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It establishes that the client must supply already-extracted reference text, which is useful pre-condition context, but it never states when to pick this over sibling tools that also surface references (e.g. get_article_references, build_citation_tree) or what inputs are unsuitable. The v1/v2 scope split is informative about roadmap rather than about invocation choices.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.