Skip to main content
Glama

get_document

Retrieve full paper details by ID. Default returns metadata only (title, authors, abstract, license, codeLinks counts) — use includeChunks=true to fetch chunk content. For specific sections or content types, use chunkContentTypes/section filters or call get_chunks instead. For long papers, prefer filtered chunk retrieval over full chunks dump. AVAILABILITY is two INDEPENDENT axes: indexingTier (none|abstract_only|full|reindexing) = whether the full text is indexed and readable via get_chunks — 'reindexing' means the document is being re-processed right now and its currently indexed chunks are STALE: do not quote them as the body and do not treat the document as abstract_only either, its state is not yet known (chunkCount shows how many); sourceAccessibility (served_by_us|external_link_only|unavailable) = how to obtain the raw source file, with sourceUrl returned whenever known. To read content: if indexingTier='full' use get_chunks; else if sourceAccessibility!='unavailable' fetch sourceUrl yourself; only 'unavailable' means no full text. canServeFile is DEPRECATED — it gates raw-PDF delivery ONLY and is NOT a content-availability signal; use indexingTier + sourceAccessibility. An identifier resolves to one specific version of a document, not to a mutable current state, so a reference cannot silently come to mean different text.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idNoDocument UUID
detailNo
offsetNoSkip the first N chunks of the ordered set. For reading a document larger than one context in successive passes. Ordering is total, so pages neither skip nor repeat.
run_idNoOptional. The active methodist run_id (as returned by the methodist diagnose / get_current_dose door). Pass it whenever you call this tool while working inside a run, so the call is attributed to that run for the §8 usage crosscheck — attribution is run-anchored, so it stays correct even if your access token refreshes mid-run. Must be YOUR run: a run_id owned by a different principal, or a non-existent run_id, is rejected.
arxivIdNoarXiv ID (e.g. 1706.03762)
searchIdNosearchId from a prior search / search_keyword / search_semantic response. Required when chunkOrder=importance.
chunkLimitNoMax chunks returned when includeChunks=true (or filter is set). Default 100. No upper bound — a document is readable in full; use offset to page through one too large for your context.
chunkOrderNo'position' (default): document order. 'importance': search relevance order — requires searchId from a prior search response; falls back to position with a note when searchId is missing or expired.position
chunkEntitiesNoSoft filter chunks by entity (case-insensitive ANY match). Implies includeChunks=true. Chunks mentioning a listed entity return first; chunks with NO entities recorded are kept (about a quarter of the corpus predates entity extraction, and dropping them would hide real matches); only chunks that HAVE entities, none of them matching, are dropped.
includeChunksNoDEFAULT FALSE — metadata only. Set true for chunk content. Combine with chunkContentTypes/chunkLimit for filtered retrieval. (search v2 changed default; pre-2026-05 v1 always returned chunks.)
chunkContentTypesNoSoft filter chunks by type. Implies includeChunks=true. Matched chunks return first; legacy chunks with NULL contentType are also included as 'unknown' tier (sorted after matches) — they are not silently dropped. Chunks with an explicit different contentType are excluded. Response includes strictMatchedChunks + unknownChunksIncluded counts.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / chunkEntities
      Added value: +{
      +  "description": "Soft filter chunks by entity (case-insensitive ANY match). Implies includeChunks=true. Chunks mentioning a listed entity return first; chunks with NO entities recorded are kept (about a quarter of the corpus predates entity extraction, and dropping them would hide real matches); only chunks that HAVE entities, none of them matching, are dropped.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
  2. Changed5 schema fields changed
    • removedInput schema / properties / chunkLimit / default
      Removed value: -20
    • changedInput schema / properties / chunkLimit / description
      Previous value: -"Max chunks returned when includeChunks=true (or filter is set)"New value: +"Max chunks returned when includeChunks=true (or filter is set). Default 100. No upper bound — a document is readable in full; use offset to page through one too large for your context."
    • removedInput schema / properties / chunkLimit / maximum
      Removed value: -200
    • removedInput schema / properties / detail / default
      Removed value: -"full"
    • addedInput schema / properties / offset
      Added value: +{
      +  "description": "Skip the first N chunks of the ordered set. For reading a document larger than one context in successive passes. Ordering is total, so pages neither skip nor repeat.",
      +  "minimum": 0,
      +  "type": "integer"
      +}
  3. Changed1 schema field changed
    • changedInput schema / properties / detail / default
      Previous value: -"standard"New value: +"full"
  4. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so unusually well: it discloses the stale-chunk hazard during 'reindexing', clarifies that canServeFile is DEPRECATED and not an availability signal, and defines two independent availability axes. It also states the immutable-version semantics of an identifier, which is non-obvious behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the default-vs-includeChunks distinction are front-loaded, and the dense detail is largely earned by the tool's complexity. It is still a very long single block with heavy ALL-CAPS emphasis, and parts of the availability-axis explanation are closer to response documentation than invocation guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the default metadata fields and the availability fields an agent must interpret. The remaining gap is the unexplained 'detail' enum and no guidance on pagination limits beyond the offset/chunkLimit hints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 91%, so the baseline is 3, but the description adds real value on top: it explains why includeChunks defaults false, when to prefer filtered retrieval, and how the availability axes gate content access. The notable miss is the 'detail' enum (minimal|standard|full), which has no description in either the schema or the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Retrieve full paper details by ID') and immediately scopes the default behavior ('metadata only') and the escape hatch ('includeChunks=true'). It explicitly names the sibling get_chunks as the alternative for section-level content, so the agent can differentiate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when/when-not routing: metadata by default, includeChunks for content, chunkContentTypes/section filters for targeted retrieval, get_chunks for section-level reads, and 'prefer filtered chunk retrieval over full chunks dump' for long papers. It also states the decision rule for reading content ('if indexingTier=full use get_chunks; else fetch sourceUrl').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.