Skip to main content
Glama

Arxiv Read Paper

arxiv_read_paper
Read-only

Fetch the full text of an arXiv paper. Tries arxiv.org/html first, falls back to ar5iv.labs.arxiv.org, and falls back again to text extracted from the PDF when neither has an HTML render — check the source field to know which one answered. Page through long papers with start and max_characters, or pass max_characters null to get the entire body in one call.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
startNoCharacter offset into the cleaned body to begin reading from. Defaults to 0. Use with max_characters to page through long papers — e.g., start=100000 with max_characters=100000 returns chars 100,000–199,999. The total length is reported as body_characters in the response.
paper_idYesarXiv paper ID (e.g., "2401.12345" or "2401.12345v2").
max_charactersNoMaximum characters of paper body to return, counted after boilerplate stripping. Defaults to 100,000; pass null to return the entire body in one call. Whole-paper reads can exceed a client tool-result size cap — math-heavy bodies run 300KB-1MB+ — so prefer the default plus start-based paging unless the full text is needed. When truncated, a notice and the total character count are included.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoPresent when the call failed. Absent on success.
startNoCharacter offset of the first character in content within the cleaned body.
titleNoPaper title (from metadata, not parsed from HTML).
sourceNoWhich upstream artifact the body was read from. arxiv_html and ar5iv are HTML renders; pdf_text is text extracted from the PDF, where prose is reliable but math, tables, and heading structure are flattened.
contentNoPaper body for the requested slice — cleaned HTML when source is arxiv_html or ar5iv, plain text when source is pdf_text. Empty when start is past body_characters.
pdf_urlNoDirect PDF download URL.
paper_idNoarXiv paper ID.
truncatedNoTrue when more body content exists past this slice (start + content.length < body_characters).
abstract_urlNoarXiv abstract page URL for attribution.
body_charactersNoCharacter count of the full cleaned body. Use with start and max_characters to page. Typically 3-4× smaller than total_characters for math-heavy HTML papers.
total_charactersNoCharacter count of the body before cleaning — the unprocessed HTML body for arxiv_html and ar5iv, and equal to body_characters for pdf_text, which needs no cleaning.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedInput schema / additionalProperties
      Added value: +false
    • changedOutput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedOutput schema / anyOf
      Added value: +[
      +  {
      +    "not": {
      +      "required": [
      +        "error"
      +      ]
      +    },
      +    "required": [
      +      "paper_id",
      +      "title",
      +      "content",
      +      "source",
      +      "truncated",
      +      "start",
      +      "total_characters",
      +      "body_characters",
      +      "pdf_url",
      +      "abstract_url"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "error"
      +    ]
      +  }
      +]
    • addedOutput schema / properties / error
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "Present when the call failed. Absent on success.",
      +  "properties": {
      +    "code": {
      +      "description": "JSON-RPC error code for this failure.",
      +      "maximum": 9007199254740991,
      +      "minimum": -9007199254740991,
      +      "type": "integer"
      +    },
      +    "data": {
      +      "additionalProperties": {},
      +      "properties": {
      +        "reason": {
      +          "description": "Machine-readable failure mode. Declared by this tool: `no_match`: Paper ID is not present in the arXiv index. `content_unavailable`: Paper exists but neither arxiv.org/html nor ar5iv has an HTML rendering and arXiv served no PDF either. `pdf_extraction_failed`: Paper has no HTML rendering and its PDF carries no text layer — an image-only or scanned submission. `version_unavailable`: A version-pinned paper_id was requested, arXiv is unreachable, and the local mirror holds only a different version — per-version reads require the live API. `rate_limited`: arXiv has throttled requests (HTTP 429 or \"Rate exceeded.\" body). `invalid_request`: arXiv rejected the metadata lookup (HTTP 4xx other than 429), e.g. malformed ID syntax. Other values are possible when a failure originates below the handler.",
      +          "examples": [
      +            "no_match",
      +            "content_unavailable",
      +            "pdf_extraction_failed",
      +            "version_unavailable",
      +            "rate_limited",
      +            "invalid_request"
      +          ],
      +          "type": "string"
      +        },
      +        "recovery": {
      +          "additionalProperties": {},
      +          "description": "Actionable next step for the caller.",
      +          "properties": {
      +            "hint": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "hint"
      +          ],
      +          "type": "object"
      +        },
      +        "retryable": {
      +          "description": "Whether retrying may succeed.",
      +          "type": "boolean"
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "message": {
      +      "description": "Human-readable description of what went wrong.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "code",
      +    "message"
      +  ],
      +  "type": "object"
      +}
    • removedOutput schema / required
      Removed value: -[
      -  "paper_id",
      -  "title",
      -  "content",
      -  "source",
      -  "truncated",
      -  "start",
      -  "total_characters",
      -  "body_characters",
      -  "pdf_url",
      -  "abstract_url"
      -]
  2. Changed10 schema fields changed
    • addedInput schema / properties / max_characters / anyOf
      Added value: +[
      +  {
      +    "maximum": 9007199254740991,
      +    "minimum": 1,
      +    "type": "integer"
      +  },
      +  {
      +    "type": "null"
      +  }
      +]
    • changedInput schema / properties / max_characters / description
      Previous value: -"Maximum characters of paper body content to return. Defaults to 100,000. HTML head/boilerplate is stripped before counting. When truncated, a notice and total character count are included."New value: +"Maximum characters of paper body to return, counted after boilerplate stripping. Defaults to 100,000; pass null to return the entire body in one call. Whole-paper reads can exceed a client tool-result size cap — math-heavy bodies run 300KB-1MB+ — so prefer the default plus start-based paging unless the full text is needed. When truncated, a notice and the total character count are included."
    • removedInput schema / properties / max_characters / maximum
      Removed value: -9007199254740991
    • removedInput schema / properties / max_characters / minimum
      Removed value: -1
    • removedInput schema / properties / max_characters / type
      Removed value: -"integer"
    • changedOutput schema / properties / body_characters / description
      Previous value: -"Character count of the full cleaned body HTML. Use with start and max_characters to page. Typically 3-4× smaller than total_characters for math-heavy papers."New value: +"Character count of the full cleaned body. Use with start and max_characters to page. Typically 3-4× smaller than total_characters for math-heavy HTML papers."
    • changedOutput schema / properties / content / description
      Previous value: -"Cleaned paper body HTML for the requested slice. Empty when start is past body_characters."New value: +"Paper body for the requested slice — cleaned HTML when source is arxiv_html or ar5iv, plain text when source is pdf_text. Empty when start is past body_characters."
    • changedOutput schema / properties / source / description
      Previous value: -"Which HTML source the content was fetched from."New value: +"Which upstream artifact the body was read from. arxiv_html and ar5iv are HTML renders; pdf_text is text extracted from the PDF, where prose is reliable but math, tables, and heading structure are flattened."
    • changedOutput schema / properties / source / enum
      Previous value: -[
      -  "arxiv_html",
      -  "ar5iv"
      -]New value: +[
      +  "arxiv_html",
      +  "ar5iv",
      +  "pdf_text"
      +]
    • changedOutput schema / properties / total_characters / description
      Previous value: -"Character count of the original unprocessed HTML body."New value: +"Character count of the body before cleaning — the unprocessed HTML body for arxiv_html and ar5iv, and equal to body_characters for pdf_text, which needs no cleaning."
  3. Changed6 schema fields changed
    • addedInput schema / properties / start
      Added value: +{
      +  "default": 0,
      +  "description": "Character offset into the cleaned body to begin reading from. Defaults to 0. Use with max_characters to page through long papers — e.g., start=100000 with max_characters=100000 returns chars 100,000–199,999. The total length is reported as body_characters in the response.",
      +  "maximum": 9007199254740991,
      +  "minimum": 0,
      +  "type": "integer"
      +}
    • changedOutput schema / properties / body_characters / description
      Previous value: -"Character count of the cleaned body HTML — what fits into max_characters. Typically 3-4× smaller than total_characters for math-heavy papers."New value: +"Character count of the full cleaned body HTML. Use with start and max_characters to page. Typically 3-4× smaller than total_characters for math-heavy papers."
    • changedOutput schema / properties / content / description
      Previous value: -"Cleaned paper body HTML, truncated to max_characters."New value: +"Cleaned paper body HTML for the requested slice. Empty when start is past body_characters."
    • addedOutput schema / properties / start
      Added value: +{
      +  "description": "Character offset of the first character in content within the cleaned body.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / truncated / description
      Previous value: -"Whether content was truncated due to max_characters."New value: +"True when more body content exists past this slice (start + content.length < body_characters)."
    • changedOutput schema / required
      Previous value: -[
      -  "paper_id",
      -  "title",
      -  "content",
      -  "source",
      -  "truncated",
      -  "total_characters",
      -  "body_characters",
      -  "pdf_url",
      -  "abstract_url"
      -]New value: +[
      +  "paper_id",
      +  "title",
      +  "content",
      +  "source",
      +  "truncated",
      +  "start",
      +  "total_characters",
      +  "body_characters",
      +  "pdf_url",
      +  "abstract_url"
      +]
  4. Changed4 schema fields changed
    • addedOutput schema / properties / body_characters
      Added value: +{
      +  "description": "Character count of the cleaned body HTML — what fits into max_characters. Typically 3-4× smaller than total_characters for math-heavy papers.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / content / description
      Previous value: -"Raw HTML content of the paper."New value: +"Cleaned paper body HTML, truncated to max_characters."
    • changedOutput schema / properties / total_characters / description
      Previous value: -"Total character count of the full (untruncated) content."New value: +"Character count of the original unprocessed HTML body."
    • changedOutput schema / required
      Previous value: -[
      -  "paper_id",
      -  "title",
      -  "content",
      -  "source",
      -  "truncated",
      -  "total_characters",
      -  "pdf_url",
      -  "abstract_url"
      -]New value: +[
      +  "paper_id",
      +  "title",
      +  "content",
      +  "source",
      +  "truncated",
      +  "total_characters",
      +  "body_characters",
      +  "pdf_url",
      +  "abstract_url"
      +]
  5. Changed3 schema fields changed
    • addedInput schema / properties / max_characters / maximum
      Added value: +9007199254740991
    • addedInput schema / properties / max_characters / minimum
      Added value: +1
    • changedInput schema / properties / max_characters / type
      Previous value: -"number"New value: +"integer"
  6. First observed

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already marks the operation as read-only; the description adds material context: the multi-source fallback order, the source field to identify which source answered, and the paging/truncation behavior. These details go well beyond what the annotation provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, then fallback behavior, then paging. Every sentence earns its place, and there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and a readOnlyHint annotation, the description covers all essential behaviors: fallback sources, source field, and paging vs. whole-body options. Remaining details like parameter bounds are in the schema, so nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description restates the paging semantics and the null option, but adds little beyond what the schema's start and max_characters descriptions already say.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Fetch the full text of an arXiv paper,' naming a specific verb and resource. It clearly distinguishes itself from siblings by focusing on full-text retrieval with a concrete fallback chain, so an agent can tell it apart from search or metadata tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly frames its role—full-text fetching—and gives concrete guidance on paging with start and max_characters, and on requesting the entire body via max_characters=null. It does not explicitly contrast with sibling tools, but the purpose statement plus sibling names make the context clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.