Skip to main content
Glama
Remus-cloud

literature-bot-mcp

by Remus-cloud

Prepare Paper for Retrieval

literature_prepare_paper
Idempotent

Download, extract, and chunk one arXiv paper by ID for evidence retrieval, caching the PDF locally and reusing it on repeat calls.

Instructions

Download, extract, and chunk one paper for evidence retrieval.

The PDF is cached in the configured local papers directory and is not overwritten. Extracted pages and chunks stay only in server memory. Repeating the call is safe and reuses the cache.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
arxiv_idYesarXiv ID returned by literature_search_papers.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
titleYes
arxiv_idYes
pdf_pathYes
page_countYes
chunk_countYes
nonempty_page_countYes
reused_memory_cacheYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that the PDF is cached in a configured local directory and never overwritten, that extracted pages/chunks live only in server memory, and that repeat calls are safe and reuse the cache. This substantiates the idempotentHint=true and readOnlyHint=false (disk cache write) annotations with concrete mechanics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the action front-loaded and secondary behavioral facts (cache location, memory-only chunks, idempotency) following. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, output-schema-backed preparation tool, the description covers the action, the caching side effect, the memory footprint, and repeat-call safety. An agent has everything needed to invoke it correctly and predict its side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single arxiv_id parameter is well documented in the schema ('arXiv ID returned by literature_search_papers'). The description adds no syntax, format, or validation detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific three-step verb chain (download, extract, chunk) on a specific resource (one paper) with a stated purpose (evidence retrieval). An agent can distinguish this preparation step from literature_search_papers and literature_retrieve_paper_evidence without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for evidence retrieval' implies this is a prerequisite for literature_retrieve_paper_evidence, giving implied usage. However, no sibling is named explicitly, and there is no statement of when not to call it (e.g. already-prepared papers) or which tool must precede it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.