Skip to main content
Glama
Remus-cloud

literature-bot-mcp

by Remus-cloud

Retrieve Evidence from a Paper

literature_retrieve_paper_evidence
Idempotent

Retrieve page-cited evidence chunks from a paper to answer content questions using concise keyword queries; returns page numbers for citations.

Instructions

Retrieve page-cited evidence chunks from one paper.

Use this before answering a question about paper content. The query must be concise English keywords because retrieval uses local BM25 rather than a model. If necessary, the paper is prepared automatically. Return the page numbers with any answer built from these results.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryYesConcise English keywords. Translate non-English questions before calling.
top_kNoReturn between 1 and 8 chunks.
arxiv_idYesarXiv ID identifying the paper to search.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryYes
titleYes
arxiv_idYes
evidenceYes
auto_preparedYes
evidence_countYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real mechanics beyond the annotations: retrieval is local BM25 rather than a model, the paper may be prepared as a side effect, and callers should carry page numbers into the answer. The auto-preparation side effect is consistent with readOnlyHint=false (a read tool that can trigger a write), so there is no contradiction, though failure modes (paper not found, no hits) go unmentioned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the action and the pre-condition, then the query-format constraint, then the side effect and the citation obligation. Nothing is redundant or padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations present and an output schema covering the return shape, the description supplies the missing operational context: when to call it, how to phrase the query, and the auto-preparation behavior. It omits edge cases such as an unknown arxiv_id or an empty result set, which keeps it just short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns credit by explaining *why* the query must be short English keywords (BM25, not a model), which is the rationale the schema omits. The top_k parameter is not addressed in prose beyond the schema's own range documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: retrieving page-cited evidence chunks, scoped to 'one paper', which implicitly separates it from the corpus-wide literature_search_papers sibling. It never names that sibling, so the differentiation is inferable rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this before answering a question about paper content" gives a clear trigger condition, and "the paper is prepared automatically" tells the agent it need not call literature_prepare_paper first. No explicit when-not or named alternative is given, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.