Skip to main content
Glama
VelvetSP

io.github.VelvetSP/web-retrieval-mcp

by VelvetSP

research_papers

Read-only

Search over 3 million arXiv AI/ML papers to find relevant literature, returning ranked results with titles, IDs, relevance scores, and abstracts. Then verify claims against full text before citing.

Instructions

Search 3M+ arXiv AI/ML papers via the Firecrawl Research Index (state-of-the-art paper recall — far better than general web search for finding the right literature). Returns ranked papers: title, arXiv id, relevance score, abstract. Then call research_paper(paper_id, query=…) to verify a claim against full text before citing.

SCOPE: arXiv-scoped, i.e. effectively AI/ML. For scholarly literature outside that scope (medicine, law, economics, humanities), use web_search(category="publication") — Exa's publications index (~350M works) covers what this one cannot.

Args: query: natural-language research query (topic, method, benchmark, author). k: number of papers (1–25, default 8).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
kNo
queryYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this tool read-only, so the description does not need to restate that. It adds useful behavioral context: the source index, arXiv scope, result ranking, and returned fields (title, arXiv id, relevance score, abstract), plus the downstream verification step. It does not discuss rate limits or error behavior, but the read-only annotation lowers the burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Information is front-loaded: the core search purpose and return shape appear first, the scope/routing caveat appears second, and parameter definitions are cleanly separated. Every sentence serves selection, workflow, or invocation without unnecessary filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only search tool with an output schema, the description covers purpose, scope, alternatives, return shape, and parameter semantics. The output schema handles detailed return structure, so nothing essential is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the input schema has no property descriptions, the Args block fully documents both parameters: query is a natural-language research query, and k is a count with an explicit 1–25 range and default 8. This fully compensates for the 0% schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Search 3M+ arXiv AI/ML papers via the Firecrawl Research Index', and explicitly states the tool is arXiv-scoped. It also distinguishes itself from siblings by describing the ranked paper list output and pointing to research_paper for verification, so an agent can select it without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit selection criteria: use this for arXiv AI/ML literature and, for medicine/law/economics/humanities, use web_search(category="publication") because that index is larger. It also tells the agent to call research_paper to verify claims before citing, making the intended workflow clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.