Skip to main content
Glama
rookslog

arXiv Discovery MCP

by rookslog

arXiv Discovery MCP

MCP server for arXiv paper discovery, triage, and monitoring with inspectable ranking.

What This Is

arXiv Discovery MCP is a research discovery substrate inspired by arxiv-sanity. It helps researchers and AI agents discover, triage, and monitor arXiv papers through the Model Context Protocol (MCP) -- exposing tools, resources, and prompts that integrate directly into Claude Desktop, Claude Code, or any MCP-compatible client.

Unlike "chat with papers" wrappers, this system provides explicit interest modeling, inspectable ranking explanations, and structured workflow state. You build an interest profile from seed papers, followed authors, and saved queries. The system uses that profile to rank search results, surface new papers, and explain why each result scored the way it did.

The project tracks content provenance and respects reuse constraints per content type. All ranking signals are transparent: you can see exactly which interest signals contributed to each paper's score.

Related MCP server: arxiv-reader-mcp

Features

MCP Tools (13)

Discovery

  • search_papers -- Full-text search with optional profile-ranked results

  • browse_recent -- Browse recent papers by arXiv category

  • find_related_papers -- Find papers related to one or more seed papers

  • get_paper -- Retrieve full metadata for a single paper

Workflow

  • triage_paper -- Mark papers as shortlisted, seen, or dismissed

  • add_to_collection -- Add papers to named collections (auto-creates)

  • create_watch -- Create a saved query that monitors for new papers

Interest & Enrichment

  • add_signal -- Add an interest signal (seed paper, followed author, etc.)

  • batch_add_signals -- Add multiple interest signals at once

  • create_profile -- Create a named interest profile

  • suggest_signals -- Get profile expansion suggestions based on usage patterns

  • enrich_paper -- Fetch citation counts, FWCI, and topics from OpenAlex

Content

  • get_content_variant -- Retrieve paper content (abstract, HTML, or PDF-to-markdown) with rights gating

MCP Resources (4)

  • paper://{arxiv_id} -- Paper metadata, triage state, enrichment, and content variants

  • collection://{slug} -- Collection contents with pagination

  • profile://{slug} -- Interest profile with all signals

  • watch://{slug}/deltas -- New papers since last check

MCP Prompts (3)

  • daily-digest -- Workflow guidance for reviewing new papers across watches

  • literature-map-from-seeds -- Workflow for building a literature map from seed papers

  • triage-shortlist -- Workflow for reviewing and triaging a collection

CLI

All MCP capabilities are mirrored in a full CLI (arxiv-mcp) for terminal workflows, scripting, and debugging.

Test Coverage

493 tests passing across ingestion, search, workflow, interest modeling, enrichment, content normalization, and MCP integration.

Prerequisites

  • Python 3.13+ (uses 3.13 language features)

  • PostgreSQL 16+ (must be running and accessible)

  • Git (for cloning the repository)

Installation

git clone https://github.com/loganrooks/arxiv-sanity-mcp.git
cd arxiv-sanity-mcp

# Create and activate a virtual environment (recommended)
python3.13 -m venv .venv
source .venv/bin/activate  # Linux/macOS
# .venv\Scripts\activate   # Windows

pip install -e .

For development (tests, linting):

pip install -e ".[dev]"
# Or manually:
pip install pytest pytest-asyncio pytest-cov pytest-timeout respx ruff

Note: The MCP server configuration below requires the absolute path to your venv's Python interpreter. You can find it with which python after activating the venv.

Database Setup

  1. Create the database user and databases:

sudo -u postgres psql -c "CREATE USER arxiv_mcp WITH PASSWORD 'arxiv_mcp_dev';"
sudo -u postgres psql -c "CREATE DATABASE arxiv_mcp OWNER arxiv_mcp;"
sudo -u postgres psql -c "CREATE DATABASE arxiv_mcp_test OWNER arxiv_mcp;"
  1. Create a .env file in the project root (or set environment variables):

DATABASE_URL=postgresql+asyncpg://arxiv_mcp:arxiv_mcp_dev@localhost:5432/arxiv_mcp
  1. Run database migrations:

alembic upgrade head

Quick Start

Once installed and the database is set up, try these commands:

# Harvest a paper by arXiv ID
arxiv-mcp harvest fetch 2301.00001

# Search for papers
arxiv-mcp search query "attention mechanism"

# Browse recent papers in a category
arxiv-mcp search browse --category cs.AI

# Create a collection
arxiv-mcp collection create "reading-list"

# Triage a paper
arxiv-mcp triage mark 2301.00001 shortlisted

MCP Server Configuration

Claude Code

Use claude mcp add-json with your venv's absolute Python path:

claude mcp add-json arxiv-discovery --scope local '{
  "command": "/absolute/path/to/arxiv-sanity-mcp/.venv/bin/python",
  "args": ["-m", "arxiv_mcp.mcp"],
  "cwd": "/absolute/path/to/arxiv-sanity-mcp",
  "env": {
    "DATABASE_URL": "postgresql+asyncpg://arxiv_mcp:arxiv_mcp_dev@localhost:5432/arxiv_mcp"
  }
}'

Replace /absolute/path/to/arxiv-sanity-mcp with the actual path to your cloned repository.

Important: Use the absolute path to the venv Python binary (e.g., /home/user/projects/arxiv-sanity-mcp/.venv/bin/python), not just python. This ensures the MCP server uses the correct environment regardless of which directory you launch Claude Code from.

Claude Desktop

Add this to your claude_desktop_config.json:

{
  "mcpServers": {
    "arxiv-discovery": {
      "command": "/absolute/path/to/arxiv-sanity-mcp/.venv/bin/python",
      "args": ["-m", "arxiv_mcp.mcp"],
      "cwd": "/absolute/path/to/arxiv-sanity-mcp",
      "env": {
        "DATABASE_URL": "postgresql+asyncpg://arxiv_mcp:arxiv_mcp_dev@localhost:5432/arxiv_mcp"
      }
    }
  }
}

Replace /absolute/path/to/arxiv-sanity-mcp with the actual path to your cloned repository.

Configuration

Variable

Required

Default

Description

DATABASE_URL

Yes

postgresql+asyncpg://arxiv_mcp:arxiv_mcp_dev@localhost:5432/arxiv_mcp

PostgreSQL connection string

OPENALEX_EMAIL

No

(empty)

Email for OpenAlex polite pool (recommended; increases rate limit from 1 to 10 req/s)

OPENALEX_API_KEY

No

(empty)

OpenAlex API key for enrichment

DEPLOYMENT_MODE

No

local

local or hosted -- controls content license enforcement

Design Documents

Architectural documentation is in the docs/ directory:

Document

Description

01 - Project Vision

Goals and product values

02 - Product Principles

Design principles and constraints

03 - Design Space

Retrieval and ranking options explored

04 - Reference Designs

Systems studied for design inspiration

05 - Architecture Hypotheses

Architectural bets and rationale

06 - MCP Surface Options

MCP interface design decisions

07 - Data Sources & Content Rights

arXiv data access and licensing

08 - Evaluation & Experiments

Testing methodology

09 - Roadmap

Development phases

10 - Open Questions

Unresolved design questions

11 - Sources

External references

Architecture Decision Records

  • ADR-0001 -- Exploration-first architecture

  • ADR-0002 -- Metadata-first, lazy enrichment

  • ADR-0003 -- License and provenance first

  • ADR-0004 -- MCP as workflow substrate

See docs/adrs/ for full details.

License

MIT

Available Tools

13 tools
add_signalC

Add a signal to an interest profile.

Signal types: seed_paper (arXiv ID), saved_query (query slug), followed_author (author name), negative_example (arXiv ID).

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNo
signal_typeYes
profile_slugYes
signal_valueYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It communicates that the tool mutates an interest profile by adding a signal, but it does not disclose whether duplicates are prevented, whether validation occurs, what side effects arise, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action, and the signal-type list is presented efficiently. Every sentence contributes useful information without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has four parameters, no annotations, and no output schema, so the description needs to cover more than just valid signal types. It omits the meaning of profile_slug, the optional reason parameter, return behavior, and any differences from batch_add_signals, making it incomplete for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds real meaning by defining signal_type allowed values and signal_value expected formats (arXiv ID, query slug, author name), but it leaves profile_slug and reason undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Add a signal to an interest profile') and clearly enumerates the four supported signal types. It is distinct from sibling tools in name and function, though it does not explicitly contrast itself with batch_add_signals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as batch_add_signals or suggest_signals. It implies usage by defining signal types, but it never states conditions, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_to_collectionB

Add a paper to a collection. Creates the collection if it doesn't exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
arxiv_idYes
collection_nameYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It does disclose a useful side effect: the collection is created if it does not exist. However, it omits other relevant behaviors such as duplicate handling, idempotency, permissions, or what happens if the paper is already in the collection.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary action and followed by the key side effect. There is no filler or redundant explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a simple two-parameter tool: it states the action and the notable auto-creation behavior. But it lacks usage context and parameter-level guidance, and no output schema or annotations fill those gaps, so it remains minimally complete rather than fully informative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. It only maps arxiv_id to 'paper' and collection_name to 'collection' conceptually, adding no format, constraints, or usage details beyond the property names already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Add a paper to a collection.' The side-effect note distinguishes it from siblings like triage_paper or add_signal, which target different workflows.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention related tools or conditions that would make this tool preferable, and 'Creates the collection if it doesn't exist' is a side effect, not usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch_add_signalsA

Add multiple signals to an interest profile in one call.

Each signal dict must have: signal_type, signal_value. Optional per-signal: reason.

Returns a summary with counts and per-signal results. Continues on individual signal errors (partial success is OK).

ParametersJSON Schema
NameRequiredDescriptionDefault
signalsYes
profile_slugYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden. It explicitly discloses partial success behavior, continued processing on individual errors, and a return summary with counts and per-signal results. It does not mention idempotency, duplicates, or batch limits, but covers the most important behavioral traits for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every line earns its place: the purpose, per-item schema, return value, and error semantics are each stated in a compact sentence or fragment. The most important information is front-loaded, and there is no redundant wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description provides enough for an agent to invoke the tool correctly: what parameters to provide, what each signal dict must contain, what the response summarizes, and how errors are handled. Minor gaps remain around supported signal_type values and profile_slug format, but these are not essential for a basic correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description compensates by documenting the per-signal required fields signal_type and signal_value, plus optional reason. It also clarifies the shape of the signals array, which the raw schema leaves as open objects. The profile_slug parameter is not directly described, but its name and the description's 'interest profile' phrase make its meaning reasonably clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: adding multiple signals to an interest profile in a single call. It also differentiates from the sibling add_signal by emphasizing 'multiple' and 'in one call', so an agent can distinguish batch behavior from singular behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'in one call' implies the batch use case versus the singular sibling add_signal, giving clear context for when to choose this tool. However, it does not explicitly state when to avoid it or name an alternative, so it stops just short of full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browse_recentA

Browse recently announced arXiv papers, optionally filtered by category.

Use time_basis to select ordering: announced, submitted, or updated. Note: Many papers may not have announced_date populated. If results are empty with the default time_basis='announced', try time_basis='submitted'.

Optionally provide profile_slug to get profile-ranked results with ranking explanations on each result.

Response shape: {"results": {"items": [...], "page_info": {...}}, "ranker_snapshot": ...}

Without profile_slug: items include triage_state and collection_slugs; ranking_explanation is null; ranker_snapshot is null.

With profile_slug: items additionally include ranking_explanation and the response includes a ranker_snapshot capturing ranker config.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
cursorNo
categoryNo
page_sizeNo
time_basisNoannounced
profile_slugNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It transparently warns that many papers may lack announced_date, suggests a fallback, and details the response shape differences based on profile_slug presence (including which fields are null without it and the inclusion of ranker_snapshot with it). This is a thorough disclosure of what the tool returns and under what conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but well-organized: it opens with the core purpose, then provides parameter guidance, then describes the response shape. Each sentence adds value, and the most critical usage advice (time_basis fallback) is placed early. It is longer than a one-liner, but the complexity of the tool justifies the length. No filler or redundancy is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters and no output schema or annotations, the description covers the essential behavior: the main purpose, the time_basis nuance, the profile_slug effect, and the response structure. It does not explain pagination (cursor, page_size) or valid category values, but these are standard and inferable. The description is sufficient for an agent to call the tool correctly in most scenarios, though edge cases like pagination are not explicitly addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning to time_basis (announced, submitted, updated) and profile_slug (profile-ranked results with explanations), and clarifies the effect on the response. However, it does not explain days, cursor, page_size, or the format of category beyond 'optionally filtered'. These parameters are left to inference from names and defaults. The description partially compensates for the schema gap but not completely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'browse' and the resource 'recently announced arXiv papers' with an optional category filter. This distinguishes it from sibling tools like search_papers and find_related_papers, which serve different retrieval needs. The title 'browse_recent' is reinforced, but the description goes beyond a tautology by specifying the arXiv scope and the optional filtering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use time_basis (announced vs submitted vs updated) and includes a practical fallback if results are empty with the default. It also explains the conditional use of profile_slug to get ranked results with explanations. While it doesn't explicitly name alternative tools, the context clearly implies that this is for browsing recent papers, not searching or related-paper lookups. The usage advice is actionable and specific.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_profileB

Create a new interest profile.

Returns the profile summary with slug, name, signal_count, and timestamps.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
negative_weightNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool persists a new profile and returns a summary with slug, name, signal_count, and timestamps. It does not cover duplicates, idempotency, or side effects, but the core mutation behavior is explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler. The primary action is front-loaded, and the return summary is a useful addition. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter creation tool, the basics are present: what is created and what is returned. However, with no annotations, no output schema, and no explanation of the optional parameter or when to use this tool, the description leaves moderate gaps that an agent would need to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not compensate. 'name' is self-explanatory, but 'negative_weight' is a non-obvious optional parameter that is never explained. The return fields mention signal_count but not how negative_weight affects it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Create') and resource ('new interest profile'), and even lists the return payload. It distinguishes itself from sibling create_watch through the noun 'profile', though it doesn't explicitly differentiate from other profile-related operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to create a profile versus using alternatives like create_watch, add_signal, or add_to_collection. There are no prerequisites, exclusions, or context clues beyond the name itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_watchA

Create a monitored search that tracks new papers matching your query.

Read the watch://{slug}/deltas resource to see papers added since your last check.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
queryYes
cadenceNodaily
categoryNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the tool creates a persistent monitored search and that results are consumed by reading watch://{slug}/deltas, including incremental papers added 'since your last check.' It does not cover every edge case, but core side effects and the consumption model are clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary action steals. The follow-up instruction about reading the deltas resource is separate and useful, with no redundant wording or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main workflow and points the agent to the correct output resource, which is valuable given no output schema. However, it does not explain where the slug comes from or clarify optional parameters, so an agent must infer several details to make a fully informed invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only mentions 'your query' in prose. It does not explain name, cadence, or category, leaving most parameters semantically under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Create a monitored search that tracks new papers matching your query.' This is distinct from one-off sibling tools like search_papers or browse_recent, which do not create a persistent, monitored activity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance about when to choose create_watch over alternatives. The second sentence describes what to do after creation (read the deltas resource), but it does not explain when a monitored watch is preferable to a direct search or browse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enrich_paperB

Trigger OpenAlex enrichment for a paper to get topics, citations, and related works.

Set refresh=True to re-enrich even if data exists within the cooldown window.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNo
arxiv_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose a cooldown window and the refresh flag's re-enrichment behavior, which is useful. However, it does not state whether this is an asynchronous trigger, what side effects occur, whether authentication is needed, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The main purpose is front-loaded, and the refresh nuance is placed in a separate, clearly scoped sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple trigger tool, the description gives enough to attempt a call: it identifies the required parameter and the purpose. Yet with no output schema, it fails to explain whether the tool returns enriched data immediately or only initiates enrichment, and it omits cooldown duration and error behavior. These are meaningful gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains refresh=True's meaning relative to the cooldown window, adding value beyond the schema. But arxiv_id is left entirely to its name and required status; no format, example, or context is given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource ('Trigger OpenAlex enrichment') and lists the expected outputs: topics, citations, and related works. It does not explicitly contrast with siblings like find_related_papers, so it lacks full differentiation, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to choose enrich_paper over alternatives such as find_related_papers, get_paper, or triage_paper. It only explains the refresh flag's effect within a cooldown window, which is operational detail, not tool-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_content_variantA

Get paper content at the requested fidelity level.

Retrieves paper content as abstract, HTML, or PDF-derived markdown. Use variant='best' (default) to get the highest-quality available format, which tries HTML first then falls back to PDF markdown.

Valid variants: 'abstract', 'html', 'pdf_markdown', 'best'

Returns content with provenance metadata (source, backend, license). For non-abstract variants, respects per-paper license restrictions.

ParametersJSON Schema
NameRequiredDescriptionDefault
variantNobest
arxiv_idYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses useful traits: best-variant resolution order (HTML first, then PDF markdown), provenance metadata in the response, and license restrictions for non-abstract variants. It does not mention rate limits or auth, but the read-only nature is strongly implied by 'Get' and the content-focused language.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded and each sentence adds value: main purpose, variant behavior, valid values, provenance metadata, and licensing. There is minor redundancy between 'Get paper content' and 'Retrieves paper content', but the overall structure is tight and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple retrieval tool with no annotations and no output schema, the description covers the critical operational details: supported variants, default behavior, fallback order, provenance, and licensing. Formatting details for arxiv_id and explicit error behavior are absent, but these are minor for a tool this straightforward.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the bare input schema. It thoroughly explains the variant parameter, including all valid values, the default, and the fallback behavior. The arxiv_id parameter is left to its self-evident name and the surrounding context, which is sufficient for a straightforward identifier.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Get paper content at the requested fidelity level.' It then enumerates exact content types (abstract, HTML, PDF-derived markdown) and distinguishes this from the sibling get_paper by focusing on content variants rather than paper metadata or search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear guidance on which variant to choose ('Use variant='best' (default)') and lists valid variants, but it does not explicitly state when to prefer this tool over siblings like get_paper or search_papers. Usage context is mostly implied by the tool's name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_paperA

Get metadata for a single paper by its arXiv ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
arxiv_idYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It communicates a read-only metadata lookup through the verb 'Get', but it does not describe what happens for invalid IDs, whether the operation has side effects, or what the returned metadata shape will be. Adequate, but not detailed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single clear sentence with no filler. It front-loads the action, scope, and required identifier, and every word contributes to understanding the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only retrieval tool with no output schema, the description is nearly sufficient. It identifies the operation, target, and input, but it could be more complete by describing return behavior or invalid-ID outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the parameter name and type with no description. The description adds that the parameter selects a single paper, but it does not give format examples, versioning guidance, or constraints, so it only partially compensates for the 0% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get metadata'), a specific resource ('a single paper'), and the exact identifier ('arXiv ID'). The word 'single' distinguishes this from sibling tools like search_papers and browse_recent, making the intent unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by its arXiv ID' implies the tool is appropriate when the agent already has an exact ID, but it does not explicitly say when to use this instead of search_papers or other alternatives. Usage context is implied rather than directly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_papersA

Search for arXiv papers by text, title, author, or category.

Returns paginated results with paper metadata and relevance scores. Use cursor from previous results to get the next page.

Optionally provide profile_slug to get profile-ranked results with ranking explanations on each result.

Response shape: {"results": {"items": [...], "page_info": {...}}, "ranker_snapshot": ...}

Without profile_slug: items include triage_state and collection_slugs; ranking_explanation is null; ranker_snapshot is null.

With profile_slug: items additionally include ranking_explanation and the response includes a ranker_snapshot capturing ranker config.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo
titleNo
authorNo
cursorNo
date_toNo
categoryNo
date_fromNo
page_sizeNo
time_basisNoannounced
profile_slugNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears the full disclosure burden, and it delivers: it explains pagination behavior, the optional profile-ranking mode, the exact response shape, and precisely how items differ with and without profile_slug. This goes well beyond the input schema and gives the agent a realistic model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then moves logically through pagination, optional ranking mode, and response shape. Each sentence adds information an agent needs, with no filler or restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers the most important runtime behavior and return-value details, including conditional fields and ranker_snapshot. Minor gaps remain, such as date filtering semantics and how multiple search parameters combine, but the tool is callable with reasonable confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add meaning for query, title, author, category, cursor, and profile_slug, but it leaves date_from, date_to, time_basis, and page_size unexplained. This is meaningful but incomplete compensation for a 10-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description immediately states a specific verb and resource: 'Search for arXiv papers', and enumerates the search dimensions (text, title, author, or category). This clearly separates it from sibling tools like find_related_papers or browse_recent, which imply different discovery mechanisms.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: how to paginate with cursor, how to opt into profile-ranked results, and what to expect in each mode. It does not explicitly name alternatives or state when not to use this tool, so it stops short of a 5, but the context is strong enough for an agent to use it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_signalsA

Generate signal suggestions for an interest profile.

Examines workflow activity (triaged papers, frequent queries, recurring authors) to suggest new signals. With auto_add=True, adds suggestions as pending signals automatically.

Returns candidates list and added_count.

ParametersJSON Schema
NameRequiredDescriptionDefault
auto_addNo
profile_slugYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It clearly states the conditional side effect: with auto_add=True it adds suggestions as pending signals automatically. It also discloses the return fields (candidates list and added_count). It doesn't cover permissions or reversibility, but for a suggestion tool the key mutating behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with a clear front-loaded purpose. It wastes no words: purpose, data sources, conditional side effect, and return values each get concise, dedicated coverage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (two params, no output schema), the description covers the essential parts: what it analyzes, the optional mutation, and what it returns. It doesn't discuss error cases or profile lookup failures, but an agent has enough to invoke it correctly for the common case.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains auto_add behavior well, noting that true triggers automatic addition of pending signals. For profile_slug, however, it only says 'for an interest profile' without explicitly mapping the parameter or clarifying slug/format, leaving some parameter meaning under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object ('Generate signal suggestions') and identifies the resource ('interest profile'). It also specifies the data sources (triaged papers, frequent queries, recurring authors), which distinguishes it from sibling add/search tools. However, it doesn't explicitly name or contrast siblings, so it falls slightly short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: you want signal suggestions derived from workflow activity, and it even lists the activity types considered. It does not, however, say when not to use this tool or mention alternatives (e.g., using add_signal or batch_add_signals for known signals), leaving the guidance implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

triage_paperB

Set the triage state for a paper.

Valid states: seen, shortlisted, dismissed, read, cite-later, archived, unseen.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes
reasonNo
arxiv_idYes

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It states that the tool mutates a paper's triage state but does not clarify whether the state overwrites the previous one, how the optional reason parameter behaves, whether the operation is reversible, or what side effects might occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the action and object appear in the first sentence, and the valid states are enumerated in a scannable list. Every sentence earns its place with no unnecessary wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with no annotations, no output schema, and sparse parameter documentation, the description is too thin to be fully actionable. It omits behavioral details and leaves the reason parameter and arxiv_id format unexplained, so an agent may not know all consequences of invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It lists valid values for the 'state' parameter, which is helpful, but it provides no explanation for 'arxiv_id' beyond the schema title and no semantics for the 'reason' parameter. Parameter understanding remains incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Set the triage state for a paper.' It clearly identifies the action and the object, and the valid state list adds meaning. However, it does not explicitly differentiate this tool from siblings, though the operation is fairly unique among them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when an agent needs to assign or change a paper's triage state, and the list of valid states gives useful context. It provides no explicit guidance on when not to use it or which sibling tool to prefer for related operations like collections or signals.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.1.0
    • First observedadd_signal
    • First observedadd_to_collection
    • First observedbatch_add_signals
    • First observedbrowse_recent
    • First observedcreate_profile
    • First observedcreate_watch
    • First observedenrich_paper
    • First observedfind_related_papers
    • First observedget_content_variant
    • First observedget_paper
    • First observedsearch_papers
    • First observedsuggest_signals
    • First observedtriage_paper

TDQS

A3.6/5.0

Scored across 13 tools

Disambiguation4/5

Most tools target distinct actions: search, browse, related-papers, get-paper, triage, watch, signals, enrichment, and content retrieval are clearly separated. The only mild ambiguity is between search_papers and browse_recent since both return ranked paper lists with optional profile_slug, but their descriptions make the query-vs-recency distinction clear.

Naming Consistency4/5

The majority follow a verb_noun pattern such as get_paper, create_profile, add_signal, and enrich_paper. Minor deviations like browse_recent (verb + adjective) and batch_add_signals break the pattern slightly, but the naming remains predictable and readable.

Tool Count5/5

13 tools is a well-scoped count for an arXiv discovery server. Each tool contributes a meaningful capability without redundancy, and the set is neither too thin nor overloaded.

Completeness3/5

The discovery side is well covered with search, browse, get, related papers, content variants, and enrichment. However, the personalization surface has notable lifecycle gaps: you can create profiles, create watches, add signals, and add to collections, but there are no remove, delete, or list operations for these resources.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers