arXiv Discovery MCP
Allows discovery, triage, and monitoring of arXiv papers with search, browsing, collections, and interest-based ranking.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@arXiv Discovery MCPsearch for papers on attention mechanisms"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
arXiv Discovery MCP
MCP server for arXiv paper discovery, triage, and monitoring with inspectable ranking.
What This Is
arXiv Discovery MCP is a research discovery substrate inspired by arxiv-sanity. It helps researchers and AI agents discover, triage, and monitor arXiv papers through the Model Context Protocol (MCP) -- exposing tools, resources, and prompts that integrate directly into Claude Desktop, Claude Code, or any MCP-compatible client.
Unlike "chat with papers" wrappers, this system provides explicit interest modeling, inspectable ranking explanations, and structured workflow state. You build an interest profile from seed papers, followed authors, and saved queries. The system uses that profile to rank search results, surface new papers, and explain why each result scored the way it did.
The project tracks content provenance and respects reuse constraints per content type. All ranking signals are transparent: you can see exactly which interest signals contributed to each paper's score.
Related MCP server: arxiv-reader-mcp
Features
MCP Tools (13)
Discovery
search_papers-- Full-text search with optional profile-ranked resultsbrowse_recent-- Browse recent papers by arXiv categoryfind_related_papers-- Find papers related to one or more seed papersget_paper-- Retrieve full metadata for a single paper
Workflow
triage_paper-- Mark papers as shortlisted, seen, or dismissedadd_to_collection-- Add papers to named collections (auto-creates)create_watch-- Create a saved query that monitors for new papers
Interest & Enrichment
add_signal-- Add an interest signal (seed paper, followed author, etc.)batch_add_signals-- Add multiple interest signals at oncecreate_profile-- Create a named interest profilesuggest_signals-- Get profile expansion suggestions based on usage patternsenrich_paper-- Fetch citation counts, FWCI, and topics from OpenAlex
Content
get_content_variant-- Retrieve paper content (abstract, HTML, or PDF-to-markdown) with rights gating
MCP Resources (4)
paper://{arxiv_id}-- Paper metadata, triage state, enrichment, and content variantscollection://{slug}-- Collection contents with paginationprofile://{slug}-- Interest profile with all signalswatch://{slug}/deltas-- New papers since last check
MCP Prompts (3)
daily-digest-- Workflow guidance for reviewing new papers across watchesliterature-map-from-seeds-- Workflow for building a literature map from seed paperstriage-shortlist-- Workflow for reviewing and triaging a collection
CLI
All MCP capabilities are mirrored in a full CLI (arxiv-mcp) for terminal workflows, scripting, and debugging.
Test Coverage
493 tests passing across ingestion, search, workflow, interest modeling, enrichment, content normalization, and MCP integration.
Prerequisites
Python 3.13+ (uses 3.13 language features)
PostgreSQL 16+ (must be running and accessible)
Git (for cloning the repository)
Installation
git clone https://github.com/loganrooks/arxiv-sanity-mcp.git
cd arxiv-sanity-mcp
# Create and activate a virtual environment (recommended)
python3.13 -m venv .venv
source .venv/bin/activate # Linux/macOS
# .venv\Scripts\activate # Windows
pip install -e .For development (tests, linting):
pip install -e ".[dev]"
# Or manually:
pip install pytest pytest-asyncio pytest-cov pytest-timeout respx ruffNote: The MCP server configuration below requires the absolute path to your venv's Python interpreter. You can find it with
which pythonafter activating the venv.
Database Setup
Create the database user and databases:
sudo -u postgres psql -c "CREATE USER arxiv_mcp WITH PASSWORD 'arxiv_mcp_dev';"
sudo -u postgres psql -c "CREATE DATABASE arxiv_mcp OWNER arxiv_mcp;"
sudo -u postgres psql -c "CREATE DATABASE arxiv_mcp_test OWNER arxiv_mcp;"Create a
.envfile in the project root (or set environment variables):
DATABASE_URL=postgresql+asyncpg://arxiv_mcp:arxiv_mcp_dev@localhost:5432/arxiv_mcpRun database migrations:
alembic upgrade headQuick Start
Once installed and the database is set up, try these commands:
# Harvest a paper by arXiv ID
arxiv-mcp harvest fetch 2301.00001
# Search for papers
arxiv-mcp search query "attention mechanism"
# Browse recent papers in a category
arxiv-mcp search browse --category cs.AI
# Create a collection
arxiv-mcp collection create "reading-list"
# Triage a paper
arxiv-mcp triage mark 2301.00001 shortlistedMCP Server Configuration
Claude Code
Use claude mcp add-json with your venv's absolute Python path:
claude mcp add-json arxiv-discovery --scope local '{
"command": "/absolute/path/to/arxiv-sanity-mcp/.venv/bin/python",
"args": ["-m", "arxiv_mcp.mcp"],
"cwd": "/absolute/path/to/arxiv-sanity-mcp",
"env": {
"DATABASE_URL": "postgresql+asyncpg://arxiv_mcp:arxiv_mcp_dev@localhost:5432/arxiv_mcp"
}
}'Replace /absolute/path/to/arxiv-sanity-mcp with the actual path to your cloned repository.
Important: Use the absolute path to the venv Python binary (e.g.,
/home/user/projects/arxiv-sanity-mcp/.venv/bin/python), not justpython. This ensures the MCP server uses the correct environment regardless of which directory you launch Claude Code from.
Claude Desktop
Add this to your claude_desktop_config.json:
{
"mcpServers": {
"arxiv-discovery": {
"command": "/absolute/path/to/arxiv-sanity-mcp/.venv/bin/python",
"args": ["-m", "arxiv_mcp.mcp"],
"cwd": "/absolute/path/to/arxiv-sanity-mcp",
"env": {
"DATABASE_URL": "postgresql+asyncpg://arxiv_mcp:arxiv_mcp_dev@localhost:5432/arxiv_mcp"
}
}
}
}Replace /absolute/path/to/arxiv-sanity-mcp with the actual path to your cloned repository.
Configuration
Variable | Required | Default | Description |
| Yes |
| PostgreSQL connection string |
| No | (empty) | Email for OpenAlex polite pool (recommended; increases rate limit from 1 to 10 req/s) |
| No | (empty) | OpenAlex API key for enrichment |
| No |
|
|
Design Documents
Architectural documentation is in the docs/ directory:
Document | Description |
Goals and product values | |
Design principles and constraints | |
Retrieval and ranking options explored | |
Systems studied for design inspiration | |
Architectural bets and rationale | |
MCP interface design decisions | |
arXiv data access and licensing | |
Testing methodology | |
Development phases | |
Unresolved design questions | |
External references |
Architecture Decision Records
ADR-0001 -- Exploration-first architecture
ADR-0002 -- Metadata-first, lazy enrichment
ADR-0003 -- License and provenance first
ADR-0004 -- MCP as workflow substrate
See docs/adrs/ for full details.
License
Available Tools
13 toolsadd_signalC
Add a signal to an interest profile.
Signal types: seed_paper (arXiv ID), saved_query (query slug), followed_author (author name), negative_example (arXiv ID).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| signal_type | Yes | ||
| profile_slug | Yes | ||
| signal_value | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It communicates that the tool mutates an interest profile by adding a signal, but it does not disclose whether duplicates are prevented, whether validation occurs, what side effects arise, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action, and the signal-type list is presented efficiently. Every sentence contributes useful information without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has four parameters, no annotations, and no output schema, so the description needs to cover more than just valid signal types. It omits the meaning of profile_slug, the optional reason parameter, return behavior, and any differences from batch_add_signals, making it incomplete for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning by defining signal_type allowed values and signal_value expected formats (arXiv ID, query slug, author name), but it leaves profile_slug and reason undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Add a signal to an interest profile') and clearly enumerates the four supported signal types. It is distinct from sibling tools in name and function, though it does not explicitly contrast itself with batch_add_signals.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as batch_add_signals or suggest_signals. It implies usage by defining signal types, but it never states conditions, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_to_collectionB
Add a paper to a collection. Creates the collection if it doesn't exist.
| Name | Required | Description | Default |
|---|---|---|---|
| arxiv_id | Yes | ||
| collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does disclose a useful side effect: the collection is created if it does not exist. However, it omits other relevant behaviors such as duplicate handling, idempotency, permissions, or what happens if the paper is already in the collection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the primary action and followed by the key side effect. There is no filler or redundant explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple two-parameter tool: it states the action and the notable auto-creation behavior. But it lacks usage context and parameter-level guidance, and no output schema or annotations fill those gaps, so it remains minimally complete rather than fully informative.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It only maps arxiv_id to 'paper' and collection_name to 'collection' conceptually, adding no format, constraints, or usage details beyond the property names already present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Add a paper to a collection.' The side-effect note distinguishes it from siblings like triage_paper or add_signal, which target different workflows.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description does not mention related tools or conditions that would make this tool preferable, and 'Creates the collection if it doesn't exist' is a side effect, not usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_add_signalsA
Add multiple signals to an interest profile in one call.
Each signal dict must have: signal_type, signal_value. Optional per-signal: reason.
Returns a summary with counts and per-signal results. Continues on individual signal errors (partial success is OK).
| Name | Required | Description | Default |
|---|---|---|---|
| signals | Yes | ||
| profile_slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. It explicitly discloses partial success behavior, continued processing on individual errors, and a return summary with counts and per-signal results. It does not mention idempotency, duplicates, or batch limits, but covers the most important behavioral traits for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every line earns its place: the purpose, per-item schema, return value, and error semantics are each stated in a compact sentence or fragment. The most important information is front-loaded, and there is no redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description provides enough for an agent to invoke the tool correctly: what parameters to provide, what each signal dict must contain, what the response summarizes, and how errors are handled. Minor gaps remain around supported signal_type values and profile_slug format, but these are not essential for a basic correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates by documenting the per-signal required fields signal_type and signal_value, plus optional reason. It also clarifies the shape of the signals array, which the raw schema leaves as open objects. The profile_slug parameter is not directly described, but its name and the description's 'interest profile' phrase make its meaning reasonably clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: adding multiple signals to an interest profile in a single call. It also differentiates from the sibling add_signal by emphasizing 'multiple' and 'in one call', so an agent can distinguish batch behavior from singular behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in one call' implies the batch use case versus the singular sibling add_signal, giving clear context for when to choose this tool. However, it does not explicitly state when to avoid it or name an alternative, so it stops just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browse_recentA
Browse recently announced arXiv papers, optionally filtered by category.
Use time_basis to select ordering: announced, submitted, or updated. Note: Many papers may not have announced_date populated. If results are empty with the default time_basis='announced', try time_basis='submitted'.
Optionally provide profile_slug to get profile-ranked results with ranking explanations on each result.
Response shape: {"results": {"items": [...], "page_info": {...}}, "ranker_snapshot": ...}
Without profile_slug: items include triage_state and collection_slugs; ranking_explanation is null; ranker_snapshot is null.
With profile_slug: items additionally include ranking_explanation and the response includes a ranker_snapshot capturing ranker config.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | ||
| cursor | No | ||
| category | No | ||
| page_size | No | ||
| time_basis | No | announced | |
| profile_slug | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It transparently warns that many papers may lack announced_date, suggests a fallback, and details the response shape differences based on profile_slug presence (including which fields are null without it and the inclusion of ranker_snapshot with it). This is a thorough disclosure of what the tool returns and under what conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed but well-organized: it opens with the core purpose, then provides parameter guidance, then describes the response shape. Each sentence adds value, and the most critical usage advice (time_basis fallback) is placed early. It is longer than a one-liner, but the complexity of the tool justifies the length. No filler or redundancy is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters and no output schema or annotations, the description covers the essential behavior: the main purpose, the time_basis nuance, the profile_slug effect, and the response structure. It does not explain pagination (cursor, page_size) or valid category values, but these are standard and inferable. The description is sufficient for an agent to call the tool correctly in most scenarios, though edge cases like pagination are not explicitly addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning to time_basis (announced, submitted, updated) and profile_slug (profile-ranked results with explanations), and clarifies the effect on the response. However, it does not explain days, cursor, page_size, or the format of category beyond 'optionally filtered'. These parameters are left to inference from names and defaults. The description partially compensates for the schema gap but not completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'browse' and the resource 'recently announced arXiv papers' with an optional category filter. This distinguishes it from sibling tools like search_papers and find_related_papers, which serve different retrieval needs. The title 'browse_recent' is reinforced, but the description goes beyond a tautology by specifying the arXiv scope and the optional filtering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use time_basis (announced vs submitted vs updated) and includes a practical fallback if results are empty with the default. It also explains the conditional use of profile_slug to get ranked results with explanations. While it doesn't explicitly name alternative tools, the context clearly implies that this is for browsing recent papers, not searching or related-paper lookups. The usage advice is actionable and specific.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_profileB
Create a new interest profile.
Returns the profile summary with slug, name, signal_count, and timestamps.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| negative_weight | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool persists a new profile and returns a summary with slug, name, signal_count, and timestamps. It does not cover duplicates, idempotency, or side effects, but the core mutation behavior is explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler. The primary action is front-loaded, and the return summary is a useful addition. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter creation tool, the basics are present: what is created and what is returned. However, with no annotations, no output schema, and no explanation of the optional parameter or when to use this tool, the description leaves moderate gaps that an agent would need to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate. 'name' is self-explanatory, but 'negative_weight' is a non-obvious optional parameter that is never explained. The return fields mention signal_count but not how negative_weight affects it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create') and resource ('new interest profile'), and even lists the return payload. It distinguishes itself from sibling create_watch through the noun 'profile', though it doesn't explicitly differentiate from other profile-related operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to create a profile versus using alternatives like create_watch, add_signal, or add_to_collection. There are no prerequisites, exclusions, or context clues beyond the name itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_watchA
Create a monitored search that tracks new papers matching your query.
Read the watch://{slug}/deltas resource to see papers added since your last check.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| query | Yes | ||
| cadence | No | daily | |
| category | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the tool creates a persistent monitored search and that results are consumed by reading watch://{slug}/deltas, including incremental papers added 'since your last check.' It does not cover every edge case, but core side effects and the consumption model are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the primary action steals. The follow-up instruction about reading the deltas resource is separate and useful, with no redundant wording or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main workflow and points the agent to the correct output resource, which is valuable given no output schema. However, it does not explain where the slug comes from or clarify optional parameters, so an agent must infer several details to make a fully informed invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only mentions 'your query' in prose. It does not explain name, cadence, or category, leaving most parameters semantically under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Create a monitored search that tracks new papers matching your query.' This is distinct from one-off sibling tools like search_papers or browse_recent, which do not create a persistent, monitored activity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance about when to choose create_watch over alternatives. The second sentence describes what to do after creation (read the deltas resource), but it does not explain when a monitored watch is preferable to a direct search or browse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enrich_paperB
Trigger OpenAlex enrichment for a paper to get topics, citations, and related works.
Set refresh=True to re-enrich even if data exists within the cooldown window.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| arxiv_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose a cooldown window and the refresh flag's re-enrichment behavior, which is useful. However, it does not state whether this is an asynchronous trigger, what side effects occur, whether authentication is needed, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The main purpose is front-loaded, and the refresh nuance is placed in a separate, clearly scoped sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple trigger tool, the description gives enough to attempt a call: it identifies the required parameter and the purpose. Yet with no output schema, it fails to explain whether the tool returns enriched data immediately or only initiates enrichment, and it omits cooldown duration and error behavior. These are meaningful gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains refresh=True's meaning relative to the cooldown window, adding value beyond the schema. But arxiv_id is left entirely to its name and required status; no format, example, or context is given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource ('Trigger OpenAlex enrichment') and lists the expected outputs: topics, citations, and related works. It does not explicitly contrast with siblings like find_related_papers, so it lacks full differentiation, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to choose enrich_paper over alternatives such as find_related_papers, get_paper, or triage_paper. It only explains the refresh flag's effect within a cooldown window, which is operational detail, not tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_content_variantA
Get paper content at the requested fidelity level.
Retrieves paper content as abstract, HTML, or PDF-derived markdown. Use variant='best' (default) to get the highest-quality available format, which tries HTML first then falls back to PDF markdown.
Valid variants: 'abstract', 'html', 'pdf_markdown', 'best'
Returns content with provenance metadata (source, backend, license). For non-abstract variants, respects per-paper license restrictions.
| Name | Required | Description | Default |
|---|---|---|---|
| variant | No | best | |
| arxiv_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses useful traits: best-variant resolution order (HTML first, then PDF markdown), provenance metadata in the response, and license restrictions for non-abstract variants. It does not mention rate limits or auth, but the read-only nature is strongly implied by 'Get' and the content-focused language.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded and each sentence adds value: main purpose, variant behavior, valid values, provenance metadata, and licensing. There is minor redundancy between 'Get paper content' and 'Retrieves paper content', but the overall structure is tight and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple retrieval tool with no annotations and no output schema, the description covers the critical operational details: supported variants, default behavior, fallback order, provenance, and licensing. Formatting details for arxiv_id and explicit error behavior are absent, but these are minor for a tool this straightforward.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare input schema. It thoroughly explains the variant parameter, including all valid values, the default, and the fallback behavior. The arxiv_id parameter is left to its self-evident name and the surrounding context, which is sufficient for a straightforward identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Get paper content at the requested fidelity level.' It then enumerates exact content types (abstract, HTML, PDF-derived markdown) and distinguishes this from the sibling get_paper by focusing on content variants rather than paper metadata or search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear guidance on which variant to choose ('Use variant='best' (default)') and lists valid variants, but it does not explicitly state when to prefer this tool over siblings like get_paper or search_papers. Usage context is mostly implied by the tool's name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_paperA
Get metadata for a single paper by its arXiv ID.
| Name | Required | Description | Default |
|---|---|---|---|
| arxiv_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It communicates a read-only metadata lookup through the verb 'Get', but it does not describe what happens for invalid IDs, whether the operation has side effects, or what the returned metadata shape will be. Adequate, but not detailed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single clear sentence with no filler. It front-loads the action, scope, and required identifier, and every word contributes to understanding the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only retrieval tool with no output schema, the description is nearly sufficient. It identifies the operation, target, and input, but it could be more complete by describing return behavior or invalid-ID outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the parameter name and type with no description. The description adds that the parameter selects a single paper, but it does not give format examples, versioning guidance, or constraints, so it only partially compensates for the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get metadata'), a specific resource ('a single paper'), and the exact identifier ('arXiv ID'). The word 'single' distinguishes this from sibling tools like search_papers and browse_recent, making the intent unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'by its arXiv ID' implies the tool is appropriate when the agent already has an exact ID, but it does not explicitly say when to use this instead of search_papers or other alternatives. Usage context is implied rather than directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersA
Search for arXiv papers by text, title, author, or category.
Returns paginated results with paper metadata and relevance scores. Use cursor from previous results to get the next page.
Optionally provide profile_slug to get profile-ranked results with ranking explanations on each result.
Response shape: {"results": {"items": [...], "page_info": {...}}, "ranker_snapshot": ...}
Without profile_slug: items include triage_state and collection_slugs; ranking_explanation is null; ranker_snapshot is null.
With profile_slug: items additionally include ranking_explanation and the response includes a ranker_snapshot capturing ranker config.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| title | No | ||
| author | No | ||
| cursor | No | ||
| date_to | No | ||
| category | No | ||
| date_from | No | ||
| page_size | No | ||
| time_basis | No | announced | |
| profile_slug | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the full disclosure burden, and it delivers: it explains pagination behavior, the optional profile-ranking mode, the exact response shape, and precisely how items differ with and without profile_slug. This goes well beyond the input schema and gives the agent a realistic model of the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then moves logically through pagination, optional ranking mode, and response shape. Each sentence adds information an agent needs, with no filler or restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers the most important runtime behavior and return-value details, including conditional fields and ranker_snapshot. Minor gaps remain, such as date filtering semantics and how multiple search parameters combine, but the tool is callable with reasonable confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaning for query, title, author, category, cursor, and profile_slug, but it leaves date_from, date_to, time_basis, and page_size unexplained. This is meaningful but incomplete compensation for a 10-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description immediately states a specific verb and resource: 'Search for arXiv papers', and enumerates the search dimensions (text, title, author, or category). This clearly separates it from sibling tools like find_related_papers or browse_recent, which imply different discovery mechanisms.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: how to paginate with cursor, how to opt into profile-ranked results, and what to expect in each mode. It does not explicitly name alternatives or state when not to use this tool, so it stops short of a 5, but the context is strong enough for an agent to use it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_signalsA
Generate signal suggestions for an interest profile.
Examines workflow activity (triaged papers, frequent queries, recurring authors) to suggest new signals. With auto_add=True, adds suggestions as pending signals automatically.
Returns candidates list and added_count.
| Name | Required | Description | Default |
|---|---|---|---|
| auto_add | No | ||
| profile_slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It clearly states the conditional side effect: with auto_add=True it adds suggestions as pending signals automatically. It also discloses the return fields (candidates list and added_count). It doesn't cover permissions or reversibility, but for a suggestion tool the key mutating behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with a clear front-loaded purpose. It wastes no words: purpose, data sources, conditional side effect, and return values each get concise, dedicated coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (two params, no output schema), the description covers the essential parts: what it analyzes, the optional mutation, and what it returns. It doesn't discuss error cases or profile lookup failures, but an agent has enough to invoke it correctly for the common case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains auto_add behavior well, noting that true triggers automatic addition of pending signals. For profile_slug, however, it only says 'for an interest profile' without explicitly mapping the parameter or clarifying slug/format, leaving some parameter meaning under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object ('Generate signal suggestions') and identifies the resource ('interest profile'). It also specifies the data sources (triaged papers, frequent queries, recurring authors), which distinguishes it from sibling add/search tools. However, it doesn't explicitly name or contrast siblings, so it falls slightly short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case: you want signal suggestions derived from workflow activity, and it even lists the activity types considered. It does not, however, say when not to use this tool or mention alternatives (e.g., using add_signal or batch_add_signals for known signals), leaving the guidance implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
triage_paperB
Set the triage state for a paper.
Valid states: seen, shortlisted, dismissed, read, cite-later, archived, unseen.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | ||
| reason | No | ||
| arxiv_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It states that the tool mutates a paper's triage state but does not clarify whether the state overwrites the previous one, how the optional reason parameter behaves, whether the operation is reversible, or what side effects might occur.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the action and object appear in the first sentence, and the valid states are enumerated in a scannable list. Every sentence earns its place with no unnecessary wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no annotations, no output schema, and sparse parameter documentation, the description is too thin to be fully actionable. It omits behavioral details and leaves the reason parameter and arxiv_id format unexplained, so an agent may not know all consequences of invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It lists valid values for the 'state' parameter, which is helpful, but it provides no explanation for 'arxiv_id' beyond the schema title and no semantics for the 'reason' parameter. Parameter understanding remains incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Set the triage state for a paper.' It clearly identifies the action and the object, and the valid state list adds meaning. However, it does not explicitly differentiate this tool from siblings, though the operation is fairly unique among them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when an agent needs to assign or change a paper's triage state, and the list of valid states gives useful context. It provides no explicit guidance on when not to use it or which sibling tool to prefer for related operations like collections or signals.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.1.0- First observed
add_signal - First observed
add_to_collection - First observed
batch_add_signals - First observed
browse_recent - First observed
create_profile - First observed
create_watch - First observed
enrich_paper - First observed
find_related_papers - First observed
get_content_variant - First observed
get_paper - First observed
search_papers - First observed
suggest_signals - First observed
triage_paper
TDQS
Scored across 13 tools
Most tools target distinct actions: search, browse, related-papers, get-paper, triage, watch, signals, enrichment, and content retrieval are clearly separated. The only mild ambiguity is between search_papers and browse_recent since both return ranked paper lists with optional profile_slug, but their descriptions make the query-vs-recency distinction clear.
The majority follow a verb_noun pattern such as get_paper, create_profile, add_signal, and enrich_paper. Minor deviations like browse_recent (verb + adjective) and batch_add_signals break the pattern slightly, but the naming remains predictable and readable.
13 tools is a well-scoped count for an arXiv discovery server. Each tool contributes a meaningful capability without redundancy, and the set is neither too thin nor overloaded.
The discovery side is well covered with search, browse, get, related papers, content variants, and enrichment. However, the personalization surface has notable lifecycle gaps: you can create profiles, create watches, add signals, and add to collections, but there are no remove, delete, or list operations for these resources.
Maintenance
Related MCP Connectors
arXiv MCP — preprint server search (free, no auth)
Academic research MCP server for paper search, citation checks, graphs, and deep research.
MCP server for Altmetric APIs - track research attention across news, policy, social media, and more
MCP server for VC pitch-deck scoring, thesis-fit matching, and deal-flow management.
Related MCP Servers
- FlicenseAqualityDmaintenanceAn MCP server that enables intelligent searching, filtering, and exporting of Software Engineering papers on arXiv with tools for querying by keywords, authors, analyzing trends, and finding related research.76-
- AlicenseAqualityCmaintenanceMCP server for searching and retrieving arXiv papers with full-text PDF extraction.52MIT
- AlicenseNot gradedqualityDmaintenanceA lightweight MCP server for querying arXiv with semantic and keyword search, supporting natural language queries and structured filters.2MIT
- FlicenseNot gradedqualityCmaintenanceA local, rule-based MCP server for searching and analyzing academic papers from arXiv. Enables paper search, ranking, smart summarization, keyword extraction, and citation generation without API keys or LLM calls.-