Paper Search MCP
This server provides an MCP-based academic paper search and download tool covering 14 platforms (arXiv, Web of Science, PubMed, Google Scholar, Crossref, Scopus, Springer, etc.).
Search papers across many platforms with filters such as year, author, journal, category, sort order, and max results
Perform platform-specific searches: arXiv, Web of Science, PubMed, bioRxiv, medRxiv, Semantic Scholar, IACR, Google Scholar, ScienceDirect, Springer, Scopus, Crossref, and Sci-Hub
Download paper PDFs from supported platforms (arXiv, bioRxiv, medRxiv, Semantic Scholar, IACR, Sci-Hub, Springer, Wiley) with optional save path
Retrieve paper metadata by DOI across platforms
Discover public publisher PDF candidates for a DOI, with optional bounded PDF verification
Get citation data, references, citing records, and related records via Semantic Scholar or Web of Science Expanded API
Check Sci-Hub mirror health status
Check platform status, capability, and API-key configuration locally
Search directly or use
platform: "all"to randomly select an efficient platformSupports optional advanced options: open-access filters, subject areas, document types, affiliation, fields of study, publication types, and detailed IACR records
Enables searching and downloading physics and computer science preprints, with support for full-text PDF downloads and metadata retrieval.
Provides access to Web of Science database with advanced search capabilities including multi-topic queries, field tag support, year range filtering, and citation sorting.
Allows paper retrieval and information lookup using DOI identifiers across multiple academic platforms with validation and security checks.
Enables searching ScienceDirect full-text database and Scopus citation database with filters for open access content, authors, and document types.
Enables comprehensive academic search across publishers with citation data, automatic filtering for peer-reviewed papers, and year range queries.
Enables searching and downloading open access papers from Springer Nature's OpenAccess API, bioRxiv, and medRxiv preprint servers.
Allows searching biomedical literature from PubMed/MEDLINE database with filters for authors, journals, publication types, and date ranges.
Provides access to the largest citation database with advanced filtering by affiliation, document type, and comprehensive citation metrics.
Provides AI-powered semantic search with citation networks, research field filtering, and direct access to open access PDFs.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Paper Search MCPfind recent papers about large language model alignment"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Paper Search MCP (Node.js)
English|δΈζ
A Node.js Model Context Protocol (MCP) server for searching and downloading academic papers from multiple sources, including arXiv, Web of Science, PubMed, Google Scholar, Sci-Hub, ScienceDirect, Springer, Wiley, Scopus, Crossref, and 14 academic platforms in total.
Related MCP server: Paper Search MCP
π Sponsors
This project is sponsored by ScrapingAnt, a web scraping service for accessing public web data.
Offer for project users: Use code ENTHUSIAST_50 for 50% off the first month of the Enthusiast plan. The discount applies to the first month only.
β¨ Key Features
π 14 Academic Platforms: arXiv, Web of Science, PubMed, Google Scholar, bioRxiv, medRxiv, Semantic Scholar, IACR ePrint, Sci-Hub, ScienceDirect, Springer Nature, Wiley, Scopus, Crossref
π§ WoS Starter + Expanded: Starter v2 by default; Expanded SR/FR and citation relationships only when explicitly selected
π Public-page access discovery: optional ScrapingAnt Extended fetching for DOI publisher pages; never used for Clarivate API or institutional login pages
π MCP Protocol Integration: Seamless integration with Claude Desktop and other AI assistants
π Unified Data Model: Standardized paper format across all platforms
β‘ High-Performance Search: Concurrent search with intelligent rate limiting
π‘οΈ Security First: DOI validation, query sanitization, injection prevention, sensitive data masking
π Type Safety: Complete TypeScript support with extended interfaces
π― Academic Papers First: Smart filtering prioritizing academic papers over books
π Smart Error Handling: Unified ErrorHandler with retry logic and platform fallback
π Supported Platforms
Platform | Search | Download | Full Text | Citations | API Key | Special Features |
Crossref | β | β | β | β | β | Default search, extensive metadata coverage |
arXiv | β | β | β | β | β | Physics/CS preprints |
Web of Science | β | β | β | β | β Required | Starter v2 default; Expanded SR/FR and relations opt-in |
PubMed | β | β | β | β | π‘ Optional | Biomedical literature |
Google Scholar | β | β | β | β | β | Direct parser or optional ScrapingAnt General endpoint |
bioRxiv | β | β | β | β | β | Biology preprints |
medRxiv | β | β | β | β | β | Medical preprints |
Semantic Scholar | β | β | β | β | π‘ Optional | AI semantic search |
IACR ePrint | β | β | β | β | β | Cryptography papers |
Sci-Hub | Opt-in | Opt-in | β | β | β | Controlled DOI-only HTML adapter; disabled by default |
ScienceDirect | β | β | β | β | β Required | Elsevier's full-text database |
Springer Nature | β | β * | β | β | β Required | Dual API: Meta v2 & OpenAccess |
Wiley | β | β | β | β | β Required | TDM API: DOI-based PDF download only |
Scopus | β | β | β | β | β Required | Largest citation database |
β Supported | β Not supported | π‘ Optional | β * Open Access only
Note: Wiley TDM API does not support keyword search. Use
search_crossrefto find Wiley articles, then usedownload_paperwithplatform="wiley"to download PDFs by DOI.
βοΈ Compliance & Ethical Use (Sci-Hub / Google Scholar)
This project includes integrations that may have legal, contractual (ToS), and ethical constraints. You are responsible for ensuring your usage complies with applicable laws, institutional policies, and thirdβparty terms.
Sci-Hub: Disabled by default and limited to an unstable DOI/mirror adapter. It does not grant access rights; enable it only for content you are legally authorized to access.
Google Scholar/ScrapingAnt: Automated fetching may trigger blocking or contractual restrictions. ScrapingAnt is used only for public Scholar or publisher pages when configured, never for WoS login, SSO, MFA, or institutional subscription pages.
π Quick Start
System Requirements
Node.js 20.18.1+ (Node.js 21 is excluded by dependency engine constraints)
npm or yarn
Installation
# Clone repository
git clone https://github.com/your-username/paper-search-mcp-nodejs.git
cd paper-search-mcp-nodejs
# Install dependencies
npm install
# Copy environment template
cp .env.example .envConfiguration
Get Web of Science API Key
Register and apply for Web of Science API access
Add API key to
.envfile
Get PubMed API Key (Optional)
Without API key: Free usage, 3 requests/second limit
With API key: 10 requests/second, more stable service
Get key: See NCBI API Keys
Configure Environment Variables
# Edit .env file # Web of Science defaults to Starter v2 and uses WOS_API_KEY. WOS_API_KEY=your_web_of_science_api_key # Optional Expanded product key. WOS_EXPANDED_API_KEY=your_expanded_key WOS_STARTER_VERSION=v2 WOS_STARTER_RPS=1 WOS_STARTER_DAILY_LIMIT=50 WOS_EXPANDED_RPS=2 # Full Record records/day; 0 means unlimited local accounting. WOS_EXPANDED_FULL_RECORD_BUDGET=0 WOS_EXPANDED_BASE_URL=https://api.clarivate.com/api/wos # Optional public-page HTML fallback; never a WoS/API proxy. # A key alone does not authorize paid retrieval; browser and residential # escalation are separate opt-ins. Invalid/missing paid configuration keeps Direct available. # Without residential authorization Publisher/Scholar default to 50 credits/operation and 10/request. # With SCRAPINGANT_ALLOW_RESIDENTIAL=true their defaults become 500/125; # explicit limits always win and a residential request still requires the residential ceiling. SCRAPINGANT_API_KEY= SCRAPINGANT_ENABLED=false SCRAPINGANT_ALLOW_BROWSER_ESCALATION=false SCRAPINGANT_ALLOW_RESIDENTIAL=false SCRAPINGANT_MAX_CREDITS_PER_OPERATION=50 SCRAPINGANT_MAX_CREDITS_PER_REQUEST=10 SCRAPINGANT_MAX_CONCURRENCY=1 SCRAPINGANT_PROXY_TYPE=datacenter # Controlled Sci-Hub adapter; disabled unless explicitly enabled. # Mirror addresses are discovered from the Sci-Hub and ooopn directory pages. # SCIHUB_MIRRORS is optional and accepts comma-separated supplemental URLs. SCIHUB_ENABLED=false SCIHUB_FETCH_MODE=fallback SCIHUB_MIRRORS= SCIHUB_HEALTHCHECK_CONCURRENCY=3 # PubMed API key (optional, recommended for better performance) PUBMED_API_KEY=your_ncbi_api_key_here # Semantic Scholar API key (optional, increases rate limits) SEMANTIC_SCHOLAR_API_KEY=your_semantic_scholar_api_key # Elsevier API key: ScienceDirect Search v2 and Scopus details ELSEVIER_API_KEY=your_elsevier_api_key # Optional dedicated key for Scopus Search API; falls back to ELSEVIER_API_KEY SCOPUS_SEARCH_API_KEY= # Springer Nature API keys (required for Springer) SPRINGER_API_KEY=your_springer_api_key # For Metadata API v2 # Optional: Separate key for OpenAccess API (if different from main key) SPRINGER_OPENACCESS_API_KEY=your_openaccess_api_key # Wiley TDM token (required for Wiley) WILEY_TDM_TOKEN=your_wiley_tdm_token
Build and Run
Method 1: NPX (Recommended for MCP)
# Direct run with npx (most common MCP deployment)
npx -y paper-search-mcp-nodejs
# Or install globally
npm install -g paper-search-mcp-nodejs
paper-search-mcpMethod 2: Local Development
# Build TypeScript code
npm run build
# Start server
npm start
# Or run in development mode
npm run devMCP Server Configuration
Add the following configuration to your Claude Desktop config file:
macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
Windows: %APPDATA%\Claude\claude_desktop_config.json
Minimal NPX Configuration (Start Here)
Use this minimal configuration to start immediately with public sources such as Crossref and arXiv. No API key or optional environment variable is required; add platform-specific keys only when you need them.
{
"mcpServers": {
"paper-search-nodejs": {
"command": "npx",
"args": ["-y", "paper-search-mcp-nodejs"]
}
}
}Complete NPX Configuration (Advanced/Optional)
Use the complete configuration below only when you need keyed platforms, optional fallbacks, or custom limits. Every value in env must be a string; replace placeholders with values from your own environment, and never commit real credentials.
{
"mcpServers": {
"paper-search-nodejs": {
"command": "npx",
"args": ["-y", "paper-search-mcp-nodejs"],
"env": {
"NODE_ENV": "production",
"LOG_LEVEL": "info",
"WOS_API_KEY": "your_web_of_science_api_key",
"WOS_STARTER_VERSION": "v2",
"WOS_STARTER_RPS": "1",
"WOS_STARTER_DAILY_LIMIT": "50",
"WOS_EXPANDED_API_KEY": "",
"WOS_EXPANDED_RPS": "2",
"WOS_EXPANDED_FULL_RECORD_BUDGET": "0",
"WOS_EXPANDED_BASE_URL": "https://api.clarivate.com/api/wos",
"PUBMED_API_KEY": "",
"SEMANTIC_SCHOLAR_API_KEY": "",
"ELSEVIER_API_KEY": "",
"SCOPUS_SEARCH_API_KEY": "",
"SPRINGER_API_KEY": "",
"SPRINGER_OPENACCESS_API_KEY": "",
"WILEY_TDM_TOKEN": "",
"CROSSREF_MAILTO": "you@example.com",
"SCRAPINGANT_API_KEY": "",
"SCRAPINGANT_ENABLED": "false",
"SCRAPINGANT_ALLOW_BROWSER_ESCALATION": "false",
"SCRAPINGANT_ALLOW_RESIDENTIAL": "false",
"SCRAPINGANT_MAX_CREDITS_PER_OPERATION": "50",
"SCRAPINGANT_MAX_CREDITS_PER_REQUEST": "10",
"SCRAPINGANT_MAX_CONCURRENCY": "1",
"SCRAPINGANT_PROXY_TYPE": "datacenter",
"SCHOLAR_PROXY": "",
"SCIHUB_ENABLED": "false",
"SCIHUB_FETCH_MODE": "fallback",
"SCIHUB_MIRRORS": "",
"SCIHUB_HEALTHCHECK_CONCURRENCY": "3",
"DEFAULT_DOWNLOAD_PATH": "./downloads",
"MAX_FILE_SIZE_MB": "100",
"RATE_LIMIT_REQUESTS_PER_MINUTE": "60",
"RATE_LIMIT_BURST": "10"
}
}
}
}SCHOLAR_PROXY is an optional explicit override for Scholar direct transport. When it is empty, Scholar uses the standard local proxy aliases (HTTPS_PROXY/HTTP_PROXY/ALL_PROXY, including lowercase forms) when configured; when set, it takes precedence and is parsed as a full HTTP(S)/SOCKS proxy URL. It does not select the ScrapingAnt paid fallback. The retrieval benchmark does not enable or replace it. WOS_EXPANDED_API_KEY, SCRAPINGANT_API_KEY, and the other optional keys may remain empty.
Local Installation Configuration
For a local build, keep the complete env object above and change only the server command:
{
"command": "node",
"args": ["/path/to/paper-search-mcp-nodejs/dist/server.js"]
}π οΈ MCP Tools
search_papers
Search academic papers across multiple platforms
// Random platform selection (default behavior)
search_papers({
query: "machine learning",
platform: "all", // Randomly selects one platform for efficiency
maxResults: 10,
year: "2023",
sortBy: "date"
})
// Search specific platform
search_papers({
query: "quantum computing",
platform: "webofscience", // Target specific platform
maxResults: 5
})Platform Selection Behavior:
platform: "crossref"(default) - Free API with extensive scholarly metadata coverageplatform: "all"- Randomly selects one platform for efficient, focused resultsSpecific platform - Searches only that platform
Available platforms:
crossref,arxiv,webofscience/wos,pubmed,biorxiv,medrxiv,semantic,iacr,googlescholar/scholar,scihub,sciencedirect,springer,scopusNote:
wileyonly supports PDF download by DOI, not keyword search
search_crossref
Search academic papers from Crossref database (default search platform)
search_crossref({
query: "machine learning",
maxResults: 10,
year: "2023",
author: "Smith",
sortBy: "relevance", // or "date", "citations"
sortOrder: "desc"
})search_arxiv
Search arXiv preprints specifically
search_arxiv({
query: "transformer neural networks",
maxResults: 10,
category: "cs.AI",
author: "Vaswani",
year: "2023",
sortBy: "date", // relevance, date, citations
sortOrder: "desc" // asc, desc
})search_webofscience
Search Web of Science database specifically
search_webofscience({
query: "CRISPR gene editing",
maxResults: 5,
year: "2022",
journal: "Nature",
apiProduct: "expanded", // omit for Starter v2
recordView: "short", // expanded only; "full" is opt-in
discoverAccess: true, // optional publisher public-page discovery
discoverAccessMaxItems: 5 // explicit bound; default is 5, range 1-100
})
get_webofscience_related_records({
uid: "WOS:000000000000001",
relation: "citing", // references, citing, or related
maxResults: 50
})search_pubmed
Search PubMed/MEDLINE biomedical literature database
search_pubmed({
query: "COVID-19 vaccine efficacy",
maxResults: 20,
year: "2023",
author: "Smith",
journal: "New England Journal of Medicine",
publicationType: ["Journal Article", "Clinical Trial"],
sortBy: "date" // relevance, date
})search_google_scholar
Search Google Scholar academic database
search_google_scholar({
query: "machine learning",
maxResults: 10,
yearLow: 2020,
yearHigh: 2023,
author: "Bengio"
})search_biorxiv / search_medrxiv
Search biology and medical preprints
search_biorxiv({
query: "CRISPR",
maxResults: 15,
days: 30,
category: "genomics" // neuroscience, genomics, etc.
})
search_medrxiv({
query: "COVID-19",
maxResults: 10,
days: 30,
category: "infectious_diseases"
})search_semantic_scholar
Search Semantic Scholar AI semantic database
search_semantic_scholar({
query: "deep learning",
maxResults: 10,
fieldsOfStudy: ["Computer Science"],
year: "2023"
})search_iacr
Search IACR ePrint cryptography archive
search_iacr({
query: "zero knowledge proof",
maxResults: 5,
fetchDetails: true
})search_scihub
Controlled, opt-in Sci-Hub DOI lookup/download; disabled by default and not an official API
// Requires SCIHUB_ENABLED=true. DOI/doi.org inputs only.
// Mirrors are discovered from https://sci-hub.mobi/en/mirrors and
// https://www.ooopn.com/tool/scihub/; SCIHUB_MIRRORS adds optional comma-separated URLs.
search_scihub({
doiOrUrl: "10.1038/nature12373",
downloadPdf: true,
savePath: "./downloads"
})search_sciencedirect
Search Elsevier ScienceDirect database
search_sciencedirect({
query: "artificial intelligence",
maxResults: 10,
year: "2023",
author: "Smith",
openAccess: true // Filter for open access articles
})search_springer
Search Springer Nature database (Metadata API v2 or OpenAccess API)
search_springer({
query: "machine learning",
maxResults: 10,
year: "2023",
openAccess: true, // Use OpenAccess API for downloadable PDFs
type: "Journal" // Filter: Journal, Book, or Chapter
})search_scopus
Search Scopus citation database
search_scopus({
query: "renewable energy",
maxResults: 10,
year: "2023",
affiliation: "MIT",
documentType: "ar" // ar=article, cp=conference, re=review
})Scopus search requests COMPLETE by default. If Elsevier explicitly denies the COMPLETE view because the key lacks that entitlement, the search enters one bounded STANDARD fallback strategy; any transient retries remain subject to the existing retry policy. The fallback omits the field override so the API can return its standard field set. Other authentication, query, rate-limit, network, and server errors are not converted into a view fallback. STANDARD may contain less enriched metadata (for example, full author, abstract, keyword, affiliation, and funding fields); a valid Scopus API key is still required. See Scopus Search API views for the fields available in each view.
check_scihub_mirrors
Check health status of Sci-Hub mirror sites
check_scihub_mirrors({
forceCheck: true // Force fresh health check
})download_paper
Download paper PDF files
download_paper({
paperId: "2106.12345", // or DOI for Sci-Hub
platform: "arxiv", // or "scihub" for Sci-Hub downloads
savePath: "./downloads"
})get_paper_by_doi
Get paper information by DOI
get_paper_by_doi({
doi: "10.1038/s41586-023-12345-6",
platform: "all"
})discover_paper_access
Discover one public publisher PDF candidate by DOI. This does not claim open-license status or complete download success; PDF verification is opt-in and bounded.
discover_paper_access({
doi: "https://doi.org/10.1038/s41586-023-12345-6",
verifyPdf: false
})get_platform_status
Check local platform capability and API-key status. This is a local diagnostic and does not validate ScrapingAnt account quota or make a paid request.
get_platform_status({})Public access and cost boundaries
discover_paper_access accepts only a DOI. It returns a bounded access state such as oa_candidate, pdf_verified, not_found, restricted, failed, or skipped; a candidate is not an open-license or complete-download claim. verifyPdf is opt-in and performs only a bounded PDF prefix check. Candidate order, source provenance, HTTP/API status, fallback attempts, and known local cost are separate evidence fields.
Paid retrieval is Direct-first and finite. Empty/parse-failed public pages may consume an explicitly authorized fallback attempt; a known permission, unsafe target, resource limit, cancellation, deadline, unknown price, or closed ledger stops paid work. Missing/invalid post-dispatch billing is recorded as unknown and consumes its local estimate, but does not independently stop a bounded retry/fallback chain; the final reported credits remain unknown until reconciled. When browser escalation is explicitly enabled, Scholar production tries ScrapingAnt browser:datacenter before the remaining paid combinations; browser is one bounded dispatch and has no retry. Returned diagnostics never include cookies, authorization, queries, raw HTML, or sensitive URLs. Public cookies are not read from configuration; Scholar session cookies, when obtained, stay on the exact Scholar HTTPS origin and are never sent to ScrapingAnt.
Use SCRAPINGANT_ENABLED=true only after obtaining deployment authorization. Scholar's browser-first fallback requires SCRAPINGANT_ALLOW_BROWSER_ESCALATION=true; SCRAPINGANT_ALLOW_BROWSER_ESCALATION and SCRAPINGANT_ALLOW_RESIDENTIAL are independent controls. Restarting with either flag disabled rolls back the corresponding combinations; it does not erase already observed local costs. A run can also be disabled by leaving the key/paid flag off. No setting synchronizes provider quota or clears an in-flight ledger.
Offline benchmark
The fixed benchmark corpus and strategy schedule are validated without network access. A complete 360-cell live run is a separately authorized external qualification, not the completion gate for each engineering fix:
# Accounting/report simulation (no production retrieval workflow)
npm run --silent benchmark:offline -- --json
# Reviewed offline production workflow beneath fixed raw fixtures (virtual clock)
npm run --silent benchmark:offline -- --workflow --json
# Optional: write encoded run-id .json and .md artifacts without overwriting existing files.
npm run --silent benchmark:offline -- --workflow --output-dir ./benchmark-artifactsThe default command validates the frozen 20 DOI/10 Scholar query corpus with an injected response-only accounting simulator. Add --workflow to run the reviewed createProductionBenchmarkCellExecutor instead: it invokes the real Publisher/Scholar business entries, providers, parsing, session, fallback, scheduler, billing bridge and PDF-prefix paths beneath independent fixed raw fixtures, using a virtual clock while preserving the production pacing rules. Both modes are explicitly offline-only; neither initializes a live provider or uses credentials for retrieval, and neither can produce live_passed. The workflow report must be identified separately from simulator totals. Live evaluation is exposed only through the separately named command below; it requires explicit --authorize-live, an exclusive explicit --run-id, a complete preflight, and the fixed safety ceiling of 15,000 credits, 1,500 HTTP dispatches, and 7,200,000ms. A failed preflight writes a privacy-safe blocked report with zero dispatches; it never shrinks the frozen matrix or bypasses configuration. The live executor uses real production transports and is not the offline fixture executor. The planned schedule is 60 production cells plus 300 independent comparison cells; unexecuted live cells remain not_run and keep their denominator. Artifact paths are exclusive so a report cannot silently overwrite an earlier run.
# Explicitly authorized live campaign; preflight blocks without complete paid/browser/residential configuration.
npm run --silent benchmark:live -- --authorize-live --authorize-scholar-proxy --run-id live-YYYYMMDD-01 --output-dir ./live-benchmark-artifacts --jsonDo not pass --authorize-live unless the campaign, provider billing, target scope, and live-capable harness have been reviewed. Pass --authorize-scholar-proxy only when the explicit SCHOLAR_PROXY endpoint has separately passed TLS/ownership review; ambient proxy aliases alone remain blocked. The command never resumes an old run, appends budget, or changes offlineOnly fixtures.
π Data Model
All platform paper data is converted to a unified format:
interface Paper {
paperId: string; // Unique identifier
title: string; // Paper title
authors: string[]; // Author list
abstract: string; // Abstract
doi: string; // DOI
publishedDate: Date; // Publication date
pdfUrl: string; // PDF link
url: string; // Paper page URL
source: string; // Source platform
citationCount?: number; // Citation count
journal?: string; // Journal name
year?: number; // Publication year
categories?: string[]; // Subject categories
keywords?: string[]; // Keywords
// ... more fields
}π§ Development
Project Structure
src/
βββ models/
β βββ Paper.ts # Paper data model
βββ platforms/
β βββ PaperSource.ts # Abstract base class
β βββ ArxivSearcher.ts # arXiv searcher
β βββ WebOfScienceSearcher.ts # Web of Science searcher
β βββ PubMedSearcher.ts # PubMed searcher
β βββ GoogleScholarSearcher.ts # Google Scholar searcher
β βββ BioRxivSearcher.ts # bioRxiv/medRxiv searcher
β βββ SemanticScholarSearcher.ts # Semantic Scholar searcher
β βββ IACRSearcher.ts # IACR ePrint searcher
β βββ SciHubSearcher.ts # Sci-Hub searcher with mirror management
β βββ ScienceDirectSearcher.ts # ScienceDirect (Elsevier) searcher
β βββ SpringerSearcher.ts # Springer Nature searcher (Meta v2 & OpenAccess APIs)
β βββ WileySearcher.ts # Wiley TDM API (DOI-based PDF download only)
β βββ ScopusSearcher.ts # Scopus citation database searcher
β βββ CrossrefSearcher.ts # Crossref API searcher (default platform)
βββ mcp/
β βββ tools.ts # MCP tool definitions
β βββ schemas.ts # Zod schemas for tool arguments
β βββ handleToolCall.ts # Tool call dispatcher
β βββ searchers.ts # Searcher initialization
βββ utils/
β βββ SecurityUtils.ts # DOI validation, query sanitization, injection prevention
β βββ PublicNetwork.ts # Public-target DNS/redirect and SSRF checks
β βββ ConcurrencyLimiter.ts # Dependency-free bounded concurrency
β βββ ErrorHandler.ts # Unified error handling with retry logic
β βββ RateLimiter.ts # Token bucket rate limiting
β βββ QuotaManager.ts # Daily quota tracking
β βββ RequestCache.ts # LRU caching for requests
β βββ PDFExtractor.ts # PDF text extraction
β βββ Logger.ts # Debug logging
βββ config/
β βββ constants.ts # Timeouts, endpoints, limits
βββ services/
β βββ CitationService.ts # Citation fetching service
β βββ WebOfScienceParser.ts # Starter/Expanded response parsers
β βββ WebOfScienceRequestService.ts # WoS retry, rate, quota, and status
β βββ PublicHttpClient.ts # Redirect-checked public HTTP
β βββ ScrapingAntFetcher.ts # General/Extended HTML API wrapper
β βββ PublicAccessDiscovery.ts # DOI publisher-page PDF discovery
βββ server.ts # MCP server main fileAdding New Platforms
Create new searcher class extending
PaperSourceImplement required abstract methods
Register new searcher in
searchers.tsAdd corresponding MCP tool in
tools.ts
Security Best Practices
All DOIs are validated before use in URLs
Query parameters are escaped to prevent injection
API keys are masked in all log output
Request timeouts prevent hanging connections
Query complexity limits prevent DoS attacks
Rate limiting and quota management prevent API abuse
Caching reduces external API calls
Testing
# Run tests
npm test
# Run linting
npm run lint
# Code formatting
npm run formatTest Coverage:
Includes WoS HTTP contracts, ScrapingAnt status/credits, public-target SSRF checks, access discovery, Sci-Hub fallback/PDF validation, and MCP schemas.
Platform searchers covered by unit and contract tests
Security utilities (DOI validation, query sanitization)
ErrorHandler (error classification, retry logic)
Rate limiting integration, QuotaManager, RequestCache
Test Suite | Coverage |
Platform Searchers | β |
SecurityUtils | β |
ErrorHandler | β |
RateLimiter & Integration | β |
QuotaManager | β |
RequestCache | β |
π Platform-Specific Features
Springer Nature Dual API System
Springer Nature provides two APIs:
Metadata API v2 (Main API)
Endpoint:
https://api.springernature.com/meta/v2/jsonSearches all Springer content (subscription + open access)
Requires API key from https://dev.springernature.com/
OpenAccess API (Optional)
Endpoint:
https://api.springernature.com/openaccess/jsonOnly searches open access content
May require separate API key or special permissions
Better for finding downloadable PDFs
// Search all Springer content
search_springer({
query: "machine learning",
maxResults: 10
})
// Search only open access papers
search_springer({
query: "COVID-19",
openAccess: true, // Uses OpenAccess API if available
maxResults: 5
})Web of Science Advanced Search
π― WoS Starter + Expanded: Starter API v2 is the default and v1 remains an explicit compatibility choice. Expanded must be selected explicitly.
API Version and product configuration:
# Starter version (default: v2; fixed for the process)
WOS_STARTER_VERSION=v2
# WOS_STARTER_VERSION=v1
# Web of Science defaults to Starter v2.
WOS_API_KEY=...
# Optional Expanded product key.
WOS_EXPANDED_API_KEY=...Starter requests use the documented /documents endpoints, page at most 50 records, and preserve unknown citation counts as null. Expanded uses its separate /api/wos contract, defaults to Short Record, and supports Full Record, references, citing, and related-record operations only when requested. The default is the current Swagger server https://api.clarivate.com/api/wos; older wos-api.clarivate.com guidance is not used.
// Multi-topic search
search_webofscience({
query: 'oriented structure',
year: '2023-2025',
sortBy: 'date',
sortOrder: 'desc',
maxResults: 10
})
// Year range filtering
search_webofscience({
query: 'machine learning',
year: '2020-2024', // Supports range format
sortBy: 'citations',
sortOrder: 'desc'
})
// Advanced query with filters
search_webofscience({
query: 'blockchain',
author: 'zhang',
journal: 'Nature',
year: '2023',
sortBy: 'date',
sortOrder: 'desc'
})
// Traditional WOS query syntax with field tags
search_webofscience({
query: 'TS="machine learning" AND PY=2023 AND DT="Article"',
maxResults: 20
})Supported Search Options:
query: Search terms (supports multi-topic)year: Single year "2023" or range "2020-2023"author: Author name filteringjournal: Journal/source filteringsortBy: Supported sort field (date,citations,relevance)sortOrder: Sort direction (asc,desc)maxResults: Maximum results (1-100; Starter fetches 50 per page)apiProduct:starter(default) or explicitexpandedrecordView: Expandedshort(default) or explicitfulldiscoverAccess: Optional publisher public-page discoverydiscoverAccessMaxItems: Discovery bound (1-100; explicit value, deployment default, then 5)
Supported WOS Field Tags (18 total):
Tag | Description | Tag | Description |
| Topic (title, abstract, keywords) |
| Title |
| Author |
| Author Identifier |
| Source/Journal |
| ISSN/ISBN |
| Publication Year |
| Final Publication Year |
| DOI |
| Date of Publication |
| Volume |
| Page |
| Issue |
| Document Type |
| PubMed ID |
| Accession Number |
| Organization |
| Source URL |
Example with Field Tags:
// Search by PMID
search_webofscience({ query: 'PMID=12345678' })
// Search by DOI
search_webofscience({ query: 'DO="10.1038/nature12373"' })
// Filter by document type
search_webofscience({ query: 'TS="CRISPR" AND DT="Review"' })
// Search specific volume/issue
search_webofscience({ query: 'SO="Nature" AND VL=580 AND CS=7805' })π§ Debugging WOS Issues:
# Enable debug logging
export NODE_ENV=development
# In CI, logDebug is enabled automatically when CI=trueGoogle Scholar Features
HTML search adapter: Uses the Scholar web page, not an official public API
Metadata and citations: Parses titles, authors, abstracts, publication years, and βCited byβ counts
Bounded retrieval: Requests at most 20 results across at most 10 pages with deduplication, dispatch-time pacing, adaptive delays, and bounded retry handling
Provider-neutral transport: Direct Scholar traffic keeps its isolated same-origin HTTPS session; source waits do not occupy the global transport slot, while ScrapingAnt is a bounded fallback for eligible transport/server failures and empty/parse-failed pages
Session privacy: Scholar cookies are created in memory for the exact Scholar origin and are never forwarded to ScrapingAnt or returned in diagnostics
No full-text authorization: PDF/library links remain publisher or institutional links
Google Scholar access: Googleβs official help says automated software should respect
robots.txtand that bulk access is not provided. Direct requests use the standard local proxy aliases (HTTPS_PROXY/HTTP_PROXY/ALL_PROXY, including lowercase forms) when configured. An explicitSCHOLAR_PROXYoverrides those aliases and is parsed as a complete HTTP(S)/SOCKS proxy URL. With the default transport, the configured ScrapingAnt fetcher is used only as a bounded backup for eligible transport failures, upstream 5xx responses, or no usable/parseable resultsβnot to bypass permission, 403/429, or CAPTCHA responses:# Optional explicit HTTP/HTTPS proxy override SCHOLAR_PROXY=http://user:pass@host:port # Optional explicit TLS-to-proxy endpoint SCHOLAR_PROXY=https://user:pass@host:port # Optional SOCKS proxy SCHOLAR_PROXY=socks://host:portRequired packages are loaded lazily (
http-proxy-agent,https-proxy-agent,socks-proxy-agent) β install the one matching your proxy type. ScrapingAnt Proxy mode is not used as a transparent substitute:SCRAPINGANT_PROXY_TYPEis only the provider proxy ceiling, and the legacySCHOLAR_PROXYpath remains independent and outside benchmark acceptance.
Semantic Scholar Features
AI-Powered Search: Semantic understanding of queries
Citation Networks: Paper relationships and influence metrics
Open Access PDFs: Direct links to freely available papers
Research Fields: Filter by specific academic disciplines
ScrapingAnt Public Fetch Layer
ScrapingAnt is an optional, paid public-page fallback: a key alone does not enable dispatch. Set SCRAPINGANT_ENABLED=true to opt in; browser and residential escalation are independent authorizations. The current Scholar browser:datacenter-first order is provisional and requires real provider capability validation; otherwise it retains static-first behavior. Without residential authorization, Publisher/Scholar defaults are 50 credits per operation and 10 credits per request; with SCRAPINGANT_ALLOW_RESIDENTIAL=true, their defaults become 500/125. Explicit valid limits always win, but residential dispatch still requires SCRAPINGANT_PROXY_TYPE=residential; a datacenter ceiling never sends residential traffic. Google Scholar uses the generic /v2/general HTML endpoint (not a dedicated Scholar API), while WoS DOI access discovery and Sci-Hub fallback use /v2/extended. The layer is never a Clarivate/WoS API proxy. DOI discovery rejects known login/SSO/Clarivate targets before dispatch where the local redirect chain is visible. Local DNS checks cannot prove the remote proxy's own redirect destination, so a discovered link is not an OA, authorization, or copyright determination. Actual usage is taken from response credit headers; local budgets are not provider billing balances. A known permission/security/resource/cancellation/deadline failure does not trigger paid fallback, and completed strategy scopes retain only finite cache metadata rather than raw provider documents. Persistent Scholar blocking remains an external provider limitation.
Public-paper and Markdown tools
download_public_paperis a separate, strict MCP tool forpublisher,googlescholar, andscihub. Publisher/Sci-Hub use a DOI inpaperId; Scholar uses a short-lived reference published by the same MCP connection's search result.get_paper_markdownis explicit opt-in only. It makes at most one static/datacenter/v2/markdownrequest and returns bounded, untrusted Markdown; Markdown is never used for Paper fields, PDF candidate extraction, or implicit search/download calls.Scholar references are in-memory only, capped at 256 entries per handler, TTL 300 seconds, non-sliding, and invalidated on handler disposal/restart. Missing, expired, or ambiguous references do not guess a URL or make a provider request.
Set
SCRAPINGANT_ALLOW_AUTHORIZED_CORPUS=trueplus canonical tokens inSCRAPINGANT_AUTHORIZED_CORPUS_PLATFORMSonly when a deployment is authorized to attempt restricted pages. This does not forward cookies, Authorization headers, institutional credentials, or MCP-supplied URLs.Provider HTML/Markdown is not a PDF authority. The new downloader re-checks public targets, MIME,
%PDF-, size, cancellation, paths, symlinks, and atomic no-clobber publication. Existingdownload_paperkeeps its name, eight-platform schema, and behavior; no Proxy mode or AI Extractor is enabled.Offline tests and the default configuration do not send live or paid requests. Live canaries require separate approval, budget, samples, and stop conditions.
Sci-Hub Features
Opt-in only: Disabled by default; accepts DOI,
doi:anddoi.orgforms onlyControlled fallback: Direct mirror lookup first, then bounded ScrapingAnt Extended HTML fallback when enabled
Mirror discovery and health monitoring: Fetches mirror lists from
https://sci-hub.mobi/en/mirrorsandhttps://www.ooopn.com/tool/scihub/, merges optionalSCIHUB_MIRRORSsupplements, then performs max-three-concurrent direct checks with caching and single-flightExplicit states: Distinguishes not-found, blocked, markup-changed, unhealthy, and transport failures
Safe PDF handling: Public-target redirect checks, MIME/magic/size validation, temporary files, and atomic replacement
Compliance notice: Does not grant access rights or claim copyright/authorization status
π License
MIT License - see LICENSE file for details.
π€ Contributing
Contributions welcome!
Fork the project
Create feature branch (
git checkout -b feature/amazing-feature)Commit changes (
git commit -m 'Add amazing feature')Push to branch (
git push origin feature/amazing-feature)Open Pull Request
π Issue Reporting
If you encounter issues, please report them at GitHub Issues.
π Acknowledgments
Original paper-search-mcp for the foundation
MCP community for the protocol standards
β If this project helps you, please give it a star!
Available Tools
21 toolscheck_scihub_mirrorsC
Check the health status of all Sci-Hub mirror sites
| Name | Required | Description | Default |
|---|---|---|---|
| forceCheck | No | Force a fresh health check even if recent data exists |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool checks health status, implying a read-only operation, but doesn't specify what 'health status' entails (e.g., uptime, response time), whether it's cached or real-time, or any rate limits. For a tool with no annotations, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence: 'Check the health status of all Sci-Hub mirror sites.' It's front-loaded with the core purpose, has zero wasted words, and is appropriately sized for the tool's simplicity, making it highly efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete. It doesn't explain what the health check returns (e.g., a list of mirrors with statuses), how results are formatted, or any behavioral nuances. For a tool that likely outputs structured data, this leaves the agent guessing about the tool's full functionality.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the parameter 'forceCheck' well-documented in the schema. The description doesn't add any meaning beyond this, as it mentions no parameters. Given the high schema coverage, the baseline score of 3 is appropriate, as the description doesn't compensate but doesn't need to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check the health status of all Sci-Hub mirror sites.' It specifies the verb ('check') and resource ('Sci-Hub mirror sites'), distinguishing it from sibling tools focused on searching or downloading papers. However, it doesn't explicitly differentiate from 'get_platform_status,' which might be a related sibling, so it's not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, timing, or comparisons with siblings like 'get_platform_status' or 'search_scihub,' leaving the agent to infer usage based on the name alone. This lack of explicit context reduces its helpfulness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discover_paper_accessA
Discover a public publisher PDF candidate for one DOI without claiming open-license or complete download success
| Name | Required | Description | Default |
|---|---|---|---|
| doi | Yes | A DOI, doi: prefix, or doi.org URL; arbitrary HTTP URLs are rejected | |
| verifyPdf | No | Optionally perform a bounded direct PDF prefix probe |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It does disclose that results are candidates and that success is not guaranteed, which is valuable. However, it does not mention failure modes, output shape, or any rate/access limitations beyond that caveat, leaving some behavior unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is a single, tightly worded sentence. It front-loads the action and resource, then adds the key caveat without any filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description plus the 100% covered schema is largely sufficient. It explains the tool's purpose, scope, and key limitation. The only notable gap is the absence of any statement about what the tool returns, though the description's wording makes this reasonably inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both doi and verifyPdf clearly. The description echoes the DOI scope but adds no new meaning about parameter formats, defaults, or the bounded probe behavior described in the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific action: discovering a public publisher PDF candidate for a single DOI. It also adds the important qualifier that it does not claim open-license or complete download success, which distinguishes it from download_paper and other access-oriented siblings without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: when you need a PDF candidate for one DOI and are not expecting guaranteed open-license or download success. It implicitly separates this from search tools by scoping to a single DOI, though it does not explicitly name alternatives or state when not to use them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_paperC
Download PDF file of an academic paper
| Name | Required | Description | Default |
|---|---|---|---|
| paperId | Yes | Paper ID (e.g., arXiv ID, DOI for Sci-Hub) | |
| platform | Yes | Platform where the paper is from | |
| savePath | No | Directory to save the PDF file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the action ('Download PDF file') but lacks critical behavioral details: whether this requires authentication, potential rate limits, file size considerations, or what happens on failure (e.g., if the paper isn't found). For a download operation with no annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded with the core action and resource, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of downloading files across multiple platforms and no annotations or output schema, the description is incomplete. It doesn't address error handling, return values (e.g., success confirmation or file path), or platform-specific behaviors, which are crucial for a tool with 3 parameters and varied platforms.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds no additional meaning about parameters beyond implying they're needed for downloading. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Download') and resource ('PDF file of an academic paper'), making the purpose unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'get_paper_by_doi' or 'search_arxiv', which might also retrieve papers but perhaps in different formats or contexts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With many sibling tools for searching and retrieving papers, it's unclear if this is the primary download method or if others like 'get_paper_by_doi' serve similar purposes. No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_citationsA
Retrieve citation data (citation count, references, venue) for a paper by DOI using Semantic Scholar
| Name | Required | Description | Default |
|---|---|---|---|
| doi | Yes | DOI (Digital Object Identifier) | |
| forceRefresh | No | Bypass the cache and fetch fresh data (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full burden. It confirms it is a read operation (no side effects), but it does not disclose caching behavior, rate limits, or error handling. The mention of Semantic Scholar implies an external dependency but no caveats. This is a modest gap given the simple nature of the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is front-loaded with the core action and output, followed by the input method and source. No wasted words, length is appropriate for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 2 parameters, no output schema, and no annotations, the description provides sufficient context: it specifies what data is returned and that input is a DOI. The forceRefresh parameter is explained in the schema, and the description implies a single-paper lookup. Missing details like potential errors or rate limits are minor for this scope, making the description complete enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters (doi, forceRefresh) already documented in the schema. The description adds minimal extra meaning beyond confirming the DOI is for the paper and that data comes from Semantic Scholar. Since the schema fully covers parameters, the baseline of 3 is appropriate; the description does not enhance parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (retrieve citation data), the specific resource (a paper by DOI), and the data fields returned (citation count, references, venue). It also names the external service (Semantic Scholar), which distinguishes it from siblings like get_paper_by_doi that likely return different metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly conveys when to use it: when you have a DOI and need citation metrics. It does not explicitly mention when not to use it or alternatives, but the phrase 'citation data' differentiates it from general paper lookups. No exclusions are given, but the context is clear enough for an agent to choose it over search tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_paper_by_doiC
Retrieve paper information using DOI from available platforms
| Name | Required | Description | Default |
|---|---|---|---|
| doi | Yes | DOI (Digital Object Identifier) | |
| platform | No | Platform to search |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. 'Retrieve' implies a read-only lookup, but the behavior across platforms is not disclosed: whether it returns the first match, merges results, or what happens when the DOI is not found on a selected platform. No error handling, fallback, or aggregation behavior is described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise, front-loaded sentence with no filler. It communicates the core action and input immediately, though 'available platforms' could be more explicit. It earns a high score for efficiency, but not top marks due to slight vagueness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a 2-parameter tool with no output schema and no annotations, so the description must do more to orient the agent. It does not explain what information will be returned, how platform selection affects the lookup, or what distinguishes this from the many sibling search tools. The agent is left to guess expected behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already described in the schema. The description mentions DOI but adds no new semantic detail beyond the parameter descriptions, such as format expectations or how 'platform' affects results. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Retrieve paper information') and the key identifier ('DOI'), which clearly distinguishes it from keyword-based sibling tools like search_papers or search_google_scholar. The phrase 'available platforms' is somewhat vague, but the core purpose is understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. It does not mention that this tool is appropriate when a DOI is already known, nor does it suggest sibling tools for keyword or title-based lookup. The agent must infer usage from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_platform_statusC
Check the status and capabilities of available academic platforms
| Name | Required | Description | Default |
|---|---|---|---|
| validate | No | Whether to validate configured API keys by making a real request (may trigger rate limits). Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While 'Check' implies a read-only operation, it doesn't specify what 'status and capabilities' includes (uptime, rate limits, authentication requirements, available features). The description mentions capabilities but doesn't detail what information is returned or any behavioral constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that communicates the core purpose without unnecessary words. It's appropriately sized for a simple status-checking tool and front-loads the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is insufficiently complete. It doesn't explain what 'status and capabilities' means in practice, what format the information is returned in, or how this differs from the many search-focused sibling tools. The agent would need to guess about the tool's behavior and output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage for the single parameter, the schema already fully documents the 'validate' parameter. The description adds no additional parameter information beyond what's in the schema, so it meets the baseline expectation but doesn't provide extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Check') and resource ('status and capabilities of available academic platforms'), making it immediately understandable. However, it doesn't explicitly distinguish this from sibling tools that focus on searching or downloading content rather than platform status checking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With many sibling tools focused on searching specific databases, there's no indication whether this should be used before attempting searches, when troubleshooting, or as a general health check. The lack of usage context is a significant gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_arxivC
Search academic papers specifically from arXiv preprint server
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| sortBy | No | Sort results by field | |
| category | No | arXiv category filter (e.g., cs.AI, physics.gen-ph) | |
| sortOrder | No | Sort order: ascending or descending | |
| maxResults | No | Maximum number of results to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. While 'Search' implies a read-only operation, the description provides no information about rate limits, authentication requirements, result format, pagination, error conditions, or what happens when no results are found. This is a significant gap for a search tool with many parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that gets straight to the point. There's no wasted language or unnecessary elaboration - it clearly communicates the core function without any fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with 7 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what kind of results to expect, how results are structured, whether there are limitations or constraints, or how this tool differs from the many other search tools available. The context demands more guidance than what's provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% description coverage, so all parameters are well-documented in the structured schema. The description doesn't add any parameter-specific information beyond what's already in the schema, which is acceptable given the comprehensive schema coverage. The baseline of 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search') and resource ('academic papers from arXiv preprint server'), making the purpose immediately understandable. However, it doesn't distinguish this tool from its many sibling search tools (like search_biorxiv, search_pubmed, etc.) beyond mentioning arXiv specifically.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to use this tool versus the many alternative search tools available on the server. There's no mention of when arXiv search is preferable to other academic databases or what makes this tool distinct from its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_biorxivB
Search bioRxiv preprint server for biology papers
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Number of days to search back (default: 30) | |
| query | Yes | Search query string | |
| category | No | Category filter (e.g., neuroscience, genomics) | |
| maxResults | No | Maximum number of results to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the action 'search' but doesn't describe what the search returns (e.g., list of papers, metadata), any rate limits, authentication needs, or error conditions. This is a significant gap for a search tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a search tool with 4 parameters, no annotations, and no output schema, the description is minimally adequate. It clarifies the domain (bioRxiv, biology papers) but lacks details on return values, behavioral traits, or usage context, which are important for effective tool selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with all parameters well-documented in the input schema (e.g., query string, maxResults range, days default, category examples). The description adds no additional parameter semantics beyond what the schema provides, so it meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'search' and the resource 'bioRxiv preprint server for biology papers', which is specific and unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'search_medrxiv' or 'search_papers', which might have overlapping domains or purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'search_medrxiv' (for medical preprints) or 'search_pubmed' (for published papers). It lacks explicit context, exclusions, or prerequisites, leaving the agent to infer usage based on the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_crossrefC
Search academic papers from Crossref database. Free API with extensive scholarly metadata coverage across publishers.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| sortBy | No | Sort results by relevance, date, or citations | |
| sortOrder | No | Sort order: ascending or descending | |
| maxResults | No | Maximum number of results to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the API is 'free' and has 'extensive coverage,' which adds some context about cost and scope, but it doesn't cover critical behaviors like rate limits, authentication needs, error handling, or response format. For a search tool with no annotation coverage, this leaves significant gaps in understanding how the tool behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, with two sentences that efficiently convey the core purpose and key features (free API, extensive coverage). There's no wasted text, and it avoids redundancy. However, it could be slightly more structured by explicitly separating purpose from behavioral context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (6 parameters, no output schema, no annotations), the description is somewhat complete but has gaps. It covers the purpose and high-level context but lacks details on behavioral traits and usage guidelines. Without an output schema, it doesn't explain return values, which is a missed opportunity to add value beyond structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, meaning all parameters are documented in the input schema. The description doesn't add any specific parameter semantics beyond what the schema provides (e.g., it doesn't explain query syntax or year format details). According to the rules, with high schema coverage (>80%), the baseline is 3 even without param info in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search academic papers from Crossref database.' It specifies the verb ('search'), resource ('academic papers'), and data source ('Crossref database'), making the intent unambiguous. However, it doesn't explicitly differentiate from sibling tools like 'search_papers' or 'search_semantic_scholar' that might also search academic papers, which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides minimal usage guidance. It mentions that the API is 'free' and has 'extensive scholarly metadata coverage,' which hints at when to use it (e.g., for broad, cost-effective searches), but it doesn't explicitly state when to choose this tool over alternatives like 'search_pubmed' or 'search_arxiv' from the sibling list. No exclusions or specific contexts are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_google_scholarC
Search Google Scholar for academic papers using web scraping
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query string | |
| author | No | Author name filter | |
| yearLow | No | Earliest publication year | |
| yearHigh | No | Latest publication year | |
| maxResults | No | Maximum number of results to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only mentions 'web scraping' without detailing behavioral traits like rate limits, authentication needs, or potential risks (e.g., blocking). It lacks information on response format, error handling, or operational constraints, leaving significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero wasteβit directly states the tool's function and method. It's appropriately sized and front-loaded, making it easy to grasp immediately without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a search tool with 5 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain return values, error cases, or how results are structured, which is critical for an agent to use the tool effectively. The 'web scraping' hint is insufficient for full context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 5 parameters. The description adds no additional meaning beyond the schema, such as query syntax examples or interactions between parameters like yearLow and yearHigh. Baseline 3 is appropriate as the schema handles parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('search') and resource ('Google Scholar for academic papers'), making the purpose understandable. However, it doesn't differentiate from sibling tools like search_arxiv or search_pubmed, which perform similar academic searches on different platforms, so it misses full distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, such as search_arxiv for physics papers or search_pubmed for medical literature. It mentions 'web scraping' but doesn't explain implications or exclusions, leaving usage context vague.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_iacrC
Search IACR ePrint Archive for cryptography papers
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query string | |
| maxResults | No | Maximum number of results to return | |
| fetchDetails | No | Fetch detailed information for each paper (slower) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'Search' but doesn't describe what the search returns (e.g., paper titles, abstracts, metadata), performance characteristics (e.g., speed implications of fetchDetails), error conditions, or authentication requirements. The phrase 'slower' in the schema hints at performance but isn't elaborated in the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any wasted words. It's front-loaded with the core action and target, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with 3 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what the search returns (e.g., list of papers with basic info), how results are formatted, or any limitations (e.g., date ranges, sorting options). The lack of output schema means the description should compensate by detailing return values, which it doesn't.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters (query, maxResults, fetchDetails) with their types and constraints. The description adds no additional parameter semantics beyond what's in the schema, maintaining the baseline score of 3 for adequate coverage through structured data alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Search') and resource ('IACR ePrint Archive for cryptography papers'), making the purpose immediately understandable. It distinguishes from general search tools by specifying the IACR ePrint Archive, though it doesn't explicitly differentiate from sibling tools like search_arxiv that also search academic repositories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With many sibling search tools (e.g., search_arxiv, search_pubmed, search_google_scholar), there's no indication that this is specifically for cryptography papers from IACR, nor any context about when it might be preferred over other search options.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_medrxivC
Search medRxiv preprint server for medical papers
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Number of days to search back (default: 30) | |
| query | Yes | Search query string | |
| category | No | Category filter (e.g., infectious_diseases, epidemiology) | |
| maxResults | No | Maximum number of results to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure but provides minimal information. It doesn't mention rate limits, authentication requirements, response format, pagination behavior, or what happens when no results are found. For a search tool with no annotation coverage, this leaves significant behavioral questions unanswered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that states the core purpose without any wasted words. It's appropriately sized for a search tool and front-loads the essential information. Every word earns its place in this concise formulation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of having 18 sibling tools (many of which are similar search tools) and no annotations or output schema, the description is insufficiently complete. It doesn't help the agent navigate the crowded tool ecosystem or understand medRxiv's specific value proposition versus other repositories. For a search tool among many alternatives, more contextual guidance is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no parameter information beyond what's already in the schema, which has 100% coverage with clear descriptions for all 4 parameters. The baseline is 3 when schema coverage is high (>80%), and the description doesn't compensate with additional context about how parameters interact or search behavior nuances.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Search') and resource ('medRxiv preprint server for medical papers'), making the purpose immediately understandable. However, it doesn't explicitly differentiate this tool from its many sibling search tools (like search_arxiv, search_biorxiv, search_pubmed, etc.), which all search different repositories. A perfect score would require distinguishing this specific medRxiv search from other similar search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to use this tool versus its many sibling search tools. With 18 sibling tools including multiple search tools for different repositories (medRxiv, arXiv, bioRxiv, PubMed, etc.), the agent receives no help in choosing between them. There's no mention of medRxiv's specific focus (medical preprints) versus other databases' scopes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_papersC
Search academic papers from multiple sources including arXiv, Web of Science, etc.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Number of days to search back (bioRxiv/medRxiv only) | |
| year | No | Year filter (e.g., "2023", "2020-2023", "2020-") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| sortBy | No | Sort results by relevance, date, or citations | |
| journal | No | Journal name filter | |
| category | No | Category filter (e.g., cs.AI for arXiv) | |
| platform | No | Platform to search (default: crossref). Options: arxiv, webofscience/wos, pubmed, biorxiv, medrxiv, semantic, iacr, googlescholar/scholar, scihub, sciencedirect, springer, scopus, crossref, or all. Note: Wiley only supports PDF download by DOI, use download_paper instead. | |
| sortOrder | No | Sort order: ascending or descending | |
| maxResults | No | Maximum number of results to return | |
| fetchDetails | No | Fetch detailed information (IACR only) | |
| fieldsOfStudy | No | Fields of study filter (Semantic Scholar only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states it 'searches' papers, implying a read-only operation, but doesn't mention any behavioral traits like rate limits, authentication needs, pagination, or what the output looks like (e.g., format, fields returned). For a tool with 12 parameters and no output schema, this leaves significant gaps in understanding how it behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that states the core functionality without unnecessary words. It's appropriately sized for a search tool, though it could be slightly more informative given the tool's complexity. The structure is front-loaded with the main purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's high complexity (12 parameters, many siblings, no output schema, and no annotations), the description is inadequate. It doesn't explain the relationship with sibling tools, output format, or behavioral constraints. For a multi-platform search tool with extensive parameters, more context is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no parameter-specific information beyond implying a multi-source search capability, which relates to the 'platform' parameter. It doesn't provide additional context like default behaviors or parameter interactions, so it meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search academic papers from multiple sources including arXiv, Web of Science, etc.' It specifies the verb ('search') and resource ('academic papers'), and mentions the multi-source capability. However, it doesn't explicitly differentiate from sibling tools like search_arxiv or search_crossref, which are more specialized versions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its many siblings (e.g., search_arxiv, search_crossref). It mentions 'multiple sources' but doesn't clarify if this is a unified search across all platforms or how it differs from using individual platform-specific tools. The only usage hint is in the input schema's platform parameter description, which notes 'Wiley only supports PDF download by DOI, use download_paper instead,' but this isn't in the main description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_pubmedC
Search biomedical literature from PubMed/MEDLINE database using NCBI E-utilities API
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Publication year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| sortBy | No | Sort results by relevance or date | |
| journal | No | Journal name filter | |
| maxResults | No | Maximum number of results to return | |
| publicationType | No | Publication type filter (e.g., ["Journal Article", "Review"]) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions the NCBI E-utilities API but doesn't describe rate limits, authentication requirements, pagination behavior, error handling, or what the response format looks like. For a search tool with no annotation coverage, this leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that clearly states the tool's purpose. There's no wasted language or unnecessary elaboration. However, it could be slightly improved by front-loading more context about when to use this specific PubMed search versus other search tools.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with 7 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what the search returns (abstracts, citations, full metadata?), doesn't mention rate limits or API constraints, and provides no guidance on usage context. The description should do more to compensate for the lack of structured metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema. According to scoring rules, when schema coverage is high (>80%), the baseline is 3 even with no param info in description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches biomedical literature from PubMed/MEDLINE using NCBI E-utilities API, providing a specific verb ('Search') and resource ('biomedical literature from PubMed/MEDLINE database'). However, it doesn't explicitly differentiate from sibling tools like search_medrxiv or search_biorxiv that also search biomedical literature, missing sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. With many sibling tools available for searching different databases (e.g., search_arxiv, search_google_scholar, search_scopus), there's no indication of when PubMed-specific searching is preferred or what makes this tool distinct from general search tools like search_papers.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_sciencedirectC
Search academic papers from Elsevier ScienceDirect database (requires API key)
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| journal | No | Journal name filter | |
| maxResults | No | Maximum number of results to return | |
| openAccess | No | Filter for open access articles only |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It only mentions the API key requirement, but doesn't describe what the search returns (e.g., metadata, abstracts, full text availability), pagination behavior, rate limits, authentication scope, or error conditions. This leaves significant gaps for a search tool with 6 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that communicates the core purpose and key requirement. It's appropriately sized and front-loaded with the main functionality, though it could potentially be more structured with separate usage notes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with 6 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain what kind of results are returned, how they're formatted, whether there's pagination, or any limitations. The API key mention is helpful but insufficient for full contextual understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 6 parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema, maintaining the baseline score of 3 for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches academic papers from the Elsevier ScienceDirect database, which is a specific verb (search) and resource (academic papers from ScienceDirect). It distinguishes from siblings like search_arxiv or search_pubmed by specifying the database source, though it doesn't explicitly contrast with all similar search tools in the list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'requires API key' which provides some context about prerequisites, but it doesn't offer guidance on when to use this tool versus alternatives like search_scopus or search_semantic_scholar. No explicit when/when-not instructions or comparison to sibling tools are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_scihubA
Controlled, opt-in Sci-Hub DOI adapter (disabled by default); accepts DOI/doi.org URLs only, tries direct mirrors before bounded ScrapingAnt fallback, and does not grant access rights.
| Name | Required | Description | Default |
|---|---|---|---|
| doiOrUrl | Yes | DOI or doi.org URL only (e.g., "10.1038/nature12373"); ordinary paper URLs are rejected | |
| savePath | No | Directory to save the PDF file (if downloadPdf is true) | |
| downloadPdf | No | Whether to download the PDF file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses that the tool is disabled by default, requires opt-in, tries direct mirrors before a bounded ScrapingAnt fallback, and grants no access rights. It omits details like what happens when disabled or what the return payload is, but the key safety-relevant traits are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence carries the operation, constraints, fallback order, and a legal/access caveat with no filler. The most important facts (opt-in, disabled, DOI-only) are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description explains input restrictions and behavioral fallback well. The main gaps are how someone enables the opt-in, what happens when disabled, and whether the tool returns a match or only a saved PDF.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents doiOrUrl, savePath, and downloadPdf. The description adds no parameter-level meaning beyond restating the DOI-only constraint, which is already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies this as a 'Sci-Hub DOI adapter' and sharply constrains its input to DOI/doi.org URLs, which distinguishes it from sibling search tools. However, it never states an explicit operation verb such as search, fetch, or download, relying on the tool name and schema to convey what it returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear input constraints and hints at an opt-in workflow, but it does not compare this tool to siblings such as get_paper_by_doi or discover_paper_access, nor does it say when to choose it over them. Usage is mostly implied: use it when you have a DOI and want Sci-Hub access.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_scopusC
Search the Scopus abstract and citation database (requires Elsevier API key)
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| journal | No | Journal name filter | |
| subject | No | Subject area filter | |
| maxResults | No | Maximum number of results (max 25 per request) | |
| openAccess | No | Filter for open access articles only | |
| affiliation | No | Institution/affiliation filter | |
| documentType | No | Document type: ar=article, cp=conference paper, re=review, bk=book, ch=chapter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It only mentions the API key requirement but doesn't describe what the search returns (abstracts, citations, metadata), pagination behavior, rate limits, authentication scope, or error conditions. For a search tool with 9 parameters, this leaves significant behavioral aspects unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that communicates the core purpose and key prerequisite without any wasted words. It's appropriately sized and front-loaded with essential information, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a 9-parameter search tool with no annotations and no output schema, the description is insufficient. It doesn't explain what results to expect (format, fields, limitations), how results are ordered, whether there's pagination, or how to interpret the various filters. The API key mention is helpful but doesn't compensate for the missing behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters thoroughly. The description adds no parameter-specific information beyond what's in the schema. The baseline score of 3 reflects adequate coverage through the schema alone, though the description contributes nothing additional about parameter usage or interactions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search') and target resource ('Scopus abstract and citation database'), making the purpose immediately understandable. It distinguishes from some siblings by specifying the Scopus database, but doesn't explicitly differentiate from other academic search tools like search_arxiv or search_pubmed beyond the database name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the many sibling search tools available. It mentions the requirement for an Elsevier API key, which is a prerequisite but doesn't help the agent choose between Scopus and alternatives like Google Scholar, PubMed, or other databases for different search scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_semantic_scholarC
Search Semantic Scholar for academic papers with citation data
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| maxResults | No | Maximum number of results to return | |
| fieldsOfStudy | No | Fields of study filter (e.g., ["Computer Science", "Biology"]) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While 'Search' implies a read-only operation, the description doesn't mention important behavioral aspects like rate limits, authentication requirements, response format, pagination behavior, or error conditions. The mention of 'citation data' hints at what's returned but doesn't provide operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that communicates the core functionality without unnecessary words. It's appropriately sized for a search tool and front-loads the essential information. Every word earns its place in this concise statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with 4 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what kind of results to expect, how citation data is presented, or how this tool differs from the many other search tools available. The lack of behavioral context and usage guidance leaves significant gaps for an AI agent trying to use this tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 4 parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema descriptions. It mentions 'citation data' which relates to output rather than input parameters. Baseline score of 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Search') and resource ('Semantic Scholar for academic papers'), and specifies the type of data returned ('with citation data'). However, it doesn't explicitly differentiate this tool from its many sibling search tools (e.g., search_arxiv, search_pubmed) beyond mentioning Semantic Scholar specifically.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the many alternative search tools available on the server (search_arxiv, search_pubmed, search_google_scholar, etc.). There's no mention of Semantic Scholar's specific strengths, coverage, or when it might be preferred over other academic search tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_springerA
Search academic papers from Springer Nature database. Uses Metadata API by default (all content) or OpenAccess API when openAccess=true (full text available). Same API key works for both.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Publication type filter | |
| year | No | Year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| journal | No | Journal name filter | |
| subject | No | Subject area filter | |
| maxResults | No | Maximum number of results to return | |
| openAccess | No | Search only open access content |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses that the same API key works for both API modes, which is useful behavioral context about authentication. However, it doesn't mention rate limits, pagination behavior, error handling, or what the response format looks like, leaving significant gaps for a search tool with 8 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in two sentences with zero waste. The first sentence states the core purpose, and the second provides important behavioral context about API modes and authentication. Every sentence earns its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with 8 parameters, no annotations, and no output schema, the description provides adequate basic information about purpose and API behavior. However, it lacks details about response format, error conditions, or performance characteristics that would be helpful given the tool's complexity and lack of structured metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters thoroughly. The description adds minimal value by mentioning the openAccess parameter's effect on API selection, but doesn't provide additional semantic context beyond what's in the schema descriptions. Baseline 3 is appropriate when schema does heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches academic papers from Springer Nature database, specifying both the action (search) and resource (academic papers). It distinguishes itself from siblings by mentioning the Springer Nature database specifically, unlike other search tools like search_arxiv or search_pubmed that target different sources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use different API modes (Metadata API vs OpenAccess API based on openAccess parameter), which helps guide usage. However, it doesn't explicitly state when to choose this tool over sibling search tools like search_google_scholar or search_scopus, missing explicit alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_webofscienceC
Search Web of Science Starter (default) or explicitly selected Expanded API
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Publication year filter (e.g., "2023", "2020-2023") | |
| query | Yes | Search query string | |
| author | No | Author name filter | |
| sortBy | No | Starter-supported sort field | |
| journal | No | Journal name filter | |
| sortOrder | No | Sort order: ascending or descending | |
| apiProduct | No | API product; Expanded must be selected explicitly | starter |
| maxResults | No | Maximum number of results to return | |
| recordView | No | Expanded record detail; defaults to short | |
| discoverAccess | No | Find public publisher PDF links using controlled direct retrieval and optional paid fallback | |
| discoverAccessMaxItems | No | Maximum result items to enrich; defaults to ACCESS_DISCOVERY_MAX_ITEMS or 5 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It only mentions that Starter is the default and Expanded must be explicitly selected. It does not mention the return format, pagination behavior, rate limits, or side effects (e.g., discoverAccess fallback). The description carries a minimal behavioral disclosure, leaving most of the tool's behavior opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the core action. It is not verbose and has no filler. However, it is so brief that it sacrifices valuable context, but that is more a completeness issue than a conciseness issue. For the length it has, it is well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 11 parameters, no annotations, and no output schema, the description is severely incomplete. It does not explain that this searches scholarly literature, what types of queries are supported, or how it compares to the many sibling search tools. An agent has almost no guidance on what to expect or how to call it correctly beyond the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters are documented in the schema. The description adds only that 'Starter (default) or explicitly selected Expanded API', which essentially repeats the schema's default and description for apiProduct. Thus, it adds minimal value beyond the schema, and the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Search') and resource ('Web of Science'), and distinguishes between the two API products (Starter vs Expanded). It is specific enough to differentiate from siblings like search_google_scholar or search_pubmed, though it does not mention the nature of the content (scholarly database).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus the many sibling search tools. It does not mention that this is for Web of Science specifically, nor does it provide exclusions or alternatives. The only hint is the name, which is insufficient for an agent to decide between search_webofscience and search_scopus or search_crossref.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.3.2- Added
discover_paper_access - Changed
get_paper_by_doi1 field changed- changed
Input schema / properties / platform / enumPrevious value: -[ - "arxiv", - "webofscience", - "all" -]New value: +[ + "arxiv", + "webofscience", + "scihub", + "all" +]
- Added
get_webofscience_related_records - Changed
search_scihub1 field changed- changed
Input schema / properties / doiOrUrl / descriptionPrevious value: -"DOI (e.g., \"10.1038/nature12373\") or full paper URL"New value: +"DOI or doi.org URL only (e.g., \"10.1038/nature12373\"); ordinary paper URLs are rejected"
- Changed
search_webofscience8 fields changed- added
Input schema / properties / apiProductAdded value: +{ + "default": "starter", + "description": "API product; Expanded must be selected explicitly", + "enum": [ + "starter", + "expanded" + ], + "type": "string" +} - added
Input schema / properties / discoverAccessAdded value: +{ + "default": false, + "description": "Find public publisher PDF links using controlled direct retrieval and optional paid fallback", + "type": "boolean" +} - added
Input schema / properties / discoverAccessMaxItemsAdded value: +{ + "description": "Maximum result items to enrich; defaults to ACCESS_DISCOVERY_MAX_ITEMS or 5", + "maximum": 100, + "minimum": 1, + "type": "number" +} - added
Input schema / properties / maxResults / defaultAdded value: +10 - changed
Input schema / properties / maxResults / maximumPrevious value: -50New value: +100 - added
Input schema / properties / recordViewAdded value: +{ + "description": "Expanded record detail; defaults to short", + "enum": [ + "short", + "full" + ], + "type": "string" +} - changed
Input schema / properties / sortBy / descriptionPrevious value: -"Sort results by field"New value: +"Starter-supported sort field" - changed
Input schema / properties / sortBy / enumPrevious value: -[ - "relevance", - "date", - "citations", - "title", - "author", - "journal" -]New value: +[ + "relevance", + "date", + "citations" +]
2 tool updates
v0.2.7-patch- Added
get_citations - Removed
search_wiley
19 tool updates
- First observed
check_scihub_mirrors - First observed
download_paper - First observed
get_paper_by_doi - First observed
get_platform_status - First observed
search_arxiv - First observed
search_biorxiv - First observed
search_crossref - First observed
search_google_scholar - First observed
search_iacr - First observed
search_medrxiv - First observed
search_papers - First observed
search_pubmed - First observed
search_sciencedirect - First observed
search_scihub - First observed
search_scopus - First observed
search_semantic_scholar - First observed
search_springer - First observed
search_webofscience - First observed
search_wiley
TDQS
Scored across 21 tools
Most tools are clearly distinguished by their target source (e.g., arXiv, PubMed, Scopus) or action (download, check, discover). Minor overlap exists between search_papers (multi-source) and individual source-specific searches, but descriptions clarify scope. The access-related tools (search_scihub, discover_paper_access) have distinct roles.
The naming pattern is predominantly verb_noun with search_* for queries, get_* for retrieval, and a few other action prefixes (download, check, discover). While consistent overall, some names are verbose (get_webofscience_related_records) and there is no uniform verb for all operations, but the convention is recognizable and predictable.
With 21 tools, the server sits at the upper edge of reasonable scope. The breadth of academic sources justifies a higher count, but it feels slightly heavy, especially with overlapping search capabilities. Each tool serves a purpose, yet a few could be consolidated without losing functionality.
The tool surface covers the core lifecycle of academic paper search and retrieval: searching across major databases, retrieving by DOI, downloading PDFs, and accessing citation data. Minor gaps exist, such as no direct full-text extraction or citation network visualization, but these are not essential for the stated purpose. The coverage is robust for search and access.
Maintenance
Related MCP Connectors
Search and download academic papers from arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, Semanticβ¦
Find academic papers across major sources like arXiv, PubMed, bioRxiv, and more. Download PDFs wheβ¦
Federated search of books and papers, BibTeX/RIS citations, open-access retrieval and reading.
Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables searching and downloading academic papers from multiple sources including arXiv, PubMed, bioRxiv, Google Scholar, and Semantic Scholar. Provides standardized tools compatible with OpenAI Deep Research and ChatGPT connectors.15MIT
- AlicenseNot gradedqualityCmaintenanceEnables searching and downloading academic papers from multiple sources including arXiv, PubMed, bioRxiv, Google Scholar, and Semantic Scholar. Provides standardized tools for research workflows and OpenAI Deep Research integration.5MIT
- FlicenseAqualityBmaintenanceEnables searching, downloading, and reading academic papers from multiple platforms including arXiv, Semantic Scholar, PubMed, bioRxiv, medRxiv, IACR, Google Scholar, RePEc/IDEAS, and Sci-Hub with PDF to Markdown conversion.297-
- AlicenseAqualityCmaintenanceEnables users to search, download, and read academic papers from multiple platforms including arXiv, PubMed, bioRxiv, Google Scholar, Semantic Scholar, and CrossRef through a unified interface.344MIT