eutils-mcp-server
eutils-mcp-server
中文 · English
MCP server exposing the nine NCBI Entrez E-utilities as eleven read-only tools.
It gives a language model the E-utilities directly, so it can search biomedical
literature, fetch sequences, and follow links between Entrez databases without a browser
and without scraping. The repository folder is named E-utilities after the API it
wraps. The package, the server name, and the GitHub repository are all
eutils-mcp-server.
Use it when you want an answer, not a dataset. A question such as "find papers about CRISPR delivery in 2024 and show me the abstracts" is the intended shape. Do not use it for bulk download: NCBI asks that bulk data mining use the local PubMed copy instead, and this server obeys NCBI's rate limit, so a large job takes days. Do not use it to change anything, because it is read-only and NCBI exposes no write path through these utilities.
Table of Contents
Background
The Entrez Programming Utilities are a set of nine server-side programs at the National Center for Biotechnology Information. They share one URL syntax and cover 38 databases, including PubMed, PMC, Protein, Nucleotide, Gene, SNP, Structure, and Taxonomy. NCBI documents them in the E-utilities manual.
That interface is stable but awkward for a model to drive directly. Every request needs the same credentials, the same encoding, and the same knowledge of which endpoints accept which output formats. Batch retrieval needs the History server, which carries state across calls. NCBI blocks an IP that exceeds its rate limit.
This server puts those details in code. One client builds every URL, paces every request, follows redirects within an allowlist, caps response size, and redacts the API key. Each tool then maps one operation to one E-utilities call and returns a result the model can read.
Two upstream facts shaped the design. EGQuery is unreachable from the public internet, and the ESummary JSON limit is 500 records rather than the 10,000 the manual states. Both are recorded, with evidence, in docs/upstream-issues.md.
Install
git clone https://github.com/iwinoid/eutils-mcp-server.git
cd eutils-mcp-server
npm install
npm run buildDependencies
Node.js 22 or newer. The version is pinned in .nvmrc.
The server has three runtime dependencies: @modelcontextprotocol/server,
fast-xml-parser, and zod. The lockfile is committed. There is no installer script and
no native build step.
Usage
Register the server with your MCP host. Add this to the host configuration, for example
claude_desktop_config.json:
{
"mcpServers": {
"eutils": {
"command": "node",
"args": ["/absolute/path/to/E-utilities/dist/index.js"],
"env": {
"NCBI_API_KEY": "your-key",
"NCBI_EMAIL": "you@example.com",
"NCBI_TOOL": "eutils-mcp-server"
}
}
}
}To keep the key out of the host configuration, read it from .env instead:
{
"mcpServers": {
"eutils": {
"command": "node",
"args": [
"--env-file-if-exists=/absolute/path/to/E-utilities/.env",
"/absolute/path/to/E-utilities/dist/index.js"
]
}
}
}Then ask the model to search. You can also call a tool directly:
npm run call eutils_esearch '{"db":"pubmed","term":"CRISPR delivery AND 2024[pdat]","retmax":3}'The server answers with a count, a UID list, the translated query, and a History handle:
# ESearch: `CRISPR delivery AND 2024[pdat]`
Database **pubmed** matched **412** records. Showing 3 starting at 0.
## UIDs
40123456, 40123457, 40123458
## History handle
{"db":"pubmed","web_env":"MCID_6aa2...","query_key":"1"}The count changes as PubMed grows. Confirm that you see a count and three UIDs.
Configuration
Every variable is optional. NCBI asks automated clients to identify themselves, and an API key raises the rate limit.
Variable | Default | Purpose |
| unset | Raises the ceiling from 3 to 10 requests per second. Get one from the Settings page of your NCBI account. |
| unset | Contact address sent with every request. NCBI uses it to warn you before an IP block. |
|
| Name that identifies this software in the NCBI logs. |
Supply the values in the env block of your host configuration, or keep them in a .env
file at the project root and launch the server with --env-file-if-exists. npm start
and npm run dev pass that flag for you. When the file is absent, Node prints
not found. Continuing without it. and the launch continues, so a missing .env cannot
break startup. The path resolves against the working directory. Give an absolute path when
your host launches the server from elsewhere.
.gitignore excludes .env. It does not exclude a host configuration file. Check that
file before you commit it.
Set NCBI_EMAIL and NCBI_TOOL, then register both with NCBI by mail to
eutilities@ncbi.nlm.nih.gov. A request that carries the values without prior
registration does not satisfy the NCBI usage policy.
Subscribe to the Entrez Utilities announcement list at the same address. It is the only NCBI channel that reports known bugs. The NCBI Insights blog reports planned changes only, and the release notes inside the E-utilities manual stop at 2015.
Rate limits
NCBI blocks an IP that exceeds its limit. All requests, including internal batches, pass through one token bucket.
3 requests per second without an API key
10 requests per second with one
NCBI asks that large jobs run at a weekend. On a weekday, run them between 21:00 and 05:00 US Eastern time.
Limits
Entrez only. The server reads what Entrez indexes. Data that lives outside Entrez is not reachable.
retmaxceilings.eutils_esearchaccepts up to 10,000.eutils_esummaryandeutils_efetchaccept up to 500 per call. A larger UID list is split into batches of 500, and the response reportsbatches.PubMed and PMC caps. ESearch reaches only the first 10,000 records of a PubMed or PMC result set. Add date filters to segment a larger set.
Truncation. The server cuts a response over 25,000 characters. The message says how to page or narrow the query.
Response ceiling. The client abandons a response body over 5 MB.
stdio only. No HTTP transport. To add one, bind
127.0.0.1and validate theOriginandHostheaders.EGQuery coverage. EGQuery itself is unreachable, so
eutils_egquerycovers 12 databases instead of 38. See below.
Known upstream issue: EGQuery
NCBI's egquery.fcgi answers with an HTTP 301 to
ext-http-eutils.linkerd.ncbi.nlm.nih.gov. That host is not published in public DNS.
Two independent DNSSEC-validating resolvers, Cloudflare and Google, both return NXDOMAIN
for the name. A control query for eutils.ncbi.nlm.nih.gov resolves normally.
Every parameter combination tried redirects: GET and POST, with and without retmode,
retmax, tool, email, a browser User-Agent, and HTTP/1.0.
An API key does not help. With a valid key, esearch returns 200 while egquery still
returns 301 in the same session. A syntactically invalid key makes egquery return
400 API key invalid instead. That result shows NCBI validates the key before it routes
the request, so the redirect is not a credentials or rate-limit decision.
Run npm run doctor to reproduce the finding on your own network. The full evidence
chain, and the other upstream defects found while building this server, are in
docs/upstream-issues.md.
eutils_egquery tries the real EGQuery first. Only a network failure starts the fallback,
and then the server counts matches with ESearch over 12 commonly used databases. The
result carries degraded: true, a reason, and a note. Read that marker as "this covers a
subset, not all 38 databases". A validation error never starts the fallback, so a bad
query cannot cost 12 extra requests.
The fallback is lazy. If NCBI repairs the endpoint, the real EGQuery returns and no code changes.
Security
The server is read-only and holds no listener. It never writes to disk. There is no port for an attacker to connect to: the only network traffic is outbound HTTPS from the server to NCBI.
Threat | Control |
Prompt injection carried by record text | The server fences NCBI record text between |
Parameter injection | The server validates |
API key leakage | The server masks |
SSRF | The base URL is a constant, not an environment setting. The client follows redirects manually, at most three hops, and every hop must end with |
Resource exhaustion | A token bucket, per-endpoint |
Malicious XML | Entity processing is off. The client strips DOCTYPE declarations and caps the body size before parsing. |
Supply chain | Three runtime dependencies. The lockfile is committed. |
Report a vulnerability through the issue tracker.
API
Eleven tools. Every tool accepts response_format of "markdown" or "json", and
defaults to markdown. Every tool reports readOnlyHint: true and
destructiveHint: false, and declares an outputSchema that its own result satisfies.
Tool | Purpose |
| List databases, or describe one database: searchable fields, links, record count |
| Search a database. Returns UIDs and a History handle |
| Upload a UID list to the NCBI History server |
| Compact summaries for a UID set: title, authors, journal, date |
| Full records: PubMed abstracts, FASTA sequences, other formats |
| Follow links between databases, for example pubmed to pmc, or gene to protein |
| Count matches across many databases at once |
| Spelling suggestion for a query |
| Resolve formatted citations to PMIDs |
| Search and download in one call |
| Follow links and download the target records in one call |
Each tool takes one of the nine E-utilities as its subject. The tool name carries the
E-utilities name after the eutils_ prefix, so eutils_esearch calls esearch.fcgi.
Parameter names follow the API, so the manual applies directly.
Working with large result sets
The server keeps no state. The History handle travels as an ordinary value, so you pass it back unchanged.
eutils_esearch(db="pubmed", term="...", retmax=0, usehistory=true)
-> { total: 16896, history: { db, web_env, query_key } }
eutils_efetch(history={...}, retstart=0, retmax=500)
eutils_efetch(history={...}, retstart=500, retmax=500)Run the inspector to read the full input and output schema of every tool:
npm run build
npx @modelcontextprotocol/inspector node dist/index.jsMaintainers
Contributing
Ask questions in the issue tracker. Pull requests are accepted.
Read CONTRIBUTING.md before you open one. It states the development commands, the requirements a pull request must meet, and how to read the coverage number.
This project follows the Contributor Covenant, version 2.1.
License
MIT © iwinoid
NCBI supplies the data. If you redistribute this software or its output, NCBI's Disclaimer and Copyright notice must be evident to users. PubMed abstracts can be protected by copyright. Redistribution beyond fair use needs the permission of the copyright holder.