Skip to main content
Glama
README.md
# eutils-mcp-server

[中文](README.zh.md) · English

[![standard-readme compliant](https://img.shields.io/badge/readme%20style-standard-brightgreen.svg?style=flat-square)](https://github.com/RichardLitt/standard-readme) [![CI](https://img.shields.io/github/actions/workflow/status/iwinoid/eutils-mcp-server/ci.yml?style=flat-square&label=CI&logo=github)](https://github.com/iwinoid/eutils-mcp-server/actions) [![LICENSE MIT](https://img.shields.io/badge/LICENSE-MIT-blue?style=flat-square)](LICENSE) [![POWERED BY DEEPSEEK](https://img.shields.io/badge/POWERED_BY-DEEPSEEK-4D6BFE?style=flat-square&logo=deepseek&logoColor=white)](https://www.deepseek.com) [![powered by dsh](https://img.shields.io/badge/powered_by-dsh-4D6BFE?style=flat-square&logo=deepseek&logoColor=white)](https://github.com/deepseek-ai/deepseek-harness)

MCP server exposing the nine NCBI Entrez E-utilities as eleven read-only tools.

It gives a language model the E-utilities directly, so it can search biomedical
literature, fetch sequences, and follow links between Entrez databases without a browser
and without scraping. The repository folder is named `E-utilities` after the API it
wraps. The package, the server name, and the GitHub repository are all
`eutils-mcp-server`.

Use it when you want an answer, not a dataset. A question such as "find papers about
CRISPR delivery in 2024 and show me the abstracts" is the intended shape. Do not use it
for bulk download: NCBI asks that bulk data mining use the local PubMed copy instead, and
this server obeys NCBI's rate limit, so a large job takes days. Do not use it to change
anything, because it is read-only and NCBI exposes no write path through these utilities.

## Table of Contents

- [Background](#background)
- [Install](#install)
- [Usage](#usage)
- [Configuration](#configuration)
- [Rate limits](#rate-limits)
- [Limits](#limits)
- [Known upstream issue: EGQuery](#known-upstream-issue-egquery)
- [Security](#security)
- [API](#api)
- [Maintainers](#maintainers)
- [Contributing](#contributing)
- [License](#license)

## Background

The Entrez Programming Utilities are a set of nine server-side programs at the National
Center for Biotechnology Information. They share one URL syntax and cover 38 databases,
including PubMed, PMC, Protein, Nucleotide, Gene, SNP, Structure, and Taxonomy. NCBI
documents them in the [E-utilities manual](https://www.ncbi.nlm.nih.gov/books/NBK25501/).

That interface is stable but awkward for a model to drive directly. Every request needs
the same credentials, the same encoding, and the same knowledge of which endpoints accept
which output formats. Batch retrieval needs the History server, which carries state
across calls. NCBI blocks an IP that exceeds its rate limit.

This server puts those details in code. One client builds every URL, paces every request,
follows redirects within an allowlist, caps response size, and redacts the API key. Each
tool then maps one operation to one E-utilities call and returns a result the model can
read.

Two upstream facts shaped the design. EGQuery is unreachable from the public internet,
and the ESummary JSON limit is 500 records rather than the 10,000 the manual states. Both
are recorded, with evidence, in [docs/upstream-issues.md](docs/upstream-issues.md).

## Install

```console
git clone https://github.com/iwinoid/eutils-mcp-server.git
cd eutils-mcp-server
npm install
npm run build
```

### Dependencies

Node.js 22 or newer. The version is pinned in `.nvmrc`.

The server has three runtime dependencies: `@modelcontextprotocol/server`,
`fast-xml-parser`, and `zod`. The lockfile is committed. There is no installer script and
no native build step.

## Usage

Register the server with your MCP host. Add this to the host configuration, for example
`claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "eutils": {
      "command": "node",
      "args": ["/absolute/path/to/E-utilities/dist/index.js"],
      "env": {
        "NCBI_API_KEY": "your-key",
        "NCBI_EMAIL": "you@example.com",
        "NCBI_TOOL": "eutils-mcp-server"
      }
    }
  }
}
```

To keep the key out of the host configuration, read it from `.env` instead:

```json
{
  "mcpServers": {
    "eutils": {
      "command": "node",
      "args": [
        "--env-file-if-exists=/absolute/path/to/E-utilities/.env",
        "/absolute/path/to/E-utilities/dist/index.js"
      ]
    }
  }
}
```

Then ask the model to search. You can also call a tool directly:

```console
npm run call eutils_esearch '{"db":"pubmed","term":"CRISPR delivery AND 2024[pdat]","retmax":3}'
```

The server answers with a count, a UID list, the translated query, and a History handle:

```text
# ESearch: `CRISPR delivery AND 2024[pdat]`

Database **pubmed** matched **412** records. Showing 3 starting at 0.

## UIDs

40123456, 40123457, 40123458

## History handle

{"db":"pubmed","web_env":"MCID_6aa2...","query_key":"1"}
```

The count changes as PubMed grows. Confirm that you see a count and three UIDs.

## Configuration

Every variable is optional. NCBI asks automated clients to identify themselves, and an
API key raises the rate limit.

| Variable       | Default             | Purpose                                                                                                                                            |
| -------------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| `NCBI_API_KEY` | unset               | Raises the ceiling from 3 to 10 requests per second. Get one from the Settings page of your [NCBI account](https://www.ncbi.nlm.nih.gov/account/). |
| `NCBI_EMAIL`   | unset               | Contact address sent with every request. NCBI uses it to warn you before an IP block.                                                              |
| `NCBI_TOOL`    | `eutils-mcp-server` | Name that identifies this software in the NCBI logs.                                                                                               |

Supply the values in the `env` block of your host configuration, or keep them in a `.env`
file at the project root and launch the server with `--env-file-if-exists`. `npm start`
and `npm run dev` pass that flag for you. When the file is absent, Node prints
`not found. Continuing without it.` and the launch continues, so a missing `.env` cannot
break startup. The path resolves against the working directory. Give an absolute path when
your host launches the server from elsewhere.

`.gitignore` excludes `.env`. It does not exclude a host configuration file. Check that
file before you commit it.

Set `NCBI_EMAIL` and `NCBI_TOOL`, then register both with NCBI by mail to
<eutilities@ncbi.nlm.nih.gov>. A request that carries the values without prior
registration does not satisfy the NCBI usage policy.

Subscribe to the Entrez Utilities announcement list at the same address. It is the only
NCBI channel that reports known bugs. The
[NCBI Insights blog](https://ncbiinsights.ncbi.nlm.nih.gov/tag/e-utilities/) reports
planned changes only, and the release notes inside the
[E-utilities manual](https://www.ncbi.nlm.nih.gov/books/NBK25501/) stop at 2015.

## Rate limits

NCBI blocks an IP that exceeds its limit. All requests, including internal batches, pass
through one token bucket.

- 3 requests per second without an API key
- 10 requests per second with one

NCBI asks that large jobs run at a weekend. On a weekday, run them between 21:00 and
05:00 US Eastern time.

## Limits

- **Entrez only.** The server reads what Entrez indexes. Data that lives outside Entrez
  is not reachable.
- **`retmax` ceilings.** `eutils_esearch` accepts up to 10,000. `eutils_esummary` and
  `eutils_efetch` accept up to 500 per call. A larger UID list is split into batches of
  500, and the response reports `batches`.
- **PubMed and PMC caps.** ESearch reaches only the first 10,000 records of a PubMed or
  PMC result set. Add date filters to segment a larger set.
- **Truncation.** The server cuts a response over 25,000 characters. The message says how
  to page or narrow the query. A single record larger than the limit keeps its head
  and tail with an omission marker, because `retstart` pages between records, not
  within one.
- **Response ceiling.** The client abandons a response body over 5 MB.
- **stdio only.** No HTTP transport. To add one, bind `127.0.0.1` and validate the
  `Origin` and `Host` headers.
- **EGQuery coverage.** EGQuery itself is unreachable, so `eutils_egquery` covers 12
  databases instead of 38. See below.

## Known upstream issue: EGQuery

NCBI's `egquery.fcgi` answers with an HTTP 301 to
`ext-http-eutils.linkerd.ncbi.nlm.nih.gov`. That host is not published in public DNS.

Two independent DNSSEC-validating resolvers, Cloudflare and Google, both return NXDOMAIN
for the name. A control query for `eutils.ncbi.nlm.nih.gov` resolves normally.

Every parameter combination tried redirects: GET and POST, with and without `retmode`,
`retmax`, `tool`, `email`, a browser User-Agent, and HTTP/1.0.

An API key does not help. With a valid key, `esearch` returns 200 while `egquery` still
returns 301 in the same session. A syntactically invalid key makes `egquery` return
`400 API key invalid` instead. That result shows NCBI validates the key before it routes
the request, so the redirect is not a credentials or rate-limit decision.

Run `npm run doctor` to reproduce the finding on your own network. The full evidence
chain, and the other upstream defects found while building this server, are in
[docs/upstream-issues.md](docs/upstream-issues.md).

`eutils_egquery` tries the real EGQuery first. Only a network failure starts the fallback,
and then the server counts matches with ESearch over 12 commonly used databases. The
result carries `degraded: true`, a reason, and a note. Read that marker as "this covers a
subset, not all 38 databases". A validation error never starts the fallback, so a bad
query cannot cost 12 extra requests.

The fallback is lazy. If NCBI repairs the endpoint, the real EGQuery returns and no code
changes.

## Security

The server is read-only and holds no listener. It never writes to disk. There is
no port for an attacker to connect to: the only network traffic is outbound HTTPS
from the server to NCBI.

| Threat                                  | Control                                                                                                                                                                                                                                                                  |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Prompt injection carried by record text | The server fences NCBI record text between `<<<EXTERNAL_NCBI_DATA` markers and labels it as data. It strips fence markers from the content, so the content cannot close the fence early. The server never writes, so it cannot become a deputy for a destructive action. |
| Parameter injection                     | The server validates `db` against a character-class guard and a 38-database allowlist. It validates UIDs, search terms, and History fields before use. It encodes every value with `URLSearchParams` and never builds a URL by concatenation.                            |
| API key leakage                         | The server masks `api_key` in every log line, error message, and response. It never echoes the constructed URL to the model. stdio logging goes to stderr only.                                                                                                          |
| SSRF                                    | The base URL is a constant, not an environment setting. The client follows redirects manually, at most three hops, and every hop must end with `.ncbi.nlm.nih.gov`.                                                                                                      |
| Resource exhaustion                     | A token bucket, per-endpoint `retmax` ceilings, a 5 MB response ceiling, a request timeout, and bounded retries with backoff.                                                                                                                                            |
| Malicious XML                           | Entity processing is off. The client strips DOCTYPE declarations and caps the body size before parsing.                                                                                                                                                                  |
| Supply chain                            | Three runtime dependencies. The lockfile is committed.                                                                                                                                                                                                                   |

Report a vulnerability through the
[issue tracker](https://github.com/iwinoid/eutils-mcp-server/issues).

## API

Eleven tools. Every tool accepts `response_format` of `"markdown"` or `"json"`, and
defaults to markdown. Every tool reports `readOnlyHint: true` and
`destructiveHint: false`, and declares an `outputSchema` that its own result satisfies.

| Tool                       | Purpose                                                                          |
| -------------------------- | -------------------------------------------------------------------------------- |
| `eutils_einfo`             | List databases, or describe one database: searchable fields, links, record count |
| `eutils_esearch`           | Search a database. Returns UIDs and a History handle                             |
| `eutils_epost`             | Upload a UID list to the NCBI History server                                     |
| `eutils_esummary`          | Compact summaries for a UID set: title, authors, journal, date                   |
| `eutils_efetch`            | Full records: PubMed abstracts, FASTA sequences, other formats                   |
| `eutils_elink`             | Follow links between databases, for example pubmed to pmc, or gene to protein    |
| `eutils_egquery`           | Count matches across many databases at once                                      |
| `eutils_espell`            | Spelling suggestion for a query                                                  |
| `eutils_ecitmatch`         | Resolve formatted citations to PMIDs                                             |
| `eutils_search_then_fetch` | Search and download in one call                                                  |
| `eutils_link_then_fetch`   | Follow links and download the target records in one call                         |

Each tool takes one of the nine E-utilities as its subject. The tool name carries the
E-utilities name after the `eutils_` prefix, so `eutils_esearch` calls `esearch.fcgi`.
Parameter names follow the API, so the manual applies directly.

### Working with large result sets

The server keeps no state. The History handle travels as an ordinary value, so you pass it
back unchanged.

```text
eutils_esearch(db="pubmed", term="...", retmax=0, usehistory=true)
  -> { total: 16896, history: { db, web_env, query_key } }

eutils_efetch(history={...}, retstart=0,   retmax=500)
eutils_efetch(history={...}, retstart=500, retmax=500)
```

Run the inspector to read the full input and output schema of every tool:

```console
npm run build
npx @modelcontextprotocol/inspector node dist/index.js
```

## Maintainers

[@iwinoid](https://github.com/iwinoid)

## Contributing

Ask questions in the [issue tracker](https://github.com/iwinoid/eutils-mcp-server/issues).
Pull requests are accepted.

Read [CONTRIBUTING.md](CONTRIBUTING.md) before you open one. It states the development
commands, the requirements a pull request must meet, and how to read the coverage number.

This project follows the [Contributor Covenant](CODE_OF_CONDUCT.md), version 2.1.

## License

[MIT](LICENSE) © iwinoid

NCBI supplies the data. If you redistribute this software or its output, NCBI's
[Disclaimer and Copyright notice](https://www.ncbi.nlm.nih.gov/About/disclaimer.html)
must be evident to users. PubMed abstracts can be protected by copyright. Redistribution
beyond fair use needs the permission of the copyright holder.

TDQS

A4.7/5.0

Scored across 11 tools

Disambiguation5/5

Each tool maps to a distinct E-utilities operation (search, fetch, summary, link, post, spell, cite-match, database listing, cross-db counts), and descriptions include explicit 'Don't use when' cross-references. The two composite tools (search_then_fetch, link_then_fetch) could overlap with manual chains, but the descriptions clearly state when to prefer each, e.g. use esummary first to screen titles.

Naming Consistency4/5

All tools share the eutils_ prefix and mostly follow the NCBI API names (esearch, efetch, elink), giving a predictable pattern. The two convenience tools use snake_case (search_then_fetch, link_then_fetch), which is readable but a slight deviation from the concatenated style.

Tool Count5/5

11 tools is well-scoped, essentially covering the standard NCBI E-utilities suite plus two pragmatic shortcuts. Each tool earns its place with no redundant entries.

Completeness5/5

The surface covers the full E-utilities lifecycle: discovery (einfo, egquery), searching (esearch), history management (epost), retrieval (esummary, efetch), linking (elink), and utilities (espell, ecitmatch), plus end-to-end combos. No obvious dead ends for the domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues