annas-mcp
An MCP server (and CLI) that searches Anna's Archive for books and papers and downloads the copies you pick to a local folder.
book_search— find books, textbooks, manuals and standards by title, author or topic; returns a page of results with hash, authors, publisher, format, size and language.article_search— find journal articles by keyword, or look up a single paper by DOI (10.…,doi:…orhttps://doi.org/…).book_download— download a chosen book by itshashplustitle, with an optionalformatfor the file extension.article_download— download a paper by exactly one ofdoiorhash, with optionaltitle/format; falls back to the SciDB PDF when fast servers don't have it.Filtering and paging —
language,content(fiction, nonfiction, comic, magazine, standards),limitandpageon searches.Timeouts — every tool accepts
timeout_seconds(searches default 60 s, downloads 1800 s, max 86400 s).Safe downloads — saved under
ANNAS_DOWNLOAD_PATH, MD5-verified, corrupt files deleted, existing files never overwritten (a suffix is added), andpathplusbytesreturned.
Enables searching for and downloading journal articles by DOI, including DOI-based lookups that may contact doi.org.
annas-mcp
An MCP server that lets Claude, Cursor, Codex and other AI clients search Anna's Archive for books and papers and download the file you pick. It also works as a command-line tool.
Tool | What it does |
| Finds books, textbooks, manuals and standards by title, author or topic |
| Finds journal articles by keyword, or looks one up by DOI |
| Downloads a book using the hash from a search result |
| Downloads an article by DOI or hash |
Use it only for material you are entitled to obtain, under the laws and terms that apply to you.
Setup
You need Node.js 18 or newer and an Anna's Archive membership:
To… | Membership |
Search | Lucky Librarian or higher |
Download | Any tier |
1. Get your credentials
Account cookie. Sign in to Anna's Archive, open your browser's developer tools and go to Application (Chrome, Edge) or Storage (Firefox, Safari) → Cookies. Copy the value of
aa_account_id2. It expires weekly; see Renewing the cookie.API key. Copy it from your Anna's Archive account page. See the API FAQ. Only needed for downloads.
2. Add the server to your client
Claude Code
claude mcp add annas-mcp \
--env ANNAS_ACCOUNT_COOKIE=your-cookie \
--env ANNAS_SECRET_KEY=your-api-key \
--env ANNAS_DOWNLOAD_PATH=/absolute/path/to/downloads \
-- npx -y annas-mcpOn native Windows (not WSL), end the command with -- cmd /c npx -y annas-mcp.
Claude Desktop, one click. Download annas-mcp.mcpb and open it, or drag it onto Settings → Extensions. Claude Desktop asks for the cookie, API key and download folder.
Claude Desktop and Cursor, by hand. Add this to Claude Desktop's config (Settings → Developer → Edit Config) or to Cursor's ~/.cursor/mcp.json:
{
"mcpServers": {
"annas-mcp": {
"command": "npx",
"args": ["-y", "annas-mcp"],
"env": {
"ANNAS_ACCOUNT_COOKIE": "your-cookie",
"ANNAS_SECRET_KEY": "your-api-key",
"ANNAS_DOWNLOAD_PATH": "/absolute/path/to/downloads"
}
}
}
}Codex. Add this to ~/.codex/config.toml:
[mcp_servers.annas-mcp]
command = "npx"
args = ["-y", "annas-mcp"]
[mcp_servers.annas-mcp.env]
ANNAS_ACCOUNT_COOKIE = "your-cookie"
ANNAS_SECRET_KEY = "your-api-key"
ANNAS_DOWNLOAD_PATH = "/absolute/path/to/downloads"Pick a permanent ANNAS_DOWNLOAD_PATH. Folders under /tmp are cleared by the operating system.
3. Try it
Restart your client and ask something like "Find the EPUB of Pride and Prejudice and download it." The AI searches, picks a copy and returns the path of the saved file.
npx downloads the latest release on first run, verifies its checksum and caches it. Later runs start from the cache and pick up new releases automatically.
Related MCP server: annas-archive-download-mcp
Renewing the cookie
The aa_account_id2 cookie expires every week. When searches start failing with [UPSTREAM_BLOCKED]:
Copy a fresh
aa_account_id2value from your browser, as in step 1.Replace
ANNAS_ACCOUNT_COOKIEin your client's config.Restart the MCP server from your client.
ANNAS_ACCOUNT_COOKIE accepts the bare value, aa_account_id2=… or a whole Cookie header copied from the Network tab. Only aa_account_id2 is ever sent. Keep the cookie and API key private.
Configuration
Variable | Needed for | Description |
| Everything | Your |
| Downloads | Your Anna's Archive API key. |
| Downloads | Absolute folder for downloaded files. Created if missing. |
| Optional | Mirror to fall back on when automatic selection fails. Defaults to |
| Optional | Set to |
| Optional | Where |
The server picks a working mirror by itself, and only among the official domains annas-archive.gl, .pk and .gd. Whatever host you set in ANNAS_BASE_URL receives your credentials, so only use one you trust.
The server and CLI also read a .env file from their working directory. Variables already set in the environment take precedence.
Tools
Every tool accepts timeout_seconds, covering the whole operation including retries. Searches default to 60 seconds and downloads to 30 minutes.
Searching
book_search and article_search take a query and return one page of results. Each result includes a hash, title, authors, format, size and language, and books also include a publisher. Optional fields:
Field | Default | Meaning |
| 10 | Results to return from the page. Raise this before asking for another page. |
| 1 | Result page. |
| Any | Two-letter code such as |
| Any |
|
Give article_search a DOI (10.1038/nature14539, doi:… or https://doi.org/…) to get that one article instead of a page of results.
Downloading
book_download takes the hash and a title. article_download takes either a doi or a hash, plus an optional title. Both return the saved file's path and size in bytes.
Filename.
titlebecomes the filename. Passformat(such asepuborpdf) to set the extension; this doesn't convert the file. Without it, the extension is detected from the download.No overwrites. An existing file is never overwritten: a name clash gets a short suffix.
Integrity. Files are checked against the archive's MD5 hash, and incomplete or corrupt downloads are deleted.
DOI fallback. When the fast download servers don't have a paper,
article_downloadby DOI falls back to the PDF on Anna's SciDB page.
Troubleshooting
Errors start with a code:
Code | What to do |
| A variable is missing or invalid. Check Configuration. |
| The archive refused access. Usually the cookie has expired: renew it. Also check your membership tier. |
| Nothing matched, or that copy has no fast download. Try another query, or another copy from the search results. |
| Fix the hash, DOI, format, page, limit or timeout. |
| Retry, or raise |
| The archive or a download server failed. Check your daily download quota and API key, then retry. |
A failed download lists every server it tried and why each one failed. A server starting up cleanly doesn't prove your credentials work; only a search or download tests them.
If the server doesn't start at all, run npx -y annas-mcp --version in a terminal to see the error. The launcher unpacks releases with the system's tar, which comes preinstalled on macOS, Windows 10 and later, and mainstream Linux distributions.
Command line
Every tool is also a command. Set the same variables (or put them in a .env file), then:
npx -y annas-mcp book-search "pride and prejudice" --language en
npx -y annas-mcp article-search 10.1038/nature14539
npx -y annas-mcp book-download <hash> "Pride and Prejudice.epub"
npx -y annas-mcp article-download 10.1038/nature14539Add --json for machine-readable output, --timeout 10m to change the timeout, and --help to any command for all of its options.
Privacy
annas-mcp runs on your machine and has no telemetry. It sends your cookie and API key only to Anna's Archive, and downloads files from the servers Anna's Archive points it to. DOI lookups may contact doi.org. The npx launcher and the MCP Bundle contact GitHub to fetch and verify release binaries. Downloaded files stay in ANNAS_DOWNLOAD_PATH; nothing else is stored apart from the cached binary. Questions: open an issue.
Contributing
See CONTRIBUTING.md for building from source, tests and releases.
Credits
The idea for this project comes from iosifache/annas-mcp, which first exposed Anna's Archive as an MCP server. This repository is a complete rewrite and shares no code with it.
License
Available Tools
4 toolsarticle_downloadDownload articleA
Download a selected paper by exactly one of doi or hash. Saves the file under ANNAS_DOWNLOAD_PATH and returns path and bytes. Optional title and format control its filename.
| Name | Required | Description | Default |
|---|---|---|---|
| doi | No | DOI, doi: identifier, or https://doi.org/ URL. Supply exactly one of doi or hash | |
| hash | No | 32-character MD5 hash from search. Supply exactly one of doi or hash | |
| title | No | Optional title override for the saved filename | |
| format | No | Optional filename extension such as pdf. Inferred when omitted; does not convert the file | |
| timeout_seconds | No | Total operation timeout in seconds, including retries. Defaults to 1800; maximum 86400 |
Output Schema
| Name | Required | Description |
|---|---|---|
| path | Yes | |
| bytes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnly=false, openWorld=true, idempotent=false, destructive=false, so the write/external-network profile is covered. The description adds genuinely useful context beyond that: the file is written under ANNAS_DOWNLOAD_PATH and the call returns path and bytes. It omits auth requirements and the long-running nature implied by the 1800s default timeout.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no waste, front-loaded with the action and the selection constraint before the side effects. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return values need not be expanded, and annotations cover the safety profile. The description is adequate for invocation, though it could have flagged prerequisites (needing a hash from search) or the potentially very long runtime for a network-bound download.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter—including the exclusivity rule for doi/hash, the filename role of title/format, and timeout_seconds semantics—is already documented. The description restates the doi/hash rule and filename roles without adding syntax or behavior the schema lacks, matching the baseline of 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (download) and resource (a paper/article) plus the identifier constraint that selects it (exactly one of doi or hash). The article vs book distinction from siblings (book_download, article_search) is immediately legible, and the side effect (saves file, returns path and bytes) is named up front.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'exactly one of doi or hash' constraint tells the agent what inputs are valid, and 'a selected paper' implies a prior article_search is needed to obtain a hash. It never explicitly names article_search as the prerequisite or states when not to use it, so conditional guidance is inferential rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
article_searchSearch articlesARead-only
Find a paper by DOI (including DOI URLs), or search journal articles by keywords. DOI lookup returns one paper; keyword search returns one page. Compare metadata before downloading by hash or DOI.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Result page starting at 1. Defaults to 1 | |
| limit | No | Maximum results from this page. Defaults to 10; raise this before requesting another page | |
| query | Yes | Title, author, topic, or (for article_search) a DOI, doi: identifier, or https://doi.org/ URL | |
| content | No | Optional content filter: book_fiction, book_nonfiction, book_unknown, book_comic, magazine, or standards_document | |
| language | No | Optional two-letter ISO 639-1 language code, for example en | |
| timeout_seconds | No | Total operation timeout in seconds, including retries. Defaults to 60; maximum 86400 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, open-world, non-destructive behavior. The description adds useful behavioral context beyond annotations: DOI lookup returns exactly one paper while keyword search returns one page, and metadata is available for comparison before download. It does not cover pagination limits or auth, but the annotations lower the bar.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose, then result cardinality, then a workflow tip. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no output schema, the description adequately covers the two retrieval modes and result cardinality, and hints at available metadata. It omits detail about what metadata fields are returned or how pagination interacts with keyword search, but the rich input schema and annotations fill most gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters, including DOI/URL forms for query and pagination defaults. The description's mention of DOI including DOI URLs repeats the schema without adding new semantic detail, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find/search), resource (paper, journal articles), and two distinct modes (DOI lookup vs keyword search). It implicitly distinguishes itself from sibling tools: journal articles/paper vs books, and 'before downloading' points to the download tools. An agent can tell what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly explains the two usage modes and their differing result cardinality (one paper vs one page), and suggests comparing metadata before downloading. However, it does not explicitly name article_download or state when to prefer this over book_search, leaving that inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
book_downloadDownload bookA
Download the selected book using its hash from search. Saves a verified file under ANNAS_DOWNLOAD_PATH and returns path and bytes. Format sets the extension; it does not convert the file.
| Name | Required | Description | Default |
|---|---|---|---|
| hash | Yes | 32-character MD5 hash from search | |
| title | Yes | Book title used for the filename | |
| format | No | Optional filename extension such as pdf or epub. Inferred when omitted; does not convert the file | |
| timeout_seconds | No | Total operation timeout in seconds, including retries. Defaults to 1800; maximum 86400 |
Output Schema
| Name | Required | Description |
|---|---|---|
| path | Yes | |
| bytes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent, open-world operation, so the safety profile is covered. The description adds genuinely new behavior: the file is saved to ANNAS_DOWNLOAD_PATH, it is verified, and the call returns path plus bytes. It omits failure/retry behavior, though timeout handling is documented in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, with the core action and its precondition front-loaded and the format caveat last. Every sentence carries information an agent needs before invoking the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema covering the returned path and byte count and annotations covering the safety profile, the description is close to complete: it explains the save location, verification, and non-converting format. It could note permission or rate-limit expectations for an open-world download, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the four parameters is already documented with format, range, and defaults. The description reinforces that format sets only the extension and that hash comes from search, but that meaning is largely duplicated from the schema rather than extended, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource (download a book) and ties the input to its source ('hash from search'), which cleanly separates it from book_search and from the sibling article_download. It stops short of explicitly contrasting itself with article_download, but the 'book' framing makes the distinction inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the precondition that the hash must come from a prior search, which routes the agent through book_search first, and adds the caveat that format only sets the extension rather than converting. It does not spell out when to prefer this over other download paths or exclusions, so it is clear context without full alternative coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
book_searchSearch booksARead-only
Search books, textbooks, manuals, and standards by title, author, or topic. Returns one page with hash, format, metadata, and description when available. Pass a selected hash to book_download.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Result page starting at 1. Defaults to 1 | |
| limit | No | Maximum results from this page. Defaults to 10; raise this before requesting another page | |
| query | Yes | Title, author, topic, or (for article_search) a DOI, doi: identifier, or https://doi.org/ URL | |
| content | No | Optional content filter: book_fiction, book_nonfiction, book_unknown, book_comic, magazine, or standards_document | |
| language | No | Optional two-letter ISO 639-1 language code, for example en | |
| timeout_seconds | No | Total operation timeout in seconds, including retries. Defaults to 60; maximum 86400 |
Output Schema
| Name | Required | Description |
|---|---|---|
| page | Yes | |
| index | No | |
| limit | Yes | |
| content | No | |
| matched | Yes | |
| results | Yes | |
| language | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds the pagination/hash workflow context, but it does not disclose rate limits, retry behavior, or result-quality caveats beyond what structured data provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the verb and resource, then the return shape, then the next-step handoff. No filler and no repetition that wastes the agent's attention.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description is not obliged to explain return values, and the hash-to-book_download handoff completes the task flow. Coverage is solid for a six-parameter search tool, with only the sibling-search distinction left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including page, limit, content, language, and timeout is already documented in the schema. The description restates the query fields ('by title, author, or topic') but adds no syntax or format detail beyond the schema's own text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search books, textbooks, manuals, and standards') and enumerates the queryable fields, so the agent knows exactly what this tool does. It does not, however, contrast itself with the sibling article_search, which is the one alternative an agent is most likely to confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies a clear workflow context: search here, then 'Pass a selected hash to book_download.' That is real routing guidance. It stops short of explicit when-not conditions or a direct comparison against article_search, so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
article_download - First observed
article_search - First observed
book_download - First observed
book_search
TDQS
Scored across 4 tools
Each tool targets a distinct resource (article vs book) and action (search vs download), so the four tools form a clean 2x2 matrix with no overlap. An agent can unambiguously pick the right tool based on the resource type and whether it needs to find or fetch content.
All four names follow the same noun_action pattern (article_download, article_search, book_download, book_search), with consistent snake_case and identical verb vocabulary across both resources. This makes the set highly predictable and easy to reason about.
Four tools is exactly the minimal complete surface for a search-and-download server spanning two resource types. Each tool earns its place with no redundancy or filler.
The domain is search-and-fetch for articles and books, and both resource types have matching search and download operations, so the full lifecycle is covered. The search tools explicitly feed into the download tools via DOI/hash, leaving no dead ends.
Maintenance
Related MCP Connectors
Archive MCP — wraps the Internet Archive APIs (free, no auth)
MCP server for Russian books search, details, and recommendation candidates.
MCP server for Project Gutenberg — 75,000+ public-domain ebooks with full plain-text retrieval.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceSelf-hosted MCP server for searching and discovering books using Anna's Archive and Goodreads datasets, enabling full-text search, ISBN/md5 lookup, similarity matching, and optional download URL retrieval.12 npm4MIT
- FlicenseNot gradedqualityDmaintenanceMCP server for searching and downloading from Anna's Archive for free via slow channels like LibGen and Sci-Net, without an API key.2-
- AlicenseNot gradedqualityFmaintenanceMCP server for AI assistants to search and download ebooks from Z-Library and Anna's Archive, configured locally with automatic setup and download path.4AGPL 3.0
- AlicenseAqualityAmaintenanceA local MCP server for managing saved Substack posts. Enables offline reading, searching, bookmarking, and unbookmarking of Substack content via CLI or MCP clients.17MIT