openalex-mcp
This server provides eight MCP tools for querying the OpenAlex scholarly graph — searching works, authors, sources, and institutions, plus citation and author-output lookups.
Search scholarly works by query, filtered by year range, author, institution, source/journal, and open-access status, with pagination and sorting.
Fetch a single work by OpenAlex ID, DOI, or PMID, returning metadata, abstract, citation count, open-access links, and topics.
Search authors by name, optionally within an institution, to find disambiguated profiles.
Fetch an author profile by OpenAlex ID or ORCID, with affiliations, work counts, citation counts, and h-index.
Search sources such as journals, repositories, and conferences, with optional type filtering.
Search institutions by name and optional country code to obtain IDs for filtering works.
Find citing works for a given work, enabling forward citation-chaining from a paper to its downstream impact.
List an author's works with year filters and sorting by citation count or publication date.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@openalex-mcpFind recent works about open access publishing"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
openalex-mcp
MCP stdio server for the OpenAlex API (global scholarly graph).
Data via OpenAlex, OurResearch.
What this is for
OpenAlex maps roughly 240 million works and the relations between them, and this is the instrument for questions about circulation rather than content.
Ask who cites a given work and from which countries and institutions; assemble an author's full output across the name variants that defeat a title search; find which institutions cluster around a problem; identify the journals and repositories where a field actually publishes. Open-access status and PDF links come back with each record.
For a historian, the useful move is tracing whether an argument travelled — out of its language, out of its discipline, out of the decade that produced it — which is a question a bibliographic catalogue cannot answer and a citation graph can.
Related MCP server: Crossref Academic MCP Server
What the receipts are for
A search you cannot re-run is a claim you cannot check. When a footnote rests on a database query, say that no article in this index uses a term before a certain year, the reader is asked to take the search on trust: which term, in which script, on what date, against which index and which version of it, and how far down the results the author went. Ordinary searching leaves none of that behind. This server leaves all of it. Every query-answering tool returns its envelope through the ledger, which appends one line to an append-only file: the term actually sent and its script, how the source matched it, how many records existed and how many came back, the diagnostics, the tool and its parameters, the server version, a timestamp, and the hash of the previous line. The hash makes the file a chain: a line cannot be altered, removed or reordered afterwards without the verifier saying so.
What that gives a researcher:
A citable search. Name the receipt in the footnote (session slug, server, date, line hash) and a reader can see exactly what was asked and run it again against the same version.
Negative findings that carry weight. "Not found" is evidence only if the search that produced it is on record, with its term, its script and its breadth.
A method section that writes itself.
openalex-mcp-ledgermanifest <folder>summarises every query a project made, by server, script and session: the disclosure a journal, a data-availability statement or a research-integrity review asks for.A record of AI-mediated research. When a model chose the term, the receipt shows the term it chose and what came back, which is the thing to disclose about work done with an assistant.
Nothing interpreted. The receipt is the source's own answer with credentials removed. The server does not summarise, rank or paraphrase, so the record is of the source, not of the tool.
Receipts are off until you name a folder (MCP_RECEIPT_DIR); each server then writes its own
<server>.jsonl inside it, and MCP_RECEIPT_SESSION stamps a project or article slug on every
line so one folder can serve several projects. openalex-mcp-ledger verify-dir <folder> checks the chains.
The mechanics, the variables and what the envelope says when nothing is deposited are in the
receipts section below.
Install
Three routes. All three give you the same server; pick by how much you want to see of it.
Python. The pip and source routes need Python 3.10 or later; 3.10, 3.12, 3.13 and 3.14 are tested in CI on Windows, macOS and Linux. The Claude Desktop bundle uses whichever of these is already installed, and has uv download one only if none is.
Getting Python
Every route needs Python 3.10 to 3.14. The Claude Desktop bundle uses one already on the machine
and has uv download one only if none is; the other routes also need the venv module, which the
official installers include.
Windows. Download the 64-bit installer from python.org/downloads and run it; tick "Add python.exe to PATH" on the first screen. Afterwards
py --version(the launcher the installer adds) orpython --versionin a new terminal should print 3.1x. If typingpythonopens the Microsoft Store instead, Windows has no Python yet: that Store page is a stub, and it is also what "'python' is not recognized" usually means.macOS. The python.org installer, or
brew install python@3.13with Homebrew. The/usr/bin/python3that Xcode's command-line tools provide may be older than 3.10;python3 --versionsays.Linux. Your distribution's package:
sudo apt install python3 python3-venvon Debian and Ubuntu,sudo dnf install python3on Fedora. Or let uv provide one (next line).Any platform, with uv. uv installs Python itself:
uv python install 3.13, thenuv venvor theuvxroute below.
One click: the Claude Desktop bundle
Download openalex-mcp-2.1.1.mcpb from the latest release and open it; Claude Desktop installs it. One bundle serves Windows, macOS (Apple Silicon and Intel) and Linux. Claude Desktop asks for OpenAlex API key and Contact email (legacy mailto) and a receipts folder at install time; the key is stored in the OS keychain.
The bundle carries the server's source and a lock file, nothing compiled, and needs no Python of its own: Claude Desktop runs it with uv, using a uv already on your PATH if there is one and otherwise the copy the app ships. On first launch uv uses a Python 3.10 or later already on the machine, downloading one only if there is none, and installs the locked libraries: roughly 40 MB, or 60 MB with an interpreter, which took 26 to 46 seconds on the author's connection; later launches take under a second. If the first launch is slow enough that Claude Desktop reports the server disconnected, restart the app: what uv already fetched is cached, and the second launch completes. Bundles before 2.1.0 vendored libraries compiled for CPython 3.12 only and failed on every other interpreter; see Troubleshooting.
From GitHub, pinned to a release
pip install "git+https://github.com/ckgerteis/openalex-mcp@v2.1.1"
# or, without an environment of your own:
uvx --from "git+https://github.com/ckgerteis/openalex-mcp@v2.1.1" openalex-mcpinstalls the openalex-mcp console script and openalex-mcp-ledger. The tag is the thing to cite; @main gets whatever is current. Then register it in Claude Desktop (below), or let install.py do that.
The whole family
pip install "git+https://github.com/ckgerteis/bibliograph-mcp@v1.0.3" && bibliograph installinstalls all six servers and registers them together — one receipts folder, credentials asked for once. See bibliograph-mcp. From a checkout of this repository, python install.py does the same for this server alone, python install.py --all for the six, on Windows, macOS and Linux; install.ps1 remains for Windows.
From source
python3 -m venv .venv
.venv/bin/pip install .On Windows:
py -3.11 -m venv .venv
.venv\Scripts\pip.exe install .Or straight from the repository, without cloning:
uvx --from "git+https://github.com/ckgerteis/openalex-mcp" openalex-mcpVerify the install:
.venv/bin/python -c "import openalex_mcp; print(openalex_mcp.__version__)"That fails loudly if the package or one of its vendored modules is missing. Do
not use openalex-mcp --help as the check: unknown arguments are ignored, the
server starts, reads end-of-input and exits 0, so it reports success whatever
the state of the code.
Installing more than this one
Six independent packages. None imports another, none depends on another, and
each installs and answers on its own — pip install . in this directory is a
complete install of this server and nothing else.
They do share three things: a response envelope, a query ledger, and — if you
run more than one — a receipts folder. install.ps1 is vendored byte-identical
into all six and handles that on Windows; install.py is its cross-platform port. Both install this server by default, because
cloning one repository is not a request for five more.
.\install.ps1 # this server
.\install.ps1 -All # all six
.\install.ps1 -Servers openalex,cinii # a chosen subsetNothing about where things go is decided for you. The script asks where to
install (the virtual environment Claude Desktop will be pointed at), which
folder receives the receipts, and which session slug to stamp on them,
offering a neutral suggestion for each that Enter accepts; run without a
terminal it does not guess, and stops unless --venv and --receipts-dir
(or --no-receipts; -VenvDir and -ReceiptsDir for install.ps1) say
so. Whatever subset you name is registered against one receipts folder, asked for
once. The script prefers a sibling checkout to the network, carries across
credentials already registered rather than asking again, leaves servers it was
not asked about alone, and stops rather than guessing where the servers already
registered disagree about the folder or the session slug. It also asserts that
ledger.py and mediation.py are byte-identical across everything it
installed, so two envelope versions cannot end up in one environment unnoticed.
Any other MCP client
Nothing here is specific to Claude. The server speaks the Model Context Protocol over stdio and nothing else: any client that can start a process and talk JSON-RPC to it (Claude Code, Cursor, VS Code and Continue, Zed, LibreChat, a script of your own using an MCP SDK) can use it. The Claude Desktop bundle and the installers are conveniences for one client; the server underneath is the same console script. Register it anywhere by giving the client the absolute path of the console script and, optionally, the environment:
{
"mcpServers": {
"openalex": {
"command": "/absolute/path/to/.venv/bin/openalex-mcp",
"env": {
"OPENALEX_API_KEY": "your key (optional)",
"MCP_RECEIPT_DIR": "/absolute/path/to/receipts",
"MCP_RECEIPT_SESSION": "project-or-article-slug"
}
}
}
}Claude Code takes the same thing on the command line:
claude mcp add openalex -- /absolute/path/to/.venv/bin/openalex-mcpOn Windows the path ends in \.venv\Scripts\openalex-mcp.exe. MCP_RECEIPT_DIR and MCP_RECEIPT_SESSION
are optional; without them the server runs and every envelope says RECEIPT_NOT_DEPOSITED. The
stdio transport is the only one: there is no HTTP endpoint to expose, and nothing to host.
Troubleshooting
"Server disconnected" is all Claude Desktop says when the server process exited before or during the handshake, whatever the reason. The reason is in the log:
Windows:
%APPDATA%\Claude\logs\mcp-server-<name>.log(the extension's display name, or the key undermcpServers), withmcp.logbeside it for the app's side of the conversation.macOS:
~/Library/Logs/Claude/mcp-server-<name>.logandmcp.log.Linux:
~/.config/Claude/logs/.
Read the last launch from the bottom up. Three shapes account for nearly every report:
A Python traceback ending in
ImportErrororModuleNotFoundError(for exampleNo module named 'pydantic_core._pydantic_core'). The interpreter started, the code was found, and a compiled library did not match that interpreter. This is what every bundle before 2.1.0 did on any Python other than 3.12. Install the current bundle, or use the pip route, which resolves wheels for the interpreter you install into.'python' is not recognized,spawn python ENOENT, or a line from the Microsoft Store: no interpreter was found on the PATH Claude Desktop constructs. Nothing of this server ran. The current bundle does not launchpythonat all; for the pip route, register the console script by absolute path as shown above.A line from uv (
error: ..., or a download that never finished): the current bundle's runtime could not build its environment, usually because the first launch had no network or ran past Claude Desktop's sixty-second limit. Restart the app; uv keeps what it fetched. A uv older than 0.5 cannot read the lock file; upgrade it or remove it so the app uses its own.
The bundle's own entry point writes one line naming the interpreter, its path and the supported range before re-raising an import failure, so a log from 2.1.0 onwards says which of these it is.
Tools
Tool | Purpose |
| Works (articles, books, datasets, theses) by term, with year, author, institution, source and open-access filters |
| One work by OpenAlex ID, DOI or PMID |
| Authors by name, optionally within an institution |
| One author by OpenAlex ID or ORCID |
| Journals, repositories and conferences by name |
| Institutions by name, optionally by country |
| Works citing a given work, most-cited first |
| An author's works, with year filter and sort |
All eight return one typed JSON response envelope — see Response format. (Releases before 2.0.0 returned formatted markdown text; that is a breaking change, not a formatting preference.)
Response format
Every tool returns one JSON response envelope, built by mediation.py and defined in response-schema.json. Schema version 2.3.0. The same module and schema are vendored byte-identically across the server family, so an envelope from one server can be read by a consumer written for another.
The envelope reports how the search was made, not only what it found:
searched_for— on search operations, the term actually sent, its detected script, and the matching mode, hoisted to the top of the envelope so a relaying client cannot drop it. Lookups (oa_get_work,oa_get_author) and identifier filters (oa_cited_by,oa_author_works) omit it: they were handed an identifier and chose no term.query—input_termsas supplied,normalizedas sent, and the detectedscript. The credential never entersparams.matching_mode—full_text_stemmedfor term searches: OpenAlex matches title, abstract and indexed full text with stemming, soresult.totalis a loose count and a high breadth is expected.filter_exactfor identifier filters;identifier_lookupfor single-record fetches.result.breadth—none,narrow(1–50),broad(51–1000),very_broad(>1000).items[]— the family's item shape. OpenAlex's ownlanguagefield decides which typed title slot a work's title lands in (ja,ko; Latin-script titles in any other language go toen). A CJK title OpenAlex marks neitherjanorko, or a han-only title with no language, is left untyped rather than guessed;extra.titlealways carries the text andextra.languagethe code. OpenAlex identifiers, citation counts, open-access flags, topics and the reconstructed abstract sit inextra; the DOI (bare, without thehttps://doi.org/prefix) inids.doi; the landing page inids.url_en; an open-access copy inids.fulltext_url. Author, source and institution records userecord_typeauthor,sourceandinstitutionwith their metrics inextra.receipt— an ISO 8601 timestamp, a SHA-256 over the normalised query and its parameters, and the DOIs returned. Works without a DOI are identified only inextra.openalex_id, which the receipt'sresult_idsdoes not yet read.attribution— the required credit line, in every response.
Diagnostic codes
Typed and closed. A diagnostic is never prose the client has to parse.
Code | Level | Meaning |
| info | Records returned; nothing to flag. |
| warning | No records for this term and filter set. Non-English titles are indexed as the publisher supplied them, so an English rendering of a Japanese or Korean title may not match. |
| warning | A lookup by identifier answered 404. |
| error | OpenAlex answered 429. Since 2026 OpenAlex meters keyless access by a per-IP daily budget as well as per-second rate; a key raises both. |
| error | The API answered, and answered with an error (or with a 200 that was not JSON). |
| error | The request did not complete. Kept distinct from |
| info | The response was not written to the query ledger, because no receipts destination is configured. |
| warning | A receipts destination is set, the write was attempted, and it did not land. |
Configuration
OPENALEX_API_KEY=your_openalex_api_key
OPENALEX_EMAIL=your_email # legacy; see belowOpenAlex retired the polite pool on 13 February 2026 and replaced it with an API
key regime; the mailto parameter it depended on is now ignored. OPENALEX_API_KEY
is the access route. OPENALEX_EMAIL is kept only as a fallback for anyone running
against a mirror that still honours mailto, and sends nothing OpenAlex reads.
Claude Desktop
Add an entry to %APPDATA%\Claude\claude_desktop_config.json under
mcpServers, pointing at the console script in the environment you installed
into. On macOS or Linux use the absolute path to .venv/bin/openalex-mcp.
{
"mcpServers": {
"openalex": {
"command": "C:\\path\\to\\.venv\\Scripts\\openalex-mcp.exe",
"env": {
"OPENALEX_API_KEY": "your_openalex_api_key"
}
}
}
}Changed in 2.0.0. Tools return the JSON envelope rather than markdown; any consumer that parsed the 1.x text must be rewritten.
Changed in 1.1.0. Earlier versions were registered by path —
"command": "…\\python.exe", "args": ["…\\server.py"]. That entry will not
start this version, because server.py is now a module inside a package rather
than a script beside its imports. Replace it with the console script above.
Restart Claude Desktop. The eight tools should appear under "openalex" in the tool list.
Query receipts
Every envelope can be deposited to an append-only, hash-chained JSONL log by
openalex_mcp.ledger. Since 2.0.0 the envelope says whether that happened: RECEIPT_NOT_DEPOSITED when no destination is set, RECEIPT_WRITE_FAILED when one is set and the write did not land. It is off unless MCP_RECEIPT_DIR (or the legacy MCP_RECEIPT_LOG) is set, and a
logging failure is swallowed rather than raised — a search matters more than
the record of it. Secrets are redacted before a line is composed.
MCP_RECEIPT_DIR=C:\path\to\receipts # a folder, not a file
MCP_RECEIPT_SESSION=project-or-article-slug
MCP_RECEIPT_STRICT=1 # optional: make logging failure raise
MCP_RECEIPT_LOG=C:\path\to\receipts.jsonl # legacy single file; ignored when _DIR is setA folder, and one file per server. MCP_RECEIPT_DIR points at a directory
and each server writes its own <server>.jsonl inside it. That is not tidiness.
Appending is read-the-last-hash-then-write, and the lock around it is a threading
lock, which holds within one process and not between several — six servers are
six processes, and two answering at the same moment will both read the same
predecessor and both claim it. Measured, not theorised: six processes writing 150
lines to one file produced fourteen forks. MCP_RECEIPT_LOG still works and is
still correct for a single server; it is the wrong shape for a family.
install.ps1 sets this up for all six and writes a README into the folder.
Verify one chain, or the whole folder:
openalex-mcp-ledger verify receipts/openalex.jsonl
openalex-mcp-ledger verify-dir receipts
openalex-mcp-ledger manifest receipts # writes receipts/manifest.jsonverify exits non-zero on failure and says which kind it found: a fork
(concurrent writers — a configuration fault, and every line is still there), a
missing line, a reordering, or tamper (a line that does not hash to
its own content). Only the last is a claim about honesty, and reporting them
alike would invite a reader to mistake one for the other. The manifest is the
object to cite: one description of the whole deposit — per-file line counts,
first and last timestamps, terminal hashes, and combined totals by server,
script and session.
Tests
.venv/bin/pip install pytest jsonschema
.venv/bin/python -m pytest -q testsThe suite runs against recorded OpenAlex responses under tests/fixtures/ (captured 2026-09-04) and validates every envelope against response-schema.json; it needs no network and no key. RUN_LIVE=1 adds one request to the live API.
MCP SDK compatibility
Runs on both mcp 1.x and 2.x. Version 2.0.0 of the SDK removed
mcp.server.fastmcp; this server imports FastMCP where it exists and falls
back to MCPServer where it does not.
License
MIT © 2026 Christopher Gerteis. Covers the server code only; it grants no
rights over OpenAlex, OurResearch data, which remains governed by that provider's
terms. OpenAlex releases its data under CC0,
so no attribution is legally required; OurResearch asks for a credit and a link where the data is
shown, and the attribution line in every envelope carries one.
Author
Dr Christopher Gerteis, SOAS University of London.
Available Tools
8 toolsoa_author_worksARead-onlyIdempotent
Works by one author, with optional year filter and sort. Returns the unified envelope. A filter on an identifier, not a term search.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the tool read-only, idempotent, and non-destructive. The description adds useful behavior beyond that—identifier filtering rather than term search and a unified envelope return—but it does not describe pagination behavior, empty-result behavior, or any special handling. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying distinct information: scope, optional filters/sort, return envelope, and the important contrast with term search. No filler or schema repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list-by-author tool whose safety profile is covered by annotations and whose return stucture is covered by the output schema, the description is nearly complete. It could be slightly stronger by naming pagination defaults or explicitly directing term queries to oa_search_works, but nothing essential to invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The visible input schema already documents author_id (OpenAlex ID or ORCID), sort_by, year_from, and year_to with descriptions; the tool description only summarizes these as 'identifier', 'optional year filter', and 'sort'. That adds useful high-level grouping but not extra per-parameter meaning, so a baseline score is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear retrieval action—works by one author—and narrows it with optional year filtering and sorting. The closing clause, 'A filter on an identifier, not a term search,' differentiates it from oa_search_works, so an agent can tell this tool from its siblings without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly signals the intended context: use when an author identifier is known and works by that author are wanted. 'Not a term search' is an explicit exclusion, but the description could have named the sibling search tools as alternatives, so it stops short of full when-to-use/when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oa_cited_byARead-onlyIdempotent
Works that cite a given work, most-cited first. Returns the unified envelope. This is a filter on an identifier, not a term search, so no searched_for headline is set.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation safe, read-only, idempotent, and non-destructive. The description adds useful behavioral details beyond those: results are ordered most-cited first, the response uses the unified envelope, and no searched_for headline is set. These additions give the agent practical expectations without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each contributing distinct information: the result set and ordering, the return envelope, and the identifier-filter caveat. There is no filler, and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and safety annotations covering read-only and idempotent behavior, the description covers the essential remaining details: ordering, envelope, and headline behavior. The only minor gap is not explicitly pointing to a search tool to use when no identifier is available, but that is a small omission given the sibling set and clear identifier framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description clarifies that the key input is an identifier rather than a search term, which complements the schema's 'OpenAlex work ID or DOI' note. However, it does little to explain page or per_page behavior beyond what the schema defaults already convey. The added semantic value is modest but real.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation: return works that cite a given work, ordered most-cited first. It explicitly distinguishes itself from term search by framing the tool as a filter on an identifier, which separates it from the search siblings. The purpose is specific, actionable, and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: when you already have a work identifier and need the works citing it. It also warns that this is not a term search and that no searched_for headline is set, which helps prevent misuse. It does not explicitly name alternative tools, but the identifier-vs-term contrast is strong enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oa_get_authorARead-onlyIdempotent
Look up one author by OpenAlex ID or ORCID. Returns the unified envelope with a single author item, or NOT_FOUND.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation safe, read-only, and idempotent. The description adds useful behavior beyond that: it returns the unified envelope with a single author item or NOT_FOUND, which helps an agent anticipate the response shape and failure mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The action, identifier types, and expected outcome (single item or NOT_FOUND) are all front-loaded and immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lookup tool with rich annotations and an output schema present, the description is complete. It covers what the tool takes, what it returns, and the main edge case (NOT_FOUND), with nothing critical missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single author_id parameter is already documented in the schema with an explicit format and example. The tool description restates that OpenAlex IDs or ORCIDs are accepted but adds no new constraints or format details, so it is adequate but not enhancing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('look up'), a specific resource ('one author'), and the accepted identifier types (OpenAlex ID or ORCID). The singular 'one author' plus the NOT_FOUND outcome clearly distinguishes this from the search-oriented sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The exact-lookup wording makes it clear this tool is for retrieving a single known author rather than searching, and the return of NOT_FOUND reinforces that it is a direct fetch. It does not explicitly name alternatives like oa_search_authors as the fallback for unknown identifiers, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oa_get_workARead-onlyIdempotent
Look up one work by OpenAlex ID, DOI, or PMID. Returns the unified envelope with a single item, or NOT_FOUND.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, open-world, and non-destructive. The description adds useful runtime behavior beyond annotations: it returns the unified envelope with a single item, or NOT_FOUND. This helps the agent anticipate the response shape and missing-record handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler: the first states the action and accepted inputs, the second states the response behavior. The most important information is front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter lookup tool with rich annotations and an output schema, the description is complete. It covers what the tool does, what inputs are accepted, what the response envelope looks like, and the not-found behavior. No critical operational information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage reported at 0%, the tool description compensates by explicitly naming all three accepted forms for work_id: OpenAlex ID, DOI, and PMID. This fully specifies the parameter's core semantics, though it does not elaborate on formatting rules such as required DOI URL structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Look up'), a specific resource ('one work'), and enumerates the accepted identifier types (OpenAlex ID, DOI, PMID). It clearly distinguishes this tool from search-oriented siblings like oa_search_works and from oa_get_author by focusing on a single work lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the tool should be used when the agent already has a known work identifier. It does not explicitly name alternatives or state when-not-to-use conditions, so it stops short of full routing guidance, but the context is sufficiently clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oa_search_authorsARead-onlyIdempotent
Search OpenAlex for authors by name. Returns the unified envelope; each item is an author record (record_type author) with affiliations, work and citation counts, and h-index in extra.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as read-only, idempotent, and non-destructive. The description adds useful behavioral detail about the unified envelope and per-author fields such as affiliations, work/citation counts, and h-index in extra. This goes beyond annotations and helps set expectations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no waste. It front-loads the core purpose and then provides the key output detail. Every sentence contributes value without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the annotations cover safety and idempotency, the input schema documents parameters, and an output schema exists, the description provides enough high-level context. The mention of the unified envelope and extra fields helps the agent interpret results. It falls just short of perfect completeness by not steering toward sibling tools, but that gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage reported at 0%, the description must compensate by explaining parameters. It only clarifies the query is an author name; it does not address page, per_page, or institution_id filtering. The nested schema contains some property descriptions, but the description itself adds minimal parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific verb 'Search' with the resource 'OpenAlex' and the scope 'authors by name', clearly distinguishing it from sibling search tools for works, sources, and institutions. It adds that results are author records with affiliations, counts, and h-index, leaving no ambiguity about what the tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when an agent needs to find authors by name, but it does not explicitly state when to prefer this over alternatives like oa_get_author for known IDs or oa_author_works for an author's works. Context is present, but no exclusions or routing guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oa_search_institutionsARead-onlyIdempotent
Search OpenAlex for institutions. Returns the unified envelope (record_type institution); useful for finding institution IDs to filter work searches.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation as read-only, open-world, and idempotent. The description adds a return-shape detail ('unified envelope (record_type institution)') that annotations don't carry, but this phrase is jargon and no behavior like pagination defaults or rate limits is disclosed. It adds some value but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core action front-loaded. The second sentence gives a concrete use case, though 'unified envelope' is cryptic and slightly weakens clarity. Overall there is no redundant text and the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only search endpoint with rich annotations and an output schema, the description supplies the key missing 'why' (finding institution IDs) and identifies the record type. It doesn't need to explain return values because an output schema exists; the only notable gap is lack of explicit sibling routing, which is partly covered by the stated use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no parameter-level guidance: it doesn't mention country_code, page, per_page, or the query string semantics beyond what is already in the schema. Since schema description coverage is effectively low and the description does not compensate for page/per_page, an agent must rely entirely on field names and the sparse schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Search') and resource ('OpenAlex institutions'), and the phrase 'useful for finding institution IDs to filter work searches' distinguishes it from sibling search tools like oa_search_works, oa_search_authors, and oa_search_sources. The agent can clearly identify what this tool does and when it is relevant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear context: use this tool to find institution IDs for filtering work searches. It does not explicitly list alternatives or when-not-to-use cases, but the intended scenario is unambiguous and practical.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oa_search_sourcesARead-onlyIdempotent
Search OpenAlex for sources: journals, repositories, conferences. Returns the unified envelope (record_type source).
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful extra context by disclosing the unified envelope return shape and the record_type of 'source', which are not visible in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the verb, resource, and scoping entity types appear in the first sentence, and the return-envelope note is a single efficient second sentence with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and rich annotations, the description is mostly adequate for a simple search tool. However, it omits guidance on query behavior and pagination, and it only indirectly documents the 'type' filter, leaving some ambiguity for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate for undocumented parameters. It only hints at the 'type' filter by listing journal/repository/conference; it does not explain the required 'query' parameter semantics or the 'page' pagination parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search'), a specific resource ('OpenAlex sources'), and clarifies the entity types covered: journals, repositories, conferences. This clearly distinguishes it from sibling tools like oa_search_works and oa_search_authors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended domain is clear: use this tool when searching OpenAlex source entities. However, the description does not explicitly say when to prefer this tool over alternatives or provide any exclusion guidance, so usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oa_search_worksARead-onlyIdempotent
Search OpenAlex for scholarly works (articles, books, datasets, theses). Returns the unified envelope.
OpenAlex matches title, abstract and indexed full text with stemming, so result.total is a loose count (matching_mode full_text_stemmed) and a high breadth is expected. Titles are typed by OpenAlex's own language field; where that is absent, a CJK title is kept only in extra.title rather than guessed into ja or ko. Filters by year, author, institution, source and open-access status narrow the set exactly.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description discloses concrete behavioral traits: stemming and full-text matching, a loose count (matching_mode full_text_stemmed), and the conservative CJK-title handling that avoids guessing a language. It also warns that high breadth is expected and that filters narrow the set exactly. This materially informs how an agent should interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense, front-loaded paragraphs deliver purpose, scope, and behavioral caveats with no filler. The first sentence alone gives enough to route the call; each subsequent sentence adds unique context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With safety covered by annotations and returns covered by an output schema, the description provides the remaining operational knowledge: matching semantics, count looseness, and exactness of filters. For a search tool of this complexity, nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions filters for year, author, institution, source, and open-access status and states they narrow results exactly, which adds a little beyond bare parameter names. However, query, pagination, sort, and ID format details are left to the schema (which already documents them). The description doesn't fully compensate for any schema-description gap, but the addition of the exact filtering behavior warrants a middle score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Search OpenAlex for scholarly works') and enumerates the covered document types (articles, books, datasets, theses). This clearly differentiates it from sibling tools like oa_get_work (single-work lookup) and oa_search_authors/oa_search_sources, so an agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for query-based scholarly search and explains the effect of filters, but it never explicitly contrasts itself with alternatives. It does not state when to prefer oa_author_works or oa_cited_by instead, nor provide when-not-to-use guidance. This leaves usage routing to inference rather than explicit direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v1.1.0- First observed
oa_author_works - First observed
oa_cited_by - First observed
oa_get_author - First observed
oa_get_work - First observed
oa_search_authors - First observed
oa_search_institutions - First observed
oa_search_sources - First observed
oa_search_works
TDQS
Scored across 8 tools
Each tool targets a distinct entity or relationship: get vs search for authors/works, plus dedicated searches for sources and institutions, and two relationship-specific work queries (cited_by, author_works). The only slight overlap is that oa_search_works, oa_cited_by, and oa_author_works all return works, but their purposes are clearly separated by identifier filters versus term search.
All tools share the oa_ prefix and use a readable snake_case pattern: get_<entity> for lookups and search_<entity> for searches. oa_cited_by and oa_author_works break the simple verb_noun pattern but are still intuitive and consistent with the relationship-focused actions.
Eight tools is a well-scoped set for a read-only scholarly API wrapper. Each tool covers a meaningful query type without redundancy or bloat.
The core OpenAlex surface is well covered: works and authors have both exact lookup and search, with sources and institutions searchable as supporting entities. Minor gaps exist such as no get-by-ID for sources/institutions and no coverage of concepts, publishers, or funders, but the primary scholarly workflows are complete.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
OpenAlex MCP — wraps the OpenAlex API (scholarly works, free, no auth)
MCP server for Altmetric APIs - track research attention across news, policy, social media, and more
Scholarly search: OpenAlex, Crossref, arXiv, OpenCitations and PubMed in one endpoint.
Auditable MCP server for PubMed, Europe PMC, ClinicalTrials.gov, and bioRxiv/medRxiv queries
Related MCP Servers
- AlicenseAqualityDmaintenanceMCP server for verifying academic citations via Semantic Scholar, OpenAlex, and CrossRef.1MIT
- AlicenseAqualityDmaintenanceMCP server enabling AI agents to search and retrieve scientific papers, citations, and author profiles from Crossref, OpenAlex, and Semantic Scholar with no API keys required.53MIT
- AlicenseAqualityDmaintenanceMCP server for the OpenAlex scholarly database, providing AI agents with tools to search and retrieve academic works, authors, and institutions via natural language queries.8MIT
- FlicenseNot gradedqualityDmaintenanceMCP server for academic research using the OpenAlex API, enabling article search, details retrieval, and author profile lookup.-