:
This server lets you query OpenEvidence (a medical evidence platform) from any MCP client using your existing logged-in browser session — no API key required.
Core Capabilities
Ask medical questions (
oe_ask): Submit questions in fire-and-forget mode (returns immediately with a pendingarticle_id) or blocking mode (wait_for_completion: true); supports follow-up questions viaoriginal_article_idRetrieve answers (
oe_article_get): Fetch completed answers by article ID, optionally blocking until ready; returns markdown answer, figures, BibTeX citations, and Crossref-validated referencesCheck auth status (
oe_auth_status): Verify your browser session is active and the relay is connectedBrowse question history (
oe_history_list): Paginate and search your OpenEvidence question history
Collections Management
List/get collections (
oe_collections_list,oe_collections_get): View all collections or fetch one with its nested article membershipsCreate collections (
oe_collections_create): Create new collections (agent-managed ones conventionally start with#)Add articles (
oe_collections_add_article): Assign a chat/article to a collection (idempotent)
Local SQLite Mirror & Auto-Sort
Initialize local DB (
oe_collections_db_init): Create a local SQLite mirror of your chat history and collectionsSync history/collections (
oe_collections_sync_history,oe_collections_sync_db): Pull full or incremental chat history and collection memberships into SQLiteList unsorted chats (
oe_collections_unsorted): Find chats not assigned to any#-prefixed collectionSummary stats (
oe_collections_summary): Get counts of chats, collections, memberships, unsorted items, and last sync timestampsAuto-classify chats (
oe_collections_classify): Predict hashtag collections for unsorted chats using log-odds-ratio signatures and keyword rules, returning a proposed planBulk apply classifications (
oe_collections_bulk_apply): Execute a classification plan — create missing#-collections and add memberships in bulk (idempotent)
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@:what's the evidence for statins in primary prevention?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
What it does
OpenEvidence protects its API with bot detection that blocks plain server requests. This project removes that problem: when your AI tool asks OpenEvidence a question, the request is run inside your own logged-in OpenEvidence browser tab — so it carries your genuine browser session and is never challenged.
A small Chromium browser extension lends its session to a localhost relay; the MCP server speaks to that relay. The extension is a generic authenticated fetch proxy — all the OpenEvidence logic stays in the local server, and your browser login is the only credential. No API key, no cookie file, no Playwright, no headless browser.
It is designed for local personal workflows where you already have lawful access to OpenEvidence. It does not bypass authentication, remove access controls, redistribute OpenEvidence content, or include any OpenEvidence data in this repository.
Related MCP server: kjlahsdjkashdjhkasdkajshd
How it works
Claude / Codex / any MCP client
│ asks a question (oe_ask)
▼
openevidence-mcp ──▶ relay daemon ──▶ browser extension ──▶ your logged-in
(stdio server) 127.0.0.1:8787 (runs the fetch) OpenEvidence tab
▲ shared · auto-spawned │
└──────────────────────── answer ◀───────────────────────────────┘Nothing navigates or pops up — the tab stays where it is. The extension only ever talks to openevidence.com and your local relay (127.0.0.1). The relay runs as a shared daemon that owns port 8787 and outlives every session, so any number of Claude/Codex sessions funnel through the one logged-in tab.
Quick start
git clone https://github.com/htlin222/openevidence-mcp.git
cd openevidence-mcp
make all # installs deps · builds the MCP server + relay extension ·
# registers the server into Claude and Codex (whichever CLI you have)Then the one manual step make all prints (a browser action that can't be scripted):
Load the extension — open
chrome://extensions(Chrome / Edge / Brave / Arc / Vivaldi / Opera) → turn on Developer mode → Load unpacked → selectextension/dist.Stay logged in to openevidence.com in that browser, and keep a tab open. That login is your authentication.
Run your AI tool. The server auto-starts the relay and connects to the extension.
Verify and go:
curl -s http://127.0.0.1:8787/health # expect {"ok":true,"connected":true,"version":1,...}Then just ask, in any MCP client: “Use OpenEvidence to answer …”. Re-run make all anytime to rebuild + re-register; make help lists every target; make kill-all stops all servers + the relay daemon.
Already installed? Update in one line:
make update # git pull latest release · rebuild server + extension · re-registerThen reload the browser extension (chrome://extensions → Reload) and reconnect /mcp. make status shows versions + live health; make uninstall removes it. (Inside an AI session you can also just say “update openevidence-mcp” — the bundled install skill runs the right step and reminds you to reload the extension.)
Clicking the extension's toolbar icon opens a built-in how-it-works page with a live connection check.
cookies.jsonand a HAR are optional — needed only for the legacyOE_MCP_RELAY_TRANSPORT=offcookie read path, the Python collections tooling, andnpm run doctor/login/smoke. See Optional cookie path.
One tab, many sessions
The relay is a standalone daemon (auto-spawned; run/inspect with npm run relay) that owns port 8787 and is shared by every MCP server. You can run OpenEvidence from any number of Claude/Codex sessions at once — they all flow through the single logged-in tab. The daemon:
outlives session restarts and is respawned automatically if it dies (a crashed daemon is replaced and the in-flight request retried);
is idempotent — a second daemon that loses the port race exits cleanly;
reports
version+pidon/health, so an upgrade replaces a stale daemon from an older build.
Asking questions
After registration, ask your MCP client in plain English and mention OpenEvidence — the agent calls oe_ask automatically.
Use OpenEvidence to answer: DLBCL frontline treatment landscape NCCN v3.2026. Include citations and BibTeX.
Use OpenEvidence to compare Pola-R-CHP vs R-CHOP in untreated DLBCL. Include trial citations and BibTeX.
Use OpenEvidence to review current evidence for SGLT2 inhibitors in HFpEF. Include citations and BibTeX.
Use OpenEvidence to find guideline-supported anticoagulation options for cancer-associated thrombosis.Fire-and-forget by default
oe_ask returns {article_id, status:"pending"} the moment the question is submitted, freeing the tab for other sessions instead of holding the call for the whole generation:
// oe_ask → returns instantly
{ "article_id": "…", "status": "pending" }
// oe_article_get(article_id) → the finished answer, when ready
{ "article_id": "…", "status": "success", "extracted_answer_raw": "…", "figures": […], "artifacts": { … } }Fetch the answer later with
oe_article_get(article_id)— passwait_for_completion: truethere to block until it's ready.For one-shot blocking (submit and wait in a single call), pass
wait_for_completion: truetooe_askitself.
A completed article returns the OpenEvidence payload and status, the article_id, the answer markdown as extracted_answer_raw, any figures, inline BibTeX as artifacts.bibtex, and saved citation files. Pass include_bibtex: false to keep the response small while still writing citations.bib to disk. Pass strip_citation_markers: true to also get extracted_answer_clean with the [1] / [1-3] reference marks removed — handy when quoting the answer into notes.
Question pacing — stay under your hourly ask limit
OpenEvidence caps how many questions you can ask per account over time. Only
oe_ask spends that budget — a new question and every follow-up is one ask.
Everything else (oe_article_get, oe_answers_search, oe_history_list,
oe_public_get, oe_auth_status, oe_health, collections) is free of it.
To avoid bursts that trip the limit, oe_ask paces submissions: it waits
(never errors) so consecutive asks are at least OE_MCP_ASK_MIN_INTERVAL_MS
apart (default 1000 ms ≈ 1 question/second). The wait is coordinated across all
MCP sessions through the shared SQLite ask_log, so several Claude/Codex windows
can't collectively exceed the rate. Every oe_ask reply carries
ask_pacing: { waited_ms, asks_last_hour }, and oe_health reports
asks_last_hour so you can see how many questions you've spent in the trailing
hour at a glance. Set OE_MCP_ASK_MIN_INTERVAL_MS=0 to disable pacing.
The biggest saver, though, is not re-asking: oe_answers_search finds a past
answer and oe_article_get serves it from cache — both cost zero questions.
Follow-up questions — continue the same conversation
Every completed answer carries follow_up_questions — the suggestions OpenEvidence renders under the answer (e.g. "How does adjuvant therapy choice differ in elderly patients?"). To ask any of them (or your own) in the same conversation thread, call oe_ask with original_article_id set to the answer's article_id:
// first answer → { article_id: "b752a1c1-…", follow_up_questions: ["How does adjuvant therapy choice differ in elderly patients?", …] }
// oe_ask({ question: "How does adjuvant therapy choice differ in elderly patients?",
// original_article_id: "b752a1c1-…" })
// → a new article that threads on the prior turn — the answer opens
// "Elderly patients with stage III colon cancer…", carrying the earlier context.You only pass original_article_id + the new question — OpenEvidence rebuilds the conversation history server-side (the new article's inputs.history is populated for you). Each follow-up answer comes with its own fresh follow_up_questions, so you can keep drilling down. original_article_id accepts a bare UUID or a full /ask/<id> URL.
Reading conversation pages by link
Someone sends you an OpenEvidence link? oe_public_get(url) fetches the server-rendered /ask/<id> page and parses it into Q&A turns (question, answer as markdown, references). Auth escalates automatically:
Relay connected → the fetch runs inside your logged-in tab, so your own private conversations work too (and DataDome never sees it).
No relay, public link → a plain anonymous fetch; conversations whose author pressed "make public" are fully server-rendered and need zero setup.
No relay, private link → falls back to
cookies.jsonwhen present; otherwise reports clearly that the conversation is private.
Page fetches count against the same account budget as API calls, so they run through the same rate limiter (60 clinical queries/min, self-throttled at 80%). oe_public_get also accepts a bare article UUID, and oe_article_get / oe_ask.original_article_id accept full /ask/ URLs too.
Local answer store — search past answers offline, skip re-fetches
Every completed answer from oe_ask / oe_article_get is upserted into a local SQLite table (answers) in the same file the collections tooling already uses (~/.openevidence-mcp/db/oe.sqlite). This gives you two things for free:
Full-text search over what you've already asked —
oe_answers_search("query")runs an FTS5 query across questions, titles, and answer bodies and returns highlighted»…«snippets. It's fully offline: no OpenEvidence traffic, no rate-limit cost. Great for "did I already look this up?" before spending a query.// oe_answers_search({ query: "CAPEOX duration neurotoxicity" }) // → { total_stored, match_count, matches: [{ article_id, title, question, snippet, url, … }] }The query is FTS5 syntax — plain words are ANDed,
"quoted phrases"match verbatim, andAND/OR/NOTwork. Malformed queries fall back to a literal token search instead of erroring.Cache hits on
oe_article_get— fetching an article you've already stored returns instantly from disk withfrom_cache: trueand zero network round-trips (it doesn't even touch the relay). Passrefresh: trueto force a re-fetch from OpenEvidence and regenerate the on-disk citation artifacts.
This store persists across reboots (unlike the citation artifacts under the OS temp dir), and only ever holds answers this MCP fetched. For your complete server-side history, use oe_history_list or the collections sync. Point it elsewhere with OE_MCP_DB_PATH; it needs Node ≥ 22 for the built-in node:sqlite (no extra dependency).
Is the pipeline up? oe_health vs oe_auth_status
Two health checks, two speeds:
oe_health— a millisecond-fast, purely local check of the relay pipeline. It reads the daemon's/healthand reports whether the daemon is up, whether the browser extension is actually polling (extension_connected), the version match, uptime, and served/errored counts — without any OpenEvidence network call. Use it to confirm the plumbing before an ask, or to diagnose a stuck relay. Each failure state carries ahint.oe_auth_status— the full round-trip: hits/api/auth/methrough the relay to confirm your login session is still valid. Slower, but it's the one that answers "am I actually logged in?".
Tools
Tool | Purpose |
| Ask a question — fire-and-forget by default (returns a pending |
| Fetch an article by id or |
| Read a conversation page from an |
| Toggle a conversation's share visibility — |
| Full-text search (SQLite FTS5) over every answer this MCP has fetched — offline, zero rate-limit cost; returns highlighted snippets |
| Millisecond-fast local check of the relay pipeline (daemon + extension) — no network call, unlike |
| Check |
| Read your OpenEvidence question history |
| List your collections |
| Get a collection (incl. nested |
| Create a collection (agent-managed names should start with |
| Add a chat to a collection |
| Create the local SQLite mirror (idempotent) |
| Pull |
| Refresh collections + memberships into SQLite |
| Chats with no |
| Counts + last sync timestamps |
| Auto-classify unsorted chats using log-odds-ratio signatures learned from your memberships + curated keyword rules |
| Mint missing |
Privacy & security
The relay listens on localhost only (
127.0.0.1:8787) — nothing is exposed to the network.The extension and server store no credentials. Your OpenEvidence session lives in your browser, as always.
The extension only acts on requests from your own local relay, and only against
openevidence.com.This repository contains connector code only — no OpenEvidence content, datasets, cookies, or account material.
Troubleshooting
Extension badge not green /
connected:false? Make sure you're logged in to openevidence.com in that browser with a tab open, then reload the extension (chrome://extensions→ ↻).Tools fail with “relay not connected”? Start your AI tool / MCP server so the daemon comes up, then
curl -s http://127.0.0.1:8787/health. If a previous build is stuck,make kill-alland reconnect the MCP server.Run the relay in one browser at a time if you've loaded the extension in several — requests go to whichever polls first.
DataDome 403 on the legacy cookie path? See Doctor below — relevant only when
OE_MCP_RELAY_TRANSPORT=off.
Register with MCP clients
make all already registers Claude Code and Codex CLI. To do it à la carte (or for other clients), register the local stdio server node /ABSOLUTE/PATH/openevidence-mcp/dist/server.js. No env is required — the browser extension is the login.
make install-claude-global # claude mcp add-json --scope user openevidence …
make install-codex-global # codex mcp add openevidence -- node dist/server.js
make install-agy-global # Antigravity CLI (agy-cli)
make install-all # all threeFor Claude Desktop, Cursor, Cline, Continue, use this mcpServers shape (see examples/ for a full-knobs reference):
{
"mcpServers": {
"openevidence": {
"command": "node",
"args": ["/ABSOLUTE/PATH/openevidence-mcp/dist/server.js"]
}
}
}Collections sync & auto-sort
scripts/collection_sort.py mirrors your chat history and collection memberships into a local SQLite (~/.openevidence-mcp/db/oe.sqlite; override with OE_MCP_DB_PATH). The routine routines/collection-sort.md walks an MCP client through syncing, surfacing unsorted chats, and applying multi-membership hashtag tags. Convention: collections whose name starts with # are agent-managed; collections without a leading hash are human-curated and never touched.
The same pipeline is exposed as the oe_collections_* MCP tools (the TS server shells out to scripts/collection_sort.py via python3, overridable with OE_MCP_PYTHON).
python scripts/collection_sort.py init
python scripts/collection_sort.py sync-history --full # first time
python scripts/collection_sort.py sync-collections
python scripts/collection_sort.py list-unsorted --json # routine reads this
python scripts/collection_sort.py summaryAccount-scoped (schema v2). Chats, collections, and memberships carry an account column (composite primary keys), so one DB can mirror multiple OpenEvidence logins without mixing them. The account is resolved from /api/auth/me; override with --account EMAIL. A v1 DB auto-migrates on first open — pre-v2 rows are tagged --legacy-account (default legacy; or OE_MCP_LEGACY_ACCOUNT).
Schedule the sync (macOS)
The classification step needs the agent in the loop, but the sync side is pure I/O — install a daily launchd job that keeps the local mirror fresh:
bash scripts/install_launchd.sh # daily 02:00 (OE_MCP_SYNC_HOUR / OE_MCP_SYNC_MINUTE)
launchctl start com.htlin.openevidence-mcp.sync # fire once to verify
tail -30 ~/.openevidence-mcp/logs/sync.log
bash scripts/install_launchd.sh --uninstall # removeThe wrapper (scripts/collection_sync_cron.sh) appends one block per run to ~/.openevidence-mcp/logs/sync.log. It takes an optional mode flag:
Mode | Behavior |
(default) | sync only — chats accumulate as |
| sync + classify; writes |
| sync + classify + bulk-apply + reconcile; fully autonomous sort |
scripts/classify.py runs offline (no API): a per-tag log-odds-ratio signature (Monroe et al. 2008) built from your memberships every run, OR'd with curated keyword rules. Validate with python scripts/classify.py validate. Tune headless use via OE_MCP_AUTO_THRESHOLD (default 12) and OE_MCP_AUTO_TOP_K (default 3); switch the job to autonomous mode with OE_MCP_SYNC_MODE=--auto bash scripts/install_launchd.sh.
Citation artifacts
Completed oe_ask / oe_article_get calls save artifacts under ${OE_MCP_ARTIFACT_DIR}/<article_id>/ (default OS temp dir + openevidence-mcp; on macOS, /tmp may resolve under /var/folders/.../T/):
File | Purpose |
| Extracted markdown answer |
| Full OpenEvidence article payload |
| Parsed structured citations |
| BibTeX bibliography |
| Post-hoc Crossref validation results |
Crossref validation: DOI citations are validated directly; non-DOI citations use a bibliographic query and are marked candidate / not_found / error. Low-similarity matches never overwrite BibTeX metadata, and sources like NCCN guidelines may stay as local OpenEvidence metadata when Crossref has no authoritative match.
These artifact files live under the OS temp dir and can be cleaned up by the system. The answer body itself is also persisted to the SQLite answers store (see Local answer store), which survives reboots and powers oe_answers_search and the oe_article_get cache.
Optional cookie path
The extension relay is the default and recommended path. For headless reads without the extension, set OE_MCP_RELAY_TRANSPORT=off to route reads (oe_history_list, oe_article_get, oe_collections_*) over a browser-exported cookies.json — asks still need the extension. This path also backs the Python collections tooling and the npm run doctor / login / smoke CLIs.
cp /path/to/browser-cookies.json ./cookies.json
make build HAR=/path/to/www.openevidence.com.har # extracts the browser fingerprint, then compiles
npm run login && npm run smokemake build extracts openevidence-fingerprint.json from the HAR when present. The fingerprint is profile-faithful — the client sends exactly the captured browser's header set, so a HAR from any browser (incl. Safari) stays coherent with the cookie that browser minted.
Doctor (legacy cookie path)
When the cookie path fails with a DataDome 403 — most often after moving to a different computer — run the doctor:
npm run doctor # static checks + a live read probe
npm run doctor -- --offline # static checks only (no network)
npm run doctor -- --json # machine-readable outputIt flags datadome-missing / -expired / -session (cookie absent / past expiry / session-scoped), fingerprint-platform-mismatch (cookie + fingerprint minted on a different OS — re-mint here), fingerprint-default (no fingerprint; built-in signature in use), and datadome-live (a live request was actually challenged). Non-zero exit on failure for CI/pre-flight.
Make targets
Run make help for the grouped, always-current list.
Target | Purpose |
| One-shot setup: deps + build server + build extension + register into Claude & Codex (skips a CLI that isn't installed) |
| Update to the latest release: |
| Versions, live relay |
| Reap orphan relay daemons + prune stale logs/temp — non-destructive (relay respawns on next use) |
| Unregister from the CLIs, stop daemons, remove |
| Grouped reference of every target |
| Stop all MCP servers + the relay daemon and free port 8787 |
| Run the standalone relay daemon in the foreground (debug) |
| Force install · type-check · unit tests · auth+history smoke (cookie path) |
| Extract the fingerprint if a HAR is given, then compile TypeScript |
| Extract the working browser fingerprint from a HAR |
| Import and verify cookies |
| Register the server with the respective CLI(s) |
Environment variables
Variable | Default | Purpose |
|
| Set |
|
|
|
|
| Relay port (must match the extension) |
|
| Relay daemon pidfile |
|
| Relay daemon log file |
|
| OpenEvidence base URL |
| OS temp dir + | Artifact output directory |
| unset | Optional Crossref polite-pool email |
|
| Set |
|
| Minimum spacing between questions ( |
|
| Poll interval when waiting for an answer |
|
| Default poll timeout |
|
| Cookie file (legacy/optional path) |
|
| Browser signature fingerprint (legacy path) |
|
| Local SQLite file: collections mirror + the |
|
| Python interpreter the bridge tools spawn |
|
| Label for pre-v2 rows on DB account migration |
Project files
extension/README.md — the browser-extension relay (and its built-in README.html how-it-works page)
README.AI.md — agent install playbook
src/server.ts — MCP tools
src/relay-server.ts — the localhost relay (extension-facing +
/relaybridge)src/relay-client.ts / src/relay-daemon.ts — shared-daemon client + standalone daemon
src/citations.ts — citation extraction, BibTeX, Crossref validation
src/doctor.ts — stale DataDome cookie diagnostics (legacy path)
docs/plans/2026-06-04-shared-relay-daemon-design.md — relay daemon design doc
examples/ — MCP client config samples
Copyright, trademark, and medical disclaimer
This project is unofficial and independent. It is not affiliated with, endorsed by, sponsored by, or approved by OpenEvidence or its owners. "OpenEvidence" and related names, logos, product names, and content remain the property of their respective owners.
This repository contains connector code only. It does not include OpenEvidence copyrighted content, proprietary datasets, model outputs, article payloads, session cookies, or account material. Your local use of this MCP server may create files such as answer.md, article.json, and citations.bib; those artifacts can contain content retrieved from or derived from your OpenEvidence account session. Treat those files as private unless you have the right to share them.
You are responsible for complying with OpenEvidence terms, institutional policies, copyright law, and any clinical data governance rules that apply to your use. Do not publish cookies, account tokens, saved article payloads, generated answers, screenshots, guideline text, or other protected/copyrighted content unless you have permission or another valid legal basis.
This software is not medical advice and is not a medical device. It is an integration tool for an MCP client. Clinicians and qualified users remain responsible for verifying outputs against authoritative sources and applying independent clinical judgment.
License and attribution
Apache-2.0. Keep LICENSE and NOTICE when redistributing.
Based on OpenEvidence MCP by Bakhtier Sizhaev: https://github.com/bakhtiersizhaev/openevidence-mcp
Available Tools
17 toolsoe_answers_searchSearch Stored Answers (local FTS)A
Full-text search (SQLite FTS5) over every answer previously fetched by oe_ask/oe_article_get — questions, titles, and answer bodies. Millisecond-fast and fully offline: no OpenEvidence traffic, no rate-limit cost. Covers only answers this MCP has fetched and stored locally; for your complete server-side history use oe_history_list. Snippets mark matches with »…«.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | FTS5 query — plain words, "quoted phrases", AND/OR/NOT. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses offline operation, local-only scope, no rate-limit cost, and the snippet marker format. Lacking are details on what happens if no results or an error occurs, but the provided behavioral notes are substantive for a read-only search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at three sentences, with the most critical information (purpose, technology, offline nature, scope, and limitation) presented upfront. Every sentence adds unique value, achieving high density without sacrifice of clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the moderate complexity of an FTS tool with two parameters and no output schema, the description covers the key aspects: what is searched, performance (millisecond-fast), offline property, and snippet format. Minor gaps exist (no mention of return structure or error handling), but overall completeness is high.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (only 'query' has a description). The description adds context about FTS5 syntax (plain words, phrases, AND/OR/NOT) and the fields searched, which extends what the schema provides. It implicitly describes the limit parameter's purpose but does not provide per-parameter enrichment beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as performing full-text search over locally stored answers from oe_ask/oe_article_get, specifying the searched fields (questions, titles, answer bodies). It distinguishes itself from the sibling tool oe_history_list by noting coverage of only local data, leaving no ambiguity about its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use (for fast offline local search) and explicitly advises against using it for complete server-side history, directing users to oe_history_list instead. However, it does not provide guidance on prerequisites or when this tool should be avoided entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_article_getOpenEvidence Article GetA
Fetch an article (answer) by id or /ask/ URL — the fetch-later half of fire-and-forget oe_ask. Returns the current status; if it is still 'pending' either retry later or pass wait_for_completion:true to block until the answer is ready. Completed answers are served from the local SQLite store when available (from_cache:true, zero network) — pass refresh:true to force a re-fetch from OpenEvidence.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | Bypass the local answers cache and re-fetch from OpenEvidence. | |
| article_id | Yes | Article UUID, or any openevidence.com/ask/<id> URL. | |
| timeout_sec | No | ||
| include_bibtex | No | ||
| save_artifacts | No | ||
| poll_interval_ms | No | ||
| crossref_validate | No | ||
| wait_for_completion | No | ||
| strip_citation_markers | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses caching behavior, pending status handling, blocking via wait_for_completion, and refresh semantics. It explains that completed answers come from local SQLite store with from_cache:true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single paragraph that efficiently conveys key information with good front-loading. Could be slightly more structured (e.g., bullet points) but is concise overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters and no output schema, the description covers core usage and behavioral nuances. Missing details on some parameters and output format, but adequate for a fetch-and-cache tool with sibling context provided externally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 22%, so the description must compensate. It explains article_id (UUID or URL), wait_for_completion, refresh, timeout_sec, poll_interval_ms, but omits include_bibtex, save_artifacts, crossref_validate, strip_citation_markers. Adds value but not enough for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches an article by id or URL, and explicitly positions it as the 'fetch-later half of fire-and-forget oe_ask', distinguishing it from siblings like oe_ask, oe_answers_search, and oe_article_set_access.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells when to use (after oe_ask) and provides options: retry, wait_for_completion, or refresh. It lacks explicit when-not-to-use statements but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_article_set_accessOpenEvidence Article Share AccessA
Set a conversation's share visibility (PATCH /api/article//access). public:true makes it link-shareable (ANYONE_WITH_LINK — anyone with the URL can read it, no login); public:false makes it private again (CREATOR_ONLY). Returns the shareable /ask/ URL. You must own the conversation, and the relay extension must be connected (only the owning browser session may change access). ⚠️ Publishing exposes the conversation to anyone on the internet with the link — do not publish anything containing PHI or medically sensitive patient information.
| Name | Required | Description | Default |
|---|---|---|---|
| public | Yes | true → ANYONE_WITH_LINK (public); false → CREATOR_ONLY (private). | |
| article_id | Yes | Article UUID or openevidence.com/ask/<id> URL of the conversation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully covers behavioral aspects: it specifies HTTP method (PATCH), explains the effect of true/false on access levels (ANYONE_WITH_LINK vs CREATOR_ONLY), notes prerequisites (ownership, relay connection), and mentions the return of a shareable URL and the security warning.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two clear sentences plus a warning, all front-loaded with the verb and resource. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description states the return value (shareable URL). It covers prerequisites, side effects, and security implications, making it complete for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds meaning beyond schema by explaining the access levels corresponding to boolean values and that article_id accepts both UUID and full URL. This enhances the agent's understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Set a conversation's share visibility', clearly stating the verb and resource. It distinguishes from siblings by specifying this is for sharing access, while tools like oe_article_get are for retrieving content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use (to change public/private status), and provides conditions: you must own the conversation and have the relay extension connected. Also warns against publishing PHI, guiding appropriate use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_askOpenEvidence AskA
Create a question. Fire-and-forget by default: returns {article_id, status:'pending'} immediately so the browser tab is free for other sessions — fetch the finished answer later with oe_article_get (optionally wait_for_completion:true). Pass wait_for_completion:true here to block and return the answer in one call. For a follow-up question pass original_article_id. Submits POST /api/article through the connected browser-extension relay (runs in your real logged-in tab, DataDome-free); the direct Node POST is deprecated and no longer attempted. Requires the relay extension to be connected (see extension/README.md).
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | ||
| timeout_sec | No | ||
| article_type | No | Ask OpenEvidence Light with citations | |
| include_bibtex | No | ||
| save_artifacts | No | ||
| disable_caching | No | ||
| poll_interval_ms | No | ||
| crossref_validate | No | ||
| original_article_id | No | Article UUID or openevidence.com/ask/<id> URL of the conversation to follow up on. | |
| wait_for_completion | No | ||
| strip_citation_markers | No | ||
| personalization_enabled | No | ||
| variant_configuration_file | No | prod |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully bears the burden. It discloses the fire-and-forget return (immediate pending), blocking with wait_for_completion, reliance on relay extension, and deprecation of direct POST. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph front-loaded with core behavior and key distinctions. It is dense and efficient, though slightly structured as a block of text. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite good coverage of core behavior and two key parameters, with 13 parameters and no output schema, the description omits details on many parameters (e.g., timeout_sec, article_type, include_bibtex) and the full output structure. It is adequate but not fully complete given complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 8% (only original_article_id described). The description adds meaning for wait_for_completion and original_article_id, but covers few of the 13 parameters. Many defaults are listed but not explained in context, so it partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it creates a question, distinguishes fire-and-forget vs. blocking modes, and explicitly differentiates from oe_article_get for fetching results. It also mentions follow-up usage via original_article_id, providing a specific verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance: when to use fire-and-forget (default) vs. wait_for_completion, how to do follow-ups, and the requirement of a connected relay extension. It implies context (browser tab free for other sessions) and notes deprecation of a direct method.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_auth_statusOpenEvidence Auth StatusA
Check if the local OpenEvidence session is valid (full network round-trip). For a fast pipeline-connectivity check use oe_health instead.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that the tool performs a full network round-trip and is non-destructive (checking validity). The behavior is simple and well-described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and alternative. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool is simple with no parameters and no output schema. Description fully covers what the tool does and how it differs from alternatives, making it complete for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so the description does not need to add parameter information. Baseline is 4, but the description is perfectly adequate given 100% schema coverage and empty input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool checks if the local OpenEvidence session is valid and distinguishes it from the sibling oe_health tool by specifying a full network round-trip versus a fast pipeline-connectivity check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use this tool (check session validity) and when to use the alternative oe_health (fast connectivity check), providing clear usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_add_articleOpenEvidence Collection Add ArticleC
Add a chat (article) to a collection. Idempotent in practice.
| Name | Required | Description | Default |
|---|---|---|---|
| article_id | Yes | ||
| collection_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description only discloses idempotency. It fails to mention permissions, error behavior, or side effects. For a mutation tool, this is insufficient transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise, with one main sentence plus a note on idempotency. However, it sacrifices completeness for brevity; a tool with two required UUID parameters would benefit from slightly more context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It does not explain return values, prerequisites, or error conditions, which are critical for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description adds no information about the parameters beyond their names. It does not explain what collection_id or article_id represent or how to obtain them, leaving the agent to rely solely on parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Add a chat (article) to a collection', which effectively conveys the tool's purpose. It distinguishes from sibling tools by the specific verb 'add', but does not explicitly differentiate from other collection manipulation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like oe_collections_bulk_apply or other collection tools. The description only mentions idempotency, which is a behavioral trait, not usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_createOpenEvidence Collection CreateB
Create a new collection. By convention, agent-managed names start with '#'.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only states 'Create a new collection' without disclosing behavioral traits such as authentication needs, idempotency, limits, or side effects. For a creation tool, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short (two sentences) and front-loaded with the main action, but it omits important details, making it too concise for a creation tool that requires more context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations, output schema, and schema descriptions, the description should compensate with more completeness. It does not cover return values, errors, or usage constraints, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, meaning the schema itself provides no explanations. The description adds only a naming convention, not field-specific semantics like what valid values are or how description is used.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Create a new collection', which clearly identifies the action and resource. Among siblings with various collection operations, 'create' is unambiguous. The naming convention hint further clarifies purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a naming convention ('agent-managed names start with #') that implies agent usage, but does not provide explicit guidance on when to use versus alternatives, nor any conditions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_db_initInit Collections SQLite MirrorA
Create the local SQLite mirror at $OE_MCP_DB_PATH (default ~/.openevidence-mcp/db/oe.sqlite). Idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the action and idempotency but does not explain behavior if the database exists, error conditions, or permissions needed. Adequate but could be more transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences totalling 20 words. Every word is necessary: the action, the path with default, and the idempotency note. No wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with no parameters and no output schema, the description provides the essential information: what it creates, where, and its idempotence. It could mention that this tool should typically be called before others, but given its simplicity, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, and schema coverage is 100% (empty). According to guidelines, 0 parameters yields a baseline of 4. The description adds no parameter information, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Create the local SQLite mirror' at a specific path, with a default provided. The verb 'Create' and resource 'SQLite mirror' are specific, and the tool is distinct from siblings which operate on collections.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives. While it mentions 'Idempotent', which suggests safe repeated use, there is no guidance on prerequisites, ordering, or when initialization is necessary. Usage is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_getOpenEvidence Collection GetA
Fetch a collection (incl. nested questions[] = membership list) by id.
| Name | Required | Description | Default |
|---|---|---|---|
| collection_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. It discloses that the result includes nested questions, which adds some behavioral context, but lacks details on side effects, permissions, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, concise and front-loaded, with no superfluous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple fetch with one parameter and no output schema, the description adequately states the purpose and mentions the nested structure, but fails to describe the full return format or confirm whether it returns all fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description should compensate. It only repeats 'by id' without explaining the uuid format or providing examples, adding minimal value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Fetch' and the resource 'collection by id', and distinguishes the tool by mentioning included nested questions (membership list), setting it apart from siblings like oe_collections_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (use when you have a collection id and want single collection with membership), but there is no explicit guidance on when to use this versus alternatives like oe_collections_list or oe_collections_summary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_listOpenEvidence Collections ListA
List all collections owned by the authenticated user.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It indicates a read-only listing operation but lacks details on pagination, sorting, or limits. For a zero-parameter tool, the behavioral disclosure is minimal but acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no extraneous words, front-loaded with the core action. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list with no parameters and no output schema, the description adequately covers the tool's purpose and scope. Could add detail about ordering or result limits, but not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has zero parameters, so schema coverage is 100%. Per rules, baseline is 4. The description adds no parameter info, which is appropriate given no parameters exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'List all collections owned by the authenticated user' uses a specific verb (list) and resource (collections) with clear scope (owned by user), distinguishing it from sibling tools like oe_collections_create or oe_collections_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives among many sibling collection tools. Does not mention scenarios where list is appropriate or when another tool like oe_collections_unsorted might be better.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_summaryCollections SummaryB
Counts (chats, collections, hashtag, memberships, unsorted) + last sync timestamps.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must cover behavioral traits. It only states what the tool counts and timestamps, but does not mention whether it is read-only, if it requires authentication, or if there are any side effects or limitations (e.g., rate limits).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—only one sentence—yet conveys the tool's core purpose. It is front-loaded with the key counts and timestamps. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks detail on the return format (e.g., structure of the counts object, how timestamps are represented). Given no output schema, this information would be helpful. Additionally, with many sibling tools, more context on distinguishing usage would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, which the schema reflects completely (100% coverage). The description adds value by listing the specific items counted (chats, collections, etc.), which is meaningful beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns counts and last sync timestamps for specific items (chats, collections, hashtag, memberships, unsorted). The verb 'summary' aligns with the tool's name. However, it does not differentiate from siblings like oe_collections_list or oe_collections_get, which could be used for similar purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. The description implies it is for getting summary data, but it does not state when not to use it or mention other tools that might be better suited for specific scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_sync_dbSync Collections + Memberships to SQLiteB
Refresh collections and memberships from the API into local SQLite. Prunes collections + memberships the server no longer reports.
| Name | Required | Description | Default |
|---|---|---|---|
| rate_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the pruning behavior, which is a destructive side effect. However, it does not mention other behavioral traits such as idempotency, authentication requirements, or the effect of the rate_seconds parameter. With no annotations, the description carries the full burden but only partially fulfills it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two short, front-loaded sentences with no fluff. Every phrase conveys essential action: refreshing and pruning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only one optional parameter and no output schema, the description is moderately complete. It explains the main action but omits details about the parameter's role and potential side effects, leaving gaps for a sync operation in a multi-tool context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention the only parameter, rate_seconds, which has 0% coverage in the schema. The parameter name and constraints suggest it controls pacing, but the description provides no clarification, leaving the agent to infer its meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it refreshes collections and memberships from API into local SQLite and prunes removed items. This distinguishes it from sibling tools like oe_collections_list or oe_collections_get, which do not perform sync or local storage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it is for syncing and pruning, but does not explicitly state when to use it versus alternatives like oe_collections_db_init or oe_collections_sync_history. No when-to-use or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_sync_historySync Chat History to SQLiteA
Paginate /api/article/list and upsert chats. Incremental by default (stops on the first all-known page).
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | ||
| pages | No | ||
| page_size | No | ||
| rate_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the incremental behavior and pagination, but does not explain what 'upsert chats' entails in terms of database mutation, idempotency, or potential side effects. The description adds moderate behavioral context but lacks specifics about the operation's safety or impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, no wasted words, and front-loads the primary action. Every sentence adds value, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four parameters and no output schema, the description is sparse. It does not cover return values, error conditions, prerequisites (e.g., whether oe_collections_db_init must be called first), or the relationship with sibling tools like oe_collections_sync_db. The lack of detail leaves significant gaps for an AI agent to use this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, yet the description only indirectly references the 'full' parameter via 'incremental by default'. Parameters like 'pages', 'page_size', and 'rate_seconds' are not explained at all. The description adds almost no meaning beyond the parameter names, which is insufficient for a tool with four parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: paginate and upsert chats, and identifies the specific API endpoint ("/api/article/list"). It also mentions the incremental behavior which distinguishes it from a full sync. This is specific and differentiates from sibling tools like oe_collections_sync_db.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions incremental mode and stopping condition (first all-known page), which implicitly suggests when to use (incremental sync) vs full sync ('full' parameter). However, it does not explicitly compare to other siblings like oe_collections_sync_db or provide when-not-to-use scenarios. The usage guidance is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_collections_unsortedList Unsorted ChatsA
Chats with no membership in any '#'-prefixed collection. Returns {unsorted_count, shown, items[]}.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| preview_chars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description does not explicitly state read-only behavior, but implies a read operation. It does disclose the return shape ({unsorted_count, shown, items[]}) which provides some insight into what the tool returns. However, it lacks details on pagination, performance, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence defining functionality followed by a brief return value structure. It is front-loaded, concise, and every element serves a purpose with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the purpose and return format fully. However, given no output schema, it could mention edge cases (e.g., empty list) or error conditions. Still, for a simple list tool, it is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has two parameters (limit, preview_chars) with defaults and ranges, but 0% schema description coverage. The description does not explain their purpose beyond what can be inferred from names. Since the coverage is low, the description should compensate, but it does not add any meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title and description clearly state the tool lists chats not in any '#'-prefixed collection. This differentiates it from sibling tools like oe_collections_list which lists collections, or oe_collections_summary which provides summaries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives. No instructions on prerequisites, limitations, or scenarios where another tool would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_healthOpenEvidence Relay HealthA
Millisecond-fast local check of the relay pipeline (daemon + browser extension) — no OpenEvidence network call. Use this to confirm the pipeline is up before oe_ask; use oe_auth_status only when you need to verify the login session itself.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses local, fast, no network call. However, it does not specify what the output looks like (e.g., boolean, error message), which would be helpful for a complete transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no wasted words. Information is front-loaded and efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, simple health check tool without output schema, the description provides all necessary context: purpose, usage guidance, and behavioral trait (local). Complete enough for agent selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters, so baseline is 4. Description adds context by explaining the tool's action beyond the empty schema, such as being local and fast.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'Millisecond-fast local check of the relay pipeline (daemon + browser extension)', specifying the verb (check) and resource (relay pipeline). It distinguishes from siblings like oe_auth_status and oe_ask.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance: use this to confirm pipeline is up before oe_ask, and use oe_auth_status only for login session verification. Provides clear context for when to use this tool vs alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_history_listOpenEvidence History ListC
List question history from OpenEvidence account.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| search | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits but only says "list." It fails to mention pagination behavior implied by limit/offset, whether it returns only the authenticated user's history, or any rate limits. The minimal description does not compensate for missing annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at one sentence, but this brevity sacrifices necessary information about parameters and usage context. It is under-specified rather than efficiently structured, failing to earn its place by omitting critical details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, no annotations, and no output schema, the description is severely incomplete. It lacks details on pagination, search functionality, return value format, and any prerequisites, making it insufficient for effective tool use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the three parameters (limit, offset, search). The agent receives no semantic help beyond the schema's basic type/constraint info, leaving the parameters largely unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states "List question history from OpenEvidence account," clearly identifying the action (list) and resource (question history). It distinguishes from sibling tools like oe_ask (ask questions) and oe_collections_list (list collections), though it could be more specific about what constitutes 'history.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description does not mention any conditions for use, prerequisites, or when not to use it, leaving the agent without contextual decision support.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oe_public_getOpenEvidence Conversation Page GetA
Read an OpenEvidence conversation from an /ask/ link and parse the page into Q&A turns (question, answer as markdown, references). Public (shared) conversations need no setup at all; your own private ones work when the relay extension is connected (your logged-in tab) or cookies.json exists. Use oe_article_get when you want the raw API payload + saved artifacts instead.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | A https://www.openevidence.com/ask/<id> URL, or the bare article UUID. | |
| strip_citation_markers | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description carries full burden. Discloses that it parses into Q&A turns (question, answer as markdown, references) and specifies authentication requirements (public vs private). Lacks explicit read-only declaration but 'Read' implies it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: three sentences covering purpose, setup, and alternative. No redundant information. Front-loaded with primary action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, description explains output format (Q&A turns with markdown and references). Covers prerequisites and distinguishes from sibling. Minor gap: no explanation of strip_citation_markers behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (url described, strip_citation_markers not). Description adds context on URL format and output type but does not explain the boolean parameter strip_citation_markers, leaving it unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it reads an OpenEvidence conversation from an /ask/<id> link and parses it into Q&A turns. Distinguishes from sibling oe_article_get which retrieves raw API payload.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains when to use vs not: 'Use oe_article_get when you want the raw API payload + saved artifacts instead.' Also details setup needs for public vs private conversations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
3 tool updates
v0.4.4- Added
oe_auth_status - Added
oe_collections_summary - Added
oe_health
9 tool updates
v0.4.3- Added
oe_answers_search - Changed
oe_article_get5 fields changed- added
Input schema / properties / article_id / descriptionAdded value: +"Article UUID, or any openevidence.com/ask/<id> URL." - removed
Input schema / properties / article_id / formatRemoved value: -"uuid" - removed
Input schema / properties / article_id / patternRemoved value: -"^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$" - added
Input schema / properties / refreshAdded value: +{ + "default": false, + "description": "Bypass the local answers cache and re-fetch from OpenEvidence.", + "type": "boolean" +} - added
Input schema / properties / strip_citation_markersAdded value: +{ + "default": false, + "type": "boolean" +}
- Added
oe_article_set_access - Changed
oe_ask4 fields changed- added
Input schema / properties / original_article_id / descriptionAdded value: +"Article UUID or openevidence.com/ask/<id> URL of the conversation to follow up on." - removed
Input schema / properties / original_article_id / formatRemoved value: -"uuid" - removed
Input schema / properties / original_article_id / patternRemoved value: -"^([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[1-8][0-9a-fA-F]{3}-[89abAB][0-9a-fA-F]{3}-[0-9a-fA-F]{12}|00000000-0000-0000-0000-000000000000|ffffffff-ffff-ffff-ffff-ffffffffffff)$" - added
Input schema / properties / strip_citation_markersAdded value: +{ + "default": false, + "type": "boolean" +}
- Removed
oe_auth_status - Removed
oe_collections_bulk_apply - Removed
oe_collections_classify - Removed
oe_collections_summary - Added
oe_public_get
2 tool updates
v0.3.0- Changed
oe_article_get3 fields changed- added
Input schema / properties / poll_interval_msAdded value: +{ + "default": 1200, + "maximum": 10000, + "minimum": 300, + "type": "integer" +} - added
Input schema / properties / timeout_secAdded value: +{ + "default": 120, + "maximum": 600, + "minimum": 5, + "type": "integer" +} - added
Input schema / properties / wait_for_completionAdded value: +{ + "default": false, + "type": "boolean" +}
- Changed
oe_ask1 field changed- changed
Input schema / properties / wait_for_completion / defaultPrevious value: -trueNew value: +false
11 tool updates
v0.2.0- Added
oe_collections_add_article - Added
oe_collections_bulk_apply - Added
oe_collections_classify - Added
oe_collections_create - Added
oe_collections_db_init - Added
oe_collections_get - Added
oe_collections_list - Added
oe_collections_summary - Added
oe_collections_sync_db - Added
oe_collections_sync_history - Added
oe_collections_unsorted
4 tool updates
v0.1.0- First observed
oe_article_get - First observed
oe_ask - First observed
oe_auth_status - First observed
oe_history_list
TDQS
Scored across 17 tools
Every tool has a clearly distinct purpose, with descriptions that eliminate ambiguity. For example, oe_health vs oe_auth_status differentiate between local pipeline check and session verification, and oe_collections tools each handle a specific aspect of collection management.
Tools follow a consistent pattern of 'oe_' prefix followed by a domain (collections, article, auth, etc.) and then a verb or noun. While not strictly verb_noun throughout (e.g., oe_health, oe_collections_summary), the pattern is predictable and readable.
17 tools is well-scoped for the OpenEvidence domain, covering authentication, health checks, collections management (9 tools), article operations, and search. Each tool serves a specific need without unnecessary duplication.
The tool surface covers the core workflows: asking questions, retrieving answers, searching, managing collections, and setting access. Minor gaps exist, such as no explicit delete tool for articles, but overall the set is comprehensive for its purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Generate, edit, and explore AI images. Flux, Imagen, LoRA identity swap, upscale, and more.
Search thousands of free SVG cut files & clipart. Free for commercial use, no attribution.
Human-made production music for sync — search by brief or reference, preview, score to picture.
Codeforces competitive programming users, contests, problems
Related MCP Servers
- -
- -
- -
- -
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/htlin222/openevidence-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server