Skip to main content
Glama

free-search-mcp

License Python MCP

free-search-mcp is a local-first Model Context Protocol server that needs no API key. It lets any LLM (Claude, GPT, a local Ollama model, …) search the web, fetch and clean up pages, and read documents, and you never sign up for a search API.

It combines ideas from several open-source MCP servers in one Python package, and adds the output shaping for LLMs and the reliability work that each of them lacked.

research("how does reciprocal rank fusion work", depth=3)
   ↓
# Research brief: how does reciprocal rank fusion work
_engines: duckduckgo, bing, anysearch · sources: 3 · ~3,400 tokens_

## Sources
- [1] Reciprocal rank fusion | Elasticsearch Reference — <https://…>
- [2] Hybrid Search Scoring (RRF) | Microsoft Learn — <https://…>
- [3] RRF explained in 4 mins — Medium — <https://…>

## Documents
…full Markdown bodies of each page, ready for the LLM to read…

That was one tool call. It returned three sources with their full text, and no API key was involved.

Quick start

You need uv and nothing else: no sign-up, no API key, no clone.

In Claude Code, inside a session:

/plugin marketplace add sweetcornna/free-search-mcp
/plugin install free-search@free-search-mcp

In Codex:

codex plugin marketplace add sweetcornna/free-search-mcp
codex plugin add free-search@free-search-mcp

(In Codex, a [mcp_servers.search] entry left in ~/.codex/config.toml by an earlier codex mcp add silently shadows the plugin's server. Remove it with codex mcp remove search.)

The plugin brings the 11 tools, the verified-research skill (how to check a snippet against its page and its date) and, in Claude Code, the free-search:quick-search agent, all pinned to one version. /plugin update free-search moves to the next release.

Optionally, install Chromium once for the browser-rendered engines (brave, startpage, zhihu, …) and JavaScript-heavy pages. Everything else works without it, and a call that needs it returns this command as its error:

uvx --from free-search-mcp playwright install chromium

Claude Desktop, other clients, a source checkout and Docker are under Install.

Related MCP server: uvxwebsearchmcp

Why this exists

Multi-engine

No API key

Smart fallback

PDF/DOCX

FTS5 cache

Filters

Trafilatura

LLM-tuned

nickclyde/duckduckgo-mcp-server

✗

✓

✗

✗

✗

✗

✗

~

mrkrsl/web-search-mcp

✓

✓

✓

✗

✗

✗

✗

~

Aas-ee/open-webSearch

✓

✓

~

✗

✗

✗

✗

~

VincentKaufmann/noapi-google-search-mcp

✗

✓

✓

✓

✓

✗

✗

~

free-search-mcp

✓

✓

✓

✓

✓

✓

✓

✓

"LLM-tuned" means Markdown-first output with token estimates, errors that name the next step, and a research() that turns a search and its fetches into one call. docs/HOW_IT_WORKS.md explains the less obvious columns.

Tools (11)

Tool

Description

search(query, ...filters)

Parallel multi-engine search, RRF-merged, title-fuzzy + host-canonical deduped, with optional extractive lead_snippet

research(question, depth?, ...filters)

One-shot: search + fetch top N + return Markdown brief

paper_graph(paper, direction?, limit?)

Walk one paper's citation graph: references, citing works ranked by influence, and Crossref retraction/correction notices. Takes a DOI, an OpenAlex ID, an exact title, or an arXiv reference (arXiv:1706.03762, an arxiv.org/abs/… URL, a 10.48550/arXiv.… DOI, or a bare id such as 1706.03762, 2401.12345v3 or hep-th/9901001 when it is the whole input)

compare(question, urls=[2..5])

Concurrent fetch of 2-5 URLs, side-by-side excerpts keyed by question

fetch(url, render?, inline?, ...)

Fetch any resource: reader-mode Markdown for pages, parsed text for documents, or a description (type/size/dimensions/sha256) for images and binaries. inline=True returns the image itself for vision models

fetch_batch(urls, ...)

Concurrent multi-URL fetch (max 20 per call)

read_doc(source, start?, length?, ...)

Parse PDF / DOCX / XLSX / PPTX / EPUB / CSV / code / zip-tar / HTML / TXT / MD with pagination

extract_structured(url, ...)

Pull JSON-LD / OpenGraph / Twitter cards / microdata via extruct. Long prose fields (articleBody and the like) are clipped, and the result says so in trimmed. This tool returns metadata; use fetch for the text

cache_search(query, limit?, ...)

FTS5 search across previously fetched pages

engines(group?)

The source tree (group, then sub-group, then engine), one line each. group is an enum of the groups that own engines, so a typo fails schema validation

download(url, ...)

Save a file to ${SEARCH_MCP_CACHE_DIR}/downloads by default; files auto-delete after 24h. Set SEARCH_MCP_DOWNLOAD_ENABLED=false to disable it.

There are also 5 MCP prompts and 2 resource templates (cache://page/{url}, cache://search/{query_hash}). A twelfth tool, ask, appears only when an answer backend is configured (see Delegating a lookup).

search and research take these filters:

Param

Values

Effect

freshness

day / week / month / year

Only results from the last N. Results with no date are kept, and freshness_note says when that is most of them. day / week also add googlenews

include_domains

["python.org", "djangoproject.com"]

Restrict to these domains

exclude_domains

["pinterest.com"]

Remove these

category

a group (paper, finance, news, software, security, weather, …) or a sub-group (paper.biomed, finance.filings, finance.fx, software.python, dataset.ml, …)

Routes to the sources that natively index that kind of thing; the enum in the tool schema lists every value

include_text

"async"

Substring required in title/snippet

exclude_text

"beginner"

Substring forbidden

max_age_hours

24

Accept a cached answer only if it is younger than this. Default 7 days; the tightest of this, the freshness window's TTL, the 6-hour news cap and the category's own cap (weather, finance.fx 1 hour; software, security 6 hours) wins

Output is Markdown by default, with provenance and a token estimate in the header. format="json" returns structured data.

How a search runs

  • A search that names no engines asks a small keyless pool: duckduckgo, bing, anysearch and mojeek. googlenews joins when recency is asked for, and so360 when the query is written in Chinese.

  • category= routes to sources that index that kind of thing (papers, filings, datasets, packages, CVEs, …). Record sources such as pypi, nvd or openmeteo also join by themselves when the question is one they can answer directly.

  • An engine that serves a CAPTCHA, a login wall or an off-topic page is benched, a reserve takes its seat, and the response says what happened. A proxy (SEARCH_MCP_PROXY) is the fix for IP gating.

  • Every result says how old it is and where its date came from, so the agent can tell a lead from a checked fact.

  • engines() prints the whole source tree. engines=["so360", "baidu"] runs exactly the engines named.

No engine that a search reaches by itself needs a key. The details, with measurements, are in docs/HOW_IT_WORKS.md; walls and proxies in docs/PROXY_AND_GATES.md; a tour of the tools with examples in docs/USAGE.md.

Install

The plugin in Quick start is the recommended path, because the server is pinned to the plugin's version and updates with it. The other routes run the same server:

Where

How

Claude Desktop

open free-search-mcp-<version>.mcpb from the latest release

MCP Registry

io.github.sweetcornna/free-search-mcp

Claude Code, without the plugin

claude mcp add search -s user -- uvx free-search-mcp

Codex, without the plugin

codex mcp add search -- uvx free-search-mcp

Any other MCP client

the stdio command uvx free-search-mcp

Docker

docker compose build, then docker compose run --rm search-mcp

A source checkout with Chromium, a smoke test and client registration in one step:

curl -LsSf https://raw.githubusercontent.com/sweetcornna/free-search-mcp/main/scripts/install.sh | bash -s -- --client claude-code

Over HTTP, uvx free-search-mcp --transport streamable-http --port 8000 serves http://127.0.0.1:8000/mcp. That endpoint has no authentication and fetches any URL for whoever reaches it, so keep it on loopback or put an authenticating proxy in front.

The JSON for Claude Desktop, Cursor, Cline, Continue and Zed, the installer's options and the plugin's token cost are in docs/INSTALL.md. Operating rules for agents are in docs/AGENT_USAGE.md.

Search on your own account: codex and antigravity

Two opt-in engines search on an account you sign in with instead of an API key. Neither is in any pool or route: a search reaches one only when the call names it, and nothing changes for anyone who never signs in.

codex

antigravity

Searches with

OpenAI's web search, the one Codex uses

Google Search, run by a Gemini model

Account

a ChatGPT plan that includes Codex

a Google account with Antigravity

Each search counts against

the plan's Codex usage

the account's Antigravity quota

The provider allows it

yes

no, see the warning below

Opens a sign-in page by itself

on first use, over stdio on a desktop

never

Full guide, with 中文速览

docs/CODEX_SEARCH.md

docs/ANTIGRAVITY_SEARCH.md

Warning: Google's Antigravity terms forbid using its sign-in from third-party tools and name suspension of the Antigravity and Gemini CLI accounts as the consequence, and Google has suspended accounts for it. Sign in to antigravity only with an account you accept that risk for.

1. Sign in once

Run the sign-in on the machine the server runs on:

uvx --from free-search-mcp search-mcp-login codex         # the ChatGPT page `codex login` shows
uvx --from free-search-mcp search-mcp-login antigravity   # prints the warning, then Google's consent page
uvx --from free-search-mcp search-mcp-login status        # account, plan, token expiry

(uv run search-mcp-login … in a source checkout.) The settings page, search-mcp-admin, has the same Sign in / 登录 buttons. Tokens are stored in ~/.config/search-mcp/oauth/ (0600) and refreshed automatically. The plugin, uvx and the Desktop bundle all read that directory, so a running server picks up a new sign-in without a restart.

  • codex can skip this step: the first search that names it opens the sign-in page and finishes once you approve (stdio on a desktop only; SEARCH_MCP_CODEX_AUTO_SIGNIN=false turns it off). A machine already signed in to the Codex CLI can reuse that sign-in read-only with search-mcp-login codex --use-codex-cli.

  • antigravity signs in with Antigravity's own OAuth client, which is not shipped in this package. The sign-in reads it from the Antigravity app installed on the machine (checked on macOS). Without Antigravity installed, set SEARCH_MCP_ANTIGRAVITY_CLIENT_ID and SEARCH_MCP_ANTIGRAVITY_CLIENT_SECRET; a machine that has signed in keeps both in ~/.config/search-mcp/oauth/antigravity.json.

2. Name the engine

search("rust 2024 edition changes", engines=["codex"])
research("what changed in python 3.14 asyncio", engines=["antigravity"])
search("rust 2024 edition changes", engines=["codex", "duckduckgo", "bing"])

In Claude Code or Codex, ask for it in words ("search this with the codex engine"). Filters apply as usual, and results are cached like any other search. Measured on 2026-09-26, a codex search took about 3 s and an antigravity search 10 to 20 s.

3. Settings (all optional)

Var

Default

Meaning

SEARCH_MCP_CODEX_MODEL, SEARCH_MCP_ANTIGRAVITY_MODEL

latest

latest follows the account's model catalogue, so a new model is used once it is listed; a model name pins one

SEARCH_MCP_CODEX_TIMEOUT, SEARCH_MCP_ANTIGRAVITY_TIMEOUT

60

seconds for one search

SEARCH_MCP_CODEX_AUTO_SIGNIN

true

open the sign-in page on first use

SEARCH_MCP_ANTIGRAVITY_CLIENT_ID, SEARCH_MCP_ANTIGRAVITY_CLIENT_SECRET

read from the Antigravity install

for a machine without Antigravity; set both or neither

SEARCH_MCP_PROXY

empty

the sign-in and the searches go through it

On a server, over SSH, or in Docker

  • Sign in with --no-browser (search-mcp-login codex --no-browser, or antigravity --no-browser), open the printed address in any browser, and approve. The browser then fails to load a 127.0.0.1:1455 or localhost:51121 address; paste that whole address into the terminal.

  • Over streamable-http, codex never opens a sign-in page by itself, since it would open on the server. Sign in there with the command above first.

  • In Docker, point SEARCH_MCP_CONFIG_DIR at a mounted volume so the tokens outlive the container, and run the sign-in inside it with --no-browser. antigravity there needs the two client variables.

search-mcp-login logout codex or logout antigravity deletes the stored tokens. To revoke Antigravity's access itself, remove it under "Third-party apps and services" in the Google account.

When it does not work

The error says

Do this

codex not configured / antigravity not configured

no sign-in is stored on this machine; run the sign-in above

port 1455 and 1457 … are in use, or port 51121

another sign-in (codex login, an open sign-in page) holds the port; close it and retry

the usage limit, quota or rate limit was reached

wait until the reset time the error gives; the keyless engines still work

needs Antigravity's OAuth client

install Antigravity, or set the two client variables

HTTP 403 no valid license from antigravity

set SEARCH_MCP_ANTIGRAVITY_VERSION to the installed Antigravity release; it can also mean Google restricted the account

has expired or was revoked

sign in again

Optional: your own API key

You do not need one, and agents should never ask for one. Five engines run only when a call names them and the operator has set their own key:

Engine

Provider

Key

brave_api

Brave Search API

SEARCH_MCP_BRAVE_API_KEY

serper

Serper

SEARCH_MCP_SERPER_API_KEY

tavily

Tavily

SEARCH_MCP_TAVILY_API_KEY

google_cse

Google Custom Search

SEARCH_MCP_GOOGLE_CSE_API_KEY + SEARCH_MCP_GOOGLE_CSE_CX

github_code

GitHub

SEARCH_MCP_GITHUB_TOKEN

Set a key as an environment variable or .env line, or on the local settings page:

uv run search-mcp-admin        # local settings / 本地设置 — http://127.0.0.1:8765

(uvx --from free-search-mcp search-mcp-admin for a plugin or uvx install.) The page is bilingual (中英双语), binds to 127.0.0.1 only, applies a saved key without a restart, and never shows a stored value again. The walkthrough per provider is docs/API_KEYS.md.

Delegating a lookup

Delegation keeps page text out of the caller's context and returns a short answer with dated sources. None of it is on by default:

You want

Use

It needs

The agent to search and read by itself

the tools, as installed

nothing

Your host's own subagents to do the lookup

the quick-search agent definition

a host with subagents

The server to dispatch through Claude Code

SEARCH_MCP_AGENT_BACKEND=claude-code

the claude CLI, logged in

The server to dispatch through Codex

SEARCH_MCP_AGENT_BACKEND=codex

the codex CLI, logged in

The server to call a model endpoint you choose

SEARCH_MCP_AGENT_BACKEND=api

an OpenAI-compatible or Anthropic-compatible URL

Your own code to supply the model

ask(question, answer_with=...) in Python

nothing else

The settings, the agent files and measured timings are in docs/DELEGATION.md.

Configuration

Nothing is required. Settings are SEARCH_MCP_* variables, read from the environment first, then ./.env, then ~/.config/search-mcp/.env (SEARCH_MCP_CONFIG_DIR moves that directory). The ones people change most:

Var

Default

Meaning

SEARCH_MCP_PROXY

empty

outbound proxy (http, https or socks5) for the engines, the browser and fetch

SEARCH_MCP_PROXY_ENGINES

empty

proxy only these engines

SEARCH_MCP_DEFAULT_ENGINES

["duckduckgo","bing","anysearch","mojeek"]

the pool a search uses when it names no engines

SEARCH_MCP_REGION

us-en

locale for engines that take one

SEARCH_MCP_CACHE_DIR

~/.cache/search-mcp

search and page cache

SEARCH_MCP_TOOLS

empty

register only these tools

Every setting is listed in docs/CONFIGURATION.md and documented in .env.example.

Development

git clone https://github.com/sweetcornna/free-search-mcp.git && cd free-search-mcp
uv sync && uv run playwright install chromium
uv run pytest -q                              # offline
SEARCH_MCP_TEST_NETWORK=1 uv run pytest -v    # live, hits the real web

Running claude inside the checkout picks up the repo's .mcp.json, which starts the working tree's server. The architecture is in docs/HOW_IT_WORKS.md, and cutting a release in docs/RELEASING.md.

Credits

This project builds on:

License

MIT. See LICENSE.

Available Tools

11 tools
compareCompare URLs side-by-sideA
Read-onlyIdempotent

Fetch 2-5 URLs concurrently and return per-URL excerpts so the LLM can compare them against a single question in one round trip.

Best for:
- Side-by-side product/feature/article comparisons.
- "Compare X to Y" or "How does A differ from B" queries.
- Triangulating a fact across multiple sources.

Not recommended for:
- >5 URLs -> use `fetch_batch`.
- 1 URL -> use `fetch`.
- Don't have URLs yet -> use `search` or `research` first.

Returns:
- markdown (default): a comparison brief with per-URL sections, each
  containing title, sitename, published date, and a smart-truncated excerpt.
- json: {question, urls, excerpts:[{url, title, excerpt, ...}],
  tokens_estimated}.

Common mistakes:
- Asking `compare` to actually answer the question: it returns material,
  the LLM does the comparison.
- Passing >5 URLs and expecting them all to fit in context: use
  `fetch_batch` for bulk reads.

Args:
    question: The comparison question the LLM will answer using the
        returned excerpts.
    urls: 2-5 absolute http(s) URLs.
    format: "markdown" (default) or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
formatNomarkdown
questionYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and idempotentHint annotations, the description discloses concurrency, smart truncation, return format specifics (markdown/json with tokens_estimated), and common mistakes (e.g., asking it to answer the question). No contradiction with annotations; it adds valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with sections (intro, best for, not recommended, returns, mistakes, args). It front-loads the core purpose and each sentence earns its place, providing high information density without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is fully self-contained: it covers when to use, when not to use, parameter details, return formats, and common pitfalls. No output schema exists, but the return description is detailed enough for an agent to handle responses correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The 'Args' section explains every parameter: question (purpose), urls (count and format), format (enum with default). This is complete and adds meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Fetch 2-5 URLs concurrently and return per-URL excerpts'), identifies the resource (URLs), and clarifies the outcome (enables comparison). It also distinguishes itself from siblings by naming fetch_batch and fetch as alternatives for different URL counts, leaving no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Best for' and 'Not recommended for' sections give concrete usage conditions: >5 URLs → fetch_batch, 1 URL → fetch, no URLs → search/research. This is direct, actionable guidance on when to use this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

downloadDownload a file to diskA
Idempotent

Save a file from a URL to a local, auto-expiring download directory.

Downloads are enabled by default and saved under
`SEARCH_MCP_CACHE_DIR/downloads`. Set `SEARCH_MCP_DOWNLOAD_ENABLED=false`
to disable them or `SEARCH_MCP_DOWNLOAD_DIR` to override the destination.

Best for:
- Keeping an actual file (installer, dataset, archive, image) rather than
  its text.
- Handing a path to another tool that needs a real file on disk.

Not recommended for:
- Reading a document's contents -> use `read_doc`, which parses it without
  touching the filesystem.
- Looking at a web page -> use `fetch`.
- Viewing an image -> use `fetch(inline=True)`.

Returns:
- markdown (default): where the file was saved, its size and type.
- json: {url, saved_path, media_type, bytes_size, sha256, expires_in_hours}.
  An expires_in_hours value of 0 means TTL cleanup is disabled.

Retention: files older than SEARCH_MCP_DOWNLOAD_TTL_HOURS (default 24) are
deleted before the next download and at startup. A value of 0 disables TTL
cleanup. Otherwise, treat the path as short-lived and copy it elsewhere if
you need to keep it.

Args:
    url: Absolute http(s) URL of the file to save.
    format: "markdown" or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
formatNomarkdown

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations, the description discloses filesystem effects, default-enabled behavior, environment variable controls, auto-expiring paths, TTL cleanup, and return details. This is substantial additive transparency consistent with readOnlyHint=false and there is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into clear labeled sections (best for, not recommended, returns, retention, args) with no filler. The core action is front-loaded, and every section earns its place by providing actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of an output schema, the description provides both markdown and JSON return shapes, including the sha256 and expiry fields. It also covers retention, configuration, and use cases, making it complete enough for an agent to invoke safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates by specifying that url must be an absolute http(s) URL and by explaining format through the Returns section. It could go slightly deeper on format semantics, but the return value details largely close the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Save a file from a URL to a local, auto-expiring download directory.' It clearly differentiates from siblings like read_doc and fetch by saying it keeps an actual file rather than parsing content or viewing a page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Best for' and 'Not recommended for' sections explicitly name alternatives: read_doc for document content, fetch for web pages, and fetch(inline=True) for images. This gives an agent unambiguous when-to-use versus when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enginesList available search enginesA
Read-onlyIdempotent

List the available sources, grouped by what they index.

Best for:
- Choosing a source deliberately: which one indexes filings, or preprints,
  or Chinese-language pages.
- Checking a name before passing it to `engines=` on `search` / `research`.

Not recommended for:
- Calling on every search: the list is static, so read it once.

Returns (markdown): a `group -> sub-group -> engine` tree, one line of
description per engine. `group="paper"` restricts it to that group.
Returns (json): `{"engines": [...names...], "taxonomy": {...},
"descriptions": {...}}`.

Prefer `category=` over `engines=`. `category="paper"` WIDENS: it routes to
one specialist per sub-group. A dotted sub-group NARROWS: `"paper.biomed"`
queries only the biomedical indexes. Naming engines explicitly turns that
routing off entirely, so reach for it only to force a specific source.

Common mistakes:
- Passing one of these names as `query`: they belong in `engines=`.
- Passing a key-only engine with no key configured; it returns an
  actionable error, not results.
ParametersJSON Schema
NameRequiredDescriptionDefault
groupNo
formatNomarkdown

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool read-only and idempotent, and the description adds genuinely useful behavioral context: the list is static and should be read once, returns can be markdown or JSON, and key-only engines produce actionable errors. This goes well beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although long, the description is tightly organized with functional headers, scannable bullet-ish sections, and every sentence adds routing or error-avoidance value. The core purpose is front-loaded, and the detailed guidance is where an agent needs it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with two optional parameters and no output schema, this description is complete: it explains when to use it, what the parameters do, what output format to expect, and likely failure modes. No critical calling information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden of explaining parameters. It clarifies `group` values, how dotted sub-groups narrow, how `category=` differs from `engines=`, and what both markdown and JSON output shapes look like. This compensates fully for the schema's lack of parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists available sources/engines grouped by what they index, which is specific and immediately distinct from siblings like `search`, `research`, and `fetch`. It names the resource and the organizing principle without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit best-for scenarios, a not-recommended scenario, and detailed routing guidance comparing `category=` vs `engines=` and group/sub-group semantics. It also lists common mistakes, giving an agent both positive and negative selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_structuredExtract structured data from a URLA
Read-onlyIdempotent

Pull JSON-LD, OpenGraph, Twitter cards, and microdata from a web page.

Best for:
- Product pages (price, currency, availability, brand, rating).
- Article pages (author, publish date, image, headline).
- Recipe / event / video pages where rich metadata IS the answer.
- Cases where `fetch` returns prose but you need fields.

Not recommended for:
- Just reading a page -> use `fetch`.
- PDFs / DOCX -> use `read_doc`.
- Pages that don't publish schema.org metadata (most blogs): you'll get
  empty lists; fall back to `fetch`.

Returns:
- json: {url, json_ld:[], microdata:[], opengraph:[], rdfa:[]}. Twitter
  card meta tags are surfaced inside the `opengraph` list.
- markdown (default): a flattened key/value view with each block printed
  as a JSON code block under its syntax heading.

Common mistakes:
- Calling on every URL "just in case": most sites have no structured
  data, and `fetch` is what you actually want.

Args:
    url: Absolute http(s) URL.
    format: "markdown" (default) or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
formatNomarkdown

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint. The description adds value beyond these: it discloses the return format details (JSON key structure, markdown flattening behavior), the Twitter-card-inside-opengraph quirk, and the open-world caveat that most pages yield empty lists. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-organized with clear section headers (Best for, Not recommended, Returns, Common mistakes, Args). It is longer than minimal, but each section earns its place and the core purpose is front-loaded. Slightly verbose but purposefully structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a multi-format extraction tool with no output schema. The description carries the return-format burden itself, listing the JSON keys, the markdown default behavior, and parameter meanings. It also sets expectations about empty results. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate — and it does. The Args section clarifies that url must be an 'Absolute http(s) URL' (a constraint not in the schema) and explains format's options and default. This adds real meaning beyond the bare schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair: 'Pull JSON-LD, OpenGraph, Twitter cards, and microdata from a web page.' It lists the exact data formats extracted, and distinguishes itself from siblings by name (fetch, read_doc), making the tool's identity unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Exceptionally explicit. It provides 'Best for' scenarios (product, article, recipe/event/video pages), 'Not recommended for' cases with named alternatives (use fetch, use read_doc), and a 'Common mistakes' warning that most sites lack structured data and fetch is preferred. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetchFetch a URL: page text, document, or resourceA
Read-onlyIdempotent

Fetch one URL: page text, or a description of a non-text resource.

Handles any http(s) resource, not just HTML:
- HTML pages -> reader-mode Markdown (nav/footer/scripts stripped).
- PDF/DOCX/XLSX/PPTX/EPUB/CSV/code/archives -> parsed text (same engine as
  `read_doc`, which you should prefer when you need pagination).
- Images, video, audio, fonts, opaque binaries -> a description
  (media type, byte size, dimensions, sha256), NOT the bytes.

Best for:
- You already have a URL (from `search`, the user, or your own knowledge)
  and need the actual page text.
- Verifying a single claim by reading the source.
- Checking what a resource IS before deciding to spend tokens on it.

Not recommended for:
- Multiple URLs at once -> use `fetch_batch` (concurrent, one round-trip).
- "Search then read top N" -> use `research` (one call, not two).
- Long documents you need to page through -> use `read_doc` (start/length).
- You don't have a URL yet -> use `search` first.

Returns:
- markdown (default): a small header (URL, render method, token count)
  plus the cleaned page body.
- json: {url, title, content, method, truncated, tokens_estimated,
  author, published_date, sitename}, plus {media_type, bytes_size, sha256,
  width, height} for non-text resources.
- With `inline=True` on an image: the image itself, viewable by a
  vision-capable model.

Common mistakes:
- Passing a search query instead of a URL.
- Using `render="http"` on a JS-only SPA: it returns near-empty content;
  use "auto" (default) or "browser".
- Setting `inline=True` on a large image out of habit. A 1MB image costs
  well over a thousand tokens; fetch it plainly first and inline only if
  the description says it's worth looking at.
- Forgetting that results are cached 7 days: use `force_refresh=True`
  or `max_age_hours=0` for a fresh pull. The header says `cached N days ago`
  when you are not looking at the live page; for deadlines, prices and
  anything else that moves, that is the cue to refresh.
- Reading `no publication date found` as "recent". It means unknown.

Args:
    url: Absolute http(s) URL.
    render: "auto" (try HTTP, fall back to stealth Chromium), "http"
        (fast, fails on JS), "browser" (slow, robust).
    force_refresh: Bypass the page cache entirely.
    max_age_hours: Treat cached pages older than this as a miss. 0 = same
        as force_refresh. None = server default TTL (7 days).
    inline: For images only. Returns the image itself instead of a
        description, so a vision-capable model can see it. Ignored for
        text resources.
    format: "markdown" or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
formatNomarkdown
inlineNo
renderNoauto
force_refreshNo
max_age_hoursNo

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint, openWorldHint, and idempotentHint, and the description adds extensive context beyond them: cache behavior with 7-day TTL and force_refresh semantics, render mode differences including caveats for JS-only SPAs, inline image token costs, and the meaning of 'no publication date found'. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but densely structured with clear headers, bullet lists, and code-style naming. Every section adds actionable information: return formats, common mistakes, and parameter semantics. No redundant fluff; the structure makes the length digestible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with six parameters and no output schema, the description is remarkably complete. It covers input requirements, output shapes for markdown and json, non-text resource descriptions, caching, render strategies, and failure-prone usage patterns. An agent has everything needed to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it delivers a detailed Args section explaining every parameter, including defaults (render='auto', format='markdown') and nuanced behavior (max_age_hours=0 means force_refresh, None means server TTL). This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Fetch one URL: page text, or a description of a non-text resource.' It clearly identifies what the tool does and differentiates it from siblings by listing what it is not (batch fetch, research, read_doc). An agent can select it correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'Best for' and 'Not recommended for' sections, naming alternatives like fetch_batch, research, read_doc, and search with precise conditions. This is exactly the level of guidance needed to route the agent to the right tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_batchFetch many URLs concurrentlyA
Read-onlyIdempotent

Fetch a list of URLs in parallel. Per-URL failures do not raise.

Best for:
- 2+ URLs you want to read in one round-trip.
- Reading the top N results of a previous `search` call.

Not recommended for:
- A single URL -> `fetch` (no list-wrapping overhead).
- "Search and then read" -> `research` collapses both into one tool call.
- PDFs/DOCX -> `read_doc` per file.

Returns:
- markdown (default): each page rendered as a Markdown section, separated
  by horizontal rules; failed URLs become inline error notes.
- json: list[dict], one entry per URL, with `error` set on failures.

Common mistakes:
- Passing a single URL inside a 1-element list: use `fetch` directly.
- Assuming an exception means the whole batch failed; check each item's
  `error` field instead.

Args:
    urls: List of absolute http(s) URLs (max 20 per call).
    render: Same as `fetch`.
    format: "markdown" or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
formatNomarkdown
renderNoauto

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the safety profile is covered. The description adds crucial behavioral details: per-URL failures do not raise exceptions, the markdown output embeds inline error notes, and the JSON format includes an error field per entry. It also warns against common mistakes, exceeding what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into labeled sections (Best for, Not recommended, Returns, Common mistakes, Args) that are easy to scan. Every sentence adds value—there is no filler. The core purpose is stated first, and the detailed subsections are appropriately sequenced.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description fully specifies the two return formats (markdown sections with inline errors, and JSON list with error fields). It covers edge cases (single URL, failed items, maximum batch size) and references sibling tools for alternatives. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden. It explains each parameter: urls as absolute http(s) with a max of 20, render as 'Same as fetch' (referencing the sibling), and format with its two allowed values. It also clarifies the failure semantics associated with the urls parameter, which the schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Fetch a list of URLs in parallel') and immediately differentiates itself from siblings by naming fetch, research, and read_doc with explicit conditions. The purpose is unambiguous and the scope is clearly bounded.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a 'Best for' section listing concrete scenarios (2+ URLs, top N results of a search) and a 'Not recommended for' section with explicit alternatives (single URL -> fetch, search-and-read -> research, PDFs/DOCX -> read_doc). This gives an agent clear decision criteria for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

paper_graphWalk a paper's citation graphA
Read-onlyIdempotent

Follow the citations of ONE paper, and check whether it still stands.

`search` finds papers that MENTION your words. This follows the edges
instead: what a specific paper built on, and what has built on it since.

Best for:
- Checking a citation before repeating it: is the DOI real, and has the
  paper been retracted or corrected?
- "What happened after this result": citing works come back ordered by how
  much the field cited them, so a 2019 paper leads to the current state of
  the art rather than to the most recent preprint about it.
- Building a reading list backwards from one good paper.

Not recommended for:
- Finding papers by topic -> `search(category="paper")`, or a sub-group
  like `"paper.biomed"` / `"paper.cs"` / `"paper.preprint"`.
- Reading the paper itself -> `read_doc` on the returned URL.

Returns:
- markdown (default): the paper with its retraction/correction notices,
  then "References" and "Cited by" sections.
- json: {paper, references, citations, notes}, where `paper.crossref`
  carries `registered` and every post-publication `notices` entry.

Common mistakes:
- Passing a topic instead of a paper. A title resolves to its single best
  match; a phrase that names no specific paper resolves to the wrong one.
- Reading an empty `citations` list as "uncited" when `notes` says the
  lookup was truncated.

Args:
    paper: DOI (`10.1145/1571941.1572114`, or a doi.org URL), an OpenAlex
        ID (`W2148972377`), or the paper's exact title.
    direction: "both", "references" (what it cites) or "citations" (what
        cites it).
    limit: Max neighbours per direction, 1-50.
    format: "markdown" or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
paperYes
formatNomarkdown
directionNoboth

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint, openWorldHint, and idempotentHint, but the description adds substantial behavioral context: retraction/correction notices, citation ordering by field impact, truncation behavior in citations list, and specific return format details. It also explains the 'notes' field for truncated lookups. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Though lengthy, the description is meticulously structured with headers, bullet lists, and code blocks. The core purpose and differentiation are front-loaded, and every section (Best for, Not recommended, Returns, Common mistakes) provides non-redundant, actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description fully explains both markdown and json return formats, including nested fields like paper.crossref and notices. It covers edge cases like truncated citations and retraction notices, ensuring an agent has all necessary information to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description fully compensates. It explains the paper parameter with concrete examples (DOI, OpenAlex ID, exact title), the direction enum with meanings, the limit range (1-50), and format with return-type implications. This goes far beyond the schema's bare definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific verb 'follow the citations' and the resource 'ONE paper', clearly differentiating from search which finds mentions. It also explains the distinction between edge-following and text-matching, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides 'Best for' and 'Not recommended for' sections, naming alternatives like search(category=...), read_doc, and giving concrete scenarios. Also includes 'Common mistakes' that warn against passing topics instead of papers, which directly guides correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_docRead a remote (or sandboxed local) documentA
Read-onlyIdempotent

Read an http(s) document (or a sandboxed local file) into Markdown.

Best for:
- Remote PDFs and DOCX from an http(s) URL (parsed locally, no remote API).
- Local PDF/DOCX/text/Markdown files, ONLY when local reads are enabled
  (see Security below).
- Paginating through a long document via `start` / `length`.

Not recommended for:
- Arbitrary HTML web pages -> `fetch` does reader-mode cleanup that this
  tool does not.
- Pages discovered through search -> `fetch` or `research`.

Security (local files are sandboxed and OFF by default):
- Local-file reads are DISABLED unless the server operator sets the
  SEARCH_MCP_DOCUMENT_ROOT env var to a directory. With it unset, a local
  path raises a "local file reads are disabled" error. Pass an http(s)
  URL instead, or ask the operator to enable the sandbox.
- When enabled, `source` must resolve INSIDE that root; relative paths
  resolve against the root (not the process CWD) and any `..` traversal
  that escapes the root is rejected. `file://` URLs are always rejected.
- Remote http(s) sources are unaffected by this setting.

Returns:
- markdown (default): rendered document text with a small header.
- json: {content, title, format, total_chars, start, returned_chars,
  truncated}. Use `total_chars` and `returned_chars` to drive pagination.

Common mistakes:
- Calling this on a normal article URL: you'll get raw HTML noise. Use
  `fetch` instead.
- Forgetting to advance `start` when paginating: next call should pass
  `start = previous_start + returned_chars`.
- Passing a negative `length` (raises an error) or a `start` past the end
  (clamped to EOF: you'll get `returned_chars == 0`, `start == total_chars`,
  and `truncated == False`, which is the signal you've paged off the end).

Args:
    source: http(s) URL, or a local path UNDER SEARCH_MCP_DOCUMENT_ROOT when
        local reads are enabled (disabled by default; see Security).
    start: Character offset to begin reading from. Default 0. Clamped into
        [0, total_chars]; a negative value is treated as 0.
    length: Max characters to return; None = read to end (still capped by
        the per-call max content size). Must be >= 0. A negative length
        is rejected with a ValueError.
    format: "markdown" or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
startNo
formatNomarkdown
lengthNo
sourceYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint, openWorldHint, and idempotentHint, but the description goes far beyond them: it details sandboxing rules (env var requirement, path traversal rejection, file:// rejection), error behavior (disabled error, negative length ValueError), clamping semantics, and the exact pagination signal (`returned_chars == 0`). This is rich behavioral context that annotations alone could never convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every section serves a purpose: 'Best for' and 'Not recommended for' give usage guidance, 'Security' explains the critical env var behavior, 'Returns' details output formats, 'Common mistakes' prevents misusage, and 'Args' is a compact parameter reference. It front-loads the core purpose in the first line and organizes details logically, earning its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a genuinely complex tool (remote vs local, pagination, security, multiple output formats), yet the description covers every aspect needed to call it correctly: input expectations, edge cases (clamping, negative length), output structure (markdown header, json fields), and how to detect end-of-document. With no output schema present, the description fully substitutes that information. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description fully compensates. The Args section explains each parameter: source's format and security constraints, start's default and clamping behavior, length's range and rejection condition, and format's enum meaning. It also provides the json return fields to drive pagination, making each parameter's role crystal clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line states a specific verb and resource ('Read an http(s) document or sandboxed local file into Markdown'), and the 'Best for' / 'Not recommended for' sections explicitly differentiate it from siblings like fetch and research. An agent can immediately tell what this tool does and what it doesn't.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There are explicit usage recommendations with named alternatives: 'Not recommended for: Arbitrary HTML web pages -> `fetch` does reader-mode cleanup...' and 'Pages discovered through search -> `fetch` or `research`.' It also explains the crucial security prerequisite for local files. No ambiguity remains.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

researchSearch and read in one callA
Read-only

One-shot research: search the web, fetch the top results, return both.

Best for:
- Open-ended questions that need finding sources AND reading them
  ("what's new with X", "summarize the controversy around Y").
- Replacing a `search` + N x `fetch` chain with one call.
- Producing a citable brief with [n]-style source references.

Not recommended for:
- You only need links -> `search` (cheaper, no fetching).
- You only need to read one URL you already have -> `fetch`.
- You want to query previously-fetched cached pages -> `cache_search`.
- Checking or expanding one paper's citations -> `paper_graph`.

Returns:
- markdown (default): a "Research brief" with a Sources index then the
  full Markdown body of each fetched document, separated by horizontal
  rules; includes a token estimate.
- json: {question, engines, sources:[{rank,title,url,snippet,...}],
  documents:[...], tokens_estimated, errors}.

Common mistakes:
- Using `depth=8` for a quick lookup: that's 8 page fetches, and 2-3 is
  almost always enough.
- Calling `research` for a known URL: that is what `fetch` is for.
- Forgetting that `fetch=False` returns sources only (much cheaper if
  the LLM only needs to pick which one to read).

Args:
    question: What you want to know, in natural language.
    depth: How many top results to fetch (1-8). 3 is a good default.
    engines: Override the engine set (see `engines()` for names). Prefer
        `category=`. Naming engines turns category routing off.
    fetch: If False, return source list without reading them.
    freshness: "day"|"week"|"month"|"year". Restricts to recent results.
        Best-effort; undated results are kept rather than dropped.
    include_domains: Restrict to these domains (e.g. ["python.org"]).
    exclude_domains: Drop results from these domains.
    category: Which KIND of source to search; a bare group widens, a dotted
        sub-group narrows. Same values as `search`; see `engines()`.
    include_text: Substring required in title or snippet (case-insensitive).
    exclude_text: Substring forbidden in title or snippet.
    use_cache: Reuse cached search/page data within TTL.
    max_age_hours: Treat cached search results AND cached page bodies older
        than this as a read miss; fresh data is always written back. 0 =
        force-refresh both the engine search and every fetched page body;
        None = server default TTL (7 days). A non-zero value is honored for
        both halves (it used to be ignored for anything but 0).
    format: "markdown" or "json".
ParametersJSON Schema
NameRequiredDescriptionDefault
depthNo
fetchNo
formatNomarkdown
enginesNo
categoryNo
questionYes
freshnessNo
use_cacheNo
exclude_textNo
include_textNo
max_age_hoursNo
exclude_domainsNo
include_domainsNo

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true, but the description goes far beyond that: it discloses fetch behavior, cache semantics (use_cache, max_age_hours with 0 vs None), freshness handling, the difference between markdown and json return formats, the effect of fetch=False, and even behavioral nuances like 'Naming engines turns category routing off'. It also explains that non-zero max_age_hours is honored for both halves, correcting a past behavior. This is rich, accurate, and does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but impeccably structured: a one-line summary, then 'Best for', 'Not recommended for', 'Returns', 'Common mistakes', and 'Args' sections. Every sentence adds value—even the 'Common mistakes' section reinforces the usage guidelines. It is front-loaded with the core purpose, and the Args section is efficiently organized with one-line-per-parameter explanations. Nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (13 parameters, no output schema) and the absence of an output schema, the description covers everything an agent needs: return formats (markdown vs json) with descriptions, parameter semantics, usage guidance, and even error handling (json 'errors' field). It also provides token estimates and source structure, making it self-contained for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries full responsibility for parameter meaning. It does this exceptionally: each of the 13 parameters is explained in natural language with defaults, effects, and interactions (e.g., depth=3 as good default, engines override category routing, include_text substring semantics). This far exceeds the schema's basic types and enums.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line 'One-shot research: search the web, fetch the top results, return both' states a specific verb+resource and a composite action. It distinguishes itself from siblings via the 'Not recommended for' section, naming 'search', 'fetch', and 'cache_search' with explicit conditions. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Best for' and 'Not recommended for' sections are explicit, concrete, and name the exact sibling tools to use instead (search, fetch, cache_search, paper_graph) with the conditions that select them. It even gives example use cases and common mistakes (e.g., using depth=8 for quick lookups), leaving no inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.12.0
    • Changedcache_search1 field changed
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "items": {
        -            "additionalProperties": true,
        -            "type": "object"
        -          },
        -          "type": "array"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "cache_searchOutput",
        -  "type": "object"
        -}New value: +null
    • Changedcompare1 field changed
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "additionalProperties": true,
        -          "type": "object"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "compareOutput",
        -  "type": "object"
        -}New value: +null
    • Changeddownload1 field changed
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "additionalProperties": true,
        -          "type": "object"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "downloadOutput",
        -  "type": "object"
        -}New value: +null
    • Changedengines3 fields changed
      • addedInput schema / properties / format
        Added value: +{
        +  "default": "markdown",
        +  "enum": [
        +    "markdown",
        +    "json"
        +  ],
        +  "title": "Format",
        +  "type": "string"
        +}
      • addedInput schema / properties / group
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "web",
        +        "news",
        +        "paper",
        +        "github",
        +        "forum",
        +        "image",
        +        "dataset",
        +        "finance",
        +        "software",
        +        "security",
        +        "reference",
        +        "weather",
        +        "docs",
        +        "gov",
        +        "stats",
        +        "calendar"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "enum": [
        +        "news",
        +        "pdf",
        +        "github",
        +        "paper",
        +        "forum",
        +        "blog",
        +        "image",
        +        "dataset",
        +        "finance",
        +        "software",
        +        "security",
        +        "reference",
        +        "weather",
        +        "docs",
        +        "gov",
        +        "stats",
        +        "calendar",
        +        "news.world",
        +        "paper.index",
        +        "paper.preprint",
        +        "paper.biomed",
        +        "paper.cs",
        +        "paper.math",
        +        "paper.openaccess",
        +        "paper.trial",
        +        "dataset.repository",
        +        "dataset.ml",
        +        "dataset.gov",
        +        "finance.filings",
        +        "finance.market",
        +        "finance.macro",
        +        "finance.fx",
        +        "finance.entity",
        +        "finance.crypto",
        +        "software.lifecycle",
        +        "software.github",
        +        "software.python",
        +        "software.node",
        +        "software.rust",
        +        "software.registry",
        +        "software.app",
        +        "security.cve",
        +        "security.package",
        +        "security.exploited",
        +        "reference.domain",
        +        "docs.web",
        +        "docs.rfc",
        +        "gov.us",
        +        "gov.uk",
        +        "stats.indicator",
        +        "calendar.holidays",
        +        "calendar.clock"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Group"
        +}
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "items": {
        -        "type": "string"
        -      },
        -      "title": "Result",
        -      "type": "array"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "enginesOutput",
        -  "type": "object"
        -}New value: +null
    • Changedextract_structured1 field changed
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "additionalProperties": true,
        -          "type": "object"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "extract_structuredOutput",
        -  "type": "object"
        -}New value: +null
    • Changedfetch_batch1 field changed
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "items": {
        -            "additionalProperties": true,
        -            "type": "object"
        -          },
        -          "type": "array"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "fetch_batchOutput",
        -  "type": "object"
        -}New value: +null
    • Addedpaper_graph
    • Changedread_doc1 field changed
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "additionalProperties": true,
        -          "type": "object"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "read_docOutput",
        -  "type": "object"
        -}New value: +null
    • Changedresearch2 fields changed
      • changedInput schema / properties / category / anyOf
        Previous value: -[
        -  {
        -    "enum": [
        -      "news",
        -      "pdf",
        -      "github",
        -      "paper",
        -      "forum",
        -      "blog",
        -      "image",
        -      "dataset"
        -    ],
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "enum": [
        +      "news",
        +      "pdf",
        +      "github",
        +      "paper",
        +      "forum",
        +      "blog",
        +      "image",
        +      "dataset",
        +      "finance",
        +      "software",
        +      "security",
        +      "reference",
        +      "weather",
        +      "docs",
        +      "gov",
        +      "stats",
        +      "calendar",
        +      "news.world",
        +      "paper.index",
        +      "paper.preprint",
        +      "paper.biomed",
        +      "paper.cs",
        +      "paper.math",
        +      "paper.openaccess",
        +      "paper.trial",
        +      "dataset.repository",
        +      "dataset.ml",
        +      "dataset.gov",
        +      "finance.filings",
        +      "finance.market",
        +      "finance.macro",
        +      "finance.fx",
        +      "finance.entity",
        +      "finance.crypto",
        +      "software.lifecycle",
        +      "software.github",
        +      "software.python",
        +      "software.node",
        +      "software.rust",
        +      "software.registry",
        +      "software.app",
        +      "security.cve",
        +      "security.package",
        +      "security.exploited",
        +      "reference.domain",
        +      "docs.web",
        +      "docs.rfc",
        +      "gov.us",
        +      "gov.uk",
        +      "stats.indicator",
        +      "calendar.holidays",
        +      "calendar.clock"
        +    ],
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "additionalProperties": true,
        -          "type": "object"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "researchOutput",
        -  "type": "object"
        -}New value: +null
    • Changedsearch2 fields changed
      • changedInput schema / properties / category / anyOf
        Previous value: -[
        -  {
        -    "enum": [
        -      "news",
        -      "pdf",
        -      "github",
        -      "paper",
        -      "forum",
        -      "blog",
        -      "image",
        -      "dataset"
        -    ],
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "enum": [
        +      "news",
        +      "pdf",
        +      "github",
        +      "paper",
        +      "forum",
        +      "blog",
        +      "image",
        +      "dataset",
        +      "finance",
        +      "software",
        +      "security",
        +      "reference",
        +      "weather",
        +      "docs",
        +      "gov",
        +      "stats",
        +      "calendar",
        +      "news.world",
        +      "paper.index",
        +      "paper.preprint",
        +      "paper.biomed",
        +      "paper.cs",
        +      "paper.math",
        +      "paper.openaccess",
        +      "paper.trial",
        +      "dataset.repository",
        +      "dataset.ml",
        +      "dataset.gov",
        +      "finance.filings",
        +      "finance.market",
        +      "finance.macro",
        +      "finance.fx",
        +      "finance.entity",
        +      "finance.crypto",
        +      "software.lifecycle",
        +      "software.github",
        +      "software.python",
        +      "software.node",
        +      "software.rust",
        +      "software.registry",
        +      "software.app",
        +      "security.cve",
        +      "security.package",
        +      "security.exploited",
        +      "reference.domain",
        +      "docs.web",
        +      "docs.rfc",
        +      "gov.us",
        +      "gov.uk",
        +      "stats.indicator",
        +      "calendar.holidays",
        +      "calendar.clock"
        +    ],
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "additionalProperties": true,
        -          "type": "object"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "searchOutput",
        -  "type": "object"
        -}New value: +null
  2. 4 tool updatesv0.9.1
    • Addeddownload
    • Changedfetch2 fields changed
      • addedInput schema / properties / inline
        Added value: +{
        +  "default": false,
        +  "title": "Inline",
        +  "type": "boolean"
        +}
      • changedOutput schema / (root)
        Previous value: -{
        -  "properties": {
        -    "result": {
        -      "anyOf": [
        -        {
        -          "type": "string"
        -        },
        -        {
        -          "additionalProperties": true,
        -          "type": "object"
        -        }
        -      ],
        -      "title": "Result"
        -    }
        -  },
        -  "required": [
        -    "result"
        -  ],
        -  "title": "fetchOutput",
        -  "type": "object"
        -}New value: +null
    • Changedresearch1 field changed
      • changedInput schema / properties / category / anyOf
        Previous value: -[
        -  {
        -    "enum": [
        -      "news",
        -      "pdf",
        -      "github",
        -      "paper",
        -      "forum",
        -      "blog"
        -    ],
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "enum": [
        +      "news",
        +      "pdf",
        +      "github",
        +      "paper",
        +      "forum",
        +      "blog",
        +      "image",
        +      "dataset"
        +    ],
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
    • Changedsearch1 field changed
      • changedInput schema / properties / category / anyOf
        Previous value: -[
        -  {
        -    "enum": [
        -      "news",
        -      "pdf",
        -      "github",
        -      "paper",
        -      "forum",
        -      "blog"
        -    ],
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "enum": [
        +      "news",
        +      "pdf",
        +      "github",
        +      "paper",
        +      "forum",
        +      "blog",
        +      "image",
        +      "dataset"
        +    ],
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
  3. 2 tool updatesv0.2.0
    • Addedcompare
    • Addedextract_structured
  4. 7 tool updatesv0.1.0
    • First observedcache_search
    • First observedengines
    • First observedfetch
    • First observedfetch_batch
    • First observedread_doc
    • First observedresearch
    • First observedsearch

TDQS

A4.8/5.0

Scored across 11 tools

Disambiguation5/5

Each tool serves a clearly distinct purpose: search for discovery, fetch for single URLs, fetch_batch for multiple URLs, read_doc for document parsing/pagination, research for search+read combo, paper_graph for citation graphs, cache_search for local cache search, engines for engine metadata, compare for side-by-side comparison, extract_structured for metadata, and download for file storage. Cross-references between tools clarify boundaries, so misselection is unlikely.

Naming Consistency4/5

Tool names are all lowercase with underscores, but the pattern isn't uniform: some are simple verbs (fetch, search, research, compare, download), while others are compound (fetch_batch, read_doc, cache_search, paper_graph, extract_structured, engines). This is still predictable and readable, with no mixing of camelCase or inconsistent verb styles, so it's just a minor deviation from a fully consistent verb_noun pattern.

Tool Count5/5

With 11 tools, the server is well-scoped for its purpose (web search, fetching, reading, and research). Each tool earns its place, covering distinct workflows without redundancy. The count is within the ideal 3-15 range, so it feels neither thin nor bloated.

Completeness5/5

The tool surface comprehensively covers the domain: discovery (search, research, engines), fetching (fetch, fetch_batch), reading (read_doc, fetch), local cache querying (cache_search), specialized tasks (paper_graph, extract_structured, compare, download), and pagination. There are no obvious gaps—any common search/read/research workflow has a dedicated tool or a clear combination, and even edge cases like metadata extraction and download are handled.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    A lightweight MCP server that enables LLMs to search the web via DuckDuckGo, search GitHub code repositories, and extract clean content from web pages in LLM-friendly formats.
    8
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    A zero-config web search and fetch MCP server for LLM agents, featuring multi-backend metasearch, persistent rolling cache, and structured error envelopes for retry-friendly interactions.
    MIT
  • A
    license
    Not graded
    quality
    F
    maintenance
    MCP server enabling local-first web search, fetch, extract, and caching with citeable excerpts, no API key required. Supports research workflows for agents and apps.
    5 npm
    1
    MIT