Cluefinch MCP
OfficialProvides web search capabilities through a configured SearXNG instance, allowing agents to run and refine queries, choose search engines and language, use Safe Search, restrict searches to or exclude domains, and detect when search engines fail to respond.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Cluefinch MCPresearch the latest MCP protocol updates and summarize key changes"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Cluefinch MCP

Deep Research infrastructure for AI agents.
Cluefinch MCP gives AI agents a complete toolkit for working with the web through the Model Context Protocol (MCP).
With Cluefinch, an agent can search the web, read pages, navigate between related sources, and carry out in-depth research.
Cluefinch MCP can make the internet part of your AI agent's workspace — from finding a single fact to carrying out complex, multi-step research across multiple sources.
Cluefinch MCP is completely free to use and requires no paid subscription. It works with AI agents whether they are powered by local LLMs or cloud-based models, and does not require a commercial search API or a subscription to a cloud-hosted Deep Research service.
Cluefinch MCP does not impose its own limits on the number of search queries or research runs.
Your agent gets the tools it needs to work effectively with the web, while you retain control over how those tools are used. Cluefinch integrates easily with MCP-compatible AI tools and fits into the workflow you already use.
What Cluefinch MCP can do
Web Search
Cluefinch MCP lets an agent search the internet through your own SearXNG instance.
The agent can formulate and refine search queries, use different search engines, restrict searches by language or domain, and issue follow-up or revised queries when needed.
Web Reading
Once a useful link is found, the agent can open the page through Cluefinch MCP and receive cleaned text ready for model processing.
Large pages do not have to be loaded into the model's context all at once. The agent can read them in chunks, continue from a specific position, and request additional context only when it is actually needed.
Source Navigation
After finding a useful page, the agent can inspect its HTTP/HTTPS links and use the source's own structure to continue the research: move through documentation sections and report chapters, follow pagination, open related pages, and reach primary materials — without returning to a search engine at every step.
Deep Research
Cluefinch MCP gives the agent the tools to carry out multi-source research workflows.
The agent can pursue several lines of inquiry at once, work with both search results and known URLs, gather material from different sources, examine the most relevant parts of long documents, and deepen the investigation as new questions emerge.
The agent remains in control of the research process: it decides what to search for next, which sources deserve closer inspection, how to interpret the collected material, and when there is enough evidence to produce an answer.
The depth of the research — from a quick product lookup to complex, multi-stage analysis — depends on the task, the model, and the user's instructions.

Related MCP server: mcp-web
Quick Start
Cluefinch MCP requires Python 3.12.4 or later.
1. Install Cluefinch MCP
python -m pip install cluefinchAfter installation, make sure the cluefinch executable is available through the PATH environment variable.
On Windows:
where.exe cluefinchOn macOS and Linux:
command -v cluefinchIf cluefinch is not found, add the directory containing the installed executable to PATH.
2. Start SearXNG
Cluefinch uses SearXNG as its search backend. If you do not already have your own SearXNG instance, the repository includes a ready-to-use local configuration example:
examples/searxng/compose.yamlexamples/searxng/settings.yml
To run the example, you need Docker with Docker Compose support.
Download these files into a separate directory and create a .env file alongside them with a random SEARXNG_SECRET.
On macOS and Linux:
printf 'SEARXNG_SECRET=%s\n' "$(openssl rand -hex 32)" > .envOn Windows PowerShell:
$secret = -join ((1..64) | ForEach-Object { '{0:x}' -f (Get-Random -Maximum 16) })
"SEARXNG_SECRET=$secret" | Set-Content -Encoding ascii .envThen start SearXNG with Docker Compose:
docker compose up -dBy default, the local SearXNG instance will be available at:
http://127.0.0.1:80813. Connect Cluefinch MCP to your AI tool
Cluefinch connects easily to popular AI tools and runs as a standard local MCP server over stdio.
Ready-to-use connection examples are provided in the next section.
Integrating with AI tools
Cluefinch MCP uses the standard local MCP stdio transport, so in most clients you only need to specify the cluefinch command.
Below are minimal examples of integrating Cluefinch with some popular AI tools. Cluefinch can also be used with other clients that support local MCP servers over stdio.
The examples below use SearXNG at http://127.0.0.1:8081, as shown in the Quick Start section above.
If your SearXNG instance does not use the default address, pass MCP_SEARCH_SEARXNG_URL through the MCP server environment in your client's configuration.
Cursor
Add Cluefinch to your project's .cursor/mcp.json or to Cursor's global MCP configuration:
{
"mcpServers": {
"cluefinch": {
"type": "stdio",
"command": "cluefinch"
}
}
}After restarting the MCP connection, Cursor will discover the Cluefinch tools and can use them in agent tasks.
Claude Code
Cluefinch can be added with a single command:
claude mcp add --scope user cluefinch -- cluefinchTo verify the connection:
claude mcp listCodex
Add Cluefinch with:
codex mcp add cluefinch -- cluefinchTo verify the connection:
codex mcp listGitHub Copilot in VS Code
Add the local MCP server to .vscode/mcp.json:
{
"servers": {
"cluefinch": {
"type": "stdio",
"command": "cluefinch"
}
}
}Cluefinch will then become available to GitHub Copilot in agent mode as a set of MCP tools.
OpenCode
Add Cluefinch to opencode.jsonc:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"cluefinch": {
"type": "local",
"command": ["cluefinch"],
"enabled": true
}
}
}Qwen Code
Add Cluefinch to ~/.qwen/settings.json for user-level configuration, or to .qwen/settings.json for a specific project:
{
"mcpServers": {
"cluefinch": {
"command": "cluefinch",
"args": []
}
}
}After restarting Qwen Code, you can verify the connection with the /mcp command.
OpenClaw
Add Cluefinch with:
openclaw mcp add cluefinch --command cluefinchOr configure it manually in openclaw.json:
{
"mcp": {
"servers": {
"cluefinch": {
"command": "cluefinch",
"transport": "stdio"
}
}
}
}To verify the connection:
openclaw mcp probe cluefinchHermes Agent
Add Cluefinch to the config.yaml file used by your active Hermes profile:
mcp_servers:
cluefinch:
command: "cluefinch"
args: []Restart Hermes after saving the configuration.
Configuring agent behavior
The agent formulates search queries, selects sources, and manages the research process using the available Cluefinch MCP tools. You can define your own rules for that process through instructions in AGENTS.md or equivalent settings in your AI tool.
Those instructions can govern both ordinary searches and deep, multi-stage research, including how individual Cluefinch MCP tools should be used.
For example, you can ask the agent to begin with a research plan, run several refinement queries through web_search, use web_links to navigate chapters and related pages, prioritize primary sources, and read large documents incrementally with web_fetch.
For multi-source collection and initial filtering, the agent can use research_collect. Your instructions can also define rules for reusing already collected context, checking conflicting sources, limiting additional search iterations, and keeping factual evidence separate from interpretation.
The repository includes an AGENTS.md file with a ready-made example of this kind of Deep Research workflow. You can use it as a starting point, simplify it for quick research, or adapt it to your own tasks, model, and output requirements.
Cluefinch MCP tools
At the MCP level, Cluefinch exposes four complementary tools that take an agent from web search to reading specific sources and then to multi-source research.
web_search
web_search runs searches through the configured SearXNG instance and returns links to potentially useful sources.
The agent can control the number of results, choose search engines and language, use Safe Search, restrict the search to a particular domain, or exclude unwanted domains. Cluefinch also reports when one or more SearXNG engines fail to respond, so the agent does not mistake an incomplete result set for a complete one.
web_fetch
web_fetch reads web pages directly. Cluefinch downloads the HTML from the specified URL, extracts the main text, converts it to Markdown, and returns only the portion the agent needs.
The agent can first request a small preview of the page to judge whether the source is useful, then continue reading only if needed. A long document can be read incrementally from a chosen position. Cluefinch returns next_start and a ready-to-use continuation action, so the agent does not need to calculate the next position manually. Along with the content, it receives the metadata needed to continue navigating the same retained version of the extracted text safely.
web_links
web_links extracts navigational HTTP/HTTPS links from an HTML page and returns them in document order. It lets the agent inspect the structure of a source it has already found: documentation sections, report chapters, pagination, appendices, primary-source references, and related pages.
This is especially useful in Deep Research. After finding a strong source, the agent can inspect its structure, open only the relevant sections with web_fetch, and then collect material from several selected sources with research_collect. This reduces unnecessary searches, helps preserve research context, and enables deeper work with primary materials.
Relative links are resolved into absolute URLs using the document's base URL. Links can be filtered by origin when needed, and large link sets can be retrieved incrementally across multiple requests.
Cluefinch also versions each retained link set. If the page's navigation changes between requests, the agent will not continue from stale positions in an outdated list.
Extracting links does not make requests to their destinations. Full outbound safety validation is applied only if the agent later decides to fetch one of those URLs.
research_collect
research_collect is designed to work with multiple sources at once. You can provide several search queries, specific URLs, or both.
Cluefinch gathers available sources, deduplicates documents by their final URLs after redirects, and identifies the most relevant passages in long documents. Each selected passage remains an exact slice of the extracted text with stable coordinates, so the agent can return to it later and request additional context when needed.
Each source receives a stable source_id, while collection problems — such as an unreachable URL, a failed search, or a duplicate final document — are reported explicitly in gaps. The agent decides whether enough material has been collected and which sources deserve deeper inspection.
Full tool schemas, parameters, limits, and response semantics are documented in docs/REFERENCE.md.
What makes Cluefinch efficient and safe
Behind the four Cluefinch MCP tools is a retrieval layer that handles long documents, repeated requests, network constraints, and safe access to external content.

Efficient use of context and tokens
Cluefinch lets the agent send only the portion of a page needed for the current task to the model, rather than loading the entire document into context.
An extracted document can be read incrementally. The agent receives the position of the next chunk and continues only when more content is actually needed. For individual research passages, it can also request more surrounding context without rereading the entire document.
Navigation uses positions in the retained extracted text, allowing the agent to return precisely to previously identified passages. This helps the model use its context window more efficiently and spend tokens only on the parts of a source that matter to the current stage of the work.
Version control for extracted text
Every retained version of extracted text receives a content_hash.
When the agent continues reading or expands a previously selected passage using expected_content_hash, Cluefinch verifies that hash. Ready-to-use actions for continuation and expansion pass it automatically. If the page content has changed and the old coordinates can no longer be considered reliable, the tool reports the change instead of returning an outdated passage from the old position.
This makes continued reading of changing sources more reliable and reduces the risk of silently mixing passages from different versions of a document.
Source provenance and traceability
Cluefinch preserves the final URL after redirects and deduplicates sources again against the final document address. As a result, different links that lead to the same material do not become separate independent sources.
Each normalized final URL receives a stable source_id, allowing the same source to be identified consistently across different stages of the research process.
For long documents, Cluefinch splits extracted text into bounded passages and ranks them for relevance with BM25. Selected passages remain exact slices of the extracted source text with coordinates, so the model receives source material that can later be revisited and expanded with additional context.
Explicit limitations instead of hidden assumptions
Cluefinch explicitly reports conditions that may affect the completeness of retrieved data.
research_collect returns gaps when some sources cannot be retrieved or processed. web_search separately reports SearXNG engines that did not respond. When reading a page, Cluefinch also distinguishes between cases where more retained text remains available and cases where the end of the extracted document was discarded because of the configured retention limit.
This makes incomplete retrieval visible to the agent instead of presenting a partial result as if it were complete. The agent still decides whether enough information has been collected to continue the analysis or produce an answer.
Caching and data reuse
Search results and extracted pages are temporarily stored in local per-process TTL/LRU caches. While a cached entry remains valid, requesting the same resource again avoids another HTTP request. Page URLs are still validated before cache reuse, which can involve DNS lookups.
Identical concurrent searches and fetches of the same page are coalesced so parallel agent actions do not create duplicate network traffic.
Caching complements context management: incremental reading helps conserve model tokens, while the local cache avoids repeatedly downloading the same data from the internet.
Safe access to external pages
An agent can receive links from search results and arbitrary websites, so Cluefinch treats every URL as potentially untrusted.
Before retrieving content, Cluefinch validates the URL scheme, hostname, DNS results, and final IP addresses. Local, private, reserved, multicast, and other unsafe addresses are blocked. Validation is repeated after redirects and again immediately before connection. Cluefinch connects to an already validated numeric IP while preserving the original hostname for HTTP and TLS.
Cluefinch also limits response and decompressed data size, redirect count, download time, concurrent network activity, and the resources used for text extraction. Requests to the same host are also spaced over time.
These measures are primarily designed to protect against SSRF and uncontrolled resource consumption. The text of a web page is still untrusted content and should not automatically be treated by an agent as an instruction.
Separation of search and web-page retrieval
Cluefinch uses SearXNG only as a search backend. It helps discover potential sources but is not used as a proxy for reading web pages.
When the agent opens a discovered URL, Cluefinch retrieves the page directly through its own protected fetch layer. Keeping discovery and retrieval separate allows security, caching, text extraction, and long-document reading to be managed independently.
SearXNG remains a separate service, while Cluefinch MCP runs locally in the user's environment and does not require its own cloud retrieval service or built-in telemetry.
Configuring Cluefinch MCP
Cluefinch MCP can be tuned to a particular environment and agent workload through environment variables prefixed with MCP_SEARCH_.
In most cases, the defaults are sufficient. The main settings are:
Setting | Default | Purpose |
|
| SearXNG address |
|
| Allowed explicit engine subset; omitting |
|
| Maximum number of results per search |
|
| Maximum number of sources |
|
| Maximum number of search queries in one |
|
| Maximum size of a single returned page fragment |
|
| Maximum amount of extracted text retained for one page |
|
| Maximum number of concurrent page fetches |
|
| Lifetime of fetched pages in the local cache, in seconds |
|
| Lifetime of search results in the local cache, in seconds |
For example, if SearXNG is running at a different address:
export MCP_SEARCH_SEARXNG_URL=http://127.0.0.1:8888The same variables can also be passed directly through the MCP client's server configuration.
The complete list of settings, defaults, and exact behavior is documented in docs/REFERENCE.md.
Current limitations
Cluefinch MCP is designed for searching and retrieving ordinary web pages over HTTP/HTTPS. The current version supports HTML/XHTML and text extraction without running a full browser.
Cluefinch MCP currently does not include:
JavaScript rendering or browser automation;
PDF text extraction;
authenticated sessions or private pages;
CAPTCHA or paywall bypass;
vector search or embedding-based retrieval;
hosted SaaS, a REST API, or built-in telemetry.
If a page depends almost entirely on JavaScript or is unavailable without authentication, the agent should look for an alternative HTML source, public documentation, a mirror, or another accessible source.
These limitations apply to the current version of Cluefinch MCP and help keep the architecture local, predictable, and under your control.
Development setup
To work with the source code, you need uv and Python 3.12.4 or later.
git clone https://github.com/cluefinch/mcp-server.git
cd mcp-server
uv python install 3.12
uv sync --lockedYou can then run the server directly from the working tree:
uv run cluefinchIf you need a local SearXNG instance for development, use the ready-made configuration example. If you already have your own SearXNG instance, simply set its address through MCP_SEARCH_SEARXNG_URL.
The main project checks are:
uv run ruff check mcp_search tests scripts
uv run ruff format --check mcp_search tests scripts
uv run pytest -q
uvx --from 'pyright==1.1.414' pyright --pythonpath .venv/bin/python mcp_search scripts
uv buildCI tests the supported Python versions on Ubuntu and Windows and also verifies the lower bounds of direct dependencies.
Detailed information about the local development environment, smoke tests, IDE setup, agent behavior evaluation, and publishable-tree requirements is available in docs/DEVELOPMENT.md.
Contributing, support, and security
If you want to propose a change, report a problem, or contribute to the project, start with CONTRIBUTING.md. It describes the requirements for pull requests, tests, and the Developer Certificate of Origin (DCO).
For usage questions and support, see SUPPORT.md.
If you discover a vulnerability or another security-related issue, do not publish the details in a regular GitHub Issue. The responsible reporting process is described in SECURITY.md.
When creating public Issues or Pull Requests, do not publish secrets, private URLs, content from non-public pages, local capture files, or other sensitive data.
License
The original Cluefinch MCP code is distributed under the Apache License 2.0.
SearXNG is used as a separate external service and remains licensed under GNU AGPL-3.0. The Apache-2.0 license for Cluefinch MCP does not apply to SearXNG, its dependencies, or the content of web pages retrieved by the agent.
Additional information about third-party components and their licenses is available in THIRD_PARTY_NOTICES.md, while SearXNG integration and distribution considerations are documented in docs/SEARXNG_COMPLIANCE.md.
See NOTICE for copyright and attribution information.
Available Tools
4 toolsresearch_collectARead-only
Collect bounded evidence from several selected or searched sources.
USE THIS for multi-source evidence gathering when you have explicit URLs,
several search formulations, or both. It searches/fetches candidate documents,
extracts literal passages and reports acquisition gaps. It does NOT crawl links
found inside those documents. If a strong source must first be navigated to a
chapter, appendix or related page, use web_links and then pass the selected
URLs here or to web_fetch.
Supply queries=["variant one", "variant two"] and/or urls=["https://..."].
topic only ranks excerpts; topic alone does not search. Explicit URLs consume
the source budget first. If they already fill max_sources, no search is run.
Remaining candidate slots are distributed round-robin across query variants.
For domain/exclude_domains filtering, use web_search first and pass selected
result URLs here.
Returns raw material, not synthesis: query_variants, sources and gaps. Each
source has stable source_id, final URL, source_type, content_hash,
text_truncated and literal excerpts with Unicode offsets. BM25 selects passages;
source_type and ranking are not credibility judgments.
For more context around an excerpt, copy excerpt.expand.arguments into the
indicated tool; URL, offset, budget and expected_content_hash are already
supplied. content_changed returns no stale-coordinate slice.
Always inspect gaps before judging coverage. They can report failed/empty
searches, unresponsive engines, blocked/unreadable pages, insufficient text,
redirect duplicates, omitted candidates and truncation. Empty gaps do not prove
topic completeness. Expected request failures return {error, hint}; individual
source failures can coexist with successful sources.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | No | Exact source URLs to fetch first; useful after web_search with domain/exclude_domains. Supply urls and/or queries. At most 100 input URLs; max_sources still controls the fetch budget. | |
| topic | No | Optional topic used only to rank excerpts alongside queries. Does not initiate a search without queries. | |
| queries | No | Your search variants, normally 2-3; at most 10. The server does not generate queries. May be combined with explicit URLs. | |
| language | No | Language/locale forwarded to search; unused in URL-only mode. | |
| time_range | No | Search recency: day, week, month or year. Does not assert a source publication date. | |
| max_sources | No | Candidate fetch budget: default 5, configured cap 10. Failed/duplicate candidates can leave fewer successful sources. Explicit URLs take priority. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only readOnlyHint and openWorldHint; the description carries the rest, disclosing budget semantics (explicit URLs consume the source budget first, round-robin distribution over remaining slots, no search run if URLs fill max_sources), failure semantics (empty gaps don't prove completeness; {error, hint} for expected failures; source failures coexisting with successes), and non-judgment caveats (BM25 ranking and source_type are not credibility judgments). That is substantive behavior beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the primary use rule before mechanics; almost every sentence carries a constraint, routing rule, or failure caveat. It is nonetheless long and dense for a collection tool, and a few clauses (excerpt expansion, content_changed) sit below the critical usage block.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description usefully summarizes the return shape (query_variants, sources, gaps, stable source_id, offsets, content_hash) and the excerpt.expand handoff, plus the content_changed signal. Combined with the usage and failure guidance, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are already documented, but the description adds cross-parameter semantics the schema cannot express: URLs are consumed before queries, remaining candidate slots are distributed round-robin across query variants, and topic alone ranks excerpts without initiating a search. Some points (max_sources cap, topic-does-not-search) restate the schema, so it is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (collect bounded evidence from selected/searched sources) and immediately bounds scope: it searches/fetches candidates, extracts literal passages, reports gaps, and explicitly does NOT crawl links found inside documents. It is clearly distinguishable from web_search, web_fetch and web_links, all three of which are named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions (multi-source gathering with explicit URLs, several query formulations, or both), a when-not rule (no crawling of in-document links), and explicit routing to siblings: use web_links first if a source must be navigated to a chapter/appendix, and use web_search first for domain/exclude_domains filtering. Alternatives and the conditions selecting them are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_fetchARead-only
Read one selected HTML document as bounded Markdown text.
USE THIS when you already know which document you want to read, preview, or
expand around a collected excerpt. If you know the source but need to find
its chapter, appendix, next page, reference, or related document first, use web_links instead.
If you do not yet know a source, use web_search.
Preview with max_chars=500 to limit model context; this does not reduce the
initial network download. When more retained text is useful, copy
continuation.arguments into the indicated tool rather than recalculating the
next offset. Null continuation means the retained text has ended.
continuation and excerpt.expand use expected_content_hash. If the retained
text version changed, content_changed returns no slice: reacquire the document
and choose coordinates again. A matching hash protects coordinates, not
origin freshness.
A page can have too little readable text for web_fetch while still exposing
useful ordinary HTML links through web_links; treat such pages as possible
navigation hubs instead of assuming the source has no usable structure.
Returns content, title, final URL, source category, Unicode-coordinate paging,
cache/hash fields and continuation. truncated means more retained text can be
continued; text_truncated means a tail was discarded by the hard retention
ceiling and cannot be recovered by pagination. Expected failures return
{error, hint}.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | HTTP(S) page URL, a search result.url or research source.url. Supports HTML/XHTML; no PDF or JavaScript rendering. | |
| max_chars | No | Maximum returned characters: default 10000, configured cap 20000. Use 500 for preview. Ready actions include a budget; smaller output saves context, not the initial download. | |
| start_offset | No | Unicode-character offset; default 0. Ready continuation/expand arguments provide it. For manual reading, use next_start or excerpt.start_char. | |
| expected_content_hash | No | Optional retained-text version guard. Ready actions supply it; mismatch returns content_changed without a slice. Omit for unguarded reading. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover readOnlyHint and openWorldHint, and the description goes well beyond them: preview does not reduce the initial download, continuation/excerpt.expand require expected_content_hash, a mismatch returns content_changed with no slice, and the hash protects coordinates rather than origin freshness. It also distinguishes truncated from text_truncated (recoverable vs unrecoverable tail) and notes expected failures return {error, hint}.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then layered usage, preview guidance, hash semantics, and return-shape notes; each paragraph carries distinct operational information. It is dense and slightly long, but almost no sentence is redundant with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter read tool with an output schema and annotations, the description covers routing, paging, truncation semantics, hash guarding, and failure shape. Even the edge case of a page with too little readable text (treat as a navigation hub, try web_links) is addressed, leaving no operational gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents each parameter; the description nevertheless adds cross-parameter semantics the schema lacks, notably that the hash guard protects coordinates rather than content freshness and that continuation.arguments should be copied rather than the next offset recalculated. It stops short of adding new per-parameter syntax, so it stays just above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one selected HTML document as bounded Markdown text') and explicitly differentiates itself from both siblings: web_links for finding a chapter/page within a known source, web_search for when no source is known. An agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('you already know which document you want to read, preview, or expand around a collected excerpt') plus two named alternatives with the exact conditions that select them. Also prescribes the preview workflow (max_chars=500) and the continuation protocol.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_linksARead-only
Navigate from a known HTML page to pages that it explicitly links.
USE THIS when you already found a useful source but need its structure:
documentation sections, report chapters, table-of-contents entries,
next/previous publication pages, appendices, references, datasets, standards,
or related pages. It is the bridge between web_search discovery and web_fetch reading.
Prefer this over another search when the desired page is plausibly
linked from the source you already have.
This is NOT a crawler or browser. It inspects ordinary <a href> links in this
one fetched HTML page only. It does not follow them, click controls, execute
JavaScript, submit forms, DNS-resolve destinations, or safety-approve them.
Selecting a returned URL for web_fetch/web_links later triggers the normal
outbound SSRF policy.
same_origin=true keeps only links with the same scheme, normalized hostname
and effective port as the final fetched page. Leave same_origin=null when
external citations or primary sources may matter; same origin does not mean
same organization, and cross-origin does not mean untrusted.
The retained list preserves document order after deterministic URL
deduplication. Fragments remain part of link identity. same_document=true
means only that the destination has the same document identity ignoring the
fragment; web_fetch still navigates text by Unicode offsets, not HTML anchors.
Filtering happens BEFORE pagination. Copy continuation.arguments to retrieve
the next retained link page. links_hash versions the whole retained ordered
list before filtering/pagination; links_changed means the old start_index must
not be reused. continuation means more retained links exist; links_truncated
means some admissible links were lost to hard page-level resource ceilings
and continuation cannot recover them.
Each link returns url, best-effort label, rel, fragment, same_origin and
same_document. An empty links list does not prove that the site has no other
pages. Expected failures return {error, hint}.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Known HTTP(S) HTML page whose exposed navigation you want to inspect. Use after finding a useful source when you need chapters, pagination, references or related pages. The page itself is fetched safely; returned destinations are not contacted until selected later. | |
| max_links | No | Maximum links returned: default 50, hard cap 100. Ready continuation preserves the effective budget. | |
| same_origin | No | Navigation filter relative to the final fetched page. true=same scheme+normalized host+effective port only; false=cross-origin only; null=keep both (default, best when external references may matter). Filtering happens before pagination. | |
| start_index | No | Zero-based index in the filtered retained link list. Ready continuation arguments provide it. | |
| expected_links_hash | No | Optional retained-link-list version guard. Ready continuations supply it; mismatch returns links_changed without using the old index. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description discloses concrete non-behavior: it does not follow links, click controls, execute JavaScript, submit forms, DNS-resolve destinations, or safety-approve them. It also explains that later selection triggers SSRF policy, describes pagination/hash/continuation semantics, and notes that an empty list does not prove absence of other pages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and usage, then moves through limitations, filtering, pagination, and output in a logical order. It is long, but the complexity of the tool justifies most of the detail; a few points repeat schema content rather than adding new value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, substantial annotations, and the presence of an output schema, the description is complete enough for correct invocation. It covers routing, safety boundaries, continuation behavior, hash guards, truncation caveats, returned fields, and error shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds semantic nuance beyond the schema for same_origin (same origin does not mean same organization; cross-origin does not mean untrusted), same_document versus fragment identity, and links_hash behavior. Much of this overlaps with the schema's own parameter descriptions, so it is helpful but not entirely additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: navigating from a known HTML page to pages it explicitly links. It distinguishes the tool from web_search and web_fetch by calling itself the bridge between discovery and reading, and explicitly contrasts itself with a crawler or browser.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance ('USE THIS when you already found a useful source but need its structure') and when-not guidance ('This is NOT a crawler or browser'). It also names the sibling alternative indirectly by saying to prefer it over another search when the target is plausibly linked from the current source.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchARead-only
Discover candidate sources through search; this does not fetch page content.
USE THIS when the needed source or URL is not yet known, when you need a new
independent source, or when search-engine discovery itself is required.
DO NOT search again merely to find another page inside a strong source when
that page is likely linked from the current page; use web_links instead.
Domain filtering example: web_search(query="asyncio", domain="python.org",
exclude_domains=["discuss.python.org"], max_results=20). Domain operators
depend on the provider. Search pagination is not supported.
Results are candidates, not verified page evidence: snippet is a search-engine
preview and result.url has not been fetched. For a candidate:
- web_links(url=result.url): inspect chapters, pagination, references or related
pages exposed by that source without crawling them;
- web_fetch(url=result.url, max_chars=500): preview/read the selected document;
- research_collect(urls=[...]): collect bounded evidence from selected URLs.
Returns results, query_used, unresponsive_engines and cached. Always inspect
unresponsive_engines: nonempty means incomplete search coverage. Expected
failures return {error, hint}.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search terms. For several query variants and fetched excerpts, use research_collect instead. | |
| domain | No | Include site:domain in the query, e.g. docs.python.org. Use a hostname, not a page URL. Can be combined with exclude_domains. | |
| engines | No | Optional engine subset: google, google cse, brave, wikipedia, wikidata. Omit to use SearXNG defaults. | |
| language | No | Language/locale such as en or ru. Omit for any language. | |
| time_range | No | Optional recency filter: day, week, month or year. | |
| max_results | No | Requested candidates. Default 10; configured cap 20. This is a result limit, not search pagination. | |
| safe_search | No | 0=off, 1=moderate, 2=strict. Omit for provider defaults. | |
| exclude_domains | No | Exclude domains with -site: operators, e.g. ['pinterest.com', 'example.org']. Provider support varies; verify returned URLs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover read-only/open-world safety, and the description adds substantial context beyond them: no pagination, results are unfetched candidates with preview-only snippets, unresponsive_engines must be inspected for incomplete coverage, and failures return {error, hint}.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and routing before examples and result caveats. The follow-up bullet list is longer than strictly necessary but each item maps to a real next-step decision, so waste is limited.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite an output schema existing, the description adds actionable value by flagging unresponsive_engines as a coverage signal. With 8 params, sibling routing, and result-interpretation guidance all covered, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 8 params (baseline 3). The description adds a concrete combined-call example (query + domain + exclude_domains + max_results) and the caveat that domain operators are provider-dependent, which is meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Discover candidate sources through search') and immediately negates the sibling behavior it does not perform ('this does not fetch page content'), which cleanly separates it from web_fetch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit USE THIS when-conditions are given (source/URL unknown, need independent source, discovery required) plus an explicit DO NOT with the alternative named ('use web_links instead'). research_collect is also routed to for multi-variant queries, so the agent never has to infer tool choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.4- First observed
research_collect - First observed
web_fetch - First observed
web_links - First observed
web_search
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: web_search for discovery, web_fetch for reading one document, web_links for navigating known HTML pages, and research_collect for multi-source evidence gathering. The descriptions explicitly state when to use each and include 'DO NOT' guidance to prevent misuse, leaving no ambiguity.
Three tools follow a web_ prefix pattern, but research_collect deviates. Verb forms are also mixed: web_search and web_fetch are verb-based, web_links is noun-based, and research_collect combines noun+verb. All are snake_case and readable, but the pattern is not fully consistent.
Four tools is well-scoped for a web research server. Each tool serves a distinct, non-redundant function—discovery, single-document reading, link navigation, and multi-source collection—so every tool earns its place without bloat or thinness.
The surface covers the core lifecycle of web research: discovery, navigation, reading, and multi-source evidence gathering. A minor gap is the absence of any crawling or bulk link-following tool, though the descriptions intentionally exclude crawling, making it a reasonable scope boundary rather than a critical omission.
Maintenance
Related MCP Connectors
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
Web scraping, Google SERP, Markdown and ChatGPT/Gemini answers as typed tools for AI agents.
171
Related MCP Servers
- AlicenseAqualityBmaintenanceLocal on-demand web search and page reading for coding agents via SearXNG, with SSRF-protected fetching and Markdown extraction.2MIT
- FlicenseAqualityCmaintenanceEnables a locally-run LLM to search the web, fetch pages as markdown, make arbitrary HTTP requests, and optionally render pages with headless Chromium.3-
- AlicenseAqualityAmaintenanceEnables AI agents to run local deep-research workflows via a single MCP stdio server, combining web search, page extraction, query-aware distillation, and caching without cloud quotas. It exposes tools for deep research, search, and single or batch URL reading.679 PyPIMIT
- AlicenseNot gradedqualityCmaintenanceProvides LLM-free search, fetch, and rerank tools over any reachable SearXNG instance, returning extracted, deduplicated, cross-encoder-reranked results with citation-ready markers and coverage signals. Lets agent hosts such as Open WebUI, Claude, or tool-calling models drive their own iterative research loop while sharing one deterministic retrieval and ranking core.MIT