bathys
Summary: Bathys is a local MCP server that gives AI agents deep web research, search, reading, and fact-checking tools without cloud quotas.
deep_research— one-call research: metasearch + parallel page reads + query-tailored distilled digest (with budgets, cache, refresh).web_search— ranked link lists with snippets; supports time range, category, engines, language, JSON output, and automatic retries on blocked/empty results.read_url— read a single page as clean markdown; JS pages via headless browser; optional query-focused distillation, exact substring search over cached raw text, and cache bypass.read_urls— batch-read up to 10 known URLs in parallel under one shared character budget; failed pages cost one line instead of a broken call.library_docs— fetch and distill up-to-date official library documentation from the primary source; supports version pinning and subpage following; repeat calls are cached/offline.source_check— deterministic, no-LLM claim verification against provided URLs or web-found sources, returning SUPPORTED/CONTRADICTED/UNCLEAR/MISSING-EVIDENCE with quoted passages.Operational extras — TTL source cache, robots ethics, metrics/footer compression stats, and swappable internal engines (SearXNG, Crawl4AI).
Uses SearXNG as the search backend to perform web searches and collect raw results, which are then distilled into concise, query-relevant answers.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@bathysDeep research: compare SQLite WAL vs PostgreSQL for 2026"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Bathys
A single local deep-research search service for AI agents. This is a standalone product, not a wrapper over someone else's services: Bathys implements the entire pipeline itself — metasearch with deduplication and resilience to blocking, two-tier extraction (an HTTP engine by default, a headless browser only for JS pages), five-stage query-tailored distillation with hard budgets, a TTL source cache, robots ethics, metrics and diagnostics. Metasearch and extraction are packaged as swappable internal engines (SearXNG, Crawl4AI) — they can be replaced, and the product remains Bathys. No cloud quotas; no LLM inside — synthesis stays with the calling agent, and distillation is deterministic (BM25).
The sonar finds the coordinates, the bathyscaphe dives for the full texts, and the distiller hoists on deck only what answers the question.
Navigation: ⚡ Quick start · 🧹 Removal · 🔌 Harness integration · 🧠 Teach your agent · 🧭 Use cases · 🛠 Tools · 📊 Token savings · 📚 Documentation · 📍 Status
⚡ Quick start
Option 1 — install script (recommended; Python ≥ 3.10):
curl -fsSL https://raw.githubusercontent.com/Korrnals/bathys/main/install.sh | bashThe script installs the package from PyPI into a private venv (~/.local/share/bathys/venv, no sudo), adds it to PATH and runs the full setup. Re-running it is a safe update.
Option 2 — pip (the same thing, done manually):
pip install bathys
bathys setupWhat bathys setup does:
Step | Action |
1 | installs the headless browser — needed only for JS pages (regular pages are read by the built-in HTTP engine) |
2 | registers the MCP server in every harness it finds (zcode, Claude, Cursor, the VS Code family and others — 14 in total, see Harness integration) |
3 | copies the researcher subagent into the found harness directories |
4 | prints the summary and hints ( |
SearXNG does not need to be installed separately — the backend starts automatically on the first search: an external instance is checked first, then docker/podman, then native mode (a clone in BATHYS_SEARXNG_HOME).
npm (Node-first environments; the wrapper installs the Python package itself):
npm install -g bathys-mcp
bathys-mcp setupFrom sources (development):
git clone https://github.com/Korrnals/bathys.git && cd bathys
python3.12 -m venv .venv && .venv/bin/pip install -e .
.venv/bin/bathys setupPinning a specific version — an install-script variable:
BATHYS_INSTALL_VERSION=0.7.0 bash install.shA minimal image without ensurepip: the script and setup bootstrap pip themselves via get-pip.py — details in docs/getting-started/install.md.
🧹 Removal
bathys uninstall # detach Bathys from all harnesses
bathys uninstall hermes zcode # pointwise, only the named ones
bathys uninstall --purge # + delete the venv, cache and datauninstall removes only the bathys entries from harness configs (a *.bathys-backup-* backup is created before any change; foreign servers and subagents are not touched). --purge additionally deletes the ~/.local/share/bathys and ~/.cache/bathys directories; remove the bathys/venv/bin line from .profile/.bashrc manually. Details and backup restore — in the runbook.
Related MCP server: myscrape
🔌 Harness integration
Automatically — the whole stack: bathys setup (see above) registers the server in every harness it finds.
Pointwise — when you need it exactly here:
bathys install # auto-detect all installed harnesses
bathys install hermes # Hermes only (a missing config will be created)
bathys install --list # all supported targets with paths
bathys install --print-config # ready-made blocks for manual pastingDetected: zcode, Claude Code, Claude Desktop, Cursor, the VS Code family (Cline / Roo Code / Kilo Code), Gemini CLI, Windsurf, Zed, opencode, goose, Hermes; each has its own format (JSON schemas and YAML outlines for goose/hermes), and writes are idempotent with a backup. For Pi (badlogic pi-mono), which has no MCP config, there is a drop-in into AGENTS.md. Custom integrations live in the integrations/ directory.
Bathys is a stdio MCP server: the mcpServers block is the same everywhere, and only the file it goes into depends on the harness. command is an absolute path to the binary (~ is not expanded inside JSON); BATHYS_SEARXNG_HOME is optional. Ready-made blocks for every client: bathys install --print-config.
{
"mcpServers": {
"bathys": {
"command": "/path/to/bathys",
"env": { "BATHYS_SEARXNG_HOME": "/path/to/searxng-home" }
}
}
}Harness | Guide |
zcode | |
Claude Code / Claude Desktop | |
Cursor | |
Any other MCP client |
🧠 Teach your agent to work effectively
Configuration is only half the job. Out of the box the harness receives an instructions-playbook (a tool-choice matrix), annotations and three strategy prompts — bathys_deep_research, bathys_source_audit, bathys_fresh_scan — so it picks Bathys tools natively. Stronger still is the profile: the agents/bathys-researcher.md subagent with three skills, to which deep research is delegated wholesale; for clients that do not surface MCP instructions, there is the agents/HARNESS-DROPIN.md drop-in for AGENTS.md / CLAUDE.md / .cursor/rules.
The step-by-step path "out of the box → subagent → drop-in" and the footer-signal table — in "Live cases", section C.
🧭 Common use cases
Technology comparison. Asked "which one to pick for heavy load in 2026?", the agent makes a single deep_research, refines the query with terms from what it found, and verifies the conclusion against two sources: one call instead of a "search + N reads" chain, and 7.5k characters reach the context instead of ~35k.
[bathys: 34 raw hits, top 8 considered · dove 3 pages · 35669 ch fetched → 7508 ch returned · 1.3s]Auditing a contested claim. "Is it true that the benchmarks for X dropped?" — the agent takes the bathys_source_audit strategy: reads the links from the discussion in batch, searches for rebuttals, and delivers a verdict with a URL for each thesis. A dead link costs one line, not a broken call.
A fresh snapshot. "What's new in Y in the last two weeks?" — the bathys_fresh_scan strategy: a search with time_range=week, batch reading, a dated summary; a stale cache HIT is cured by a single refresh=true.
A full walkthrough of all the cases — user-facing, autonomous agents and operations — with live dialogues: docs/getting-started/cases.md.
🛠 Tools
Tool | What it does |
| searches, reads the top sources in parallel, returns a merged query-tailored distillate. The first call for any research question. |
| a ranked list of links with snippets, without page content; |
| reads a page (including text PDFs); with |
| up-to-date library documentation from the primary source, distilled to the question; repeat calls are free (cache). |
| batch-reads up to 10 known pages; the budget is split across the successes, and a failed page costs one line, not a broken call. |
| deterministic claim verification without an LLM: sources → parallel dives → verdict |
Search is narrowed by the shared time_range, category, engines, language filters. The live footer of a response shows compression and cache: [bathys: 41 raw hits, top 3 considered · dove 3 pages · 35669 ch fetched → 7508 ch returned · 3.2s].
📊 Token savings
The scale is honest and character-based; tokens ≈ chars/4; every figure is taken from the footer of a real call.
Call | From the network | To the agent | Compression |
| 37,549 chars | 1,986 chars | 18.9× |
| 17,063 chars | 2,325 chars | 7.3× |
| 35,669 chars | 7,508 chars | 4.7× |
Two-tier extraction — regular pages are read by the built-in HTTP engine (milliseconds, no browser) and JS shells by headless Chromium; then query-tailored distillation with hard character budgets.
Source cache before distillation — SQLite stores the raw text, so re-reading a page from a different angle is free and requires no network.
Zero cloud quotas —
deep_researchreplaces a "search + N reads" chain — that is, N+1 quota charges — with a single local call.
The methodology and thresholds — in docs/operations/metrics.md.
📚 Documentation
Section | What's inside | Who it's for |
The hub: the documentation tree and three reading routes | everyone — the entry point | |
Installation · configuration (20 env) · wiring up · live cases | a newcomer | |
Wiring overview · zcode · Claude Code · Cursor · any MCP client | when wiring up a harness | |
a contributor | ||
an integrator | ||
operations | ||
Charter · features · roadmap · competitors | the product owner | |
Six accepted architectural decisions | a contributor | |
docs authors |
Outside docs/: agents/ — the subagent, skills, drop-in · integrations/ — custom integrations (hermes, pi, zcode) · install.sh — the install script · npm/bathys-mcp/ — the NPM wrapper · tests/ — unit tests · CHANGELOG.md — release history.
📍 Status
0.14.1. Release history: v0.2 "Result Quality" (retries, engine health), v0.3 "Parity with Tavily" (read_urls, JSON mode), v0.4 "Operations" (robots ethics, metrics, bathys-doctor), v0.5 "Identity & Harness" (repositioning, prompts, subagent), v0.6 "Native Install" (bathys install), v0.7 "Ship & Setup" (two-tier extraction, bathys setup, the one-liner, uninstall), v0.8 "Engine Orchestration", v0.9 "Borrowed Ideas", v0.10–0.11 "Library Docs" (Phases 1–2), v0.12 "Deep Verdicts" (source_check + Phase 3), v0.13 "GitHub Tier & Deep Audit", v0.14 "Backlog Closed" — summaries in CHANGELOG.md.
Repository: github.com/Korrnals/bathys. Packages published: PyPI bathys (pip install) and bathys-mcp on npm (npm install -g); the install one-liner is above. CI and releases run through the cluster release conveyor (Korrnals/release-pipeline: verify gates → build → SHA256 → SBOM → cosign + GPG signatures → attach to the GitHub Release; all 15 releases shipped this way — the pre-policy versions retroactively re-attached their signature sets). .github/workflows/ci.yml remains the gate description; GitHub Actions itself is disabled account-wide (billing lock), so its failing runs are expected noise. Before 1.0: the first verified run of the Docker image (the Dockerfile ships with the package; the pipeline's docker build phase needs a runner-image update first).
🙏 Acknowledgements
Bathys stands on the shoulders of outstanding open projects — thanks to their authors and communities:
SearXNG — the metasearch engine (AGPL-3.0): Bathys runs it as a separate process and talks to it over a local JSON API; its sources are not modified and are not distributed inside the package.
Crawl4AI — the browser tier of extraction (Apache-2.0).
MCP Python SDK (MIT), httpx (BSD-3), Playwright (Apache-2.0).
Full attributions and license terms for each component — in THIRD_PARTY_NOTICES.md.
⚖️ License
Bathys code is MIT. The components that Bathys installs and uses are licensed separately and listed in THIRD_PARTY_NOTICES.md (in particular, SearXNG is under AGPL-3.0, with its terms honored).
Available Tools
6 toolsdeep_researchARead-only
Search the web AND read the top sources in one shot.
Runs SearXNG metasearch, dives into the top max_sources pages with a real
browser, distills each page down to passages relevant to query, and
returns one merged digest. Best first call for any research question.
Args:
query: research question or keywords (RU/EN both fine)
max_sources: how many top hits to read in full (1-6)
max_results: how many search hits to consider (1-20)
per_source_chars: per-source character budget (300-8000)
time_range: "day" | "week" | "month" | "year"
category: searxng category, e.g. "general", "news", "science", "it"
language: result language, e.g. "ru", "en", "ru-RU"
refresh: ignore cache and re-fetch search results and pages
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| refresh | No | ||
| category | No | ||
| language | No | ||
| time_range | No | ||
| max_results | No | ||
| max_sources | No | ||
| per_source_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is transparent about the internal process: it runs SearXNG metasearch, opens pages with a real browser, distills relevant passages, and returns a merged digest. It also explains the refresh parameter's cache behavior, and the readOnlyHint annotation matches the read-only nature of the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a headline summary followed by a compact parameter list. It conveys substantial detail without unnecessary fluff, keeping every sentence purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a fairly complex tool, the description covers the search engine, browser-based reading, per-source distillation, merged output, and cache refresh behavior. This is sufficient for an agent to understand what the tool does and what to expect as a result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though the schema has no per-parameter descriptions, the tool description compensates by explaining every parameter: query, max_sources, max_results, per_source_chars, time_range, category, language, and refresh. It includes value ranges and concrete examples, which is more informative than the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a crisp verb-resource statement: 'Search the web AND read the top sources in one shot.' It clearly differentiates this combined tool from the individual sibling tools by emphasizing the one-shot search-and-read workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Best first call for any research question,' giving a strong and direct recommendation for when to use the tool. This makes the usage context clear without needing to infer from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
library_docsARead-only
Fetch up-to-date official documentation for a library and distill it under your question.
Context7-style, but local and unlimited: the docs site is resolved from a
built-in index (or one live web search), fetched from the primary source,
and distilled to passages relevant to query. Repeated questions about
the same library are instant, offline and free (raw-page cache).
Args:
library: library name, e.g. "fastapi", "react", "postgresql", "crawl4ai"
query: your concrete question about the library (used for distillation)
max_chars: output character budget (300-20000)
refresh: re-fetch the docs page even if cached
subpages: when the docs home is navigational, follow this many
same-site subpages ranked by query relevance (0 disables)
version: pin docs to this version (tag, e.g. "0.115.0", "v3", branch
name). Works for GitHub-backed libraries: docs come from that
exact tag on raw.githubusercontent.com. Doc sites are shown at
their latest with an honest note; wrong/missing tag on GitHub
also falls back to latest with a note in the answer
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| library | Yes | ||
| refresh | No | ||
| version | No | ||
| subpages | No | ||
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint and openWorldHint, so the description builds on that with valuable behavioral details: caching behavior, version pinning fallback ('wrong/missing tag on GitHub also falls back to latest with a note'), subpages navigation logic, and the fact that output is distilled passages rather than full docs. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but well-structured: a clear opening sentence, a Context7-style comparison, a bullet list of parameters, and notes on version behavior. Every sentence contributes meaning; it is detailed but not wasteful. Slight redundancy ('local and unlimited' vs caching) but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 params, output schema present), the description is remarkably complete. It covers parameter semantics, edge cases (version fallback), caching, distillation, and output constraints. An agent has all necessary information to call the tool correctly, including what to expect in the answer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It explains every parameter: library with examples, query for distillation, max_chars as character budget, refresh for cache bypass, subpages with '0 disables', and version with tag examples and fallback behavior. This fully compensates for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch up-to-date official documentation for a library and distill it under your question.' It clearly identifies the tool's unique function (distillation) and distinguishes it from general web search or URL reading by its focus on library docs and local caching.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong usage context: it explains the built-in index, caching benefits ('instant, offline and free'), and the distillation process. It mentions 'Context7-style' as a comparison but does not explicitly name sibling tools or state when NOT to use them. However, the context is clear enough for an agent to infer appropriate usage for library documentation queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_urlARead-only
Read one web page; return its main content as clean, budgeted markdown.
JS-rendered pages are handled by a real headless browser. Boilerplate
(nav, footer, ads) is stripped; if query is given, only passages relevant
to it are returned. Pages are cached — re-reads with a different query are
instant and cost no network.
Args:
url: absolute http(s) URL
query: optional focus; return only passages relevant to it
max_chars: output character budget (300-50000)
refresh: ignore cache and re-fetch the page
find: search the cached RAW text for this exact substring (case-insensitive):
returns matches with counters and context, no network needed. Requires
the page to have been read before; combine with query for first reads.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| find | No | ||
| query | No | ||
| refresh | No | ||
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring readOnlyHint and openWorldHint, the description adds substantial behavioral detail: headless browser handling, boilerplate stripping, query filtering, caching with no network on re-reads, and exact find semantics with counters and context. This goes well beyond the annotations and fully discloses side effects like cache usage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-organized with a clear intro, a behavior paragraph, and an args list. It slightly repeats 'query' behavior between the intro and the args list, but every sentence earns its place. The front-loaded purpose and structured parameter explanations make it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, no schema descriptions, and an output schema available, the description fully equips an agent: it covers parameter semantics, caching, network behavior, and the specific semantics of find. The presence of an output schema means return value details don't need duplication. Nothing necessary for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains every parameter: url format, query filtering, max_chars budget with range, refresh ignoring cache, and find's case-insensitive exact match with requirements (must have been read before). This adds rich meaning the schema completely lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Read one web page; return its main content as clean, budgeted markdown.' This clearly differentiates it from sibling tools like read_urls (plural) and web_search. The additional behaviors (JS rendering, boilerplate stripping) make the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context for when to use parameters (query for relevance, refresh to bypass cache, find for cached text) but never explicitly states when to prefer this tool over siblings like read_urls or web_search. Storage and caching behavior indirectly imply single-page reads, but no alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_urlsARead-only
Read several known web pages in one call under one shared character budget.
Pages are fetched in parallel (JS-rendered, boilerplate-stripped) and the combined total_chars budget is split evenly between the pages that came back. Prefer this over N read_url calls when you already hold the URLs: one round-trip, one budget, and a failed page costs one line instead of a failed call. Args: urls: 1-10 absolute http(s) URLs; duplicates (after utm/fragment cleanup) are merged, extras beyond 10 are reported in a Skipped line query: optional focus; each page is distilled to passages relevant to it total_chars: combined output budget across all sections (300-30000) refresh: ignore cache and re-fetch every page
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | ||
| query | No | ||
| refresh | No | ||
| total_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint is consistent with the read operation, and the description adds useful behavioral details: parallel fetching, JS rendering, boilerplate stripping, even budget splitting, duplicate merging, and skipped extras. Failure behavior is also disclosed as a failed page costing one line rather than a failed call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficient and well structured, with the high-level behavior, use case, and parameter details each getting their own section. There is minor repetition of ideas such as 'one call' and 'one round-trip', but nothing that significantly hurts clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for an agent to select and invoke the tool correctly: it covers purpose, use case, parameter semantics, failure behavior, and important operational details. The presence of an output schema and readOnly/openWorld annotations reduces the need for additional output or side-effect explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema itself lacks property descriptions, the tool description fully compensates by explaining each parameter: urls constraints and deduplication, query as optional focus, total_chars as the combined budget, and refresh as cache bypass. This goes well beyond the bare schema and gives an agent everything needed to set parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states precisely what the tool does: reads several known web pages in one call under a shared character budget. This clearly distinguishes it from sibling tools like read_url, web_search, and deep_research by emphasizing the batch and known-URL aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly recommends this tool over repeated read_url calls when the URLs are already known, and gives concrete reasons: one round-trip, one budget, and partial failure handling. This gives an agent clear, actionable selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
source_checkARead-only
Deterministically verify a claim against web sources; return a verdict with quoted passages.
No LLM involved: sources are read (yours via urls, or found by a web
search on the claim), distilled under the claim, and scored lexically —
polar markers (with negation handling) decide SUPPORTED / CONTRADICTED /
UNCLEAR / MISSING-EVIDENCE. Best for checking a fact, assertion or rumour
when you need a reproducible verdict with citations, not a narrative.
Args:
claim: the statement to verify, in your own words (RU/EN both fine)
urls: optional 1-10 http(s) URLs to check against; without them the
sources are found by a web search on the claim
max_sources: how many sources to consider (1-10)
refresh: ignore cache and re-fetch the search results and pages
| Name | Required | Description | Default |
|---|---|---|---|
| urls | No | ||
| claim | Yes | ||
| refresh | No | ||
| max_sources | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description fully discloses behavior: no LLM involvement, extraction from user-provided or search-found sources, lexical scoring with negation handling, possible verdict categories, use of cache, and the refresh flag to bypass caching. This is rich and accurate context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized with a clear summary first, then behavior, then a compact Args list. Every sentence adds value, and no information is repeated from the schema. It is detailed without being bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexityched, the description is complete: it explains the deterministic method, verdict types, source selection behavior, all parameters, and the appropriate use case. The output schema exists, so return-value details do not need to be repeated in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates fully by documenting all four parameters in the Args block. It adds constraints not present in the schema (e.g., urls 1-10, max_sources 1-10, refresh meaning) and clarifies supported languages for claim.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Deterministically verify a claim against web sources; return a verdict with quoted passages.' It also distinguishes itself from siblings by emphasizing reproducibility, lexical scoring, and verdicts rather than narrative or open-ended research.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the intended use case: 'Best for checking a fact, assertion or rumour when you need a reproducible verdict with citations, not a narrative.' It also explains when to supply URLs versus rely on web search. It does not name sibling tools explicitly, but the 'not a narrative' exclusion gives clear decision context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchARead-only
Search the web via SearXNG metasearch; return a compact ranked link list.
Returns title, URL and a short snippet per hit — no page content. Empty or blocked results are retried automatically with other engine sets. To actually read pages, call read_url; to do both at once, call deep_research. Args: query: search query (natural language or keywords) max_results: 1-20 time_range: "day" | "week" | "month" | "year" category: e.g. "general", "news", "science", "it", "files" engines: comma-separated engine names, e.g. "google,bing,duckduckgo" language: e.g. "ru", "en", "ru-RU" refresh: ignore cache and re-run the search as_json: return pure machine-readable JSON {query, count, hits[], answer?} instead of the human-friendly list (no footer line)
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| as_json | No | ||
| engines | No | ||
| refresh | No | ||
| category | No | ||
| language | No | ||
| time_range | No | ||
| max_results | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that results are compact, that page content is not returned, and that empty/blocked results are retried automatically. The readOnlyHint annotation covers side effects; no contradictory claims are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized with a short overview, behavior note, and per-parameter list. There is slight redundancy between 'compact ranked link list' and 'Returns title, URL and a short snippet per hit,' but overall it is efficient and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool, the description covers the main use case, output format, retry behavior, and key parameters. It does not explain the human-friendly list format in detail, but the schema and sibling-tool references provide enough context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema itself has no parameter descriptions, but the Args section adds concise semantics for every parameter, including examples for category, engines, language, and time_range. It also clarifies the as_json return shape with the JSON placeholder.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Search the web via SearXNG metasearch') and explicitly identifies the output shape ('compact ranked link list'). It also distinguishes itself from siblings by name and behavior, making selection clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use this tool versus alternatives: 'To actually read pages, call read_url; to do both at once, call deep_research.' This removes ambiguity about whether the tool returns page content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.8.0- Added
library_docs - Changed
read_url1 field changed- added
Input schema / properties / findAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Find" +}
- Added
source_check
4 tool updates
v0.5.0- First observed
deep_research - First observed
read_url - First observed
read_urls - First observed
web_search
TDQS
Scored across 6 tools
Each tool targets a distinct workflow: search-only, read-one, read-many, combined deep research, library docs, and claim verification. The boundaries between read_url, read_urls, and deep_research are explicitly documented, so an agent can select without ambiguity.
Names mix conventions: read_url and read_urls use verb_noun, while deep_research, web_search, library_docs, and source_check lead with nouns or adjectives. There is no consistent verb-first pattern, making the set feel uneven despite all being snake_case.
Six tools cover the research/reading domain without redundancy or bloat; each adds a distinct capability, from search to verification. The count is well-scoped for a specialized server.
The surface covers the full research loop: search, read single, read many, fetch docs, and verify claims. No critical gap exists for the stated purpose of deep web research.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Free web search for AI agents. No API key required. Hosted MCP in active development.
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides local-first web intelligence over MCP with tools for search, fetch, crawl, extract, cache, find-similar, research, and autonomous agent loops, requiring no API keys.10813 npm5,410AGPL 3.0
- AlicenseNot gradedqualityCmaintenanceA self-contained web-research MCP server that lets local LLM agents search, fetch, and synthesize web content using tools like web_search, web_fetch, and web_research.2MIT
- AlicenseNot gradedqualityFmaintenanceMCP server enabling local-first web search, fetch, extract, and caching with citeable excerpts, no API key required. Supports research workflows for agents and apps.5 npm1MIT
- FlicenseAqualityBmaintenanceExposes web search and page fetching tools via the MCP protocol, allowing integration with AI editors like Cursor for autonomous research workflows.2-