Skip to main content
Glama

Bathys

A single local deep-research search service for AI agents. This is a standalone product, not a wrapper over someone else's services: Bathys implements the entire pipeline itself — metasearch with deduplication and resilience to blocking, two-tier extraction (an HTTP engine by default, a headless browser only for JS pages), five-stage query-tailored distillation with hard budgets, a TTL source cache, robots ethics, metrics and diagnostics. Metasearch and extraction are packaged as swappable internal engines (SearXNG, Crawl4AI) — they can be replaced, and the product remains Bathys. No cloud quotas; no LLM inside — synthesis stays with the calling agent, and distillation is deterministic (BM25).

The sonar finds the coordinates, the bathyscaphe dives for the full texts, and the distiller hoists on deck only what answers the question.

python version mcp license

English | Russian

Navigation: ⚡ Quick start · 🧹 Removal · 🔌 Harness integration · 🧠 Teach your agent · 🧭 Use cases · 🛠 Tools · 📊 Token savings · 📚 Documentation · 📍 Status

⚡ Quick start

Option 1 — install script (recommended; Python ≥ 3.10):

curl -fsSL https://raw.githubusercontent.com/Korrnals/bathys/main/install.sh | bash

The script installs the package from PyPI into a private venv (~/.local/share/bathys/venv, no sudo), adds it to PATH and runs the full setup. Re-running it is a safe update.

Option 2 — pip (the same thing, done manually):

pip install bathys
bathys setup

What bathys setup does:

Step

Action

1

installs the headless browser — needed only for JS pages (regular pages are read by the built-in HTTP engine)

2

registers the MCP server in every harness it finds (zcode, Claude, Cursor, the VS Code family and others — 14 in total, see Harness integration)

3

copies the researcher subagent into the found harness directories

4

prints the summary and hints (bathys doctor — self-diagnostics)

SearXNG does not need to be installed separately — the backend starts automatically on the first search: an external instance is checked first, then docker/podman, then native mode (a clone in BATHYS_SEARXNG_HOME).

npm (Node-first environments; the wrapper installs the Python package itself):

npm install -g bathys-mcp
bathys-mcp setup

From sources (development):

git clone https://github.com/Korrnals/bathys.git && cd bathys
python3.12 -m venv .venv && .venv/bin/pip install -e .
.venv/bin/bathys setup

Pinning a specific version — an install-script variable:

BATHYS_INSTALL_VERSION=0.7.0 bash install.sh

A minimal image without ensurepip: the script and setup bootstrap pip themselves via get-pip.py — details in docs/getting-started/install.md.

🧹 Removal

bathys uninstall               # detach Bathys from all harnesses
bathys uninstall hermes zcode   # pointwise, only the named ones
bathys uninstall --purge        # + delete the venv, cache and data

uninstall removes only the bathys entries from harness configs (a *.bathys-backup-* backup is created before any change; foreign servers and subagents are not touched). --purge additionally deletes the ~/.local/share/bathys and ~/.cache/bathys directories; remove the bathys/venv/bin line from .profile/.bashrc manually. Details and backup restore — in the runbook.

Related MCP server: myscrape

🔌 Harness integration

Automatically — the whole stack: bathys setup (see above) registers the server in every harness it finds.

Pointwise — when you need it exactly here:

bathys install                # auto-detect all installed harnesses
bathys install hermes         # Hermes only (a missing config will be created)
bathys install --list         # all supported targets with paths
bathys install --print-config # ready-made blocks for manual pasting

Detected: zcode, Claude Code, Claude Desktop, Cursor, the VS Code family (Cline / Roo Code / Kilo Code), Gemini CLI, Windsurf, Zed, opencode, goose, Hermes; each has its own format (JSON schemas and YAML outlines for goose/hermes), and writes are idempotent with a backup. For Pi (badlogic pi-mono), which has no MCP config, there is a drop-in into AGENTS.md. Custom integrations live in the integrations/ directory.

Bathys is a stdio MCP server: the mcpServers block is the same everywhere, and only the file it goes into depends on the harness. command is an absolute path to the binary (~ is not expanded inside JSON); BATHYS_SEARXNG_HOME is optional. Ready-made blocks for every client: bathys install --print-config.

{
  "mcpServers": {
    "bathys": {
      "command": "/path/to/bathys",
      "env": { "BATHYS_SEARXNG_HOME": "/path/to/searxng-home" }
    }
  }
}

🧠 Teach your agent to work effectively

Configuration is only half the job. Out of the box the harness receives an instructions-playbook (a tool-choice matrix), annotations and three strategy prompts — bathys_deep_research, bathys_source_audit, bathys_fresh_scan — so it picks Bathys tools natively. Stronger still is the profile: the agents/bathys-researcher.md subagent with three skills, to which deep research is delegated wholesale; for clients that do not surface MCP instructions, there is the agents/HARNESS-DROPIN.md drop-in for AGENTS.md / CLAUDE.md / .cursor/rules.

The step-by-step path "out of the box → subagent → drop-in" and the footer-signal table — in "Live cases", section C.

🧭 Common use cases

Technology comparison. Asked "which one to pick for heavy load in 2026?", the agent makes a single deep_research, refines the query with terms from what it found, and verifies the conclusion against two sources: one call instead of a "search + N reads" chain, and 7.5k characters reach the context instead of ~35k.

[bathys: 34 raw hits, top 8 considered · dove 3 pages · 35669 ch fetched → 7508 ch returned · 1.3s]

Auditing a contested claim. "Is it true that the benchmarks for X dropped?" — the agent takes the bathys_source_audit strategy: reads the links from the discussion in batch, searches for rebuttals, and delivers a verdict with a URL for each thesis. A dead link costs one line, not a broken call.

A fresh snapshot. "What's new in Y in the last two weeks?" — the bathys_fresh_scan strategy: a search with time_range=week, batch reading, a dated summary; a stale cache HIT is cured by a single refresh=true.

A full walkthrough of all the cases — user-facing, autonomous agents and operations — with live dialogues: docs/getting-started/cases.md.

🛠 Tools

Tool

What it does

deep_research(query, max_sources=3, …)

searches, reads the top sources in parallel, returns a merged query-tailored distillate. The first call for any research question.

web_search(query, max_results=8, …)

a ranked list of links with snippets, without page content; as_json=true — clean JSON for programs.

read_url(url, query=None, find=None)

reads a page (including text PDFs); with query — relevant passages; with find — exact search over the source cache without the network.

library_docs(library, query)

up-to-date library documentation from the primary source, distilled to the question; repeat calls are free (cache).

read_urls(urls, query=None, total_chars=12000)

batch-reads up to 10 known pages; the budget is split across the successes, and a failed page costs one line, not a broken call.

source_check(claim, urls?)

deterministic claim verification without an LLM: sources → parallel dives → verdict SUPPORTED/CONTRADICTED/UNCLEAR/MISSING-EVIDENCE with confidence; source classes, a domain-independence cap.

Search is narrowed by the shared time_range, category, engines, language filters. The live footer of a response shows compression and cache: [bathys: 41 raw hits, top 3 considered · dove 3 pages · 35669 ch fetched → 7508 ch returned · 3.2s].

📊 Token savings

The scale is honest and character-based; tokens ≈ chars/4; every figure is taken from the footer of a real call.

Call

From the network

To the agent

Compression

web_search

37,549 chars

1,986 chars

18.9×

read_url

17,063 chars

2,325 chars

7.3×

deep_research (3 pages)

35,669 chars

7,508 chars

4.7×

  • Two-tier extraction — regular pages are read by the built-in HTTP engine (milliseconds, no browser) and JS shells by headless Chromium; then query-tailored distillation with hard character budgets.

  • Source cache before distillation — SQLite stores the raw text, so re-reading a page from a different angle is free and requires no network.

  • Zero cloud quotas — deep_research replaces a "search + N reads" chain — that is, N+1 quota charges — with a single local call.

The methodology and thresholds — in docs/operations/metrics.md.

📚 Documentation

Section

What's inside

Who it's for

docs/index.md

The hub: the documentation tree and three reading routes

everyone — the entry point

getting-started

Installation · configuration (20 env) · wiring up · live cases

a newcomer

integrations

Wiring overview · zcode · Claude Code · Cursor · any MCP client

when wiring up a harness

architecture

Components · the cleaning pipeline · data flows

a contributor

contracts

Tools · output formats · modules · configuration

an integrator

operations

Runbook · token-savings metrics

operations

product

Charter · features · roadmap · competitors

the product owner

adr

Six accepted architectural decisions

a contributor

meta

Docs style guide · glossary

docs authors

Outside docs/: agents/ — the subagent, skills, drop-in · integrations/ — custom integrations (hermes, pi, zcode) · install.sh — the install script · npm/bathys-mcp/ — the NPM wrapper · tests/ — unit tests · CHANGELOG.md — release history.

📍 Status

0.14.1. Release history: v0.2 "Result Quality" (retries, engine health), v0.3 "Parity with Tavily" (read_urls, JSON mode), v0.4 "Operations" (robots ethics, metrics, bathys-doctor), v0.5 "Identity & Harness" (repositioning, prompts, subagent), v0.6 "Native Install" (bathys install), v0.7 "Ship & Setup" (two-tier extraction, bathys setup, the one-liner, uninstall), v0.8 "Engine Orchestration", v0.9 "Borrowed Ideas", v0.10–0.11 "Library Docs" (Phases 1–2), v0.12 "Deep Verdicts" (source_check + Phase 3), v0.13 "GitHub Tier & Deep Audit", v0.14 "Backlog Closed" — summaries in CHANGELOG.md.

Repository: github.com/Korrnals/bathys. Packages published: PyPI bathys (pip install) and bathys-mcp on npm (npm install -g); the install one-liner is above. CI and releases run through the cluster release conveyor (Korrnals/release-pipeline: verify gates → build → SHA256 → SBOM → cosign + GPG signatures → attach to the GitHub Release; all 15 releases shipped this way — the pre-policy versions retroactively re-attached their signature sets). .github/workflows/ci.yml remains the gate description; GitHub Actions itself is disabled account-wide (billing lock), so its failing runs are expected noise. Before 1.0: the first verified run of the Docker image (the Dockerfile ships with the package; the pipeline's docker build phase needs a runner-image update first).

🙏 Acknowledgements

Bathys stands on the shoulders of outstanding open projects — thanks to their authors and communities:

  • SearXNG — the metasearch engine (AGPL-3.0): Bathys runs it as a separate process and talks to it over a local JSON API; its sources are not modified and are not distributed inside the package.

  • Crawl4AI — the browser tier of extraction (Apache-2.0).

  • MCP Python SDK (MIT), httpx (BSD-3), Playwright (Apache-2.0).

Full attributions and license terms for each component — in THIRD_PARTY_NOTICES.md.

⚖️ License

Bathys code is MIT. The components that Bathys installs and uses are licensed separately and listed in THIRD_PARTY_NOTICES.md (in particular, SearXNG is under AGPL-3.0, with its terms honored).

Available Tools

6 tools
deep_researchA
Read-only

Search the web AND read the top sources in one shot.

Runs SearXNG metasearch, dives into the top max_sources pages with a real browser, distills each page down to passages relevant to query, and returns one merged digest. Best first call for any research question. Args: query: research question or keywords (RU/EN both fine) max_sources: how many top hits to read in full (1-6) max_results: how many search hits to consider (1-20) per_source_chars: per-source character budget (300-8000) time_range: "day" | "week" | "month" | "year" category: searxng category, e.g. "general", "news", "science", "it" language: result language, e.g. "ru", "en", "ru-RU" refresh: ignore cache and re-fetch search results and pages

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
refreshNo
categoryNo
languageNo
time_rangeNo
max_resultsNo
max_sourcesNo
per_source_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description is transparent about the internal process: it runs SearXNG metasearch, opens pages with a real browser, distills relevant passages, and returns a merged digest. It also explains the refresh parameter's cache behavior, and the readOnlyHint annotation matches the read-only nature of the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a headline summary followed by a compact parameter list. It conveys substantial detail without unnecessary fluff, keeping every sentence purposeful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a fairly complex tool, the description covers the search engine, browser-based reading, per-source distillation, merged output, and cache refresh behavior. This is sufficient for an agent to understand what the tool does and what to expect as a result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though the schema has no per-parameter descriptions, the tool description compensates by explaining every parameter: query, max_sources, max_results, per_source_chars, time_range, category, language, and refresh. It includes value ranges and concrete examples, which is more informative than the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a crisp verb-resource statement: 'Search the web AND read the top sources in one shot.' It clearly differentiates this combined tool from the individual sibling tools by emphasizing the one-shot search-and-read workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Best first call for any research question,' giving a strong and direct recommendation for when to use the tool. This makes the usage context clear without needing to infer from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

library_docsA
Read-only

Fetch up-to-date official documentation for a library and distill it under your question.

Context7-style, but local and unlimited: the docs site is resolved from a built-in index (or one live web search), fetched from the primary source, and distilled to passages relevant to query. Repeated questions about the same library are instant, offline and free (raw-page cache). Args: library: library name, e.g. "fastapi", "react", "postgresql", "crawl4ai" query: your concrete question about the library (used for distillation) max_chars: output character budget (300-20000) refresh: re-fetch the docs page even if cached subpages: when the docs home is navigational, follow this many same-site subpages ranked by query relevance (0 disables) version: pin docs to this version (tag, e.g. "0.115.0", "v3", branch name). Works for GitHub-backed libraries: docs come from that exact tag on raw.githubusercontent.com. Doc sites are shown at their latest with an honest note; wrong/missing tag on GitHub also falls back to latest with a note in the answer

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
libraryYes
refreshNo
versionNo
subpagesNo
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint and openWorldHint, so the description builds on that with valuable behavioral details: caching behavior, version pinning fallback ('wrong/missing tag on GitHub also falls back to latest with a note'), subpages navigation logic, and the fact that output is distilled passages rather than full docs. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but well-structured: a clear opening sentence, a Context7-style comparison, a bullet list of parameters, and notes on version behavior. Every sentence contributes meaning; it is detailed but not wasteful. Slight redundancy ('local and unlimited' vs caching) but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 params, output schema present), the description is remarkably complete. It covers parameter semantics, edge cases (version fallback), caching, distillation, and output constraints. An agent has all necessary information to call the tool correctly, including what to expect in the answer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the full burden. It explains every parameter: library with examples, query for distillation, max_chars as character budget, refresh for cache bypass, subpages with '0 disables', and version with tag examples and fallback behavior. This fully compensates for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Fetch up-to-date official documentation for a library and distill it under your question.' It clearly identifies the tool's unique function (distillation) and distinguishes it from general web search or URL reading by its focus on library docs and local caching.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong usage context: it explains the built-in index, caching benefits ('instant, offline and free'), and the distillation process. It mentions 'Context7-style' as a comparison but does not explicitly name sibling tools or state when NOT to use them. However, the context is clear enough for an agent to infer appropriate usage for library documentation queries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_urlA
Read-only

Read one web page; return its main content as clean, budgeted markdown.

JS-rendered pages are handled by a real headless browser. Boilerplate (nav, footer, ads) is stripped; if query is given, only passages relevant to it are returned. Pages are cached — re-reads with a different query are instant and cost no network. Args: url: absolute http(s) URL query: optional focus; return only passages relevant to it max_chars: output character budget (300-50000) refresh: ignore cache and re-fetch the page find: search the cached RAW text for this exact substring (case-insensitive): returns matches with counters and context, no network needed. Requires the page to have been read before; combine with query for first reads.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
findNo
queryNo
refreshNo
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint and openWorldHint, the description adds substantial behavioral detail: headless browser handling, boilerplate stripping, query filtering, caching with no network on re-reads, and exact find semantics with counters and context. This goes well beyond the annotations and fully discloses side effects like cache usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized with a clear intro, a behavior paragraph, and an args list. It slightly repeats 'query' behavior between the intro and the args list, but every sentence earns its place. The front-loaded purpose and structured parameter explanations make it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, no schema descriptions, and an output schema available, the description fully equips an agent: it covers parameter semantics, caching, network behavior, and the specific semantics of find. The presence of an output schema means return value details don't need duplication. Nothing necessary for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden. It explains every parameter: url format, query filtering, max_chars budget with range, refresh ignoring cache, and find's case-insensitive exact match with requirements (must have been read before). This adds rich meaning the schema completely lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Read one web page; return its main content as clean, budgeted markdown.' This clearly differentiates it from sibling tools like read_urls (plural) and web_search. The additional behaviors (JS rendering, boilerplate stripping) make the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context for when to use parameters (query for relevance, refresh to bypass cache, find for cached text) but never explicitly states when to prefer this tool over siblings like read_urls or web_search. Storage and caching behavior indirectly imply single-page reads, but no alternatives or exclusions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_urlsA
Read-only

Read several known web pages in one call under one shared character budget.

Pages are fetched in parallel (JS-rendered, boilerplate-stripped) and the combined total_chars budget is split evenly between the pages that came back. Prefer this over N read_url calls when you already hold the URLs: one round-trip, one budget, and a failed page costs one line instead of a failed call. Args: urls: 1-10 absolute http(s) URLs; duplicates (after utm/fragment cleanup) are merged, extras beyond 10 are reported in a Skipped line query: optional focus; each page is distilled to passages relevant to it total_chars: combined output budget across all sections (300-30000) refresh: ignore cache and re-fetch every page

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
queryNo
refreshNo
total_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation readOnlyHint is consistent with the read operation, and the description adds useful behavioral details: parallel fetching, JS rendering, boilerplate stripping, even budget splitting, duplicate merging, and skipped extras. Failure behavior is also disclosed as a failed page costing one line rather than a failed call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficient and well structured, with the high-level behavior, use case, and parameter details each getting their own section. There is minor repetition of ideas such as 'one call' and 'one round-trip', but nothing that significantly hurts clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for an agent to select and invoke the tool correctly: it covers purpose, use case, parameter semantics, failure behavior, and important operational details. The presence of an output schema and readOnly/openWorld annotations reduces the need for additional output or side-effect explanation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema itself lacks property descriptions, the tool description fully compensates by explaining each parameter: urls constraints and deduplication, query as optional focus, total_chars as the combined budget, and refresh as cache bypass. This goes well beyond the bare schema and gives an agent everything needed to set parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool does: reads several known web pages in one call under a shared character budget. This clearly distinguishes it from sibling tools like read_url, web_search, and deep_research by emphasizing the batch and known-URL aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly recommends this tool over repeated read_url calls when the URLs are already known, and gives concrete reasons: one round-trip, one budget, and partial failure handling. This gives an agent clear, actionable selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

source_checkA
Read-only

Deterministically verify a claim against web sources; return a verdict with quoted passages.

No LLM involved: sources are read (yours via urls, or found by a web search on the claim), distilled under the claim, and scored lexically — polar markers (with negation handling) decide SUPPORTED / CONTRADICTED / UNCLEAR / MISSING-EVIDENCE. Best for checking a fact, assertion or rumour when you need a reproducible verdict with citations, not a narrative. Args: claim: the statement to verify, in your own words (RU/EN both fine) urls: optional 1-10 http(s) URLs to check against; without them the sources are found by a web search on the claim max_sources: how many sources to consider (1-10) refresh: ignore cache and re-fetch the search results and pages

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsNo
claimYes
refreshNo
max_sourcesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description fully discloses behavior: no LLM involvement, extraction from user-provided or search-found sources, lexical scoring with negation handling, possible verdict categories, use of cache, and the refresh flag to bypass caching. This is rich and accurate context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized with a clear summary first, then behavior, then a compact Args list. Every sentence adds value, and no information is repeated from the schema. It is detailed without being bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexityched, the description is complete: it explains the deterministic method, verdict types, source selection behavior, all parameters, and the appropriate use case. The output schema exists, so return-value details do not need to be repeated in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates fully by documenting all four parameters in the Args block. It adds constraints not present in the schema (e.g., urls 1-10, max_sources 1-10, refresh meaning) and clarifies supported languages for claim.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Deterministically verify a claim against web sources; return a verdict with quoted passages.' It also distinguishes itself from siblings by emphasizing reproducibility, lexical scoring, and verdicts rather than narrative or open-ended research.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states the intended use case: 'Best for checking a fact, assertion or rumour when you need a reproducible verdict with citations, not a narrative.' It also explains when to supply URLs versus rely on web search. It does not name sibling tools explicitly, but the 'not a narrative' exclusion gives clear decision context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.8.0
    • Addedlibrary_docs
    • Changedread_url1 field changed
      • addedInput schema / properties / find
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Find"
        +}
    • Addedsource_check
  2. 4 tool updatesv0.5.0
    • First observeddeep_research
    • First observedread_url
    • First observedread_urls
    • First observedweb_search

TDQS

A4.5/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct workflow: search-only, read-one, read-many, combined deep research, library docs, and claim verification. The boundaries between read_url, read_urls, and deep_research are explicitly documented, so an agent can select without ambiguity.

Naming Consistency2/5

Names mix conventions: read_url and read_urls use verb_noun, while deep_research, web_search, library_docs, and source_check lead with nouns or adjectives. There is no consistent verb-first pattern, making the set feel uneven despite all being snake_case.

Tool Count5/5

Six tools cover the research/reading domain without redundancy or bloat; each adds a distinct capability, from search to verification. The count is well-scoped for a specialized server.

Completeness5/5

The surface covers the full research loop: search, read single, read many, fetch docs, and verify claims. No critical gap exists for the stated purpose of deep web research.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Provides local-first web intelligence over MCP with tools for search, fetch, crawl, extract, cache, find-similar, research, and autonomous agent loops, requiring no API keys.
    10
    813 npm
    5,410
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    A self-contained web-research MCP server that lets local LLM agents search, fetch, and synthesize web content using tools like web_search, web_fetch, and web_research.
    2
    MIT
  • A
    license
    Not graded
    quality
    F
    maintenance
    MCP server enabling local-first web search, fetch, extract, and caching with citeable excerpts, no API key required. Supports research workflows for agents and apps.
    5 npm
    1
    MIT