camoufox-research
Camoufox Research is an MCP server that gives an AI agent a real anti-detect browser for end-to-end web research — search, read, interact, extract, and export without browser code.
Search the web (DuckDuckGo, cached, paginated, snippets) and academic sources (arXiv/Semantic Scholar).
Run deep research: one-shot
researchor background campaigns (research_start,research_status,research_report,research_resume) targeting N distinct domains with quality ranking and LLM planning.Read pages:
fetch_page(single, 24h cache, delta),batch_fetch(10–50 URLs parallel),crawl(BFS),map_site,sitemap,rss, andread_document(PDF/DOCX/XLSX).Extract structured data:
extractvia CSS/XPath/LLM,table_extractto CSV,check_links,page_difffor change monitoring.Interact via live sessions: navigate, click, type, fill forms, upload/download, tabs, scroll, JS eval, network/console inspection, wait for elements.
Capture visuals:
snapshot(compact interactive-element tree with refs) andscreenshot(PNG, Set-of-Mark).Manage state: profiles (cookies/localStorage), runtime proxy switching.
Export and report:
exportto JSON/CSV/MD;citation_pack(verified sources),citation_report(MD file),research_digest(summaries + verification),research_critic(claim review).Observe and route:
stats,tool_usage,tool_hint,service_route,ping.
Enables web searching via DuckDuckGo through an anti-detect browser, providing search results with titles, URLs, and optional snippets.
Allows reading RSS/Atom feeds to retrieve posts with title, link, and date from blogs, news sites, and changelogs.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@camoufox-researchFind the latest news on NVIDIA stock and give a brief summary"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Camoufox Research
An MCP server that gives your AI agent a real browser — search the web, read JS/SPA pages, click, fill forms, crawl sites, extract tables, save files, watch for changes. Runs on the anti-detect Camoufox (Firefox), so pages see a normal browser instead of a headless bot.
По-русски: MCP-сервер для веб-ресёрча. Даёт агенту живой браузер — поиск, чтение тяжёлых страниц, клики, формы, сбор данных. Ставится одной командой.

46 tools total, 18 in the default profile — a profile is just the set of
tool groups you switch on (--caps); fewer tools make the agent choose better.
AI agent ──MCP──▶ Camoufox Research ──▶ Camoufox (anti-detect Firefox) ──▶ webWhy
Most MCP browser servers hand an agent isolated actions. This one is a complete research toolkit in a single server — the agent searches, reads, interacts and exports without you writing browser code.
🔎 Search — live web search that keeps working when a source is blocked, plus deep
researchthat gathers 10+ sources across many distinct sites in one call🌐 Browse — JS/SPA text, live sessions with tabs, clicks, forms, uploads
👁️ See pages — Set-of-Mark screenshots and a compact snapshot tree with
refs📊 Extract — CSS/XPath fields, tables → CSV, PDF/DOCX/XLSX, export JSON/MD
Related MCP server: camofox-browser-mcp
30-second demo
"Find all pricing pages on this site, extract the prices and save them to CSV."
Agent
├─ map_site discover every /pricing page
├─ crawl read them (cached)
├─ extract {"plan": "css:.plan", "price": "css:.price"}
└─ export format=csv → prices.csvNo browser-automation code — just a sentence to your agent.
Install (one command)
Requirements: Python 3.10–3.13 and git (Windows: also PowerShell 7). The installer creates a venv, installs the package from the clone, downloads the browser once, registers the MCP server and checks the handshake. There is no PyPI package — the code comes from this repo or Docker.
OS | One command (from the clone root) |
Linux (Fedora/Ubuntu/Debian/Arch/…) |
|
macOS |
|
Windows (native, PowerShell 7) |
|
Docker (any OS with Docker) |
|
No clone, one command (uv) or the prebuilt image
uv installs the package straight from this repo — no clone, no PyPI:
uv tool install git+https://github.com/aidvizhhub/camoufox-research
camoufox-research # stdio MCP server; browser downloads on first runOne-shot without installing (uvx) — handy for an MCP client:
uvx --from git+https://github.com/aidvizhhub/camoufox-research camoufox-researchThe image is prebuilt in GHCR (browser already inside, nothing to fetch):
docker run -i --rm -v camoufox-data:/data ghcr.io/aidvizhhub/camoufox-research:latestserver.json carries the MCP Registry metadata and points at that image. There
is no PyPI package by design — git, uv, or Docker.
Preview first — it changes nothing:
bash scripts/install.sh --dry-run # plan for THIS machine
pwsh -NoProfile -File scripts\install.ps1 -WhatIf # Windows planNo clone yet? The same installer pulls the repo to ~/camoufox-research:
curl -fsSL https://raw.githubusercontent.com/aidvizhhub/camoufox-research/main/scripts/install.sh | bashWindows: irm …/install.ps1 | iex. Full matrix, flags and troubleshooting —
docs/install-crossplatform.md; native Windows
details — docs/install-windows.md.
Check after install:
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py # caps → handshake → tools count
opencode mcp list # opencode: camoufox connected
claude mcp list # Claude Code; Claude Desktop / Cursor — status in their UIUpdate, or remove (cache and browser stay):
bash scripts/install/update_mcp.sh # pull → pip → reconnect
bash scripts/install.sh --reinstall # reinstall the package from the clone
bash scripts/install.sh --uninstall # venv + MCP entry; cache/browser keptBy hand, if you already have a clone and want full control:
python3 -m venv ~/.venvs/camoufox-research
~/.venvs/camoufox-research/bin/pip install .
~/.venvs/camoufox-research/bin/python -m camoufox fetch # browser, oncepip install . installs the code on disk — not from PyPI. Optional extras:
pip install '.[geoip]' adds IP-based geolocation, locale and timezone for
proxied runs (only active with a proxy and CAMOUFOX_GEOIP=1). The common env
vars (CAMOUFOX_*: cache dir, timeouts, caps, report dir) are documented in
configs/example.env; the advanced switches (stealth, SSRF,
budgets, LLM planner, transports) live in the code — list them all with
grep -rhoE 'CAMOUFOX_[A-Z0-9_]+' camoufox_research/ | sort -u.
Connect to MCP
The installer writes the client config for you. To do it by hand, the canon is
opencode v2: the config is ~/.config/opencode/opencode.jsonc and servers
live under mcp.servers.<name> — not mcp.<name>, and disabled — not
enabled (the old v1 shape is silently ignored by v2).
OpenCode (~/.config/opencode/opencode.jsonc): copy-ready section —
mcp/config/opencode.jsonc.example.
Claude Desktop and Cursor use their own shape (mcpServers + command string +
env); copy-ready files are in mcp/config/
(claude_desktop_config.json.example, cursor_mcp.json.example). Trying it from
sources still means installing once (the MCP SDK lives in the venv); then run the
console script from that venv:
~/.venvs/camoufox-research/bin/camoufox-research. Don't launch mcp/server.py
from the repo root — the local mcp/ folder shadows the SDK and the import
fails.
Installed with uv (no clone)? Use as command (same opencode v2 shape):
["uvx", "--from", "git+https://github.com/aidvizhhub/camoufox-research", "camoufox-research"].
Check — CLI opencode mcp list → ✓ camoufox connected; inside a session the
/mcps command shows the same status. A one-page walkthrough (path, paste,
troubleshooting) is in
mcp/config/opencode-connect.md.
Tools (46)
Tools are grouped, and the groups double as --caps profiles (see below).
Group | Tools | Count |
|
| 3 |
|
| 13 |
|
| 26 |
|
| 2 |
always on |
| 2 |
A few things worth knowing:
Search runs a chain of independent engines: one that is blocked or returns nothing cools down and the query moves on; results are cached per engine+query, so a blocked source never takes
web_searchdown. For engines that need a proxy there is a pool (CAMOUFOX_SEARCH_PROXY/CAMOUFOX_SEARCH_PROXIES).fetch_pagereuses a 24 h cache — a repeat costs no network.delta=Truereturns a marker when nothing changed (saves answer tokens) but still does a fresh fetch — use it for "what changed", not for cheap repeats.snapshotreturns a ~2–5 KB YAML tree of interactive elements with arefon each — click byref, no fragile selectors.researchruns a deep hunt in one call: query expansion, quality ranking (docs/GitHub/arXiv first), domain dedup, and honestpartialwhen the goal isn't reached. For a single fact, useweb_search.exportwrites JSON / CSV / Markdown to disk.
Tool profiles (fewer tools, better choice)
46 tools in one prompt degrade an agent's tool choice (industry rule of thumb:
past ~40 it drops). Pick the groups you need — the same idea as Playwright MCP's
--caps:
camoufox-research --caps research,browser # or env CAMOUFOX_CAPSProfile | Tools | What's inside |
| 18 | search + reading/extraction |
| 44 | + live tab, forms, network, files |
| 20 | + |
| 46 | the full registry |
ping/stats are always available. CAMOUFOX_TOOLS_ONLY / CAMOUFOX_TOOL_HIDE
still apply on top. Invalid group → warning, valid groups still load.
MCP resources & prompts
Resources:
camoufox://stats,camoufox://cache,camoufox://session,camoufox://info,camoufox://health,camoufox://searchPrompts:
research_plan,extract_schema,monitor_page
Transports
stdio (default), streamable-http (stateless; the remote option), sse
(legacy, kept for the 2026-07-28 spec's 12-month window):
camoufox-research --transport http --port 8833
CAMOUFOX_PORT=8833 camoufox-research --transport httpTroubleshooting
Unknown tool usually means the client is still talking to an old server — a
restart after a code change fixes it. A one-command, read-only diagnosis:
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py # human-readable
~/.venvs/camoufox-research/bin/python scripts/checks/mcp_probe.py --json # machine-readableIt shows python / repo → caps → protocol handshake → how many tools
tools/list returns → package version → search-watchdog pulse. A small
tools=N or a failed handshake means stale code in the venv → reinstall from
the clone (bash scripts/install.sh --reinstall) and reconnect (API
disconnect/connect, not kill).
Honest numbers
There is no single head-to-head here. The three numbers below come from one third-party benchmark run; ours come from separate runs on a smaller sample. They sit next to each other for context — not as the same measurement.
Reference (not our run): fastCRW diagnose_3way.py
scoring (phrases > 20 chars, recall >= 0.3 = found) on the public dataset
firecrawl/scrape-content-dataset-v1 — measured 2026-05-08 over the full 819 URLs:
Tool | Truth-recall |
fastCRW | 63.7% (522/819) |
Crawl4AI | 60.0% (491/819) |
Firecrawl | 56.0% (459/819) |
Our own number is deliberately kept out of that table. Our script runs a
30-URL prefix of the same dataset, so the comparison is indirect — different
sample size and a different date. At n=30 the 95% confidence interval is roughly
±18 p.p., wider than the gap between all the tools above. Two runs on 2026-09-29
gave 36.7% (11/30) and 40.0% (12/30), and a re-run within 24h measures the
page cache rather than the web. Read it as a smoke number for the extractor, not
as a leaderboard position. Reproduce with
python scripts/dev/bench_truth_recall.py --sample N.
Documentation
Doc | What's there |
guide for the agent: profiles, call order, limits, pitfalls | |
recipes for fast / full / deep research, measured chars and seconds | |
what was verified live and by what, plus an honest "what was NOT verified" | |
install matrix: Linux / macOS / Windows / Docker | |
native Windows (PowerShell 7) install notes | |
layers, modules, state, extension points, debts | |
verified landmines ("what not to step on") — index in EXPERIENCE.md | |
how to report a vulnerability (private advisory), scope, supported versions |
Development
See CONTRIBUTING.md: layout, how to add a tool, the live-smoke
ritual and the unit tests (pip install -e ".[dev]" → pytest -q). CI runs
pytest, py_compile + import + an MCP stdio smoke on Python 3.10–3.13, plus
the lint/filesize/version gates; full browser flows are checked locally.
License
MIT — see LICENSE.
Dependencies: camoufox (see its repo), mcp (MIT), trafilatura (GPL-3.0; обязательна — read_document/table_extract и article_only извлекают ею, без неё чтение деградирует до текста всего body).
Available Tools
18 toolsbatch_fetchПакетное чтениеARead-onlyIdempotent
Открывает НЕСКОЛЬКО URL в одном браузере — для глубокого ресёрча на 30-50 источников одним вызовом вместо серии холодных стартов. Кэш: уже посещённые URL возвращаются мгновенно, без браузера. Rate limit между переходами защищает от капчи. Батч ≥8 URL — параллельно (пул потоков, свой браузер на поток); число воркеров автоопределяется по ресурсам машины (слабый ПК — 1-2, мощный — 3-4), max_parallel — явное ограничение. Возвращает тексты с разделителями '--- URL: ...'. article_only=True — извлечь текст статьи (Trafilatura), без меню и баннеров. Пример: batch_fetch(urls=["https://docs.python.org/3/", "https://opencode.ai/docs/"], max_chars=20000, article_only=True) ХОЧУ ПОЛНОТУ → max_chars=20000..100000 + article_only=True + max_parallel=4; экономия контекста → max_chars=4000 при 30+ URL. Замер 21.09 (8 URL разных доменов, холодный кэш): 4000 → 32 598 симв. за 36.4с (все 8 обрезаны ровно по 4000), 12000 → 96 598 симв. за 30.3с, 20000 → 154 449 симв. за 45.9с — время НЕ растёт от объёма: страница и так читается до потолка кэша (100k), max_chars режет только ОТВЕТ. КОГДА: читать 10-50 URL одним вызовом (глубокий ресёрч после research(queries=[...]) или crawl/map_site). НЕ КОГДА: 1-2 страницы → fetch_page; URL ещё не собраны → research, sitemap, map_site сначала.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | ||
| max_chars | No | ||
| article_only | No | ||
| max_parallel | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive, openWorld), and the description adds substantial non-structured behavior: cache hits return instantly without a browser, inter-navigation rate limiting guards against CAPTCHA, batches of ≥8 URLs run in a thread pool with per-thread browsers, and max_chars truncates only the response, not the crawl. This is richer than typical annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: purpose, then cache/rate-limit/parallel behavior, then an example, then when/not-when. The dated benchmark timings (36.4s/30.3s/45.9s) are somewhat verbose, but they justify the tuning advice and are not pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet it still notes the '--- URL: ...' text separators. Given four parameters, no schema descriptions, and a parallel/cached execution model, the description supplies everything an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry all four parameters — and it does: urls (scale guidance), max_chars (concrete ranges 4000-100000 with measured tradeoffs), article_only (Trafilatura extraction excluding menus/banners), and max_parallel (explicit cap on auto-detected workers). A worked example demonstrates combined usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Открывает НЕСКОЛЬКО URL в одном браузере') plus the scale it targets (30-50 sources, one call instead of cold starts). An agent can immediately distinguish it from fetch_page and the other fetch-adjacent siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains explicit 'КОГДА' and 'НЕ КОГДА' sections: use for 10-50 URLs after research/crawl/map_site; use fetch_page for 1-2 pages; gather URLs first via research, sitemap, map_site. Alternatives and exclusion conditions are named outright.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_linksПроверка ссылокARead-onlyIdempotent
КОГДА: перед публикацией или после редизайна — найти битые ссылки на странице. ЧТО: HTTP-статусы собранных ссылок, отчёт вида «[404] URL». timeout — срок НА КАЖДУЮ ссылку (15с). Обычный путь (urllib) идёт ОГРАНИЧЕННЫМ ПУЛОМ потоков: CAMOUFOX_LINK_WORKERS (дефолт 8, потолок 16) — 50 ссылок ≈ минута, а не ~12. НО при поднятом живом браузере (serve) запросы идут через его sync-Playwright — он потокопривязан, поэтому пул принудительно 1 (последовательно). Бюджет вызова мост поднимает сам (900с) — таймаут MCP-клиента ставь ≥900с. НЕ: жди параллельности, когда поднят живой браузер — там строго по одной (грабли: 4 потока → 15/15 error); внешние ссылки → internal_only=False.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| timeout | No | ||
| max_links | No | ||
| internal_only | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (read-only, open-world, idempotent), and the description adds substantial behavior beyond them: per-link timeout semantics, bounded worker pool via CAMOUFOX_LINK_WORKERS (default 8, cap 16), forced sequential execution under live-browser mode, and the 900s call budget with client-timeout advice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded WHEN/WHAT/NOT structure makes it scannable, and the concurrency caveat plus timeout budget genuinely earn their space. It is dense with operational detail, but none of it is padding given the described pitfalls.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return format need not be restated, and the description covers the runtime constraints, failure mode, and budget that an agent needs to invoke this correctly. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it explains the two non-obvious params: timeout is per-link (15s), not total, and internal_only controls whether external links are checked. max_links is only implied via the 50-link/min figure, so it is slightly short of complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: find broken links on a page and report HTTP statuses as «[404] URL». That clearly distinguishes it from extract_links (which only gathers URLs) without needing to name the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit WHEN (before publishing or after a redesign) and explicit NOT (don't expect parallelism when a live browser is up; external links require internal_only=False). The conditions that select behavior are spelled out, including the observed failure mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crawlОбход сайтаARead-onlyIdempotent
КОГДА: собрать содержимое САЙТА целиком (BFS по внутренним ссылкам) для синтеза, а не одну страницу. ЧТО: тексты страниц (depth ≤ max_depth, всего ≤ max_pages) с разделителями '--- URL:'; кэш делает повторный обход дешёвым. Идёт долго (десятки переходов) — таймаут MCP-клиента ставь ≥900с. НЕ: нужны только URL → map_site / sitemap; сайт огромный → sitemap + pattern и уже потом crawl нужного раздела; одна страница → fetch_page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| pattern | No | ||
| max_chars | No | ||
| max_depth | No | ||
| max_pages | No | ||
| article_only | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld, and the description adds operational traits they cannot express: caching makes repeat crawls cheap, a crawl takes tens of transitions, and the MCP client timeout should be ≥900s. That is exactly the kind of cost/latency disclosure an agent needs before committing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The КОГДА/ЧТО/НЕ structure front-loads purpose, then output shape and cost, then exclusions. Every line is load-bearing with no filler, despite covering a fairly complex tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output-schema existence relieves it of return-value detail, and it adds timeout/caching context and alternatives. The remaining gap is the unexplained max_chars, article_only and pattern semantics, which leaves a 6-param tool partially under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it only covers url (implied), max_depth and max_pages ('depth ≤ max_depth, всего ≤ max_pages') plus a passing nod to pattern. max_chars and article_only are never explained, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The 'ЧТО' line names the resource (страницы сайта) and the mechanism (BFS по внутренним ссылкам), explicitly scoping it to a whole site rather than one page. An agent can distinguish it from fetch_page and map_site from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'КОГДА' and 'НЕ' sections give explicit when-to-use and when-not-to-use with named alternatives: URL-only needs → map_site/sitemap, huge site → sitemap + pattern then crawl a section, single page → fetch_page. This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exportЭкспорт результатаADestructive
КОГДА: результат нужен файлом на диске — CSV для таблиц, JSON для автоматизации, MD для отчёта. ЧТО: файл (json/csv/md) по своему path или авто в ~/.cache/camoufox-research/exports/. НЕ: результат идёт в разговор и синтез → отдай текст как есть.
| Name | Required | Description | Default |
|---|---|---|---|
| data | Yes | ||
| path | No | ||
| format | No | json |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so the agent knows this is a write operation. The description usefully adds the default output location (~/.cache/camoufox-research/exports/) when no path is given. However, it does not explain the destructive semantics the annotation implies — e.g., whether it overwrites an existing file at a user-supplied path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three labeled lines (WHEN/WHAT/NOT) with zero filler and the usage condition front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description covers usage, formats, and default destination adequately for a simple 3-param tool. The remaining gap is destructive/overwrite behavior, which the annotation hints at but the description never addresses.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It explains the format choices (json/csv/md) and the path behavior (own path or auto cache location), which compensates for two of three params. It does not clarify the required `data` parameter, and omits the 'markdown' enum alias present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb+resource framing (export a result to a file) and names the concrete output formats (CSV for tables, JSON for automation, MD for report). It distinguishes this tool from the conversation/synthesis path and from the many fetch/research siblings by scoping it to disk output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The explicit WHEN/WHAT/NOT structure states when to use it (result needed as a file on disk), what it produces, and when NOT to use it (result goes into conversation and synthesis → pass text as-is). This is exactly the routing guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extractИзвлечение полейARead-onlyIdempotent
КОГДА: нужны конкретные поля страницы — при стабильной вёрстке селекторами (CSS/XPath) или из текста через llm=True, когда вёрстка хрупкая и селекторы отваливаются. ЧТО: schema — JSON-объект (можно строкой для совместимости): {"поле": "css:.price"} или {"поле": {"selector": ".price", "attr": "text|href|src"}}; llm=True — {"поле": "подсказка"} и нужен LLM (DeepSeek/Ollama), иначе честный ответ «недоступен». НЕ: нужен сплошной текст → fetch_page / batch_fetch; таблицы → table_extract; не знаешь селектор → сначала snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| llm | No | ||
| url | Yes | ||
| schema | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/openWorld/idempotent/non-destructive, but the description adds real behavioral context beyond them: llm=True requires an LLM backend (DeepSeek/Ollama) and otherwise returns an honest 'unavailable' answer, and schema is accepted as a JSON string for compatibility. Minor gaps remain (rate limits, whether extraction is bounded per page).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded labeled sections (КОГДА / ЧТО / НЕ) pack the routing decision, format spec, and exclusions into a few dense lines with no filler. Every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described; the description instead covers selection criteria, parameter encoding, backend dependency, and sibling alternatives, which is complete for a 3-param extraction tool. The reference to 'snapshot' (absent from the sibling list) is the only nit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does: it documents the schema parameter's css:/XPath prefix syntax, the {selector, attr: text|href|src} object form, and the llm hint form, plus llm=True's backend dependency. Only url is left unexplained, which is self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The 'ЧТО' block states a precise verb+resource (extract named page fields) and immediately distinguishes the two modes (selector-based vs llm=True). It is trivially separable from siblings like fetch_page or table_extract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'КОГДА' and 'НЕ' blocks give explicit when-to-use, when-not, and name the concrete alternatives (fetch_page / batch_fetch for raw text, table_extract for tables, snapshot when the selector is unknown). Routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_linksСсылки страницыBRead-onlyIdempotent
Собирает ссылки страницы (фильтр по подстроке pattern).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| pattern | No | ||
| max_links | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds only the substring-filtering behavior; it says nothing about fetching live pages, rate limits, or failure modes on unreachable URLs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the resource named first and the filter qualifier second; nothing is wasted, though brevity here borders on under-specification rather than efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the read-only/idempotent profile. For a simple three-parameter tool this is minimally viable, but the undocumented url and max_links parameters and absent sibling routing leave real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it explains only one of three parameters (pattern as a substring filter). The required url and the max_links cap (default 20) are left entirely undocumented in both the schema and the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it collects the links of a page, with an optional substring filter. That is clear enough for an agent to know what it returns, though it never distinguishes itself from near-siblings like check_links, extract, or crawl.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of the alternatives in the sibling set (e.g. check_links for validation, crawl for whole-site traversal). The parenthetical filter hint implies one capability but does not tell the agent when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_pageЧтение страницыARead-onlyIdempotent
Текст страницы без HTML-мусора (статьи, доки, README). Кэш на сутки. article_only=True — текст статьи (Trafilatura), fallback — весь body. delta=True — delta-чтение: если контент не изменился с прошлого раза, вернёт маркер '[delta: ...]' вместо текста (не тратим токены на повтор). max_chars режет только ОТВЕТ, а не чтение: страница кладётся в кэш целиком (потолок 100k), поэтому повтор с большим max_chars мгновенный и бесплатный по времени (замер 21.09: asyncio-task.html = 46 094 симв. в кэше; 12000 = 26% текста, 40000 = 87%, все ответы 0.0с). КОГДА: прочитать 1 страницу (JS/SPA — тоже) чистым текстом. НЕ КОГДА: страниц 10+ → batch_fetch; нужны поля по схеме → extract; повторное чтение → delta=True; нужен клик/ввод → session_start.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| delta | No | ||
| max_chars | No | ||
| article_only | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive, open-world behavior, but the description adds material context beyond them: a ~24h cache, how delta avoids re-emitting unchanged content, and that max_chars truncates only the response while the full page (up to 100k) is cached. That directly affects token and retry decisions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then semantics, then the when/when-not routing. Dense but nearly every clause earns its place; the parenthetical benchmark measurement (21.09 timings, 26%/87% figures) is more illustrative detail than an agent strictly needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return structure needn't be documented, yet the description still explains the one non-obvious return case (delta marker). With caching, truncation, extraction fallback and sibling routing all covered, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden and does: it defines delta (marker instead of text on unchanged content), article_only (Trafilatura article extraction with full-body fallback), and max_chars (response-only truncation with cache reuse). Only url is left to self-evidence, which is reasonable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (read one page as clean text, JS/SPA included) and immediately differentiates itself from siblings: batch_fetch for 10+ pages, extract for schema fields, session_start for interaction. No sibling overlap is left ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit КОГДА / НЕ КОГДА block with named alternatives for each exclusion case (batch_fetch, extract, delta=True, session_start). This is the strongest form of routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_siteКарта сайтаARead-onlyIdempotent
КОГДА: понять структуру сайта — какие разделы и страницы есть, без чтения содержимого. ЧТО: ссылки того же домена со стартовой страницы (до max_links), pattern — фильтр по URL. НЕ: нужны тексты → crawl; есть sitemap.xml → sitemap (полнее и дешевле); одна страница → fetch_page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| pattern | No | ||
| max_links | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds genuine context beyond them: it does not read page content, it stays on the same domain, and it caps results at max_links, plus a relative cost hint versus sitemap ('дешевле'). No return format is described, but that is a minor gap given the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three labelled lines, front-loaded with the goal, then the mechanism, then the exclusions. No filler sentences; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. For a read-only mapping tool with rich annotations, the WHEN/WHAT/NOT triad plus the sibling routing makes the definition complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter burden. It does explain pattern ('фильтр по URL') and max_links ('до max_links'), but url is only implied and there is no syntax guidance for pattern (regex vs substring) or what the cap default means. Partial compensation for a 0% coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The ЧТО line gives a concrete verb+resource: 'ссылки того же домена со стартовой страницы (до max_links)', and the КОГДА line frames the goal ('понять структуру сайта... без чтения содержимого'). It is immediately distinguishable from crawl, sitemap and fetch_page.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit WHEN/WHAT/NOT structure names the alternatives by tool name and the condition that selects each: 'нужны тексты → crawl; есть sitemap.xml → sitemap (полнее и дешевле); одна страница → fetch_page.' Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_diffДифф страницыARead-onlyIdempotent
КОГДА: страницу уже читали, и нужно узнать, что ИЗМЕНИЛОСЬ — цены, доки, новости (мониторинг, второй и далее заходы). ЧТО: дифф свежего чтения с прошлым из кэша, «что поменялось». НЕ: страница читается впервые — сравнивать не с чем, сначала fetch_page; нужна полная текстовая версия → fetch_page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld, non-destructive. The description adds the critical behavioral constraint not in annotations: it diffs against a cached prior read and is meaningless without one. It doesn't discuss caching expiry or what happens when no cache exists, but the core dependency is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The WHEN/WHAT/NOT structure front-loads the deciding condition and packs scope, mechanism, and exclusions into three tight lines with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a monitoring/diff tool the description covers the deciding context, the cache prerequisite, and the fallback, and an output schema exists so return values need no explanation. The only real gap is the undocumented max_chars parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions neither parameter. The required url is obvious from context, but max_chars (default 6000) is left entirely undefined in both schema and description, so an agent gets no guidance on the output-size knob.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('дифф свежего чтения с прошлым из кэша') and explicitly frames what it produces ('что поменялось'). It cleanly distinguishes itself from fetch_page, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is organized as explicit WHEN / WHAT / NOT sections. It names the prerequisite (page already read), the exclusion (first read → fetch_page), and the alternative for full text (fetch_page). This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_searchПоиск статейARead-onlyIdempotent
Поиск научных статей: arXiv + Semantic Scholar (бесплатные API, без ключей). Возвращает статьи с годом/авторами/цитатами — первоисточники (tier 0), которых общий поиск почти не видит (паттерн индустрии: vertical index / arxiv-канал рядом с вебом). Кэш на сутки. sources — какие индексы брать (arxiv/semantic/ crossref/wiki); пусто — arxiv+semantic. Пример: paper_search("deep research agents", sources=["arxiv"])
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| sources | No | ||
| max_results | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld/idempotent/no-destructive, so safety is covered. The description adds real behavioral context beyond them: no API keys needed, results include year/authors/citation counts, and a one-day cache that affects freshness of repeated queries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then defaults, then an example — a sensible order. Some parenthetical industry commentary ('паттерн индустрии: vertical index…') is filler that could be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the description still summarizes the shape (year/authors/citations). Combined with the documented sources semantics and cache behavior, an agent has enough to call this correctly; only max_results remains unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the load. It documents sources effectively (allowed values arxiv/semantic/crossref/wiki, default behavior, and an example), but max_results and the query field are left entirely to the schema, so it doesn't fully compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Поиск научных статей') and enumerates the two backends (arXiv + Semantic Scholar). It explicitly contrasts its scope with general search ('которых общий поиск почти не видит'), letting an agent distinguish it from the web_search sibling without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: free APIs without keys, 24h cache, and how to select indexes via sources with defaults (empty = arxiv+semantic) plus a concrete example call. It does not name web_search explicitly as the alternative to avoid, so routing is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pingПроверка связиARead-onlyIdempotent
Проверка связи: возвращает pong.
Это ПРИКЛАДНОЙ health-тул, а не протокольный ping (его в спеке
2026-07-28 убрали). Имя оставлено публичным контрактом — менять нельзя.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered by structured data. The description adds non-obvious context the annotations cannot: that this is an application-level health probe rather than the removed protocol `ping`, and that the name is a frozen public contract. That is genuine value beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core behavior ('returns pong') is front-loaded in the first clause. The remaining notes about the spec version and the fixed name are slightly meta and consume space an agent doesn't strictly need, but they serve a real purpose in preventing misuse of the name, so they largely earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a trivial zero-parameter health check with a full output schema and rich annotations, the description supplies everything relevant: behavior, semantics of the name, and safety is covered structurally. Only the 'when to call me' trigger is left implicit, which is the one minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify, and it introduces no misleading parameter hints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete behavior — a health check that returns 'pong' — and explicitly distinguishes itself from the protocol-level `ping`, which prevents an agent from mistaking it for a network diagnostic. What the tool does is unambiguous. It doesn't need to differentiate from siblings since none of the listed tools (crawl, fetch_page, web_search) are comparable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It identifies itself as an application health tool, which implies the use case (connectivity/liveness check), but it never states when an agent should call it — e.g. before a batch of operations, or to verify server availability. No alternatives or exclusions are named because none exist among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_documentЧтение документаARead-onlyIdempotent
КОГДА: нужен документ (PDF/DOCX/XLSX) — отчёт, прайс, спецификация; такие файлы часто попадаются ссылками с сайтов. ЧТО: текст документа; source — URL или локальный путь (pypdf / python-docx / openpyxl, до max_chars). Локальный путь — только из каталога загрузок сервера (туда кладёт session_download); свой путь — вентиль CAMOUFOX_ALLOW_ANY_PATH=1. НЕ: HTML-страница → fetch_page (дешевле); старые .doc/.xls не читаются — сначала libreoffice --convert-to docx/xlsx.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, it discloses real operational constraints: local paths are restricted to the server download directory populated by session_download, and arbitrary paths require the CAMOUFOX_ALLOW_ANY_PATH=1 gate. It also notes that old .doc/.xls formats fail and output is truncated at max_chars. This is substantial context, though permission/auth details for remote URLs are not covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The КОГДА/ЧТО/НЕ front-loading is efficient and scannable, and every block earns its place. The parenthetical library names and format-conversion aside add density that a reader must parse, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value description is unnecessary, and the description covers the remaining gaps: input constraints, path restrictions, format limitations, and the cheaper alternative for HTML. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does: 'source' is defined as a URL or local path with the directory restriction, and max_chars is framed as the truncation ceiling (default 6000 implied by schema). It adds meaning beyond the bare schema, though it does not describe accepted URL schemes or path syntax in detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The 'ЧТО' (WHAT) block states a specific verb and resource: extracts the text of a document (PDF/DOCX/XLSX), naming the underlying libraries. It explicitly distinguishes itself from the sibling fetch_page for HTML pages, so an agent can separate it from the other extraction tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The КОГДА/НЕ structure gives explicit when-to-use ('нужен документ... отчёт, прайс, спецификация') and when-not ('HTML-страница → fetch_page (дешевле)'), naming the alternative and the reason it is preferred. It also routes legacy .doc/.xls users to a conversion step first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
researchГлубокий ресёрчARead-onlyIdempotent
Когда: нужно СЕЙЧАС 10+ источников на тему одним вызовом — глубокий разбор без фонового ожидания. Один факт, новость или точный URL — это web_search. Что: сервер ищет по КАЖДОЙ формулировке (queries), дедуплицирует URL и отдаёт список со сниппетами; fetch_top>0 сразу читает топ-N текстов. Результат кэшируется на сутки.
По умолчанию хватит queries (2-4 формулировки) и fetch_top=3..5. Остальное — тонкая настройка, всё рабочее, но трогать не обязательно: max_chars режет только ответ; max_parallel — воркеры чтения; target_domains=N — цель по РАЗНЫМ доменам (20 = двадцать сайтов, доборка волнами); domains_limit=K — не больше K с одного домена; expand=True — переформулировки («X comparison», «X documentation»); terms_wave=True — вторая волна из редких термов первой (паттерн Open Deep Research); quality_first=True — доки/GitHub/arXiv первыми; fetch_all=True — тексты ВСЕХ отобранных, а не топ-N (30 источников × 12k ≈ 90k токенов); as_json=True — машинный JSON: meta (счётчики, follow-up запросы), sources (title/url/domain/tier/tier_label/snippet), texts, notes (тот же объект в structuredContent, content — прежняя строка); academic=True — вертикальный arXiv + Semantic Scholar канал (tier 0, без ключей); llm_planner=True — LLM-планировщик follow-up (DeepSeek/Ollama, нужен DEEPSEEK_API_KEY или OLLAMA_HOST, иначе пропуск); mode="быстро"|"полно"|"глубоко" (fast/full/deep) — пресет одним словом, трогает только ручки на дефолте (явный аргумент сильнее пресета). Пример глубокого разбора: research(queries=["deep research agents"], target_domains=20, domains_limit=2, expand=True, terms_wave=True, quality_first=True, academic=True, llm_planner=True, fetch_all=True, as_json=True, max_results_per_query=6).
Замер 21.09 (5 источников): fetch_top=0 — 0 текстов и 2 822 симв. за 5.5с; fetch_top=3 + max_chars=12000 — 3 текста и 32 432 симв. за 23.1с. ⏱ Долгий: один вызов идёт до ~15 мин (внутренний таймаут 900с, столько же ставь таймауту MCP-клиента) — ждать ответа, не поллить.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| expand | No | ||
| as_json | No | ||
| queries | Yes | ||
| academic | No | ||
| fetch_all | No | ||
| fetch_top | No | ||
| max_chars | No | ||
| terms_wave | No | ||
| llm_planner | No | ||
| article_only | No | ||
| max_parallel | No | ||
| domains_limit | No | ||
| quality_first | No | ||
| target_domains | No | ||
| max_results_per_query | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| meta | No | |
| notes | No | |
| texts | No | |
| result | No | |
| sources | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld, and non-destructive behavior. The description adds major behavioral context beyond annotations: daily caching, long runtime up to ~15 minutes with a 900s internal timeout, a warning not to poll, API key requirements for llm_planner, and token-cost implications for fetch_all. No contradiction with the read-only/idempotent annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the length is largely justified by the 16-parameter surface and the need to document behavior in the absence of schema descriptions. It is front-loaded with when/what, but the benchmark measurement and some tightly packed parameter shorthand make it slightly heavier than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity, zero schema description coverage, and the fact that an output schema exists, the description is complete. It covers selection guidance, defaults, advanced parameters, return shape for as_json, timeout expectations, and key dependencies, so an agent has what it needs to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter semantics, and it does so almost completely. It explains queries, fetch_top, max_chars, max_parallel, target_domains, domains_limit, expand, terms_wave, quality_first, fetch_all, as_json, academic, llm_planner, and mode, with defaults and effects. Only article_only is left unmentioned, which is a minor omission in an otherwise thorough namespace.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: deep research, retrieving 10+ sources on a topic in one call, with no background wait. It explicitly distinguishes from the sibling web_search for single facts, news, or exact URLs, so an agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use ('нужно СЕЙЧАС 10+ источников...') and when-not-to-use ('Один факт, новость или точный URL — это web_search') guidance. It also provides default parameter advice and an advanced example, leaving no routing ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rssRssARead-onlyIdempotent
КОГДА: следить за обновлениями — новости, блог, changelog, релизы (лента отдаёт даты одним вызовом). ЧТО: посты RSS/Atom: title, link, дата (до limit). НЕ: у URL обычная HTML-страница → fetch_page; следить за лентой целиком → crawl/map_site.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful context beyond them — that the feed returns dates in one call and lists returned fields (title, link, date) — but does not discuss rate limits or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight labeled blocks (WHEN/WHAT/NOT) with zero filler, front-loading the trigger before the exclusions. Every line earns its place and is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required. Combined with rich when/when-not guidance and sibling routing, the description gives an agent everything needed to invoke this two-parameter read tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters. The description partially compensates by clarifying 'до limit' (up to limit results), giving the limit parameter meaning, but 'url' is left entirely to the schema and no format/behavior detail is added. With only partial compensation for the coverage gap, this sits at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The ЧТО block states a specific verb and resource: fetch RSS/Atom posts with title, link, and date fields. It is clearly distinguishable from siblings like fetch_page (HTML pages) and crawl/map_site (whole-feed monitoring), so an agent can pick it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The КОГДА block gives explicit positive triggers (news, blog, changelog, releases) and the НЕ block names two concrete alternatives with the conditions that select them: fetch_page for ordinary HTML, crawl/map_site for full-feed monitoring. This is textbook when/when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sitemapSitemapARead-onlyIdempotent
КОГДА: нужен полный список страниц сайта — фид для crawl или проверка «что вообще есть на домене». ЧТО: URL из sitemap.xml (+ .xml.gz и вложенные sitemapindex), до max_links. НЕ: у сайта нет sitemap → map_site (ссылки со страницы); тексты → crawl.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_links | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context by disclosing supported sitemap sources and the max_links cap, though it does not mention auth requirements or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The KOGDA/CHTO/NE structure is compact and front-loaded. Each line earns its place, with no redundant explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The rich annotations and output schema mean return format and safety notes are covered. The description handles usage and alternatives well, but the undocumented url parameter keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that max_links caps the number of URLs, but the required url parameter is never explicitly defined or given a format, leaving a significant semantic gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The WHAT clause states a specific resource and scope: URLs from sitemap.xml, .xml.gz, and nested sitemapindex, up to max_links. It explicitly distinguishes the tool from siblings by naming map_site and crawl for different needs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The KOGDA and NE sections give explicit when-to-use conditions and alternatives: use for a full page list, crawl feed, or domain check; if no sitemap use map_site; for texts use crawl. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statsСтатистика вызововARead-onlyIdempotent
КОГДА: понять, какие тулы реально работают и что падает — аудит
вызовов и отладка клиента.
ЧТО: по каждому тулу счётчик вызовов, среднее время, ошибки, плюс
последние вызовы (секреты замаскированы). В профилях caps виден
ВСЕГДА (ALWAYS_ON); явные CAMOUFOX_TOOLS_ONLY / CAMOUFOX_TOOL_HIDE
сильнее и могут его убрать.
НЕ: это не метрика «какими тулами пользуются» (её даёт
~/.cache/camoufox-research/tool_usage.json) и
не здоровье сервера (его даёт ping).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the safety profile (readOnly/idempotent/non-destructive). The description adds genuinely new behavioral context: secrets are masked in returned call data, and availability rules (ALWAYS_ON in caps profiles, overridable by CAMOUFOX_TOOLS_ONLY / CAMOUFOX_TOOL_HIDE). This is exactly the kind of trait annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the WHEN/WHAT/NOT labels and no filler. The profile-visibility clause is somewhat dense but conveys a real constraint rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be spelled out, and the description covers purpose, alternatives, masking, and availability. The only gap is the undocumented `limit` parameter, which leaves the agent guessing at how much data is returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single `limit` parameter (default 20) is completely unexplained in the description. Since coverage is below 50%, the description is expected to compensate, but it says nothing about what limit controls (recent calls shown?) or its bounds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (per-tool call statistics) with concrete contents: call counter, average time, errors, and recent calls. The НЕ block explicitly distinguishes it from two named alternatives (tool_usage.json and ping), so an agent can identify it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Uses an explicit WHEN/WHAT/NOT structure. 'КОГДА' names the use case (auditing which tools work and what fails, client debugging), and 'НЕ' names the alternatives for the adjacent questions (usage metrics via tool_usage.json, server health via ping). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
table_extractИзвлечение таблицARead-onlyIdempotent
КОГДА: на странице есть — прайсы, характеристики, сравнения. ЧТО: CSV-текст таблиц (до max_tables) по CSS-селектору. НЕ: нужных данных в таблице нет → extract; таблиц нет вовсе → fetch_page; JS-грид на div'ах → extract / snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| selector | No | table | |
| max_tables | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context by specifying that output is CSV text and that extraction is capped by max_tables, though it does not discuss pagination or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured into WHEN, WHAT, and NOT sections, is front-loaded with the triggering condition, and contains no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple extraction task, rich annotations, and an output schema, the description covers routing and basic behavior well. Parameter titles and defaults in the schema fill most remaining gaps, though explicit url semantics are still absent from the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that selector is a CSS selector and that max_tables limits the number of extracted tables, but it does not explicitly explain the required url parameter or the default values already present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: extract table content as CSV text using a CSS selector. It also distinguishes the tool from siblings by naming cases where extract, fetch_page, or snapshot are preferable instead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The WHEN/WHAT/NOT structure explicitly states when to use this tool, when not to use it, and which alternative tool to use in each excluded case. This leaves little ambiguity for an agent choosing among extract, fetch_page, and table_extract.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchВеб-поискARead-onlyIdempotent
Когда: БЫСТРЫЙ ОТВЕТ — один запрос, факт, новость, точный URL. Что: номера, заголовки и URL из DuckDuckGo (анти-детект браузер), кэш на сутки; include_snippets — сниппет под URL, pages>1 — пагинация. engines — явный список движков для диагностики (["brave"], ["ddg_html","bing"]), пусто — цепочка из env CAMOUFOX_SEARCH_ENGINES; transport — "auto"/"http"/"browser" (HTTP-фолбэк быстрее, браузер нужен JS/consent-движкам); proxy — разовый прокси на этот вызов ('host:port', 'user:pass@host:port', 'socks5://host:port'); без них поведение прежнее. Не когда: нужно 10+ разных сайтов → research (сейчас); научные статьи → paper_search.
| Name | Required | Description | Default |
|---|---|---|---|
| pages | No | ||
| proxy | No | ||
| query | Yes | ||
| engines | No | ||
| transport | No | auto | |
| max_results | No | ||
| include_snippets | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent/non-destructive/open-world, and the description adds real context beyond them: anti-detect browser backend, 24h cache, engines defaulting to the CAMOUFOX_SEARCH_ENGINES chain, HTTP fallback being faster, and browser transport required for JS/consent engines.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with Когда/Что/Не когда and no filler sentences, but the 'Что' block is a dense run-on with nested parentheticals and example values that takes effort to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations cover safety. The description supplies usage routing, backend/cache behavior, and most parameter meanings; only max_results semantics remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the burden — and it explains include_snippets, pages, engines (with example values and the empty-chain default), transport modes, and per-call proxy formats. max_results is never mentioned, leaving one of seven parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (быстрый веб-поиск через DuckDuckGo) and what it returns (номера, заголовки, URL). It explicitly distinguishes itself from siblings by naming 'research' and 'paper_search' as different tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Uses an explicit Когда / Не когда structure: one query, fact, news, exact URL is the target; 10+ different sites → research, scientific articles → paper_search. The routing decision is fully specified with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
50 tool updates
v0.21.1- Changed
batch_fetch1 field changed- changed
Input schema / properties / max_chars / defaultPrevious value: -4000New value: +12000
- Removed
browser_click - Removed
browser_navigate - Removed
browser_type - Removed
citation_pack - Removed
citation_report - Changed
export1 field changed- added
Input schema / properties / format / enumAdded value: +[ + "json", + "csv", + "md", + "markdown" +]
- Changed
extract2 fields changed- added
Input schema / properties / schema / anyOfAdded value: +[ + { + "type": "string" + }, + { + "additionalProperties": true, + "type": "object" + } +] - removed
Input schema / properties / schema / typeRemoved value: -"string"
- Changed
fetch_page1 field changed- changed
Input schema / properties / max_chars / defaultPrevious value: -6000New value: +12000
- Changed
paper_search3 fields changed- added
Input schema / properties / sources / anyOfAdded value: +[ + { + "items": { + "enum": [ + "arxiv", + "semantic", + "crossref", + "wiki" + ], + "type": "string" + }, + "type": "array" + }, + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / sources / defaultPrevious value: -"arxiv,semantic"New value: +null - removed
Input schema / properties / sources / typeRemoved value: -"string"
- Removed
profile_load - Removed
profile_save - Changed
research10 fields changed- changed
Input schema / properties / fetch_top / defaultPrevious value: -0New value: +3 - changed
Input schema / properties / max_chars / defaultPrevious value: -4000New value: +12000 - added
Input schema / properties / modeAdded value: +{ + "default": "", + "enum": [ + "", + "быстро", + "fast", + "quick", + "lite", + "полно", + "full", + "normal", + "глубоко", + "deep", + "max" + ], + "title": "Mode", + "type": "string" +} - added
Output schema / properties / metaAdded value: +{ + "type": "object" +} - added
Output schema / properties / notesAdded value: +{ + "type": "array" +} - removed
Output schema / properties / result / titleRemoved value: -"Result" - added
Output schema / properties / sourcesAdded value: +{ + "type": "array" +} - added
Output schema / properties / textsAdded value: +{ + "type": "array" +} - removed
Output schema / requiredRemoved value: -[ - "result" -] - removed
Output schema / titleRemoved value: -"researchOutput"
- Removed
research_critic - Removed
research_digest - Removed
research_index - Removed
research_report - Removed
research_resume - Removed
research_start - Removed
research_status - Removed
screenshot - Removed
service_route - Removed
session_back - Removed
session_block - Removed
session_click - Removed
session_console - Removed
session_download - Removed
session_end - Removed
session_eval - Removed
session_form_fill - Removed
session_key_press - Removed
session_links - Removed
session_navigate - Removed
session_network - Removed
session_resize - Removed
session_scroll - Removed
session_select_option - Removed
session_start - Removed
session_status - Removed
session_tabs - Removed
session_text - Removed
session_type - Removed
session_unblock - Removed
session_upload - Removed
session_wait_for - Removed
set_proxy - Removed
snapshot - Removed
tool_hint - Removed
tool_usage - Changed
web_search3 fields changed- added
Input schema / properties / enginesAdded value: +{ + "anyOf": [ + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Engines" +} - added
Input schema / properties / proxyAdded value: +{ + "default": "", + "title": "Proxy", + "type": "string" +} - added
Input schema / properties / transportAdded value: +{ + "default": "auto", + "enum": [ + "auto", + "http", + "browser" + ], + "title": "Transport", + "type": "string" +}
15 tool updates
v0.18.1- Added
citation_pack - Added
citation_report - Changed
extract1 field changed- added
Input schema / properties / llmAdded value: +{ + "default": false, + "title": "Llm", + "type": "boolean" +}
- Added
paper_search - Changed
research9 fields changed- added
Input schema / properties / academicAdded value: +{ + "default": false, + "title": "Academic", + "type": "boolean" +} - added
Input schema / properties / as_jsonAdded value: +{ + "default": false, + "title": "As Json", + "type": "boolean" +} - added
Input schema / properties / domains_limitAdded value: +{ + "default": 0, + "title": "Domains Limit", + "type": "integer" +} - added
Input schema / properties / expandAdded value: +{ + "default": false, + "title": "Expand", + "type": "boolean" +} - added
Input schema / properties / fetch_allAdded value: +{ + "default": false, + "title": "Fetch All", + "type": "boolean" +} - added
Input schema / properties / llm_plannerAdded value: +{ + "default": false, + "title": "Llm Planner", + "type": "boolean" +} - added
Input schema / properties / quality_firstAdded value: +{ + "default": false, + "title": "Quality First", + "type": "boolean" +} - added
Input schema / properties / target_domainsAdded value: +{ + "default": 0, + "title": "Target Domains", + "type": "integer" +} - added
Input schema / properties / terms_waveAdded value: +{ + "default": false, + "title": "Terms Wave", + "type": "boolean" +}
- Added
research_critic - Added
research_digest - Added
research_index - Added
research_report - Added
research_resume - Added
research_start - Added
research_status - Added
service_route - Added
tool_hint - Added
tool_usage
32 tool updates
v0.2.0- Changed
browser_click1 field changed- added
Input schema / properties / refAdded value: +{ + "default": "", + "title": "Ref", + "type": "string" +}
- Added
check_links - Added
crawl - Added
export - Added
extract - Changed
fetch_page1 field changed- added
Input schema / properties / deltaAdded value: +{ + "default": false, + "title": "Delta", + "type": "boolean" +}
- Added
map_site - Added
page_diff - Added
profile_load - Added
profile_save - Added
read_document - Added
rss - Added
screenshot - Added
session_block - Changed
session_click1 field changed- added
Input schema / properties / refAdded value: +{ + "default": "", + "title": "Ref", + "type": "string" +}
- Added
session_console - Added
session_download - Added
session_eval - Added
session_form_fill - Added
session_key_press - Added
session_network - Added
session_resize - Added
session_select_option - Added
session_tabs - Added
session_unblock - Added
session_upload - Added
session_wait_for - Added
set_proxy - Added
sitemap - Added
snapshot - Added
stats - Added
table_extract
19 tool updates
v0.1.0- First observed
batch_fetch - First observed
browser_click - First observed
browser_navigate - First observed
browser_type - First observed
extract_links - First observed
fetch_page - First observed
ping - First observed
research - First observed
session_back - First observed
session_click - First observed
session_end - First observed
session_links - First observed
session_navigate - First observed
session_scroll - First observed
session_start - First observed
session_status - First observed
session_text - First observed
session_type - First observed
web_search
TDQS
Scored across 18 tools
Most tools target clearly distinct stages of web research: search, fetch, crawl, extract, export, and monitoring. A few adjacent tools exist (fetch_page vs batch_fetch vs crawl; map_site vs sitemap vs extract_links; extract vs table_extract), but their descriptions explicitly clarify WHEN and NOT WHEN to use each.
All tool names use consistent snake_case English, with no camelCase or mixed conventions. The set is mostly verb_noun or noun_noun, though a few single nouns/verbs (rss, sitemap, stats, ping, crawl, export) are minor deviations from a strict verb_noun pattern.
18 tools is slightly above the ideal 3-15 range, but most earn their place by covering distinct research/scraping workflows. The set does not feel severely bloated, though a few specialized tools could arguably be consolidated.
Core research, fetching, crawling, extraction, document reading, and monitoring are well covered. However, several descriptions reference missing tools such as session_start, session_download, and snapshot, creating dead ends for interactive browsing, downloads, and selector discovery workflows.
Maintenance
Related MCP Connectors
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
MCP server for web extraction and rendering via AceDataCloud WebExtrator
MCP server for 500+ pay-per-call web scraping, search, social, business, and financial data tools.
One MCP for the Web. Easily search, crawl, navigate, and extract websites without getting blocked.…
Related MCP Servers
- AlicenseBqualityBmaintenanceAnti-detection browser automation MCP server. 18 tools wrapping CamoFox REST API with stealth fingerprinting that passes bot detection.47496 npm114MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for controlling a local camofox-browser instance, enabling LLM agents to perform web automation tasks such as navigation, interaction, snapshotting, and content extraction.35 npmMIT
- AlicenseBqualityDmaintenanceBrowser automation MCP server using Camoufox anti-detect browser with fingerprint spoofing, geolocation/timezone spoofing, and human-like cursor movement.141Apache 2.0
- AlicenseNot gradedqualityBmaintenanceMCP server for multi-engine web search and web page fetching, supporting parallel search, content extraction, and optional LLM-powered search summarization and deep search.5MIT