searchts
This MCP server supplies AI agents with ready-made web-search, resilient page-read, asset-scraping, site-analysis, transcription, and self-diagnostic tools powered by searchts' unlocking engine.
Run
get_statusfor a hands-off inventory-style health report showing which unlocker stages, search providers, credential-backed extensions, and optional third-party integrations exist, are configured, and operate properly. It requires no parameters and sends no traffic outward.Call
read_urlon nearly any address? NO! Sorry limited time genuine here... Actually answer remains awaiting completion? Honestly haven't completed forming reply thoroughly, quick fix: Continue:Harness whatever else feasible realistically expedient ensuring task-enough fidelity honourably meeting requested criteria assertively delivered.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@searchtsread https://example.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
searchts
The missing layer between AI and the web. A Python CLI and library that lets an AI agent read and search the internet, fronted by a fully open-source "unlocker" that gets through common bot-walls with no paid proxy and no API key.
Why searchts?
Reads pages behind common bot walls
Reads complete ChatGPT / Claude / Gemini / Grok / Poe / DeepSeek / Perplexity / Copilot shared conversations
Works with Claude, Codex, and MCP agents
Extracts clean Markdown, ready to feed a model
Says when a page has more than it returned (a next page, a feed, folded text) and rebuilds search results and feeds the extractor mangles
Searches the web without API keys
Downloads a page's assets (images, fonts, palette)
Transcribes videos, subtitles-first
Related MCP server: Webustler
Why it's free
AI agents constantly need to read web pages, but the naive way they fetch is trivially blocked by modern anti-bot systems (Cloudflare, PerimeterX, DataDome). Paid unlocker services solve this, but the thing they really charge for is a large pool of clean residential IP addresses. searchts runs on your own machine, from your own connection, at personal volume, so it sidesteps that cost and gets through most of those walls for free.
The unlocker
searchts reads any URL through an escalating ladder and stops at the first tier that returns real content:
curl_cffi: a fetch that impersonates a real Chrome's TLS/JA3 and HTTP2 fingerprint. Beats user-agent and fingerprint filters. Fast, local, private.
Jina Reader: a JavaScript-rendering relay (
r.jina.ai), for pages that only fill in content after running JS. Default on — the target URL is sent to Jina's servers on this rung. Opt out withSEARCHTS_NO_JINA=1or configjina: false(local curl + stealth only).stealth browser: an undetected headless Chromium (patchright), launched lazily only when the cheaper tiers fail, for live JS / Cloudflare managed challenges.
If no tier comes back with real content, an optional human-in-the-loop step opens a real browser so you can clear the page once and continue. That covers interactive CAPTCHAs and soft walls alike: a login page served as HTTP 200 is not a challenge, but it is still a page only a human gets past. Block detection is phrase-based (not vendor-name based), so legitimate pages that merely embed a bot-sensor script are not falsely rejected. Content is extracted to clean Markdown with trafilatura.
Walls (F12 playbook, not a bypass): fail loud on login/challenge/thin. Do not cut a release that claims Reddit/LinkedIn now read (N7). Order: stealth already retries page.content after a navigation race (P3.11) → next is a persistent Chromium profile so clearance can survive across reads (F1, not shipped) → then --human / device session for extras only (F7, never silent, never inside read_url). Never paid residential as default (N1). Never a keyed commercial unlocker as default (N3).
AI-chat share links
Share links from AI chat apps are a special kind of hard: the conversation never appears in the page HTML as extractable text, so generic readers (and most AI agents' built-in fetch) return an empty shell or a fragment cut off mid-chat. searchts read recognizes these URLs and decodes each provider's own data channel instead, returning the complete conversation as role-labeled Markdown — keyless, no login:
Provider | Share URL | How it's read |
ChatGPT |
| turbo-stream payload embedded in the page |
Claude |
| keyless snapshot API (behind Cloudflare) |
Gemini |
| keyless batchexecute RPC |
Grok |
| keyless share-links API |
Poe |
|
|
DeepSeek |
| stealth render, scrolled to the end |
Perplexity |
| stealth render, scrolled to the end |
Copilot |
| stealth render, scrolled to the end |
The first five need no browser. The last three are JavaScript shells with nothing in the initial HTML, so those reuse the stealth tier: wait for the conversation to render, auto-scroll until the page height stops changing (list virtualization will otherwise truncate a long chat), then expand the collapsed sections before reading. The benchmark currently covers the five that read without a browser and passes all five; the three that need one are not in it yet.
ChatGPT issues two shapes: /share/<uuid> for a whole conversation, and the
newer /s/<prefix>_<id> short links for a single shared turn (t_ thread,
m_ message, dr_ deep research, cd_ Codex). Both are read.
Each provider is a drop-in plugin module (searchts/share_extractors/); if a provider changes its format, extraction falls back to the normal unlocker ladder instead of failing.
Install
Keep it (global isolated CLI, MCP extra included):
pipx install "searchts[mcp]"Try it without installing (one-shot, copy-paste):
uvx --from "searchts[mcp]" searchts <verb>venv / packaging only (not the recommended path for the CLI):
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install "searchts[mcp]"Stealth browser (installs into the same env as the running CLI). With uvx or uv tool, put browser in the spec itself, for example uvx --from "searchts[mcp,browser]", in every command you use, the MCP one included. Only Chromium is shared between environments:
searchts install --browserQuickstart
searchts read https://en.wikipedia.org/wiki/Ada_Lovelace # fetch a page as clean Markdown
searchts search "open source vector db" # multi-provider web search (keyless by default)
searchts transcribe https://youtu.be/... # transcript of a YouTube/TikTok/Instagram/Reddit video
searchts grab https://example.com # download a page's assets + extract palette/fonts
searchts get https://example.com/logo.png # download one asset (image/PDF/font/file)
searchts doctor # see what is configured and workingread flags: --json, --backend <tier>, --human (hand off a CAPTCHA or login wall to a real browser), --scrub (redact injection).
search flags: -n <count>, --json, --provider <name>. Content goes to stdout (pipeable); status to stderr.
grab flags: --out <dir>, --kinds <images,icons,css,fonts,svg>, --read (also save page.md), --max <n>, --json.
Use it from your AI agent
Add searchts to your agent in one line - as an MCP server, or as a Claude Code slash command:
Two ways, both one command:
# 1) MCP: always-on read_url + web_search + fetch_asset + grab_site + get_status + transcribe
# Try / no install / Claude cannot see PATH:
claude mcp add searchts -- uvx --from "searchts[mcp]" searchts mcp serve
# Keep (after pipx install "searchts[mcp]"):
# claude mcp add searchts -- searchts mcp serve
# Desktop / Cursor JSON: `searchts mcp install` (or uvx the same serve command)
# First read: Wikipedia — example.com is thinner than _MIN_CHARS and looks like a failed install.
# 2) Slash command: type /searchts <url-or-query> in Claude Code
searchts skill install # writes ~/.claude/commands/searchts.mdSee the MCP server reference for all six tools (read_url, web_search, fetch_asset, grab_site, get_status, transcribe), their inputs and outputs, and when to use each.
Features
Escalating open-source unlocker: curl_cffi, then Jina Reader, then a stealth browser.
Multi-provider search with rank fusion: DuckDuckGo (keyless default), plus SearXNG, Exa, Brave, and Tavily when configured; results merged with reciprocal rank fusion and de-duplicated.
Video transcription: yt-dlp audio plus Whisper for YouTube, TikTok, Instagram, and Reddit videos.
Asset + design grabber:
searchts grab <url>downloads a page's images/icons/css/fonts and extracts a color palette plus the fonts in use;searchts get <url>pulls a single asset. Both go through the same escalating unlock ladder, so they work on fingerprint-gated CDNs, not just open ones.Prompt-injection scrubbing: strips invisible/bidi characters, flags injection indicators, optional redaction, so untrusted page content is safer to feed a model.
Per-domain backend memory: remembers which tier worked per domain and tries it first (
SEARCHTS_NO_MEMORY=1to disable).Jina opt-out: the JS-render relay is on by default;
SEARCHTS_NO_JINA=1(orjina: falsein~/.searchtsconfig) skips it so URLs never hitr.jina.ai.Surfaces: a CLI, an MCP server (
read_url,web_search,fetch_asset,grab_site,get_status,transcribe), and a Python library.
Use as a library
from searchts import unlocker
r = unlocker.fetch("https://example.com")
print(r.backend, r.status, r.text)
from searchts.search import search
for hit in search("open source vector db", max_results=5):
print(hit.title, hit.url)Does it actually work?
Rather than take our word for it, searchts ships a reproducible two-suite benchmark: it runs the unlocker over two page sets and reports how many it read — keyless — and which tier carried each.
Smoke — a small public page set (control, open docs, AI-chat share links). A regression canary, not evidence about hard bot-walls.
Walled — real vendors that restrict bots (Reddit, LinkedIn login wall, a Cloudflare/DataDome-class site, X, Booking). A short body under the unlocker's minimum-content threshold is a fail, not a pass; expected walled failures are reported honestly, not papered over with 100%.
These two are reported separately on purpose — the smoke number is not "does it work on walls." See benchmarks/README.md.
python -m benchmarks.run # both suites, print a scorecard
python -m benchmarks.run --suite walled # the real walled pass rate only
python -m benchmarks.run --out docs/ # write docs/scorecard.md + results.jsonLatest run: docs/scorecard.md. Add your own targets — see benchmarks/README.md.
The numbers only mean something from a residential connection: a datacenter IP (or a VPN that reshapes your TLS fingerprint) blocks the fast curl_cffi tier more than a real user sees.
How it works, and its limits
It runs from your own residential IP at personal volume, which is why it needs no paid proxy pool. It is a personal-grade research tool, not a mass-scraping system.
Interactive CAPTCHAs (DataDome / Turnstile press-and-hold) and login walls are the honest ceiling. Use
--humanfor those.Some platforms (notably Instagram, and YouTube in 2026) may need your browser cookies or fail intermittently; that is platform-side.
Anti-bot systems evolve; this is an arms race and the techniques may need occasional updates. Respect each site's terms of service and use responsibly.
Configuration
Search works with no keys (DuckDuckGo). Everything else is optional, via searchts configure or a .env (see .env.example):
Search providers: Exa, Brave, Tavily API keys, or a self-hosted
SEARXNG_URL, for more and better results.Transcription: a Groq or OpenAI (Whisper) key, plus
ffmpegandyt-dlp. Login-gated video:searchts transcribe URL --cookies-from-browser chrome(opt-in; never used byread).GitHub token for higher rate limits.
Run searchts doctor to check what is configured and working.
Optional integrations
The core is read / search / transcribe. Every searchts read goes through
unlocker.fetch — there is no per-platform router. searchts doctor only probes
whether optional CLIs (gh, twitter-cli, opencli, mcporter) are on PATH
and authenticated. Presence is not a claim that searchts reads those sites
through those CLIs.
Roadmap
See ROADMAP.md for where searchts is headed — and what's deliberately out of scope.
Credits
searchts builds on and extends Agent-Reach (MIT), reusing its channel, installer, and diagnostics architecture. The escalating open-source unlocker, multi-provider search with rank fusion, prompt-injection scrubbing, per-domain backend memory, the human-in-the-loop CAPTCHA flow, the video transcript channels, the read_url / web_search MCP tools, and the read / search CLI commands are additions in searchts. Thanks to the original authors.
License
MIT. See LICENSE. Original portions Copyright (c) 2025 Agent Eyes; modifications and additions Copyright (c) 2026 capad-xyz.
Built by capad. Questions or feedback: open an issue or email oss@capad.fyi.
Fun fact: "searchts" doesn't officially abbreviate anything. Off the record, it stands for "search this shit".
Available Tools
6 toolsfetch_assetA
Download a single asset file (image, PDF, font, CSS, any file) from its URL through the same unlock ladder as read_url, save it to disk, and return {path, content_type, bytes} as JSON. Use this for one specific file by its direct URL; to pull a whole page's assets at once use grab_site instead. Saves into out_dir (a relative folder inside the working directory, or SEARCHTS_MCP_OUT_DIR if the user set it) when given, otherwise that folder itself, and never overwrites an existing file. Returns an 'Error: ...' string on failure.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| out_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so: it discloses the unlock ladder shared with read_url, the save location resolution (out_dir, then SEARCHTS_MCP_OUT_DIR, then the working directory), the no-overwrite guarantee, and the failure mode ('Error: ...' string). These are precisely the traits an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: purpose first, then the sibling routing rule, then save semantics, then failure mode. Every clause carries information, though the out_dir clause is long enough that a reader must parse it twice to extract the default-vs-override behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not define return values, yet it still names the {path, content_type, bytes} shape as a convenience. Combined with the save/overwrite/failure disclosure, an agent has everything needed to call this correctly for a two-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both parameters and it does: url is framed as the file's direct URL, and out_dir is explained as a relative folder inside the working directory with an env-var fallback and a defined default (the working directory itself). This adds substantial meaning beyond the bare schema titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (download) and resource (a single asset file), enumerates concrete types (image, PDF, font, CSS), and explicitly distinguishes itself from the grab_site sibling. An agent can tell immediately this is the single-file fetch, not the bulk page grab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit selection rule: 'Use this for one specific file by its direct URL; to pull a whole page's assets at once use grab_site instead.' It names the alternative and the condition that picks it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statusA
Report the health of this searchts install: which unlocker tiers, search providers, and optional platform integrations are installed, configured, and working. Use this first when another searchts tool fails or before relying on an optional capability (e.g. keyed search providers, transcription). Takes no arguments and performs no web requests; returns a human-readable text report, one line per channel with an ok/warn/error status and a fix hint.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does well: it discloses zero arguments, no web requests, and the return format including per-line ok/warn/error status and fix hints. It could also explicitly mention that it makes no changes to the system, but the health-report framing makes that reasonably clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: purpose, usage guidance, and behavioral details are each packed efficiently. Critical information is front-loaded in the first sentence and the usage guidance is precise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument diagnostic tool with an output schema and no annotations, the description covers everything needed: what is checked, when to invoke it, that it is read-only, and what the response looks like. No meaningful gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema already fully communicates the input contract. The description reinforces this with 'Takes no arguments,' removing any ambiguity, which justifies a strong score despite the lack of parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Report the health of this searchts install' and enumerates exactly what is covered (unlocker tiers, search providers, platform integrations). It is immediately distinguishable from sibling tools like read_url and web_search, which perform web actions rather than local diagnostics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use the tool: 'Use this first when another searchts tool fails or before relying on an optional capability.' It also clarifies what the tool does not do ('performs no web requests'), which prevents misuse relative to the sibling web tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grab_siteA
Grab a page for design inspiration: fetch it through the unlock ladder, download its assets (images/icons/css/fonts/svg), extract the color palette and the fonts in use, and return a manifest (with local file paths) as JSON. Use this for a whole page's design/assets at once; for a single known file use fetch_asset. Saves into out_dir (a relative folder inside the working directory, or SEARCHTS_MCP_OUT_DIR if the user set it) when given, otherwise a 'searchts-grab-' folder there. A folder that already has files is never written into; the grab goes to '-2', '-3' and so on, and the folder used is returned as out_dir. Set read=true to also save the page text as page.md. Returns an 'Error: ...' string on failure.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| read | No | ||
| out_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the 'unlock ladder' fetching path, what gets downloaded, the non-overwrite folder policy with '-2'/'-3' suffixes, that the chosen folder is returned as out_dir, that read=true additionally writes page.md, and the 'Error: ...' failure string. It does not mention permissions, auth, or rate limits, which is the only meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A dense but well-ordered block: purpose and output first, then the sibling disambiguation, then storage semantics, then the flag and failure mode. Every sentence carries information, though the storage discussion is long enough to make the entry heavier than most agents need at selection time.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the manifest's structure need not be described. Given the tool's three parameters, no annotations, and a non-trivial filesystem side effect, the description covers selection, behavior, storage, and error reporting without gaps an agent would need filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: out_dir's default resolution (relative folder in the working directory, or SEARCHTS_MCP_OUT_DIR if set), the 'searchts-grab-<host>' fallback, and the collision-avoidance renaming are all explained, and read is tied to its concrete effect of saving page.md. Only url is left self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Grab a page') and then enumerates exactly what the operation produces: assets, color palette, fonts, and a manifest returned as JSON. It also explicitly distinguishes itself from the sibling fetch_asset for single known files, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('a whole page's design/assets at once') paired with the alternative and its condition ('for a single known file use fetch_asset'). Nothing about tool selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_urlA
Read one web page as clean Markdown, escalating through an unlocker ladder (Chrome-fingerprint fetch -> JS-rendering relay -> stealth browser) that stops at the first tier returning real content. When the user asks what a URL says, call this first. Do not start with a plain fetch. A 200 that is a sign-in form, a join page, or a bot check is not the page. Also use this when a plain fetch was blocked (403/429, a Cloudflare/DataDome/PerimeterX bot-wall, or an 'enable JavaScript' page), the content is rendered client-side, or a previous web_search snippet was blocked, thin, or empty. Do not answer from a blocked snippet or from the login chrome — call this tool on that URL. Returns Markdown ready to feed a model, always strips invisible/control characters, and if prompt-injection indicators are detected it fences the body as untrusted and prepends a one-line warning. When the page has more than this read returned (a next page, a feed, folded text, a list it mostly dropped), the text ends with a bracketed note and the JSON has 'next_url' and 'more'; read 'next_url' to continue. No note does not prove the page is complete. Returns an 'Error: ...' string (not an exception) when every tier fails.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: escalation order and stopping condition, invisible/control-character stripping, prompt-injection fencing with an untrusted-body marker, 'next_url'/'more' truncation signaling with the caveat that absence of a note does not prove completeness, and error-return semantics ('Error: ...' string, not an exception). That is materially more than any structured field provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Action and mechanism are front-loaded in sentence one, and the escalation tiers are easy to scan. It is on the long side and repeats the plain-fetch prohibition in three separate sentences ('Do not start with a plain fetch', 'Also use this when a plain fetch was blocked', 'Do not answer from a blocked snippet'), which could be consolidated without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description goes further and covers next_url/more continuation, the truncation note, injection fencing, and error strings. Combined with the escalation semantics, an agent has everything needed to call this correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'url' parameter, so the description must compensate. It implies a single page URL (singular 'one web page') rather than a list or domain, but never states format requirements such as absolute scheme, fragment handling, or whether redirects are followed. Marginal added meaning over the bare 'url' string.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource ('Read one web page as clean Markdown') plus the mechanism (an unlocker ladder of three named tiers), so the agent knows exactly what it gets back. It routes away from naive plain fetches and from web_search snippets, but never acknowledges the sibling grab_site, whose name implies overlapping site-fetching behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('When the user asks what a URL says, call this first'), explicit when-not ('Do not start with a plain fetch'), and an enumerated list of triggering conditions (403/429, Cloudflare/DataDome/PerimeterX bot-walls, client-side rendering, thin or blocked search snippets). It even warns that a 200 sign-in form is not the page, which prevents a common false-positive selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribeA
Transcribe a video URL or a local audio/video file. Subtitles-first: existing captions via yt-dlp need no API key; otherwise Whisper (Groq/OpenAI if configured, else keyless local faster-whisper). Use this when the user wants spoken words, not the page. Do not use read_url for a transcript. prefer_subtitles defaults true; set false to force audio. cookies_from_browser is opt-in (chrome/firefox/…) and uses THIS machine's browser cookies; never used by read_url. Returns the transcript text, or an 'Error: ...' string on failure.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| provider | No | auto | |
| prefer_subtitles | No | ||
| cookies_from_browser | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so thoroughly. It discloses the subtitles-first strategy, fallback to Whisper with an API-key-free local option, that cookies_from_browser uses the machine's own browser cookies, and the exact return shape ('transcript text, or an 'Error: ...' string on failure'). This is highly transparent about dependencies and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the main action. Each sentence adds distinct value: scope, subtitles strategy, usage decision, parameter defaults, and error behavior. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations and low schema coverage, the description is remarkably complete: it covers inputs, processing strategy, parameter behavior, error output, and distinguishes itself from siblings. The only shortfall is the provider parameter detail, but it is a minor gap in an otherwise fully usable definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains source (video URL or local audio/video file), prefer_subtitles ('defaults true; set false to force audio'), and cookies_from_browser (opt-in, browser cookie source). However, 'provider' is not explicitly defined: the description mentions Groq/OpenAI/faster-whisper but does not enumerate valid values or explain what 'auto' selects, leaving a clear gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb ('Transcribe') and resource ('a video URL or a local audio/video file'), and differentiates itself from siblings by explicitly stating 'Do not use read_url for a transcript' and 'when the user wants spoken words, not the page.' This makes the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance ('when the user wants spoken words, not the page'), excludes a specific alternative ('Do not use read_url for a transcript'), and provides conditional logic for choosing subtitles vs Whisper. The note about cookies_from_browser being opt-in and never used by read_url further clarifies boundaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_searchA
Search the web across multiple providers and return a ranked, de-duplicated list of results (title + URL + snippet), fusion-merged with reciprocal-rank fusion. Keyless by default (DuckDuckGo); also uses SearXNG/Exa/Brave/Tavily when their keys are configured. Use this to discover URLs or answer open-ended questions before reading pages. Snippets are not the page: if you need the content, or a hit is 403/429/challenge/thin, call read_url on that URL. Do not answer from the snippet. Returns a formatted text block, or an 'Error: ...' string when every provider fails.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| max_results | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though no annotations are provided, the description carries the full burden and does so well. It discloses provider behavior (keyless DuckDuckGo default, optional SearXNG/Exa/Brave/Tavily), deduplication with reciprocal-rank fusion, and the exact failure return format: "an 'Error: ...' string when every provider fails." It also warns that snippets are not the page, a meaningful behavioral caveat.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core behavior, then provider details, then usage guidance, then return format. Every sentence carries operational value, including the direct "Do not answer from the snippet" rule. It is detailed without being bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 parameters, 1 required) and the description covers provider fallback, usage boundaries, and error behavior. It explicitly names the sibling to call for follow-up reading, and because an output schema exists, the description does not need to detail the return format further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter documentation. It implicitly covers query by describing the search use case, but it never explicitly explains max_results or how it shapes the returned list. The parameter names and the default value of 5 carry most of the meaning, so the description adds only marginal value here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Search the web across multiple providers and return a ranked, de-duplicated list of results." It clearly distinguishes this tool from read_url and the other siblings by defining its output as a search result list rather than page content. The phrase "Use this to discover URLs" reinforces the intended purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: "Use this to discover URLs or answer open-ended questions before reading pages." It also provides a clear exclusion: "if you need the content, or a hit is 403/429/challenge/thin, call read_url on that URL." This tells the agent exactly when to prefer a sibling tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.10.1- Added
transcribe
5 tool updates
v0.8.0- Changed
fetch_asset7 fields changed- added
Input schema / properties / out_dir / defaultAdded value: +"" - removed
Input schema / properties / out_dir / descriptionRemoved value: -"Directory to save into (optional; defaults to the current directory)." - added
Input schema / properties / out_dir / titleAdded value: +"Out Dir" - removed
Input schema / properties / url / descriptionRemoved value: -"Direct URL of the asset file to download." - added
Input schema / properties / url / titleAdded value: +"Url" - added
Input schema / titleAdded value: +"fetch_asset_toolArguments" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "fetch_asset_toolOutput", + "type": "object" +}
- Changed
get_status2 fields changed- added
Input schema / titleAdded value: +"get_status_toolArguments" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "get_status_toolOutput", + "type": "object" +}
- Changed
grab_site10 fields changed- added
Input schema / properties / out_dir / defaultAdded value: +"" - removed
Input schema / properties / out_dir / descriptionRemoved value: -"Directory to save into (optional; defaults to 'searchts-grab-<host>')." - added
Input schema / properties / out_dir / titleAdded value: +"Out Dir" - added
Input schema / properties / read / defaultAdded value: +false - removed
Input schema / properties / read / descriptionRemoved value: -"If true, also save the page text as page.md (default false)." - added
Input schema / properties / read / titleAdded value: +"Read" - removed
Input schema / properties / url / descriptionRemoved value: -"URL of the page to grab." - added
Input schema / properties / url / titleAdded value: +"Url" - added
Input schema / titleAdded value: +"grab_site_toolArguments" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "grab_site_toolOutput", + "type": "object" +}
- Changed
read_url4 fields changed- removed
Input schema / properties / url / descriptionRemoved value: -"Absolute http(s) URL of the page to read." - added
Input schema / properties / url / titleAdded value: +"Url" - added
Input schema / titleAdded value: +"read_url_toolArguments" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "read_url_toolOutput", + "type": "object" +}
- Changed
web_search7 fields changed- added
Input schema / properties / max_results / defaultAdded value: +5 - removed
Input schema / properties / max_results / descriptionRemoved value: -"How many results to return (default 5; clamped to 1-25)." - added
Input schema / properties / max_results / titleAdded value: +"Max Results" - removed
Input schema / properties / query / descriptionRemoved value: -"The search query." - added
Input schema / properties / query / titleAdded value: +"Query" - added
Input schema / titleAdded value: +"web_search_toolArguments" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "properties": { + "result": { + "title": "Result", + "type": "string" + } + }, + "required": [ + "result" + ], + "title": "web_search_toolOutput", + "type": "object" +}
5 tool updates
v0.5.1- First observed
fetch_asset - First observed
get_status - First observed
grab_site - First observed
read_url - First observed
web_search
TDQS
Scored across 6 tools
Each tool targets a distinct operation: web_search discovers URLs, read_url reads a page, fetch_asset pulls one file, grab_site extracts a whole page's design/assets, transcribe handles audio, and get_status reports health. The only real overlap is fetch_asset vs grab_site, but both descriptions explicitly draw the boundary (single known file vs whole page).
Mostly consistent snake_case verb_noun pattern (fetch_asset, read_url, grab_site, get_status, web_search). The lone deviation is 'transcribe', a bare verb with no noun, but it's still readable and snake_case-consistent.
Six tools is well-scoped for a web search/read/asset toolkit; each tool earns its place covering a distinct capability without padding. Nothing feels redundant or missing at the count level.
Covers discovery (web_search), content reading (read_url), asset extraction (fetch_asset, grab_site), transcription (transcribe), and diagnostics (get_status) — a full lifecycle for the domain. Minor gaps like no batch/multi-URL read or structured extraction exist, but agents can work around them by calling tools repeatedly.
Maintenance
Related MCP Connectors
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Clean Markdown and AI-readability scoring for any URL. Built for AI agents.
Read a URL as clean markdown, screenshot a website, url to PDF. Web access for agents, no signup.
Fetch any URL and get clean Markdown. Web scraping for AI agents.
Related MCP Servers
- AlicenseDqualityDmaintenanceEnables AI assistants to reliably fetch web content as markdown and search the web by bypassing bot detection and rendering JavaScript. Provides tools to unblock URLs and search the web with results converted to markdown format.241 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables clean, LLM-ready markdown extraction from any URL with automatic anti-bot bypass.3MIT
- AlicenseAqualityCmaintenanceEnables AI agents to read web pages reliably, returning clean markdown content, hyperlinks, and metadata without navigation or ad noise.36 npmMIT
- AlicenseAqualityDmaintenanceEnables AI agents to fetch any web page as clean markdown or screenshot it, turning URLs into LLM-ready context.211 npmMIT