PyScrappy
PyScrappy is a web scraping MCP server that provides AI agents with 22 tools to retrieve structured data from across the web — from general URLs to specialized platforms.
General Web Scraping
scrape_url— Scrape any URL for structured text, links, images, tables, and metadata; supports CSS selectors, pagination, and optional JavaScript rendering
Data & Research
scrape_wikipedia— Fetch Wikipedia articles in full, paragraph, or header modescrape_stock— Get Yahoo Finance stock quotes, historical price data, and company profilesscrape_news— Fetch articles from RSS/Atom feeds, auto-discover feeds from a news site, or extract full text from a single article URLsearch_images— Search for images and return URLs and metadata (default engine: Bing)search_youtube— Search YouTube videos and return titles, channels, links, and metadatasearch_linkedin_jobs— Search public LinkedIn job postings by keyword and locationsearch_github— Search GitHub repositories by query, sortable by stars, forks, or recencysearch_hackernews— Search Hacker News stories by relevance or datesearch_books— Search books via Open Library by title, author, or free textget_weather— Get current weather (temperature, humidity, wind, condition) for any location (no API key required)get_crypto— Get cryptocurrency prices, market cap, and 24h change via CoinGeckoconvert_currency— Get exchange rates and convert amounts between currenciesdefine_word— Look up English word definitions, part of speech, and usage exampleslookup_movie— Look up movie/TV info from IMDB via the OMDb API (requiresOMDB_API_KEY)
E-Commerce
search_amazon— Search Amazon products and return title, price, rating, and imagesearch_newegg— Search Newegg for electronics and computer hardwaresearch_ikea— Search IKEA furniture and home products with per-country pricing
Food Delivery
search_ubereats— List Uber Eats restaurants delivering in a cityget_ubereats_menu— Retrieve a full menu (items and prices) for a specific Uber Eats restaurantscrape_zomato— Search restaurants on Zomato by city (with optional cuisine/name filter)
Entertainment
search_soundcloud— Search SoundCloud tracks (uses browser backend for JS rendering)
Scrapes Amazon marketplace for product listings, including titles, prices, and details.
Scrapes GitHub repositories or profiles (details not fully shown in excerpt but listed as built-in scraper).
Scrapes IKEA product search results per country, including prices and details.
Fetches movie and TV information via OMDb API, such as title, year, rating, and genre.
Scrapes Newegg electronics and hardware product listings.
Scrapes RSS feeds (e.g., news articles) from any provided feed URL.
Searches SoundCloud for tracks and returns metadata like title and plays.
Scrapes Uber Eats restaurant listings and menus by city and locale.
Scrapes Wikipedia articles, summaries, and infoboxes by query.
Searches YouTube for videos and returns metadata such as title, views, and URL.
Scrapes Zomato restaurant listings by city.
PyScrappy is an AI-native web scraping toolkit that turns websites into structured, LLM-ready data. Use it as a Python library or expose it as an MCP server for AI agents.
📖 Documentation: pyscrappy.vercel.app
Key features
Generic scraper — give it any URL, get back structured text, links, images, tables, and metadata
LLM-ready output —
.to_markdown()turns any result into clean Markdown; also.to_json()and.to_dataframe()MCP server — expose the scrapers as tools for AI agents (Claude, Cursor, local LLMs, …)
JS rendering — optional Playwright backend for JavaScript-heavy sites
Custom selectors — pass CSS selectors to extract exactly what you need
Chainable
Selector— navigate HTML directly with CSS/XPath,find_all,find_by_text, andfind_similar(Scrapy/BeautifulSoup-style)Adaptive (self-healing) selectors — remember an element and relocate it by similarity when a site changes its markup, so scrapers don't silently break
Concurrent scraping —
scrape_many/scrape_allrun scrapes in parallelSitemap crawling — enumerate and scrape a whole site from its
sitemap.xml(index + gzip aware)Proxy & scraping-API support — route through a proxy or ScraperAPI/ScrapeOps for blocked sites
TLS-fingerprint impersonation —
impersonate="chrome"gets past anti-bot filters that block plain clients (optionalcurl_cffibackend)Command-line extract —
pyscrappy extract <url> out.mdscrapes a URL straight to a file, no codeRetry & rate-limiting — built-in exponential backoff and per-domain rate limiting
Type-safe — full type hints,
py.typedmarker20+ built-in scrapers — Wikipedia, IMDB, stocks, news, GitHub, Amazon/IKEA, YouTube, and more
Related MCP server: mcp-firecrawl
Installation
pip install pyscrappyOptional extras:
# Browser support (for JS-rendered pages)
pip install 'pyscrappy[browser]'
playwright install chromium
# DataFrame support
pip install 'pyscrappy[dataframe]'
# MCP server (use PyScrappy's scrapers as AI-agent tools)
pip install 'pyscrappy[mcp]'
# Stealth (TLS-fingerprint impersonation to bypass anti-bot filters)
pip install 'pyscrappy[stealth]'
# Parquet / Excel export (ScrapeResult.to_parquet() / .to_excel())
pip install 'pyscrappy[parquet]'
pip install 'pyscrappy[excel]'
# Everything
pip install 'pyscrappy[all]'For AI agents
PyScrappy ships an MCP server that exposes its scrapers as tools, so an agent (Claude, Cursor, an OpenAI agent, a local LLM) can pull structured web data from any URL and hand it straight to the model:
AI agent ──MCP tool call──▶ PyScrappy ──fetch + extract──▶ Any website
▲ │
└────────────── clean Markdown / JSON ◀───────────────────────┘pip install 'pyscrappy[mcp]'
claude mcp add pyscrappy pyscrappy-mcpThen just ask: "use pyscrappy to summarize the latest headlines from bbc.com." See MCP server for the full setup and tool list.
Local models (Ollama), no MCP host needed
Ollama can't talk MCP on its own, so normally you'd run a host (Goose, Cline, …) in between. PyScrappy skips that with a built-in agent that talks to Ollama directly and lets a local model call the scrapers as tools:
pip install 'pyscrappy[mcp]' # needs Python 3.10+
pyscrappy chat --model qwen2.5 "what's the current AAPL quote?"It exposes the same 22 tools as the MCP server. The only requirement is a model
that supports tool calling (Llama 3.1, Qwen 2.5, Mistral, …); how well it
picks the right tool is up to the model. Point it at a remote Ollama with
--host, and pass -v to see each tool call.
MCP server (use PyScrappy from an AI agent)
PyScrappy ships an optional Model Context Protocol server, so an AI agent (e.g. Claude) can call PyScrappy's scrapers as tools and get structured web data back.
pip install 'pyscrappy[mcp]'The MCP extra installs the standalone fastmcp package and requires Python 3.10
or newer. On Python 3.9 the core scraping library still works, but the MCP server
is unavailable.
This installs the pyscrappy-mcp command. It uses stdio by default for local MCP
clients; Streamable HTTP and legacy SSE are available for remote deployments:
pyscrappy-mcp # stdio (default)
pyscrappy-mcp --http # Streamable HTTP
pyscrappy-mcp --sse # legacy SSEYou can also run the stdio server with python -m pyscrappy.mcp.
Register with Claude Code
claude mcp add pyscrappy pyscrappy-mcpRegister with Claude Desktop
Add to your claude_desktop_config.json and restart the app:
{
"mcpServers": {
"pyscrappy": {
"command": "pyscrappy-mcp"
}
}
}Tip: Claude Desktop does not inherit your shell
PATH. Ifpyscrappy-mcpis not found, use the absolute path to the command (e.g. the one printed bywhich pyscrappy-mcp).
Available tools
The server exposes 20+ tools. The most common ones are scrape_url (any
URL → text, links, images, tables, metadata), scrape_wikipedia,
scrape_stock, scrape_news, and search_github — plus many more
covering image/YouTube/LinkedIn/Hacker News/book search, weather, crypto,
currency, dictionary, Amazon/Newegg/IKEA/SoundCloud, IMDB, and Zomato/Uber Eats.
To see the full, live list, ask the agent to call the list_available_scrapers
tool, or from a shell:
python -c "from pyscrappy import list_scrapers; print(', '.join(sorted(list_scrapers())))"The lookup_movie tool needs a free OMDb API
key. Pass it to the server through your MCP client config, e.g. for Claude Desktop:
{
"mcpServers": {
"pyscrappy": {
"command": "pyscrappy-mcp",
"env": { "OMDB_API_KEY": "your-key" }
}
}
}Once registered, just ask the agent naturally, e.g. "use pyscrappy to get the latest headlines from bbc.co.uk and the AAPL stock quote."
Built-in scrapers
PyScrappy ships 24 built-in scrapers, and every one that works without a proxy is also exposed as an MCP tool.
A few of them:
GenericScraper— scrape any URL with auto-extraction (text, links, images, tables, metadata)Data / research —
WikipediaScraper,StockScraper(Yahoo Finance),NewsScraper(RSS/Atom),GitHubScraper,HackerNewsScraper, plus weather, crypto, currency, dictionary, image, LinkedIn-jobs, and book searchE-commerce —
AmazonScraper,NeweggScraper,IKEAScraperSocial / media / food —
YouTubeScraper, SoundCloud, Zomato, Uber Eats (Instagram / Twitter / Spotify also ship, but are blocked and need a proxy)
…and many more. To see the full, live list:
python -c "from pyscrappy import list_scrapers; print(', '.join(sorted(list_scrapers())))"IMDBScraper (lookup_movie) is the one exception that needs a key — a free
OMDb OMDB_API_KEY (see the
MCP config above for how to pass it).
Plugins
PyScrappy is extensible: you can add your own scrapers, and third parties can
ship them as standalone pyscrappy-<name> packages. A registered scraper works
everywhere a built-in does, including the MCP server and the pyscrappy chat
agent, with no change to PyScrappy core.
In your own code — register with the decorator:
from pyscrappy import BaseScraper, register_scraper, get_scraper
from pyscrappy.core.models import ScrapeResult, ScrapeMetadata
@register_scraper("reddit")
class RedditScraper(BaseScraper):
def scrape(self, subreddit: str, **kwargs) -> ScrapeResult:
data = self.fetch_and_parse(f"https://old.reddit.com/r/{subreddit}/.json")
# ... build a list of dicts ...
return ScrapeResult(data=[...], metadata=ScrapeMetadata(scraper="reddit"))
get_scraper("reddit")().scrape(subreddit="python")As a distributable package — advertise an entry point in your
pyproject.toml, and PyScrappy discovers it once your package is installed:
[project.entry-points."pyscrappy.scrapers"]
reddit = "pyscrappy_reddit:RedditScraper"After pip install pyscrappy-reddit, the scraper shows up in
list_scrapers(), and an AI agent can call it via the scrape_with MCP tool —
no core change required.
First-class MCP tools (optional). Add an mcp_tools mapping and your scraper
becomes a dedicated, typed MCP tool instead of only being reachable through the
generic scrape_with — its schema is derived from the method signature, so
agents get proper named arguments:
@register_scraper("reddit")
class RedditScraper(BaseScraper):
mcp_tools = {"search_reddit": "scrape"} # tool name -> method
def scrape(self, subreddit: str, sort: str = "hot") -> ScrapeResult:
...See the plugin template for a complete, copyable starting point, and the plugin guide for the full walkthrough.
Quick start
Scrape any URL → clean, LLM-ready Markdown
from pyscrappy import scrape
result = scrape("https://en.wikipedia.org/wiki/Web_scraping")
print(result.to_markdown()) # feed straight to an LLM
# ...or result.to_json() / result.to_dataframe()
# Write to a file — format inferred from the extension:
result.save("out.json") # .json .csv .md .ndjson .yaml .parquet .xlsx
# (.parquet needs pyscrappy[parquet]; .xlsx needs pyscrappy[excel])Prefer raw fields? Every result is a ScrapeResult with .data (a list of
dicts):
print(result.data[0]["metadata"]["title"])
print(result.data[0]["text"]["word_count"])Custom CSS selectors
from pyscrappy import GenericScraper
with GenericScraper() as gs:
result = gs.scrape(
url="https://news.ycombinator.com",
selectors={"title": ".titleline a", "score": ".score"},
)
for item in result.data:
print(item["title"], item.get("score", ""))Navigate HTML with Selector
When you want to traverse markup directly (Scrapy/BeautifulSoup-style) rather than
get back structured dicts, use Selector:
from pyscrappy import Selector
page = Selector(html) # or navigate any HTML string
page.css(".title::text").getall() # CSS with ::text / ::attr(name)
page.xpath("//a/@href").getall() # XPath (elements, text(), @attr)
page.find_all("h2", class_="title") # BeautifulSoup-style search
page.find_by_text("Add to cart", tag="button") # search by text content
first = page.css(".product")[0]
first.css(".price::text").get() # chainable
first.find_similar() # sibling elements shaped like this onecss() / xpath() return a SelectorList with .get() / .getall() / .text().
find_similar() locates elements with the same tag and overlapping classes, handy
for pulling every card/row once you've found one.
Adaptive (self-healing) selectors
A hard-coded CSS selector silently breaks the day a site changes its markup. Adaptive selectors survive that: save a fingerprint of the element the first time, and if the selector later matches nothing, relocate it by structural and textual similarity instead of returning empty.
from pyscrappy import Selector
# First run: match normally and remember this element under an id.
page = Selector(html_v1, url="https://shop.example.com")
price = page.css(".price", auto_save=True, adaptive_id="price").get()
# Later, after a redesign renamed ".price" — heal instead of breaking.
# `expect` is an optional contract: the relocated element must satisfy it,
# so a good structural score can't smuggle in the wrong field.
page = Selector(html_v2, url="https://shop.example.com")
result = page.css(
".price",
adaptive=True,
adaptive_id="price",
expect=lambda s: s.text().startswith("$"),
)
print(result.get(), "→ confidence:", result.adaptive_confidence)How the relocation decides — and where it's stronger than a naive similarity match:
Weighted signals, not a flat average. A stable
id/data-*hook counts far more than a sibling-tag list, so weak signals can't outvote strong ones.Anchor-relative. It remembers the nearest stable ancestor (an id'd /
data-*container) and depth, so it survives layout reshuffles that move absolute positions.Volatility-aware text. Prices, dates, and counts are down-weighted, so healing stays reliable on exactly the fields that change most between scrapes.
Confidence-scored.
SelectorList.adaptive_confidence(0-100) tells you how sure the relocation was;threshold=sets the minimum to accept.Contract-enforced (opt-in). Pass
expect=<callable>to require the healed element to satisfy an invariant (e.g. "text looks like a price"). A heal that clears the threshold but fails the contract is rejected, so structural similarity alone never redefines what a field means.
A heal is a change to what a selector resolves to, so every accepted heal is
recorded. The store keeps an append-only audit log (adaptive.heal.ndjson
beside the fingerprint store) with the confidence, the runner-up gap, and the
before/after fingerprint, readable via store.heal_log() — so drift stays
observable instead of being silently absorbed. For an at-a-glance summary,
store.heal_report() aggregates the log into one row per selector (heal count,
latest/lowest/average confidence, when it last healed), sorted most-healed first
— so the selectors that have drifted the most, and the shakiest relocations
(lowest confidence), surface at the top for a human to review.
Fingerprints persist in a small JSON store (~/.pyscrappy/adaptive.json by
default, or $PYSCRAPPY_HOME), namespaced by site so the same adaptive_id on
two sites never collides. Adaptive is entirely opt-in: without adaptive=True, a
broken selector still just returns empty, exactly as before.
Site-specific scrapers
Every built-in scraper follows the same pattern — instantiate, scrape(...),
read result.data (or .to_dataframe() / .to_markdown()):
from pyscrappy import WikipediaScraper
with WikipediaScraper() as ws:
result = ws.scrape(query="Python (programming language)", mode="summary")
print(result.data[0]["text"])Each scraper has its own arguments (Wikipedia, stocks, IMDB, news, YouTube, Amazon/Newegg/IKEA, Uber Eats, and more — see the full list). For per-scraper arguments and examples, see the documentation.
From the command line
Scrape a URL straight to a file without writing any code — the output format is inferred from the file extension:
pyscrappy extract https://example.com out.md # clean Markdown
pyscrappy extract https://example.com out.json # structured JSON
pyscrappy extract https://example.com out.txt # extracted page text
pyscrappy extract https://example.com out.html # raw fetched HTML
# Narrow to elements matching a CSS selector, or render JS first:
pyscrappy extract https://example.com items.txt --css-selector ".product"
pyscrappy extract https://example.com page.md --render-jsConfiguration
from pyscrappy import ScraperConfig, GenericScraper
config = ScraperConfig(
timeout=20.0, # request timeout in seconds
max_retries=3, # retry failed requests
retry_jitter=True, # spread exponential retries to avoid lockstep traffic
rate_limit=2.0, # seconds between requests per domain
proxy="http://...", # proxy URL, or a list to rotate through
scraper_api=None, # route via a scraping-API service (see below)
headless=True, # browser runs headless
render_js="auto", # auto-detect if JS rendering is needed
cache_ttl=0, # response cache TTL in seconds (0 = disabled)
cache_dir=None, # also persist the cache to disk (survives restarts)
cache_dir_max_size=512, # max live entries kept on disk before oldest are pruned
impersonate=None, # e.g. "chrome" to spoof a browser's TLS fingerprint (see below)
)
with GenericScraper(config) as gs:
result = gs.scrape(url="https://example.com")Proxies and blocked sites
Some sites (e.g. eBay, Instagram, Twitter/X, Spotify) block direct automated requests. PyScrappy supports two ways to get through them.
A proxy (or a rotating list) — applies to both the HTTP and browser backends:
from pyscrappy import ScraperConfig, AmazonScraper
# Single proxy
config = ScraperConfig(proxy="http://user:pass@host:port")
# Rotating list (one picked per request)
config = ScraperConfig(proxy=["http://p1:8080", "http://p2:8080"])A scraping-API service (ScraperAPI, ScrapeOps, ScrapingBee) — routes requests through the service, which handles proxies and anti-bot challenges for you:
config = ScraperConfig(scraper_api={
"provider": "scraperapi", # or "scrapeops", "scrapingbee"
"api_key": "YOUR_KEY",
"render_js": True, # optional
})
# Now any scraper works through the service, unchanged:
with AmazonScraper(config) as scraper:
result = scraper.scrape(query="laptop")This is the reliable way to use the scrapers marked "needs proxy" above.
TLS-fingerprint impersonation — many anti-bot systems block a plain HTTP
client by its TLS/JA3 fingerprint before serving any content. Set impersonate
to mimic a real browser's fingerprint and get past that class of block without a
headless browser:
from pyscrappy import ScraperConfig, GenericScraper
# needs the optional extra: pip install 'pyscrappy[stealth]'
config = ScraperConfig(impersonate="chrome") # or "chrome124", "safari", "firefox"
with GenericScraper(config) as gs:
result = gs.scrape("https://example.com")Impersonation works on both the sync and async paths (async uses
curl_cffi's AsyncSession), so you can combine stealth with high-throughput
async scraping. All the usual retry, rate-limiting, caching, and robots handling
still apply.
import asyncio
from pyscrappy import scrape_async, ScraperConfig
async def main():
cfg = ScraperConfig(impersonate="chrome")
return await scrape_async("https://example.com", config=cfg)
asyncio.run(main())Concurrent scraping
Scraping is I/O-bound, so running several scrapes at once parallelizes the
network waits. scrape_many runs one scraper over many inputs; scrape_all
runs a mix of scrapers together. Both preserve input order.
from pyscrappy import scrape_many, scrape_all, AmazonScraper, WikipediaScraper, NewsScraper
# One scraper, many queries, concurrently:
results = scrape_many(AmazonScraper, [{"query": "laptop"}, {"query": "phone"}])
# Different scrapers at once:
results = scrape_all([
lambda: WikipediaScraper().scrape(query="Python"),
lambda: NewsScraper().scrape(feed_url="https://rss.nytimes.com/services/xml/rss/nyt/World.xml"),
])Sitemap crawling
Pagination follows next-page links; a sitemap enumerates a whole site's URLs
directly. GenericScraper can read /sitemap.xml (discovered from robots.txt
Sitemap: directives, or the conventional path), follow a <sitemapindex> into
its child sitemaps, and scrape every listed page.
from pyscrappy import GenericScraper
with GenericScraper() as gs:
# Just enumerate the URLs:
urls = gs.sitemap_urls("https://example.com") # -> list[str]
# Or fetch + extract each, concurrently, into one result:
result = gs.scrape_sitemap("https://example.com", max_urls=100)
print(len(result.data), "pages scraped")Handles <urlset> leaves and <sitemapindex> files (recursing one level),
gzip-compressed sitemaps (.xml.gz), and de-duplicates URLs. Fetches go through
the usual rate-limiting, caching, proxy, and stealth machinery, and the fan-out
reuses scrape_all. max_urls caps the crawl (a sitemap can list tens of
thousands of URLs, so it's required for scrape_sitemap).
Response caching
Set cache_ttl to a positive number of seconds to cache successful GET
responses. Repeated requests for the same URL (and query params) within the TTL
are served from cache, skipping both the network and the rate limiter. Caching
is disabled by default (cache_ttl=0).
from pyscrappy import WikipediaScraper
from pyscrappy import ScraperConfig
config = ScraperConfig(cache_ttl=300) # cache for 5 minutes
with WikipediaScraper(config) as ws:
ws.scrape(query="Python") # fetched over the network
ws.scrape(query="Python") # served from cacheThe cache is in memory and shared across scraper instances in the same process
(so it also speeds up repeated calls through the MCP server), and is cleared
when the process exits. Call HttpClient.clear_cache() to empty it manually.
It is LRU-bounded: at most cache_max_size live entries (default 512),
with the least-recently-used entry evicted once the cap is reached. So a
long-running process (e.g. the MCP server) that fetches many distinct URLs stays
bounded rather than growing until restart. Raise or lower the cap as needed:
config = ScraperConfig(cache_ttl=300, cache_max_size=2000)Persistent (on-disk) cache. Set cache_dir to also persist responses to
disk, so cache hits survive across process restarts and separate runs — useful
for re-running a scrape or a CLI job without re-fetching. The in-memory cache
still fronts it for speed; a disk hit is promoted back into memory.
config = ScraperConfig(cache_ttl=3600, cache_dir="~/.cache/pyscrappy")The on-disk cache is bounded too: each write prunes expired entries and trims the
oldest past cache_dir_max_size (default 512), so a cache_dir doesn't grow
one file per distinct URL forever.
clear_cache() empties the in-memory cache; the on-disk cache persists by
design — delete its cache_dir to clear it.
Observability hooks
For long crawls, pass lightweight callbacks to watch requests live (progress bars, metrics) without turning on logging:
config = ScraperConfig(
on_request=lambda url: print("GET", url), # before a network fetch
on_retry=lambda url, attempt, delay, err: print("retry", attempt, url),
on_cache_hit=lambda url: print("cached", url), # served from cache
)on_request(url)fires once before a URL is fetched (not on a cache hit).on_retry(url, attempt, delay, error)fires before each backoff sleep.on_cache_hit(url)fires when a request is served from cache.
All three are best-effort: a callback that raises is logged at debug and never breaks the scrape. They fire on both the sync and async paths.
Dependencies
Required: httpx, beautifulsoup4, lxml
Optional: playwright (JS rendering), pandas (DataFrames), fastmcp
(MCP server, Python 3.10+)
License
Contributing
All contributions welcome. See Issues.
This package is for educational and research purposes.
Available Tools
24 toolsconvert_currencyConvert CurrencyA
Fetch live exchange rates and convert an amount from one currency to others.
Returns a dict with the base currency, the amount converted, and a mapping of each target currency code to its converted value and unit exchange rate (e.g. {"base": "USD", "amount": 100, "results": {"EUR": {"rate": 0.92, "value": 92.0}}}).
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Target currency codes as a comma-separated string; omit or leave empty to return rates for all available currencies. Example: "EUR,GBP". Default: None (all rates). | |
| base | No | Base currency code as a 3-letter ISO 4217 string. Example: "USD". Default: "USD". | USD |
| amount | No | Amount of the base currency to convert, as a number (int or float). Example: 100. Default: 1. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It reveals that rates are 'live' and provides the exact return structure, which is useful. However, it does not disclose potential pitfalls such as invalid currency codes, handling of unknown target codes, rate staleness, or external API dependencies. For an unannotated tool, this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: the first states the core action concisely; the second handily illustrates the return structure with a concrete example. No redundant wording, front-loaded purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter, all-optional tool, the description plus rich schema fully cover how to call it and what to expect. The return example doubles as a de facto output schema. The only missing piece is error behavior, but that is not strictly required for a straightforward conversion tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with detailed descriptions for 'to', 'base', and 'amount', including examples and defaults. The description adds no extra parameter meaning beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Fetch live exchange rates and convert') and resource (currencies), clearly distinguishing it from all sibling tools, which are spam/scrapers/search tools. No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no direct sibling for currency conversion, so no alternative routing is needed. The description clearly implies this is the tool for any currency-conversion need. It lacks explicit 'when not to use' guidance, but the scope is obvious from the unique functionality.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
define_wordDefine WordA
Look up an English word and return its dictionary entry: definitions, part(s) of speech, and example sentences.
Fetches from an online dictionary data source, so a network connection is required. Read-only with no side effects. If the word is not found (misspelled or not in the dictionary), returns an empty result or a not-found response rather than raising.
Returns a structured entry for the word, typically containing the word itself, one or more part-of-speech groupings, and for each a list of definitions with optional example sentences.
| Name | Required | Description | Default |
|---|---|---|---|
| word | Yes | The English word to define, as a string. Example: "serendipity". No default (required). Single words only; not phrases or non-English terms. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It explicitly notes that the tool is read-only with no side effects, requires a network connection, and handles not-found cases by returning an empty or not-found response instead of raising an error. It also describes the structure of the return value. This is substantial coverage for a simple lookup tool, though it could mention rate limits or authentication if applicable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short paragraphs with no fluff. The purpose is front-loaded in the first sentence, and the second paragraph adds essential behavioral details (network, error handling, return structure). Every sentence earns its place, and the length is appropriate for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and a structured output schema. The description covers the purpose, behavioral notes, and a summary of the return value. Given that an output schema exists, the description does not need to detail every field. It is complete enough for an agent to understand what the tool does and how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage and already fully describes the 'word' parameter, including its type, example, required status, and constraints (single words only). The tool description does not add any additional meaning beyond what the schema already provides, so it meets the baseline of 3 for a high-coverage schema. No further elaboration is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (look up an English word) and the resource (dictionary entry), listing specific content: definitions, parts of speech, and example sentences. It is distinct from the sibling tools, which are all web scraping or search tools, so there is no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives, though the sibling list makes it obvious it is for dictionary lookups. It does provide a constraint in the schema parameter (single words only, not phrases or non-English terms), but there is no explicit routing guidance such as 'use this when you need a definition rather than scraping a page.' The intended use is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cryptoGet CryptoA
Fetch live cryptocurrency market data and return a list of coin records, each with fields: id, symbol, name, current price (in vs_currency), market cap, and 24h price change (percent).
Fetches from a live crypto market data API over the network, so results reflect current prices and require internet access; no local state is read or written. When query is omitted, returns the top coins ranked by market cap. If no coins match the query, returns an empty list rather than raising.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | String of comma-separated coin ids. Example: "bitcoin, ethereum". Default None (returns top coins by market cap). | |
| max_results | No | Integer maximum number of coins to return. Example: 10. Default 20. | |
| vs_currency | No | String fiat or quote currency code for prices, lowercase. Example: "usd". Allowed: any currency supported by the data source, e.g. "usd", "eur", "gbp", "jpy". Default "usd". | usd |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the tool makes a network call, requires internet, reads no local state, returns top coins when query is omitted, and returns an empty list rather than throwing on no matches. This meaningfully sets expectations, though it could also mention failure behavior on network errors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused paragraphs with no filler. The return shape is front-loaded, the network/read-only behavior is stated next, and edge-case behavior is briefly covered. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only lookup tool with three well-documented optional parameters and an output schema, the description is complete. It covers result shape, defaults, edge cases, and operational context without needing to restate schema details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema coverage is 100%, so the baseline is 3. The description adds value by clarifying that query is a comma-separated id list, that omission means top coins by market cap, and that prices are expressed in vs_currency. It complements rather than restates the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('fetch') with a clear resource ('live cryptocurrency market data') and precisely enumerates the returned coin fields. It is immediately distinct from sibling tools like convert_currency or search_news, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes clear context: use this when you need current cryptocurrency prices, top coins by market cap, or coin lookups by id. It does not explicitly name alternatives or state when not to use it, but the intended domain is unmistakable from the first sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_weatherGet WeatherA
Fetch the current weather conditions for a named place and return a dict with keys: temperature (number, degrees Celsius), humidity (number, percent), wind_speed (number, wind speed), condition (str, e.g. "Clear", "Rain"), and location (str, the resolved place name).
Makes a live network call to an external weather provider on each invocation, so results reflect real-time conditions and require internet access. If the location cannot be resolved or the provider returns no match, the tool returns an empty result (or an error field) rather than raising.
| Name | Required | Description | Default |
|---|---|---|---|
| location | Yes | String naming the place to look up; a city name, optionally with a region or country to disambiguate. Example: "Tokyo, Japan". No default; this parameter is required. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does well by stating that each call makes a live network request, depends on an external provider, needs internet access, and returns an empty/error result rather than raising. It could also mention latency or rate limits, but the core runtime behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: first sentence states the action and return contract, second sentence covers runtime behavior and error handling. Every sentence earns its place, with no repetition of the title or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description is nearly complete: it covers invocation, return format, live-network behavior, internet dependency, and failure handling. The only notable gap is explicit guidance on when to prefer this over sibling tools, which is already captured under usage guidelines.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the parameter description already explains the expected string format, optional region/country disambiguation, and that it is required. The tool description adds only general context about 'named place' and the resolved place name in the response, so it does not significantly extend the schema's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Fetch the current weather conditions for a named place') and goes further to enumerate the exact return keys and their types. This clearly distinguishes it from the sibling search and scraper tools, even without an explicit comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the core use case clear: retrieve current weather for a named location. However, it provides no explicit guidance about when to choose this tool over siblings or when not to use it, leaving tool-selection routing mostly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_available_scrapersList Available ScrapersA
List every scraper registered with this server and return their names for use with scrape_with.
Reads the server's in-process scraper registry, which includes built-in scrapers plus any installed third-party pyscrappy-* plugin packages that self-register on import. No network or browser access is performed and no state is changed. If no scrapers are registered, returns an empty list.
Returns: list[str]: Scraper name identifiers (for example ["amazon", "flipkart", "youtube"]), each usable as the scraper argument to scrape_with. Empty list when none are registered.
Usage Guidelines: Call this first to discover valid scraper names, then pass a returned name to scrape_with; use it to confirm a plugin registered correctly after installing a pyscrappy-* package.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that no network or browser access is performed, no state is changed, and that the tool reads an in-process registry. It also covers the empty-list edge case and plugin self-registration behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a purpose statement, behavioral details, return format, and usage guidelines. It is slightly redundant (e.g., empty list behavior appears twice, and the connection to scrape_with is stated more than once), but remains focused and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (zero parameters, output schema available), the description is fully complete. It explains what values are returned, how they should be used, what happens when no scrapers exist, and what side effects (none) to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics to add beyond the schema. The description still confirms the absence of parameters by describing the tool as a simple registry listing, and it clearly explains what the returned values mean.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('List every scraper registered with this server') and states the results are for use with scrape_with. This clearly distinguishes it from sibling scraping tools like scrape_url and scrape_with, which actually perform scrapes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs the agent to call this first to discover valid scraper names and then pass a returned name to scrape_with. It also gives a concrete verification use case (confirming a plugin registered), leaving no doubt about when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_movieLookup MovieA
Look up movie and TV data from IMDB via the OMDb API and return a JSON-serializable dict; a title search returns {"results": [...]} with each item holding title, year, imdb_id, and type, while an IMDB-id lookup returns a single record with full details (plot, ratings, cast, runtime, genre).
Reads over the network from the OMDb HTTP API; no browser is needed and nothing is written or cached. Requires a free OMDb API key in the OMDB_API_KEY environment variable (get one at https://www.omdbapi.com/apikey.aspx); if it is missing the tool returns {"error": ...} explaining how to set it. When the query matches nothing, it returns an empty results list rather than raising.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | String. A title to search for (e.g. "inception"), or an IMDB id starting with "tt" for a direct single-record lookup (e.g. "tt1375666"). Required, no default. | |
| max_pages | No | Integer. Number of search-result pages to fetch at 10 results per page; only applies to title searches and is ignored for IMDB-id lookups (e.g. 3). Default 1. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and meets it well: it discloses network reads, that nothing is written or cached, the external API key requirement, the error format for a missing key, and the empty-list behavior on no matches. This is substantial transparency for a read-only external API call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed paragraphs, with the core lookup behavior and response shape front-loaded, followed by essential operational details (network, key, error behavior). No filler or redundant restating of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a network-backed lookup tool: it covers both query modes, return shapes, API key setup, failure behavior for missing key and no matches, and explicitly notes it is non-destructive. The presence of an output schema further covers return details, and the description adds what schema/annotations cannot.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description reinforces that 'query' can be a title or an 'tt' IMDB id and that the response shape differs by input, but the schema already documents these details. The description adds no meaning beyond the schema for the parameters themselves.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Look up') and resource ('movie and TV data from IMDB via the OMDb API'), and clarifies the return type and shape. It stands apart from sibling scraping/search tools by describing an HTTP API lookup rather than browser scraping, so an agent can distinguish it without inspecting others.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use the tool and how to choose between title search and IMDB-id lookup, including behavior on no matches and missing API key. It does not explicitly name alternatives or state when not to use it, but the 'no browser is needed' line implies a contrast with scraping siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_newsScrape NewsA
Fetch news articles from an RSS/Atom feed, a news site (feed auto-discovered), or a single article, and return a list of article dicts (typically: title, url, published date, author, summary, and full text where available).
Provide exactly one of feed_url, site_url, or article_url. Fetches live content over the network at call time; results are not cached. article_url returns one article; feed_url and site_url return up to max_articles. Returns an empty list if the feed/site yields no articles or if a feed cannot be discovered or parsed.
| Name | Required | Description | Default |
|---|---|---|---|
| feed_url | No | String, direct URL to an RSS/Atom feed. Example: "https://example.com/rss.xml". Default: None. | |
| site_url | No | String, news site homepage URL whose feed is auto-discovered. Example: "https://example.com". Default: None. | |
| article_url | No | String, single article URL to extract full text from. Example: "https://example.com/2026/news-story". Default: None. | |
| max_articles | No | Integer, max articles to return for feed_url or site_url; ignored for article_url. Example: 20. Default: 50. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that fetches are live, results are not cached, and explains return behavior (one article vs. up to max_articles, empty list on failure). It also mentions network usage and potential non-return scenarios. This is comprehensive for a scraping tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two well-structured sentences. The first sentence states the primary purpose and output format, front-loading the core action. The second sentence provides usage constraints and behavioral details. There is no redundancy or fluff; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all essential aspects for correct invocation: input options, exclusivity constraint, return behavior, and failure handling. It accounts for network usage and caching. The presence of an output schema covers return format specifics. No critical information is missing for an agent to use this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds crucial semantic value beyond the schema: it enforces mutual exclusivity of feed_url, site_url, and article_url, and clarifies that max_articles only applies to feed_url and site_url. These are not apparent from the schema alone, so the description elevates the score to 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches news articles from RSS/Atom feeds, news sites, or single articles, and returns a list of article dicts. It specifies the verb, resource, and output format, and differentiates from generic siblings like scrape_url by focusing on news content and feed auto-discovery.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage instructions: 'Provide exactly one of feed_url, site_url, or article_url.' It explains the behavior of each mode and the max_articles limit. It doesn't explicitly mention when not to use this tool versus alternatives, but the purpose is unambiguous enough that an agent can select it for news-related scraping tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_stockScrape StockA
Fetch stock market data from Yahoo Finance and return it as a dict.
The returned shape depends on mode:
"quote": {"symbol", "currency", "exchange", "price", "previous_close", "volume", "day_high", "day_low", "fifty_two_week_high", "fifty_two_week_low"}.
"history": {"symbol", "period", "rows": [{"date", "open", "high", "low", "close", "volume"}, ...]}.
"profile": {"symbol", "name", "currency", "exchange", "market", "timezone", "instrument_type"}.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | String selecting what to fetch; one of "quote", "history", "profile". Example: "quote". Default: "quote". | quote |
| period | No | String history window, used only when mode="history" and ignored otherwise; one of "1d", "5d", "1mo", "3mo", "6mo", "1y", "2y", "5y", "10y", "ytd", "max". Example: "1y". Default: "1mo". | 1mo |
| symbol | Yes | Ticker symbol as a string. Example: "AAPL". No default (required). | |
| interval | No | Candle size for history bars, used only when mode="history"; one of "1d", "1wk", "1mo". Example: "1wk". Default: "1d". | 1d |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the external data source, the dict return type, and how the output varies by mode. However, it does not mention whether the operation is read-only, how errors or invalid symbols are handled, or any rate-limit/network caveats typical of a scraping tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, well-structured, and front-loaded with the core purpose. The mode-by-mode output list is dense but easy to scan, and every sentence contributes useful information without drifting into fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's modest complexity, the description covers the essential call contract: parameters, mode behavior, and return structure. It lacks only explicit guidance on edge cases like invalid symbols, repeated calls, and rate limits, but the output schema and full parameter coverage make the definition sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by tying each mode to a concrete output shape and clarifying that period and interval are history-specific, which complements the schema's parameter descriptions rather than merely repeating them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch stock market data from Yahoo Finance and return it as a dict.' This clearly distinguishes the tool from crypto, weather, and search siblings. The mode-dependent return shapes add precision, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly communicates it is for stock market data from Yahoo Finance, which implies when to use it. However, it does not explicitly name alternatives or state 'use get_crypto for crypto' or 'do not use for non-stock assets', so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_urlScrape UrlA
Scrape any HTTP(S) URL and return a ScrapeToolResult whose data holds one object per page containing extracted text (with word_count), links, images, tables, and page metadata.
Fetches the page over the network and parses the HTML; no data is stored or mutated. By default it makes a plain static HTTP request, so pages built client-side with JavaScript come back nearly empty. When that is detected, the returned errors list gets a hint to retry with render_js=true; set render_js=true to render with a headless browser instead (requires the pyscrappy[browser] extra). On empty or failed results, data is [], count is 0, and errors describes the problem rather than raising.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | String, the page URL to scrape including scheme, e.g. "https://example.com/products". Required, no default. | |
| max_pages | No | Integer, follow "next"-style pagination up to this many pages, e.g. 3. Default 1 (scrape only the given URL). | |
| render_js | No | Boolean, render JavaScript with a headless browser backend, e.g. True. Default False; allowed values True or False, and True needs the pyscrappy[browser] extra installed. | |
| selectors | No | Optional dict mapping output field name to CSS selector to extract specific values into each data item, e.g. {"title": "h1", "price": ".amount"}. Default None (returns only the standard text/links/images/tables/metadata). |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does it thoroughly: network fetch, no storage/mutation, static vs JS rendering behavior, the error hint, and the exact empty-result shape (data [], count 0, errors) instead of raising. This is exemplary disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: first sentence states purpose and output shape, and the rest covers behavior, failure mode, and the JS-rendering option. Every sentence earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the input schema is fully documented, the description supplies the missing behavioral context: network effects, JS caveat, dependency requirement, and failure semantics. Nothing needed to invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents url, max_pages, render_js, and selectors with defaults and examples. The tool description reinforces the render_js behavior but adds little meaning beyond the schema, matching the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Scrape any HTTP(S) URL and return...' which names a specific verb, a clear resource, and the general-purpose scope. This distinguishes it from the specialized scrapers in the sibling list without requiring a tool-by-tool comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete guidance for the main decision point: use the default static request unless the page is client-side rendered, in which case set render_js=true. It does not explicitly call out sibling tools as alternatives or state when not to use this tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_wikipediaScrape WikipediaA
Fetch a Wikipedia article by title or search term and return its text content.
Makes a live network request to Wikipedia, resolving the query to the best-matching article and extracting its body. The shape of the returned text depends on mode: "full" returns the entire article as one string; "paragraphs" returns the article split into a list of paragraph strings; "headers" returns a list of the article's section heading strings (its table of contents). If no article matches the query, an empty result is returned (empty string for "full", empty list for "paragraphs" or "headers").
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | String, one of "full", "paragraphs", or "headers". Selects the return shape as described above. Example: "paragraphs". No default (required). | full |
| query | Yes | String. Article title or search term. Example: "Model Context Protocol". No default (required). |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses that a live network request is made, that the query is resolved to the best-matching article, and that empty results are returned when no match exists. It could add error-handling or rate-limit context, but the core behavioral traits are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded: purpose first, then network behavior, then mode semantics, then edge-case behavior. Every sentence adds necessary information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the parameter schema covers both parameters, the description is largely complete. It explains the mode-dependent return shapes and empty-result behavior. It could mention network failure or default-mode behavior, but overall it gives an agent enough to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds useful meaning by explaining each mode's return shape and the no-match behavior. However, it does not clarify the default for `mode`, and the schema itself is internally inconsistent (default 'full' vs 'No default (required)'), so the description misses a chance to resolve that ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch a Wikipedia article by title or search term and return its text content.' This clearly distinguishes it from sibling scrapers like scrape_url or scrape_stock, and the mode details further specify what 'text content' means.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the intended use clear: retrieving Wikipedia article content by title or search term. It does not explicitly name alternatives or state when not to use it, but the Wikipedia-specific scope is unambiguous enough that an agent can select it correctly among the listed siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_withScrape WithA
Run any registered scraper (built-in or plugin) by name and return that scraper's raw scrape() output.
This is the generic dispatch entry point for scrapers that lack a dedicated tool, notably third-party plugins. It looks up the scraper in the registry, calls its scrape() method with the given args, and returns whatever that scraper returns (typically a dict or list of records; exact shape is scraper-specific). Side effects and requirements (network requests, browser/headless rendering, auth) depend entirely on the target scraper. If name is not a registered scraper, it raises an error rather than returning empty; if the scraper runs but finds nothing, it returns that scraper's empty result (e.g. an empty list).
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Dict of keyword arguments forwarded to the named scraper's scrape() method; required keys depend on that scraper. Example: {"query": "Alan Turing", "lang": "en"}. No default (required). | |
| name | Yes | String, the scraper's registered name from list_available_scrapers. Example: "wikipedia". No default (required). |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does an excellent job. It discloses that side effects depend on the target scraper, mentions network/browser/auth possibilities, distinguishes 'not found' errors from empty results, and states that returns are scraper-specific.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise yet complete, with the core purpose front-loaded and supporting details about lookup, return shape, side effects, and error behavior presented in a logical order. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a generic dispatch tool with no annotations, the description covers all critical operational aspects: how it dispatches, what it returns, what side effects may occur, and how errors differ from empty results. The output schema handles return-value details, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage of both parameters, including examples and forwarding semantics. The description reinforces that args are forwarded to the scraper, but it does not add meaning beyond what the schema already contains, so the high-coverage baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that this tool runs any registered scraper by name and returns its raw output, using a specific verb and resource. It also distinguishes itself as the generic dispatch entry point for scrapers without a dedicated tool, setting it apart from sibling tools like scrape_wikipedia and scrape_stock.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use this for scrapers that lack a dedicated tool, especially third-party plugins. It does not name the dedicated alternatives individually, but the sibling list provides those options and the intended scoping is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_zomatoScrape ZomatoA
Search Zomato for restaurants in a city and return a list of restaurant records.
Scrapes Zomato's public restaurant listings over the network for the given city, optionally filtered by a cuisine or name term. Each result is a dict with fields such as name, cuisine, rating, price_for_two, address, and url; the exact keys depend on what Zomato exposes for each listing. Returns a list of these dicts ordered as Zomato ranks them, capped at max_results. Returns an empty list if the city is unknown or no restaurants match the query. Requires outbound network access; results reflect live Zomato data at call time and may vary between calls.
| Name | Required | Description | Default |
|---|---|---|---|
| city | Yes | String city name to search within. Example: "Bangalore". Required, no default. | |
| query | No | Optional string cuisine or restaurant search term to filter results. Example: "biryani". Defaults to None (returns all restaurants for the city). | |
| max_results | No | Integer maximum number of restaurants to return. Example: 20. Defaults to 50. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full transparency burden and does well: it discloses outbound network dependency, live-data variability between calls, result ordering, capping by max_results, and the empty-list behavior. It does not mention potential rate limits or blocking, but the provided behavioral caveats are substantial and useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded, with the core purpose in the first sentence followed by behavioral details. There is mild redundancy between 'Search Zomato...' and 'Scrapes Zomato's public restaurant listings...', but overall every sentence earns its place and no extraneous content appears.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and the tool's straightforward parameter set, the description is complete enough: it covers return shape, ordering, capping, unknown-city behavior, no-match behavior, and network requirements. Minor gaps like pagination or error modes are not critical for this tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minimal meaning beyond the schema: it restates the query filter as 'cuisine or name term' and says results are 'capped at max_results', but most parameter meaning (examples, defaults, types) already lives in the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource: 'Search Zomato for restaurants in a city and return a list of restaurant records.' This clearly differentiates the tool from siblings like scrape_wikipedia, scrape_stock, and search_ubereats by naming the exact platform, data type, and action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this when you need Zomato restaurant listings for a city, optionally filtered by cuisine/name. However, it never explicitly contrasts this tool with alternatives like scrape_url or search_ubereats, nor states when not to use it. The usage context is clear but the decision guidance is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_amazonSearch AmazonA
Scrape Amazon search results for a query and return a list of matching products, each with its title, price, rating, and image URL.
This performs a live network scrape of Amazon's public search results pages (no login, no API key). It has no side effects beyond outgoing HTTP requests. Results reflect Amazon's current listings and may vary by region, availability, and anti-bot throttling. Returns a list of dicts, one per product, each shaped as {"title": str, "price": str, "rating": str, "image": str}; fields that Amazon omits for a listing come back as empty strings or None. Returns an empty list when the query yields no products or when scraping is blocked.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Product search phrase, as a string. Example: "wireless headphones". No default (required). | |
| max_pages | No | Number of result pages to scrape, as an integer; higher values return more products but take longer and raise the chance of throttling. Example: 3. Default: 1. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden of behavioral disclosure. It explicitly states this is a live network scrape, requires no login or API key, has no side effects beyond HTTP requests, may be affected by region, availability, and anti-bot throttling, and details the exact return shape including empty string/None treatment and empty list on failure. This is comprehensive and exceeds typical descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, then adds behavioral context and return format details in a logical, dense second paragraph. Every sentence contributes value—no filler, no repetition of schema fields. It is suitably concise for the complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and an output schema (which already captures return structure), the description covers all essential operational context: authentication, side effects, variability, throttling, empty results, and failure modes. It leaves no critical gap for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already fully documents both parameters. The description adds no additional parameter-level meaning; it only says 'for a query' in passing. Baseline of 3 is appropriate because the schema handles parameter documentation, and the description does not need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Scrape Amazon search results for a query' and clearly states the return payload (list of products with title, price, rating, image URL). This distinguishes it from siblings like search_newegg or search_books by explicitly targeting Amazon, and from generic scrape tools by specifying the query-driven search use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context that this tool is for Amazon search listings, including live network behavior and no authentication. It does not explicitly name alternative tools or exclusions, but the specificity of 'Amazon search results' makes the usage context obvious. It falls short of a 5 because it never explicitly says 'use this instead of X' or 'do not use when Y'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_booksSearch BooksA
Search books by title, author, or free text via the Open Library search API and return a list of matching book records.
Queries Open Library over the network (no browser or authentication required). Returns a list of dicts, each typically containing: title (str), author_names (list of str), first_publish_year (int or None), edition_count (int), and the Open Library work key (str, e.g. "/works/OL45804W"). Fields missing upstream are omitted or None. Returns an empty list when the query matches nothing. Read-only: no local files or state are modified.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Title, author, or free-text search string. Type: string. Example: "the hobbit tolkien". Required, no default. | |
| max_results | No | Maximum number of books to return. Type: integer. Example: 10. Default: 20. Allowed: any positive integer. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral disclosure burden. It states that the tool queries the network, requires no browser/auth, returns empty lists for no matches, omits missing fields or uses None, and is read-only with no local file/state modification. This is unusually complete for a search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action, then provides targeted details about network behavior, return structure, empty results, and read-only semantics. Every sentence adds value for an agent deciding whether and how to call the tool; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter search tool with an output schema, the description covers the important invocation context: what it searches, where it queries, what it returns, how missing fields are handled, the empty-list case, and that it is non-destructive. Nothing essential is missing for correct selection and use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description repeats that the query can be by title, author, or free text, but adds no parameter-specific meaning beyond what the schema already documents for 'query' and 'max_results'. It neither compensates for a gap nor introduces new confusion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Search books by title, author, or free text via the Open Library search API and return a list of matching book records.' It names the exact upstream source (Open Library) and the book domain, which clearly differentiates it from sibling tools like search_images, search_youtube, or scrape_url.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it is a network-based search for book records, requires no browser or authentication, and is read-only. It does not explicitly name alternative tools or state when not to use it, but the domain specificity and the large set of clearly unrelated siblings make the intended usage obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_githubSearch GithubA
Search GitHub for public repositories and return a list of repository records.
Queries the GitHub search API over the network and returns a list of dicts, each with: name (str), owner (str), stars (int), description (str), and language (str). Results are ordered per the sort argument. Returns an empty list when no repository matches the query. Requires network access; may be subject to GitHub API rate limits.
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | String ordering for results; one of "best-match", "stars", "forks", or "updated". Example: "stars". Default "best-match". | best-match |
| query | Yes | String search expression using GitHub search syntax, including qualifiers like "language:" or "stars:". Example: "web scraping language:python". No default (required). | |
| max_results | No | Integer maximum number of repositories to return. Example: 10. Default 20. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does it well: it discloses network access, GitHub API rate-limit susceptibility, return format with field names and types, ordering by the sort argument, and empty-list behavior. It stops short of covering error handling or authentication requirements, but it is substantially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the core purpose, and the second adds behavioral and result details without filler. Every sentence earns its place, and there is no redundant or vague wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a public search tool with fully documented parameters, the description covers the main call semantics, result shape, no-match case, and network/rate-limit caveats. It does not enumerate alternatives or error behavior, but those are secondary for a no-annotation search endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already fully documents query, sort, and max_results with defaults and examples. The description adds only that results are ordered per the sort argument, which mostly restates the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Search GitHub for public repositories and return a list of repository records.' This clearly distinguishes the tool from sibling search tools like search_images or search_youtube by naming the GitHub domain and describing the repository record output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The GitHub-specific scope makes the intended use clear, and the description explicitly notes it queries the GitHub search API, so an agent can infer when to choose this tool. It does not explicitly name alternatives or exclusion conditions, but no misleading usage signals are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_hackernewsSearch HackernewsA
Search Hacker News stories and return a list of matching story dicts, each with title, url, points, author (username), and num_comments (comment count).
Queries the public Hacker News search index (Algolia HN API) over the network; makes no local changes. Returns an empty list when nothing matches or the query is empty.
| Name | Required | Description | Default |
|---|---|---|---|
| by | No | Result ordering; type: string; one of "relevance" or "date" ("date" sorts most recent first); example: "date"; default: "relevance". | relevance |
| tags | No | Algolia HN tag filter; type: string; common values "story", "comment", "show_hn", "ask_hn", "poll", "job"; example: "show_hn"; default: "story". | story |
| query | Yes | Search terms to match against story titles and text; type: string; example: "rust async runtime"; no default (required). | |
| max_results | No | Maximum number of stories to return; type: integer; example: 10; default: 20. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a good job: it discloses that the tool is a read-only network query ('makes no local changes', 'queries the public Hacker News search index'), and it specifies the empty-list behavior for empty queries and no matches. It omits potential rate limits or error behavior, but those are not critical for a simple search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences with no fluff. It front-loads the purpose and return shape, then gives network/read-only behavior, then the empty-result edge case. Every sentence contributes distinct, useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with full parameter schema coverage and an output schema, the description is complete enough: it explains result field names, clarifies network access and safety, and defines edge behavior. An agent can confidently select and invoke this tool without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All four parameters are already fully documented in the input schema with descriptions, defaults, and examples (100% schema coverage), so the baseline is 3. The description adds no additional meaning for by, tags, query, or max_results beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search Hacker News stories and return a list of matching story dicts'. This clearly distinguishes it from sibling search tools like search_images, search_github, and search_books by naming the exact content domain and result format.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes clear usage context by stating this searches the public Hacker News index over the network. It does not explicitly name alternatives or exclusions, but the tool's resource is unique among siblings, so a 'when to use' statement is mostly implicit in the name and first sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_ikeaSearch IkeaA
Search IKEA's online catalog for furniture and home products, returning a list of product dicts with fields name, type, price, and rating.
Scrapes the IKEA store website for the given country at call time, so results require network access and reflect that store's live listings. Prices, availability, and currency are per-country and per-language. Returns an empty list if the query matches no products.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | String language code for that store's listings; must be a language the chosen country's store supports, e.g. "en" for "us"/"gb" or "de" for "de". Example "de". Default "en". | en |
| query | Yes | String search term for the product name or type, e.g. "desk" or "bookshelf". Required, no default. | |
| country | No | String two-letter IKEA store country code that sets pricing and availability; allowed values are IKEA market codes such as "us", "gb", "de", "fr", "se". Example "gb". Default "us". | us |
| max_results | No | Integer cap on the number of products returned; e.g. 10. Default 24. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does it well: it discloses live scraping at call time, network dependency, per-country/per-language pricing and availability, and the empty-list behavior for no matches. This goes well beyond what the input schema or tool name convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences: purpose and return shape, live-scraping behavior with country/language context, and empty-result behavior. Every sentence adds distinct value, and the most important identifying information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no annotations, the description covers the key operational caveats: live network scraping, per-country pricing, and empty results. An output schema exists and the return fields are named, so return-value details are adequately covered. Minor gaps remain around failure modes (e.g., network errors or blocked scraping), preventing a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents query, country, lang, and max_results in detail. The description adds useful context about country-specific live pricing, but it does not need to repeat parameter-level details; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search IKEA's online catalog for furniture and home products.' It also names the return shape (list of product dicts with fields), which clearly distinguishes it from sibling tools like search_amazon or scrape_url.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes when this tool is appropriate: for IKEA-specific product searches with live store data. It does not explicitly name alternatives or exclusion conditions, but the context is clear enough that an agent can route to it versus the other catalog and search tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_imagesSearch ImagesA
Search the web for images and return a list of result objects with image URLs and metadata.
Each result is a dict with keys: "url" (direct link to the image), "thumbnail" (small preview), "title" (caption or alt text), "source_page" (page the image was found on), "width", and "height" (pixels). Every engine returns this same key set; fields a given engine can't provide are empty ("" for text, null for width/height — e.g. the Google path fills only "url", "title", and "source_page"). Results are returned in the engine's relevance order.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search terms as a string, e.g. "golden gate bridge". Required, no default. | |
| engine | No | String naming the search engine, one of "bing" or "google", e.g. "google". Defaults to "bing". | bing |
| max_images | No | Integer cap on the number of results returned, e.g. 10. Defaults to 20. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly explains the return format (list of dicts with specific keys), per-engine variations (Google fills only url, title, source_page), handling of missing fields (empty strings or null), and result ordering. No contradictions; all behavioral traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficient and well-structured: a one-sentence purpose followed by a focused paragraph on output format. Every sentence adds necessary detail, and the most important information (what the tool does and what it returns) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple engines, varied output fields), the description is complete. It explains the result object keys, engine-specific behavior, missing-field conventions, and ordering, covering everything an agent needs to invoke and interpret results correctly. An output schema is present, so return-value details are not missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description does not add significant meaning beyond the schema for parameters; it mostly restates what the schema already documents (e.g., query is a string, engine defaults to bing, max_images caps results). No additional parameter detail is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Search the web for images') and a clear resource (images), and differentiates from sibling tools by focusing exclusively on image results with URLs and metadata. It is unambiguous and not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly conveys that this tool is for image searches, which implies when to use it (when images are needed) but does not explicitly mention alternatives or when not to use it. It provides clear context but lacks explicit exclusionary guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_linkedin_jobsSearch Linkedin JobsA
Search LinkedIn public job postings and return a list of matched jobs.
Scrapes LinkedIn's public job search results over the network (no login required) and returns a list of dicts, one per posting, typically with keys: title, company, location, url, and posted_date. Returns an empty list if no postings match or the query yields no results. Live web scraping, so results reflect LinkedIn at call time and may vary between runs.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Job title or keywords as a string. Example: "machine learning engineer". No default (required). | |
| location | No | Location filter as a string; city, region, or country. Example: "London" or "United Kingdom". No default (required). | |
| max_pages | No | Number of result pages to scrape, as an integer. Each extra page adds jobs but more scraping time. Example: 2. Default: 1. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses live network scraping, no-login operation, call-time variability, the expected return structure ('list of dicts... typically with keys'), and empty-list behavior. This is far more transparent than most tool descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three focused sentences with the core purpose front-loaded. Every sentence earns its place by adding behavioral or outcome detail without repeating schema content or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, the description still supplies the key context an agent needs: operation type, network dependence, authentication requirement, result format, and empty-result behavior. The output schema is also marked as present, and pagination is handled by the schema. The description is complete enough for correct selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the input schema already documents query, location, and max_pages with examples and defaults. The prose description adds no additional parameter-level meaning, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States an unambiguous action ('Search LinkedIn public job postings') and a concrete outcome ('return a list of matched jobs'). This clearly differentiates from generic scrapers and sibling search tools by naming the specific resource and domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: the tool is for public LinkedIn job postings, works over the network with no login required, and returns live results. It does not explicitly name alternatives or exclusion conditions, but the domain-specific scope leaves little ambiguity about when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_neweggSearch NeweggA
Search Newegg for electronics and computer hardware, returning a list of product dicts each with title, price, product_url, image_url, rating, and item_number.
Live-scrapes Newegg search result pages over the network; requires outbound internet access and returns an empty list if no products match or the page structure cannot be parsed. Read-only, with no side effects beyond the outbound HTTP requests.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Product search terms, given as a string. Example: "graphics card". No default (required). | |
| max_pages | No | Number of result pages to scrape, given as an integer; higher values return more products but take longer. Example: 3. Default 1. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description takes the full burden and succeeds: it discloses that this is a live network scrape, requires outbound internet, is read-only with no side effects beyond HTTP, and returns an empty list on no matches or parse failure. That goes beyond the schema and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences: the first defines purpose and return shape, the second covers operational behavior and edge cases. No filler, redundant detail, or wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 2-parameter tool with full schema coverage and an output schema, the description covers purpose, return fields, network requirement, read-only nature, and empty-list behavior. The only minor gaps are operational details like rate limits or anti-bot behavior, which are not critical for calling the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'query' and 'max_pages' with examples and effects. The tool description adds no additional parameter semantics beyond what the schema provides, earning the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Search'), a specific resource ('Newegg'), and a domain ('electronics and computer hardware'), plus the exact shape of the returned product dicts. This clearly distinguishes it from sibling scrapers/search tools by site and product focus.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when you need product data from Newegg's search results. It does not explicitly state when to choose this over alternatives like search_amazon or scrape_url, nor does it mention exclusions. The name and siblings make the context reasonably inferable, but explicit guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_soundcloudSearch SoundcloudA
Search SoundCloud for tracks and return a list of track dicts, each with keys: title (str), artist (str), plays (int), likes (int), and url (str, the track page URL).
Renders SoundCloud's JavaScript search results with a browser backend (Playwright/Selenium), so it requires the pyscrappy[browser] extra to be installed and launches a headless browser per call. This makes it slower and heavier than the HTTP-based search tools. Results reflect SoundCloud's live public search at call time; no login or API key is used. Returns an empty list if the query matches no tracks.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query string. Example: "lofi beats". No default (required). | |
| max_results | No | Maximum number of tracks to return, as an integer. Example: 10. Default 20. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly discloses that the tool launches a headless browser, requires an extra dependency, is slower/heavier, shows live public SoundCloud results, needs no login/API key, and returns an empty list for no matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and efficiently structured. The purpose and return format are front-loaded, then caveats follow in a second paragraph. Every sentence contributes meaningful information with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no shown output schema, the description is complete: it explains prerequisites, behavior, performance trade-offs authentication requirements, and empty-result behavior. An agent has enough information to decide whether to call it and what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents both parameters with examples and defaults. The description adds little about parameter behavior beyond the schema's existing coverage, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search SoundCloud for tracks') and describes the exact return shape as track dicts with known keys. It also distinguishes itself from HTTP-based search tools by noting its browser-backend rendering, so an agent can tell it apart from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use this tool: it is slower and heavier than HTTP-based search toolsais and requires the pyscrappy[browser] extra, implying it should be chosen only when SoundCloud's JavaScript-rendered results are required. It does not explicitly name alternatives or state hard exclusions, but the guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_ubereatsSearch UbereatsA
Search Uber Eats for restaurants delivering in a given city, returning a ScrapeToolResult envelope whose data is a list of restaurant objects (typically name, eta, delivery fee, and store url).
Fetches live listings from Uber Eats over the network at call time; no API key is required. The data list is capped at max_results and each item's store url is the input for get_ubereats_menu. Alongside data, the envelope carries count, scraper, source_urls, and errors (non-fatal issues, each with a url and message). If the city is unrecognized or no restaurants are found, data is an empty list and count is 0.
| Name | Required | Description | Default |
|---|---|---|---|
| city | Yes | City name to search, as a string. Example: "London". No default (required). | |
| max_results | No | Maximum number of restaurants to return, as an integer. Example: 10. Default 30. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so well. It discloses that the tool makes a live network call, requires no API key, caps results at max_results, returns non-fatal errors in an envelope, and handles unrecognized cities with empty data and count 0. This is rich, accurate behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise paragraphs with the purpose front-loaded in the first sentence. Every sentence adds value: output shape, network behavior, cap, error envelope, and edge cases. There is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers input, output, error handling, edge cases, and the follow-on tool (get_ubereats_menu). It even explains the ScrapeToolResult envelope's fields. For a 2-parameter search tool, this is comprehensive and leaves no significant gap for an agent making the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces that max_results caps the list and mentions edge-case behavior for the city parameter, but it does not materially add beyond the schema's parameter descriptions. No parameter descriptions are missing, but the description's contribution is marginal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Search Uber Eats for restaurants delivering in a given city'. It also describes the return envelope (ScrapeToolResult with restaurant objects), which clearly differentiates it from siblings like scrape_zomato or search_amazon. The mention of get_ubereats_menu as the consumer of the returned URLs further clarifies its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes its usage context clear: it is for discovering Uber Eats restaurants in a city and provides the store URL for a downstream get_ubereats_menu call. It does not explicitly contrast with alternative scraping/search tools or state when not to use it, but the Uber Eats specificity and the downstream tool reference give enough guidance for the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_youtubeSearch YoutubeA
Search YouTube for videos matching a query and return a list of matching videos with their metadata.
Performs a live YouTube search over the network, so results reflect current YouTube data and may vary between calls; it is read-only and has no side effects. Returns a list of video objects, each typically containing: title (str), channel (str), url/link (str) to the video, video_id (str), duration (str), view_count (int), and published/upload date (str). Returns an empty list when the query matches no videos.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search text, as a string. Example: "model context protocol tutorial". No default (required). | |
| max_results | No | Maximum number of videos to return, as an integer. Example: 10. Default 20. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| count | No | |
| errors | No | |
| scraper | No | |
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses live network search, variability between calls, read-only nature with no side effects, and empty-list behavior on no matches. This is strong behavioral context, though it omits potential rate limits or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the main purpose front-loaded. The second sentence packs useful behavioral and return details without filler. The enumeration of return fields is slightly verbose but valuable, keeping the overall structure efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter search tool, the description covers purpose, network behavior, read-only status, return shape, and empty-case handling. With an output schema present, the return-field listing is redundant but not problematic. Missing edge cases like errors or pagination are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents query and max_results with types, examples, and defaults. The description adds no parameter-specific meaning beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Search YouTube for videos matching a query and return a list of matching videos with their metadata.' This clearly distinguishes it from sibling tools like search_images or search_hackernews by naming YouTube directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance, and no mention of alternative tools. The purpose is implied by the name and description, but an agent must infer when to select this over other search tools. The live-network and read-only context hints at behavior, not usage selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.4.6- Changed
convert_currency2 fields changed- changed
Input schema / properties / base / descriptionPrevious value: -"Base currency code as a 3-letter ISO 4217 string. Example: \"USD\". No default (required)."New value: +"Base currency code as a 3-letter ISO 4217 string. Example: \"USD\". Default: \"USD\"." - changed
Input schema / properties / to / descriptionPrevious value: -"Target currency codes as a comma-separated string; omit or leave empty to return rates for all available currencies. Example: \"EUR,GBP\". Default: \"\" (all rates)."New value: +"Target currency codes as a comma-separated string; omit or leave empty to return rates for all available currencies. Example: \"EUR,GBP\". Default: None (all rates)."
- Changed
scrape_stock2 fields changed- added
Input schema / properties / intervalAdded value: +{ + "default": "1d", + "description": "Candle size for history bars, used only when mode=\"history\"; one of \"1d\", \"1wk\", \"1mo\". Example: \"1wk\". Default: \"1d\".", + "type": "string" +} - changed
Input schema / properties / mode / descriptionPrevious value: -"String selecting what to fetch; one of \"quote\", \"history\", \"profile\". Example: \"quote\". No default (required)."New value: +"String selecting what to fetch; one of \"quote\", \"history\", \"profile\". Example: \"quote\". Default: \"quote\"."
- Changed
search_hackernews1 field changed- added
Input schema / properties / tagsAdded value: +{ + "default": "story", + "description": "Algolia HN tag filter; type: string; common values \"story\", \"comment\", \"show_hn\", \"ask_hn\", \"poll\", \"job\"; example: \"show_hn\"; default: \"story\".", + "type": "string" +}
- Changed
search_images1 field changed- changed
Input schema / properties / engine / descriptionPrevious value: -"String naming the search engine, one of \"bing\", \"google\", or \"duckduckgo\", e.g. \"google\". Defaults to \"bing\"."New value: +"String naming the search engine, one of \"bing\" or \"google\", e.g. \"google\". Defaults to \"bing\"."
23 tool updates
- Changed
convert_currency3 fields changed- changed
Input schema / properties / amount / descriptionPrevious value: -"Amount of base currency to convert (default 1)."New value: +"Amount of the base currency to convert, as a number (int or float). Example: 100. Default: 1." - changed
Input schema / properties / base / descriptionPrevious value: -"Base currency code, e.g. \"USD\"."New value: +"Base currency code as a 3-letter ISO 4217 string. Example: \"USD\". No default (required)." - changed
Input schema / properties / to / descriptionPrevious value: -"Comma-separated target codes (e.g. \"EUR,GBP\"). Omit for all rates."New value: +"Target currency codes as a comma-separated string; omit or leave empty to return rates for all available currencies. Example: \"EUR,GBP\". Default: \"\" (all rates)."
- Changed
define_word1 field changed- changed
Input schema / properties / word / descriptionPrevious value: -"The word to define."New value: +"The English word to define, as a string. Example: \"serendipity\". No default (required). Single words only; not phrases or non-English terms."
- Changed
get_crypto3 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"Max coins to return (default 20)."New value: +"Integer maximum number of coins to return. Example: 10. Default 20." - changed
Input schema / properties / query / descriptionPrevious value: -"Comma-separated coins (e.g. \"bitcoin, ethereum\"). Omit for top coins."New value: +"String of comma-separated coin ids. Example: \"bitcoin, ethereum\". Default None (returns top coins by market cap)." - changed
Input schema / properties / vs_currency / descriptionPrevious value: -"Fiat currency for prices, e.g. \"usd\", \"eur\" (default \"usd\")."New value: +"String fiat or quote currency code for prices, lowercase. Example: \"usd\". Allowed: any currency supported by the data source, e.g. \"usd\", \"eur\", \"gbp\", \"jpy\". Default \"usd\"."
- Changed
get_ubereats_menu1 field changed- changed
Input schema / properties / store_url / descriptionPrevious value: -"A store URL from a search_ubereats result's \"url\" field."New value: +"URL string of the Uber Eats store page, taken from a search_ubereats result's \"url\" field. Example: \"https://www.ubereats.com/store/some-restaurant/abc123\". No default (required)."
- Changed
get_weather1 field changed- changed
Input schema / properties / location / descriptionPrevious value: -"Place name, e.g. \"London\" or \"Tokyo, Japan\"."New value: +"String naming the place to look up; a city name, optionally with a region or country to disambiguate. Example: \"Tokyo, Japan\". No default; this parameter is required."
- Changed
lookup_movie2 fields changed- changed
Input schema / properties / max_pages / descriptionPrevious value: -"Pages of search results to fetch, 10 per page (title search)."New value: +"Integer. Number of search-result pages to fetch at 10 results per page; only applies to title searches and is ignored for IMDB-id lookups (e.g. 3). Default 1." - changed
Input schema / properties / query / descriptionPrevious value: -"A title to search for (e.g. \"inception\"), or an IMDB id\n(e.g. \"tt1375666\") for a direct lookup."New value: +"String. A title to search for (e.g. \"inception\"), or an IMDB id starting with \"tt\" for a direct single-record lookup (e.g. \"tt1375666\"). Required, no default."
- Changed
scrape_news4 fields changed- changed
Input schema / properties / article_url / descriptionPrevious value: -"A single article URL to extract full text from."New value: +"String, single article URL to extract full text from. Example: \"https://example.com/2026/news-story\". Default: None." - changed
Input schema / properties / feed_url / descriptionPrevious value: -"Direct URL to an RSS/Atom feed."New value: +"String, direct URL to an RSS/Atom feed. Example: \"https://example.com/rss.xml\". Default: None." - changed
Input schema / properties / max_articles / descriptionPrevious value: -"Max articles to return from a feed (default 50)."New value: +"Integer, max articles to return for feed_url or site_url; ignored for article_url. Example: 20. Default: 50." - changed
Input schema / properties / site_url / descriptionPrevious value: -"News site URL — its feed is auto-discovered."New value: +"String, news site homepage URL whose feed is auto-discovered. Example: \"https://example.com\". Default: None."
- Changed
scrape_stock3 fields changed- changed
Input schema / properties / mode / descriptionPrevious value: -"\"quote\", \"history\", or \"profile\"."New value: +"String selecting what to fetch; one of \"quote\", \"history\", \"profile\". Example: \"quote\". No default (required)." - changed
Input schema / properties / period / descriptionPrevious value: -"History window when mode=\"history\", e.g. \"1mo\", \"1y\"."New value: +"String history window, used only when mode=\"history\" and ignored otherwise; one of \"1d\", \"5d\", \"1mo\", \"3mo\", \"6mo\", \"1y\", \"2y\", \"5y\", \"10y\", \"ytd\", \"max\". Example: \"1y\". Default: \"1mo\"." - changed
Input schema / properties / symbol / descriptionPrevious value: -"Ticker symbol, e.g. \"AAPL\", \"GOOGL\"."New value: +"Ticker symbol as a string. Example: \"AAPL\". No default (required)."
- Changed
scrape_url4 fields changed- changed
Input schema / properties / max_pages / descriptionPrevious value: -"Follow pagination up to this many pages (default 1)."New value: +"Integer, follow \"next\"-style pagination up to this many pages, e.g. 3. Default 1 (scrape only the given URL)." - changed
Input schema / properties / render_js / descriptionPrevious value: -"Render JavaScript with a browser backend (needs pyscrappy[browser])."New value: +"Boolean, render JavaScript with a headless browser backend, e.g. True. Default False; allowed values True or False, and True needs the pyscrappy[browser] extra installed." - changed
Input schema / properties / selectors / descriptionPrevious value: -"Optional CSS selectors, e.g. {\"title\": \"h1\", \"price\": \".amount\"}."New value: +"Optional dict mapping output field name to CSS selector to extract specific values into each data item, e.g. {\"title\": \"h1\", \"price\": \".amount\"}. Default None (returns only the standard text/links/images/tables/metadata)." - changed
Input schema / properties / url / descriptionPrevious value: -"The page to scrape."New value: +"String, the page URL to scrape including scheme, e.g. \"https://example.com/products\". Required, no default."
- Changed
scrape_wikipedia2 fields changed- changed
Input schema / properties / mode / descriptionPrevious value: -"\"full\", \"paragraphs\", or \"headers\"."New value: +"String, one of \"full\", \"paragraphs\", or \"headers\". Selects the return shape as described above. Example: \"paragraphs\". No default (required)." - changed
Input schema / properties / query / descriptionPrevious value: -"Article title or search term, e.g. \"Model Context Protocol\"."New value: +"String. Article title or search term. Example: \"Model Context Protocol\". No default (required)."
- Changed
scrape_with2 fields changed- changed
Input schema / properties / args / descriptionPrevious value: -"Keyword arguments passed to that scraper's `scrape()` method."New value: +"Dict of keyword arguments forwarded to the named scraper's scrape() method; required keys depend on that scraper. Example: {\"query\": \"Alan Turing\", \"lang\": \"en\"}. No default (required)." - changed
Input schema / properties / name / descriptionPrevious value: -"A scraper name from `list_available_scrapers`, e.g. \"wikipedia\"."New value: +"String, the scraper's registered name from list_available_scrapers. Example: \"wikipedia\". No default (required)."
- Changed
scrape_zomato3 fields changed- changed
Input schema / properties / city / descriptionPrevious value: -"City name, e.g. \"Bangalore\"."New value: +"String city name to search within. Example: \"Bangalore\". Required, no default." - changed
Input schema / properties / max_results / descriptionPrevious value: -"Maximum number of restaurants to return (default 50)."New value: +"Integer maximum number of restaurants to return. Example: 20. Defaults to 50." - changed
Input schema / properties / query / descriptionPrevious value: -"Optional cuisine or restaurant search term."New value: +"Optional string cuisine or restaurant search term to filter results. Example: \"biryani\". Defaults to None (returns all restaurants for the city)."
- Changed
search_amazon2 fields changed- changed
Input schema / properties / max_pages / descriptionPrevious value: -"Number of result pages to scrape (default 1)."New value: +"Number of result pages to scrape, as an integer; higher values return more products but take longer and raise the chance of throttling. Example: 3. Default: 1." - changed
Input schema / properties / query / descriptionPrevious value: -"Product search query, e.g. \"wireless headphones\"."New value: +"Product search phrase, as a string. Example: \"wireless headphones\". No default (required)."
- Changed
search_books2 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"Max books to return (default 20)."New value: +"Maximum number of books to return. Type: integer. Example: 10. Default: 20. Allowed: any positive integer." - changed
Input schema / properties / query / descriptionPrevious value: -"Title, author, or free-text search."New value: +"Title, author, or free-text search string. Type: string. Example: \"the hobbit tolkien\". Required, no default."
- Changed
search_github3 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"Max repositories to return (default 20)."New value: +"Integer maximum number of repositories to return. Example: 10. Default 20." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query, e.g. \"web scraping language:python\"."New value: +"String search expression using GitHub search syntax, including qualifiers like \"language:\" or \"stars:\". Example: \"web scraping language:python\". No default (required)." - changed
Input schema / properties / sort / descriptionPrevious value: -"\"best-match\" (default), \"stars\", \"forks\", or \"updated\"."New value: +"String ordering for results; one of \"best-match\", \"stars\", \"forks\", or \"updated\". Example: \"stars\". Default \"best-match\"."
- Changed
search_hackernews3 fields changed- changed
Input schema / properties / by / descriptionPrevious value: -"\"relevance\" (default) or \"date\" (most recent first)."New value: +"Result ordering; type: string; one of \"relevance\" or \"date\" (\"date\" sorts most recent first); example: \"date\"; default: \"relevance\"." - changed
Input schema / properties / max_results / descriptionPrevious value: -"Max stories to return (default 20)."New value: +"Maximum number of stories to return; type: integer; example: 10; default: 20." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query."New value: +"Search terms to match against story titles and text; type: string; example: \"rust async runtime\"; no default (required)."
- Changed
search_ikea4 fields changed- changed
Input schema / properties / country / descriptionPrevious value: -"IKEA store country code, e.g. \"us\", \"gb\", \"de\" (default \"us\")."New value: +"String two-letter IKEA store country code that sets pricing and availability; allowed values are IKEA market codes such as \"us\", \"gb\", \"de\", \"fr\", \"se\". Example \"gb\". Default \"us\"." - changed
Input schema / properties / lang / descriptionPrevious value: -"Language code for that store, e.g. \"en\", \"de\" (default \"en\")."New value: +"String language code for that store's listings; must be a language the chosen country's store supports, e.g. \"en\" for \"us\"/\"gb\" or \"de\" for \"de\". Example \"de\". Default \"en\"." - changed
Input schema / properties / max_results / descriptionPrevious value: -"Maximum number of products to return (default 24)."New value: +"Integer cap on the number of products returned; e.g. 10. Default 24." - changed
Input schema / properties / query / descriptionPrevious value: -"Product search query, e.g. \"desk\" or \"bookshelf\"."New value: +"String search term for the product name or type, e.g. \"desk\" or \"bookshelf\". Required, no default."
- Changed
search_images3 fields changed- changed
Input schema / properties / engine / descriptionPrevious value: -"Search engine to use (default \"bing\")."New value: +"String naming the search engine, one of \"bing\", \"google\", or \"duckduckgo\", e.g. \"google\". Defaults to \"bing\"." - changed
Input schema / properties / max_images / descriptionPrevious value: -"Maximum number of image results (default 20)."New value: +"Integer cap on the number of results returned, e.g. 10. Defaults to 20." - changed
Input schema / properties / query / descriptionPrevious value: -"Image search query, e.g. \"golden gate bridge\"."New value: +"Search terms as a string, e.g. \"golden gate bridge\". Required, no default."
- Changed
search_linkedin_jobs3 fields changed- changed
Input schema / properties / location / descriptionPrevious value: -"Location filter, e.g. \"London\" or \"United Kingdom\"."New value: +"Location filter as a string; city, region, or country. Example: \"London\" or \"United Kingdom\". No default (required)." - changed
Input schema / properties / max_pages / descriptionPrevious value: -"Pages of results to scrape (default 1)."New value: +"Number of result pages to scrape, as an integer. Each extra page adds jobs but more scraping time. Example: 2. Default: 1." - changed
Input schema / properties / query / descriptionPrevious value: -"Job title or keywords, e.g. \"machine learning engineer\"."New value: +"Job title or keywords as a string. Example: \"machine learning engineer\". No default (required)."
- Changed
search_newegg2 fields changed- changed
Input schema / properties / max_pages / descriptionPrevious value: -"Number of result pages to scrape (default 1)."New value: +"Number of result pages to scrape, given as an integer; higher values return more products but take longer. Example: 3. Default 1." - changed
Input schema / properties / query / descriptionPrevious value: -"Product search query, e.g. \"graphics card\"."New value: +"Product search terms, given as a string. Example: \"graphics card\". No default (required)."
- Changed
search_soundcloud2 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"Maximum number of tracks to return (default 20)."New value: +"Maximum number of tracks to return, as an integer. Example: 10. Default 20." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query, e.g. \"lofi beats\"."New value: +"Search query string. Example: \"lofi beats\". No default (required)."
- Changed
search_ubereats2 fields changed- changed
Input schema / properties / city / descriptionPrevious value: -"City name, e.g. \"London\"."New value: +"City name to search, as a string. Example: \"London\". No default (required)." - changed
Input schema / properties / max_results / descriptionPrevious value: -"Maximum restaurants to return (default 30)."New value: +"Maximum number of restaurants to return, as an integer. Example: 10. Default 30."
- Changed
search_youtube2 fields changed- changed
Input schema / properties / max_results / descriptionPrevious value: -"Maximum number of videos to return (default 20)."New value: +"Maximum number of videos to return, as an integer. Example: 10. Default 20." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query, e.g. \"model context protocol tutorial\"."New value: +"The search text, as a string. Example: \"model context protocol tutorial\". No default (required)."
24 tool updates
v1.3.3- Changed
convert_currency20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / amount / descriptionAdded value: +"Amount of base currency to convert (default 1)." - removed
Input schema / properties / amount / titleRemoved value: -"Amount" - added
Input schema / properties / base / descriptionAdded value: +"Base currency code, e.g. \"USD\"." - removed
Input schema / properties / base / titleRemoved value: -"Base" - added
Input schema / properties / to / descriptionAdded value: +"Comma-separated target codes (e.g. \"EUR,GBP\"). Omit for all rates." - removed
Input schema / properties / to / titleRemoved value: -"To" - removed
Input schema / titleRemoved value: -"convert_currencyArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
define_word16 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / word / descriptionAdded value: +"The word to define." - removed
Input schema / properties / word / titleRemoved value: -"Word" - removed
Input schema / titleRemoved value: -"define_wordArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
get_crypto20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_results / descriptionAdded value: +"Max coins to return (default 20)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Comma-separated coins (e.g. \"bitcoin, ethereum\"). Omit for top coins." - removed
Input schema / properties / query / titleRemoved value: -"Query" - added
Input schema / properties / vs_currency / descriptionAdded value: +"Fiat currency for prices, e.g. \"usd\", \"eur\" (default \"usd\")." - removed
Input schema / properties / vs_currency / titleRemoved value: -"Vs Currency" - removed
Input schema / titleRemoved value: -"get_cryptoArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
get_ubereats_menu16 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / store_url / descriptionAdded value: +"A store URL from a search_ubereats result's \"url\" field." - removed
Input schema / properties / store_url / titleRemoved value: -"Store Url" - removed
Input schema / titleRemoved value: -"get_ubereats_menuArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
get_weather16 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / location / descriptionAdded value: +"Place name, e.g. \"London\" or \"Tokyo, Japan\"." - removed
Input schema / properties / location / titleRemoved value: -"Location" - removed
Input schema / titleRemoved value: -"get_weatherArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
list_available_scrapers3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - removed
Input schema / titleRemoved value: -"list_available_scrapersArguments" - removed
Output schema / titleRemoved value: -"list_available_scrapersDictOutput"
- Changed
lookup_movie18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_pages / descriptionAdded value: +"Pages of search results to fetch, 10 per page (title search)." - removed
Input schema / properties / max_pages / titleRemoved value: -"Max Pages" - added
Input schema / properties / query / descriptionAdded value: +"A title to search for (e.g. \"inception\"), or an IMDB id\n(e.g. \"tt1375666\") for a direct lookup." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"lookup_movieArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
scrape_news22 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / article_url / descriptionAdded value: +"A single article URL to extract full text from." - removed
Input schema / properties / article_url / titleRemoved value: -"Article Url" - added
Input schema / properties / feed_url / descriptionAdded value: +"Direct URL to an RSS/Atom feed." - removed
Input schema / properties / feed_url / titleRemoved value: -"Feed Url" - added
Input schema / properties / max_articles / descriptionAdded value: +"Max articles to return from a feed (default 50)." - removed
Input schema / properties / max_articles / titleRemoved value: -"Max Articles" - added
Input schema / properties / site_url / descriptionAdded value: +"News site URL — its feed is auto-discovered." - removed
Input schema / properties / site_url / titleRemoved value: -"Site Url" - removed
Input schema / titleRemoved value: -"scrape_newsArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
scrape_stock20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / mode / descriptionAdded value: +"\"quote\", \"history\", or \"profile\"." - removed
Input schema / properties / mode / titleRemoved value: -"Mode" - added
Input schema / properties / period / descriptionAdded value: +"History window when mode=\"history\", e.g. \"1mo\", \"1y\"." - removed
Input schema / properties / period / titleRemoved value: -"Period" - added
Input schema / properties / symbol / descriptionAdded value: +"Ticker symbol, e.g. \"AAPL\", \"GOOGL\"." - removed
Input schema / properties / symbol / titleRemoved value: -"Symbol" - removed
Input schema / titleRemoved value: -"scrape_stockArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
scrape_url22 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_pages / descriptionAdded value: +"Follow pagination up to this many pages (default 1)." - removed
Input schema / properties / max_pages / titleRemoved value: -"Max Pages" - added
Input schema / properties / render_js / descriptionAdded value: +"Render JavaScript with a browser backend (needs pyscrappy[browser])." - removed
Input schema / properties / render_js / titleRemoved value: -"Render Js" - added
Input schema / properties / selectors / descriptionAdded value: +"Optional CSS selectors, e.g. {\"title\": \"h1\", \"price\": \".amount\"}." - removed
Input schema / properties / selectors / titleRemoved value: -"Selectors" - added
Input schema / properties / url / descriptionAdded value: +"The page to scrape." - removed
Input schema / properties / url / titleRemoved value: -"Url" - removed
Input schema / titleRemoved value: -"scrape_urlArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
scrape_wikipedia18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / mode / descriptionAdded value: +"\"full\", \"paragraphs\", or \"headers\"." - removed
Input schema / properties / mode / titleRemoved value: -"Mode" - added
Input schema / properties / query / descriptionAdded value: +"Article title or search term, e.g. \"Model Context Protocol\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"scrape_wikipediaArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
scrape_with18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / args / descriptionAdded value: +"Keyword arguments passed to that scraper's `scrape()` method." - removed
Input schema / properties / args / titleRemoved value: -"Args" - added
Input schema / properties / name / descriptionAdded value: +"A scraper name from `list_available_scrapers`, e.g. \"wikipedia\"." - removed
Input schema / properties / name / titleRemoved value: -"Name" - removed
Input schema / titleRemoved value: -"scrape_withArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
scrape_zomato20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / city / descriptionAdded value: +"City name, e.g. \"Bangalore\"." - removed
Input schema / properties / city / titleRemoved value: -"City" - added
Input schema / properties / max_results / descriptionAdded value: +"Maximum number of restaurants to return (default 50)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Optional cuisine or restaurant search term." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"scrape_zomatoArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_amazon18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_pages / descriptionAdded value: +"Number of result pages to scrape (default 1)." - removed
Input schema / properties / max_pages / titleRemoved value: -"Max Pages" - added
Input schema / properties / query / descriptionAdded value: +"Product search query, e.g. \"wireless headphones\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_amazonArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_books18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_results / descriptionAdded value: +"Max books to return (default 20)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Title, author, or free-text search." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_booksArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_github20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_results / descriptionAdded value: +"Max repositories to return (default 20)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Search query, e.g. \"web scraping language:python\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - added
Input schema / properties / sort / descriptionAdded value: +"\"best-match\" (default), \"stars\", \"forks\", or \"updated\"." - removed
Input schema / properties / sort / titleRemoved value: -"Sort" - removed
Input schema / titleRemoved value: -"search_githubArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_hackernews20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / by / descriptionAdded value: +"\"relevance\" (default) or \"date\" (most recent first)." - removed
Input schema / properties / by / titleRemoved value: -"By" - added
Input schema / properties / max_results / descriptionAdded value: +"Max stories to return (default 20)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Search query." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_hackernewsArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_ikea22 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / country / descriptionAdded value: +"IKEA store country code, e.g. \"us\", \"gb\", \"de\" (default \"us\")." - removed
Input schema / properties / country / titleRemoved value: -"Country" - added
Input schema / properties / lang / descriptionAdded value: +"Language code for that store, e.g. \"en\", \"de\" (default \"en\")." - removed
Input schema / properties / lang / titleRemoved value: -"Lang" - added
Input schema / properties / max_results / descriptionAdded value: +"Maximum number of products to return (default 24)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Product search query, e.g. \"desk\" or \"bookshelf\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_ikeaArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_images20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / engine / descriptionAdded value: +"Search engine to use (default \"bing\")." - removed
Input schema / properties / engine / titleRemoved value: -"Engine" - added
Input schema / properties / max_images / descriptionAdded value: +"Maximum number of image results (default 20)." - removed
Input schema / properties / max_images / titleRemoved value: -"Max Images" - added
Input schema / properties / query / descriptionAdded value: +"Image search query, e.g. \"golden gate bridge\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_imagesArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_linkedin_jobs20 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / location / descriptionAdded value: +"Location filter, e.g. \"London\" or \"United Kingdom\"." - removed
Input schema / properties / location / titleRemoved value: -"Location" - added
Input schema / properties / max_pages / descriptionAdded value: +"Pages of results to scrape (default 1)." - removed
Input schema / properties / max_pages / titleRemoved value: -"Max Pages" - added
Input schema / properties / query / descriptionAdded value: +"Job title or keywords, e.g. \"machine learning engineer\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_linkedin_jobsArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_newegg18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_pages / descriptionAdded value: +"Number of result pages to scrape (default 1)." - removed
Input schema / properties / max_pages / titleRemoved value: -"Max Pages" - added
Input schema / properties / query / descriptionAdded value: +"Product search query, e.g. \"graphics card\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_neweggArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_soundcloud18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_results / descriptionAdded value: +"Maximum number of tracks to return (default 20)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Search query, e.g. \"lofi beats\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_soundcloudArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_ubereats18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / city / descriptionAdded value: +"City name, e.g. \"London\"." - removed
Input schema / properties / city / titleRemoved value: -"City" - added
Input schema / properties / max_results / descriptionAdded value: +"Maximum restaurants to return (default 30)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - removed
Input schema / titleRemoved value: -"search_ubereatsArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
- Changed
search_youtube18 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / max_results / descriptionAdded value: +"Maximum number of videos to return (default 20)." - removed
Input schema / properties / max_results / titleRemoved value: -"Max Results" - added
Input schema / properties / query / descriptionAdded value: +"Search query, e.g. \"model context protocol tutorial\"." - removed
Input schema / properties / query / titleRemoved value: -"Query" - removed
Input schema / titleRemoved value: -"search_youtubeArguments" - removed
Output schema / $defsRemoved value: -{ - "ToolError": { - "description": "A non-fatal problem encountered while scraping.", - "properties": { - "message": { - "title": "Message", - "type": "string" - }, - "url": { - "title": "Url", - "type": "string" - } - }, - "required": [ - "url", - "message" - ], - "title": "ToolError", - "type": "object" - } -} - removed
Output schema / properties / count / titleRemoved value: -"Count" - removed
Output schema / properties / data / titleRemoved value: -"Data" - removed
Output schema / properties / errors / items / $refRemoved value: -"#/$defs/ToolError" - added
Output schema / properties / errors / items / descriptionAdded value: +"A non-fatal problem encountered while scraping." - added
Output schema / properties / errors / items / propertiesAdded value: +{ + "message": { + "type": "string" + }, + "url": { + "type": "string" + } +} - added
Output schema / properties / errors / items / requiredAdded value: +[ + "url", + "message" +] - added
Output schema / properties / errors / items / typeAdded value: +"object" - removed
Output schema / properties / errors / titleRemoved value: -"Errors" - removed
Output schema / properties / scraper / titleRemoved value: -"Scraper" - removed
Output schema / properties / source_urls / titleRemoved value: -"Source Urls" - removed
Output schema / titleRemoved value: -"ScrapeToolResult"
5 tool updates
v1.3.0- Added
convert_currency - Added
list_available_scrapers - Added
scrape_stock - Added
scrape_with - Added
search_hackernews
4 tool updates
v1.2.0- Removed
scrape_stock - Added
scrape_wikipedia - Added
search_ikea - Added
search_linkedin_jobs
TDQS
Scored across 24 tools
Most tools are clearly separated by target source or data type, so an agent can usually pick the right one from the name alone. The main ambiguity is that scrape_with is a generic dispatch that can overlap with the specialized scrapers, and a few search/scrape pairs (e.g. search_ubereats vs scrape_zomato) serve similar purposes but descriptions disambiguate them.
Names follow readable snake_case verb_noun structure, but the verbs are mixed across three patterns: search_*, scrape_*, and get_* with occasional one-offs like lookup_movie, define_word, and scrape_with. This is consistent within clusters but not predictable across the whole toolset.
At 24 tools, the set is in the heavy range and feels like a broad aggregation of scrapers rather than a tightly scoped server. The count is defensible for a multi-site scraping toolkit, but it is large enough that navigation and selection become a real burden.
The server covers the core scraper lifecycle well: generic URL scraping, named-source scrapers, plugin discovery, and generic dispatch via scrape_with. Minor gaps exist, such as no runtime scraper registration and no detailed arg discovery for registered plugin scrapers, but the main workflows are functional.
Maintenance
Related MCP Connectors
Structured web research tool for AI agents: search, fetch and shape web data into the JSON schema…
Web data for agents: YouTube transcripts, screenshots, Google News, WHOIS, jobs, tech stack, more.
Public web data for AI agents: scrape any page, plus Google, TikTok, Instagram and Amazon as JSON.
Unblocking and fresh web data for agents: URL to Markdown, YouTube, Maps, Amazon, jobs. Pay per call
Related MCP Servers
- AlicenseAqualityAmaintenanceWeb scraping, crawling, and structured data extraction for AI agents. 5 tools: scrape (clean markdown from any URL), crawl (entire sites), map (discover URLs), extract (structured JSON), and search. 833ms avg latency, single binary, self-hostable.81,105AGPL 3.0
- AlicenseNot gradedqualityBmaintenanceWeb scraping and search MCP server that wraps Firecrawl API for URL discovery and web search with optional content retrieval.140 npm1MIT
- AlicenseNot gradedqualityDmaintenanceFetches web pages and converts them to markdown for LLM consumption, supporting chunked reading and raw content extraction.MIT

Scout MCP Serverofficial
AlicenseAqualityDmaintenanceProvides coding agents with live web capabilities including web search, scraping to Markdown, structured extraction, crawling, screenshots, and company lookup, all with zero dependencies.81MIT