Skip to main content
Glama
Maybeyes111

google-scrape-mcp

by Maybeyes111

google-scrape-mcp

MCP server for Google Search via pure scraping — no API key required. Two engines work together: a fast HTTP path (curl_cffi with real Chrome TLS impersonation) and a Camoufox headless browser fallback that actually runs JavaScript when Google challenges the request.

Built for agents that need Google results reliably and honestly: every response carries an explicit status, and the server never fabricates results.

Features

  • Unified search tool + 9 tabs: web, images, videos, news, books, shopping, scholar, patents, and AI Mode.

  • Cookie bootstrap fast path: a browser homepage warm-up mints fresh session cookies; subsequent HTTP searches pass in 0.3–2.9s (vs 5–15s full browser render). Cookies are cached for 5 minutes.

  • Browser fallback that obeys JS: persistent Camoufox profile, human-like pacing, consent handling, challenge wait/reload — everything a real visitor does.

  • Profile rotation: when a browser identity gets burned (repeated blocks), it is archived and a fresh profile takes over automatically.

  • Adaptive cooldowns: HTTP is retried per endpoint family with escalating backoff; browser blocks trigger a fail-fast cooldown so a rate-limit period does not cost 100s per call.

  • Proxy pool: rotation, per-proxy cooldown, proven-only usage for browser fallback, plus a curation tool that validates proxies against /search.

  • Forensics & research harness: blocked pages are sampled to disk with metadata; probe measures block types per endpoint/engine so you can tune policies with data instead of guesses.

  • RSS/JSON surfaces (news, patents, trends, suggest, translate, finance) that are essentially never blocked and are preferred for reliability.

Related MCP server: Google Search Tool

Install

pip install -e .            # HTTP engine only
pip install -e ".[browser]" # + Camoufox browser engine
camoufox fetch              # download the Camoufox browser once

Core dependencies: fastmcp, curl_cffi, beautifulsoup4, lxml. Optional: camoufox (browser engine).

Run

google-scrape-mcp           # stdio transport (for MCP clients)

MCP client config:

{
  "mcpServers": {
    "google-scrape": {
      "command": "google-scrape-mcp"
    }
  }
}

Tools (20)

Tool

Source

Notes

google_search

unified: tab = web/images/videos/news/books/shopping/scholar/patents/ai, page 1–10

dispatcher to the tools below

google_web_search

organic results, featured snippet, related searches

HTTP fast path, browser fallback

google_image_search

direct image URLs, page URLs, dimensions

fast path OK

google_video_search

tbm=vid

fast path OK

google_books_search

tbm=bks

fast path OK

google_shopping_search

product title, price, was-price, merchant, rating

fast path when possible, browser fallback

google_news_search / google_news_homepage

News RSS

always live, never blocked

google_scholar_search

papers, authors/venue, citations, PDF links

HTTP 429 → browser

google_scholar_cited_by

Scholar cites=

extra protection: often blocked (reported honestly)

google_patents_search

Patents XHR JSON

always live

google_finance_quote

stocks/forex/crypto (USD-IDR, BBCA:IDX)

always live

google_translate

unofficial gtx endpoint

always live

google_suggest

autocomplete

always live

google_trends_daily

Trends RSS

always live

google_trends_interest

explore → multiline widgetdata

HTTP often 401 → browser fallback

google_ai_mode

AI Mode (udm=50): synthesized answer + sources

browser (JS streams the answer)

google_crawl

read any URL: title, meta, text, outbound links

live

google_help

agent usage guide (same as AGENT_GUIDE.md)

live

google_status

health check: endpoints, proxies, cache, cooldowns

live

All tools return {"status": "ok" | "blocked" | "rate_limited" | "limited" | "empty" | "error", ...}. Blocked means blocked — no fake results.

AI agents: read AGENT_GUIDE.md, or call the google_help tool at runtime for the same guidance.

Engines

Every search tool accepts engine:

  • auto (default) — HTTP fast path with bootstrap cookies, then browser.

  • http — direct HTTP only (challenge-prone on flagged IPs).

  • proxy — force the proxy pool.

  • browser — Camoufox headless render (~5–15s).

  1. A browser warm-up visits google.com (persistent profile) and mints fresh session cookies (NID, AEC, SNID, GSP, …).

  2. Those cookies are sent over plain HTTP with coherent navigation metadata (Referer + Sec-Fetch-Site: same-origin) from a clean session jar.

  3. Surfaces that pass over HTTP: web, images, videos, books, shopping. Scholar (429) and AI Mode (answer is streamed by JS) stay on the browser.

Cookies are cached in ~/.cache/google-scrape-mcp/bootstrap_cookies.json with a 5-minute TTL (GOOGLE_SCRAPE_COOKIE_TTL).

Proxy pool

Accepted line formats: http://host:port, https://host:port, socks5://host:port (socks5h normalized), socks4://host:port, with optional user:pass@, plus bare host:port and host:port:user:pass.

Sources (priority order):

  1. GOOGLE_SCRAPE_PROXIES — inline list (comma/space/newline separated) or a file path (starting with /, ~, or ending in .txt).

  2. GOOGLE_SCRAPE_PROXY_FILES — colon-separated file paths.

  3. Defaults when present: ~/.config/google-scrape/proxies.txt and ~/.cache/google-scrape-mcp/proxies_curated.txt (max 500 lines/file).

Failed proxies get a cooldown (GOOGLE_SCRAPE_PROXY_COOLDOWN, default 600s). Browser fallback only uses proxies that have succeeded before (proven), so dead datacenter pools don't waste minutes.

Curate your own pool (validates liveness and /search capability):

python3 -m google_scrape_mcp.curate --limit 100 --workers 8

Environment variables

Env

Default

Purpose

GOOGLE_SCRAPE_PROXY_MODE

auto

auto (direct→proxy), off, always

GOOGLE_SCRAPE_PROXY_COOLDOWN

600

failed-proxy cooldown (s)

GOOGLE_SCRAPE_PROXY_CA

—

extra CA bundle for self-signed HTTPS proxies

GOOGLE_SCRAPE_MIN_INTERVAL

1.0

min delay between requests (s)

GOOGLE_SCRAPE_TIMEOUT

25

request timeout (s)

GOOGLE_SCRAPE_RETRIES

3

proxy attempts per request

GOOGLE_SCRAPE_CACHE_TTL

600

response cache TTL (0 disables)

GOOGLE_SCRAPE_CACHE_MAX

256

max cache entries

GOOGLE_SCRAPE_IMPERSONATE

chrome120

curl_cffi TLS target

GOOGLE_SCRAPE_JS_COOLDOWN

600

base HTTP cooldown after a block (doubles per level, capped at 1h)

GOOGLE_SCRAPE_COOKIE_TTL

300

bootstrap cookie lifetime (s)

GOOGLE_SCRAPE_BROWSER_COOLDOWN

180

base browser cooldown after a block (doubles per level)

GOOGLE_SCRAPE_FORENSICS

1

save blocked-page samples

GOOGLE_SCRAPE_NO_SELFHEAL

—

disable automatic camoufox fetch when the browser is missing

Research & forensics

  • RESEARCH.md — measured findings: why raw HTTP /search always gets a JS challenge, why cookies alone don't help, how the bootstrap fast path works, surface-by-surface results, profile burnout.

  • python3 -m google_scrape_mcp.probe --quick|--full|--tls|--analyze — controlled block-measurement harness with human-like spacing.

  • Blocked pages are sampled (HTML + metadata) under ~/.cache/google-scrape-mcp/forensics/, with counters in block_stats.json (surfaced by google_status).

Known limitations

  • google_scholar_cited_by is frequently blocked (Google protects that endpoint harder). Use the cited_by count from google_scholar_search.

  • Shopping product URLs are rendered on click; the tool returns title, price, was-price, merchant and rating.

  • No Maps/local search, reverse image search, AI Overview (inline) or flights.

  • Datacenter proxy pools are mostly useless for Google; prefer residential.

License

MIT — see LICENSE.

Available Tools

20 tools
google_ai_modeGoogle Ai ModeA

Google Mode AI (udm=50): jawaban sintesis + sumber, cocok untuk pertanyaan langsung. Bentuk hasil: answer + sources.

engine: auto | http | proxy | browser (lihat google_web_search). Halaman Mode AI hampir selalu butuh JS; engine auto akan memakai browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
glNous
hlNoen
queryYes
engineNoauto

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers real behavioral context: it discloses the result shape (answer + sources) and the important operational trait that AI Mode pages almost always require JS, so engine=auto will select the browser. Rate limits, auth needs, and latency are not covered, but the JS/engine caveat is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and output shape are front-loaded, and the engine note is separated into its own block. It is compact and largely waste-free, with only minor redundancy around the engine explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be detailed, and the engine/JS behavior is covered. However, the locale parameters gl and hl are undocumented in both description and schema, leaving a real gap for a 4-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for 4 parameters, so the description must compensate. It explains the engine parameter's valid values (auto|http|proxy|browser) and their behavioral implication, but leaves gl, hl, and query entirely undocumented, so its coverage is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and mode (Google AI Mode, udm=50) and states the output type (synthesis answer + sources) and intended use (direct questions). It distinguishes itself from generic search siblings by emphasizing synthesized answers rather than link lists, though it never explicitly names a contrasting sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies usage with 'cocok untuk pertanyaan langsung' (suitable for direct questions), which signals when this tool fits, but offers no explicit when-not-to-use guidance or named alternative. The engine reference to google_web_search is a helpful pointer but does not route between the two tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_crawlGoogle CrawlC

Crawl any URL (e.g. a search result): title, meta, text, outbound links.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_charsNo
include_linksNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It hints at the extraction scope via the returned field list, but says nothing about permissions, rate limits, redirect/robots handling, or side effects of crawling an arbitrary URL.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the core action front-loaded and zero filler. It is efficient, though its brevity is partly under-specification rather than deliberate tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, but with zero annotation coverage and 0% parameter documentation the definition leaves the two optional controls and the crawling behavior entirely undefined for a tool that fetches arbitrary external URLs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all 3 parameters. The description only obliquely maps 'outbound links' to include_links and never mentions max_chars (which defaults to 8000) or the meaning/units of these controls, so it fails to compensate for the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a specific verb ('Crawl') with a resource ('any URL') and enumerates the extracted artifacts (title, meta, text, outbound links). This clearly separates it from the sibling search/quote/translate tools, though it does not name a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Crawl any URL (e.g. a search result)' implies the context of fetching page content behind a link, but gives no explicit when-to-use/when-not guidance or alternatives. An agent must infer that this is the follow-up to a search result rather than a search itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_finance_quoteGoogle Finance QuoteB

Scrape Google Finance quote page: price, prev close, day high/low, change.

ParametersJSON Schema
NameRequiredDescriptionDefault
hlNoen
tickerYes
exchangeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing beyond the resource being scraped. It says nothing about rate limits, blocking/anti-bot behavior, auth needs, or whether the ticker must correspond to a known exchange listing - all material for a scraper.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action, the target, and the payload with zero filler. Nothing to trim and nothing important buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but the definition leaves the two optional parameters entirely undocumented and omits all behavioral context for a scraper with no annotations. An agent knows roughly what it does but not how to call it safely or completely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across all three parameters. The description hints only at the subject of the scrape (a quote page) and does not explain 'ticker' format, what 'exchange' does when the ticker is ambiguous, or that 'hl' controls language.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Scrape'), a specific resource ('Google Finance quote page'), and enumerates the returned fields (price, prev close, day high/low, change). No sibling tool covers finance quotes, so an agent can route here unambiguously from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus any alternative, nor any exclusions or prerequisites. The use case (fetching a stock quote) is only inferable from the tool name and the listed fields, not from explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_helpGoogle HelpA

Agent-facing usage guide: tool map, engine semantics, status contract and safe usage patterns. Read this before experimenting.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the topics covered (including a 'status contract' and 'safe usage patterns'), which signals it is a read-only reference, but it never explicitly states that it has no side effects or how the returned guidance should be consumed. For a zero-parameter read tool the risk is low, so this gap is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a trailing imperative; every clause (tool map, engine semantics, status contract, safe patterns) adds distinct information and none is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and an output schema present, the description needn't explain return values, and it adequately frames what the call yields and when to make it. Nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the calibration this defaults to a baseline of 4. The description adds no parameter detail, but there is nothing to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (an agent-facing usage guide) and enumerates its contents: tool map, engine semantics, status contract, and safe usage patterns. This clearly separates it from all the search/quote/translate siblings, which perform actual lookups rather than explaining the toolkit. It stops short of a crisp 'returns documentation about X' verb, keeping it just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Read this before experimenting' gives explicit when-to-use guidance pointing at the pre-use/onboarding moment, which is exactly when an agent would want a tool map. It doesn't name an alternative or an exclusion, but for a meta help tool no real alternative exists, so the omission is minor.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_news_homepageGoogle News HomepageC

Scrape Google News homepage RSS: top headlines right now.

ParametersJSON Schema
NameRequiredDescriptionDefault
glNoUS
hlNoen-US
ceidNoUS:en

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full behavioral burden, and it does not. It implies a read-only RSS scrape and hints at real-time freshness ('right now'), but says nothing about rate limits, auth requirements, result volume, or pagination for a scraping tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is efficient. Given the undocumented locale parameters, though, the brevity shades into under-specification rather than ideal density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return values need no explanation, but three undocumented parameters and zero annotations leave the definition materially incomplete for a tool whose behavior depends on locale inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three parameters (gl, hl, ceid) have 0% schema description coverage and no enums, and the description adds nothing about them. Locale/region codes like 'US:en' are non-obvious, and the description leaves the agent unable to tell that these control geo/language of the feed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scrape) and resource (Google News homepage RSS) plus what it returns ('top headlines right now'). It implicitly separates itself from google_news_search by naming the homepage, but never explicitly routes between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no conditions, and no mention of the obvious alternative google_news_search for query-based retrieval. The agent must infer that this is the un-parameterized 'front page' variant.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_scholar_cited_byGoogle Scholar Cited ByC

Scrape papers citing a Scholar cluster (from cluster_id in scholar results).

engine: auto | http | proxy | browser (lihat google_web_search).

ParametersJSON Schema
NameRequiredDescriptionDefault
hlNoen
numNo
startNo
engineNoauto
cluster_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says 'scrape' but discloses nothing about rate limits, anti-bot behavior, pagination semantics, or auth requirements; the engine note merely delegates to another tool rather than describing behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no padding and the core purpose front-loaded. However, the second sentence is a mixed-language fragment ('engine: auto | http | proxy | browser (lihat google_web_search)') that adds a reference without explaining it, weakening structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, but with 5 parameters at 0% schema coverage and zero annotations the description leaves too much unspecified for an agent to call the tool correctly across hl/num/start/engine.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 5 parameters, so the description must compensate. It only hints at cluster_id's origin and lists engine values via a cross-reference; hl, num, and start remain completely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('scrape') and resource ('papers citing a Scholar cluster') and clarifies the source of the required cluster_id, which separates it from sibling google_scholar_search. It is clear but does not explicitly contrast itself with the other Scholar tool beyond the parenthetical.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note that cluster_id comes 'from cluster_id in scholar results' implies the workflow (first run a Scholar search, then use this), but there is no explicit when-to-use/when-not guidance or named alternative for obtaining citation data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_statusGoogle StatusA

Health check: which Google endpoints are reachable from this IP right now, plus proxy pool, cache, cooldowns and forensic counters.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the content of the report: endpoint reachability, proxy pool, cache, cooldowns and forensic counters. It implies a read-only diagnostic but does not explicitly state that it mutates nothing or that no auth/quota consumption is involved, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the 'Health check' framing and then the specific report contents. Every clause adds information and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not enumerate return values, yet it helpfully previews the report sections. For a zero-param diagnostic tool this is nearly complete; the only omission is an explicit statement that it is a safe, non-mutating probe.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document and the baseline is 4. Nothing in the description is needed to explain inputs, and it introduces none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: a health check reporting which Google endpoints are reachable from the caller's IP, extended by the diagnostic surfaces it includes (proxy pool, cache, cooldowns, forensic counters). This is clearly distinguishable from every sibling, all of which fetch search or data results rather than report connectivity state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Health check' implies a diagnostic use case, and no sibling overlaps, so usage is reasonably inferable. However, the description never explicitly says when to call this (e.g. before other Google calls, or when a sibling fails), leaving the trigger condition to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_suggestGoogle SuggestC

Google autocomplete suggestions (no block).

ParametersJSON Schema
NameRequiredDescriptionDefault
hlNoen
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. '(no block)' is vague and may hint at a behavioral trait, but it does not clarify read-only status, auth needs, rate limits, or response behavior. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded with no filler. However, the parenthetical '(no block)' is unclear and the overall phrase is too terse to be truly helpful given the lack of structured context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Still, with no annotations and 0% schema description coverage, the description should explain parameter meanings and basic behavior, which it does not. It is incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no meaning for either parameter. It does not explain the required 'query' parameter or the optional 'hl' language parameter, leaving the agent with only the schema's bare type information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource clearly: 'Google autocomplete suggestions.' It distinguishes the tool from siblings like google_search and google_web_search by focusing on autocomplete rather than search results. However, it lacks a specific verb and the '(no block)' qualifier is cryptic, so it is not fully unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives are provided. The agent can infer that this tool returns autocomplete suggestions, but there is no explicit context for selecting it over other search-related siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_translateGoogle TranslateB

Translate text via Google's unofficial gtx endpoint (no key).

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
sourceNoauto
targetNoen

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that the endpoint is unofficial and requires no API key, which tells the agent about auth expectations and reliability risk. However, it says nothing about rate limits, error/failure behavior, or supported language handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core action front-loaded and no wasted words. Nothing to trim and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the tool is simple. Still, the undocumented language-code parameters and absence of any usage or behavioral guidance leave real gaps for an agent invoking it unaided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions 'text', 'source', or 'target'. Critically, it does not clarify that source/target are language codes or what format they take, nor the meaning of the 'auto' default — the schema field names alone leave ambiguity that the description should have resolved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Translate text') and identifies the mechanism used (Google's unofficial gtx endpoint). It is self-evidently distinct from the sibling tools, which are all search/finance/trend lookups, though it doesn't explicitly contrast itself with any of them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, no prerequisites, and no exclusions. The purpose is inferable from the name, but the description provides no routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv0.7.2
    • First observedgoogle_ai_mode
    • First observedgoogle_books_search
    • First observedgoogle_crawl
    • First observedgoogle_finance_quote
    • First observedgoogle_help
    • First observedgoogle_image_search
    • First observedgoogle_news_homepage
    • First observedgoogle_news_search
    • First observedgoogle_patents_search
    • First observedgoogle_scholar_cited_by
    • First observedgoogle_scholar_search
    • First observedgoogle_search
    • First observedgoogle_shopping_search
    • First observedgoogle_status
    • First observedgoogle_suggest
    • First observedgoogle_translate
    • First observedgoogle_trends_daily
    • First observedgoogle_trends_interest
    • First observedgoogle_video_search
    • First observedgoogle_web_search

TDQS

B3/5.0

Scored across 20 tools

Disambiguation2/5

google_search promises one tool for all Google tabs, yet separate web/image/video/books/shopping/news/scholar/patents search tools duplicate those same tabs, creating unclear boundaries. Descriptions explain some distinctions, but an agent can easily misselect between the unified and specialized tools.

Naming Consistency5/5

All names use a google_ prefix with snake_case, following a consistent service_action or service_property pattern. The minor variation in action vocabulary is predictable and readable.

Tool Count3/5

20 tools is borderline heavy. The redundancy between unified google_search and specialized vertical search tools means several tools may not earn their place, though the breadth of Google properties accounts for many.

Completeness4/5

Covers major Google verticals including web, images, video, books, shopping, news, scholar, patents, trends, finance, and translate, plus crawl/help/status. Some Google properties like Maps, Jobs, and Flights are missing, but the core scraping surface is strong.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Enables free web searching using Google search results with no API keys required, returning structured results with titles, URLs, and descriptions.
    1
    36 npm
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to perform real-time Google searches with anti-bot protection. Bypasses search engine restrictions using advanced browser automation to extract search results locally without requiring paid API services.
    6
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to perform Google searches and retrieve structured JSON results for 10 types including web, images, news, and shopping with live prices. Offers 1,000 free searches per month with no credit card required.
    1
    38 npm
    MIT