Skip to main content
Glama
Maybeyes111

google-scrape-mcp

by Maybeyes111

google-scrape-mcp

License: MIT Python 3.10+ CodeQL Bandit Dependabot MCP

google-scrape-mcp

MCP server that gives AI agents real Google results with no API key: an HTTP fast path (0.3–2.9s searches when cookies are warm), a Camoufox browser fallback that actually runs the JS challenges, adaptive cooldowns derived from probing, and honest status codes instead of fabricated results.

live demo

Full-quality recording: assets/demo_real.mp4 (24s). What you see in the recording: web 1.1s, images 0.8s, scholar 7.1s, AI Mode 9.2s (2,548 chars), USD/IDR quote 0.5s.

1. What it does

  • Searches Google across web, images, videos, news, books, shopping, scholar, patents and AI Mode through one MCP server.

  • Answers each request with status: ok | blocked | rate_limited | limited | empty | error. When Google blocks, it says so. It never invents results.

  • Keeps a browser profile warm so HTML pages that need JavaScript (scholar, AI Mode) come back complete.

  • Exposes RSS/JSON endpoints (news, patents, trends, suggest, translate, finance) that are never blocked, for when reliability matters more than coverage.

Related MCP server: Google Search Tool

2. Install

pip install -e .             # HTTP engine only
pip install -e ".[browser]"  # + Camoufox browser engine
camoufox fetch               # download the Camoufox browser once

Core dependencies: fastmcp, curl_cffi, beautifulsoup4, lxml, defusedxml. Optional: camoufox for the browser engine.

Platform support, honestly:

Component

Platforms

Notes

HTTP engine (curl_cffi)

Linux, macOS, Windows (x86_64 wheels; arm64 where published)

no browser needed

Browser engine (camoufox)

Linux x86_64; upstream also ships macOS (Intel/Apple Silicon) and Windows x86_64 builds

~660 MB download; ARM Linux and Android are not supported by upstream builds and need manual work. On those targets use the HTTP engine only and expect more blocked responses

3. MCP client config

{
  "mcpServers": {
    "google-scrape": {
      "command": "google-scrape-mcp"
    }
  }
}

Run google-scrape-mcp directly for stdio transport. Agents can call the google_help tool for a runtime usage guide, or read AGENT_GUIDE.md.

4. Tools (22)

Tool

Source

Notes

google_search

unified: tab = web/images/videos/news/books/shopping/scholar/patents/ai, page 1–10

dispatcher for the tools below

google_web_search

organic results, featured snippet, related searches

HTTP fast path, browser fallback

google_image_search

direct image URLs, page URLs, dimensions

fast path

google_video_search

tbm=vid

fast path

google_books_search

tbm=bks

fast path

google_shopping_search

title, price, was-price, merchant, rating

fast path when the markup allows, browser otherwise

google_news_search / google_news_homepage

News RSS

never blocked

google_scholar_search

papers, venue, citations, PDF links

browser (HTTP 429)

google_scholar_cited_by

Scholar cites=

extra protection, often blocked and reported as such

google_patents_search

Patents XHR JSON

never blocked

google_finance_quote

stocks, forex, crypto. USDIDR, USD/IDR are normalized to USD-IDR; stocks use ticker + exchange (BBCA:IDX)

never blocked

google_fx_rate

fast FX from the SERP converter widget (1 USD to IDR)

returns fx_rate {rate, from, to}

google_kurs_bi

official Bank Indonesia transaction rates (sell/buy, 26 currencies)

browser render of the BI table, currency filter optional

google_translate

unofficial gtx endpoint

never blocked

google_suggest

autocomplete

never blocked

google_trends_daily

Trends RSS

never blocked

google_trends_interest

explore into multiline widgetdata

HTTP often 401, browser fallback

google_ai_mode

AI Mode (udm=50), synthesized answer + sources

browser, JS streams the answer

google_crawl

read any URL: title, meta, text, links

live

google_help

agent usage guide

live

google_status

endpoints, proxies, cache, cooldowns, forensics

live

Every search tool takes an engine argument:

  • auto (default): HTTP first, browser when needed.

  • http: direct only. Fast, challenge-prone on flagged IPs.

  • proxy: force the proxy pool.

  • browser: Camoufox headless render, 5–15s.

The fast path works like this. A browser homepage visit mints fresh session cookies (NID, AEC, SNID, GSP). Plain HTTP /search with those cookies, a coherent Referer/Sec-Fetch-Site and a clean cookie jar passes in 0.3–2.9s. Surfaces proven to pass over HTTP: web, images, videos, books, shopping. Scholar (HTTP 429) and AI Mode (answer is streamed by JS) stay on the browser. Cookies are cached for 5 minutes and refreshed in the background while tools are in use, so searches rarely pay the warm-up cost.

Behavior notes behind this are in RESEARCH.md.

6. Performance

Operation

Before

Now

Web search (end-to-end)

~1.0s

0.3–0.7s

Browser warm-up + cookies

16.4s

5.7s

Scholar search

11.3s

4.2–4.4s

AI Mode answer

11.1s, sometimes truncated

9–11s, complete

SERP parse + /goto resolution

sequential

parallel, ~0.5s

Mechanisms: one priority job for warm-up plus cookies, adaptive warm-up skipping (GOOGLE_SCRAPE_WARM_TTL), selector and text-stability waits instead of fixed sleeps, parallel /goto resolution, a background bootstrap keeper, and a 1.5–3.5s human-like gap between browser jobs.

7. Anti-block research

The repository ships the methodology, not just the result:

  • python3 -m google_scrape_mcp.probe --quick|--full|--tls|--analyze runs a controlled endpoint/engine matrix with human-like spacing.

  • Blocked pages are sampled to ~/.cache/google-scrape-mcp/forensics/ with counters in block_stats.json, surfaced by google_status.

  • Adaptive cooldowns are per endpoint family (search/scholar/finance), with escalating backoff and a fail-fast browser cooldown. Burned browser profiles are rotated automatically.

Measured conclusions, including why raw HTTP /search always gets a JS challenge and why cookies alone do not fix it, are in RESEARCH.md.

8. Proxy pool

Accepted formats: http://host:port, https://host:port, socks5://host:port (socks5h normalized), socks4://host:port, with optional user:pass@, plus bare host:port and host:port:user:pass.

Sources, in priority order:

  1. GOOGLE_SCRAPE_PROXIES: inline list (comma/space/newline separated) or a file path (starting with /, ~, or ending in .txt).

  2. GOOGLE_SCRAPE_PROXY_FILES: colon-separated file paths.

  3. Defaults when present: ~/.config/google-scrape/proxies.txt and ~/.cache/google-scrape-mcp/proxies_curated.txt.

Free lists can work here, but only when combined with the bootstrap cookies: with fresh cookies, clean proxies can pass /search; without them everything gets a JS challenge. curate is cookie-aware, clears the cookie jar before probing, limits search concurrency (session cookies plus parallel IPs looks anomalous), and writes the fastest proxies first.

python3 -m google_scrape_mcp.curate --limit 100 --workers 8

Failed proxies get a cooldown. Browser fallback only uses proxies that have succeeded before, so a dead pool never burns minutes.

9. Environment variables

Variable

Default

Purpose

GOOGLE_SCRAPE_PROXY_MODE

auto

auto, off, or always

GOOGLE_SCRAPE_PROXY_COOLDOWN

600

failed-proxy cooldown (s)

GOOGLE_SCRAPE_PROXY_CA

extra CA bundle for self-signed HTTPS proxies

GOOGLE_SCRAPE_MIN_INTERVAL

1.0

minimum delay between requests (s)

GOOGLE_SCRAPE_TIMEOUT

25

request timeout (s)

GOOGLE_SCRAPE_RETRIES

3

proxy attempts per request

GOOGLE_SCRAPE_CACHE_TTL

600

response cache TTL, 0 disables

GOOGLE_SCRAPE_CACHE_MAX

256

max cache entries

GOOGLE_SCRAPE_IMPERSONATE

chrome120

curl_cffi TLS target

GOOGLE_SCRAPE_JS_COOLDOWN

600

base HTTP cooldown after a block, doubles per level, capped at 1h

GOOGLE_SCRAPE_COOKIE_TTL

300

bootstrap cookie lifetime (s)

GOOGLE_SCRAPE_BROWSER_COOLDOWN

180

base browser cooldown after a block

GOOGLE_SCRAPE_WARM_TTL

600

skip the homepage warm-up if warmed more recently

GOOGLE_SCRAPE_JOB_GAP

1.5-3.5

human-like gap between browser jobs (s)

GOOGLE_SCRAPE_KEEPER

1

background bootstrap keeper, 0 disables

GOOGLE_SCRAPE_FORENSICS

1

save blocked-page samples

GOOGLE_SCRAPE_NO_SELFHEAL

disable automatic camoufox fetch when missing

10. Security

  • CodeQL (Python) and Bandit run on every push and weekly.

  • Dependabot keeps dependencies and GitHub Actions updated, grouped weekly.

  • Secret scanning with push protection is enabled.

  • main is protected against force pushes and deletion, including admins.

  • XML feeds are parsed with defusedxml when available.

  • Reporting: see SECURITY.md.

11. Project layout

src/google_scrape_mcp/
  server.py      MCP tools, engine orchestration, status
  client.py      HTTP path: sessions, cache, cooldowns, bootstrap cookies
  browser.py     Camoufox worker: persistent profile, challenges, rotation
  parsers.py     SERP/RSS/JSON parsing, AI Mode cleanup, shopping cards
  proxies.py     proxy pool: parsing, rotation, cooldowns, proven-only usage
  forensics.py   block taxonomy samples and counters
  probe.py       block-monitoring harness
  curate.py      proxy curation (liveness + /search capability)

Docs: AGENT_GUIDE.md for agents, RESEARCH.md for the anti-block notes.

12. Known limitations

  • google_scholar_cited_by is often blocked (Google protects that endpoint harder). The cited_by count from google_scholar_search still works.

  • Shopping does not expose product URLs in the initial HTML; links render on click.

  • No Maps/local search, reverse image search, inline AI Overview, or flights.

  • Featured snippets for FX queries can be noisy; fx_rate is the clean field.

13. Reliability: this is a cat-and-mouse game

This is an unofficial scraper fighting Google's bot detection. Pretending otherwise would be dishonest. Expect the following:

  • Raw HTTP /search from a flagged IP is challenged almost every time; the fast path exists because fresh browser cookies are attached. Browser profiles can burn out, and they get rotated, but sustained volume without good proxies will hit blocked regularly.

  • Free and datacenter proxy pools are mostly dead or rejected. Residential proxies work better and cost money.

  • Browser sessions are heavy (hundreds of MB, 5-15s per render) and need the Camoufox binary; the fast path exists precisely to avoid them when possible.

  • RSS/JSON surfaces (news, patents, trends, suggest, translate, finance quotes) are the stable part. Search surfaces are the fragile part.

  • Results vary with IP reputation, region and time of day.

If you need guaranteed, stable Google results, use an official API. This project trades that stability for zero cost and no key.

  • Scraping Google Search violates Google's Terms of Service. You run this at your own risk. The MIT license grants no rights beyond the code and takes no responsibility for your use.

  • Rate limits, IP reputation and any account consequences are yours to manage. Do not use this for abuse or to attack services.

  • Treat every proxy as untrusted. A proxy operator can log, inject or modify traffic; a malicious one can serve you forged content. Free proxies are the worst case. Never route credentials or sensitive traffic through pool proxies, and read curated lists as "reachable", not "safe". Prefer residential/VPS proxies you control.

15. License

MIT, see LICENSE.

Available Tools

22 tools
google_ai_modeGoogle Ai ModeA
Read-onlyIdempotent

Google Mode AI (udm=50): jawaban sintesis + sumber, cocok untuk pertanyaan langsung. Bentuk hasil: answer + sources.

engine: auto | http | proxy | browser (lihat google_web_search). Halaman Mode AI hampir selalu butuh JS; engine auto akan memakai browser.

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
glNous
hlNoen
queryYes
engineNoauto

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations covering the safety profile (readOnly, idempotent, openWorld, non-destructive), the description adds genuinely novel context: Google rate-limiting, the fact that blocked results require waiting or switching tools, and that the AI Mode page nearly always needs JS so engine=auto falls back to browser. That is real behavioral value beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and result shape, then engine notes, then rate-limit guidance — three tight blocks with no filler. Slightly fragmented by mixing languages and a parenthetical cross-reference, but every line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description needn't explain return values, and it is complete on rate-limiting, engine selection, and JS fallback. The main gap is that the localization parameters (gl/hl) are undocumented anywhere, leaving an agent guessing about their meaning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% with 4 parameters, so the description carries the burden. It usefully documents the engine parameter's values (auto|http|proxy|browser) and the JS/browser fallback behavior, which the schema does not express. But gl, hl, and query are left entirely unexplained, so compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource with scope: a Google AI Mode (udm=50) synthesis search returning 'answer + sources', positioned for 'pertanyaan langsung' (direct questions). It distinguishes itself contextually from generic listings and references the sibling google_web_search for engine semantics. It stops short of an explicit contrast with the many other search siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives implied usage ('cocok untuk pertanyaan langsung') and explicit failure-path guidance ('when the result is blocked, wait or switch to an RSS-based tool'), which names an alternative class of tool. However there is no clear when-not-to-use statement or routing rule against the concrete siblings (google_web_search, google_search).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_crawlGoogle CrawlA
Read-onlyIdempotent

Crawl any URL (e.g. a search result): title, meta, text, outbound links.

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_charsNo
include_linksNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, open-world, non-destructive behavior, so the safety profile is covered. The description adds genuinely new operational context the annotations cannot convey: rate limiting by Google, avoiding rapid repeats, and what to do when a result is blocked.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no filler. The primary capability is front-loaded and the operational caveat follows immediately after.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be restated, and the rate-limit/blocked behavior is well covered. However, with 0% parameter description coverage across three parameters, the definition is incomplete for correct invocation, particularly for max_chars and include_links.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so nothing in the structured data explains url, max_chars, or include_links. The description names output fields but never defines max_chars (a default of 8000) or how include_links affects the returned outbound links, leaving two of three parameters unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Crawl) and resource (any URL), and enumerates exactly what comes back: title, meta, text, outbound links. That output list also implicitly separates it from the sibling search tools, which return query results rather than page content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete use case ("e.g. a search result") and an explicit fallback path for the blocked case (wait, or switch to an RSS-based tool). It never names a specific sibling tool or states when not to use this one, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_finance_quoteGoogle Finance QuoteA
Read-onlyIdempotent

Scrape Google Finance quote: saham, forex, kripto.

Ticker: 'BBCA' + exchange 'IDX', 'AAPL' + 'NASDAQ', atau pair forex 'USD-IDR' (USDIDR/USD/IDR juga diterima, dinormalisasi otomatis).

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
hlNoen
tickerYes
exchangeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld, so the safety profile is covered. The description adds genuinely new behavioral context: Google imposes rate limits, rapid repeats risk blocking, and blocked results require waiting or an alternative tool. That's real operational guidance beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded blocks: purpose, ticker syntax, rate-limit caveat. No wasted sentences, though the mixed Indonesian/English phrasing adds mild reading friction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be described. With that covered, the description supplies the missing input-format guidance and the rate-limit behavior. Main residual gap is not clarifying its relationship to the overlapping google_fx_rate sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load, and it does: it explains ticker format with concrete examples per exchange ('BBCA'+'IDX', 'AAPL'+'NASDAQ') and forex pair normalization ('USD-IDR', 'USDIDR/USD/IDR' accepted). It leaves 'hl' and 'exchange' defaults unmentioned, hence not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Scrape Google Finance quote') and enumerates covered asset classes (stocks, forex, crypto). It's clear what it does, though it doesn't explicitly distinguish itself from the sibling google_fx_rate, which presumably overlaps on forex quotes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a useful rate-limit workaround ('wait or switch to an RSS-based tool'), which is a strong when-blocked alternative. But it never states when to use this over google_fx_rate for currency pairs, leaving sibling routing ambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_fx_rateGoogle Fx RateA
Read-onlyIdempotent

Kurs langsung dari widget konverter SERP (mis. "1 USD to IDR").

Hasil: fx_rate {rate, from, to, formatted} + results biasa. Lebih tahan terhadap perubahan layout daripada halaman finance, karena membaea widget konverter yang muncul di SERP. Cocok untuk kurs cepat; untuk kurs resmi BI pakai google_kurs_bi.

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
glNous
hlNoen
baseNoUSD
quoteNoIDR
amountNo
engineNoauto

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, open-world), but the description adds real operational context: Google rate-limiting, the blocked-result failure mode, and why it is more layout-resilient than finance pages. It stops short of describing pagination or latency, but the added context is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Short and front-loaded: purpose, return shape, differentiation, and the rate-limit caveat each occupy about one sentence. Only the final rate-limit sentence is slightly tangential to selection, but it earns its place as an operational warning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be detailed, and the routing guidance is complete. However, with six undocumented parameters and no schema descriptions, an agent lacks guidance on locale/engine options, which keeps this from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across six parameters, so the description carries the full burden. The example '1 USD to IDR' loosely implies base/quote/amount, but gl, hl, and engine are never explained and no value formats or defaults are clarified beyond the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('live rate from the SERP converter widget') with a concrete example ('1 USD to IDR') and names the output field fx_rate. It also distinguishes itself from the sibling google_kurs_bi, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('quick rates') and when not to ('for official BI rates use google_kurs_bi'), plus a fallback path when results are blocked (wait or switch to an RSS-based tool). Both the alternative and the exclusions are stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_helpGoogle HelpA
Read-onlyIdempotent

Agent-facing usage guide: tool map, engine semantics, status contract and safe usage patterns. Read this before experimenting.

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, openWorld and non-destructive. The description adds genuinely new behavioral context: Google-side rate limiting, avoiding rapid repeats, and the wait-or-switch R2B fallback when blocked. That is meaningful operational guidance beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the tool's identity and content, then the critical rate-limit guidance. No filler; every clause carries information the agent needs before experimenting.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary. Given the tool is a documentation/help endpoint, the description adequately signals its scope (tool map, engine semantics, status contract, safe usage) and the key operational caveat.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify beyond the schema. Baseline 4 applies since no parameter semantics are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific purpose: an 'Agent-facing usage guide' covering tool map, engine semantics, status contract and safe usage patterns. This clearly separates it from the sibling search/fetch tools, though it does not name which sibling it complements. Verb+resource are concrete and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Read this before experimenting' gives explicit timing for invocation, and the rate-limit sentence routes the agent to an alternative ('switch to an RSS-based tool') when results are blocked. It provides clear context and a fallback path, but does not enumerate when-not-to-use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_kurs_biGoogle Kurs BiA
Read-onlyIdempotent

Kurs Transaksi Bank Indonesia (Jual/Beli) dari tabel resmi BI.

currency opsional: 'USD', 'EUR', dst (kosong = semua). Halaman BI JS-heavy, jadi engine auto memakai browser; engine 'http' dicoba dulu sebagai jalur murah.

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
engineNoauto
currencyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, openWorld), and the description adds genuinely useful behavior: the page is JS-heavy so 'auto' routes to a browser, 'http' is tried first as a cheap path, and Google rate-limits/blocking can occur. That fallback ordering is not derivable from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the resource identity and followed by parameter and rate-limit notes. Slightly informal/mixed-language phrasing but no wasted filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the annotations cover the safety profile. The description supplies the missing rate-limit caveat and engine fallback, so an agent has enough to invoke it correctly; only sibling differentiation and engine enum values are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden and mostly does: currency examples and the empty-string default, plus engine behavior ('http' cheap first, browser for JS-heavy pages). It never enumerates the valid engine values, leaving some ambiguity for a 0%-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource: Bank Indonesia transaction rates (buy/sell) from the official BI table. Clear verb+resource, though it never explicitly distinguishes itself from sibling google_fx_rate, which an agent might reasonably confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides operational guidance (wait or switch to an RSS-based tool when blocked) but gives no explicit when-to-use-this vs google_fx_rate or other rate tools. Currency semantics ('kosong = semua') are implied usage rather than selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_news_homepageGoogle News HomepageA
Read-onlyIdempotent

Scrape Google News homepage RSS: top headlines right now.

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
glNoUS
hlNoen-US
ceidNoUS:en

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the read-only, idempotent, non-destructive, open-world profile. The description adds a trait annotations cannot express: Google rate-limits this endpoint, so rapid repeats may be blocked and recovery requires waiting or switching tools. That is materially useful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, purpose front-loaded, no filler. The second sentence is dedicated to the one non-obvious operational risk, so every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values and annotations cover safety, so those are not needed here. The remaining gap is the unexplained locale parameters, which matter for correct invocation of a region-varying feed but are absent from both the description and the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three locale parameters (gl, hl, ceid) have 0% schema description coverage, and the description says nothing about them. For a region/language-sensitive endpoint, the agent gets no explanation of what gl, hl, or ceid control or how they interact, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scrape Google News homepage RSS') plus the payload ('top headlines right now'), which cleanly distinguishes it from the sibling google_news_search by scope (homepage feed vs. query). It does not name that sibling explicitly, so differentiation is implicit rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a conditional fallback ('when the result is blocked, wait or switch to an RSS-based tool'), which is genuine usage guidance. However it never states when to prefer this over google_news_search, and the 'RSS-based tool' alternative is left unnamed, so routing remains inferential.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_scholar_cited_byGoogle Scholar Cited ByB
Read-onlyIdempotent

Scrape papers citing a Scholar cluster (from cluster_id in scholar results).

engine: auto | http | proxy | browser (lihat google_web_search).

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
hlNoen
numNo
startNo
engineNoauto
cluster_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), and the description adds real behavioral context beyond them: the engine modes and the fact that Google rate-limits these calls and may return blocked results. It does not describe pagination behavior or result volume, but the added characteristics are substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then engine options and the rate-limit caveat, with no filler. The mixed-language cross-reference '(lihat google_web_search)' is slightly opaque to an English reader but costs little space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the rate-limit warning is a genuine completion of the picture. What is missing is any explanation of the pagination parameters an agent must set to page through citing papers, which is a meaningful gap for a scraper tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden, and it only partially does: it explains cluster_id's source and lists engine options. The pagination parameters (num, start) and hl are never explained in either place, leaving three of five parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('scrape') and resource ('papers citing a Scholar cluster') and clarifies the input's origin ('cluster_id from scholar results'), which cleanly separates it from google_scholar_search. It stops short of explicitly naming that sibling as the tool to call first, so it is clear but not fully differentiated in prose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives implied context via the parenthetical pointing to cluster_id from scholar results, and the rate-limit note ('avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool') is useful operating guidance. However, it never states explicitly that you must run google_scholar_search first, nor names a concrete sibling alternative for fallback.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_statusGoogle StatusA
Read-onlyIdempotent

Health check: which Google endpoints are reachable from this IP right now, plus proxy pool, cache, cooldowns and forensic counters.

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds real value beyond that: Google-side rate limiting, the meaning of a blocked result, remediation steps, and the diagnostic payload returned (proxy pool, cache, cooldowns, counters).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the first front-loads what the tool reports, the second front-loads the operational constraint and remedy. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema already present, the description only needs to convey purpose, rate-limit behavior, and failure handling - all of which it does. An agent has everything required to call it and interpret a blocked result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters and the schema is fully covered, so there is nothing for the description to clarify. Baseline applies; the sentence about blocked results usefully explains result-state semantics rather than argument semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: a health check of which Google endpoints are reachable from the current IP, and enumerates the returned diagnostics (proxy pool, cache, cooldowns, forensic counters). This cleanly distinguishes it from every search/crawl sibling in the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-not-to-do guidance - 'avoid rapid repeats' and, when blocked, 'wait or switch to an RSS-based tool' - which names a fallback class. It lacks an explicit positive trigger (e.g. 'call this first to diagnose reachability before running a search'), so it stops just short of a full when/when-not pairing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_suggestGoogle SuggestA
Read-onlyIdempotent

Google autocomplete suggestions (no block).

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
hlNoen
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld, so the bar is lower, and the description still adds genuinely useful context: Google rate-limits this endpoint and results can come back blocked, with a recommended recovery path. It does not describe the shape or reliability of results beyond the blocking caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the capability front-loaded and the operational caveat second; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, but with 0% schema coverage the description leaves both parameters undefined and never clarifies what 'no block' means. Adequate for the operation itself, incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, and the description says nothing about either. In particular 'hl' (default 'en') is never identified as a language code nor is the expected format given for 'query', so the description does not compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (Google autocomplete suggestions) with a qualifier, so an agent can tell it apart from the many sibling search tools. The parenthetical '(no block)' is cryptic and not explained anywhere, which slightly muddies the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives behavioral usage guidance ('avoid rapid repeats', 'when blocked, wait or switch to an RSS-based tool'), which implies when the tool fails and what to do. However, it never says when to prefer this autocomplete tool over sibling search tools like google_search or google_web_search, so selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

google_translateGoogle TranslateA
Read-onlyIdempotent

Translate text via Google's unofficial gtx endpoint (no key).

Rate-limited by Google: avoid rapid repeats; when the result is blocked, wait or switch to an RSS-based tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
sourceNoauto
targetNoen

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower; the description adds genuinely useful context by disclosing that the endpoint is unofficial, keyless, rate-limited, and can return blocked results. It stops short of describing retry timing or response shape, but that is a modest gap given the annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the core purpose and the endpoint caveat are front-loaded before the rate-limit warning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the failure behavior is covered. However, with three parameters at 0% schema coverage, the description is not complete enough for an agent to know how to fill source/target correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description says nothing about text, source, or target. It never explains that source defaults to 'auto' or that target defaults to 'en', nor the expected language-code format, so the undocumented parameters are left entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Translate) and resource (text) plus the underlying mechanism (Google's unofficial gtx endpoint, no key). No sibling does translation, so it is trivially distinguishable from the search/finance/trends tools listed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives operational guidance: avoid rapid repeats, and when results are blocked, wait or switch to an RSS-based tool. This is clear failure-mode routing, but it never says when to prefer this over any listed sibling or what to do for other translation needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.8.0
    • Addedgoogle_fx_rate
    • Addedgoogle_kurs_bi
  2. 20 tool updatesv0.7.2
    • First observedgoogle_ai_mode
    • First observedgoogle_books_search
    • First observedgoogle_crawl
    • First observedgoogle_finance_quote
    • First observedgoogle_help
    • First observedgoogle_image_search
    • First observedgoogle_news_homepage
    • First observedgoogle_news_search
    • First observedgoogle_patents_search
    • First observedgoogle_scholar_cited_by
    • First observedgoogle_scholar_search
    • First observedgoogle_search
    • First observedgoogle_shopping_search
    • First observedgoogle_status
    • First observedgoogle_suggest
    • First observedgoogle_translate
    • First observedgoogle_trends_daily
    • First observedgoogle_trends_interest
    • First observedgoogle_video_search
    • First observedgoogle_web_search

TDQS

B3.4/5.0

Scored across 22 tools

Disambiguation2/5

Many tools duplicate functionality: google_search can perform all the specialized searches (web, images, news, etc.), so it's unclear when to use a specialized tool versus the unified one. Descriptions mention engine and rate limits but don't resolve the core overlap.

Naming Consistency4/5

All tools share the google_ prefix and snake_case, which is consistent. However, patterns vary (some are verb-based like google_translate, others are noun-based like google_finance_quote), so minor deviations exist.

Tool Count3/5

22 tools is on the high end for a single server; while many verticals are distinct, the unified search tool makes several redundant, leading to over-provisioning.

Completeness4/5

The surface covers most major Google search verticals (web, images, video, news, books, shopping, scholar, patents, finance, trends, AI) plus utility tools. Missing verticals like maps, jobs, or local could be minor gaps for some use cases, but core scraping needs are met.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Enables free web searching using Google search results with no API keys required, returning structured results with titles, URLs, and descriptions.
    1
    32 npm
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to perform real-time Google searches with anti-bot protection. Bypasses search engine restrictions using advanced browser automation to extract search results locally without requiring paid API services.
    6
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to perform Google searches and retrieve structured JSON results for 10 types including web, images, news, and shopping with live prices. Offers 1,000 free searches per month with no credit card required.
    1
    38 npm
    MIT