google-scrape-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| GOOGLE_SCRAPE_PROXIES | No | inline list (comma/space/newline separated) or a file path (starting with `/`, `~`, or ending in `.txt`) | |
| GOOGLE_SCRAPE_RETRIES | No | proxy attempts per request | 3 |
| GOOGLE_SCRAPE_TIMEOUT | No | request timeout (s) | 25 |
| GOOGLE_SCRAPE_PROXY_CA | No | extra CA bundle for self-signed HTTPS proxies | |
| GOOGLE_SCRAPE_CACHE_MAX | No | max cache entries | 256 |
| GOOGLE_SCRAPE_CACHE_TTL | No | response cache TTL (`0` disables) | 600 |
| GOOGLE_SCRAPE_FORENSICS | No | save blocked-page samples | 1 |
| GOOGLE_SCRAPE_COOKIE_TTL | No | bootstrap cookie lifetime (s) | 300 |
| GOOGLE_SCRAPE_PROXY_MODE | No | auto (direct→proxy), off, always | auto |
| GOOGLE_SCRAPE_IMPERSONATE | No | curl_cffi TLS target | chrome120 |
| GOOGLE_SCRAPE_JS_COOLDOWN | No | base HTTP cooldown after a block (doubles per level, capped at 1h) | 600 |
| GOOGLE_SCRAPE_NO_SELFHEAL | No | disable automatic `camoufox fetch` when the browser is missing | |
| GOOGLE_SCRAPE_PROXY_FILES | No | colon-separated file paths | |
| GOOGLE_SCRAPE_MIN_INTERVAL | No | min delay between requests (s) | 1.0 |
| GOOGLE_SCRAPE_PROXY_COOLDOWN | No | failed-proxy cooldown (s) | 600 |
| GOOGLE_SCRAPE_BROWSER_COOLDOWN | No | base browser cooldown after a block (doubles per level) | 180 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| logging | {} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| google_web_searchB | Scrape Google Web Search: organic results, featured snippet, related searches. engine: auto (HTTP cepat, fallback browser bila diblokir) | http (tanpa browser) | proxy (paksa pool proxy) | browser (langsung Camoufox headless, tanpa proxy/API key). |
| google_image_searchC | Scrape Google Images (tbm=isch): direct image URLs, page URLs, sizes. engine: auto | http | proxy | browser (lihat google_web_search). |
| google_video_searchC | Scrape Google Video Search (tbm=vid): title, URL, snippet. engine: auto | http | proxy | browser (lihat google_web_search). |
| google_books_searchC | Scrape Google Books Search (tbm=bks): title, URL, snippet. engine: auto | http | proxy | browser (lihat google_web_search). |
| google_shopping_searchC | Scrape Google Shopping (tbm=shop): product title, URL, snippet. engine: auto | http | proxy | browser (lihat google_web_search). |
| google_news_searchB | Scrape Google News: RSS dulu (no block); fallback tab News renderan bila RSS gagal. engine: auto | http (RSS saja) | proxy | browser (tab News via Camoufox). |
| google_news_homepageC | Scrape Google News homepage RSS: top headlines right now. |
| google_scholar_searchC | Scrape Google Scholar: papers, authors/venue, citations, PDF links. start: 0-based result offset (unified google_search computes it from page). engine: auto | http | proxy | browser (lihat google_web_search). |
| google_scholar_cited_byC | Scrape papers citing a Scholar cluster (from cluster_id in scholar results). engine: auto | http | proxy | browser (lihat google_web_search). |
| google_patents_searchC | Scrape Google Patents XHR JSON: title, number, inventors, dates, PDF. page: 1-based (1 = first page). NOTE: page number lives INSIDE the url= query string (page=N), not as an outer param. |
| google_finance_quoteB | Scrape Google Finance quote page: price, prev close, day high/low, change. |
| google_translateB | Translate text via Google's unofficial gtx endpoint (no key). |
| google_suggestC | Google autocomplete suggestions (no block). |
| google_trends_dailyC | Google daily trending searches RSS (no block): topic, traffic, links. |
| google_trends_interestA | Google Trends interest-over-time via internal explore API (no key). keywords: comma-separated, max 5 (e.g. "opencode, cursor, windsurf"). timeframe: e.g. 'now 7-d', 'today 12-m', 'today 5-y', 'all'. engine: auto (HTTP → fallback browser) | http (HTTP saja) | browser. NOTE: endpoint widgetdata Trends sering menolak request non-browser (HTTP 400/401) walau token explore valid — bila itu terjadi, engine auto memakai Camoufox: fetch dijalankan dari dalam halaman Trends sehingga cookies/fingerprint ikut. Bila tetap gagal, status "limited". |
| google_ai_modeA | Google Mode AI (udm=50): jawaban sintesis + sumber, cocok untuk pertanyaan langsung. Bentuk hasil: answer + sources. engine: auto | http | proxy | browser (lihat google_web_search). Halaman Mode AI hampir selalu butuh JS; engine auto akan memakai browser. |
| google_searchB | One search tool for all Google tabs + pagination (pages 1-10). tab: web | images | videos | news | books | shopping | scholar | patents | ai. page: 1-10 (each page = next result set; news slices the RSS feed). num: results per page. engine: auto (HTTP cepat, fallback browser bila diblokir) | http (tanpa browser) | proxy (paksa pool proxy) | browser (langsung Camoufox). |
| google_crawlC | Crawl any URL (e.g. a search result): title, meta, text, outbound links. |
| google_helpA | Agent-facing usage guide: tool map, engine semantics, status contract and safe usage patterns. Read this before experimenting. |
| google_statusA | Health check: which Google endpoints are reachable from this IP right now, plus proxy pool, cache, cooldowns and forensic counters. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 20 tools
google_search promises one tool for all Google tabs, yet separate web/image/video/books/shopping/news/scholar/patents search tools duplicate those same tabs, creating unclear boundaries. Descriptions explain some distinctions, but an agent can easily misselect between the unified and specialized tools.
All names use a google_ prefix with snake_case, following a consistent service_action or service_property pattern. The minor variation in action vocabulary is predictable and readable.
20 tools is borderline heavy. The redundancy between unified google_search and specialized vertical search tools means several tools may not earn their place, though the breadth of Google properties accounts for many.
Covers major Google verticals including web, images, video, books, shopping, news, scholar, patents, trends, finance, and translate, plus crawl/help/status. Some Google properties like Maps, Jobs, and Flights are missing, but the core scraping surface is strong.