Skip to main content
Glama
Maybeyes111

google-scrape-mcp

by Maybeyes111

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
GOOGLE_SCRAPE_PROXIESNoinline list (comma/space/newline separated) or a file path (starting with `/`, `~`, or ending in `.txt`)
GOOGLE_SCRAPE_RETRIESNoproxy attempts per request3
GOOGLE_SCRAPE_TIMEOUTNorequest timeout (s)25
GOOGLE_SCRAPE_PROXY_CANoextra CA bundle for self-signed HTTPS proxies
GOOGLE_SCRAPE_CACHE_MAXNomax cache entries256
GOOGLE_SCRAPE_CACHE_TTLNoresponse cache TTL (`0` disables)600
GOOGLE_SCRAPE_FORENSICSNosave blocked-page samples1
GOOGLE_SCRAPE_COOKIE_TTLNobootstrap cookie lifetime (s)300
GOOGLE_SCRAPE_PROXY_MODENoauto (direct→proxy), off, alwaysauto
GOOGLE_SCRAPE_IMPERSONATENocurl_cffi TLS targetchrome120
GOOGLE_SCRAPE_JS_COOLDOWNNobase HTTP cooldown after a block (doubles per level, capped at 1h)600
GOOGLE_SCRAPE_NO_SELFHEALNodisable automatic `camoufox fetch` when the browser is missing
GOOGLE_SCRAPE_PROXY_FILESNocolon-separated file paths
GOOGLE_SCRAPE_MIN_INTERVALNomin delay between requests (s)1.0
GOOGLE_SCRAPE_PROXY_COOLDOWNNofailed-proxy cooldown (s)600
GOOGLE_SCRAPE_BROWSER_COOLDOWNNobase browser cooldown after a block (doubles per level)180

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
logging
{}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}

Tools

Functions exposed to the LLM to take actions

NameDescription
google_web_searchB

Scrape Google Web Search: organic results, featured snippet, related searches.

engine: auto (HTTP cepat, fallback browser bila diblokir) | http (tanpa browser) | proxy (paksa pool proxy) | browser (langsung Camoufox headless, tanpa proxy/API key).

google_image_searchC

Scrape Google Images (tbm=isch): direct image URLs, page URLs, sizes.

engine: auto | http | proxy | browser (lihat google_web_search).

google_video_searchC

Scrape Google Video Search (tbm=vid): title, URL, snippet.

engine: auto | http | proxy | browser (lihat google_web_search).

google_books_searchC

Scrape Google Books Search (tbm=bks): title, URL, snippet.

engine: auto | http | proxy | browser (lihat google_web_search).

google_shopping_searchC

Scrape Google Shopping (tbm=shop): product title, URL, snippet.

engine: auto | http | proxy | browser (lihat google_web_search).

google_news_searchB

Scrape Google News: RSS dulu (no block); fallback tab News renderan bila RSS gagal.

engine: auto | http (RSS saja) | proxy | browser (tab News via Camoufox).

google_news_homepageC

Scrape Google News homepage RSS: top headlines right now.

google_scholar_searchC

Scrape Google Scholar: papers, authors/venue, citations, PDF links.

start: 0-based result offset (unified google_search computes it from page). engine: auto | http | proxy | browser (lihat google_web_search).

google_scholar_cited_byC

Scrape papers citing a Scholar cluster (from cluster_id in scholar results).

engine: auto | http | proxy | browser (lihat google_web_search).

google_patents_searchC

Scrape Google Patents XHR JSON: title, number, inventors, dates, PDF.

page: 1-based (1 = first page). NOTE: page number lives INSIDE the url= query string (page=N), not as an outer param.

google_finance_quoteB

Scrape Google Finance quote page: price, prev close, day high/low, change.

google_translateB

Translate text via Google's unofficial gtx endpoint (no key).

google_suggestC

Google autocomplete suggestions (no block).

google_trends_dailyC

Google daily trending searches RSS (no block): topic, traffic, links.

google_trends_interestA

Google Trends interest-over-time via internal explore API (no key).

keywords: comma-separated, max 5 (e.g. "opencode, cursor, windsurf"). timeframe: e.g. 'now 7-d', 'today 12-m', 'today 5-y', 'all'. engine: auto (HTTP → fallback browser) | http (HTTP saja) | browser.

NOTE: endpoint widgetdata Trends sering menolak request non-browser (HTTP 400/401) walau token explore valid — bila itu terjadi, engine auto memakai Camoufox: fetch dijalankan dari dalam halaman Trends sehingga cookies/fingerprint ikut. Bila tetap gagal, status "limited".

google_ai_modeA

Google Mode AI (udm=50): jawaban sintesis + sumber, cocok untuk pertanyaan langsung. Bentuk hasil: answer + sources.

engine: auto | http | proxy | browser (lihat google_web_search). Halaman Mode AI hampir selalu butuh JS; engine auto akan memakai browser.

google_searchB

One search tool for all Google tabs + pagination (pages 1-10).

tab: web | images | videos | news | books | shopping | scholar | patents | ai. page: 1-10 (each page = next result set; news slices the RSS feed). num: results per page. engine: auto (HTTP cepat, fallback browser bila diblokir) | http (tanpa browser) | proxy (paksa pool proxy) | browser (langsung Camoufox).

google_crawlC

Crawl any URL (e.g. a search result): title, meta, text, outbound links.

google_helpA

Agent-facing usage guide: tool map, engine semantics, status contract and safe usage patterns. Read this before experimenting.

google_statusA

Health check: which Google endpoints are reachable from this IP right now, plus proxy pool, cache, cooldowns and forensic counters.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3/5.0

Scored across 20 tools

Disambiguation2/5

google_search promises one tool for all Google tabs, yet separate web/image/video/books/shopping/news/scholar/patents search tools duplicate those same tabs, creating unclear boundaries. Descriptions explain some distinctions, but an agent can easily misselect between the unified and specialized tools.

Naming Consistency5/5

All names use a google_ prefix with snake_case, following a consistent service_action or service_property pattern. The minor variation in action vocabulary is predictable and readable.

Tool Count3/5

20 tools is borderline heavy. The redundancy between unified google_search and specialized vertical search tools means several tools may not earn their place, though the breadth of Google properties accounts for many.

Completeness4/5

Covers major Google verticals including web, images, video, books, shopping, news, scholar, patents, trends, finance, and translate, plus crawl/help/status. Some Google properties like Maps, Jobs, and Flights are missing, but the core scraping surface is strong.

Maintenance

ActivityMaintained
ResponsivenessNo issues