InfinityScrape MCP
Scans Codeforces for public user profiles and competitive programming activity as part of multi-platform OSINT reconnaissance.
Scans Dev.to for public developer profiles and posts as part of multi-platform OSINT reconnaissance.
Provides real-time web searches via DuckDuckGo, including text search and automatic scraping of top results.
Scans GitHub for public user, team, and repository information as part of multi-platform OSINT reconnaissance.
Scans GitLab for public user, group, and project information as part of multi-platform OSINT reconnaissance.
Scans Google Scholar for scholarly profiles and publications as part of multi-platform OSINT reconnaissance.
Scans Kaggle for public user profiles and dataset/competition activity as part of multi-platform OSINT reconnaissance.
Scans LeetCode for public user profiles and coding activity as part of multi-platform OSINT reconnaissance.
Scans Medium for public author profiles and articles as part of multi-platform OSINT reconnaissance.
Resolves addresses to GPS coordinates, street/postcode details, and administrative boundaries using OpenStreetMap.
Scans Reddit for public user profiles and activity as part of multi-platform OSINT reconnaissance.
Scans ResearchGate for public researcher profiles and publications as part of multi-platform OSINT reconnaissance.
Scans Substack for public author profiles and newsletters as part of multi-platform OSINT reconnaissance.
Extracts YouTube video, Shorts, and live transcripts with timestamps directly over HTTP, without requiring video downloads or local GPU models.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@InfinityScrape MCPScrape the content from https://example.com and summarize it, bypassing anti-bot checks."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ InfinityScrape MCP: World-Class Web Scraping, Dynamic SPA Rendering & 25-Tool OSINT Intelligence Suite
InfinityScrape MCP is a standalone, production-grade Model Context Protocol (MCP) server engineered to provide AI models (Open WebUI, Claude 3.7, DeepSeek-R1/V3, Antigravity AI, Cursor, LM Studio) with unlimited, high-speed, anti-bot resilient web scraping, dynamic SPA rendering, DuckDuckGo web search, Wayback Machine time-travel, instant YouTube transcription, and precision OSINT / GEOINT location intelligence.
๐ Table of Contents
Related MCP server: FineData MCP Server
๐ Why InfinityScrape MCP?
Standard web scrapers often fail on modern websites due to Cloudflare challenges, heavy client-side JavaScript rendering, intrusive cookie consent modals, and rate limits. InfinityScrape solves these problems out-of-the-box:
Dual-Engine Scraping Architecture:
Fast TLS Engine (
primp+httpx): Mimics real Chrome/Safari browser TLS/JA3 fingerprints and HTTP/2 headers to bypass Cloudflare and Akamai challenges in<100ms.Dynamic Headless Browser (
Playwright Chromium): Renders complex SPAs (React, Vue, Next.js, Angular), performs infinite scrolling, clicks elements, and executes custom JavaScript.
Singleton
BrowserPool(~1.8s SPA Renders):Persistent Chromium process with ephemeral context isolation that completely eliminates cold-start latency and avoids memory leaks.
Military-Grade Anti-Bot Evasions (100% Undetectable):
Cleanly deletes
navigator.webdriverfrom prototype (get: () => undefined).Authentic WebGL Vendor & Renderer spoofing (
Google Inc. (NVIDIA)&NVIDIA GeForce RTX 3080).Hardware concurrency (8 cores), device memory (8GB), and authentic
window.chromeruntime emulation.Realistic Chrome PDF Viewer plugins and NaCl mimeTypes.
3-Tier Multi-Engine Search Failover (Anti-429 Resilience):
Cascades automatically from
DuckDuckGo APIโก๏ธDuckDuckGo HTML Liteโก๏ธBing HTML Fallback. Zero API keys, zero 429 rate-limits.
Cloudflare Turnstile & Interstitial Auto-Solver:
Detects Turnstile challenge iframes and executes automated coordinate jitter to bypass interstitials.
Network-Level Ad & Tracker Elimination:
Intercepts and aborts network calls to 35+ ad networks and tracking scripts (
doubleclick,criteo,outbrain,google-analytics) before they download, cutting page load time by ~300% and memory usage by 70%.Automatically detects and decomposes OneTrust, Cookiebot, and sticky overlay popups.
Wayback Machine Time-Travel:
Query internet archive history for any URL across custom date ranges to track competitor pricing changes, deleted pages, and historical copy.
Zero-GPU Instant YouTube Transcriber:
Extracts complete video/shorts/live transcripts with timestamps (
[MM:SS]) in<300msdirectly via HTTP streams without downloading video or requiring local GPU Whisper models.
Deep Recursive Documentation Crawler:
Asynchronous Breadth-First-Search (BFS) crawler with domain locking and path prefix filtering to aggregate entire documentation trees into unified Markdown.
State-of-the-Art Public OSINT & GEOINT Reconnaissance:
Multi-Signal Confidence Scoring (0% - 100%): Evaluates Name + City + Street + PIN + Org + Role correlation to rank discovered dossiers.
25+ Global Platform Scanners: Scans GitHub, GitLab, StackOverflow, Kaggle, HuggingFace, LeetCode, Codeforces, Dev.to, Medium, Substack, Google Scholar, ResearchGate, Reddit, etc.
OpenStreetMap GEOINT: Resolves global addresses down to street/postcode level with GPS coordinates and administrative boundaries.
SQLite Persistent Caching Layer:
In-memory and SQLite-backed local cache for instant
0msresponses on repeat lookups with configurable TTL.
โก Competitive Comparison
Feature / Capability | Standard MCP Scrapers | Cloud Scraping APIs | InfinityScrape MCP |
Cost & API Keys | Free (Basic) | Paid ($20 - $200/mo) | 100% Free / Zero API Keys |
Cloudflare / Akamai TLS Bypass | โ Fails / 403 | โ Yes | โ
Built-in ( |
Dynamic SPAs & Infinite Scroll | โ Limited | โ Yes | โ
Built-in ( |
Real-Time Web Search & Dorking | โ No | โ ๏ธ Extra Cost | โ Built-in (DuckDuckGo & Dorks) |
Wayback Historical Snapshots | โ No | โ No | โ Built-in (Archive API) |
Network-Level Ad & Popup Stripping | โ No | โ ๏ธ Partial | โ Built-in (35+ domains) |
Zero-GPU YouTube Transcripts | โ No | โ No | โ Built-in (<300ms) |
Online PDF Page-by-Page Parser | โ No | โ ๏ธ Extra Cost | โ
Built-in ( |
Deep Documentation Crawler | โ No | โ ๏ธ Extra Cost | โ Built-in (Async BFS) |
25+ Platform OSINT & Geocoding | โ No | โ No | โ Built-in (0-100% Confidence) |
OpenAPI 3.1.0 REST Bridge (Port 8000) | โ No | โ ๏ธ Proprietary | โ Built-in (FastAPI /docs) |
๐๏ธ Architectural Overview
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ AI Clients: Open WebUI / Claude Desktop / Cursor โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ โ
โผ โผ
[OpenAPI Bridge (Port 8000)] [Stdio JSON-RPC 2.0 Server]
FastAPI /docs & /openapi.json (server.py)
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโ
โผ โผ โผ โผ
[Fast TLS Engine] [Playwright Engine] [OSINT / GEOINT] [Search & Media]
โข primp JA3/TLS โข Stealth Chromium โข 25+ Platform โข DuckDuckGo Search
โข HTTP/2 Headers โข Ad/Tracker Blocker Scanners โข Wayback Snapshots
โข <100ms Execution โข Infinite Scroll โข OpenStreetMap โข YouTube (<300ms)
โข Auto-Dismiss CMPs โข Reverse Geocoding โข Remote PDF Parser
โข Match Confidence
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SQLite Caching Layer (0ms) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ๐ Quick Start & 1-Click Installation
1. Automated Setup
# Clone the repository
git clone https://github.com/virajverse/infinity-scraper.git
cd infinity-scraper
# Create virtual environment & install
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
pip install -e .
playwright install chromium2. Launch FastAPI Bridge (Port 8000)
python openapi_bridge.pyInteractive Swagger Docs:
http://127.0.0.1:8000/docsOpenAPI 3.1.0 Schema:
http://127.0.0.1:8000/openapi.json
๐ AI Client Integration
1. Open WebUI (FastAPI Bridge on Port 8000)
Ensure the bridge is running (
python openapi_bridge.py).In Open WebUI, navigate to Workspace -> Tools -> Add Tool.
Import from URL:
http://127.0.0.1:8000/openapi.jsonor useinfinity_scraper_suite.All 25 tools are instantly accessible to your agents!
2. Antigravity AI / Claude Desktop (Native Stdio)
Add to your mcp_config.json:
{
"mcpServers": {
"infinity-scraper": {
"command": "python",
"args": ["-m", "infinity_scraper.server"],
"env": {
"PYTHONUNBUFFERED": "1"
}
}
}
}๐ ๏ธ Complete 25-Tool Reference Catalog
1. Anti-Bot Web Scraping, Dynamic SPAs & Crawlers (6 Tools)
Tool | Description |
| Production-grade scraping into clean, ad-free Markdown with auto-engine switching (Fast TLS -> Playwright Chromium fallback). |
| Dynamic SPA rendering via Singleton Playwright BrowserPool (~1.8s) with Cloudflare Turnstile auto-bypass and infinite scroll. |
| Zero-selector semantic extraction mapping custom JSON schemas with auto JSON-LD & OpenGraph meta fallback. |
| Asynchronous recursive BFS documentation crawler with depth limits, domain locking, and path prefix filters. |
| Extracts HTML tables into structured Markdown/JSON datasets and aggregates rich page metadata. |
| High-throughput concurrent scraping of multiple URLs with configurable concurrency limits. |
2. Real-Time Web Search, Research & Archive OSINT (4 Tools)
Tool | Description |
| 3-tier resilient real-time web search (DDGS API โก๏ธ DDG HTML Lite โก๏ธ Bing HTML Fallback) without API keys or 429 rate limits. |
| End-to-end autonomous web research pipeline: executes queries and scrapes top results into a synthesized research report. |
| Advanced boolean dorking engine supporting |
| Queries historical archive snapshots, CDX timestamps, and past versions of deleted or updated web pages. |
3. Media, Document & RAG Extraction (4 Tools)
Tool | Description |
| Streams remote online PDF files page-by-page via TLS into clean Markdown text without local disk bloat. |
| Zero-GPU, sub-300ms transcript extraction with timestamps ( |
| Extracts camera hardware specifications, timestamps, and GPS coordinates with direct Google Maps navigation links. |
| Semantic text and Markdown chunker with token boundary optimization for LLM RAG pipelines and vector stores. |
4. Deep Public OSINT & Entity Reconnaissance (4 Tools)
Tool | Description |
| Cross-correlates developer platforms, academic registries, and corporate filings into an entity intelligence dossier. |
| Scans 25+ global developer, creator, and tech platforms to map digital handles and alias footprints. |
| Extracts Reddit discussions, original post content, and nested comment trees into structured Markdown. |
| Ingests RSS and Atom feeds for real-time news tracking, competitor updates, and blog monitoring. |
5. Infrastructure, Domain & Network Reconnaissance (5 Tools)
Tool | Description |
| Audits domain WHOIS, RDAP records, registrar info, and SSL/TLS certificate chains. |
| Deeply fingerprints website frontend frameworks, backend stacks, CDNs, and analytics trackers. |
| Gathers IP geolocation, Autonomous System Number (ASN), ISP, and network routing data. |
| Discovers hidden subdomains, internal staging servers, and API routes via Certificate Transparency logs (crt.sh). |
| Resolves and audits DNS records (A, AAAA, MX, TXT, NS, CNAME) including mail security SPF/DKIM verification. |
6. Precision GEOINT & Location Intelligence (2 Tools)
Tool | Description |
| Forward geocoding of landmarks, streets, and addresses via OpenStreetMap/Nominatim down to postal code and GPS coordinates. |
| Hierarchical entity location dorking correlating business entities with city, street, and postal landmarks. |
๐ป Command-Line Interface (CLI)
InfinityScrape provides a built-in CLI for quick terminal testing:
# Scrape a webpage into Markdown
infinity-scrape scrape "https://news.ycombinator.com" --format markdown
# Search DuckDuckGo from the terminal
infinity-scrape search "Generative Engine Optimization 2026" --limit 5
# Extract YouTube Transcript
infinity-scrape youtube "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
# OSINT Persona Lookup
infinity-scrape osint --name "Linus Torvalds" --platforms github,gitlab๐งช Running Automated Tests
# Run unit and integration tests
pytest tests/ -v๐ License & Authors
This server cannot be deployed
Maintenance
Related MCP Connectors
Direct access to 60+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.
- GoroOAuthai.usegoro
62 real-world tools for agents: search, scraping, social, enrichment, image, video, voice.
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
Live web access for agents: scrape, SERP search, crawl/map, 74 collectors, datasets, proxies.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to perform undetectable browser automation that bypasses Cloudflare, antibots, and social media blocks. Provides 105 tools for element extraction, network debugging, and real-world web scraping with a 98.7% success rate on protected sites.2,060MIT
- AlicenseAqualityBmaintenanceEnables AI agents to scrape any website by providing tools for JavaScript rendering, antibot bypass, and automatic captcha solving. It supports synchronous, asynchronous, and batch scraping operations with built-in proxy rotation.5207 npmMIT

ScrapeLab MCPofficial
AlicenseNot gradedqualityDmaintenanceEnables undetectable web scraping and browser automation for AI agents with 84 tools including stealth navigation, element extraction, network interception, and auto cookie consent dismissal. Bypasses anti-bot systems like Cloudflare and DataDome while providing LLM-ready markdown output and full Chrome DevTools Protocol access.MIT
HasData MCP Serverofficial
AlicenseAqualityAmaintenanceDirect access to 40+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.8636MIT