Skip to main content
Glama

๐ŸŒ InfinityScrape MCP: World-Class Web Scraping & Deep OSINT Intelligence Suite

License: MIT Python 3.10+ Protocol: MCP Zero-Cloud-API Zero-GPU

InfinityScrape MCP is a standalone, production-grade Model Context Protocol (MCP) server engineered to provide AI models (LM Studio, Claude Desktop, Cursor, Open WebUI, Antigravity AI) with unlimited, high-speed, anti-bot resilient web scraping, dynamic SPA rendering, instant YouTube transcription, and precision OSINT / GEOINT location intelligence.


๐Ÿ“‘ Table of Contents


Related MCP server: FineData MCP Server

๐ŸŒŸ Why InfinityScrape MCP?

Standard web scrapers often fail on modern websites due to Cloudflare challenges, heavy client-side JavaScript rendering, intrusive cookie consent modals, and rate limits. InfinityScrape solves these problems out-of-the-box:

  1. Dual-Engine Architecture:

    • Fast TLS Engine (primp + httpx): Mimics real Chrome/Safari browser TLS/JA3 fingerprints and HTTP/2 headers to bypass Cloudflare and Akamai challenges in <100ms.

    • Dynamic Headless Browser (Playwright Chromium): Renders complex SPAs (React, Vue, Next.js, Angular), performs infinite scrolling, clicks elements, and executes custom JavaScript.

  2. Network-Level Ad & Tracker Elimination:

    • Intercepts and aborts network calls to 35+ ad networks and tracking scripts (doubleclick, criteo, outbrain, google-analytics) before they download, cutting page load time by ~300% and memory usage by 70%.

    • Automatically detects and decomposes OneTrust, Cookiebot, and sticky overlay popups.

  3. Zero-GPU Instant YouTube Transcriber:

    • Extracts complete video/shorts/live transcripts with timestamps ([MM:SS]) in <300ms directly via HTTP streams without downloading video or requiring local GPU Whisper models.

  4. Deep Recursive Documentation Crawler:

    • Asynchronous Breadth-First-Search (BFS) crawler with domain locking and path prefix filtering to aggregate entire documentation trees into unified Markdown.

  5. State-of-the-Art Public OSINT & GEOINT Reconnaissance:

    • Multi-Signal Confidence Scoring (0% - 100%): Evaluates Name + City + Street + PIN + Org + Role correlation to rank discovered dossiers.

    • 25+ Global Platform Scanners: Scans GitHub, GitLab, StackOverflow, Kaggle, HuggingFace, LeetCode, Codeforces, Dev.to, Medium, Substack, Google Scholar, ResearchGate, Reddit, etc.

    • OpenStreetMap GEOINT: Resolves global addresses down to street/postcode level with GPS coordinates and administrative boundaries.

  6. SQLite Persistent Caching Layer:

    • In-memory and SQLite-backed local cache for instant 0ms responses on repeat lookups with configurable TTL.


โšก Competitive Comparison

Feature / Capability

Standard MCP Scrapers

Cloud Scraping APIs

InfinityScrape MCP

Cost & API Keys

Free (Basic)

Paid ($20 - $200/mo)

100% Free / Zero API Keys

Cloudflare / Akamai TLS Bypass

โŒ Fails / 403

โœ… Yes

โœ… Built-in (primp JA3)

Dynamic SPAs & Infinite Scroll

โŒ Limited

โœ… Yes

โœ… Built-in (playwright)

Network-Level Ad & Popup Stripping

โŒ No

โš ๏ธ Partial

โœ… Built-in (35+ domains)

Zero-GPU YouTube Transcripts

โŒ No

โŒ No

โœ… Built-in (<300ms)

Online PDF Page-by-Page Parser

โŒ No

โš ๏ธ Extra Cost

โœ… Built-in (pypdf)

Deep Documentation Crawler

โŒ No

โš ๏ธ Extra Cost

โœ… Built-in (Async BFS)

25+ Platform OSINT & Geocoding

โŒ No

โŒ No

โœ… Built-in (0-100% Confidence)

Local SQLite 0ms Caching

โŒ No

โŒ No

โœ… Built-in (Auto TTL)


๐Ÿ—๏ธ Architectural Overview

                      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                      โ”‚    AI Client (LM Studio / Claude / Cursor)     โ”‚
                      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                              โ”‚ JSON-RPC 2.0 (Stdio)
                                              โ–ผ
                      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                      โ”‚          InfinityScrape MCP Server             โ”‚
                      โ”‚                  (server.py)                   โ”‚
                      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                              โ”‚                โ”‚               โ”‚
            โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”Œโ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
            โ–ผ                   โ–ผ     โ–ผ                 โ–ผ    โ–ผ                  โ–ผ
     [Fast TLS Engine]     [Playwright Engine]   [OSINT / GEOINT]   [Media & PDF Engines]
     โ€ข primp JA3/TLS       โ€ข Stealth Chromium    โ€ข 25+ Platform     โ€ข YouTube (<300ms)
     โ€ข HTTP/2 Stealth      โ€ข Network Ad Blocker    Scanners         โ€ข Remote PDF Stream
     โ€ข <100ms Execution    โ€ข Infinite Scroll     โ€ข OpenStreetMap    โ€ข Table Markdownify
                           โ€ข Auto-Dismiss CMPs   โ€ข Match Confidence
                                    โ”‚
                                    โ–ผ
                      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                      โ”‚ SQLite Caching Layer (0ms TTL) โ”‚
                      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿš€ Quick Start & 1-Click Installation

Prerequisites

  • Python 3.10, 3.11, or 3.12+ installed.

  • Windows, macOS, or Linux.

1-Click Setup:

On Windows:

Double-click install.bat or run in PowerShell:

.\install.bat

On Linux / macOS:

chmod +x install.sh
./install.sh

Manual Setup (Any Platform):

# 1. Create virtual environment
python -m venv .venv

# 2. Activate virtual environment
# Windows: .venv\Scripts\activate | Linux/Mac: source .venv/bin/activate

# 3. Install requirements & Playwright browser
pip install -r requirements.txt
playwright install chromium

๐Ÿ”Œ AI Client Integration

1. LM Studio (v0.3+)

Go to Settings โž” Developer โž” MCP Servers โž” Edit Config and add:

{
  "mcpServers": {
    "infinity-scraper": {
      "command": "C:/path/to/infinity-scraper/.venv/Scripts/python.exe",
      "args": [
        "-m",
        "infinity_scraper.server"
      ],
      "cwd": "C:/path/to/infinity-scraper",
      "env": {
        "PYTHONUNBUFFERED": "1"
      }
    }
  }
}

2. Claude Desktop

Edit %APPDATA%\Claude\claude_desktop_config.json (Windows) or ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):

{
  "mcpServers": {
    "infinity-scraper": {
      "command": "C:/path/to/infinity-scraper/.venv/Scripts/python.exe",
      "args": [
        "-m",
        "infinity_scraper.server"
      ]
    }
  }
}

3. Cursor IDE

In Cursor Settings โž” Features โž” MCP Servers โž” Add New MCP Server:

  • Name: infinity-scraper

  • Type: command

  • Command: C:/path/to/infinity-scraper/.venv/Scripts/python.exe -m infinity_scraper.server


๐Ÿ› ๏ธ Complete 23-Tool Reference Guide

1. Web Scraping & Content Crawling

Tool

Purpose

Key Parameters

scrape_url

Universal scraping with auto-upgrading TLS-to-Browser engine.

url, engine='auto', strip_ads=True, use_cache=True

scrape_dynamic

Headless browser for SPAs, infinite scrolls, and click actions.

url, scroll_depth=3, click_selector, wait_seconds

search_web

Real-time internet search via DuckDuckGo.

query, max_results=5, search_type='text'

search_and_scrape

Searches web and automatically scrapes top results into a cited report.

query, max_results=4

deep_crawl

Recursive async BFS site crawler for full documentation trees.

start_url, max_pages=10, max_depth=2, path_prefix

scrape_batch

Concurrently scrape multiple URLs in parallel.

urls (list), concurrency=4

extract_schema

Extract targeted fields using a CSS selector map to JSON.

url, schema={"title": "h1", "price": ".price"}

extract_structured

Extract JSON-LD, OpenGraph metadata, and HTML tables.

url, extract_tables=True

optimize_rag_chunks

Semantic RAG chunker & token optimizer for massive pages.

text_or_markdown, max_chunk_chars=2000


2. Media, Social, Video & Document Parsers

Tool

Purpose

Key Parameters

get_youtube_transcript

Zero-GPU YouTube video transcript extraction with timestamps.

url, with_timestamps=True, languages

extract_pdf

Streams and extracts remote online PDF documents page-by-page.

url, max_pages=20

extract_image_exif

Extracts camera specs, timestamps, and GPS geotags from photos.

image_url

extract_reddit_thread_tool

Ingests Reddit posts, scores, and nested comment dialogues.

url, max_comments=25

extract_rss_feed_tool

Real-time RSS/Atom feed parser for blogs, Substack, and news.

feed_url, max_items=10


3. Deep Public OSINT & Entity Reconnaissance

Tool

Purpose

Key Parameters

osint_deep_public_recon

Multi-domain open web profile scraper & confidence-ranked dossier builder.

name, location, street_or_locality, postal_code, organization, role_or_keywords, exclude_terms, time_range

osint_geoint_lookup

Global OpenStreetMap forward geocoding & administrative breakdown.

location_query, country_code

osint_location_entity_search

Hierarchical location drill-down search matrix with negative filters.

entity_name, country, city, street_or_landmark, postal_code

osint_username_check

Scans username presence across 26 coding, academic, and creative networks.

username

osint_search

Precision search dorking (site:, filetype:pdf, intitle:, -exclude).

query, site, filetype, exclude_terms


4. Technical, Domain & Network Intelligence

Tool

Purpose

Key Parameters

osint_domain_recon

Inspects domain SSL/TLS certificate validity, DNS, and RDAP/WHOIS.

domain_or_url

osint_tech_stack

Detects frontend frameworks (React, Next.js, Vue), CMS, CDN, and servers.

url

osint_ip_lookup

Public IP Geolocation, ASN, ISP, and Organization intel.

ip_or_host

osint_wayback_time_machine

Historical time-travel & deleted webpage snapshot scraper.

url, timestamp, list_snapshots

osint_subdomain_enumeration

Certificate Transparency subdomains discovery in <1 sec.

domain, limit=50

osint_dns_audit

Deep DNS records (IPv4, IPv6, MX) infrastructure audit.

domain


๐Ÿง  Autonomous AI Agent Playbook

InfinityScrape includes an advanced Cognitive Reasoning Framework (skills/infinity-scraper/SKILL.md) that teaches autonomous AI agents how to:

  • Dynamically deconstruct user prompts into search and location clues.

  • Multi-tool chain across tools (e.g. Search โž” Filter โž” Batch Scrape or Geocode โž” Locality Dork โž” Profile Extraction).

  • Auto-escalate from fast TLS to Playwright headless browser when encountering dynamic React single-page apps.

๐Ÿ‘‰ Read the full agent playbook: skills/infinity-scraper/SKILL.md


๐Ÿ’ป Command-Line Interface (CLI)

You can also use InfinityScrape directly from your terminal:

# Scrape a URL to Markdown
python -m infinity_scraper.cli scrape "https://example.com"

# Scrape dynamic SPA with infinite scroll
python -m infinity_scraper.cli scrape "https://news.ycombinator.com" --browser --scroll 3

# Live search and auto-scrape top results
python -m infinity_scraper.cli search "Quantum computing breakthroughs" --scrape --max 4

# Crawl documentation tree
python -m infinity_scraper.cli crawl "https://docs.python.org/3/library/asyncio.html" --pages 5 --depth 2

# Extract remote PDF
python -m infinity_scraper.cli pdf "https://example.com/report.pdf" --pages 10

๐Ÿงช Running Automated Tests

Run the comprehensive unit and integration test suite:

python -m tests.test_scraper

Test Coverage:

  • โœ… Fast TLS Impersonator

  • โœ… Playwright Dynamic Browser

  • โœ… DuckDuckGo Live Search

  • โœ… HTML Table to Markdown Converter

  • โœ… Recursive BFS Documentation Crawler

  • โœ… Zero-GPU YouTube Transcript Extraction

  • โœ… OSINT SSL, IP Intel & 25+ Platform Presence Check


This project is licensed under the MIT License (with Mandatory Attribution & DMCA Enforcement).

IMPORTANT

Mandatory Attribution Notice:

  • You are free to use, modify, and integrate this project for commercial or personal use.

  • However, the original Author Attribution and Copyright notice MUST be preserved in all copies, forks, or derivative distributions.

  • Removing the author's name/credits and re-uploading/pushing to GitHub as your own work is strictly prohibited and constitutes copyright infringement. Any infringing repository is subject to immediate GitHub DMCA Takedown & Repository Deletion and legal enforcement.

F
license - not found
Not graded
quality - not tested
C
maintenance

Maintenance

โ€“Maintainers
โ€“Response time
โ€“Release cycle
โ€“Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to perform undetectable browser automation that bypasses Cloudflare, antibots, and social media blocks. Provides 105 tools for element extraction, network debugging, and real-world web scraping with a 98.7% success rate on protected sites.
    1,589
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables undetectable web scraping and browser automation for AI agents with 84 tools including stealth navigation, element extraction, network interception, and auto cookie consent dismissal. Bypasses anti-bot systems like Cloudflare and DataDome while providing LLM-ready markdown output and full Chrome DevTools Protocol access.
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Direct access to 40+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.
    2
    42
    4
    MIT

View all related MCP servers

Related MCP Connectors

  • Give your agent live data from Twitter, Reddit, the web and GitHub. No API keys, no scraping stack.

  • Reliable web access for AI agents: smart HTTP, rotating proxies, and full-browser rendering.

  • Enable language models to perform advanced AI-powered web scraping with enterprise-grade reliabiliโ€ฆ

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/virajverse/infinity-scraper'

If you have feedback or need assistance with the MCP directory API, please join our Discord server