InfinityScrape MCP
🌐 InfinityScrape MCP: ワールドクラスのWebスクレイピング&ディープOSINTインテリジェンススイート
InfinityScrape MCP は、AIモデル(LM Studio、Claude Desktop、Cursor、Open WebUI、Antigravity AI)に無制限・高速・アンチボット耐性のあるWebスクレイピング、動的SPAレンダリング、即時YouTube文字起こし、高精度なOSINT / GEOINT位置情報インテリジェンスを提供するために設計された、スタンドアロンでプロダクショングレードのModel Context Protocol(MCP)サーバーです。
📑 目次
Related MCP server: FineData MCP Server
🌟 InfinityScrape MCPが選ばれる理由
標準的なWebスクレイパーは、Cloudflareチャレンジ、大量のクライアントサイドJavaScriptレンダリング、邪魔なCookie同意モーダル、レート制限などにより、現代のWebサイトではしばしば失敗します。InfinityScrapeはこれらの問題を追加設定なしで解決します:
デュアルエンジンアーキテクチャ:
高速TLSエンジン(
primp+httpx): 実際のChrome/SafariブラウザのTLS/JA3フィンガープリントとHTTP/2ヘッダーを模倣し、CloudflareやAkamaiのチャレンジを<100msでバイパスします。動的ヘッドレスブラウザ(
Playwright Chromium): 複雑なSPA(React、Vue、Next.js、Angular)をレンダリングし、無限スクロールの実行、要素のクリック、カスタムJavaScriptの実行を行います。
ネットワークレベルでの広告&トラッカー排除:
35以上の広告ネットワークとトラッキングスクリプト(
doubleclick、criteo、outbrain、google-analytics)へのネットワークコールを、ダウンロードされる前にインターセプトして中断し、ページ読み込み時間を約300%、メモリ使用量を**70%**削減します。OneTrust、Cookiebot、スティッキーオーバーレイポップアップを自動検出して分解します。
Zero-GPU即時YouTube文字起こし:
動画/ショート/ライブの完全な文字起こしを、タイムスタンプ(
[MM:SS])付きで<300ms以内に抽出します。動画をダウンロードしたり、ローカルのGPU Whisperモデルを必要としたりせず、HTTPストリーム経由で直接取得します。
深層再帰型ドキュメントクローラー:
ドメインロックとパスプレフィックスフィルタリングを備えた非同期の幅優先探索(BFS)クローラーで、ドキュメントツリー全体を統合Markdownに集約します。
最先端のパブリックOSINT&GEOINT偵察:
マルチシグナル信頼度スコアリング(0%〜100%): 氏名+都市+通り+PIN+組織+役職の相関を評価し、発見された人物ファイルをランク付けします。
25以上のグローバルプラットフォームスキャナー: GitHub、GitLab、StackOverflow、Kaggle、HuggingFace、LeetCode、Codeforces、Dev.to、Medium、Substack、Google Scholar、ResearchGate、Redditなどをスキャンします。
OpenStreetMap GEOINT: GPS座標と行政境界を含め、通り/郵便番号レベルまでのグローバルな住所を解決します。
SQLite永続キャッシュレイヤー:
インメモリとSQLiteバックアップのローカルキャッシュにより、繰り返しのルックアップでは即時
0msの応答を実現し、TTLも設定可能です。
⚡ 競合比較
機能 / 能力 | 標準MCPスクレイパー | クラウドスクレイピングAPI | InfinityScrape MCP |
コスト&APIキー | 無料(基本) | 有料(月額$20〜$200) | 100%無料 / APIキー不要 |
Cloudflare / Akamai TLSバイパス | ❌ 失敗 / 403 | ✅ 対応 | ✅ 内蔵( |
動的SPA&無限スクロール | ❌ 制限あり | ✅ 対応 | ✅ 内蔵( |
ネットワークレベルでの広告&ポップアップ除去 | ❌ なし | ⚠️ 一部対応 | ✅ 内蔵(35以上のドメイン) |
Zero-GPU YouTube文字起こし | ❌ なし | ❌ なし | ✅ 内蔵(<300ms) |
オンラインPDFページ単位パーサー | ❌ なし | ⚠️ 追加費用 | ✅ 内蔵( |
深層ドキュメントクローラー | ❌ なし | ⚠️ 追加費用 | ✅ 内蔵(非同期BFS) |
25以上のプラットフォームOSINT&ジオコーディング | ❌ なし | ❌ なし | ✅ 内蔵(0〜100%信頼度) |
ローカルSQLite 0msキャッシュ | ❌ なし | ❌ なし | ✅ 内蔵(自動TTL) |
🏗️ アーキテクチャ概要
┌────────────────────────────────────────────────┐
│ AI Client (LM Studio / Claude / Cursor) │
└───────────────────────┬────────────────────────┘
│ JSON-RPC 2.0 (Stdio)
▼
┌────────────────────────────────────────────────┐
│ InfinityScrape MCP Server │
│ (server.py) │
└───────┬────────────────┬───────────────┬───────┘
│ │ │
┌─────────────────┴─┐ ┌────────┴────────┐ ┌─┴────────────────┐
▼ ▼ ▼ ▼ ▼ ▼
[Fast TLS Engine] [Playwright Engine] [OSINT / GEOINT] [Media & PDF Engines]
• primp JA3/TLS • Stealth Chromium • 25+ Platform • YouTube (<300ms)
• HTTP/2 Stealth • Network Ad Blocker Scanners • Remote PDF Stream
• <100ms Execution • Infinite Scroll • OpenStreetMap • Table Markdownify
• Auto-Dismiss CMPs • Match Confidence
│
▼
┌────────────────────────────────┐
│ SQLite Caching Layer (0ms TTL) │
└────────────────────────────────┘🚀 クイックスタート&ワンクリックインストール
前提条件
Python 3.10、3.11、または3.12+がインストールされていること。
Windows、macOS、またはLinux。
ワンクリックセットアップ:
Windowsの場合:
install.batをダブルクリックするか、PowerShellで以下を実行します:
.\install.batLinux / macOSの場合:
chmod +x install.sh
./install.sh手動セットアップ(全プラットフォーム):
# 1. Create virtual environment
python -m venv .venv
# 2. Activate virtual environment
# Windows: .venv\Scripts\activate | Linux/Mac: source .venv/bin/activate
# 3. Install requirements & Playwright browser
pip install -r requirements.txt
playwright install chromium🔌 AIクライアント連携
1. LM Studio(v0.3+)
Settings ➔ Developer ➔ MCP Servers ➔ Edit Config に移動し、以下を追加します:
{
"mcpServers": {
"infinity-scraper": {
"command": "C:/path/to/infinity-scraper/.venv/Scripts/python.exe",
"args": [
"-m",
"infinity_scraper.server"
],
"cwd": "C:/path/to/infinity-scraper",
"env": {
"PYTHONUNBUFFERED": "1"
}
}
}
}2. Claude Desktop
%APPDATA%\Claude\claude_desktop_config.json(Windows)または~/Library/Application Support/Claude/claude_desktop_config.json(macOS)を編集します:
{
"mcpServers": {
"infinity-scraper": {
"command": "C:/path/to/infinity-scraper/.venv/Scripts/python.exe",
"args": [
"-m",
"infinity_scraper.server"
]
}
}
}3. Cursor IDE
CursorのSettings ➔ Features ➔ MCP Servers ➔ Add New MCP Server で設定します:
Name:
infinity-scraperType:
commandCommand:
C:/path/to/infinity-scraper/.venv/Scripts/python.exe -m infinity_scraper.server
🛠️ 全23ツールリファレンスガイド
1. Webスクレイピング&コンテンツクローリング
ツール | 目的 | 主要パラメータ |
| TLSからブラウザエンジンへ自動アップグレードするユニバーサルスクレイピング。 |
|
| SPA、無限スクロール、クリック操作に対応したヘッドレスブラウザ。 |
|
| DuckDuckGoによるリアルタイムのインターネット検索。 |
|
| Webを検索し、上位結果を自動的にスクレイピングして引用付きレポートにまとめます。 |
|
| 完全なドキュメントツリーのための再帰的・非同期BFSサイトクローラー。 |
|
| 複数のURLを並行して同時にスクレイピングします。 |
|
| CSSセレクターマップを使用して対象フィールドを抽出しJSONに変換します。 |
|
| JSON-LD、OpenGraphメタデータ、HTMLテーブルを抽出します。 |
|
| 巨大なページのためのセマンティックRAGチャンカー&トークン最適化ツール。 |
|
2. メディア、ソーシャル、ビデオ&ドキュメントパーサー
ツール | 目的 | 主要パラメータ |
| タイムスタンプ付きのZero-GPU YouTube動画文字起こし抽出。 |
|
| リモートのオンラインPDFドキュメントをページごとにストリーミング抽出します。 |
|
| 写真からカメラ仕様、タイムスタンプ、GPSジオタグを抽出します。 |
|
| Redditの投稿、スコア、ネストされたコメント対話を取り込みます。 |
|
| ブログ、Substack、ニュース向けのリアルタイムRSS/Atomフィードパーサー。 |
|
3. ディープパブリックOSINT&エンティティ偵察
ツール | 目的 | 主要パラメータ |
| マルチドメインのオープンWebプロフィールスクレイパー&信頼度ランキング付き人物ファイルビルダー。 |
|
| グローバルなOpenStreetMapフォワードジオコーディング&行政区分の分解。 |
|
| ネガティブフィルター付きの階層型ロケーション詳細検索マトリックス。 |
|
| 26のコーディング、学術、クリエイティブネットワークにおけるユーザー名の存在をスキャンします。 |
|
| 精密な検索ドーキング( |
|
4. テクニカル、ドメイン&ネットワークインテリジェンス
ツール | 目的 | 主要パラメータ |
| ドメインのSSL/TLS証明書の有効性、DNS、RDAP/WHOISを検査します。 |
|
| フロントエンドフレームワーク(React、Next.js、Vue)、CMS、CDN、サーバーを検出します。 |
|
| パブリックIPの地理位置情報、ASN、ISP、組織インテル。 |
|
| 過去へのタイムトラベル&削除されたウェブページのスナップショットを取得するスクレイパー。 |
|
| Certificate Transparencyによるサブドメイン探索を1秒未満で実行。 |
|
| 詳細なDNSレコード(IPv4、IPv6、MX)のインフラ監査。 |
|
🧠 自律型AIエージェントプレイブック
InfinityScrapeには、自律型AIエージェントに以下の方法を教える高度な**Cognitive Reasoning Framework(skills/infinity-scraper/SKILL.md)**が含まれています:
ユーザーのプロンプトを検索の手がかりと場所の手がかりに動的に分解します。
ツール横断のマルチツールチェーン(例:
Search ➔ Filter ➔ Batch ScrapeやGeocode ➔ Locality Dork ➔ Profile Extraction)。動的なReactシングルページアプリに遭遇した場合、高速TLSからPlaywrightヘッドレスブラウザへ自動エスカレーションします。
👉 エージェントのプレイブック全文を読む:skills/infinity-scraper/SKILL.md
💻 コマンドラインインターフェース(CLI)
また、InfinityScrapeはターミナルから直接使用することもできます:
# Scrape a URL to Markdown
python -m infinity_scraper.cli scrape "https://example.com"
# Scrape dynamic SPA with infinite scroll
python -m infinity_scraper.cli scrape "https://news.ycombinator.com" --browser --scroll 3
# Live search and auto-scrape top results
python -m infinity_scraper.cli search "Quantum computing breakthroughs" --scrape --max 4
# Crawl documentation tree
python -m infinity_scraper.cli crawl "https://docs.python.org/3/library/asyncio.html" --pages 5 --depth 2
# Extract remote PDF
python -m infinity_scraper.cli pdf "https://example.com/report.pdf" --pages 10🧪 自動テストの実行
包括的なユニットテストおよび統合テストスイートを実行します:
python -m tests.test_scraperテストカバレッジ:
✅ 高速TLSインパーソネーター
✅ Playwright動的ブラウザ
✅ DuckDuckGoライブ検索
✅ HTMLテーブル→Markdownコンバーター
✅ 再帰的BFSドキュメントクローラー
✅ ゼロGPU YouTube文字起こし抽出
✅ OSINT SSL、IPインテル&25以上のプラットフォーム存在確認
📄 ライセンスと著作権保護
このプロジェクトは**MITライセンス(必須の帰属表示&DMCA執行付き)**の下でライセンス供与されています。
[!IMPORTANT] 必須の帰属表示に関する通知:
このプロジェクトは、商用または個人使用のために自由に使用、変更、統合することができます。
ただし、元の著者帰属表示および著作権表示は、すべてのコピー、フォーク、または派生的配布物において必ず保持されなければなりません。
著者の名前・クレジットを削除し、自身の作品としてGitHubに再アップロード/プッシュすることは固く禁止されており、著作権侵害を構成します。侵害するリポジトリは、即時のGitHub DMCAテイクダウン&リポジトリ削除および法的執行の対象となります。
著者/作成者: VirajVerse
完全な法的条件については
LICENSEを参照してください。
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to perform undetectable browser automation that bypasses Cloudflare, antibots, and social media blocks. Provides 105 tools for element extraction, network debugging, and real-world web scraping with a 98.7% success rate on protected sites.1,589MIT
- AlicenseAqualityBmaintenanceEnables AI agents to scrape any website by providing tools for JavaScript rendering, antibot bypass, and automatic captcha solving. It supports synchronous, asynchronous, and batch scraping operations with built-in proxy rotation.5207MIT

ScrapeLab MCPofficial
AlicenseNot gradedqualityDmaintenanceEnables undetectable web scraping and browser automation for AI agents with 84 tools including stealth navigation, element extraction, network interception, and auto cookie consent dismissal. Bypasses anti-bot systems like Cloudflare and DataDome while providing LLM-ready markdown output and full Chrome DevTools Protocol access.MIT
HasData MCP Serverofficial
AlicenseAqualityAmaintenanceDirect access to 40+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.2424MIT
Related MCP Connectors
Give your agent live data from Twitter, Reddit, the web and GitHub. No API keys, no scraping stack.
Reliable web access for AI agents: smart HTTP, rotating proxies, and full-browser rendering.
Enable language models to perform advanced AI-powered web scraping with enterprise-grade reliabili…
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/virajverse/infinity-scraper'
If you have feedback or need assistance with the MCP directory API, please join our Discord server