MCP Web Research Server
MCP ディープウェブリサーチサーバー (v0.3.0)
高度な Web 調査のためのモデル コンテキスト プロトコル (MCP) サーバー。
最新の変更点
ウェブページコンテンツを直接抽出するための visit_page ツールを追加しました
MCPタイムアウト制限内で動作するように最適化されたパフォーマンス
デフォルトのmaxDepthとmaxBranchingパラメータを削減
ページの読み込み効率の向上
プロセス全体にタイムアウトチェックを追加しました
タイムアウト時のエラー処理の強化
このプロジェクトは、 mzxraiによるmcp-webresearchのフォークであり、ディープウェブリサーチ機能のための追加機能が強化されています。私たちは、元の作成者たちの基礎的な作業に感謝いたします。
インテリジェントな検索キューイング、強化されたコンテンツ抽出、詳細な調査機能により、Claude にリアルタイムの情報を取り込みます。
Related MCP server: MCP Web Research Server
特徴
インテリジェント検索キューシステム
レート制限付きのバッチ検索操作
進捗状況を追跡するキュー管理
エラー回復と自動再試行
検索結果の重複排除
強化されたコンテンツ抽出
TF-IDFベースの関連性スコアリング
キーワード近接分析
コンテンツセクションの重み付け
読みやすさスコア
HTML構造解析の改善
構造化データ抽出
コンテンツの整理とフォーマットの改善
コア機能
Google検索統合
ウェブページコンテンツの抽出
研究セッションの追跡
フォーマットが改善されたマークダウン変換
前提条件
Node.js >= 18 (
npmとnpxを含む)
インストール
グローバルインストール(推奨)
# Install globally using npm
npm install -g mcp-deepwebresearch
# Or using yarn
yarn global add mcp-deepwebresearch
# Or using pnpm
pnpm add -g mcp-deepwebresearchローカルプロジェクトのインストール
# Using npm
npm install mcp-deepwebresearch
# Using yarn
yarn add mcp-deepwebresearch
# Using pnpm
pnpm add mcp-deepwebresearchクロードデスクトップ統合
パッケージをインストールした後、 claude_desktop_config.jsonに次のエントリを追加します。
ウィンドウズ
{
"mcpServers": {
"deepwebresearch": {
"command": "mcp-deepwebresearch",
"args": []
}
}
}場所: %APPDATA%\Claude\claude_desktop_config.json
macOS
{
"mcpServers": {
"deepwebresearch": {
"command": "mcp-deepwebresearch",
"args": []
}
}
}場所: ~/Library/Application Support/Claude/claude_desktop_config.json
この設定により、Claude Desktop は必要に応じて Web リサーチ MCP サーバーを自動的に起動できるようになります。
初回セットアップ
インストール後、このコマンドを実行して必要なブラウザ依存関係をインストールします。
npx playwright install chromium使用法
Claude とチャットを開始し、Web リサーチに役立つプロンプトを送信するだけです。より詳細な Web リサーチ向けにカスタマイズされた既成のプロンプトが必要な場合は、このパッケージで提供されているagentic-researchプロンプトをご利用ください。Claude Desktop でこのプロンプトにアクセスするには、チャット入力欄のペーパークリップアイコンをクリックし、 Choose an integration → deepwebresearch → agentic-researchを選択します。
ツール
deep_researchコンテンツ分析による包括的な調査を実施
引数:
{ topic: string; maxDepth?: number; // default: 2 maxBranching?: number; // default: 3 timeout?: number; // default: 55000 (55 seconds) minRelevanceScore?: number; // default: 0.7 }戻り値:
{ findings: { mainTopics: Array<{name: string, importance: number}>; keyInsights: Array<{text: string, confidence: number}>; sources: Array<{url: string, credibilityScore: number}>; }; progress: { completedSteps: number; totalSteps: number; processedUrls: number; }; timing: { started: string; completed?: string; duration?: number; operations?: { parallelSearch?: number; deduplication?: number; topResultsProcessing?: number; remainingResultsProcessing?: number; total?: number; }; }; }
parallel_searchインテリジェントなキューイングで複数の Google 検索を並行して実行します
引数:
{ queries: string[], maxParallel?: number }注: 信頼性の高いパフォーマンスを確保するために、maxParallel は 5 に制限されています。
visit_pageウェブページにアクセスしてコンテンツを抽出する
引数:
{ url: string }戻り値:
{ url: string; title: string; content: string; // Markdown formatted content }
プロンプト
agentic-research
クロードが徹底的なウェブリサーチを行うのに役立つガイド付きリサーチプロンプト。このプロンプトは、クロードに以下の指示を与えます。
トピックの状況を理解するために、まずは広範囲な検索から始めましょう
高品質で信頼できる情報源を優先する
調査結果に基づいて研究の方向性を繰り返し改善する
情報を提供し、インタラクティブに研究を進めることができます
常にURLでソースを引用する
設定オプション
サーバーは環境変数を通じて設定できます:
MAX_PARALLEL_SEARCHES: 同時検索の最大数(デフォルト: 5)SEARCH_DELAY_MS: 検索間の遅延(ミリ秒単位)(デフォルト: 200)MAX_RETRIES: 失敗したリクエストの再試行回数(デフォルト: 3)TIMEOUT_MS: リクエストタイムアウト(ミリ秒)(デフォルト: 55000)LOG_LEVEL: ログレベル(デフォルト: 'info')
エラー処理
よくある問題
レート制限
症状: 「リクエストが多すぎます」というエラー
解決策:
SEARCH_DELAY_MSを増やすか、MAX_PARALLEL_SEARCHESを減らす
ネットワークタイムアウト
症状: 「リクエストがタイムアウトしました」というエラー
解決策: リクエストが60秒のMCPタイムアウト内に完了することを確認する
ブラウザの問題
症状: 「ブラウザの起動に失敗しました」というエラー
解決策: Playwright が正しくインストールされていることを確認してください (
npx playwright install)
デバッグ
これはベータ版ソフトウェアです。問題が発生した場合は、以下の手順に従ってください。
Claude Desktop の MCP ログを確認します。
# On macOS tail -n 20 -f ~/Library/Logs/Claude/mcp*.log # On Windows Get-Content -Path "$env:APPDATA\Claude\logs\mcp*.log" -Tail 20 -Waitデバッグ ログを有効にする:
export LOG_LEVEL=debug
発達
設定
# Install dependencies
pnpm install
# Build the project
pnpm build
# Watch for changes
pnpm watch
# Run in development mode
pnpm devテスト
# Run all tests
pnpm test
# Run tests in watch mode
pnpm test:watch
# Run tests with coverage
pnpm test:coverageコード品質
# Run linter
pnpm lint
# Fix linting issues
pnpm lint:fix
# Type check
pnpm type-check貢献
リポジトリをフォークする
機能ブランチを作成します(
git checkout -b feature/amazing-feature)変更をコミットします (
git commit -m 'Add some amazing feature')ブランチにプッシュする (
git push origin feature/amazing-feature)プルリクエストを開く
コーディング標準
TypeScriptのベストプラクティスに従う
テストカバレッジを80%以上維持する
新しい機能とAPIを文書化する
重要な変更についてはCHANGELOG.mdを更新してください
セマンティックバージョニングに従う
パフォーマンスに関する考慮事項
可能な場合はバッチ操作を使用する
適切なエラー処理と再試行を実装する
大規模なデータセットでのメモリ使用量を考慮する
適切な場合に結果をキャッシュする
大容量コンテンツにはストリーミングを使用する
要件
Node.js >= 18
Playwright (依存関係として自動的にインストールされます)
検証済みプラットフォーム
[x] macOS
[x] ウィンドウズ
[ ] リナックス
ライセンス
マサチューセッツ工科大学
クレジット
このプロジェクトは、 mzxraiによるmcp-webresearchの優れた成果を基盤としています。オリジナルのコードベースは、私たちの強化された機能と性能の基盤となりました。
著者
Available Tools
3 toolsdeep_researchC
Perform deep research on a topic with content extraction and analysis
| Name | Required | Description | Default |
|---|---|---|---|
| maxBranching | No | Maximum number of related paths to explore | |
| maxDepth | No | Maximum depth of related content exploration | |
| minRelevanceScore | No | Minimum relevance score for including content | |
| timeout | No | Research timeout in milliseconds | |
| topic | Yes | Research topic or question |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'content extraction and analysis' but fails to detail critical aspects such as execution time, resource usage, error handling, or output format. This leaves significant gaps in understanding how the tool behaves beyond its basic purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's front-loaded with the core action and avoids redundancy, making it highly concise and well-structured for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a 'deep research' tool with 5 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what 'deep research' entails, how results are returned, or any behavioral constraints, leaving the agent with inadequate information for effective use in a broader context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, meaning all parameters are documented in the schema. The description adds no additional semantic context about parameters beyond implying 'deep research' involves branching and depth. This meets the baseline for high schema coverage but doesn't enhance understanding of parameter roles or interactions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Perform deep research on a topic with content extraction and analysis,' which specifies the verb (perform deep research) and resource (topic) with additional capabilities (content extraction and analysis). However, it doesn't explicitly differentiate from sibling tools like 'parallel_search' or 'visit_page,' which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'parallel_search' or 'visit_page.' It lacks any context about appropriate scenarios, prerequisites, or exclusions, leaving the agent with minimal direction for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parallel_searchC
Perform multiple Google searches in parallel
| Name | Required | Description | Default |
|---|---|---|---|
| maxParallel | No | Maximum number of parallel searches | |
| queries | Yes | Array of search queries to execute in parallel |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'parallel' execution but doesn't explain what that entails operationally (e.g., concurrency limits, error handling, or performance implications). It also omits critical details like authentication needs, rate limits, or whether this is a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with a single, clear sentence that directly states the tool's function. There is no wasted language or unnecessary elaboration, making it easy to parse and understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is insufficient for a tool that performs parallel operations. It doesn't address key behavioral aspects like error handling, result format, or limitations of parallel execution, leaving significant gaps in understanding how to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, providing clear documentation for both parameters. The description adds minimal value beyond the schema by implying the tool handles multiple queries simultaneously, but doesn't elaborate on parameter interactions or usage nuances beyond what's already in the structured data.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('perform') and resource ('Google searches'), and specifies the parallel execution aspect. However, it doesn't explicitly differentiate from sibling tools like 'deep_research' or 'visit_page', which might have overlapping search functionality but different approaches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'deep_research' or 'visit_page'. It doesn't specify scenarios where parallel searching is preferred over sequential or deeper research methods, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visit_pageC
Visit a webpage and extract its content
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to visit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'visit a webpage and extract its content', which implies a read operation, but doesn't specify details like authentication needs, rate limits, error handling, or what 'extract content' entails (e.g., HTML, text, metadata). For a tool with no annotations, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded with the core action ('visit a webpage') and purpose ('extract its content'), making it easy to understand quickly. Every part of the sentence earns its place by conveying essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (a web interaction tool with potential behavioral nuances) and the lack of annotations and output schema, the description is incomplete. It doesn't cover what 'extract content' means in terms of output format, error cases, or limitations. For a tool that interacts with external webpages, more context is needed to ensure proper usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'url' parameter clearly documented as 'URL to visit'. The description adds no additional meaning beyond this, as it doesn't elaborate on URL format constraints or extraction specifics. With high schema coverage, the baseline score of 3 is appropriate, as the description doesn't compensate but doesn't need to given the schema's clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('visit') and resource ('webpage'), and specifies the action ('extract its content'). However, it doesn't differentiate this tool from potential sibling tools like 'deep_research' or 'parallel_search', which might have overlapping functionality. The description is not tautological but lacks sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any context, prerequisites, or exclusions, and doesn't reference sibling tools like 'deep_research' or 'parallel_search' that might be related. Usage is implied only by the tool's name and description, with no explicit guidelines.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
deep_research - First observed
parallel_search - First observed
visit_page
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: deep_research is for comprehensive topic analysis, parallel_search is for multi-query search execution, and visit_page is for single-page content extraction. There is no overlap in functionality, making tool selection unambiguous for an agent.
All tools follow a consistent snake_case verb_noun pattern (deep_research, parallel_search, visit_page) with clear action-oriented names. The naming scheme is predictable and readable throughout the set.
With only 3 tools, the set feels thin for a 'Web Research Server' domain, lacking operations like search filtering, result summarization, or citation management. While the tools cover core actions, the count is borderline minimal for comprehensive research workflows.
The tools cover basic research steps (search, page access, analysis), but there are notable gaps: no ability to refine searches, save results, compare sources, or handle authentication. This limits agents to a linear workflow without advanced research capabilities.
Maintenance
Related MCP Connectors
Live AI-native web search with citations. One tool for every MCP client. Flat per-request pricing.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Web MCP: scrape/crawl sites, web search, brand assets, app stores, YouTube, Reddit, Hacker News.
The Remote MCP server acts as a standardized bridge between LLM applications (like Claude, ChatGPT, and Cursor) and external services, enabling AI agents to access external tools and resources. Its primary capability is providing a centralized search tool to discover other MCP servers and their respective tools. Unlike local implementations, it runs remotely with OAuth authentication and permission controls for security.
Related MCP Servers
- AlicenseBqualityFmaintenanceA Model Context Protocol (MCP) server for web research. Bring real-time info into Claude and easily research any topic.31,370 npm298MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables Claude to perform web research by integrating Google search, extracting webpage content, and capturing screenshots.31,370 npm20MIT
- AlicenseAqualityCmaintenanceA Model Context Protocol server that enables Claude to perform web research by integrating Google search, extracting webpage content, and capturing screenshots in real-time.41,370 npm9MIT
- AlicenseBqualityDmaintenanceA server that integrates with Claude Desktop to enable real-time web research capabilities, allowing users to search Google, extract webpage content, and capture screenshots directly from conversations.31,370 npmMIT