MCP Web Research Server
MCP 딥웹 연구 서버(v0.3.0)
고급 웹 연구를 위한 MCP(Model Context Protocol) 서버입니다.
최신 변경 사항
웹페이지 콘텐츠를 직접 추출하기 위한 visit_page 도구가 추가되었습니다.
MCP 시간 초과 제한 내에서 작동하도록 성능이 최적화되었습니다.
기본 maxDepth 및 maxBranching 매개변수가 축소되었습니다.
향상된 페이지 로딩 효율성
프로세스 전반에 걸쳐 시간 초과 확인이 추가되었습니다.
시간 초과에 대한 향상된 오류 처리
이 프로젝트는 mzxrai 가 개발한 mcp-webresearch 의 포크(fork)로, 딥웹 리서치 기능을 위한 추가 기능이 강화되었습니다. 초기 개발에 기여해 주신 개발자분들께 감사드립니다.
지능형 검색 대기열, 향상된 콘텐츠 추출, 심층 조사 기능을 통해 클로드에 실시간 정보를 제공합니다.
Related MCP server: MCP Web Research Server
특징
지능형 검색 대기열 시스템
속도 제한을 사용한 일괄 검색 작업
진행 상황 추적을 통한 대기열 관리
오류 복구 및 자동 재시도
검색 결과 중복 제거
향상된 콘텐츠 추출
TF-IDF 기반 관련성 점수
키워드 근접 분석
콘텐츠 섹션 가중치
가독성 점수
개선된 HTML 구조 파싱
구조화된 데이터 추출
더 나은 콘텐츠 정리 및 서식
핵심 기능
Google 검색 통합
웹페이지 콘텐츠 추출
연구 세션 추적
향상된 서식을 사용한 마크다운 변환
필수 조건
Node.js >= 18 (
npm및npx포함)
설치
글로벌 설치(권장)
지엑스피1
로컬 프로젝트 설치
# Using npm
npm install mcp-deepwebresearch
# Using yarn
yarn add mcp-deepwebresearch
# Using pnpm
pnpm add mcp-deepwebresearchClaude 데스크톱 통합
패키지를 설치한 후 claude_desktop_config.json 에 다음 항목을 추가하세요.
윈도우
{
"mcpServers": {
"deepwebresearch": {
"command": "mcp-deepwebresearch",
"args": []
}
}
}위치: %APPDATA%\Claude\claude_desktop_config.json
맥OS
{
"mcpServers": {
"deepwebresearch": {
"command": "mcp-deepwebresearch",
"args": []
}
}
}위치: ~/Library/Application Support/Claude/claude_desktop_config.json
이 구성을 사용하면 Claude Desktop이 필요할 때 자동으로 웹 연구 MCP 서버를 시작할 수 있습니다.
첫 번째 설정
설치 후 다음 명령을 실행하여 필요한 브라우저 종속성을 설치합니다.
npx playwright install chromium용법
Claude와 채팅을 시작하고 웹 리서치에 도움이 될 만한 메시지를 보내세요. 심층적인 웹 리서치에 맞춰 미리 만들어진 프롬프트를 원하시면 이 패키지를 통해 제공되는 agentic-research 프롬프트를 사용하실 수 있습니다. Claude 데스크톱에서 채팅 입력란의 페이퍼클립 아이콘을 클릭한 다음, Choose an integration → deepwebresearch → agentic-research 선택하여 해당 프롬프트에 접속하세요.
도구
deep_research콘텐츠 분석을 통해 포괄적인 연구를 수행합니다.
인수:
{ topic: string; maxDepth?: number; // default: 2 maxBranching?: number; // default: 3 timeout?: number; // default: 55000 (55 seconds) minRelevanceScore?: number; // default: 0.7 }보고:
{ findings: { mainTopics: Array<{name: string, importance: number}>; keyInsights: Array<{text: string, confidence: number}>; sources: Array<{url: string, credibilityScore: number}>; }; progress: { completedSteps: number; totalSteps: number; processedUrls: number; }; timing: { started: string; completed?: string; duration?: number; operations?: { parallelSearch?: number; deduplication?: number; topResultsProcessing?: number; remainingResultsProcessing?: number; total?: number; }; }; }
parallel_search지능형 대기열을 통해 여러 Google 검색을 병렬로 수행합니다.
인수:
{ queries: string[], maxParallel?: number }참고: maxParallel은 안정적인 성능을 보장하기 위해 5로 제한됩니다.
visit_page웹 페이지를 방문하여 콘텐츠 추출
인수:
{ url: string }보고:
{ url: string; title: string; content: string; // Markdown formatted content }
프롬프트
agentic-research
클로드가 철저한 웹 조사를 수행하는 데 도움이 되는 가이드 조사 프롬프트입니다. 이 프롬프트는 클로드에게 다음을 지시합니다.
주제 환경을 이해하려면 광범위한 검색으로 시작하세요.
고품질의 권위 있는 출처를 우선시하세요
연구 결과를 바탕으로 연구 방향을 반복적으로 개선합니다.
귀하에게 정보를 제공하고 대화형으로 연구를 안내합니다.
항상 URL을 사용하여 출처를 인용하세요
구성 옵션
서버는 환경 변수를 통해 구성할 수 있습니다.
MAX_PARALLEL_SEARCHES: 동시 검색 최대 수 (기본값: 5)SEARCH_DELAY_MS: 검색 간 지연 시간(밀리초) (기본값: 200)MAX_RETRIES: 실패한 요청에 대한 재시도 횟수(기본값: 3)TIMEOUT_MS: 요청 시간 초과(밀리초) (기본값: 55000)LOG_LEVEL: 로깅 레벨(기본값: 'info')
오류 처리
일반적인 문제
속도 제한
증상: "요청이 너무 많습니다" 오류
해결 방법:
SEARCH_DELAY_MS늘리거나MAX_PARALLEL_SEARCHES줄이세요.
네트워크 시간 초과
증상: "요청 시간 초과" 오류
솔루션: 요청이 60초 MCP 시간 초과 내에 완료되도록 보장합니다.
브라우저 문제
증상: "브라우저를 시작하지 못했습니다" 오류
해결 방법: Playwright가 제대로 설치되었는지 확인하세요(
npx playwright install)
디버깅
베타 소프트웨어입니다. 문제가 발생할 경우:
Claude Desktop의 MCP 로그를 확인하세요.
# On macOS tail -n 20 -f ~/Library/Logs/Claude/mcp*.log # On Windows Get-Content -Path "$env:APPDATA\Claude\logs\mcp*.log" -Tail 20 -Wait디버그 로깅 활성화:
export LOG_LEVEL=debug
개발
설정
# Install dependencies
pnpm install
# Build the project
pnpm build
# Watch for changes
pnpm watch
# Run in development mode
pnpm dev테스트
# Run all tests
pnpm test
# Run tests in watch mode
pnpm test:watch
# Run tests with coverage
pnpm test:coverage코드 품질
# Run linter
pnpm lint
# Fix linting issues
pnpm lint:fix
# Type check
pnpm type-check기여하다
저장소를 포크하세요
기능 브랜치를 생성합니다(
git checkout -b feature/amazing-feature)변경 사항을 커밋하세요(
git commit -m 'Add some amazing feature')브랜치에 푸시(
git push origin feature/amazing-feature)풀 리퀘스트 열기
코딩 표준
TypeScript 모범 사례를 따르세요
테스트 커버리지를 80% 이상으로 유지
새로운 기능 및 API 문서화
중요한 변경 사항이 있는 경우 CHANGELOG.md를 업데이트하세요.
의미적 버전 관리를 따르세요
성능 고려 사항
가능한 경우 일괄 작업을 사용하세요
적절한 오류 처리 및 재시도를 구현합니다.
대용량 데이터 세트의 메모리 사용량을 고려하세요
적절한 경우 캐시 결과
대용량 콘텐츠에 스트리밍을 사용하세요
요구 사항
노드.js >= 18
Playwright(종속성으로 자동 설치됨)
검증된 플랫폼
[x] 맥OS
[x] 윈도우
[ ] 리눅스
특허
MIT
크레딧
이 프로젝트는 mzxrai 가 개발한 mcp-webresearch 의 훌륭한 작업을 기반으로 합니다. 원본 코드베이스는 향상된 기능과 성능의 기반을 제공했습니다.
작가
Available Tools
3 toolsdeep_researchC
Perform deep research on a topic with content extraction and analysis
| Name | Required | Description | Default |
|---|---|---|---|
| maxBranching | No | Maximum number of related paths to explore | |
| maxDepth | No | Maximum depth of related content exploration | |
| minRelevanceScore | No | Minimum relevance score for including content | |
| timeout | No | Research timeout in milliseconds | |
| topic | Yes | Research topic or question |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'content extraction and analysis' but fails to detail critical aspects such as execution time, resource usage, error handling, or output format. This leaves significant gaps in understanding how the tool behaves beyond its basic purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's front-loaded with the core action and avoids redundancy, making it highly concise and well-structured for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a 'deep research' tool with 5 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what 'deep research' entails, how results are returned, or any behavioral constraints, leaving the agent with inadequate information for effective use in a broader context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, meaning all parameters are documented in the schema. The description adds no additional semantic context about parameters beyond implying 'deep research' involves branching and depth. This meets the baseline for high schema coverage but doesn't enhance understanding of parameter roles or interactions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Perform deep research on a topic with content extraction and analysis,' which specifies the verb (perform deep research) and resource (topic) with additional capabilities (content extraction and analysis). However, it doesn't explicitly differentiate from sibling tools like 'parallel_search' or 'visit_page,' which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'parallel_search' or 'visit_page.' It lacks any context about appropriate scenarios, prerequisites, or exclusions, leaving the agent with minimal direction for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parallel_searchC
Perform multiple Google searches in parallel
| Name | Required | Description | Default |
|---|---|---|---|
| maxParallel | No | Maximum number of parallel searches | |
| queries | Yes | Array of search queries to execute in parallel |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'parallel' execution but doesn't explain what that entails operationally (e.g., concurrency limits, error handling, or performance implications). It also omits critical details like authentication needs, rate limits, or whether this is a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with a single, clear sentence that directly states the tool's function. There is no wasted language or unnecessary elaboration, making it easy to parse and understand at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is insufficient for a tool that performs parallel operations. It doesn't address key behavioral aspects like error handling, result format, or limitations of parallel execution, leaving significant gaps in understanding how to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, providing clear documentation for both parameters. The description adds minimal value beyond the schema by implying the tool handles multiple queries simultaneously, but doesn't elaborate on parameter interactions or usage nuances beyond what's already in the structured data.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('perform') and resource ('Google searches'), and specifies the parallel execution aspect. However, it doesn't explicitly differentiate from sibling tools like 'deep_research' or 'visit_page', which might have overlapping search functionality but different approaches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'deep_research' or 'visit_page'. It doesn't specify scenarios where parallel searching is preferred over sequential or deeper research methods, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visit_pageC
Visit a webpage and extract its content
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to visit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'visit a webpage and extract its content', which implies a read operation, but doesn't specify details like authentication needs, rate limits, error handling, or what 'extract content' entails (e.g., HTML, text, metadata). For a tool with no annotations, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded with the core action ('visit a webpage') and purpose ('extract its content'), making it easy to understand quickly. Every part of the sentence earns its place by conveying essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (a web interaction tool with potential behavioral nuances) and the lack of annotations and output schema, the description is incomplete. It doesn't cover what 'extract content' means in terms of output format, error cases, or limitations. For a tool that interacts with external webpages, more context is needed to ensure proper usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'url' parameter clearly documented as 'URL to visit'. The description adds no additional meaning beyond this, as it doesn't elaborate on URL format constraints or extraction specifics. With high schema coverage, the baseline score of 3 is appropriate, as the description doesn't compensate but doesn't need to given the schema's clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('visit') and resource ('webpage'), and specifies the action ('extract its content'). However, it doesn't differentiate this tool from potential sibling tools like 'deep_research' or 'parallel_search', which might have overlapping functionality. The description is not tautological but lacks sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any context, prerequisites, or exclusions, and doesn't reference sibling tools like 'deep_research' or 'parallel_search' that might be related. Usage is implied only by the tool's name and description, with no explicit guidelines.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
deep_research - First observed
parallel_search - First observed
visit_page
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: deep_research is for comprehensive topic analysis, parallel_search is for multi-query search execution, and visit_page is for single-page content extraction. There is no overlap in functionality, making tool selection unambiguous for an agent.
All tools follow a consistent snake_case verb_noun pattern (deep_research, parallel_search, visit_page) with clear action-oriented names. The naming scheme is predictable and readable throughout the set.
With only 3 tools, the set feels thin for a 'Web Research Server' domain, lacking operations like search filtering, result summarization, or citation management. While the tools cover core actions, the count is borderline minimal for comprehensive research workflows.
The tools cover basic research steps (search, page access, analysis), but there are notable gaps: no ability to refine searches, save results, compare sources, or handle authentication. This limits agents to a linear workflow without advanced research capabilities.
Maintenance
Related MCP Connectors
Live AI-native web search with citations. One tool for every MCP client. Flat per-request pricing.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Web MCP: scrape/crawl sites, web search, brand assets, app stores, YouTube, Reddit, Hacker News.
The Remote MCP server acts as a standardized bridge between LLM applications (like Claude, ChatGPT, and Cursor) and external services, enabling AI agents to access external tools and resources. Its primary capability is providing a centralized search tool to discover other MCP servers and their respective tools. Unlike local implementations, it runs remotely with OAuth authentication and permission controls for security.
Related MCP Servers
- AlicenseBqualityFmaintenanceA Model Context Protocol (MCP) server for web research. Bring real-time info into Claude and easily research any topic.31,370 npm298MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables Claude to perform web research by integrating Google search, extracting webpage content, and capturing screenshots.31,370 npm20MIT
- AlicenseAqualityCmaintenanceA Model Context Protocol server that enables Claude to perform web research by integrating Google search, extracting webpage content, and capturing screenshots in real-time.41,370 npm9MIT
- AlicenseBqualityDmaintenanceA server that integrates with Claude Desktop to enable real-time web research capabilities, allowing users to search Google, extract webpage content, and capture screenshots directly from conversations.31,370 npmMIT