Skip to main content
Glama

✨ Features

  • 🔍 DuckDuckGo Search: Fast and privacy-focused web search capability

  • 📄 Content Extraction: Clean, readable text extraction from web pages

  • 🚀 Parallel Processing: Support for extracting content from multiple URLs simultaneously

  • 💾 Memory Optimization: Smart memory management to prevent application crashes

  • ⏱️ Rate Limiting: Intelligent request throttling to avoid API blocks

  • 🛡️ Error Handling: Robust error handling for reliable operation

Related MCP server: Web Search MCP Server

📦 Installation

Installing via Smithery

To install Web Scout for Claude Desktop automatically via Smithery:

npx -y @smithery/cli install @pinkpixel-dev/web-scout-mcp --client claude

Global Installation

npm install -g @pinkpixel/web-scout-mcp

Local Installation

npm install @pinkpixel/web-scout-mcp

🚀 Usage

Command Line

After installing globally, run:

web-scout-mcp

With MCP Clients

Add this to your MCP client's config.json (Claude Desktop, Cursor, etc.):

{
  "mcpServers": {
    "web-scout": {
      "command": "npx",
      "args": [
        "-y",
        "@pinkpixel/web-scout-mcp@latest"
      ]
    }
  }
}

Environment Variables

Set the WEB_SCOUT_DISABLE_AUTOSTART=1 environment variable when embedding the package and calling createServer() yourself. By default running the published entrypoint (for example node dist/index.js or npx @pinkpixel/web-scout-mcp) automatically bootstraps the stdio transport.

🧰 Tools

The server provides the following MCP tools:

Initiates a web search query using the DuckDuckGo search engine and returns a well-structured list of findings.

Input:

  • query (string): The search query string

  • maxResults (number, optional): Maximum number of results to return (default: 10)

Example:

{
  "query": "latest advancements in AI",
  "maxResults": 5
}

Output: A formatted list of search results with titles, URLs, and snippets.

📄 UrlContentExtractor

Fetches and extracts clean, readable content from web pages by removing unnecessary elements like scripts, styles, and navigation.

Input:

  • url: Either a single URL string or an array of URL strings

Example (single URL):

{
  "url": "https://example.com/article"
}

Example (multiple URLs):

{
  "url": [
    "https://example.com/article1",
    "https://example.com/article2"
  ]
}

Output: Extracted text content from the specified URL(s).

🛠️ Development

# Clone the repository
git clone https://github.com/pinkpixel-dev/web-scout-mcp.git
cd web-scout-mcp

# Install dependencies
npm install

# Build
npm run build

# Run
npm start

📚 Documentation

For more detailed information about the project, check out these resources:

📋 Requirements

  • Node.js >= 18.0.0

  • npm or yarn

📄 License

This project is licensed under the Apache 2.0 License.

Available Tools

2 tools
DuckDuckGoWebSearchCInspect

Initiates a web search query using the DuckDuckGo search engine and returns a well-structured list of findings. Input the keywords, question, or topic you want to search for using DuckDuckGo as your query. Input the maximum number of search entries you'd like to receive using maxResults - defaults to 10 if not provided.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch query string
maxResultsNoMaximum number of results to return (default: 10)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the tool 'returns a well-structured list of findings' and defaults maxResults to 10, but lacks details on rate limits, authentication needs, error handling, or what 'well-structured' entails. For a search tool with no annotation coverage, this leaves significant gaps in understanding its operational behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized with two sentences that efficiently cover the tool's function and parameters. It is front-loaded with the core purpose, though the second sentence could be slightly more streamlined by avoiding repetition of 'Input'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (2 parameters, no output schema, no annotations), the description is adequate but incomplete. It covers the basic purpose and parameters but lacks details on output format, error cases, and behavioral traits. Without annotations or an output schema, more context on what 'well-structured list' means would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both parameters (query and maxResults) fully documented in the schema. The description adds minimal value beyond the schema: it reiterates that query is for 'keywords, question, or topic' and notes the default for maxResults, which is already in the schema. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Initiates a web search query using the DuckDuckGo search engine and returns a well-structured list of findings.' This specifies the verb ('initiates a web search'), resource ('web search query'), and engine ('DuckDuckGo'), distinguishing it from the sibling tool UrlContentExtractor. However, it doesn't explicitly contrast with the sibling beyond mentioning the engine.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like UrlContentExtractor. It mentions the query input and maxResults default but offers no context about appropriate use cases, prerequisites, or exclusions. Usage is implied through parameter descriptions but not explicitly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

UrlContentExtractorCInspect

Fetches and extracts content from a given webpage URL. Input the URL of the webpage you want to extract content from as a string using the url parameter. You can also input an array of URLs to fetch content from multiple pages at once.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL or list of URLs to fetch

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool fetches and extracts content but lacks details on potential issues like rate limits, authentication needs, error handling, or what 'extracts content' entails (e.g., text, HTML, metadata). This leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with the core purpose stated first. Both sentences are relevant, but the second sentence could be slightly more concise by combining the single and multiple URL explanations without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete for a tool that performs web content extraction. It doesn't explain what 'extracts content' means in terms of output format, potential limitations (e.g., JavaScript-rendered content), or error scenarios, leaving the agent with insufficient context for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already fully documents the 'url' parameter as a string or array of URIs. The description adds minimal value by restating this in plain language without providing additional context, such as URL format constraints or performance implications of array inputs, aligning with the baseline score for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('fetches and extracts content') and resource ('from a given webpage URL'), making it easy to understand what it does. However, it doesn't explicitly differentiate from its sibling tool DuckDuckGoWebSearch, which likely serves a different search-oriented purpose rather than direct content extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, such as DuckDuckGoWebSearch. It mentions the ability to handle single or multiple URLs but doesn't clarify scenarios where one might prefer this over other tools or when it's inappropriate to use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 2 tool updates
    • First observedDuckDuckGoWebSearch
    • First observedUrlContentExtractor

TDQS

B3/5.0
Disambiguation5/5

The two tools have completely distinct purposes: DuckDuckGoWebSearch performs web searches to find URLs, while UrlContentExtractor extracts content from specific URLs. There is no overlap in functionality or ambiguity about when to use each tool.

Naming Consistency4/5

Both tools use descriptive, multi-word names that clearly indicate their function. While not following a strict verb_noun pattern, they maintain readability and consistency in style. The minor deviation from perfect pattern consistency prevents a score of 5.

Tool Count2/5

With only 2 tools for a web search and content extraction server, the surface feels thin and incomplete. A typical web search server would benefit from additional tools like advanced search filters, result pagination, or content analysis utilities to provide more comprehensive coverage.

Completeness2/5

While the basic search-to-extract workflow is covered, there are significant gaps in the web search domain. Missing operations include search result filtering, handling pagination, saving/search history, content summarization, or image/video search capabilities that would be expected in a complete web search toolkit.

Maintenance

ActivitySlowing
ResponsivenessSyncing

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables web search across multiple search engines (DuckDuckGo, Bing, Startpage) with parallel execution and result deduplication. Also provides web page content extraction capabilities.
    2
    -
  • A
    license
    B
    quality
    D
    maintenance
    Enables web search through DuckDuckGo and webpage content fetching with intelligent text extraction. Features built-in rate limiting and LLM-optimized result formatting for seamless integration with language models.
    2
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables web search without API keys using DuckDuckGo and Bing search engines, and retrieves webpage content. Supports multiple search engines simultaneously with privacy protection and asynchronous processing.
    2
    9
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/pinkpixel-dev/web-scout-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server