Skip to main content
Glama
guhcostan
by guhcostan

Web Search MCP

GitHub Sponsors

Minimal MCP server that can search the web and extract readable page content, similar to Cursor's built-in web context.

Features

  • search_web: Query the web (DuckDuckGo HTML) and return result URLs and titles

  • fetch_page: Fetch any URL and extract readable content using Mozilla Readability + JSDOM

Requirements

  • Node.js 20+ (recommended: 20.18.1+)

Install

npm install

Run (stdio)

npm start

Install globally

npm i -g @guhcostan/web-search-mcp

Then reference the binary web-search-mcp.

Integrate with Cursor (MCP)

Add to your Cursor MCP settings:

{
  "mcpServers": {
    "web-search-mcp": { "command": "web-search-mcp" }
  }
}

Alternatively, without global install, use npx:

{
  "mcpServers": {
    "web-search-mcp": {
      "command": "npx",
      "args": ["-y", "@guhcostan/web-search-mcp@latest"]
    }
  }
}

Tools

  • search_web

    • input:

      • query (string, required)

      • limit (number, optional, 1–10, default 5)

    • output: array of { url: string; title?: string; snippet?: string }

  • fetch_page

    • input:

      • url (string URL, required)

    • output: { url: string; title?: string; content: string }

Development

Type-check, lint and tests:

npm run check

Run individually:

npm run build
npm run lint
npm test

Notes

  • Web search uses DuckDuckGo HTML; results may vary and are HTML-scraped (no API key required)

  • Be mindful of target site terms of use and robots policies when fetching pages

License

MIT

Available Tools

2 tools
fetch_pageC

Fetch a page and extract its readable content and title using Readability.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'extract its readable content and title using Readability', which implies processing and transformation, but doesn't cover critical aspects like error handling, rate limits, authentication needs, or what happens with invalid URLs. For a tool that interacts with external resources, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded with the core action ('Fetch a page') and includes essential details ('extract its readable content and title using Readability'), making it easy to parse quickly. Every part of the sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of fetching and processing web pages, with no annotations, no output schema, and minimal schema coverage, the description is incomplete. It doesn't explain the return format (e.g., structured data with content and title), error conditions, or dependencies like network access. This leaves the agent with insufficient information to use the tool effectively in varied contexts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 1 parameter with 0% description coverage, so the description must compensate. It doesn't explicitly mention the 'url' parameter or provide any details beyond what the schema indicates (e.g., format expectations or examples). Since the schema is minimal and self-explanatory for a URL, the description adds no extra semantic value, meeting the baseline for adequate but not helpful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Fetch a page and extract its readable content and title using Readability.' This specifies the verb ('fetch'), resource ('page'), and extraction method ('Readability'), making it easy to understand what the tool does. However, it doesn't explicitly differentiate from the sibling tool 'search_web', which might have overlapping functionality for web content retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'search_web'. It lacks context about use cases, prerequisites, or exclusions, leaving the agent to infer usage based on the tool name and description alone. This omission reduces its effectiveness in helping the agent select the right tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_webC

Search the web and return a list of result URLs and titles. Uses DuckDuckGo HTML.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
limitNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It mentions the search engine (DuckDuckGo) and output format, but lacks critical behavioral details: whether this requires authentication, rate limits, network dependencies, error handling, or pagination behavior. The description doesn't contradict annotations (none exist), but provides minimal behavioral context for a tool that performs external web searches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately concise with two sentences that convey the core functionality and implementation detail. It's front-loaded with the main purpose. The second sentence about DuckDuckGo adds useful context without being verbose. No wasted words, though it could be slightly more structured with explicit parameter mentions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a web search tool with 2 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain the search scope, result format details beyond 'URLs and titles', error conditions, or performance characteristics. The mention of DuckDuckGo provides some context, but critical information about how results are returned, sorted, or filtered is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no parameter-specific information beyond what the schema provides. With 0% schema description coverage, both parameters (query and limit) are undocumented in both schema and description. However, the description implies the 'query' parameter through 'Search the web' and 'limit' through 'return a list', but doesn't explain their semantics, formats, or constraints. This meets the baseline 3 since the schema covers parameter structure, though the description doesn't compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Search the web and return a list of result URLs and titles.' It specifies the verb (search), resource (web), and output format (URLs and titles). However, it doesn't explicitly differentiate from its sibling 'fetch_page', which likely retrieves a specific page rather than performing a search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It mentions 'Uses DuckDuckGo HTML' which hints at the search engine used, but doesn't explain when to choose this over 'fetch_page' or other search methods. No explicit when/when-not instructions or alternative tool references are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.5
    • First observedfetch_page
    • First observedsearch_web

TDQS

B3.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: fetch_page retrieves and extracts content from a specific URL, while search_web performs a general web search to find relevant URLs. There is no overlap or ambiguity between them, as one is for accessing known pages and the other is for discovering new ones.

Naming Consistency5/5

Both tools follow a consistent verb_noun naming pattern (fetch_page and search_web), using clear, descriptive verbs that align with their functions. The naming is uniform and predictable across the set.

Tool Count3/5

With only two tools, the server feels thin for a web search domain, as it lacks operations like advanced search filtering, result pagination, or handling different search engines. While the tools cover basic fetch and search, the count is borderline minimal for the apparent scope.

Completeness3/5

The tools provide core search and fetch capabilities, but there are notable gaps: no ability to refine searches (e.g., by date or site), manage search history, or handle errors like rate limits. This could lead to agent workarounds for more complex tasks.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers