Skip to main content
Glama
wxs-lang

webscout-mcp

by wxs-lang

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
WEBSCOUT_AI_MODELNoAI model to use. Default: qwen2.5:7bqwen2.5:7b
WEBSCOUT_CACHE_TTLNoCache TTL in seconds. Default: 36003600
WEBSCOUT_VECTOR_DBNoVector database to use. Default: chromachroma
WEBSCOUT_AI_API_KEYNoAPI key for AI backend (e.g., OpenAI, Doubao). Required for non-Ollama backends.
WEBSCOUT_AI_BACKENDNoAI backend to use for content understanding (ollama, openai, doubao, or custom). Default: ollamaollama
WEBSCOUT_BROWSER_TYPENoBrowser type for Playwright automation (chromium, firefox, webkit). Default: chromiumchromium
WEBSCOUT_SETUP_OLLAMANoInstall and configure Ollama during setup. Default: falsefalse
WEBSCOUT_CACHE_ENABLEDNoEnable caching. Default: truetrue
WEBSCOUT_SETUP_CHROMADBNoInstall ChromaDB during setup. Default: falsefalse
WEBSCOUT_EMBEDDING_MODELNoEmbedding model name. Default: BAAI/bge-small-zh-v1.5BAAI/bge-small-zh-v1.5
WEBSCOUT_BROWSER_HEADLESSNoRun browser in headless mode. Default: truetrue
WEBSCOUT_MONITOR_INTERVALNoDefault monitoring check interval in seconds. Default: 300300
WEBSCOUT_SETUP_PLAYWRIGHTNoInstall Playwright and Chromium during setup. Default: truetrue
WEBSCOUT_EMBEDDING_BACKENDNoEmbedding backend for vector search (local, openai, custom). Default: locallocal
WEBSCOUT_MONITOR_MIN_CHANGENoMinimum change size to trigger monitoring alerts. Default: 1010
WEBSCOUT_RATE_LIMIT_ENABLEDNoEnable per-domain rate limiting. Default: truetrue
WEBSCOUT_SEARCH_MAX_RESULTSNoMaximum number of search results to return. Default: 1010
WEBSCOUT_BROWSER_BLOCK_MEDIANoBlock media resources (images, videos) in browser automation. Default: truetrue
WEBSCOUT_SEARCH_DEFAULT_BACKENDNoDefault search backend (e.g., bing, duckduckgo, google, brave). Default: bingbing

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
web_searchB

Search the web and return structured results.

SearchService uses health-based provider ranking, deterministic recovery classification, and automatic fallback across registered search providers. Empty/invalid queries are rejected without network calls. No API key required for the default HTML providers.

web_fetchA

Fetch a URL and return one window of its extracted main content.

Progressive delivery: each call returns at most max_chars chars (default 8000). If the response contains continuation.has_more=true, more content is available locally. To read the next window, call this SAME tool again with the SAME url/extract/output_format/max_chars and set start_char to the previous continuation.next_start_char. Follow-up windows are served from a local snapshot (no new HTTP fetch, extraction, or browser). Stop when has_more is false. You do not need to read the whole page unless the task requires it.

web_crawlB

Crawl a website starting from a seed URL, respecting depth and page limits.

Fetches pages concurrently within each depth level. Respects robots.txt by default.

web_extractC

Extract structured data from a web page using CSS selectors.

cache_statsA

Return cache statistics: entry count, total size, TTL, and limits.

cache_clearA

Clear all cached entries. Returns the number of entries deleted.

search_healthA

Get health report for all search backends.

Returns overall health score, per-backend status (healthy/degraded/open/half-open), circuit breaker state, and request statistics. Use this to diagnose search failures.

Reports the active SearchService health first, with legacy SearchEngine health included as fallback reference during the migration period.

metadata_extractA

Extract metadata from a web page.

Extracts JSON-LD, OpenGraph, Twitter Cards, article metadata, images, links, and other structured metadata from the page.

rss_parseA

Parse an RSS or Atom feed and return its entries.

Fetches and parses RSS 2.0, RSS 1.0, and Atom feeds. Returns feed title, description, link, and a list of entries with title, link, description, publication date, and author.

content_qualityA

Analyze content quality of a web page.

Evaluates readability scores (Flesch-Kincaid, Gunning Fog), keyword density, content structure, metadata quality, and provides actionable suggestions for improvement.

broken_linksA

Check for broken links on a web page.

Extracts all links from the page, checks each one's HTTP status, and returns a report with broken links, redirects, and link statistics.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.5/5.0

Scored across 11 tools

Disambiguation4/5

Most tools have clearly distinct purposes, e.g., web_search, web_fetch, web_crawl, rss_parse. A small amount of overlap exists between web_extract and metadata_extract, but their descriptions clarify the difference between CSS-selector extraction and metadata extraction.

Naming Consistency3/5

Names are readable and mostly use snake_case, but the pattern is mixed: some are object_verb (web_fetch, rss_parse, cache_clear), while others are noun phrases (broken_links, content_quality, search_health). The web_ prefix helps, but there is no consistent verb_noun convention.

Tool Count5/5

11 tools is well within the ideal range and each tool serves a distinct web-research purpose, from searching and fetching to crawling, parsing, extracting, and cache management. No tool feels redundant or unnecessary.

Completeness4/5

The tool surface covers the core web-scouting workflow well: search, fetch, crawl, parse RSS, extract structured data and metadata, check links, and assess quality. Minor gaps exist, such as no search pagination or browser-level automation, but agents can generally complete tasks without dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues