Scout MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Scout MCPget the weight and dimensions from https://mikrotik.com/product/RB433AH"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Scout MCP
Scout is an MCP server for building a web crawler that gets better with use. Its job is to get good information from a website.
Scout is an HonestBot. It scouts the open web without tricks and never pretends to be a person:
Every request carries one plain user agent:
ScoutMCP/0.1.0 (+https://github.com/sireggserroneous/scout-mcp). Scout refuses to start if you set a user agent that looks like a browser.It obeys
robots.txtandCrawl-delay, and sends at most one request per second to a host.It does not solve CAPTCHAs, rotate identities, use proxies, or log in for you.
If a site says no, Scout reports that and suggests an honest next step. It does not try to get around the refusal.
Install
Claude Code (plugin: MCP server and the /scout skill):
/plugin marketplace add sireggserroneous/scout-mcp
/plugin install scout@scout-mcpThen type /scout followed by the url you want to scout:
/scout https://mikrotik.com
/scout https://mikrotik.com/product/RB433AH weight|dimensionsAny other MCP client. Scout needs uv. Add this to the client's MCP config:
{ "mcpServers": { "scout": { "command": "uvx",
"args": ["--from", "git+https://github.com/sireggserroneous/scout-mcp", "scout-mcp"] } } }Pages built by JavaScript. Scout needs a headless browser to read these. Add "--with", "playwright" to the
uvx args (in Claude Code, the plugin's .mcp.json), then download Chromium once:
uvx playwright install chromiumIf you already run a self-hosted Firecrawl or
crawl4ai, set FIRECRAWL_URL or CRAWL4AI_URL and Scout adds it as a
reader. It sends the same honest user agent through every reader.
Installing Scout gives you three things at once:
A crawler. It reads sitemaps and walks a site's links.
A reader. It fetches pages directly, through a headless browser, or through your Firecrawl or crawl4ai.
A recipe book. It keeps every route it learns.
You do not have to work out each site by hand.
Related MCP server: mcp-web-search
How it works: a maze, a Markov chain, and a register of rewrite rules
Reaching a page is a maze. One site answers a plain fetch. Another is an empty JavaScript shell until a browser
renders it. A third keeps its details on /specifications, one hop past the url you were given. Most crawlers
retry these dead ends every time.
Scout treats each way through as a move:
Readers:
direct,firecrawl,crawl4ai,browserurl_rewrite: a regex on the full url, such as a print view or a public archive copylisting_path: where a site keeps its index of things, such as/products,/catalogor/docsdetail_suffix: where the details live, such as/specificationsor/specs
The idea is borrowed from Markov chains and from MarkovJunior-style rewrite rules. The state is where Scout has got to on the url. Each move is a transition to a new state. The end of the maze is the goal state: good information. Good information means real content (not a wall, a login page or an error page) that carries what you asked for, when you asked for something.
flowchart LR
U[url + want] --> R{robots.txt allows?}
R -- no --> E1[ROBOTS_DISALLOWED + next steps]
R -- yes --> O[host's ordered moves<br/>winner first, recent failures last]
O --> J{good information?}
J -- yes --> W[return page<br/>learn the winner]
J -- no --> N[next move: reader → url_rewrite → detail_suffix → on-page link]
N --> J
N -- maze exhausted --> E2[coded error + next steps<br/>logged for whoever evolves the register]There are two kinds of memory, both in one SQLite file:
Each host keeps an ordered list of readers. The reader that last worked goes first. A reader that failed there goes to the back for six hours. It is never dropped, because a block today may be gone tomorrow. The second time you reach an excalidraw.com page, Scout skips the plain fetch that gave it an empty shell and goes straight to the browser.
The moves register stores rewrite rules as data. Each rule has a scope (
*, a domain, or a host regex) and a score, its Laplace win rate(wins + 1) / (tries + 2). A move nobody has tried scores 0.5, so it gets a fair try. A move that wins 3 of 3 scores 0.8. A move that loses its first five tries is retired, but it stays in the table as evidence. Scout tries register moves only after its built-in moves have failed, so a new move has to earn its place.
Anyone can add a move: you, an agent, or an evaluator reading the failure log. Scout scores each new move on real traffic, so the moves that work rise to the top and the rest drop away. That is the sense in which the crawler evolves.
Point SCOUT_DB at a shared path and a team shares one memory. A route one person learns is the first move the next
person tries.
Errors that tell an AI what to do next
When Scout fails, it returns an error code, an explanation, and next steps that are concrete tool calls, plus the
tried trail it took through the maze. Any capable model can start from there:
{
"ok": false,
"error": {
"code": "WANT_MISS",
"message": "Reached real content at https://mikrotik.com/product/RB433AH (7336 chars) but nothing matched want='zzqqxx', and the sub-pages Scout tried did not either.",
"next_steps": [
"`want` is a regex matched case-insensitively against the page text: loosen it (e.g. 'weight|kg|lbs') or drop it to read the page.",
"Read `closest.markdown` below: the words on the page may differ from the ones you asked for.",
"If this site keeps details on a sub-page, teach Scout: moves(action='propose', kind='detail_suffix', scope='mikrotik.com', spec={'suffix': '/specifications'})",
"site_map('https://mikrotik.com') — pick a real url from the site's own list instead of guessing"
]
},
"tried": [{"reader": "direct", "outcome": "want_miss", "chars": 7336},
{"reader": "direct", "outcome": "not_found", "via": "detail_suffix #12 /specifications"}],
"closest": {"title": "MikroTik · RB433AH", "markdown": "…"}
}Code | Meaning | What the next steps point to |
| robots.txt says no | an official API or feed, allowed pages from |
| the host answered 429 | wait for |
| a bot challenge or consent wall | an official API, a public archive copy as a |
| 401/403/451 to an honest bot | other paths from |
| a login or paywall | the site's API with your own key (outside Scout), public copies |
| 404, or an error page |
|
| the page is built by script | installing the browser reader, the site's JSON endpoints, a print view |
| nothing readable came back | a rendering reader, |
| real content, but not what you wanted | loosening |
| transport problems | the cause (for example a missing intermediate certificate), retrying later |
| the map came up empty |
|
Every failure is also written to a log (recipe() with no host shows it). That log is where the next move to add
should come from.
Tools
Tool | What it does |
| The whole run: reach the page, map the site's url patterns, find its listing, and return what was learned |
| Reach one page and judge whether it is good information |
| The site's own urls (sitemaps from robots.txt, then a polite link walk) and their url patterns |
| List, propose or retire rewrite rules |
| What Scout knows about a host: reader order, routes, moves, recent failures |
want is a case-insensitive regex. Pages longer than 6,000 characters come back as an excerpt, plus windows around
each want match. Pass full=true to get the whole page.
Configuration
Variable | Default | Purpose |
|
| the memory; share the path to share what is learned |
|
| your own bot name and contact; browser-like values are rejected |
|
| minimum seconds between requests to one host |
| unset | add a Firecrawl reader |
| unset | add a crawl4ai reader |
Develop
uv venv && uv pip install -e '.[browser]'
python -m scout_mcp.scout # self-checks, no networkScout is one module (scout_mcp/scout.py) and a thin MCP wrapper (scout_mcp/server.py). The seed moves live in
scout_mcp/recipes.json.
MIT licensed.
This server cannot be deployed
Maintenance
Related MCP Connectors
Web scraping for agents. Point it at a URL and it returns the page as clean markdown, JavaScript-rendered pages included. Point it at a site and it maps the URLs or crawls the section you need in the background, a few pages at a time so results fit in the conversation. Search the web and read full pages, extract fields with a JSON schema you define (validated, never invented), read a store's catalogue or a blog's posts from the platform's own feed, and check whether a page has changed. Failed requests cost nothing. The free plan includes 1,500 credits a month.
Agent-readiness scanner (0-5 score), robots.txt + llms.txt generators, managed agent enablement.
Deterministic web intake and data utilities for autonomous agents.
Reliable web access for AI agents: smart HTTP, rotating proxies, and full-browser rendering.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables AI agents to read web pages reliably, returning clean markdown content, hyperlinks, and metadata without navigation or ad noise.36 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to search the web, read and extract content from webpages, fetch JSON from REST APIs, and collect links while bypassing anti-bot protections.51 npmISC
- AlicenseNot gradedqualityCmaintenanceEnables agents to extract article text from public URLs or HTML, verify results against acceptance contracts, and fall back across extractors while ranking future attempts using local observed outcomes.MIT
- AlicenseNot gradedqualityAmaintenanceEnables agents to save reusable extraction recipes for public web pages, inspect page structures, and fetch fresh JSON or TSV data over plain HTTP, with a headless browser fallback when needed.977 npmMIT