servo-fetch
servo-fetch is a self-contained browser engine server (powered by Servo) offering a suite of tools for fetching, crawling, and interacting with web pages via MCP or HTTP API. It leverages real JavaScript execution and CSS layout without requiring Chromium or an API key. Key capabilities:
Fetch: Retrieve content from a single URL as Markdown, JSON, HTML, plain text, or accessibility tree, with support for CSS selector extraction, pagination (
start_index), settle time for SPAs, and timeouts.Batch Fetch: Submit up to 20 URLs in parallel, returning results in completion order with per-URL options, inline failure reporting, and Markdown/JSON output.
Crawl: Perform a BFS crawl of a website (up to 500 pages, depth 10), respecting
robots.txt, with include/exclude URL glob patterns and full JavaScript execution/CSS layout for accurate content extraction.Map: Quickly discover site URLs via sitemaps and link extraction without rendering, supporting up to 100,000 URLs,
robots.txtcompliance, and glob filtering.Execute JavaScript: Evaluate a JavaScript expression on a loaded page and obtain the result alongside console messages.
Screenshot: Capture viewport or full-page PNG screenshots using a software renderer (no GPU required).
Advanced features include automatic boilerplate removal (navbars, sidebars, footers, cookie banners, modals, hidden elements) and security measures such as blocking private IPs, stripping URL credentials, disabling redirects, and output sanitization.
servo-fetch embeds the Servo browser engine. It executes JavaScript, computes CSS layout, captures screenshots with a software renderer, and extracts clean content — available as a CLI, a Rust library, a Python SDK, and a Node.js SDK.
# CLI
servo-fetch "https://example.com" # clean Markdown
servo-fetch "https://example.com" --format png -o page.png # PNG screenshot// Rust
let md = servo_fetch::markdown("https://example.com").await?;# Python
page = servo_fetch.fetch("https://example.com")
print(page.markdown)// Node.js
import { fetch } from "servo-fetch";
const md = await fetch("https://example.com");Why servo-fetch
Zero dependencies — single binary, no Chromium, no API key
Real JS execution — SpiderMonkey runs JavaScript, parallel CSS engine computes layout
Layout- and visibility-aware extraction — strips navbars, sidebars, footers by rendered position, plus cookie banners, modals, and CSS-hidden content (
opacity:0,aria-hidden, sr-only)Schema-driven JSON — declarative CSS-selector schema pulls structured data
Parallel batch fetch — multiple URLs fetched concurrently
Isolated browser sessions — one-use worker process per session keeps cookies and storage fully separated
Site crawling — BFS link traversal with robots.txt, same-site scope, and rate limiting
URL discovery — sitemap-based URL mapping without rendering (fast, lightweight)
Screenshots without GPU — software renderer captures PNG/full-page screenshots anywhere
Accessibility tree — AccessKit integration with roles, names, and bounding boxes
Agent-ready — drop-in web tool for AI agents: a built-in MCP server, or wrap the Python API as a tool in any agent framework
Related MCP server: markdown-for-agents-mcp
Performance and quality
Apple M3 Pro, versus Playwright (the typical AI-agent stack):
Benchmark | servo-fetch | playwright:optimized |
Time — static-small | ~231 ms | ~645 ms |
Time — spa-heavy | ~331 ms | ~798 ms |
Memory (peak RSS) | 51–64 MB | 300–328 MB |
Extraction quality: mean word-F1 0.819 vs Readability's 0.728 across
eight page-type fixtures, with without[] boilerplate removal at 95.0%
vs 78.6%. Direct-binary engine peers (chrome-headless-shell, Lightpanda,
curl) are opt-in.
Methodology, three-axis breakdown, per-fixture F1, and raw JSON:
benchmarks/README.md +
benchmarks/results/.
Install
Interface | Install | Docs |
CLI |
| |
Rust |
| |
Python |
| |
Node.js |
|
cargo binstall servo-fetch-cli # prebuilt binary
cargo install servo-fetch-cli # build from sourceOr download from GitHub Releases.
Linux — install runtime deps and use xvfb-run on headless servers:
sudo apt install -y libegl1 libfontconfig1 libfreetype6
xvfb-run --auto-servernum servo-fetch "https://example.com"Windows — cargo binstall does not copy sidecar files (cargo-binstall#353), so the installed servo-fetch.exe fails at startup with a missing libEGL.dll. Download the .zip from Releases instead — it bundles libEGL.dll and libGLESv2.dll.
macOS — no extra setup needed.
Quick Start
CLI
servo-fetch "https://example.com" # Markdown (default)
servo-fetch "https://example.com" --format json # Structured JSON
servo-fetch "https://example.com" --format png -o page.png # PNG screenshot
servo-fetch "https://example.com" --js "document.title" # Run JavaScript
servo-fetch "https://example.com" --schema schema.json # Schema-driven JSON
servo-fetch "https://example.com" --cookies cookies.txt # Send session cookies
servo-fetch "https://example.com" -H "X-Api-Key: KEY" # Custom request header
servo-fetch URL1 URL2 URL3 # Parallel batch
servo-fetch "https://example.com" --output page.md # Save to a single file
servo-fetch URL1 URL2 --output-dir ./out/ # Save each URL to its own file
servo-fetch crawl "https://docs.example.com" --limit 20 # Crawl a site
servo-fetch crawl URL --output-dir ./pages/ # Save each crawled page to its own file
servo-fetch map "https://example.com" # Discover URLs via sitemap
servo-fetch mcp # MCP server (stdio)
servo-fetch serve # HTTP API serverFull CLI reference → servo-fetch-cli
Rust
cargo add servo-fetch// URL → Markdown in one line (async by default; use `blocking::*` for sync)
let md = servo_fetch::markdown("https://example.com").await?;
// Fetch with options
use servo_fetch::{fetch, FetchOptions};
use std::time::Duration;
let page = fetch(&FetchOptions::new("https://example.com").timeout(Duration::from_secs(60))).await?;
println!("{}", page.html);
let md = page.markdown()?;
// Crawl a site
servo_fetch::crawl_each(
&servo_fetch::CrawlOptions::new("https://docs.example.com")
.limit(100)
.user_agent("MyBot/1.0"),
|result| match &result.outcome {
Ok(page) => println!("{}: {} chars", result.url, page.content.len()),
Err(e) => eprintln!("{}: {e}", result.url),
},
).await?;
// Discover URLs via sitemap (no rendering)
let urls = servo_fetch::map(
&servo_fetch::MapOptions::new("https://example.com").limit(1000),
).await?;
for u in &urls {
println!("{}", u.url);
}Full API reference → servo-fetch
Python
Requires Python 3.11 or later.
pip install servo-fetchimport servo_fetch
page = servo_fetch.fetch("https://example.com")
print(page.markdown)
# Schema extraction
from servo_fetch import Schema, Field
schema = Schema(
base_selector=".product",
fields=[
Field(name="title", selector="h2", type="text"),
Field(name="price", selector=".price", type="text"),
],
)
page = servo_fetch.fetch("https://shop.example.com", schema=schema)
print(page.extracted)Full API reference → bindings/python
Node.js
npm install servo-fetchimport { fetch, crawl } from "servo-fetch";
const md = await fetch("https://example.com");
for await (const page of crawl("https://docs.example.com", { limit: 50 })) {
if (page.ok) console.log(page.url, page.title);
}Or run the bundled CLI without installing:
npx servo-fetch "https://example.com"Full API reference → bindings/node
MCP Server
Built-in Model Context Protocol server with six tools: fetch,
batch_fetch, crawl, map, screenshot, and execute_js.
{
"mcpServers": {
"servo-fetch": {
"command": "servo-fetch",
"args": ["mcp"]
}
}
}Streamable HTTP: servo-fetch mcp --port 8080
Full MCP tool reference → servo-fetch-cli README
Prefer in-process tools? Wrap the Python API as agent tools — see bindings/python/examples/strands_agent.py.
HTTP API
REST endpoints for containerized deployments and HTTP clients:
servo-fetch serve # 127.0.0.1:3000
servo-fetch serve --host 0.0.0.0 --port 80 # expose to network
curl -X POST http://127.0.0.1:3000/v1/fetch \
-H 'content-type: application/json' \
-d '{"url":"https://example.com"}'Endpoints: GET /health, GET /version, POST /v1/fetch, POST /v1/batch_fetch, POST /v1/screenshot, POST /v1/execute_js, POST /v1/crawl, POST /v1/map.
Full HTTP API reference → servo-fetch-cli README
Docker
Multi-arch image on GitHub Container Registry (linux/amd64, linux/arm64):
docker run --rm -p 3000:3000 ghcr.io/konippi/servo-fetch:latest
curl -X POST http://127.0.0.1:3000/v1/fetch \
-H 'content-type: application/json' \
-d '{"url":"https://example.com"}'Runs as non-root (UID 1001). Images are signed with cosign (keyless) and published with SLSA provenance and SBOM attestations.
Agent Skills
servo-fetch ships with an Agent Skills package for AI coding agents:
npx skills add https://github.com/konippi/servo-fetch/tree/main/skills/servo-fetchSecurity
servo-fetch blocks all private and reserved IP ranges (RFC 6890), strips credentials from URLs, disables HTTP redirects to prevent SSRF bypass, and sanitizes all output against terminal escape injection (CVE-2021-42574). See SECURITY.md for details.
Limitations
Sites behind CAPTCHAs are not supported.
Contributing
See CONTRIBUTING.md for development setup, commit conventions, and PR guidelines.
License
MIT OR Apache-2.0
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- AlicenseAqualityAmaintenanceUltra-fast web fetcher and MCP server written in Rust. Fetches any URL as clean Markdown with HTTP/3, JavaScript rendering, anti-fingerprinting, browser cookie authentication from Chrome/Firefox/Brave, and 1Password integration.810MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for AI agents -- fetch any URL with full JavaScript rendering (Playwright/Chromium) and convert to clean, token-efficient markdown. Works on React, Vue, Angular, and any JS-heavy page. Includes web search, batch fetching, binary file download, LRU cache, SSRF protection, and structured output.13MIT
- AlicenseNot gradedqualityBmaintenanceA lightweight MCP server for parsing HTML, fetching URLs, rendering terminal-style screenshots, and executing JavaScript on static HTML without external dependencies.3MIT
- FlicenseNot gradedqualityBmaintenanceMCP server that retrieves bot-unfriendly page content as Markdown and screenshots as vision-ready image tiles, escalating through increasingly sophisticated extraction tiers (plain HTTP, trafilatura, TLS impersonation, headless Chromium) only as needed.
Related MCP Connectors
Scrape, crawl, map & search the web. Open-source, self-hostable Rust crawler & search for AI agents.
Free remote MCP server for fetching public web pages through a rotating proxy pool.
Zenrows MCP server — Fetch, Extract, Batch, and Browser Sessions for AI coding assistants
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/konippi/servo-fetch'
If you have feedback or need assistance with the MCP directory API, please join our Discord server