web-speed-oss
Web Speed
Web Speed 解决了 AI 智能体的信噪比问题。现代网页是为人类视觉优化的(混乱的 HTML、复杂的布局、重度依赖 JS 的界面),而 Web Speed 将这种混乱转化为确定性的、Token 高效的结构化地图,专为高吞吐量的智能体集群设计。
内部无 AI。 没有 anthropic,没有 openai,没有任何 LLM 依赖。所有的解释工作都在调用智能体中完成。
为什么存在
问题 | Web Speed 解决方案 |
原始 HTML 包含超过 150,000 个字符的脚本、样式和 SVG 噪声 | 去除所有非结构化内容 → 最高可减少 97% 的 Token |
LLM 会产生元素 ID 幻觉,并在原始 DOM 中遗漏交互点 | 返回冻结的结构化地图 — 存在即所见,绝无虚构 |
自定义爬虫在不同站点上会失效 | 确定性协议 — 网页上每个站点都采用相同的 JSON 格式 |
智能体必须一次往返地重新发现页面 |
|
Related MCP server: Delta-MCP
工具
工具 | 描述 |
| 完整的结构化地图:标题、导航、内容链接、表单、表格、文本、元数据 |
| 提交表单(GET 或 POST),返回结果页面的地图 |
| 从根 URL 开始爬取,返回所有页面的组合地图 |
| 针对匹配 CSS 选择器的节点提供深度结构化数据 |
| 即时页面分类 — |
| 删除缓存的地图,以便下一次调用获取最新数据 |
安装
Mac / Linux
cd web-interpreter
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txtWindows
cd web-interpreter
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt运行
在本地使用 MCP 检查器进行开发:
mcp dev server.py直接通过 stdio 运行(MCP 客户端启动它的方式):
python server.py在 Claude / Cowork 中注册
添加到 ~/Library/Application Support/Claude/claude_desktop_config.json (Mac) 或 Windows 上的等效路径:
{
"mcpServers": {
"web-speed": {
"command": "/absolute/path/to/web-interpreter/venv/bin/python",
"args": ["/absolute/path/to/web-interpreter/server.py"]
}
}
}然后退出并重新启动 Claude Desktop / Cowork。这六个工具将出现在 web-speed MCP 服务器下。
输出模式
interpret_page
{
"url": "https://example.com/",
"fetched_at": "2025-01-01T12:00:00Z",
"page_type": "other",
"title": "Example Domain",
"description": "",
"headings": [
{ "level": 1, "text": "Example Domain" }
],
"navigation": [
{ "label": "Home", "url": "https://example.com/", "location": "header" }
],
"content_links": {
"total": 47,
"truncated": false,
"items": [
{ "label": "More information...", "url": "https://www.iana.org/domains/example" }
]
},
"forms": [
{
"id": "search",
"action": "https://example.com/search",
"method": "GET",
"fields": [
{
"name": "q",
"type": "text",
"label": "Search",
"placeholder": "Search...",
"required": false,
"value": ""
},
{
"name": "_csrf",
"type": "hidden",
"label": "",
"placeholder": "",
"required": false,
"value": "abc123"
}
]
}
],
"tables": [
{
"id": "results",
"headers": ["Name", "Price", "Stock"],
"rows": [["Widget A", "$9.99", "In stock"]]
}
],
"text_blocks": [
{ "tag": "p", "text": "This domain is for use in illustrative examples." }
],
"metadata": {
"lang": "en",
"canonical": "",
"open_graph": { "title": "", "description": "", "image": "" }
}
}关键字段:
navigation— 语义化 nav/header/footer 元素内的链接(站点外壳、菜单)。上限为 60 个。content_links— 页面主体内的链接(文章、搜索结果、列表)。始终包含total,以便即使在截断为 60 个时也能知道实际数量。forms— 每个表单及其所有字段,CSRF 令牌在隐藏字段value中原样保留。page_type— 根据结构推断:密码字段 →login,大量项目/链接 →listing,带有段落的<article>→article,表单 →form,大部分为链接 →navigation。
page_type
轻量级 — 仅返回分类。页面缓存时即时返回。
{
"url": "https://example.com/login",
"fetched_at": "2025-01-01T12:00:00Z",
"page_type": "login",
"title": "Sign In"
}submit_form
输出形状与 interpret_page 相同,用于提交后服务器跳转到的页面。
{
"url": "https://example.com/login",
"method": "POST",
"fields": {
"email": "user@example.com",
"password": "hunter2",
"_csrf": "abc123"
}
}CSRF 令牌原样放入 fields 中 — 从上一次 interpret_page 调用中 forms 数组的隐藏字段中提取。
inspect_element
针对匹配 CSS 选择器的节点提供深度结构化数据。上限为 25 个元素。
{
"url": "https://example.com/shop",
"selector": ".product-card",
"matched": 48,
"truncated": true,
"elements": [
{
"tag": "div",
"id": "product-42",
"classes": ["product-card", "featured"],
"text": "Widget Pro $49.99 Add to cart",
"attributes": { "id": "product-42" },
"links": [{ "label": "Add to cart", "url": "https://example.com/cart/add/42" }],
"fields": [],
"children": [
{ "tag": "h3", "text": "Widget Pro" },
{ "tag": "span", "text": "$49.99" },
{ "tag": "a", "text": "Add to cart", "href": "https://example.com/cart/add/42" }
]
}
]
}示例选择器:#login-form, .product-card, table.results tbody tr, nav a, [data-testid="price"]
site_map
{
"root_url": "https://example.com",
"crawled_at": "2025-01-01T12:00:00Z",
"total_pages": 8,
"pages": [
{
"url": "https://example.com",
"title": "Home",
"page_type": "navigation",
"depth": 0,
"links_to": ["https://example.com/about", "https://example.com/contact"]
}
],
"all_forms": [
{
"found_on": "https://example.com/contact",
"id": "contact",
"action": "https://example.com/contact/submit",
"method": "POST",
"fields": [
{ "name": "email", "type": "email", "label": "Your email", "placeholder": "", "required": true, "value": "" },
{ "name": "message", "type": "textarea", "label": "Message", "placeholder": "", "required": true, "value": "" }
]
}
],
"all_navigation": [
{ "label": "About", "url": "https://example.com/about" },
{ "label": "Contact", "url": "https://example.com/contact" }
]
}invalidate_cache
{ "url": "https://example.com", "invalidated": true }错误
工具从不抛出异常。失败时:
{
"error": true,
"code": "FETCH_FAILED | PARSE_FAILED | TIMEOUT | NOT_HTML",
"message": "human-readable explanation",
"url": "https://example.com/broken"
}智能体应如何使用输出
浏览站点:
读取 navigation 获取站点外壳(菜单、页眉、页脚),读取 content_links 获取页面主体链接。content_links.total 会告诉你即使列表被截断,实际存在多少链接。选择符合你目标的链接并调用 interpret_page。
提交表单:
读取 forms。每个字段都有 name(发送内容)、type(期望的数据类型)、label/placeholder(用途)、required 和 value。隐藏字段(type: "hidden")携带 CSRF 令牌 — 将其 value 原样传回。构建一个扁平的 name → value 字典并调用 submit_form。
提交前分类:
当你需要分支逻辑(例如,这是登录页面还是仪表板?)而不想支付完整的 interpret_page 费用时,先调用 page_type。
深入组件:
在地图中看到了表格但想要单独的行?看到了产品列表但想要每张卡片的链接和价格?使用 CSS 选择器调用 inspect_element,无需重新加载整个页面即可获取这些特定节点的结构化详细信息。
预先规划多步工作流:
开始前调用 site_map。你可以获得每个页面的标题、类型、深度和外链,以及整个站点的所有表单 — 你可以规划整个工作流(找到登录表单、找到数据输入页面、找到提交端点),而无需进行单次往返。
page_type 是信号,而非保证:
分类是启发式的。渲染空 HTML 外壳的 JS 单页应用(SPA)通常会被归类为 other — 在 JavaScript 运行之前,密码字段不在 HTML 中。将 page_type 视为快速过滤器,然后根据实际的 forms 和 headings 进行验证。
共享注册表同步
默认情况下,你的 OSS 服务器构建的每个新页面地图都会异步贡献给 api.getwebspeed.io 的 Web Speed 共享注册表。这是一个众包飞轮 — 每个获取 URL 的智能体都会将其添加到全局缓存中,因此任何地方的下一个智能体都能获得即时响应。
这是默认开启的,而非默认关闭。 默认开启是因为贡献者越多,每个人的智能体运行速度就越快。
共享内容
仅结构化页面数据:
页面类型、标题、描述
标题、导航链接、内容链接
表单字段名称、类型和标签(无值)
表格、文本块
Open Graph 元数据
绝不共享: Cookie、会话令牌、表单值、JS 渲染的地图(可能包含特定于会话的登录状态)。
禁用同步
在启动服务器之前设置环境变量:
WEB_SPEED_REGISTRY_SYNC=false python server.py或者在你的 MCP 客户端配置中:
{
"mcpServers": {
"web-speed": {
"command": "/path/to/venv/bin/python",
"args": ["/path/to/server.py"],
"env": {
"WEB_SPEED_REGISTRY_SYNC": "false"
}
}
}
}指向自托管注册表
如果你运行自己的托管实例,请将同步指向那里:
WEB_SPEED_REGISTRY_URL=https://your-instance.example.com python server.py同步行为
即发即弃:贡献在后台发送。无论 ping 是否成功,你的智能体请求都会全速完成。
仅在缓存未命中时:本地 24 小时磁盘缓存中已有的地图不会重新发送。
失败静默处理:网络错误、超时和服务器拒绝仅在 DEBUG 级别记录,绝不会浮现给智能体。
架构
URL ──▶ fetcher.py (httpx: 10s timeout, 5 redirects, Chrome UA
▼ OR Playwright headless Chromium for js=true)
cleaner.py (BeautifulSoup/lxml: strip noise, split nav vs content
▼ links, filter layout tables, deduplicate text blocks,
structured map infer page_type, detect auth_gated)
▼
cache.py (24h TTL, MD5 keyed JSON files in ./cache/)
▼
registry_sync.py (fire-and-forget POST to api.getwebspeed.io/v1/contribute)
▼
server.py (FastMCP: 8 tools over stdio)没有 AI。没有解释。智能体才是大脑。
已知限制
JS 渲染的 SPA:通过 JavaScript(React、Vue、Angular)加载内容的页面仅返回预渲染的 HTML 外壳。由 JS 注入的密码字段、搜索结果和导航将会缺失。对可见内容使用
inspect_element,并针对重度 SPA 的目标配合浏览器自动化工具使用。page_type启发式算法:分类是结构化的且快速的,但并非万无一失。带有许多内部链接的营销页面可能被归类为listing;带有电子邮件字段但没有密码的页面不会被归类为login。缓存为本地磁盘:
./cache/目录是本地的。在多进程或分布式部署中,缓存条目不会在实例间共享。对于共享缓存,请将cache.py替换为 Redis 或 Memcached 后端。未强制执行速率限制:Web Speed 不会限制出站请求。对于高容量智能体集群,请在服务器前放置一个速率限制代理(例如 Cloudflare、nginx)。
This server cannot be deployed
Maintenance
Related MCP Connectors
Agentic identity trust: precision decisioning, cryptographic release tokens, hash-chained proof
Paid token risk and security intelligence for AI agents over MCP with x402 payments.
The MCP gateway with an EU-hosted, persistent memory layer that shrinks your token bill.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Related MCP Servers
- AlicenseAqualityAmaintenanceThe MCP for the Web Speed Agent SDK that enables post-auth agents.1936 PyPI3GPL 3.0
- AlicenseNot gradedqualityCmaintenanceToken-efficient MCP reimplementation with progressive tool discovery, result handling, and compact wire encoding, reducing token usage by up to 89% on tool definitions.1MIT
- AlicenseBqualityBmaintenanceEnables AI agents to access design system tokens and component contracts through MCP, reducing token usage and ensuring consistency.29MIT
- AlicenseNot gradedqualityDmaintenanceConsolidates code understanding, documentation, browser automation, memory, and knowledge graph into a single MCP server with progressive discovery for up to 98% token reduction.Apache 2.0