Skip to main content
Glama

Web Speed

Web Speed 解决了 AI 智能体的信噪比问题。现代网页是为人类视觉优化的(混乱的 HTML、复杂的布局、重度依赖 JS 的界面),而 Web Speed 将这种混乱转化为确定性的、Token 高效的结构化地图,专为高吞吐量的智能体集群设计。

内部无 AI。 没有 anthropic,没有 openai,没有任何 LLM 依赖。所有的解释工作都在调用智能体中完成。


为什么存在

问题

Web Speed 解决方案

原始 HTML 包含超过 150,000 个字符的脚本、样式和 SVG 噪声

去除所有非结构化内容 → 最高可减少 97% 的 Token

LLM 会产生元素 ID 幻觉,并在原始 DOM 中遗漏交互点

返回冻结的结构化地图 — 存在即所见,绝无虚构

自定义爬虫在不同站点上会失效

确定性协议 — 网页上每个站点都采用相同的 JSON 格式

智能体必须一次往返地重新发现页面

site_map 可在一次调用中爬取整个域名


Related MCP server: Delta-MCP

工具

工具

描述

interpret_page

完整的结构化地图:标题、导航、内容链接、表单、表格、文本、元数据

submit_form

提交表单(GET 或 POST),返回结果页面的地图

site_map

从根 URL 开始爬取,返回所有页面的组合地图

inspect_element

针对匹配 CSS 选择器的节点提供深度结构化数据

page_type

即时页面分类 — login(登录)、listing(列表)、article(文章)、form(表单)、navigation(导航)、other(其他)

invalidate_cache

删除缓存的地图,以便下一次调用获取最新数据


安装

Mac / Linux

cd web-interpreter
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Windows

cd web-interpreter
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt

运行

在本地使用 MCP 检查器进行开发:

mcp dev server.py

直接通过 stdio 运行(MCP 客户端启动它的方式):

python server.py

在 Claude / Cowork 中注册

添加到 ~/Library/Application Support/Claude/claude_desktop_config.json (Mac) 或 Windows 上的等效路径:

{
  "mcpServers": {
    "web-speed": {
      "command": "/absolute/path/to/web-interpreter/venv/bin/python",
      "args": ["/absolute/path/to/web-interpreter/server.py"]
    }
  }
}

然后退出并重新启动 Claude Desktop / Cowork。这六个工具将出现在 web-speed MCP 服务器下。


输出模式

interpret_page

{
  "url": "https://example.com/",
  "fetched_at": "2025-01-01T12:00:00Z",
  "page_type": "other",
  "title": "Example Domain",
  "description": "",
  "headings": [
    { "level": 1, "text": "Example Domain" }
  ],
  "navigation": [
    { "label": "Home", "url": "https://example.com/", "location": "header" }
  ],
  "content_links": {
    "total": 47,
    "truncated": false,
    "items": [
      { "label": "More information...", "url": "https://www.iana.org/domains/example" }
    ]
  },
  "forms": [
    {
      "id": "search",
      "action": "https://example.com/search",
      "method": "GET",
      "fields": [
        {
          "name": "q",
          "type": "text",
          "label": "Search",
          "placeholder": "Search...",
          "required": false,
          "value": ""
        },
        {
          "name": "_csrf",
          "type": "hidden",
          "label": "",
          "placeholder": "",
          "required": false,
          "value": "abc123"
        }
      ]
    }
  ],
  "tables": [
    {
      "id": "results",
      "headers": ["Name", "Price", "Stock"],
      "rows": [["Widget A", "$9.99", "In stock"]]
    }
  ],
  "text_blocks": [
    { "tag": "p", "text": "This domain is for use in illustrative examples." }
  ],
  "metadata": {
    "lang": "en",
    "canonical": "",
    "open_graph": { "title": "", "description": "", "image": "" }
  }
}

关键字段:

  • navigation — 语义化 nav/header/footer 元素内的链接(站点外壳、菜单)。上限为 60 个。

  • content_links — 页面主体内的链接(文章、搜索结果、列表)。始终包含 total,以便即使在截断为 60 个时也能知道实际数量。

  • forms — 每个表单及其所有字段,CSRF 令牌在隐藏字段 value 中原样保留。

  • page_type — 根据结构推断:密码字段 → login,大量项目/链接 → listing,带有段落的 <article> → article,表单 → form,大部分为链接 → navigation。


page_type

轻量级 — 仅返回分类。页面缓存时即时返回。

{
  "url": "https://example.com/login",
  "fetched_at": "2025-01-01T12:00:00Z",
  "page_type": "login",
  "title": "Sign In"
}

submit_form

输出形状与 interpret_page 相同,用于提交后服务器跳转到的页面。

{
  "url": "https://example.com/login",
  "method": "POST",
  "fields": {
    "email": "user@example.com",
    "password": "hunter2",
    "_csrf": "abc123"
  }
}

CSRF 令牌原样放入 fields 中 — 从上一次 interpret_page 调用中 forms 数组的隐藏字段中提取。


inspect_element

针对匹配 CSS 选择器的节点提供深度结构化数据。上限为 25 个元素。

{
  "url": "https://example.com/shop",
  "selector": ".product-card",
  "matched": 48,
  "truncated": true,
  "elements": [
    {
      "tag": "div",
      "id": "product-42",
      "classes": ["product-card", "featured"],
      "text": "Widget Pro $49.99 Add to cart",
      "attributes": { "id": "product-42" },
      "links": [{ "label": "Add to cart", "url": "https://example.com/cart/add/42" }],
      "fields": [],
      "children": [
        { "tag": "h3", "text": "Widget Pro" },
        { "tag": "span", "text": "$49.99" },
        { "tag": "a", "text": "Add to cart", "href": "https://example.com/cart/add/42" }
      ]
    }
  ]
}

示例选择器:#login-form, .product-card, table.results tbody tr, nav a, [data-testid="price"]


site_map

{
  "root_url": "https://example.com",
  "crawled_at": "2025-01-01T12:00:00Z",
  "total_pages": 8,
  "pages": [
    {
      "url": "https://example.com",
      "title": "Home",
      "page_type": "navigation",
      "depth": 0,
      "links_to": ["https://example.com/about", "https://example.com/contact"]
    }
  ],
  "all_forms": [
    {
      "found_on": "https://example.com/contact",
      "id": "contact",
      "action": "https://example.com/contact/submit",
      "method": "POST",
      "fields": [
        { "name": "email", "type": "email", "label": "Your email", "placeholder": "", "required": true, "value": "" },
        { "name": "message", "type": "textarea", "label": "Message", "placeholder": "", "required": true, "value": "" }
      ]
    }
  ],
  "all_navigation": [
    { "label": "About", "url": "https://example.com/about" },
    { "label": "Contact", "url": "https://example.com/contact" }
  ]
}

invalidate_cache

{ "url": "https://example.com", "invalidated": true }

错误

工具从不抛出异常。失败时:

{
  "error": true,
  "code": "FETCH_FAILED | PARSE_FAILED | TIMEOUT | NOT_HTML",
  "message": "human-readable explanation",
  "url": "https://example.com/broken"
}

智能体应如何使用输出

浏览站点: 读取 navigation 获取站点外壳(菜单、页眉、页脚),读取 content_links 获取页面主体链接。content_links.total 会告诉你即使列表被截断,实际存在多少链接。选择符合你目标的链接并调用 interpret_page。

提交表单: 读取 forms。每个字段都有 name(发送内容)、type(期望的数据类型)、label/placeholder(用途)、required 和 value。隐藏字段(type: "hidden")携带 CSRF 令牌 — 将其 value 原样传回。构建一个扁平的 name → value 字典并调用 submit_form。

提交前分类: 当你需要分支逻辑(例如,这是登录页面还是仪表板?)而不想支付完整的 interpret_page 费用时,先调用 page_type。

深入组件: 在地图中看到了表格但想要单独的行?看到了产品列表但想要每张卡片的链接和价格?使用 CSS 选择器调用 inspect_element,无需重新加载整个页面即可获取这些特定节点的结构化详细信息。

预先规划多步工作流: 开始前调用 site_map。你可以获得每个页面的标题、类型、深度和外链,以及整个站点的所有表单 — 你可以规划整个工作流(找到登录表单、找到数据输入页面、找到提交端点),而无需进行单次往返。

page_type 是信号,而非保证: 分类是启发式的。渲染空 HTML 外壳的 JS 单页应用(SPA)通常会被归类为 other — 在 JavaScript 运行之前,密码字段不在 HTML 中。将 page_type 视为快速过滤器,然后根据实际的 forms 和 headings 进行验证。


共享注册表同步

默认情况下,你的 OSS 服务器构建的每个新页面地图都会异步贡献给 api.getwebspeed.io 的 Web Speed 共享注册表。这是一个众包飞轮 — 每个获取 URL 的智能体都会将其添加到全局缓存中,因此任何地方的下一个智能体都能获得即时响应。

这是默认开启的,而非默认关闭。 默认开启是因为贡献者越多,每个人的智能体运行速度就越快。

共享内容

仅结构化页面数据:

  • 页面类型、标题、描述

  • 标题、导航链接、内容链接

  • 表单字段名称、类型和标签(无值)

  • 表格、文本块

  • Open Graph 元数据

绝不共享: Cookie、会话令牌、表单值、JS 渲染的地图(可能包含特定于会话的登录状态)。

禁用同步

在启动服务器之前设置环境变量:

WEB_SPEED_REGISTRY_SYNC=false python server.py

或者在你的 MCP 客户端配置中:

{
  "mcpServers": {
    "web-speed": {
      "command": "/path/to/venv/bin/python",
      "args": ["/path/to/server.py"],
      "env": {
        "WEB_SPEED_REGISTRY_SYNC": "false"
      }
    }
  }
}

指向自托管注册表

如果你运行自己的托管实例,请将同步指向那里:

WEB_SPEED_REGISTRY_URL=https://your-instance.example.com python server.py

同步行为

  • 即发即弃:贡献在后台发送。无论 ping 是否成功,你的智能体请求都会全速完成。

  • 仅在缓存未命中时:本地 24 小时磁盘缓存中已有的地图不会重新发送。

  • 失败静默处理:网络错误、超时和服务器拒绝仅在 DEBUG 级别记录,绝不会浮现给智能体。


架构

URL  ──▶  fetcher.py         (httpx: 10s timeout, 5 redirects, Chrome UA
                ▼             OR Playwright headless Chromium for js=true)
          cleaner.py         (BeautifulSoup/lxml: strip noise, split nav vs content
                ▼             links, filter layout tables, deduplicate text blocks,
          structured map      infer page_type, detect auth_gated)
                ▼
          cache.py           (24h TTL, MD5 keyed JSON files in ./cache/)
                ▼
          registry_sync.py   (fire-and-forget POST to api.getwebspeed.io/v1/contribute)
                ▼
          server.py          (FastMCP: 8 tools over stdio)

没有 AI。没有解释。智能体才是大脑。


已知限制

  • JS 渲染的 SPA:通过 JavaScript(React、Vue、Angular)加载内容的页面仅返回预渲染的 HTML 外壳。由 JS 注入的密码字段、搜索结果和导航将会缺失。对可见内容使用 inspect_element,并针对重度 SPA 的目标配合浏览器自动化工具使用。

  • page_type 启发式算法:分类是结构化的且快速的,但并非万无一失。带有许多内部链接的营销页面可能被归类为 listing;带有电子邮件字段但没有密码的页面不会被归类为 login。

  • 缓存为本地磁盘:./cache/ 目录是本地的。在多进程或分布式部署中,缓存条目不会在实例间共享。对于共享缓存,请将 cache.py 替换为 Redis 或 Memcached 后端。

  • 未强制执行速率限制:Web Speed 不会限制出站请求。对于高容量智能体集群,请在服务器前放置一个速率限制代理(例如 Cloudflare、nginx)。

Related MCP Connectors

Related MCP Servers