Skip to main content
Glama

wasp-mcp

Web Agent Semantic Protocol — MCP 服务器

wasp-mcp 是一个 模型上下文协议 (Model Context Protocol) 服务器,允许 Claude(或任何 MCP 客户端)通过高效、结构感知的检索查询任意网页。WASP 不会将原始 HTML 直接塞入上下文窗口,而是根据页面的标题构建一个轻量级的结构索引(即 manifest),然后仅获取与查询相关的部分内容。

结果:以极低的 Token 成本获取基于真实页面内容的回答。

请参阅 WASP 白皮书 获取完整协议规范。


工作原理

每个网页都有两个有用的层级:

  1. 结构 — 构成目录的标题和部分锚点。体积小,索引成本低。

  2. 内容 — 每个标题下的文本。完整发送成本高昂;且大部分内容与特定查询无关。

WASP 利用这种拆分,通过两级流水线进行处理:

Tier 1 — get_manifest(url)
  ↓ Try GET /.well-known/wasp.json (site-native manifest, 3 s timeout)
  ↓ Fall back: fetch HTML → parse headings → generate manifest client-side
  → Returns: structured index (headings, anchors, depth, token estimates)

Tier 2 — fetch_chunk(url, anchor)
  ↓ Resolve anchor → DOM element (getElementById → querySelector → fuzzy match)
  ↓ Extract section text via Range API / heading-sibling walk
  → Returns: plain-text body of that section only

query_page(url, query)
  ↓ get_manifest → score chunks by keyword match → fetch_chunk for top results
  ↓ Build numbered [1. Heading] context → call Claude API → inline [N] citations
  → Returns: { answer, sources[] }

对典型的教职员工个人资料进行原始全页抓取大约需要 16,700 个 Token。通过 WASP 进行相同的查询仅需约 2,700 个 Token —— 减少了 6 倍。


Related MCP server: MCP Web Research Server

安装

要求: Node.js ≥ 18,一个 Anthropic API 密钥。

git clone https://github.com/seanfeeney/wasp-mcp
cd wasp-mcp
npm install
npm run build

设置您的 API 密钥:

export ANTHROPIC_API_KEY=sk-ant-...

运行服务器(stdio 传输,适用于 Claude Desktop / Claude Code):

node dist/index.js

添加到 Claude Code

在您的 Claude Code 项目配置中将 wasp-mcp 添加为本地 MCP 服务器:

claude mcp add wasp -- node /absolute/path/to/wasp-mcp/dist/index.js

或者手动编辑 .claude/settings.json:

{
  "mcpServers": {
    "wasp": {
      "command": "node",
      "args": ["/absolute/path/to/wasp-mcp/dist/index.js"],
      "env": {
        "ANTHROPIC_API_KEY": "sk-ant-..."
      }
    }
  }
}

保存后重启 Claude Code。确认服务器已启动:

/mcp

MCP 工具

get_manifest

获取 URL 的结构索引。首先尝试站点自身的 /.well-known/wasp.json;如果失败,则回退到从获取的 HTML 中进行客户端 DOM 生成。

参数

名称

类型

必需

描述

url

string

是

页面的完整 URL

示例

get_manifest("https://engineering.tamu.edu/cse/profiles/aklappenecker.html")
{
  "wasp": "1.0",
  "url": "https://engineering.tamu.edu/cse/profiles/aklappenecker.html",
  "title": "Andreas Klappenecker — Texas A&M CSE",
  "summary": "Faculty profile for Andreas Klappenecker.",
  "keywords": ["quantum computing", "cryptography", "image processing"],
  "chunks": [
    { "id": "chunk_001", "heading": "Andreas Klappenecker", "anchor": "#wasp-001", "depth": 1, "tokens": 5, "order": 1 },
    { "id": "chunk_002", "heading": "Research Interests",   "anchor": "#wasp-002", "depth": 2, "tokens": 4, "order": 2 },
    { "id": "chunk_003", "heading": "Selected Publications","anchor": "#wasp-003", "depth": 2, "tokens": 5, "order": 3 }
  ],
  "generated": "client"
}

fetch_chunk

检索由锚点标识的单个部分的纯文本正文。锚点解析使用三阶段回退:getElementById → querySelector → 模糊标题匹配。

参数

名称

类型

必需

描述

url

string

是

页面 URL(用于缓存查找;如果未缓存则重新获取)

anchor

string

是

来自 manifest 的 CSS 锚点字符串(例如 "#research-interests")

示例

fetch_chunk(
  "https://engineering.tamu.edu/cse/profiles/aklappenecker.html",
  "#wasp-002"
)
Quantum computing, image processing, cryptography.

query_page

完整的端到端检索:构建 manifest,根据查询对分块进行评分,获取相关部分的正文,调用 Claude,并返回带引用的回答。

参数

名称

类型

必需

描述

url

string

是

要查询的页面

query

string

是

自然语言问题

provider

string

否

"claude" (默认)

"openai"

"ollama"

示例

query_page(
  "https://engineering.tamu.edu/cse/profiles/aklappenecker.html",
  "What are this professor's research interests?"
)
{
  "answer": "Professor Klappenecker's research interests are quantum computing [1], image processing [1], and cryptography [1].",
  "sources": [
    { "heading": "Research Interests", "anchor": "#wasp-002" }
  ]
}

Token 效率

方法

发送给 LLM 的 Token

示例页面

原始 HTML 抓取

~16,700

TAMU 教职员工资料

WASP query_page

~2,700

同一页面,同一查询

减少量

6.1×

Token 节省量随页面长度增加而增加。一个 50,000 Token 的文档页面在仅有 2-3 个部分相关时,可能会减少 20-40 倍的 Token。


项目结构

wasp-mcp/
  index.ts        MCP server entry — registers tools
  manifest.ts     get_manifest() — discovery + DOM generation
  chunks.ts       fetch_chunk() — anchor resolution + text extraction
  retrieval.ts    query_page() — scoring, enrichment, LLM call
  providers.ts    claude / openai / ollama provider adapters
  cache.ts        In-memory URL → { manifest, html } cache with TTL
  types.ts        Shared TypeScript types

许可证

MIT © Sean Feeney, 2026

Available Tools

3 tools
fetch_chunkA

Fetch the plain-text content of a specific section of a webpage by its CSS anchor. Call get_manifest first to discover available anchors. Uses a three-stage anchor resolution: getElementById → querySelector → fuzzy heading match.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesFully-qualified URL of the webpage
anchorYesCSS anchor of the target section (e.g. "#introduction" or "#wasp-003")

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Reveals the three-stage anchor resolution (getElementById → querySelector → fuzzy heading match), which adds transparency beyond basic description. However, no annotations are provided, and the description omits error handling, permission needs, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences with no redundancy. Purpose is front-loaded, and every sentence adds value (purpose, prerequisite, resolution algorithm).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the key usage pattern (prerequisite, resolution logic) for a 2-parameter read operation. Lacks details on return format (only says 'plain-text content') and potential edge cases, but overall sufficient for typical use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes both parameters (url and anchor) with 100% coverage. The description adds context about anchor being a CSS anchor and the resolution stages, but no significant extra meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it fetches plain-text content of a webpage section by CSS anchor. Distinguishes from siblings by mentioning the prerequisite get_manifest and the specific anchor resolution method.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call get_manifest first to discover anchors, and describes the three-stage resolution process. However, does not specify when not to use the tool or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_manifestA

Fetch the WASP structural index for a webpage. Returns a manifest with the page title, summary, keywords, language, and a list of heading sections (chunks) with their anchors and token estimates. Checks /.well-known/wasp.json first (native manifest); falls back to DOM-generated manifest if not found.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesFully-qualified URL of the webpage to index

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses the fallback mechanism (native vs DOM-generated manifest) and return fields. However, it omits error conditions, permissions, or rate limits, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences efficiently convey purpose, returns, and fallback. No redundant information. Every sentence adds value, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description fully explains what is returned (title, summary, keywords, etc.) and how the tool behaves (fallback check). For a simple one-param tool with no output schema, this is complete and informative.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and schema already describes 'url' as 'Fully-qualified URL of the webpage to index'. The description adds no new parameter semantics beyond the schema, meeting the baseline for full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Fetch the WASP structural index for a webpage' using a specific verb and resource, and enumerates return fields (title, summary, keywords, etc.). It distinguishes from sibling tools (fetch_chunk, query_page) by focusing on structural indexing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (when you need the structural index) and explains fallback behavior, but does not explicitly contrast with siblings or state when not to use. Agents can infer usage, but direct guidance is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_pageA

Ask a natural-language question about a webpage. Internally runs the full WASP two-tier retrieval pipeline: fetch manifest → score relevant chunks → fetch chunk content → call Claude API → return answer with inline citations. Requires ANTHROPIC_API_KEY environment variable (or OPENAI_API_KEY for provider=openai).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesFully-qualified URL of the webpage to query
queryYesNatural-language question to answer about the page
providerNoLLM provider to use (default: "claude"). Requires corresponding API key env var.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description adequately discloses the internal pipeline steps (fetch manifest, score chunks, etc.) and the need for API keys. However, it does not explicitly state side effects (none expected for a query) or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise: two sentences. The first sentence front-loads the primary purpose, and the second adds necessary context about the pipeline and requirements. No redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description mentions the return type ('answer with inline citations'). It covers the essential aspects for a query tool, though it could mention potential error cases or timeout behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, and the tool description does not add additional meaning beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Ask a natural-language question about a webpage.' This is a specific verb+resource combination that distinguishes it from sibling tools like fetch_chunk and get_manifest.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions required API keys ('Requires ANTHROPIC_API_KEY environment variable (or OPENAI_API_KEY for provider=openai)') but does not provide guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites beyond keys.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedfetch_chunk
    • First observedget_manifest
    • First observedquery_page

TDQS

A4.4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: get_manifest retrieves the page index, fetch_chunk gets content of a specific section, and query_page performs a full Q&A pipeline. No functional overlap.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (get_manifest, fetch_chunk, query_page), making the API predictable and easy to navigate.

Tool Count5/5

Three tools is appropriate for the server's purpose of indexing and querying webpages. Each tool earns its place and there are no superfluous or missing functions.

Completeness5/5

The set covers the full pipeline: manifest retrieval (get_manifest), targeted content access (fetch_chunk), and high-level question answering (query_page). No obvious gaps for the intended functionality.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers