Skip to main content
Glama
guanyuyan
by guanyuyan

web_extract

Extract structured JSON data from any webpage by defining field selectors. Provide a URL and a schema to receive only the specified fields, ready for database storage.

Instructions

抓取网页并按字段 schema 抽取结构化 JSON(字段级数据,可入库)。

适合需要"数据而非整页正文"的场景:给 URL 和字段规则,直接返回字段值, 而不是一段 compact 文本。抓取链路(分级反爬、登录态、合规)与 web_fetch 一致。

字段规则示例: {"fields": { "title": {"selector": "h1", "type": "text"}, "first_heading": {"selector": "h1", "type": "text"}, "main_link": {"selector": "a", "type": "attr", "attr": "href"}, "link_count": {"selector": "a", "type": "count"}, "tags": {"selector": ".tag", "type": "list", "list_key": "text"} }} 字段类型:text(默认,节点归一化文本)/ attr(需配 attr,取属性值)/ count(匹配节点数)/ list(取所有匹配节点,list_key 决定取值方式: text / text_trimmed / attr / html)。

Args: url: 目标网页完整 URL。 schema: 字段抽取规则字典,见上方示例。

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
schemaYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.2.0

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses that extraction uses the same pipeline as web_fetch (tiered anti-crawling, login state, compliance), that it returns field-level data rather than full-page text, and details field-type semantics such as count and list behavior. It does not cover error handling or rate limits, but it provides meaningful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and usage, then uses a compact example and a concise field-type reference. It is somewhat long, but the length is justified because the schema parameter requires documentation beyond what the input schema provides. The structure is logical and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with only two parameters, an output schema, and no annotations, the description covers the essential usage, parameter construction, and behavioral context. It explains when to use it, how to define extraction rules, and what kind of result to expect. Nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It defines url as the target complete URL and provides an extensive breakdown of the schema parameter, including a full JSON example and detailed semantics for text, attr, count, and list field types with list_key options. This more than covers the minimal input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: it scrapes a webpage and extracts structured JSON by field schema. It explicitly contrasts with web_fetch ('directly returns field values, not a compact text'), so an agent can distinguish it from the main sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a direct use-case signal ('适合需要数据而非整页正文的场景') and explains that the scraping chain is the same as web_fetch. It does not explicitly enumerate when not to use it or mention web_batch, but the context is clear enough for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools