web_extract
Extract structured JSON data from any webpage by defining field selectors. Provide a URL and a schema to receive only the specified fields, ready for database storage.
Instructions
抓取网页并按字段 schema 抽取结构化 JSON(字段级数据,可入库)。
适合需要"数据而非整页正文"的场景:给 URL 和字段规则,直接返回字段值, 而不是一段 compact 文本。抓取链路(分级反爬、登录态、合规)与 web_fetch 一致。
字段规则示例: {"fields": { "title": {"selector": "h1", "type": "text"}, "first_heading": {"selector": "h1", "type": "text"}, "main_link": {"selector": "a", "type": "attr", "attr": "href"}, "link_count": {"selector": "a", "type": "count"}, "tags": {"selector": ".tag", "type": "list", "list_key": "text"} }} 字段类型:text(默认,节点归一化文本)/ attr(需配 attr,取属性值)/ count(匹配节点数)/ list(取所有匹配节点,list_key 决定取值方式: text / text_trimmed / attr / html)。
Args: url: 目标网页完整 URL。 schema: 字段抽取规则字典,见上方示例。
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| schema | Yes |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |