Skip to main content
Glama

extract

Get structured data from web pages by defining a CSS-based JSON schema. Target specific elements like tables, lists, or product info and receive parsed JSON.

Instructions

Scrape + structured extraction using a CSS-based JSON schema.

Use when the user wants structured data (tables, lists, product info) extracted from a page. Define a CSS schema to target specific elements.

The schema is a JsonCssExtractionStrategy schema: { "name": "PageItems", "baseSelector": "div.item", "fields": [{"name": "title", "selector": "h2", "type": "text"}, ...] }

Returns parsed JSON in data.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
preferNoauto
schemaYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.7.1

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden. It discloses that this is a scraping/extraction operation and says it returns parsed JSON in `data`. But it omits practical behavioral details like dynamic-content handling, page-size limits, error behavior, and what the `prefer` option controls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded. The first sentence captures the essence, and the rest provides a concrete example and return-shape note without wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema example and return mention make the tool callable, but with no annotations and zero parameter descriptions, important usage context is missing—especially around `prefer` and dynamic page behavior. It is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does a good job for the `schema` parameter by providing a concrete JsonCssExtractionStrategy example, but `prefer` is not explained at all, and `url` is only meaningful by convention.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation ('Scrape + structured extraction') with a clear mechanism (CSS-based JSON schema) and names concrete use cases (tables, lists, product info). This distinguishes it from siblings like scrape or crawl, which are for less structured data collection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool: when the user wants structured data extracted from a page. However, it does not explicitly name alternatives or state when not to use it, so the contrast with sibling tools is mostly implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.