Skip to main content
Glama

Extract schema-validated fields

tab_extract_structured
Read-only

Extract up to 30 typed fields from a live page using CSS selectors, validate them against JSON Schema draft-07, and return each value with source URL, selector, and quote provenance.

Instructions

Read named fields from observed DOM using selectors, validate strict JSON Schema draft-07, and return per-field source URL/selector/quote provenance. Supports text, attributes, current non-sensitive values, typed scalars and arrays. At most 30 fields, 20 matches per field, 100 total. Does not call a model or infer missing facts; hidden/password controls and truncated evidence are rejected.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fieldsYes
schemaYes
session_idYesSession ID returned by tab_open.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the read-only and non-destructive annotations, the description adds meaningful behavioral detail: hard limits (30 fields, 20 matches per field, 100 total), explicit rejection of hidden/password/truncated evidence, and the guarantee that no model or inference is used. This gives an agent a strong, accurate picture of what will and won't happen.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and efficient: the core action and return value are front-loaded, followed by supported modes, limits, and explicit non-behaviors. Every sentence conveys essential operational information without fluff or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficiently complete for a 3-parameter tool with no output schema: it explains what is returned (per-field source URL, selector, quote provenance), what constraints apply, and what is rejected. An agent can reasonably decide whether to call it and understand the expected result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, so the description must compensate. It does: it clarifies that 'fields' can use text, attribute, and value modes, supports typed scalars/arrays, and imposes field-count limits. It doesn't explain the 'required' or 'multiple' flags explicitly, but the schema already enumerates those and the description's added semantics are useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read'), a precise resource ('named fields from observed DOM using selectors'), and the differentiating value-add: JSON Schema draft-07 validation and per-field provenance. This clearly separates it from siblings like tab_extract and tab_find by emphasizing schema validation, typed extraction, and evidence provenance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys when to use this tool: when you need schema-validated, selector-based field extraction with provenance, and not when you need model-based inference ('Does not call a model or infer missing facts'). It does not explicitly name alternative sibling tools or contrast them, but the context is clear and the exclusions (hidden/password controls, truncated evidence) are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.