Skip to main content
Glama

extract_structured_data

Read-onlyIdempotent

PAID CAPABILITY ($0.08 USDC per successful schema-conforming extraction via x402 v2). Fetches one public page and returns JSON fields extracted by an LLM against your own JSON Schema, re-validated against that schema before return; non-conforming output returns 422 and never settles. Single page only — no crawling or JavaScript rendering; page content is truncated to 8000 characters. This MCP call validates the target and returns the canonical x402 HTTP handoff; payment and the result are exchanged at POST https://api.santosautomation.com/v1/extract/structured with {"url": "…", "schema": {...}} — POST only, because a JSON Schema does not fit in a query string. No account or API key is required.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesA publicly reachable HTTP or HTTPS page.
schemaNoOptional here — the handoff is returned either way. Required on the paid POST: a self-contained JSON Schema (type: object, no $ref) describing the fields to extract. Max 4000 characters.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesExact URL to request, then pay for and retry.
methodYesHTTP method to use for the paid request.
networkYesCAIP-2 chain id, eip155:8453 (Base mainnet).
settlesNoWhen funds move.
protocolYesx402-v2
price_usdcYesPrice in USDC for one successful call.
payment_requiredYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed15 schema fields changed
    • changedInput schema / properties / schema / description
      Previous value: -"Self-contained JSON Schema (type: object, no $ref) describing the fields to extract. Max 4000 characters."New value: +"Optional here — the handoff is returned either way. Required on the paid POST: a self-contained JSON Schema (type: object, no $ref) describing the fields to extract. Max 4000 characters."
    • removedInput schema / properties / token
      Removed value: -{
      -  "description": "Optional verified-email token. Without it the daily free quota is keyed on the caller IP, which every caller behind that address shares — hosted agents should pass a token so each user gets their own allowance. Obtain one via POST /api/leads/verify/request then /confirm; valid 30 days.",
      -  "type": "string"
      -}
    • changedInput schema / required
      Previous value: -[
      -  "url",
      -  "schema"
      -]New value: +[
      +  "url"
      +]
    • removedOutput schema / properties / data
      Removed value: -{
      -  "description": "Fields extracted against the caller's own JSON Schema, re-validated before return.",
      -  "type": "object"
      -}
    • addedOutput schema / properties / method
      Added value: +{
      +  "description": "HTTP method to use for the paid request.",
      +  "type": "string"
      +}
    • removedOutput schema / properties / model
      Removed value: -{
      -  "type": "string"
      -}
    • addedOutput schema / properties / network
      Added value: +{
      +  "description": "CAIP-2 chain id, eip155:8453 (Base mainnet).",
      +  "type": "string"
      +}
    • addedOutput schema / properties / payment_required
      Added value: +{
      +  "const": true,
      +  "type": "boolean"
      +}
    • addedOutput schema / properties / price_usdc
      Added value: +{
      +  "description": "Price in USDC for one successful call.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / protocol
      Added value: +{
      +  "description": "x402-v2",
      +  "type": "string"
      +}
    • removedOutput schema / properties / schema_version
      Removed value: -{
      -  "type": "string"
      -}
    • addedOutput schema / properties / settles
      Added value: +{
      +  "description": "When funds move.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / url / description
      Added value: +"Exact URL to request, then pay for and retry."
    • removedOutput schema / properties / word_count
      Removed value: -{
      -  "type": "integer"
      -}
    • changedOutput schema / required
      Previous value: -[
      -  "data"
      -]New value: +[
      +  "payment_required",
      +  "protocol",
      +  "method",
      +  "url",
      +  "price_usdc",
      +  "network"
      +]
  2. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "properties": {
      +    "data": {
      +      "description": "Fields extracted against the caller's own JSON Schema, re-validated before return.",
      +      "type": "object"
      +    },
      +    "model": {
      +      "type": "string"
      +    },
      +    "schema_version": {
      +      "type": "string"
      +    },
      +    "url": {
      +      "format": "uri",
      +      "type": "string"
      +    },
      +    "word_count": {
      +      "type": "integer"
      +    }
      +  },
      +  "required": [
      +    "data"
      +  ],
      +  "type": "object"
      +}
  3. Changed1 schema field changed
    • addedInput schema / properties / token
      Added value: +{
      +  "description": "Optional verified-email token. Without it the daily free quota is keyed on the caller IP, which every caller behind that address shares — hosted agents should pass a token so each user gets their own allowance. Obtain one via POST /api/leads/verify/request then /confirm; valid 30 days.",
      +  "type": "string"
      +}
  4. Added

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds substantial behavioral context beyond annotations: the paid capability and exact price, the 8000-character truncation, the re-validation against the schema, the 422 non-conforming output behavior, and the fact that the MCP call only validates and returns a handoff while payment/result happen at the POST endpoint. This is rich, non-obvious behavior that an agent needs to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: pricing, behavior, constraints, handoff details, and auth requirements are all packed into a compact block. It is front-loaded with the most critical fact (paid capability) and ends with the auth note. It could arguably be split into clearer sections, but it is not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (paid x402 flow, schema validation, handoff vs. actual extraction), the description is remarkably complete. It covers the payment model, the exact POST endpoint and payload shape, the validation behavior, the truncation limit, the single-page constraint, and the auth requirement. The output schema exists, so return values need not be spelled out. Nothing an agent needs to decide whether to call this tool and how to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds meaning by explaining that 'schema' is optional on the MCP call but required on the paid POST, and by adding constraints (self-contained, type: object, no $ref, max 4000 characters) that go beyond the schema's own description. It also clarifies that 'url' must be publicly reachable. This is above the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Fetches one public page and returns JSON fields extracted by an LLM against your own JSON Schema'), names the resource (a public page + user-provided JSON Schema), and clearly distinguishes itself from siblings by noting it is single-page only, with no crawling or JavaScript rendering. It also names the paid POST endpoint, which differentiates it from free sibling tools like extract_page_markdown or summarize.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it (when you need schema-conforming structured extraction from a single public page) and what it is not for ('no crawling or JavaScript rendering', 'Single page only'). It also gives the exact HTTP handoff and POST-only constraint, and notes no account/API key is required. This is strong routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources