Skip to main content
Glama

Extract Clean

extract_clean
Read-only

Clean page extraction from any public URL: title, meta description, main text, JSON-LD, language. Optional schema keeps only the listed top-level keys. Apify (JS-rendered) with direct-fetch fallback. Pay per call with USDC or USDT on Base via x402 – no API key, no signup. 3 free calls per wallet.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesPublic URL to extract.
schemaNoOptional JSON schema for the output.
walletNoYour EVM wallet address (0x...). Unlocks the free tier (3 free calls per wallet).
api_keyNoOptional AMR enterprise API key.
x_paymentNoOptional signed x402 payment payload (base64 JSON) for USDC/USDT on Base.
instructionsNoOptional natural-language hints.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataYesExtracted page data (title, description, text, json_ld, language, ...).
tokenNoToken actually used for payment (USDC, USDT or FREE).USDC
routingYesHow the request was routed (provider, fallback, cache, latency).
cost_usdcYesPrice charged in stablecoin units (USDC/USDT 1:1).
confidenceYesConfidence score between 0 and 1.
fetched_atNoUTC timestamp when the data was fetched.
sources_checkedYesNumber of providers consulted.
freshness_secondsYesAge of the underlying data in seconds (0 = fetched live).
x_payment_responseNoBase64-encoded x402 settlement receipt (mirrors the X-PAYMENT-RESPONSE header). Only present when the call was paid with `x_payment`.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "description": "Response for ``POST /extract-clean``.",
      +  "properties": {
      +    "confidence": {
      +      "description": "Confidence score between 0 and 1.",
      +      "maximum": 1,
      +      "minimum": 0,
      +      "title": "Confidence",
      +      "type": "number"
      +    },
      +    "cost_usdc": {
      +      "description": "Price charged in stablecoin units (USDC/USDT 1:1).",
      +      "title": "Cost Usdc",
      +      "type": "number"
      +    },
      +    "data": {
      +      "additionalProperties": true,
      +      "description": "Extracted page data (title, description, text, json_ld, language, ...).",
      +      "title": "Data",
      +      "type": "object"
      +    },
      +    "fetched_at": {
      +      "description": "UTC timestamp when the data was fetched.",
      +      "format": "date-time",
      +      "title": "Fetched At",
      +      "type": "string"
      +    },
      +    "freshness_seconds": {
      +      "description": "Age of the underlying data in seconds (0 = fetched live).",
      +      "title": "Freshness Seconds",
      +      "type": "integer"
      +    },
      +    "routing": {
      +      "description": "How the request was routed (provider, fallback, cache, latency).",
      +      "properties": {
      +        "cache_hit": {
      +          "default": false,
      +          "description": "True if the result was served from cache.",
      +          "title": "Cache Hit",
      +          "type": "boolean"
      +        },
      +        "fallback_used": {
      +          "default": false,
      +          "description": "True if the primary provider failed and a fallback served the result.",
      +          "title": "Fallback Used",
      +          "type": "boolean"
      +        },
      +        "latency_ms": {
      +          "default": 0,
      +          "description": "End-to-end routing latency in milliseconds.",
      +          "title": "Latency Ms",
      +          "type": "integer"
      +        },
      +        "provider": {
      +          "description": "Name of the upstream provider that served the result.",
      +          "title": "Provider",
      +          "type": "string"
      +        },
      +        "providers_tried": {
      +          "description": "Providers attempted, in order.",
      +          "items": {
      +            "type": "string"
      +          },
      +          "title": "Providers Tried",
      +          "type": "array"
      +        }
      +      },
      +      "required": [
      +        "provider"
      +      ],
      +      "title": "RoutingMeta",
      +      "type": "object"
      +    },
      +    "sources_checked": {
      +      "description": "Number of providers consulted.",
      +      "title": "Sources Checked",
      +      "type": "integer"
      +    },
      +    "token": {
      +      "default": "USDC",
      +      "description": "Token actually used for payment (USDC, USDT or FREE).",
      +      "title": "Token",
      +      "type": "string"
      +    },
      +    "x_payment_response": {
      +      "description": "Base64-encoded x402 settlement receipt (mirrors the X-PAYMENT-RESPONSE header). Only present when the call was paid with `x_payment`.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "data",
      +    "sources_checked",
      +    "freshness_seconds",
      +    "confidence",
      +    "cost_usdc",
      +    "routing"
      +  ],
      +  "type": "object"
      +}
  2. Changed2 schema fields changed
    • addedInput schema / properties / url / format
      Added value: +"uri"
    • changedInput schema / properties / wallet / description
      Previous value: -"Your EVM wallet address (0x...). Unlocks the free tier (3 calls per wallet)."New value: +"Your EVM wallet address (0x...). Unlocks the free tier (3 free calls per wallet)."
  3. First observed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=true, so the safety profile is already covered. The description adds valuable behavioral context beyond annotations: it discloses the Apify (JS-rendered) with direct-fetch fallback behavior, the pay-per-call model with USDC/USDT on Base via x402, no API key/signup requirement, and the 3-free-calls-per-wallet limit. This is exactly the kind of context that helps an agent anticipate side effects (cost) and prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero waste. The core function is front-loaded, followed by the optional filter, then the rendering/fallback mechanism, then payment/free-tier details. Every sentence earns its place and no information is repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a read-only extraction tool with an output schema present. It covers the core behavior, the optional filtering, the rendering mechanism, and the payment model. The only minor gap is that it doesn't describe the output format in detail, but the output schema exists and the description names the extracted fields. The payment/free-tier context is a strong addition that most tool descriptions lack.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 6 parameters. The description adds meaning by explaining the `schema` parameter's effect ('keeps only the listed top-level keys') and the `wallet` parameter's role in unlocking the free tier. It doesn't repeat parameter names verbatim but adds behavioral context that the schema lacks. Baseline 3 is exceeded because the description clarifies the two most important parameters' semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('extract'), a specific resource ('clean page extraction from any public URL'), and enumerates the exact fields returned (title, meta description, main text, JSON-LD, language). It also distinguishes itself from siblings by mentioning the optional `schema` filter and the Apify/fallback mechanism. An agent can tell this apart from html_to_markdown, web_crawl, and url_status without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when to use it: for clean page extraction from any public URL, with optional schema filtering. It doesn't explicitly name alternatives or exclusions (e.g., 'use html_to_markdown for raw HTML'), but the context is clear enough that an agent can infer the use case. The payment/free-tier info also helps the agent decide whether to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.