Skip to main content
Glama
mysleekdesigns

CrawlForge MCP Server

extract_structured

Read-onlyIdempotent

Extract structured data from any webpage using a JSON schema when you know the fields but not the selectors. Returns clean JSON for product details, job listings, and event data.

Instructions

Use this to get a specific data shape from a page using a JSON schema when you can describe the fields but not their selectors - product details, job listings, event data. Uses an LLM by default; falls back to CSS selectors when no LLM is configured. Not for pages with stable markup whose selectors you know (scrape_structured, no LLM, cheaper). Cost: 3 credits. Example: extract_structured({url: "https://jobs.example.com/post/123", schema: {properties: {title: {type:"string"}, salary: {type:"string"}}, required:["title"]}})

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to extract structured data from
promptNoNatural language instructions for extraction
schemaYesJSON schema defining the data structure to extract
llmConfigNoLLM provider configuration for AI-powered extraction
user_agentNoOverride the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.
selectorHintsNoCSS selector hints to guide extraction
respect_robotsNoRespect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.
verify_numbersNoNumeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page.
fallbackToSelectorsNoFall back to CSS selector extraction if LLM is unavailable

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNo
dataNoExtracted fields matching the requested schema
_costNoCost-transparency metadata (D3.5), present when injected into the text copy of the result
errorNo
successNoFalse when the extraction errored, or a required field came back missing, empty, or in the wrong shape
confidenceNo
provenanceNo
validationNo
schema_usedNo
processingTimeNo
extractionNotesNo
extraction_methodNo"llm" | "css_fallback" | "keyword_fallback" | "none"

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv6.5.0
    • changedOutput schema / properties / success / description
      Previous value: -"False when the extraction errored or a required field came back missing or empty"New value: +"False when the extraction errored, or a required field came back missing, empty, or in the wrong shape"
  2. Changed11 schema fields changedv6.0.0
    • removedInput schema / additionalProperties
      Removed value: -false
    • removedInput schema / properties / llmConfig / additionalProperties
      Removed value: -false
    • removedInput schema / properties / schema / additionalProperties
      Removed value: -false
    • addedInput schema / properties / schema / properties / properties / propertyNames
      Added value: +{
      +  "type": "string"
      +}
    • addedInput schema / properties / selectorHints / propertyNames
      Added value: +{
      +  "type": "string"
      +}
    • changedOutput schema / properties / _cost / additionalProperties
      Previous value: -trueNew value: +{}
    • addedOutput schema / properties / data / propertyNames
      Added value: +{
      +  "type": "string"
      +}
    • changedOutput schema / properties / provenance / additionalProperties
      Previous value: -trueNew value: +{}
    • changedOutput schema / properties / provenance / properties / unverified / items / additionalProperties
      Previous value: -trueNew value: +{}
    • addedOutput schema / properties / schema_used / propertyNames
      Added value: +{
      +  "type": "string"
      +}
    • changedOutput schema / properties / validation / additionalProperties
      Previous value: -trueNew value: +{}
  3. Changed5 schema fields changedv5.4.0
    • addedInput schema / properties / respect_robots
      Added value: +{
      +  "description": "Respect the target site's robots.txt (default: true). Setting this to false is honoured, returns a warning in the response, and is recorded against your API key — it is your decision, not a silent default.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / user_agent
      Added value: +{
      +  "description": "Override the outbound User-Agent. CrawlForge identifies itself honestly by default; use this only for targets you have your own agreement with.",
      +  "type": "string"
      +}
    • addedInput schema / properties / verify_numbers
      Added value: +{
      +  "default": true,
      +  "description": "Numeric provenance guard (default: true): every price or numeric value the LLM returns must appear literally in the page source, else it is returned as null with a reason in `provenance.unverified`. Set false to get the model's raw numbers back, including ones it derived (a count, a sum, a total) rather than read off the page.",
      +  "type": "boolean"
      +}
    • addedOutput schema / properties / provenance
      Added value: +{
      +  "additionalProperties": true,
      +  "properties": {
      +    "enabled": {
      +      "description": "Whether the numeric provenance guard ran",
      +      "type": "boolean"
      +    },
      +    "nulled": {
      +      "description": "Numeric values replaced with null because the source does not contain them",
      +      "type": "number"
      +    },
      +    "skipped": {
      +      "description": "\"empty_source\" when there was nothing to check against",
      +      "type": "string"
      +    },
      +    "unverified": {
      +      "items": {
      +        "additionalProperties": true,
      +        "properties": {
      +          "path": {
      +            "description": "Path to the field, e.g. configurations[2].price",
      +            "type": "string"
      +          },
      +          "reason": {
      +            "description": "\"not_found_in_source\"",
      +            "type": "string"
      +          },
      +          "value": {
      +            "description": "The value that was removed"
      +          }
      +        },
      +        "type": "object"
      +      },
      +      "type": "array"
      +    },
      +    "verified": {
      +      "description": "Numeric values found literally in the page source",
      +      "type": "number"
      +    }
      +  },
      +  "type": "object"
      +}
    • addedOutput schema / properties / success
      Added value: +{
      +  "description": "False when the extraction errored or a required field came back missing or empty",
      +  "type": "boolean"
      +}
  4. Changed1 schema field changedv5.1.0
    • changedOutput schema / properties / extraction_method / description
      Previous value: -"\"llm\" | \"css_fallback\" | \"none\""New value: +"\"llm\" | \"css_fallback\" | \"keyword_fallback\" | \"none\""
  5. Changed2 schema fields changedv5.0.4
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "$schema": "https://json-schema.org/draft/2020-12/schema",
      +  "additionalProperties": false,
      +  "properties": {
      +    "_cost": {
      +      "additionalProperties": true,
      +      "description": "Cost-transparency metadata (D3.5), present when injected into the text copy of the result",
      +      "properties": {
      +        "actual": {
      +          "description": "Credits actually charged (0 in creator mode, half-rate on error)",
      +          "type": "number"
      +        },
      +        "projected": {
      +          "description": "Credits projected for this call before execution",
      +          "type": "number"
      +        },
      +        "projection_note": {
      +          "description": "Human-readable note about how the cost was projected",
      +          "type": "string"
      +        },
      +        "remaining_credits": {
      +          "description": "Credits remaining on the account after this call, if known",
      +          "type": [
      +            "number",
      +            "null"
      +          ]
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "confidence": {
      +      "type": "number"
      +    },
      +    "data": {
      +      "additionalProperties": {},
      +      "description": "Extracted fields matching the requested schema",
      +      "type": "object"
      +    },
      +    "error": {
      +      "type": "string"
      +    },
      +    "extractionNotes": {
      +      "items": {
      +        "type": "string"
      +      },
      +      "type": "array"
      +    },
      +    "extraction_method": {
      +      "description": "\"llm\" | \"css_fallback\" | \"none\"",
      +      "type": "string"
      +    },
      +    "processingTime": {
      +      "type": "number"
      +    },
      +    "schema_used": {
      +      "additionalProperties": {},
      +      "type": "object"
      +    },
      +    "url": {
      +      "type": "string"
      +    },
      +    "validation": {
      +      "additionalProperties": true,
      +      "properties": {
      +        "errors": {
      +          "items": {
      +            "type": "string"
      +          },
      +          "type": "array"
      +        },
      +        "valid": {
      +          "type": "boolean"
      +        }
      +      },
      +      "type": "object"
      +    }
      +  },
      +  "type": "object"
      +}
  6. First observedv4.10.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already cover readOnly, idempotent, and non-destructive traits. The description adds genuinely useful behavioral disclosure beyond those: 'Uses an LLM by default; falls back to CSS selectors when no LLM is configured' and 'Cost: 3 credits.' There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear purpose and uses only three sentences plus an example. Every sentence earns its place: the when-to-use rule, a priority alternative, the cost, and a complete invocation example. Nothing is redundant or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description does not need to explain return values. It fully covers how to select this tool versus alternatives, how to invoke it (example), the cost, the default LLM behavior, and its fallback path. Combined with the 100% schema coverage, this is a complete definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the example in the description demonstrates the `url` and `schema` parameters in context. However, the description mostly restates the schema's own parameter descriptions and does not add deeper explanation for parameters like `llmConfig`, `selectorHints`, or `verify_numbers` beyond what the input schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description's first sentence states exactly what the tool does: 'Use this to get a specific data shape from a page using a JSON schema.' It gives the precise condition 'when you can describe the fields but not their selectors' and cites concrete use cases (product details, job listings, event data), which clearly distinguishes it from the sibling scrape_structured.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly sets when to use it ('when you can describe the fields but not their selectors') and when not to use it ('Not for pages with stable markup whose selectors you know') and directly names the alternative: scrape_structured. It also notes the default LLM use and the fallback path, so an agent is not left to guess.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.