Skip to main content
Glama

Extract Structured Data

extract_structured
Read-onlyIdempotent

Extract typed fields from document text using a caller-defined schema. Uses a quality AI model with retry logic. Use when you need specific data points from a document rather than full text. For invoices with known fields, parse_invoice (prebuilt schema) may be simpler. For general summarization, use summarize_document instead. Schema format: { "field_name": "type hint or description" } — e.g. { "contract_date": "ISO date", "party_a": "string", "penalty_usd": "number" }. Returns: { data: { : value }, data_cited: { : { value, confidence: "high"|"medium"|"low", citations: [{ quote, paragraphs[] }] } } } Example prompts:

  • "Extract the contract date, parties, and penalty amount from this agreement."

  • "Pull the vendor name, PO number, and total from this document."

  • "Get me all named fields from this form using my custom schema."

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
textYesDocument text to extract from. Obtain via extract_text or extract_url. Example: "This Service Agreement is entered into on 2025-03-15 between ACME Corp and Beta Inc..."
schemaYesField map: describe each field you want extracted with a type hint. Example: { "total_usd": "number", "vendor": "string", "invoice_date": "ISO date YYYY-MM-DD" }
max_tokensNoInput length cap (1 token ≈ 4 chars). Default ~2500 tokens. Truncates input, not output. Example: 3000

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataYes
data_citedYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added
  2. Removed
  3. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "$schema": "http://json-schema.org/draft-07/schema#",
      +  "additionalProperties": false,
      +  "properties": {
      +    "data": {
      +      "additionalProperties": {},
      +      "propertyNames": {
      +        "type": "string"
      +      },
      +      "type": "object"
      +    },
      +    "data_cited": {
      +      "additionalProperties": {
      +        "additionalProperties": false,
      +        "properties": {
      +          "citations": {
      +            "items": {
      +              "additionalProperties": false,
      +              "properties": {
      +                "artifact": {
      +                  "type": "string"
      +                },
      +                "char_end": {
      +                  "type": "number"
      +                },
      +                "char_start": {
      +                  "type": "number"
      +                },
      +                "chunk_id": {
      +                  "type": "string"
      +                },
      +                "confidence": {
      +                  "enum": [
      +                    "high",
      +                    "medium",
      +                    "low"
      +                  ],
      +                  "type": "string"
      +                },
      +                "page": {
      +                  "type": "number"
      +                },
      +                "paragraphs": {
      +                  "items": {
      +                    "type": "number"
      +                  },
      +                  "type": "array"
      +                },
      +                "quote": {
      +                  "type": "string"
      +                }
      +              },
      +              "required": [
      +                "quote",
      +                "paragraphs",
      +                "confidence"
      +              ],
      +              "type": "object"
      +            },
      +            "type": "array"
      +          },
      +          "confidence": {
      +            "enum": [
      +              "high",
      +              "medium",
      +              "low"
      +            ],
      +            "type": "string"
      +          },
      +          "value": {}
      +        },
      +        "required": [
      +          "value",
      +          "confidence",
      +          "citations"
      +        ],
      +        "type": "object"
      +      },
      +      "propertyNames": {
      +        "type": "string"
      +      },
      +      "type": "object"
      +    }
      +  },
      +  "required": [
      +    "data",
      +    "data_cited"
      +  ],
      +  "type": "object"
      +}
  4. Changed3 schema fields changed
    • changedInput schema / properties / max_tokens / description
      Previous value: -"Input length cap (1 token ≈ 4 chars). Default ~2500 tokens. Truncates input, not output."New value: +"Input length cap (1 token ≈ 4 chars). Default ~2500 tokens. Truncates input, not output. Example: 3000"
    • changedInput schema / properties / schema / description
      Previous value: -"Field map: { \"field_name\": \"type hint\" }, e.g. { \"total_usd\": \"number\", \"vendor\": \"string\", \"invoice_date\": \"ISO date YYYY-MM-DD\" }"New value: +"Field map: describe each field you want extracted with a type hint. Example: { \"total_usd\": \"number\", \"vendor\": \"string\", \"invoice_date\": \"ISO date YYYY-MM-DD\" }"
    • changedInput schema / properties / text / description
      Previous value: -"Document text to extract from. Obtain via extract_text or extract_url."New value: +"Document text to extract from. Obtain via extract_text or extract_url. Example: \"This Service Agreement is entered into on 2025-03-15 between ACME Corp and Beta Inc...\""
  5. First observed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/idempotent profile, so the description earns credit by adding non-schema behavior: the quality AI model and retry logic. It also discloses the confidence/citation behavior of results, though the return shape is largely echoed from the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and well-sectioned, with purpose and routing before schema/return detail. The return-shape block and three example prompts are somewhat redundant given a formal output schema and a fully covered input schema, so it is slightly longer than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested-object, caller-schema-driven tool this covers the essentials: routing, schema construction, and result shape. Minor gap is the absence of failure/partial-extraction behavior (what happens when a field cannot be found), which matters more than the return-format text that the output schema already supplies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), and the description still adds real value by explaining the caller-defined schema map format with a concrete example and by clarifying schema is a field-to-type-hint map rather than a JSON Schema document. It does not extend to the max_tokens truncation semantics, which live only in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Extract typed fields from document text') and scopes it with 'using a caller-defined schema', which cleanly separates it from sibling extractors. It explicitly names the confusable siblings it is not (parse_invoice, summarize_document).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit condition ('when you need specific data points rather than full text') plus two named alternatives with the conditions that select them: parse_invoice for known invoice fields, summarize_document for summarization. No inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources