Skip to main content
Glama

Extract Data from Document

talonic_extract

Extract structured, schema-validated JSON from a document: PDF, scan, image, DOCX, email, or photo. Returns the requested fields with per-field confidence scores.

USE WHEN: the user asks to extract data from a document, turn a PDF into JSON, pull fields from a file, or parse a scan / form / statement / receipt / report, for common (invoice, contract) or unusual document types. NOT FOR: full plain text (use talonic_to_markdown) · finding documents (use talonic_search / talonic_filter). BY NAME: if the user names a file, call talonic_search first to get its document_id, then call this. ARGS: define the fields you want with inline schema (JSON Schema, e.g. {type:'object',properties:{vendor_name:{type:'string'}}}) OR a saved schema_id, not both. Don't know the fields yet? Set auto_schema:true to let Talonic discover them (open capture) and return a suggested schema you can refine. Provide EXACTLY ONE document source: document_id (cheapest, a workspace doc), file_url (public URL), or file_data+filename (small local files only). COST: each call uses credits; talonic_get_balance shows the remaining balance. RETURNS: data (the JSON), confidence.overall and confidence.fields (treat <0.7 as needs review), document metadata, extraction_id.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
schemaNoInline schema definition. REQUIRED unless `schema_id` is provided. Recommended: full JSON Schema {type:'object', properties:{...}}. Also accepted: flat key-type map {field_name:'string', amount:'number'}. Mutually exclusive with `schema_id`.
file_urlNoURL to a document file. The Talonic API fetches it server-side. Use this for documents already on the public web.
filenameNoOriginal filename including extension, e.g. 'invoice.pdf'. Used to infer MIME type when uploading via `file_data`. Required when `file_data` is provided.
file_dataNoBase64-encoded file bytes. Recommended path when the agent already has the file in memory (e.g., the user attached a PDF to the conversation). Pair with `filename` so MIME type can be inferred. Works regardless of where the file lives on disk.
file_pathNoLocal path to a document file. Only works if the MCP server has read access to that path. In sandboxed chat clients (Claude Desktop, Cowork) where uploads land in a host-owned directory, use `file_data` instead.
schema_idNoID of a saved schema. REQUIRED unless `schema` is provided. Accepts UUID or SCH-XXXXXXXX short id from talonic_list_schemas. Mutually exclusive with `schema`.
auto_schemaNoOpen capture: when true, extract WITHOUT providing a schema — Talonic discovers the document's fields and returns them plus a suggested schema you can refine and reuse. Use this when you don't yet know the fields. Mutually exclusive with `schema` and `schema_id`.
document_idNoID of a document already in the workspace, to re-extract with a new schema.
instructionsNoNatural-language guidance for the extractor, e.g. 'Focus on the billing section. Amounts are in EUR.'
include_markdownNoInclude OCR-converted markdown in the response alongside structured data.
include_provenanceNoInclude per-field provenance (source_text, section, page) showing where each value was found in the document.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
costNoPer-call cost and post-call balance, parsed from the X-Talonic-* response headers. `null` for non-extract calls; not always present on legacy clients.
dataYesThe extracted structured data, shape determined by the schema.
linksNoURLs for self, document, and human-readable dashboard view.
schemaNoSchema metadata: which schema was used and how it can be saved.
statusYesExtraction status (e.g. 'complete').
documentYesMetadata about the ingested document.
markdownNoOCR-converted markdown. Present only when `include_markdown: true`.
confidenceNoExtraction confidence. Treat fields below ~0.7 as needing human review.
processingNoProcessing metadata: duration, pages processed, region.
provenanceNoPer-field source evidence (source_text, section, page). Present only when `include_provenance: true`.
request_idNoServer-assigned request ID for support and debugging.
extraction_idYesStable identifier for this extraction.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

Score is being calculated.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources