Process a business document end-to-end
process_documentExtract invoices, contracts, and ID documents from text, Markdown, URLs, or images into typed JSON with validation in one call.
Instructions
Single-call pipeline: provided text/Markdown or Mistral OCR → classify (if kind=auto) → typed extraction → validation. source.type=text skips OCR and Files uploads. Text with kind=generic makes no API calls; classification and typed extraction use Mistral chat. Results expose extraction_source. Provided text has null ocr_confidence and page_count; ocr_text contains the supplied text unchanged. Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.
Kinds: contract | invoice | id_document | generic. Use kind=auto to let the server classify.
Returns a discriminated union — switch on kind to access typed fields.
Validation checks schema and, for OCR sources, OCR confidence; not factual or accounting accuracy.
Typed extraction rejects text longer than 60000 characters rather than truncating it.
Cache keys include source, kind, page limit, endpoint, models and pipeline version. Override location with MISTRAL_MCP_CACHE_DIR. Override mode with options.cache. Default cache mode is 'read_write' EXCEPT for kind=id_document (auto-bypass to avoid persisting PII). Set options.cache='read_write' explicitly to opt in for id documents.
options.maxPages and options.minOcrConfidence apply only to OCR sources. The confidence floor defaults to 0.3. Below the floor the
tool returns isError. Missing or partial confidence scores also return isError;
use mistral_ocr directly if you need raw OCR without a confidence guarantee.
0.3 is a conservative starting point, not a measured one: calibrate it for your
corpus with npm run eval:docs.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Extraction task. auto classifies the document; generic returns text without typed extraction. Text source with generic makes no API calls. | auto |
| source | Yes | Already extracted text/Markdown, or an OCR source: remote URL, uploaded file ID, or inline image. | |
| options | No | Page selection, OCR confidence floor and local cache policy. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| dob | No | ||
| kind | Yes | ||
| name | No | ||
| total | No | ||
| expiry | No | ||
| vendor | No | ||
| clauses | No | ||
| country | No | ||
| parties | No | ||
| summary | No | ||
| currency | No | ||
| due_date | No | ||
| ocr_text | Yes | Text used for extraction: provided text unchanged, or Markdown returned by Mistral OCR. | |
| anomalies | No | ||
| cache_hit | Yes | ||
| key_dates | No | ||
| source_id | Yes | ||
| line_items | No | ||
| page_count | Yes | Pages processed by Mistral OCR. Null for provided text, whose pagination is unknown. | |
| risk_score | No | ||
| document_type | No | ||
| ocr_confidence | Yes | Mean Mistral OCR page confidence. Null for provided text; never an extraction accuracy score. | |
| structured_text | No | ||
| pipeline_version | Yes | ||
| extraction_source | Yes | How the input text was obtained. provided_text is supplied by the caller, not verified by OCR. | |
| total_duration_ms | Yes |