Skip to main content
Glama
Swih

Mistral MCP — Document Extraction

by Swih

Process a business document end-to-end

process_document
Read-onlyIdempotent

Extract invoices, contracts, and ID documents from text, Markdown, URLs, or images into typed JSON with validation in one call.

Instructions

Single-call pipeline: provided text/Markdown or Mistral OCR → classify (if kind=auto) → typed extraction → validation. source.type=text skips OCR and Files uploads. Text with kind=generic makes no API calls; classification and typed extraction use Mistral chat. Results expose extraction_source. Provided text has null ocr_confidence and page_count; ocr_text contains the supplied text unchanged. Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.

Kinds: contract | invoice | id_document | generic. Use kind=auto to let the server classify. Returns a discriminated union — switch on kind to access typed fields. Validation checks schema and, for OCR sources, OCR confidence; not factual or accounting accuracy. Typed extraction rejects text longer than 60000 characters rather than truncating it.

Cache keys include source, kind, page limit, endpoint, models and pipeline version. Override location with MISTRAL_MCP_CACHE_DIR. Override mode with options.cache. Default cache mode is 'read_write' EXCEPT for kind=id_document (auto-bypass to avoid persisting PII). Set options.cache='read_write' explicitly to opt in for id documents.

options.maxPages and options.minOcrConfidence apply only to OCR sources. The confidence floor defaults to 0.3. Below the floor the tool returns isError. Missing or partial confidence scores also return isError; use mistral_ocr directly if you need raw OCR without a confidence guarantee. 0.3 is a conservative starting point, not a measured one: calibrate it for your corpus with npm run eval:docs.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
kindNoExtraction task. auto classifies the document; generic returns text without typed extraction. Text source with generic makes no API calls.auto
sourceYesAlready extracted text/Markdown, or an OCR source: remote URL, uploaded file ID, or inline image.
optionsNoPage selection, OCR confidence floor and local cache policy.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dobNo
kindYes
nameNo
totalNo
expiryNo
vendorNo
clausesNo
countryNo
partiesNo
summaryNo
currencyNo
due_dateNo
ocr_textYesText used for extraction: provided text unchanged, or Markdown returned by Mistral OCR.
anomaliesNo
cache_hitYes
key_datesNo
source_idYes
line_itemsNo
page_countYesPages processed by Mistral OCR. Null for provided text, whose pagination is unknown.
risk_scoreNo
document_typeNo
ocr_confidenceYesMean Mistral OCR page confidence. Null for provided text; never an extraction accuracy score.
structured_textNo
pipeline_versionYes
extraction_sourceYesHow the input text was obtained. provided_text is supplied by the caller, not verified by OCR.
total_duration_msYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.8.3

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description goes well beyond them: cache key composition and the id_document PII auto-bypass, the isError conditions around OCR confidence, that validation is schema-only (not factual), and the 60000-character rejection behavior. This is unusually rich behavioral disclosure for a read-only processing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the pipeline summary, then grouped by concerns (sources, kinds, returns, validation, cache, options). Length is defensible for a 3-param nested pipeline tool, but a few statements restate schema facts (the 60000 limit, the 0.3 floor, maxPages/minOcrConfidence scoping), which is mild redundancy given a fully described schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description still usefully explains the discriminated-union return (switch on kind, extraction_source, ocr_confidence semantics). Combined with the caveats about confidence floors and cache bypass, an agent has everything needed to call and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds framing the schema does not: that maxPages/minOcrConfidence apply only to OCR sources (and how they interact with provided text), that text sources yield null ocr_confidence/page_count, and that the 0.3 floor is a placeholder to calibrate. It adds value beyond the field-level docs, though several details (60000 chars, 0.3 default) are duplicated from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete pipeline with specific stages (classify → typed extraction → validation) and the input modalities it accepts. It also explicitly positions itself against siblings: 'Replaces the manual chain of mistral_ocr + mistral_chat + JSON parsing.' An agent can distinguish it from mistral_ocr and mistral_chat without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing guidance: kind=auto lets the server classify, source.type=text skips OCR and Files uploads, text+generic makes no API calls, and 'use mistral_ocr directly if you need raw OCR without a confidence guarantee.' It names the alternative and the condition that selects it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.