Skip to main content
Glama

Recognise a passport or ID document

scan_document
Idempotent

Read passports, national ID cards, and driver's licences from an image and return the printed fields, MRZ checks, and authenticity results as structured JSON.

Instructions

Recognise a passport, national ID card or driver's licence from a photo or scan and return what is printed on it as structured JSON. Inputs: the image as image_base64 (always available), image_path (a local file, and only inside the directory DOC_CHEAP_IMAGE_ROOT names) or image_url (https, on a public address); plus the optional expect_country, return_portrait, retain_hours, reference and idempotency_key. Output: a Scan object – meta (id, status, billed, confidence, timing), document (kind, issuing country, number, series, date of issue, date of expiry, whether it has expired and how many days are left), holder (given names, surname, date of birth, sex, nationality), fields (every field read off the printed page, each with its own confidence), mrz (whether the machine-readable zone checks out, why not when it does not, and its lines exactly as read), images, quality and authenticity – plus a one-line summary of the same result. Calls POST /v1/scans. Cost: it bills one credit ($0.01) only when a document is recognised; an unreadable image, an empty frame or an unsupported type costs nothing, and meta.billed says which happened. Without a key, the public sandbox key is used. It gives 10 free recognised documents per address in all, and at most 10 requests per address an hour, whatever their answer. Registering gives 100 free documents every month. Use it whenever someone hands over an identity document and wants it read, transcribed, or checked against what they claim – a name, a document number, a date of birth or an expiry date.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
image_urlNohttps: URL of an image on a public internet address, which the server fetches (25 MB maximum).
referenceNoYour own correlation string, echoed back in the result.
image_pathNoPath to a local image file, inside the directory named by DOC_CHEAP_IMAGE_ROOT. Disabled unless that variable is set; send image_base64 instead.
image_base64NoThe document image as base64 (a data: URL is also accepted).
retain_hoursNoHours the result stays readable via GET /v1/scans/{id} (0 = store nothing). Omit it to use the account's own history-retention setting.
expect_countryNoISO 3166-1 alpha-3 country you expect, or omit for any.
idempotency_keyNoMakes a retried scan return the first result instead of charging again.
return_portraitNoWhether to include the holder photograph crop, images.main_photo (default true).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
mrzYes
metaYes
fieldsYesEvery field the engine extracted off the printed document, re-keyed to our vocabulary – the open set. Always present; empty when nothing was extracted. A field read in more than one language appears once per language, so `name` repeats and only `id` is unique.
holderYes
imagesYes
qualityYes
documentYes
authenticityYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.3.7

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds extensive behavioral context beyond the annotations: billing rules (one credit only when a document is recognised, free for unreadable images), rate limits (10 free per address, 10 requests per address per hour), sandbox key fallback, retention via retain_hours, and idempotency key behaviour. Annotations already declare openWorld and idempotent, but the description makes the practical implications concrete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and inputs, then covers output, cost, and rate limits in a logical order. It is lengthy and spends several sentences explaining return values that are already covered by the output schema, but the extra detail is mostly relevant behavioural context rather than pure waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with eight parameters, an output schema, and non-trivial billing and rate-limit rules, the description covers everything an agent needs: input options, output shape summary, cost model, authentication fallback, retention, and the core use case. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description lists the input options and frames the optional parameters, but largely repeats what the schema already documents (e.g. image_path directory restriction, https public URL for image_url, retain_hours semantics). It does not add syntax or format details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (recognise) and resource (passport, national ID card, driver's licence from a photo or scan) and clearly distinguishes itself from the unrelated siblings check_balance and search_docs. An agent can tell exactly what the tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use it whenever someone hands over an identity document and wants it read, transcribed, or checked against what they claim' — a clear usage context. It does not state when not to use it or name alternatives, but for this tool the alternatives are not relevant given the sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools