scantobill-mcp-server
This server integrates ScanToBill's AI-powered OCR with MCP-compatible AI clients (e.g., Claude Desktop) to extract structured JSON data from business documents and manage extraction history.
Extract structured data (
extract_document): Submit a public URL (PDF, JPEG, PNG, TIFF, WebP) to extract fields like vendor, customer, line items, totals, VAT, and currency from invoices, receipts, contracts, IDs, and other documents. Specify document type (auto, inventory, general, pos, contract, ID) and language (Arabic, English, French, Spanish, or auto). Each extraction consumes one quota unit.List past extractions (
list_extractions): Retrieve history sorted newest first, filterable by date range, document type, status, and limit (up to 100).Fetch specific extraction (
get_extraction): Retrieve full metadata (provider, status, timestamp) by ID.Easy authentication via CLI flag or environment variable, with a free tier available.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@scantobill-mcp-serverExtract the invoice data from this URL: https://example.com/invoice.pdf"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ScanToBill MCP Server
What is this?
ScanToBill is an AI-powered OCR platform for Arabic and English business documents — invoices, receipts, contracts, and IDs. This package connects ScanToBill to Claude Desktop (or any MCP-compatible AI client) so you can extract invoice data just by sharing a URL with Claude.
You: "Extract the line items from this invoice: https://example.com/invoice.pdf"
Claude: Here's what I found:
- Consulting services: SAR 5,000
- VAT (15%): SAR 750
- Total: SAR 5,750
Vendor: Al-Noor Trading Co.
Date: 2026-07-15
Invoice #: INV-2026-0042Related MCP server: Mistral OCR MCP Server
Quick start
npx @scantobill/mcp-server --api-key ocr_sk_YOUR_KEYGet a free API key at app.scantobill.com/keys — 20 extractions/month, no credit card.
Claude Desktop setup
1. Open the config file
OS | Path |
macOS |
|
Windows |
|
2. Add this entry
{
"mcpServers": {
"scantobill": {
"command": "npx",
"args": ["-y", "@scantobill/mcp-server", "--api-key", "ocr_sk_YOUR_KEY"]
}
}
}3. Restart Claude Desktop — done.
Available tools
Tool | Description |
| Extract structured JSON from an invoice, receipt, or business document at a public URL |
| List past extractions with optional date, type, and status filters |
| Fetch full metadata for a specific extraction by ID |
extract_document
url (required) Public URL — PDF, JPG, PNG, or TIFF
document_type invoice · receipt · contract · id (default: invoice)
language ar · en · auto (default: auto)list_extractions
limit Number of results (default: 10, max: 50)
status pending · processing · done · failed
from ISO 8601 start date
to ISO 8601 end dateget_extraction
id (required) Extraction ID returned by extract_documentAuthentication
# CLI flag
npx @scantobill/mcp-server --api-key ocr_sk_YOUR_KEY
# Environment variable
SCANTOBILL_API_KEY=ocr_sk_YOUR_KEY npx @scantobill/mcp-serverSupported documents
Type | Arabic | English | Mixed |
Invoice (فاتورة) | ✅ | ✅ | ✅ |
Receipt (إيصال) | ✅ | ✅ | ✅ |
Contract (عقد) | ✅ | ✅ | ✅ |
National ID | ✅ | ✅ | — |
VAT invoice (ZATCA) | ✅ | — | — |
Pricing
Plan | Extractions/month | Price |
Free | 20 | $0 |
Starter | 500 | $15 |
Pro | 5,000 | $49 |
Enterprise | Unlimited | Custom |
License
MIT © ScanToBill
Available Tools
3 toolsextract_documentA
Extract structured data from an invoice, receipt, commercial register, VAT certificate, or other business document at a public URL. Returns JSON with all detected fields including vendor, customer, line items, totals, VAT, and currency. Each extraction consumes one unit from the authenticated user's plan quota.
| Name | Required | Description | Default |
|---|---|---|---|
| language | No | Language hint (ar=Arabic, en=English, fr=French, es=Spanish). Omit for automatic detection. | |
| document_url | Yes | Public URL of the document to extract (JPEG, PNG, WebP, or PDF). Must be publicly accessible — no authentication required to fetch. | |
| document_type | No | Document subtype hint. Use 'inventory' for itemised invoices with line items, 'general' for service invoices, 'pos' for POS receipts, 'auto' to let the API decide (default). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description adds valuable behavioral context: 'Each extraction consumes one unit from the authenticated user's plan quota' and 'Returns JSON with all detected fields...'. This goes beyond the schema by revealing a side effect (quota usage) and the output structure. It does not contradict any annotations (none given), though it could also disclose whether the operation is asynchronous or how errors are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinctly written in three sentences: purpose, output, and quota cost. It front-loads the primary action, contains no redundant text, and each sentence contributes substantive information. This is an example of concise, well-structured writing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers core functionality, return format, and quota implications. However, it fails to explain the relationship with sibling tools – specifically whether the extraction is synchronous or if the result must be retrieved later via get_extraction or list_extractions. This omission is significant for an agent deciding on the correct sequence of tool calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes all three parameters (100% coverage), so the baseline for this dimension is 3. The description reinforces that document_url must be public and lists output fields, but does not add any additional semantic meaning for the language or document_type parameters beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies 'Extract structured data from an invoice, receipt, commercial register, VAT certificate, or other business document at a public URL' – a clear verb+resource. It implies a distinct action from the retrieval-oriented siblings (get_extraction, list_extractions) by describing the creation of an extraction and its quota consumption, but does not explicitly differentiate by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for extracting data from public documents, but provides no explicit guidance on when to use this tool versus the sibling tools. It mentions the requirement of a public URL and quota consumption, yet lacks clear 'use this for X, use get_extraction for Y' instructions, leaving the agent to infer the division of responsibilities.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_extractionA
Get the metadata for a specific past extraction by its log ID. Returns the document type, OCR provider used, status, and timestamp.
| Name | Required | Description | Default |
|---|---|---|---|
| extraction_id | Yes | The extraction log ID returned by list_extractions. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the transparency burden. It does disclose the return fields (document type, OCR provider, status, timestamp), which is helpful. However, it does not describe edge-case behavior such as what happens when an invalid or nonexistent extraction_id is provided (e.g., error, null response), nor does it state any read-only guarantee or prerequisite beyond having the ID.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that gets straight to the point. It states the action, the resource, the identifier, and the expected return fields without any filler or redundant information. Front-loaded and efficiently sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no annotations, no output schema), the description is fairly complete. It explains what the tool does and explicitly lists the returned fields, which is essential since there is no output schema. It lacks only edge-case behavior (e.g., error on bad ID) and explicit exclusions, but those are minor for a simple read operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for the single parameter extraction_id, so the baseline is 3. The description adds no additional semantics beyond restating that it's a log ID, and the schema already explains its origin from list_extractions. The description's phrase 'by its log ID' aligns with the schema without adding new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get'), the resource ('metadata for a specific past extraction'), and the identifier ('by its log ID'). It explicitly differentiates from sibling tools: 'list_extractions' lists all extractions, while this tool retrieves a specific one by ID.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need metadata for a single extraction identified by its log ID, as opposed to listing all extractions. The schema description for extraction_id reinforces that it is returned by list_extractions, providing clear context. It does not explicitly state exclusions or alternative tools, but the 'specific' vs. 'list' distinction is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_extractionsA
List past extraction log entries for the authenticated user. Returns extraction metadata (ID, type, provider, date, status) sorted newest first. Use get_extraction for the full details of a specific entry.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (1–100). Default: 20. | |
| to_date | No | Return only extractions on or before this date (YYYY-MM-DD). | |
| from_date | No | Return only extractions on or after this date (YYYY-MM-DD). | |
| document_type | No | Filter by document type: auto, inventory, general, or pos. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It states the return format (metadata fields), sorting order (newest first), and user scoping, which are important behavioral traits. It does not mention pagination or error behavior, but for a read-only list operation this is adequate. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, then a valuable pointer to the sibling. No redundant or extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description explains the return metadata and sorting. It also provides the key alternative for deeper details. The optional filters are fully described in the schema, making this description sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for all four optional parameters, so the schema handles parameter semantics. The description does not add any parameter-specific details beyond the schema, which is acceptable per baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'List' with the resource 'past extraction log entries' and scope 'for the authenticated user'. It also distinguishes from the sibling get_extraction by noting that get_extraction provides full details, making the purpose of this tool clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names get_extraction as the alternative for full details, providing clear guidance on when to use this tool (listing metadata) vs. sibling. It also implies that this tool is for log/history retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.1- First observed
extract_document - First observed
get_extraction - First observed
list_extractions
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: extracting a new document, listing past extractions, and fetching a specific extraction's metadata. No overlap or ambiguity exists.
All tool names follow the same verb_noun pattern with lowercase and underscores: get_extraction, extract_document, list_extractions. Consistent and predictable.
Three tools is slightly minimal but appropriate for a focused extraction service. Each tool serves a distinct need in the extraction workflow, though additional operations like delete or re-extract could be added.
The set covers the core lifecycle: create (extract_document), list (list_extractions), and get (get_extraction). Missing update/delete operations, but extractions are likely immutable, so this is acceptable. A tool to check quota would be a minor addition.
Maintenance
Related MCP Connectors
MCP server unifying ERPs, CRMs, APIs and knowledge base for Claude, ChatGPT and Gemini.
Hosted MCP server: convert PDFs to clean, LLM-ready Markdown with tables, formulas and OCR.
Hosted Amazon Seller and Vendor MCP server for Claude, ChatGPT, Cursor, Codex, Gemini, Copilot.
Hosted Amazon Seller Central and Amazon Ads MCP server for Claude, ChatGPT, Cursor, and agents.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceA Python MCP server for invoice and receipt processing that uses OCR technology to extract data from PDFs and images, offering AI assistants the ability to process, extract text from, and merge invoice documents.2-
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables Claude to perform OCR on local files using Mistral AI's document processing capabilities. It converts documents and images into markdown format for seamless analysis and interaction.-
- FlicenseNot gradedqualityDmaintenanceAn AI-powered MCP server that extracts structured data from Indian identity documents (Aadhaar, Passport, PAN, Driving License) using OCR, enabling Claude Desktop to read and process document images locally.-
- AlicenseNot gradedqualityDmaintenanceMCP server for ReceiptConverter that allows AI assistants to parse any receipt or invoice image/PDF into structured JSON with a single tool call.5MIT