Skip to main content
Glama
nooot77
by nooot77

ScanToBill MCP Server


What is this?

ScanToBill is an AI-powered OCR platform for Arabic and English business documents — invoices, receipts, contracts, and IDs. This package connects ScanToBill to Claude Desktop (or any MCP-compatible AI client) so you can extract invoice data just by sharing a URL with Claude.

You: "Extract the line items from this invoice: https://example.com/invoice.pdf"

Claude: Here's what I found:
  - Consulting services: SAR 5,000
  - VAT (15%): SAR 750
  - Total: SAR 5,750

  Vendor: Al-Noor Trading Co.
  Date: 2026-07-15
  Invoice #: INV-2026-0042

Related MCP server: Mistral OCR MCP Server

Quick start

npx @scantobill/mcp-server --api-key ocr_sk_YOUR_KEY

Get a free API key at app.scantobill.com/keys — 20 extractions/month, no credit card.


Claude Desktop setup

1. Open the config file

OS

Path

macOS

~/Library/Application Support/Claude/claude_desktop_config.json

Windows

%APPDATA%\Claude\claude_desktop_config.json

2. Add this entry

{
  "mcpServers": {
    "scantobill": {
      "command": "npx",
      "args": ["-y", "@scantobill/mcp-server", "--api-key", "ocr_sk_YOUR_KEY"]
    }
  }
}

3. Restart Claude Desktop — done.


Available tools

Tool

Description

extract_document

Extract structured JSON from an invoice, receipt, or business document at a public URL

list_extractions

List past extractions with optional date, type, and status filters

get_extraction

Fetch full metadata for a specific extraction by ID

extract_document

url            (required) Public URL — PDF, JPG, PNG, or TIFF
document_type  invoice · receipt · contract · id  (default: invoice)
language       ar · en · auto  (default: auto)

list_extractions

limit    Number of results (default: 10, max: 50)
status   pending · processing · done · failed
from     ISO 8601 start date
to       ISO 8601 end date

get_extraction

id    (required) Extraction ID returned by extract_document

Authentication

# CLI flag
npx @scantobill/mcp-server --api-key ocr_sk_YOUR_KEY

# Environment variable
SCANTOBILL_API_KEY=ocr_sk_YOUR_KEY npx @scantobill/mcp-server

Supported documents

Type

Arabic

English

Mixed

Invoice (فاتورة)

Receipt (إيصال)

Contract (عقد)

National ID

VAT invoice (ZATCA)


Pricing

Plan

Extractions/month

Price

Free

20

$0

Starter

500

$15

Pro

5,000

$49

Enterprise

Unlimited

Custom


License

MIT © ScanToBill

Available Tools

3 tools
extract_documentA

Extract structured data from an invoice, receipt, commercial register, VAT certificate, or other business document at a public URL. Returns JSON with all detected fields including vendor, customer, line items, totals, VAT, and currency. Each extraction consumes one unit from the authenticated user's plan quota.

ParametersJSON Schema
NameRequiredDescriptionDefault
languageNoLanguage hint (ar=Arabic, en=English, fr=French, es=Spanish). Omit for automatic detection.
document_urlYesPublic URL of the document to extract (JPEG, PNG, WebP, or PDF). Must be publicly accessible — no authentication required to fetch.
document_typeNoDocument subtype hint. Use 'inventory' for itemised invoices with line items, 'general' for service invoices, 'pos' for POS receipts, 'auto' to let the API decide (default).

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description adds valuable behavioral context: 'Each extraction consumes one unit from the authenticated user's plan quota' and 'Returns JSON with all detected fields...'. This goes beyond the schema by revealing a side effect (quota usage) and the output structure. It does not contradict any annotations (none given), though it could also disclose whether the operation is asynchronous or how errors are surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinctly written in three sentences: purpose, output, and quota cost. It front-loads the primary action, contains no redundant text, and each sentence contributes substantive information. This is an example of concise, well-structured writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description covers core functionality, return format, and quota implications. However, it fails to explain the relationship with sibling tools – specifically whether the extraction is synchronous or if the result must be retrieved later via get_extraction or list_extractions. This omission is significant for an agent deciding on the correct sequence of tool calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes all three parameters (100% coverage), so the baseline for this dimension is 3. The description reinforces that document_url must be public and lists output fields, but does not add any additional semantic meaning for the language or document_type parameters beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies 'Extract structured data from an invoice, receipt, commercial register, VAT certificate, or other business document at a public URL' – a clear verb+resource. It implies a distinct action from the retrieval-oriented siblings (get_extraction, list_extractions) by describing the creation of an extraction and its quota consumption, but does not explicitly differentiate by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for extracting data from public documents, but provides no explicit guidance on when to use this tool versus the sibling tools. It mentions the requirement of a public URL and quota consumption, yet lacks clear 'use this for X, use get_extraction for Y' instructions, leaving the agent to infer the division of responsibilities.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_extractionA

Get the metadata for a specific past extraction by its log ID. Returns the document type, OCR provider used, status, and timestamp.

ParametersJSON Schema
NameRequiredDescriptionDefault
extraction_idYesThe extraction log ID returned by list_extractions.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the transparency burden. It does disclose the return fields (document type, OCR provider, status, timestamp), which is helpful. However, it does not describe edge-case behavior such as what happens when an invalid or nonexistent extraction_id is provided (e.g., error, null response), nor does it state any read-only guarantee or prerequisite beyond having the ID.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that gets straight to the point. It states the action, the resource, the identifier, and the expected return fields without any filler or redundant information. Front-loaded and efficiently sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no annotations, no output schema), the description is fairly complete. It explains what the tool does and explicitly lists the returned fields, which is essential since there is no output schema. It lacks only edge-case behavior (e.g., error on bad ID) and explicit exclusions, but those are minor for a simple read operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for the single parameter extraction_id, so the baseline is 3. The description adds no additional semantics beyond restating that it's a log ID, and the schema already explains its origin from list_extractions. The description's phrase 'by its log ID' aligns with the schema without adding new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get'), the resource ('metadata for a specific past extraction'), and the identifier ('by its log ID'). It explicitly differentiates from sibling tools: 'list_extractions' lists all extractions, while this tool retrieves a specific one by ID.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when you need metadata for a single extraction identified by its log ID, as opposed to listing all extractions. The schema description for extraction_id reinforces that it is returned by list_extractions, providing clear context. It does not explicitly state exclusions or alternative tools, but the 'specific' vs. 'list' distinction is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_extractionsA

List past extraction log entries for the authenticated user. Returns extraction metadata (ID, type, provider, date, status) sorted newest first. Use get_extraction for the full details of a specific entry.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of results to return (1–100). Default: 20.
to_dateNoReturn only extractions on or before this date (YYYY-MM-DD).
from_dateNoReturn only extractions on or after this date (YYYY-MM-DD).
document_typeNoFilter by document type: auto, inventory, general, or pos.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It states the return format (metadata fields), sorting order (newest first), and user scoping, which are important behavioral traits. It does not mention pagination or error behavior, but for a read-only list operation this is adequate. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose, then a valuable pointer to the sibling. No redundant or extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description explains the return metadata and sorting. It also provides the key alternative for deeper details. The optional filters are fully described in the schema, making this description sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for all four optional parameters, so the schema handles parameter semantics. The description does not add any parameter-specific details beyond the schema, which is acceptable per baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'List' with the resource 'past extraction log entries' and scope 'for the authenticated user'. It also distinguishes from the sibling get_extraction by noting that get_extraction provides full details, making the purpose of this tool clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly names get_extraction as the alternative for full details, providing clear guidance on when to use this tool (listing metadata) vs. sibling. It also implies that this tool is for log/history retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.1
    • First observedextract_document
    • First observedget_extraction
    • First observedlist_extractions

TDQS

A4.1/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: extracting a new document, listing past extractions, and fetching a specific extraction's metadata. No overlap or ambiguity exists.

Naming Consistency5/5

All tool names follow the same verb_noun pattern with lowercase and underscores: get_extraction, extract_document, list_extractions. Consistent and predictable.

Tool Count4/5

Three tools is slightly minimal but appropriate for a focused extraction service. Each tool serves a distinct need in the extraction workflow, though additional operations like delete or re-extract could be added.

Completeness4/5

The set covers the core lifecycle: create (extract_document), list (list_extractions), and get (get_extraction). Missing update/delete operations, but extractions are likely immutable, so this is acceptable. A tool to check quota would be a minor addition.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

Appeared in Searches