Document & FinTech Parser MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Document & FinTech Parser MCPextract the multi-page revenue table from annual_report.pdf as JSON"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Document & FinTech Parser MCP
Resilient streaming PDF table extraction, high-res page rasterization, AcroForm flattening, and OCR bounding box dewarp repair.
Built specifically for FinTech developers, legal AI agents, PDF ingestion pipelines, and enterprise automation bots.
⚡ Quickstart
Smithery Install
smithery skill add whambammy/document-fintech-parser-mcpClaude Desktop / Cursor (claude_desktop_config.json)
{
"mcpServers": {
"document-fintech-parser-mcp": {
"command": "npx",
"args": ["-y", "@whambammy/document-fintech-parser-mcp"],
"env": {
"PAYMENT_WALLET": "0x9793E7269b3301893318dEa8338576Ba612F39B3",
"BASE_RPC_URL": "https://mainnet.base.org"
}
}
}
}Related MCP server: document-to-json-mcp
🛠️ Included Tools
Tool Name | Price (USDC) | Capability |
| $0.040 | Extracts multi-page financial tables from complex PDFs with merged cells, ruled/unruled borders, and wrapped column baselines without misaligning cells. |
| $0.035 | Rasterizes complex vector PDF pages into crisp, 300 DPI antialiased WebP/PNG images optimized for multi-modal vision LLMs with zero text clipping. |
| $0.035 | Performs geometric perspective dewarping for photographed and scanned paper documents, correcting camera skew, tilt, and binding curvature. |
| $0.035 | Flattens interactive AcroForms and dynamic XML XFA forms into static, immutable PDF pages with 100% field content preservation for optical validation. |
🔄 End-to-End Workflow
A financial agent receives a scanned loan agreement -> rasterizes pages at 300 DPI -> repairs warped OCR bounding boxes -> extracts financial tables into structured JSON -> flattens any interactive signature forms.
💰 The x402 Base L2 Micropayment Protocol
When an agent invokes a tool without payment, the server responds with a deterministic HTTP 402 Payment Required challenge containing:
Target tool price in USDC
Base Native USDC Contract:
0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913Recipient payout wallet address
Single-use cryptographic nonce
Once broadcasted on Base L2, resubmitting with paymentSignature unlocks deterministic execution.
📄 License
MIT License. Created by Whambammy.
Available Tools
4 toolsocr_bounding_box_dewarp_repairA
Performs geometric perspective dewarping for photographed and scanned paper documents, correcting camera skew, tilt, and binding curvature. (0.035 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It mentions a cost (0.035 USDC) but does not explain what happens to the input, what output is returned, whether any data is persisted, or any side effects. The transformation behavior is implied by purpose but not explicitly disclosed as a behavioral trait.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the core purpose and adds the cost. It is not verbose, and every word adds value. Minor deduction because the cost might have been better placed elsewhere, but overall structure is good.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex image-processing tool with no output schema and no annotations, the description is incomplete. It does not specify required input format, output representation, size limits, or any preconditions. The agent would have to guess how to construct the payload and what to expect in return, making it inadequate for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both parameters have descriptions), so the baseline is 3. The description adds no parameter-specific guidance; it does not clarify what the 'payload' should contain (e.g., image data format) or how paymentSignature relates to the operation. Thus it stays at baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Performs geometric perspective dewarping'), the target resource ('photographed and scanned paper documents'), and the specific corrections ('camera skew, tilt, and binding curvature'). This is specific enough to distinguish it from sibling tools, none of which mention dewarping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description defines a clear usage context: it is for photographed or scanned paper documents that need geometric correction. This implicitly tells an agent when to use it, though it does not name alternatives or state when not to use it, which would push it to 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_form_xfa_acroform_flattenerA
Flattens interactive AcroForms and dynamic XML XFA forms into static, immutable PDF pages with 100% field content preservation for optical validation. (0.035 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It communicates that the output is static and immutable, guarantees 100% field content preservation, and discloses the 0.035 USDC cost. It does not detail edge-case behavior, but the key transformation is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one efficiently packed sentence with no redundant words. It front-loads the core action, then adds the preservation guarantee and pricing context, all of which are useful for tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is sufficient for an agent to recognize what the tool does and that it costs money, but the payload structure is opaque and there is no output schema to clarify return behavior. An agent may know to select this tool but would need external knowledge to construct a correct payload.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but both parameter descriptions are generic ('Input parameters or JSON string payload', 'Base L2 USDC micropayment signature'). The tool description adds no further meaning about what the payload must contain, so it neither improves nor worsens the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Flattens') and names precise resources (interactive AcroForms, dynamic XML XFA forms) plus the resulting output (static, immutable PDF pages). It clearly distinguishes this tool from form-parsing siblings like parse_pdf_form_fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for optical validation' implies a use case, but the description does not explicitly state when to prefer this tool over alternatives such as parse_pdf_form_fields or pdf_page_rasterizer_highres. No when-to-use or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_page_rasterizer_highresA
Rasterizes complex vector PDF pages into crisp, 300 DPI antialiased WebP/PNG images optimized for multi-modal vision LLMs with zero text clipping. (0.035 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the output format, resolution, anti-aliasing, and a cost of 0.035 USDC on Base L2, which gives some insight into the payment requirement. However, it does not explain how to provide the payment signature, whether paymentSignature is mandatory, or what happens on failure (e.g., invalid PDF, payment rejection). The description adds value but lacks depth on transactional and error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence followed by a parenthetical cost note. It front-loads the primary purpose and includes key details (resolution, format, anti-aliasing, target use case) without any wasted words. The structure is efficient and easily scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a payment aspect and an opaque payload parameter, yet the description does not explain how to construct the payload (e.g., file path, base64, URL) or which page(s) to rasterize. It also omits any return value specification, since there is no output schema. The cost is mentioned but the payment mechanism and requirement are not clarified. These are significant gaps that could prevent an agent from calling the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (payload and paymentSignature), so the baseline is 3. The description does not add any additional meaning to the parameters beyond what the schema states. The payload parameter is described generically as 'Input parameters or JSON string payload for the tool execution,' which does not specify how to pass the PDF content or page selection. The description does not compensate for this ambiguity, so it remains at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb (rasterizes), a resource (complex vector PDF pages), and the output characteristics (300 DPI antialiased WebP/PNG images). It also distinguishes itself from sibling tools like convert_svg_to_png and compress_image_webp by focusing on high-resolution PDF rasterization for vision LLMs. The purpose is unambiguous and immediately differentiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool: for high-quality rasterization of complex PDF pages into images suited for multi-modal vision LLMs. It implies the need for high resolution and zero text clipping, but it does not explicitly mention alternatives or conditions when not to use it. Since there are no exclusions or named alternative tools, it falls short of a 5 but clearly guides usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_table_stream_extractor_resilientA
Extracts multi-page financial tables from complex PDFs with merged cells, ruled/unruled borders, and wrapped column baselines without misaligning cells. (0.040 USDC on Base L2)
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | Input parameters or JSON string payload for the tool execution | |
| paymentSignature | No | Base L2 USDC micropayment signature or transaction hash for x402 settlement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does provide a behavioral promise ('without misaligning cells') and exposes the cost, but it omits input format, return shape, side effects, or limitations. Some useful behavioral context, but clear gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence defines the capability and scope, followed by a useful cost note. There is no fluff, and every element contributes to the agent's decision-making.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is high-complexity, but the description does not explain how to provide the PDF, what output structure to expect, or any usage constraints. With a generic payload schema and no output schema or annotations, the agent cannot reliably construct a correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the payload parameter is described generically as 'Input parameters or JSON string payload'. The description adds no specificity about what the payload must contain or how the PDF should be passed, so it stays at the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Extracts'), a clear resource ('multi-page financial tables from complex PDFs'), and lists distinguishing features like merged cells, ruled/unruled borders, and wrapped columns. This clearly separates it from siblings like parse_pdf_form_fields or extract_tables_from_markdown.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: complex multi-page financial PDFs with merged cells and varied borders. It also mentions the 0.040 USDC cost, which is useful. However, it does not explicitly name alternatives or state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
ocr_bounding_box_dewarp_repair - First observed
pdf_form_xfa_acroform_flattener - First observed
pdf_page_rasterizer_highres - First observed
pdf_table_stream_extractor_resilient
TDQS
Scored across 4 tools
Each tool performs a distinct operation on documents: table extraction, page rasterization, dewarping/repair, and form flattening. Boundaries are clear from the descriptions, though the rasterizer vs. table extractor split (both PDF-reading, output-different) requires reading descriptions to pick correctly.
All four names are snake_case with a consistent pattern of domain prefix + operation + qualifier (pdf_table_stream_extractor_resilient, pdf_page_rasterizer_highres). Minor deviation: one uses an 'ocr_' prefix instead of 'pdf_', and names are unusually verbose.
Four tools is lean but reasonable for a specialized document-parsing pipeline where each tool handles a distinct transformation. It sits at the low end, leaving little room for composition flexibility.
Covers table extraction, rasterization, dewarping, and form flattening, but a 'Document & FinTech Parser' would be expected to also offer plain text/content extraction, metadata or key-value parsing, and format conversion. Agents needing raw text or structured JSON output would hit a dead end.
Maintenance
Related MCP Connectors
Generate and read PDFs for AI agents: a generate_pdf and a read_pdf tool, priced per document.
Pay-per-call PDF to text, RSS/Atom to JSON, sitemap to URLs for agents. x402 USDC, Base/Algorand.
Pay-per-call (x402/USDC-Base) web + crypto data tools for AI agents: audit, extract, crypto, DeFi.
Pay-per-use tool API for AI agents. Free tier, x402 USDC micropayments, or API key.
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server for document intelligence via x402 micropayments. 6 tools: document analysis, invoice extraction, screenshot data, alt text, PII detection, sentiment analysis. Pay-per-use with USDC on Base — no API keys needed.61 npmMIT
- AlicenseAqualityDmaintenanceConvert PDFs to structured JSON. Extract invoices, bank statements, contracts, and more. Pay per call via x402 USDC.58MIT
- AlicenseAqualityBmaintenanceVerifiable document intelligence for AI agents. Extract text, tables, and structured data from PDFs and URLs. Summarize, answer questions, check claims, and translate — all with cited evidence. Store tamper-evident evidence bundles with cryptographic signatures and on-chain attestation via Base L2. Cross-document semantic search and Q&A across named collections. Pay per call with USDC2221 npm1MIT
- AlicenseNot gradedqualityAmaintenanceExtracts text and tables from PDFs for AI agents via MCP, enabling structured data retrieval from invoices, reports, and statements.78 PyPI1MIT