Document & FinTech Parser MCP
<p align="center">
<img src="./assets/logo.png" width="130" height="130" alt="Document & FinTech Parser MCP Logo" />
</p>
# Document & FinTech Parser MCP
[](https://smithery.ai)
[](https://modelcontextprotocol.io)
[](https://base.org)
[](#included-tools)
[](LICENSE)
**Resilient streaming PDF table extraction, high-res page rasterization, AcroForm flattening, and OCR bounding box dewarp repair.**
Built specifically for FinTech developers, legal AI agents, PDF ingestion pipelines, and enterprise automation bots.
---
## ⚡ Quickstart
### Smithery Install
```bash
smithery skill add whambammy/document-fintech-parser-mcp
```
### Claude Desktop / Cursor (`claude_desktop_config.json`)
```json
{
"mcpServers": {
"document-fintech-parser-mcp": {
"command": "npx",
"args": ["-y", "@whambammy/document-fintech-parser-mcp"],
"env": {
"PAYMENT_WALLET": "0x9793E7269b3301893318dEa8338576Ba612F39B3",
"BASE_RPC_URL": "https://mainnet.base.org"
}
}
}
}
```
---
## 🛠️ Included Tools
| Tool Name | Price (USDC) | Capability |
| :--- | :---: | :--- |
| `pdf_table_stream_extractor_resilient` | $0.040 | Extracts multi-page financial tables from complex PDFs with merged cells, ruled/unruled borders, and wrapped column baselines without misaligning cells. |
| `pdf_page_rasterizer_highres` | $0.035 | Rasterizes complex vector PDF pages into crisp, 300 DPI antialiased WebP/PNG images optimized for multi-modal vision LLMs with zero text clipping. |
| `ocr_bounding_box_dewarp_repair` | $0.035 | Performs geometric perspective dewarping for photographed and scanned paper documents, correcting camera skew, tilt, and binding curvature. |
| `pdf_form_xfa_acroform_flattener` | $0.035 | Flattens interactive AcroForms and dynamic XML XFA forms into static, immutable PDF pages with 100% field content preservation for optical validation. |
---
## 🔄 End-to-End Workflow
A financial agent receives a scanned loan agreement -> rasterizes pages at 300 DPI -> repairs warped OCR bounding boxes -> extracts financial tables into structured JSON -> flattens any interactive signature forms.
---
## 💰 The x402 Base L2 Micropayment Protocol
When an agent invokes a tool without payment, the server responds with a deterministic `HTTP 402 Payment Required` challenge containing:
- Target tool price in USDC
- Base Native USDC Contract: `0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913`
- Recipient payout wallet address
- Single-use cryptographic nonce
Once broadcasted on Base L2, resubmitting with `paymentSignature` unlocks deterministic execution.
---
## 📄 License
MIT License. Created by [Whambammy](https://github.com/Whambammy).
TDQS
Scored across 4 tools
Each tool performs a distinct operation on documents: table extraction, page rasterization, dewarping/repair, and form flattening. Boundaries are clear from the descriptions, though the rasterizer vs. table extractor split (both PDF-reading, output-different) requires reading descriptions to pick correctly.
All four names are snake_case with a consistent pattern of domain prefix + operation + qualifier (pdf_table_stream_extractor_resilient, pdf_page_rasterizer_highres). Minor deviation: one uses an 'ocr_' prefix instead of 'pdf_', and names are unusually verbose.
Four tools is lean but reasonable for a specialized document-parsing pipeline where each tool handles a distinct transformation. It sits at the low end, leaving little room for composition flexibility.
Covers table extraction, rasterization, dewarping, and form flattening, but a 'Document & FinTech Parser' would be expected to also offer plain text/content extraction, metadata or key-value parsing, and format conversion. Agents needing raw text or structured JSON output would hit a dead end.