Skip to main content
Glama
628,416 tools. Updated 2026-10-02 00:23

"Tools and methods for extracting data from PDF files" matching MCP tools:

  • Concatenate two or more PDFs into a single PDF, in the order supplied, and return the merged file. Over MCP the PDF is never inlined: it comes back as a stored URL that stays retrievable for about 24 hours. Page content is copied unchanged — it does not compress (use compress_pdf) or select pages (use split_pdf). Fewer than two files is rejected. Every input must already be a PDF; a photo handed to this tool is rejected rather than converted. Use images_to_pdf when any input is a picture — it takes PDFs alongside them and splices both in one pass. There is no upload channel over MCP: pass `fileUrls`, an array of URLs in razi.pro's own storage; third-party URLs are rejected. Obtain such a URL by uploading the file over the REST API first (POST /api/v1/tools/execute with the file attached). Over the REST API the files may instead be attached as multipart/form-data. Limited to 20 merges per hour per IP.
    ConnectorNo auth
  • Concatenate two or more PDFs into a single PDF, in the order supplied, and return the merged file. Over MCP the PDF is never inlined: it comes back as a stored URL that stays retrievable for about 24 hours. Page content is copied unchanged — it does not compress (use compress_pdf) or select pages (use split_pdf). Fewer than two files is rejected. Every input must already be a PDF; a photo handed to this tool is rejected rather than converted. Use images_to_pdf when any input is a picture — it takes PDFs alongside them and splices both in one pass. There is no upload channel over MCP: pass `fileUrls`, an array of URLs in razi.pro's own storage; third-party URLs are rejected. Obtain such a URL by uploading the file over the REST API first (POST /api/v1/tools/execute with the file attached). Over the REST API the files may instead be attached as multipart/form-data. Limited to 20 merges per hour per IP.
    ConnectorNo auth
  • Upload media and get back a reusable media_id. Two modes: (1) pass `url` to upload from a publicly accessible URL (preferred for anything over a few MB), or (2) pass `data` (base64-encoded file bytes) plus `mime_type` to upload bytes directly from the model context. Supports images (PNG/JPEG), videos (MP4/MOV), and PDFs (application/pdf). A PDF returns a document-kind media_id — pass it to create_post on a LinkedIn account to publish a native LinkedIn document post (PDF carousel); set platform_configurations.linkedin.document_title to control the title. Use the returned media_id with the `media` param on create_post/update_post. HEIC/HEIF images are not supported — convert to JPEG or PNG first. Direct `data` uploads are capped at 3MB raw because of serverless request-body limits — for larger files, host at a public URL and use `url` mode, or upload via the dashboard.
    ConnectorOAuth
  • Upload a PDF file to Formify. Returns a fileId for use with create_document, create_draft or create_link. Provide exactly ONE of these inputs: (1) `fileReference` — a file reference that the host application supplies when the user attaches a file and the host supports passing files to tools. Pass it exactly as provided; the server downloads the bytes itself. Never construct one by hand and never put a local path in it. (2) `url` — a PDF reachable over HTTPS. Works in every client; use it whenever a URL is available. (3) `uploadId` — for an attached file when the host does not pass file references but you can execute shell commands: call request_file_upload_url first, run its curlCommand, then pass the uploadId here. (4) `file` — the file's bytes as base64 together with fileName, for an attached file when none of the above is possible. Practical ceiling is roughly 30–50 kB of PDF: base64 is a third larger than the file, and a 300 kB PDF would take more output than a model can emit in one message. IMPORTANT: Do NOT use base64 unless you can guarantee you are passing the complete, untruncated file content. AI assistants routinely truncate large strings, which silently corrupts the file and causes upload failures. If the user attached a file but no fileReference reached this tool, retry the call once before choosing another path. Max size: 50 MB. The PDF must not be password-protected or contain digital signatures from other services.
    ConnectorOAuth
  • Export GAAP financial reports for data room / diligence: trial_balance, balance_sheet, income_statement, cash_flow. format=json (default) returns report JSON from /api/ledger/reports; format=csv returns CSV text; format=pdf returns printable HTML as base64 (open in browser → Print → Save as PDF — same as Accounting → Reports). Requires ledger.view. Use ledger_list_books for bookId (primary GAAP book). Point-in-time reports use asOfDate; income_statement and cash_flow use fromDate + toDate (default: YTD through asOfDate/toDate). cash_flow and income_statement line amounts are period activity for that window. Current holdings or amounts owed: ledger_trial_balance.
    ConnectorOAuth
  • Low-level Telegram API (MTProto) invoke for methods not wrapped by other tools. Dangerous methods require allow_dangerous=true. Success: API result dict or normalized error. PII and credential-shaped fields (phone, access_hash) are dropped from a successful result by default; pass include_sensitive=true for the raw payload. A bare message id needs a chat binding: requests with no peer field (messages.GetMessages, messages.DeleteMessages) are refused, because a bare id resolves against an arbitrary dialog. Use channels.GetMessages or messages.GetHistory, which carry the binding. messages.GetHistory cannot address a forum topic (no thread_id/top_msg_id in the schema, and channels.GetHistory does not exist) -- use messages.Search with top_msg_id, or the high-level get_messages with reply_to_id. Full documentation: https://github.com/leshchenko1979/fast-mcp-telegram/blob/main/docs/Tools-Reference.md
    Connector
    Destructive
    No auth

Matching MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables structured extraction of methods and reproducibility heuristics from academic papers, allowing AI agents to obtain metadata, full text, structured methods, code repository discovery, and a no-clone reproducibility verdict from a paper URL.
    8
    36 PyPI
    MIT

Matching MCP Connectors

  • Free PDF tools for AI agents: merge, split, rotate, watermark, page numbers, metadata, flatten.

  • Decision Layer for AI Agents — 58+ tools, Advisor, MCP. Free key: POST /v1/register {}.

  • Low-level Telegram API (MTProto) invoke for methods not wrapped by other tools. Dangerous methods require allow_dangerous=true. Success: API result dict or normalized error. PII and credential-shaped fields (phone, access_hash) are dropped from a successful result by default; pass include_sensitive=true for the raw payload. A bare message id needs a chat binding: requests with no peer field (messages.GetMessages, messages.DeleteMessages) are refused, because a bare id resolves against an arbitrary dialog. Use channels.GetMessages or messages.GetHistory, which carry the binding. messages.GetHistory cannot address a forum topic (no thread_id/top_msg_id in the schema, and channels.GetHistory does not exist) -- use messages.Search with top_msg_id, or the high-level get_messages with reply_to_id. Full documentation: https://github.com/leshchenko1979/fast-mcp-telegram/blob/main/docs/Tools-Reference.md
    Connector
    Destructive
    No auth
  • Convert HTML or Markdown to a pixel-perfect PDF. Returns JSON: { url } — a temporary download URL (valid ~1 hour). Great for generating invoices, reports, receipts, or formatted documents programmatically. Supports full HTML/CSS including tables, images (base64 or URL), and inline styles. For Markdown input, set format='markdown'. 50 sats per conversion. Use convert_file instead for converting existing files between formats (e.g., DOCX→PDF). Pay per request with Bitcoin Lightning — no API key or signup needed. Requires create_payment with toolName='convert_html_to_pdf'.
    ConnectorNo auth
  • READ-ONLY: returns generated source code as text and writes nothing to disk, creates no project and runs no command. Generates an idiomatic @imqueue/rpc service (an IMQService subclass with @expose()d, JSDoc-typed methods) plus a bootstrap that starts it. Provide the methods you want, or omit them for a starter template. Any non-primitive parameter or return type also gets a types.ts with the required @classType()/@property() declarations — without those the generated client types it `any`, which compiles. Use create_service (local install only) if you want files actually written.
    ConnectorNo auth
  • Converts a document to markdown or plain text: pass a public URL or the file itself as base64, and get back the content with headings, tables and lists preserved, at a fraction of the tokens that rendered pages cost. Use it when a harness has no native reader for the format — .docx, .xlsx, .odt and .numbers rarely have one — when a document is only a URL away, or when a long PDF's text matters and its layout does not. Handles PDF (.pdf), Word (.docx), Excel (.xlsx, .xlsm, .xlsb, .xls), OpenDocument (.odt, .ods), Apple Numbers, CSV, HTML, XML, and plain-text formats such as .txt and .md. The format is detected from magic bytes, not trusted from the file name, so a PDF served from a .php URL still converts. Two honest limits: a scanned PDF with no text layer has nothing to extract (this is conversion, not OCR), and legacy binary .doc and .ppt files are not readable — resave them as .docx or .pptx. Images are refused rather than described. Documents up to 10 MB.
    ConnectorNo auth
  • Combine a visual PDF you already have with CII XML into one Factur-X PDF/A-3. Use for an existing visual invoice; use generate_invoice to render invoice data or produce UBL. Match the PDF's parties, lines and totals to xml yourself: check examines only XML and cannot detect disagreement with the visible PDF. The embedded profile comes from xml; neither check nor language changes it. language does not translate the PDF. Invalid input or failed rules return a tool error with no file; rule errors include ids. For a findings report use validate_invoice. Success returns a JSON summary (profile, size_bytes, warning count) and a base64 PDF/A-3 resource. Requires a plan key; success consumes one document, errors consume none. Nothing is stored or sent.
    ConnectorAPI key
  • ⚠ COSTS LLM CREDITS on the NexusTrade account — spins up an Aurora agent via Router V5 classification + ReAct execution loops, billed per token. **Manual approval required**: do NOT call unless the user explicitly asked to launch an Aurora agent. For strategy creation/backtesting/analysis prefer no-LLM tools: structured create_portfolio (pass full IPortfolio JSON), backtest_portfolio, query_backtest_history, query_*, fetch_portfolios. Create a new autonomous Aurora agent using the same body shape as POST /api/agent. When maxIterations or automationMode are omitted, applies the user's saved ChatSettings. Agent models are product-locked (openai/gpt-5.6-luna planner, openai/gpt-6-luna executor, and the platform tool-role defaults) and cannot be overridden. Pass attachment_ids from upload_chat_attachment (READY) to bind files onto the last user message — same as the web FILES tab. Use this for a method-brief PDF plus a short analyze/report request.
    ConnectorOAuth
  • Render user-provided receipt text as a PDF transcription and attach it to a QuickBooks transaction. Intended for receipts that exist only as text, such as an email body or copied order confirmation. The PDF is visibly marked as a transcription, not an original document. Requires confirmed_by_user: true after the user explicitly requests a transcription, and full access on the connection. The receipt fields contain the original text and, when applicable, the actual email sender and subject. Does not transfer original uploaded files.
    ConnectorOAuth
  • Merge two or more PDFs into one, in the order given, and return the result. Sources are read only and never modified. Returns the merged PDF inline as an application/pdf resource. This tool only concatenates whole documents: it does not reorder, rotate or delete pages within them, and it does not fill forms — use pdf_fill for a fillable form and pdf_invoice to build a document from data. An unreadable or rejected input fails with an error, so a partial merge is never returned.
    ConnectorNo auth
  • Merge two or more PDFs into one, in the order given, and return the result. Sources are read only and never modified. Returns the merged PDF inline as an application/pdf resource. This tool only concatenates whole documents: it does not reorder, rotate or delete pages within them, and it does not fill forms — use pdf_fill for a fillable form and pdf_invoice to build a document from data. An unreadable or rejected input fails with an error, so a partial merge is never returned.
    ConnectorNo auth
  • Convert any document to another format without storing a template. Supports 100+ input/output format combinations: Office documents, PDFs, images, web pages, spreadsheets, and more. The source file can be a local path, a URL, or a base64 string. Carbone tags are PRESERVED, not resolved: converting a template keeps every {d.field} intact, so this is also how you proof a template in another format (DOCX template → PDF, or DOCX → ODT while it stays a template). Use render_document instead when you need data injection ({d.field} tags resolved), translations, or batch generation. Common conversions: DOCX → PDF (file: "report.docx", convertTo: "pdf"; add converter: "I" for the fastest DOCX→PDF path), XLSX → PDF (file: "data.xlsx", convertTo: "pdf"), PPTX → PDF (file: "slides.pptx", convertTo: "pdf", converter: "O" for best fidelity), HTML → PDF (file: "page.html", convertTo: "pdf", converter: "C" for full CSS/JS rendering), DOCX → HTML (file: "doc.docx", convertTo: "html"), XLSX → CSV (file: "sheet.xlsx", convertTo: "csv"), PDF → PNG (file: "doc.pdf", convertTo: "png"), PPTX → PNG (first slide as image), MD → PDF (file: "readme.md", convertTo: "pdf").
    ConnectorNo auth
  • Upload multiple PDF files from ChatGPT file attachments. Use this when the user provides multiple file attachments in ChatGPT. Downloads each PDF from its signed URL and stores it. Returns session_id and a list of job_ids. Like upload_pdf, this ONLY works on hosts that resolve chat attachments for you (ChatGPT). On Claude and other MCP clients, call create_upload_page instead. Never invent or guess a download_url or file_id. MANDATORY WORKFLOW before calling this tool: 1. ALWAYS call check_upload_status FIRST — even if you think the files are new. 2. Only include files confirmed absent from check_upload_status. If ALL files are already uploaded, skip batch_upload_pdf entirely and reuse the existing job_ids. 3. Reuse job_ids from already_uploaded — do NOT re-upload those files. Skipping step 1 and calling batch_upload_pdf directly is FORBIDDEN. After batch_upload_pdf completes: if the user requested a comparison, call 'compare_pdfs' with the returned job_ids immediately.
    ConnectorNo auth
  • Convert HTML and CSS to a PDF document using the WeasyPrint rendering engine. Supports every PDF/A archival level, PDF/UA accessibility and the PDF/X print standards. Best for professional documents: invoices, reports, certificates, contracts, and accessible documents. Also produces **fillable PDF forms** — set pdfForms to true. Send a complete HTML document including <html>, <head> with <style>, and <body> tags. Page geometry comes from the document's own CSS @page rule unless paperSize or orientation is set explicitly. Returns a temporary download URL for the generated PDF (valid for 30 minutes). Requires a paid PdfBroker.io plan (Starter or above). EU-first defaults: A4 paper, Portrait orientation when neither the document nor the caller says otherwise.
    ConnectorNo auth
  • Attach every concrete output — notes, drafts, results, files, links — as an artifact so it's part of the task record, not just chat. Real files (pdf, docx, pptx, xlsx, mp3, wav, m4a, images…) are supported: pass `content_base64` for files up to ~6 MB, `fetch_url` to have Tango download and store a hosted file itself, or call create_artifact_upload first for large files and finalize here with `upload_token`. `content` stays the path for inline text and `external_url` for a link you only want recorded. Reference artifact ids in complete_task's evidence_artifact_ids. If a lease is active, Tango attributes the artifact to the lease holder. Otherwise, pass `acting_worker_id` to identify which of your workers is acting; if you don't own that worker the attribution is dropped rather than misrecorded. Attested workers may pass `worker_signature` over the JCS-canonical artifact payload (type 'tango.artifact'); an invalid signature rejects the call and nothing is stored. Delegated workers are signed for automatically. API reference: https://tango.applayer.io/docs/api/tools/add_artifact
    ConnectorOAuth
  • Download a file from a public http(s) URL and store it in the user's Second Brain as a file object — use when the user shares a direct link to a PDF, image, spreadsheet, or other file and asks to save, download, or keep it. The saved file shows up with their uploads and can be read afterwards with read_file. Not for web pages (that is read_web_page with save=true) and not for files behind a sign-in. Files over 50MB are refused.
    ConnectorNo auth
  • Combine 2+ PDFs into a single PDF, in the order given. Inputs are file_ids (from prior tools) or base64 PDFs. Returns a file_id (for chaining) and a ~1h download URL.
    ConnectorNo auth