flatmark
Server Details
Document to Markdown MCP server: PDF, Word, PowerPoint, Excel and HTML, with OCR for large files.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- flatmark-dev/flatmark
- GitHub Stars
- 0
TDQS
Scored across 3 tools
The three tools have distinct roles: convert_document for immediate small conversions, submit_conversion_job for queued OCR/large-file conversions, and get_conversion_job for polling status. Descriptions explicitly guide when to use which conversion path, eliminating ambiguity.
All tool names follow a consistent verb_noun pattern in snake_case (convert_document, get_conversion_job, submit_conversion_job). The shared 'conversion_job' suffix for the asynchronous pair reinforces the pattern.
Three tools are well-scoped for a focused document conversion service: one synchronous converter, one asynchronous submitter, and one status checker. Each tool earns its place without redundancy.
The surface covers the full conversion lifecycle: direct small-file conversion, queued large-file/OCR conversion, and job result retrieval. Optional webhooks and result URLs are supported, with no obvious missing operations for the stated domain.
Available Tools
3 toolsconvert_documentConvert a document to Markdown in one callARead-onlyInspect
Convert a document up to 8 MB to Markdown and return it directly.
Pass the file as base64 in content_base64, or as a download link in
file, and its name in filename. The extension decides the type: .pdf,
.docx, .pptx, .xlsx, .html, .htm or .txt. Returns markdown and meta
(filename, length in characters). PDFs come back as plain text
without headings and without OCR. For scans, tables and page numbers,
use submit_conversion_job. Works without an API key at the anonymous
rate limit.
Credits: 1 per call.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | The file as a download link, instead of `content_base64`. In ChatGPT, pass the user's uploaded file here. Other clients can set `download_url` to a public https URL. | |
| filename | Yes | The file name with its extension (.pdf, .docx, .pptx, .xlsx, .html, .htm or .txt). The extension decides how the file is read. | |
| content_base64 | No | The file content, base64-encoded. Up to 8 MB before encoding. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, but the description adds substantial context beyond them: the 8 MB size cap, supported extensions, PDF plain-text limitations without headings or OCR, returned fields, anonymous auth behavior, and credit cost. No annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and size limit, then structured details on inputs, outputs, limitations, and alternatives. Every sentence contributes unique operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description need not define return structure but still names the markdown and meta outputs. It covers size limits, supported formats, PDF caveats, alternative tool routing, authentication, and cost, leaving no meaningful gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters. The description adds useful selection semantics by explaining the two input modes: base64 via content_base64 or a download link via file, plus that filename's extension determines the parser.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: convert a document up to 8 MB to Markdown and return it directly. It clearly distinguishes itself from submit_conversion_job by naming that sibling for scans, tables, and page numbers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly covers when to use this tool versus the alternative, including the exact conditions for choosing submit_conversion_job. It also states the anonymous rate-limit availability and the 1-credit cost.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_conversion_jobCheck a queued conversionARead-onlyInspect
Check a queued conversion job you submitted.
Returns job_id and status (queued, running, succeeded or
failed). A succeeded job includes the markdown when it is under
100 KB, otherwise a result_url to download it (also free). A failed job
includes error, one sentence with the reason. Results are deleted
7 days after submission. An expired result is reported in error.
Requires an API key.
Credits: free. This tool is not metered.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The `job_id` returned by `submit_conversion_job`. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only covering readOnly/openWorld, the description carries real behavioral weight: it enumerates all four status values, explains the 100 KB inline-vs-result_url branch, the failure error contract, the 7-day retention and expiry reporting, and the API-key requirement. This is exactly the context an agent needs to interpret responses.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then well-organized status/return/retention details. Slightly redundant at the end, where 'Credits: free.' and 'This tool is not metered.' say nearly the same thing, costing some tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single required parameter, existing output schema, and read-only annotations, nothing an agent needs is missing: status semantics, size-based return branch, failure contract, expiry, auth, and cost are all covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single job_id parameter and its description points back to submit_conversion_job, so the schema already does the work. The prose adds no format or syntax detail beyond that, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Check a queued conversion job you submitted') and immediately distinguishes itself from submit_conversion_job by scoping to an already-submitted job. An agent can pick this tool over its siblings without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'you submitted' and the job_id's provenance ('returned by submit_conversion_job') imply the post-submission polling context clearly. It does not, however, state explicit alternatives or when-not-to-use conditions, e.g. that it should be polled repeatedly until terminal status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_conversion_jobQueue a document for conversion with OCR and tablesAInspect
Queue a document up to 25 MB and 200 pages for conversion with OCR and tables.
Pass the file as base64 in content_base64, or as a download link in
file, and its name in filename (.pdf, .docx, .pptx, .xlsx, .html,
.htm or .txt). Returns job_id, status and result_url. Call
get_conversion_job with the job_id until status is succeeded or
failed. A conversion may take up to about 2 minutes.
Optional webhook_url: a public http or https URL that receives a POST
with the job's final state. The request carries
X-Appkit-Signature: sha256=<hex>, an HMAC-SHA256 of the raw body keyed
with the returned webhook_secret (shown only once). 3 delivery attempts.
Requires an API key. The credits are refunded if the job fails. The uploaded file is deleted when the job finishes. Results are kept for 7 days after submission.
Credits: 10 per call.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | The file as a download link, instead of `content_base64`. In ChatGPT, pass the user's uploaded file here. Other clients can set `download_url` to a public https URL. | |
| filename | Yes | The file name with its extension (.pdf, .docx, .pptx, .xlsx, .html, .htm or .txt). The extension decides how the file is read. | |
| webhook_url | No | Optional. A public http or https URL that receives the job's final state. | |
| content_base64 | No | The file content, base64-encoded. Up to 25 MB before encoding. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag non-idempotent, open-world, non-destructive, non-read-only; the description goes far beyond that with limits, the ~2 minute duration, API-key requirement, credit cost and refund-on-failure policy, file deletion on completion, 7-day result retention, and webhook signature/HMAC/secret-shown-once/3-retry semantics. This is exactly the kind of behavior an agent must know before issuing a paid, non-idempotent job.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five short paragraphs are front-loaded with the key constraints and each covers a distinct concern (submission, polling, webhook, lifecycle, cost). Minor redundancy exists where the 25 MB limit and allowed extensions restate the schema, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite an output schema existing (so return fields need no explanation), the description still names the returned `job_id`/`status`/`result_url` to anchor the polling loop. Combined with limits, auth, cost, retention and webhook contract, nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description adds genuine value by explaining the two mutually exclusive ingestion paths (`content_base64` vs. `file`) and the webhook's signature and retry contract. The 25 MB limit and extension list partially repeat schema text, keeping it just below a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource (queue a document for conversion) with concrete scope limits (25 MB, 200 pages) and the processing features (OCR, tables). It also routes the agent forward to `get_conversion_job`, which cleanly separates this async queueing tool from its polling sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear operational workflow (submit, then poll `get_conversion_job` until `succeeded`/`failed`, or supply a `webhook_url`) and names the follow-up tool explicitly. It never states when to choose this over the sibling `convert_document`, so the alternative-selection guidance is incomplete, but the usage context itself is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
convert_document - First observed
get_conversion_job - First observed
submit_conversion_job
Related MCP Connectors
Document-to-Markdown MCP server — convert PDF, Office and HTML into LLM-ready Markdown.
Document conversion MCP server: PDF to Markdown, image OCR, spreadsheet parsing.
PDF, Word, PowerPoint, Excel, HTML, EPUB to Markdown: OCR, page ranges, tables, RAG chunking
Hosted MCP server: convert PDFs to clean, LLM-ready Markdown with tables, formulas and OCR.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceConverts PDF, Word, and Excel documents to Markdown with image extraction and header/footer removal via MCP or REST API.51MIT
- FlicenseNot gradedqualityCmaintenanceConverts .docx files to Markdown, with optional image extraction and HTML table conversion, accessible via MCP server or Python API.-
- FlicenseNot gradedqualityDmaintenanceConverts documents (PDF, DOCX, images, etc.) to Markdown using Microsoft's Markitdown library, with no local setup required. Integrates with AI agents via MCP for seamless document conversion.1-
- AlicenseNot gradedqualityCmaintenanceMCP server for converting various file formats (PDF, DOCX, images, audio, etc.) to Markdown using Microsoft MarkItDown, with support for large files and Cyrillic text.1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.