flatmark
Server Details
Document to Markdown MCP server: PDF, Word, PowerPoint, Excel and HTML, with OCR for large files.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- flatmark-dev/flatmark
- GitHub Stars
- 0
TDQS
Scored across 3 tools
The two conversion tools are clearly differentiated by scope: convert_document for small, fast, no-OCR synchronous conversions and submit_conversion_job for large, OCR/table-capable asynchronous jobs. get_conversion_job is a distinct polling endpoint. An agent can easily choose based on file size, features, or API key availability.
All tools use snake_case with a consistent verb_noun structure: convert_document, get_conversion_job, submit_conversion_job. The pattern is predictable and unambiguous, making the tool set easy to scan and remember.
Three tools perfectly match the minimal service surface: one synchronous conversion, one asynchronous job submission, and one job status check. No redundancy or bloat, and each tool earns its place.
Core conversion lifecycle is covered (sync convert, async submit, status poll). Missing job cancellation or listing, but these are minor gaps that don't block typical agent workflows and results auto-delete after 7 days.
Available Tools
3 toolsconvert_documentConvert a document to Markdown in one callARead-onlyInspect
Convert a document up to 8 MB to Markdown and return it directly.
Pass the file as base64 in content_base64 and its name in filename.
The extension decides the type: .pdf, .docx, .pptx, .xlsx, .html, .htm
or .txt. Returns markdown and meta (filename, length in
characters). PDFs come back as plain text without headings and without
OCR. For scans, tables and page numbers, use submit_conversion_job.
Works without an API key at the anonymous rate limit.
Credits: 1 per call.
| Name | Required | Description | Default |
|---|---|---|---|
| filename | Yes | The file name with its extension (.pdf, .docx, .pptx, .xlsx, .html, .htm or .txt). The extension decides how the file is read. | |
| content_base64 | Yes | The file content, base64-encoded. Up to 8 MB before encoding. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=true and openWorldHint=false annotations, it discloses the 8 MB size limit, that PDFs return as plain text without headings or OCR, the credit cost, and that authentication is optional. These are exactly the behavioral traits an agent needs that annotations do not cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short paragraphs, front-loaded with the action and input contract, then caveats, then alternatives, then cost. Every sentence carries actionable information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, yet the description notes the `markdown` and `meta` fields briefly without bloat. All constraints, alternatives, auth behavior and cost are covered for a two-parameter converter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented, and the description largely restates that (base64 content, filename, extension drives type). The 8 MB-before-encoding detail duplicates the schema, so no meaningful parameter value is added beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Convert a document ... to Markdown') plus scope ('up to 8 MB') and delivery mode ('return it directly'). It distinguishes itself from the sibling submit_conversion_job by naming the cases the sibling handles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use this tool for the listed extensions, but switch to submit_conversion_job for scans, tables and page numbers. It also notes the operation works without an API key at the anonymous rate limit, which is useful invocation context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_conversion_jobCheck a queued conversionARead-onlyInspect
Check a queued conversion job you submitted.
Returns job_id and status (queued, running, succeeded or
failed). A succeeded job includes the markdown when it is under
100 KB, otherwise a result_url to download it (also free). A failed job
includes error, one sentence with the reason. Results are deleted
7 days after submission. An expired result is reported in error.
Requires an API key.
Credits: free. This tool is not metered.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The `job_id` returned by `submit_conversion_job`. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only covering readOnly/openWorld, the description carries real behavioral weight: it enumerates all four status values, explains the 100 KB inline-vs-result_url branch, the failure error contract, the 7-day retention and expiry reporting, and the API-key requirement. This is exactly the context an agent needs to interpret responses.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then well-organized status/return/retention details. Slightly redundant at the end, where 'Credits: free.' and 'This tool is not metered.' say nearly the same thing, costing some tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single required parameter, existing output schema, and read-only annotations, nothing an agent needs is missing: status semantics, size-based return branch, failure contract, expiry, auth, and cost are all covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single job_id parameter and its description points back to submit_conversion_job, so the schema already does the work. The prose adds no format or syntax detail beyond that, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Check a queued conversion job you submitted') and immediately distinguishes itself from submit_conversion_job by scoping to an already-submitted job. An agent can pick this tool over its siblings without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'you submitted' and the job_id's provenance ('returned by submit_conversion_job') imply the post-submission polling context clearly. It does not, however, state explicit alternatives or when-not-to-use conditions, e.g. that it should be polled repeatedly until terminal status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_conversion_jobQueue a document for conversion with OCR and tablesAInspect
Queue a document up to 25 MB and 200 pages for conversion with OCR and tables.
Pass the file as base64 in content_base64 and its name in filename
(.pdf, .docx, .pptx, .xlsx, .html, .htm or .txt). Returns job_id,
status and result_url. Call get_conversion_job with the job_id
until status is succeeded or failed. A conversion may take up to
about 2 minutes.
Optional webhook_url: a public http or https URL that receives a POST
with the job's final state. The request carries
X-Appkit-Signature: sha256=<hex>, an HMAC-SHA256 of the raw body keyed
with the returned webhook_secret (shown only once). 3 delivery attempts.
Requires an API key. The credits are refunded if the job fails. The uploaded file is deleted when the job finishes. Results are kept for 7 days after submission.
Credits: 10 per call.
| Name | Required | Description | Default |
|---|---|---|---|
| filename | Yes | The file name with its extension (.pdf, .docx, .pptx, .xlsx, .html, .htm or .txt). The extension decides how the file is read. | |
| webhook_url | No | Optional. A public http or https URL that receives the job's final state. | |
| content_base64 | Yes | The file content, base64-encoded. Up to 25 MB before encoding. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which only declare non-read-only, non-idempotent, open-world, non-destructive), the description discloses credit cost and refund-on-failure, API key requirement, that credits are consumed per call, file deletion on completion, 7-day result retention, webhook signing scheme with HMAC-SHA256 and one-time secret, and 3 delivery attempts. This is exactly the behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and limits, then structured into parameters, webhook behavior, and operational facts. It is a bit long and duplicates the extension list and filename semantics that already live in the schema, but almost every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a job-submitting tool with an output schema, the description covers everything an agent needs: input format, size/page caps, polling contract, webhook verification, cost, failure refund, and retention. Even though returns are documented by the output schema, summarizing `job_id`/`status`/`result_url` aids the polling instruction.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning: it restates the accepted extensions in context, explains the webhook_url delivery contract (POST, signature header, retries) and ties `job_id` to the polling loop. Slight redundancy with the schema's own extension list, but the added semantics are genuine.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb (Queue) plus the resource (a document) and scope (OCR and tables, size/page limits). It is clear what the tool does, but it never distinguishes itself from the sibling `convert_document`, so an agent cannot tell the two apart from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the workflow explicitly: call `get_conversion_job` with the returned `job_id` until `status` is `succeeded` or `failed`, and gives a ~2 minute expectation. What is missing is guidance on when to choose this tool over the sibling `convert_document`.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
convert_document - First observed
get_conversion_job - First observed
submit_conversion_job
Related MCP Connectors
Document-to-Markdown MCP server — convert PDF, Office and HTML into LLM-ready Markdown.
Document conversion MCP server: PDF to Markdown, image OCR, spreadsheet parsing.
PDF, Word, PowerPoint, Excel, HTML, EPUB to Markdown: OCR, page ranges, tables, RAG chunking
Hosted MCP server: convert PDFs to clean, LLM-ready Markdown with tables, formulas and OCR.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceConverts PDF, Word, and Excel documents to Markdown with image extraction and header/footer removal via MCP or REST API.50MIT
- FlicenseNot gradedqualityCmaintenanceConverts .docx files to Markdown, with optional image extraction and HTML table conversion, accessible via MCP server or Python API.-
- FlicenseNot gradedqualityDmaintenanceConverts documents (PDF, DOCX, images, etc.) to Markdown using Microsoft's Markitdown library, with no local setup required. Integrates with AI agents via MCP for seamless document conversion.1-
- AlicenseNot gradedqualityCmaintenanceMCP server for converting various file formats (PDF, DOCX, images, audio, etc.) to Markdown using Microsoft MarkItDown, with support for large files and Cyrillic text.1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.