docbridge
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| DOCBRIDGE_ROOTS | No | Lists the folders docbridge may read and write, separated by the system's path separator (`:` on macOS and Linux, `;` on Windows); by default it is the user's home folder. | |
| DOCBRIDGE_PANDOC | No | The Pandoc executable. Pandoc >= 2.15 is needed only for markdown_to_docx and docx_to_markdown; docbridge finds it on PATH if not set. | |
| DOCBRIDGE_SEARCH_TIMEOUT_S | No | The time bound on one search, default 2 s; a stopped search returns search_timeout, never a partial result. | 2 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| docbridge_contractA | The vocabulary every response uses: statuses, reasons, error codes, which axes a global PASS covers per operation, what each format can carry, the declared text normalization, what 'verbatim' means per format, and whether Pandoc is available. Call this before interpreting a report. [docbridge schema 0.2.4] |
| document_outlineA | The structural map of a PDF/DOCX/MD/TXT in a few hundred characters: unit kind and count (pages, lines or paragraphs), headings from the file's own markup or PDF bookmarks (never inferred), tables, per-page text-layer status, unreadable units. Use it to decide what to read. [docbridge schema 0.2.4] |
| document_searchA | Exact search inside a PDF/DOCX/MD/TXT. Returns each match verbatim with its page/line/paragraph, character offsets, surrounding context and a match_id for document_read_around. Whitespace in the query matches any line break or spacing. Units that cannot be read are listed: no match there is not evidence of absence. [docbridge schema 0.2.4] |
| document_readA | Verbatim text of units start..end (PDF pages, MD/TXT lines, DOCX paragraphs; end omitted = to the end), up to max_chars. Never summarized or filtered. The excerpt carries offsets and SHA-256; if cut, truncated=true, remaining_chars says how much is left and next_cursor continues exactly. [docbridge schema 0.2.4] |
| document_read_aroundA | Verbatim text around a search match: the match's units plus |
| get_reportA | The full validation report behind a report_id (held for the last 200 operations of this server session), or one section of it: axes, structure_checks, differences, warnings, unsupported_checks. [docbridge schema 0.2.4] |
| document_extractA | A mechanical inventory of one PDF/DOCX/MD/TXT - not a substitute for reading it. Counts always; lists only for the kinds asked (numbers, identifiers, dates, urls, emails, headings, tables, links, emphasis, images, pages, blocks, warnings, unsupported, normalized_text), paged by offset/max_items with the rest stated. output_path writes everything to JSON. [docbridge schema 0.2.4] |
| pdf_mergeA | Merge PDFs, in the order given, into one PDF. Bookmarks, links, annotations and form fields are carried. The report checks page count, and that every output page is pixel-identical (rendered), text-identical and rotation-identical to the source page it came from. [docbridge schema 0.2.4] |
| pdf_splitA | Split a PDF into one file per page range ('1-3', '5', '7-' to the end) or one file per page (each_page=true). Files are named _p-.pdf. The report checks each output's page count and that every page is identical to its source page. [docbridge schema 0.2.4] |
| images_to_pdfA | One PDF page per image (JPEG, PNG, TIFF, BMP, GIF, WEBP). JPEGs are embedded byte-for-byte; EXIF orientation and the optional clockwise rotations are applied by page placement, never by re-encoding. The report checks bytes or pixels (colour and alpha separately) and orientation. [docbridge schema 0.2.4] |
| pdf_to_markdownA | Markdown from a PDF's text layer, with page markers. Headings, lists and tables are INFERRED (reported NOT_CHECKED); the text, numbers and identifiers are validated against the PDF. Pages with no text layer are marked, never guessed (no OCR). [docbridge schema 0.2.4] |
| docx_to_markdownA | Markdown (CommonMark + pipe tables + strikethrough) from a DOCX, via sandboxed Pandoc. The report checks text, numbers, identifiers, links, headings, lists, tables and bold/italic, and lists what Markdown cannot carry (headers/footers, comments, equations, images). [docbridge schema 0.2.4] |
| markdown_to_docxA | DOCX from Markdown (CommonMark + pipe tables + strikethrough) via sandboxed Pandoc: headings, paragraphs, bold/italic, lists, tables, links, code blocks. Images are never fetched (alt text is kept). The report checks text, numbers, identifiers and structure. [docbridge schema 0.2.4] |
| txt_to_docxA | DOCX from plain text: one paragraph per line, tabs kept, form feeds as page breaks. Characters DOCX cannot store are refused with their positions, never dropped. The report checks text and numbers. [docbridge schema 0.2.4] |
| document_compareA | Compare two documents (any of PDF, DOCX, MD, TXT) through their canonical forms: missing, added and changed text, numbers and identifiers, structure and page-count differences, and what could not be compared. Read-only. [docbridge schema 0.2.4] |
| validate_conversionA | Given a source document and a converted one (made by anything), run every check that is possible for the format pair and return the validation report: text, numeric and identifier fidelity, structure, and every property NOT_CHECKED or UNSUPPORTED. Writes nothing unless report_path is set. [docbridge schema 0.2.4] |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 16 tools
Most tools target clearly distinct operations: read, read_around, outline, search, extract each occupy a different slice of document inspection, and merge/split/convert tools are format-specific. Minor overlap exists between document_compare and validate_conversion (both diff documents) and between document_read and document_read_around, but descriptions clarify boundaries well.
Names follow readable, domain-specific patterns: document_* for inspection, <format>_to_<format> for conversions, and verb_noun for get_report/validate_conversion. It is not one single verb_noun convention throughout, but each family is internally consistent and predictable.
16 tools is slightly on the heavy side but justified by the broad domain (inspection, extraction, four conversion routes, merge/split, compare, validate). Each tool addresses a distinct format or operation, so few feel redundant.
The surface covers a full read-search-extract-convert-validate lifecycle across PDF/DOCX/MD/TXT plus merge/split and image input. Minor gaps remain: no direct pdf_to_docx or markdown_to_pdf route (only via markdown), but core workflows and validation are well covered.