Skip to main content
Glama

docbridge

An MCP server that converts documents — PDF, DOCX, Markdown, plain text, images — and returns, with every conversion, a report of what was checked, what failed, what was never checked, and what the formats cannot carry. It also reads long documents cheaply without weakening them: outline, exact search and verbatim reads.

NOT_CHECKED is never treated as PASS. A check that could not run, or that met content docbridge cannot read, is reported as NOT_CHECKED or UNSUPPORTED and is never counted as a pass. Silent corruption is treated as worse than a refused conversion: a changed digit in an organisation number (923 456 781 → 923 456 718), an amount (€4,999 → €4,599) or a date (14/07/2024 → 14/07/2025) is a FAIL, reported with both values, their locations and their context.

docbridge went through four rounds of adversarial audit. Every confirmed finding — 21 across the four rounds — was fixed and is covered by regression tests. 282 tests pass.

Version 0.2.4 · schema 0.2.4 · Python ≥ 3.10 · licence AGPL-3.0

Tools

Tool

Does

docbridge_contract

the vocabulary: statuses, reasons, error codes, what each format can carry — call it first

document_outline

the map of a document in a few hundred characters: pages/lines/paragraphs, headings from the file itself, tables, unreadable pages

document_search

exact search with page/line, offsets, context and a match_id; says which pages it could not search

document_read

verbatim text of a page/line/paragraph range, capped; a cut is stated and a cursor continues it exactly

document_read_around

the pages (or paragraphs, lines) around a search match, verbatim

get_report

the full report behind a report_id

pdf_merge

PDFs → one PDF, every page proven identical to its source page

pdf_split

PDF → one file per range, or per page

images_to_pdf

images → PDF, JPEG bytes unchanged, rotation without re-encoding

pdf_to_markdown

PDF text layer → Markdown with page markers (no OCR)

docx_to_markdown

DOCX → Markdown (Pandoc, sandboxed)

markdown_to_docx

Markdown → DOCX (Pandoc, sandboxed)

txt_to_docx

plain text → DOCX, one paragraph per line

document_extract

a mechanical inventory: counts, then only the lists asked for (numbers, dates, URLs, identifiers, tables, ...), paged

document_compare

two documents compared through their canonical forms

validate_conversion

the full report for a conversion made by anything

Related MCP server: PDF MCP Flow

Supported conversions

From

To

Tool

Notes

PDF

Markdown

pdf_to_markdown

the PDF's text layer, with page markers; no OCR

DOCX

Markdown

docx_to_markdown

via Pandoc

Markdown

DOCX

markdown_to_docx

via Pandoc

plain text

DOCX

txt_to_docx

one paragraph per line

images

PDF

images_to_pdf

JPEG bytes embedded unchanged; rotation without re-encoding

several PDFs

one PDF

pdf_merge

every page proven identical to its source page

PDF

several PDFs

pdf_split

by ranges, or one file per page

A conversion made by any other tool can be checked with validate_conversion, and any two documents compared with document_compare, across PDF, DOCX, Markdown and plain text.

Reading a report

By default a tool returns report_summary and a report_id; detail="full" or get_report(report_id) gives the whole report. The summary never states a different verdict.

operation_completed          the file was written            (not a verdict)
status                       PASS | FAIL | NOT_CHECKED | UNSUPPORTED
required_axes                what a PASS covers for this operation
status_basis                 why, in one sentence
differences                  each change: axis, kind, before, after, location, context
unsupported_checks           everything that was not established

NOT_CHECKED and UNSUPPORTED are never passes. Full semantics: DOCUMENT_MCP_CONTRACT.md and VALIDATION_MODEL.md.

Evidence preservation and lossless reading

docbridge may reduce the volume of its own responses, but it never alters, summarizes, paraphrases, semantically filters, or silently truncates source evidence.

  • Reads are verbatim. Every excerpt carries the file's SHA-256, its offsets and the SHA-256 of the returned text.

  • A cut is always stated: truncated, remaining_chars, and a next_cursor that continues at exactly the next character. A cursor or match_id issued for another version of the file is refused with stale_reference.

  • Content docbridge cannot read — a page with no text layer, an equation, an embedded object — becomes an UNKNOWN sentinel that matches nothing, so no comparison can pass across it.

  • Sources are never overwritten; outputs are written atomically.

  • Every path must be absolute and inside the allowed roots (DOCBRIDGE_ROOTS).

  • A PDF's text means its text layer, not what the page shows. There is no OCR.

document_outline(path)                      the map: units, headings, tables, unreadable pages
document_search(path, "tax residence")      exact matches, each with a match_id and its page
document_read_around(path, match_id)        the surrounding pages, verbatim
document_read(path, start=12, end=16)       any range, verbatim; follow next_cursor if truncated

Install

Python ≥ 3.10. Pandoc ≥ 2.15 is needed only for markdown_to_docx and docx_to_markdown; docbridge finds it on PATH, or set DOCBRIDGE_PANDOC to the executable. Every other tool works without it.

git clone https://github.com/Mormolykos/docbridge.git
cd docbridge
python -m venv .venv
.venv/bin/pip install -e ".[dev]"      # Windows: .venv\Scripts\pip install -e ".[dev]"

requirements.lock pins the exact environment the test suite ran in.

MCP server setup

docbridge is a stdio MCP server: python -m docbridge.server, or the docbridge-mcp command the package installs. For clients that use the common mcpServers configuration:

{
  "mcpServers": {
    "docbridge": {
      "command": "/path/to/docbridge/.venv/bin/python",
      "args": ["-m", "docbridge.server"],
      "env": { "DOCBRIDGE_ROOTS": "/path/to/your/documents" }
    }
  }
}

On Windows, command is C:\\path\\to\\docbridge\\.venv\\Scripts\\python.exe. DOCBRIDGE_ROOTS lists the folders docbridge may read and write, separated by the system's path separator (: on macOS and Linux, ; on Windows); by default it is the user's home folder. Optional: DOCBRIDGE_PANDOC (the Pandoc executable) and DOCBRIDGE_SEARCH_TIMEOUT_S (the time bound on one search, default 2 s; a stopped search returns search_timeout, never a partial result).

Tests

.venv/bin/python -m pytest              # 282 tests, including the property fuzz
.venv/bin/python scripts/e2e_fresh.py   # a fresh end-to-end run over the real MCP protocol

DOCBRIDGE_FUZZ_N sets the property-fuzz iterations (default 300). Pandoc-dependent tests skip with a stated reason if Pandoc is absent — a skip is not a pass.

Documentation

  • DOCUMENT_MCP_CONTRACT.md — every tool, argument, status, reason and error code.

  • VALIDATION_MODEL.md — what each check compares, the declared domains, the known limits, and the four audit rounds with their findings and repairs.

Layout

src/docbridge/
  contract.py      statuses, reasons, error codes, required axes, format capabilities
  normalize.py     the declared domain: tokens, numbers, identifiers
  canonical.py     the one shape every format is read into
  reading.py       lossless reading: outline, search, verbatim reads, cursors
  _search_worker.py  one search in its own process, so it can be stopped (time bound)
  extract/         readers: text.py, markdown.py, html.py (raw HTML in Markdown), docx.py, pdf.py
  compare.py       token views, sequence alignment, structure checks
  validate.py      builds a report and the global verdict
  report.py        the report as typed models (the output schema)
  convert/         pandoc.py, pdf_md.py, txt_docx.py, pdf_ops.py, images.py
  paths.py         path policy, atomic writes
  tools.py         the sixteen tools as plain functions
  server.py        the MCP server
tests/             programmatic fixtures (builders.py) and 16 test files
scripts/e2e_fresh.py

Licence

AGPL-3.0 — see LICENSE. docbridge imports PyMuPDF, which is licensed under the AGPL-3.0.

Available Tools

16 tools
docbridge_contractA
Read-onlyIdempotent

The vocabulary every response uses: statuses, reasons, error codes, which axes a global PASS covers per operation, what each format can carry, the declared text normalization, what 'verbatim' means per format, and whether Pandoc is available. Call this before interpreting a report. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
contractYes
disclaimerNo
schema_versionYes
operation_completedNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds value beyond this: it discloses that the tool is a static, versioned contract ('schema 0.2.4') and that it should be invoked as a prerequisite before interpreting reports, which is behavioral context the annotations do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource definition, followed by the actionable 'Call this before interpreting a report' directive and a version tag. The mid-sentence enumeration is lengthy but each item ('what verbatim means per format', 'whether Pandoc is available') earns its place by telling the agent what questions this tool answers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and it does not attempt to. For a zero-parameter introspection tool it fully characterizes what the contract covers and when to consult it, leaving no gap that would cause an incorrect call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

This tool takes zero parameters, so per the rubric the baseline is 4. The description correctly offers no parameter guidance because none is needed, and schema coverage is 100%, so nothing is left ambiguous about invocation inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (the vocabulary/contract every response uses) and enumerates its contents (statuses, reasons, error codes, format capabilities, normalization, verbatim semantics, Pandoc availability). The verb is only implied via 'Call this', but an agent can immediately tell this is a schema/introspection tool, clearly distinct from the operational siblings like document_read or pdf_merge.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit when-to-use directive: 'Call this before interpreting a report.' This gives a clear sequencing condition rather than leaving usage to inference. It does not name an alternative or an exclusion, but no sibling serves the same introspection role, so the practical guidance is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

document_compareA
Read-onlyIdempotent

Compare two documents (any of PDF, DOCX, MD, TXT) through their canonical forms: missing, added and changed text, numbers and identifiers, structure and page-count differences, and what could not be compared. Read-only. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
left_pathYesAbsolute file path.
right_pathYesAbsolute file path.
left_encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
right_encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
max_differencesNoDifferences kept in the full report; the rest are counted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description still adds real context by disclosing that comparison happens on canonical forms and that uncomparable content is reported, plus a schema-version tag; it does not repeat destructive details, but that is unnecessary here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence front-loads the action and the artifact list, followed by a short read-only note and version tag. No wasted prose, though the long enumeration of comparison outputs is slightly heavy for one sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, 100% schema coverage, an output schema, and full annotations, the definition covers what an agent needs to call it correctly. The only shortfall is the absence of routing guidance between this tool and near-neighbors such as document_read or get_report.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the `detail` enum and encoding behavior, so the schema carries the parameter burden. The description adds no syntax or format guidance for left_path/right_path/encodings/max_differences beyond what the schema already states, matching the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (Compare) and resource (two documents of PDF/DOCX/MD/TXT) and spells out the comparison dimensions produced: missing/added/changed text, numbers and identifiers, structure, page-count, and uncomparable content. This is clearly distinct from siblings like document_search, document_read, or document_outline.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes that this is a read-only comparison but never states when to prefer it over siblings such as document_search or document_read, nor any preconditions or exclusions. The `detail` parameter's link to get_report is only explained in the schema, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

document_extractA
Idempotent

A mechanical inventory of one PDF/DOCX/MD/TXT - not a substitute for reading it. Counts always; lists only for the kinds asked (numbers, identifiers, dates, urls, emails, headings, tables, links, emphasis, images, pages, blocks, warnings, unsupported, normalized_text), paged by offset/max_items with the rest stated. output_path writes everything to JSON. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
kindsNo
offsetNo
encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
max_itemsNo
overwriteNo
input_pathYesAbsolute file path.
output_pathNo
text_offsetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
documentNo
disclaimerNo
written_toNo
schema_versionYes
operation_completedYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false with destructiveHint=false and idempotent=true; the description explains why by noting output_path writes everything to JSON, and it discloses the truncation contract ('counts always; lists only for the kinds asked... paged by offset/max_items with the rest stated'). It stops short of covering overwrite behavior or what happens on encoding mismatches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core scope decision in the first clause, then the mechanics. The long parenthetical of kinds is dense but it is enumerating real capability; only the trailing schema-version tag feels like filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the paging/truncation contract is stated. For an 8-parameter tool the remaining gaps are overwrite conflict behavior and the text_offset/offset distinction, which are minor against the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, so the description must compensate; it does partially by explaining the kinds enumeration and the offset/max_items paging interplay, plus output_path's JSON dump. It adds nothing for overwrite, text_offset, or input_path semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource ('mechanical inventory of one PDF/DOCX/MD/TXT') and immediately distinguishes itself from adjacent tools by declaring it is 'not a substitute for reading it', which routes the agent away from document_read. The enumerated outputs make the scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'not a substitute for reading it' framing implies when to prefer document_read, and the paging note ('the rest stated') implies a browse-large-doc scenario. However no alternatives are named explicitly and no preconditions are given, so the guidance is contextual rather than directive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

document_outlineA
Read-onlyIdempotent

The structural map of a PDF/DOCX/MD/TXT in a few hundred characters: unit kind and count (pages, lines or paragraphs), headings from the file's own markup or PDF bookmarks (never inferred), tables, per-page text-layer status, unreadable units. Use it to decide what to read. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
max_itemsNo
input_pathYesAbsolute file path.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
sourceNo
outlineNo
invariantNo
schema_versionYes
operation_completedYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive, closed-world behavior, and the description adds real context beyond that: the output is a few hundred characters, headings are 'never inferred' (only from file markup or PDF bookmarks), and unreadable units and per-page text-layer status are surfaced. It does not mention failure modes for unsupported/corrupt files, but the annotation baseline is already covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense, front-loaded sentence describes the payload, followed by a one-line usage cue and a schema version tag. Every clause carries information, though the enumeration is long enough that a reader must parse carefully.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, yet the description still conveys the shape and size of the output plus its reliability caveats. The only material gap is the undocumented max_items parameter and what happens when the outline exceeds it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: encoding and input_path are documented in the schema, but max_items has no description anywhere and the tool description never explains it (truncation behavior at the default 200 / max 10000). The description adds no parameter-level meaning beyond what the schema already supplies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (structural map of a PDF/DOCX/MD/TXT) and enumerates exactly what it returns: unit kind/count, headings from native markup or PDF bookmarks, tables, per-page text-layer status, and unreadable units. The clause 'Use it to decide what to read' cleanly separates it from document_read and document_extract.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear selection context – run this reconnaissance step to decide what to read next – which implies the read/extract siblings are downstream actions. It stops short of naming an alternative or a when-not condition (e.g., when the file is already small enough to read whole).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

document_readA
Read-onlyIdempotent

Verbatim text of units start..end (PDF pages, MD/TXT lines, DOCX paragraphs; end omitted = to the end), up to max_chars. Never summarized or filtered. The excerpt carries offsets and SHA-256; if cut, truncated=true, remaining_chars says how much is left and next_cursor continues exactly. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
startNo
cursorNo
encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
max_charsNoMost characters returned; a cut is always stated, with a cursor to continue.
input_pathYesAbsolute file path.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
matchNo
sourceNo
excerptNo
invariantNo
truncatedNoTrue if anything in the requested range was not returned.
next_cursorNoPass as `cursor` to document_read to continue exactly.
before_cursorNo
schema_versionYes
remaining_charsNoCharacters of the requested range after the excerpt.
requested_unitsNo
unreadable_unitsNoUnreadable units inside this excerpt.
operation_completedYes
omitted_before_charsNo
range_unreadable_unitsNoEvery unreadable unit of the whole requested range, returned or not.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered; the description then adds genuinely valuable behavior beyond that – truncation semantics (truncated=true, remaining_chars, next_cursor continues exactly), SHA-256 offsets on the excerpt, and the 'docbridge never guesses' encoding stance. It stops short of describing error modes, but this is well above the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph with no filler; the core purpose and range semantics are front-loaded, followed by truncation/continuation detail. The parenthetical format list and trailing schema version tag slightly compress readability but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since an output schema exists, the description needn't enumerate return fields, and it appropriately focuses on the truncation/cursor contract that governs paging. For a six-parameter, multi-format read tool this is nearly complete; only the absence of any sibling routing leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%, so start, end, and cursor carry no schema-level docs; the description compensates by defining the units of start/end (pages/lines/paragraphs), the 'end omitted = to the end' default, and cursor continuation semantics. It adds little about encoding or input_path, but those are already documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Verbatim text of units start..end') and clarifies what a 'unit' means per format (PDF pages, MD/TXT lines, DOCX paragraphs), so the agent knows exactly what is returned. It does not explicitly differentiate itself from the sibling document_read_around or document_search, which is the only gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Never summarized or filtered' implicitly contrasts this with a summarizing/search tool, and the start..end/cursor mechanics imply the range-read use case. However, no sibling is named and there is no explicit when-to-use-this-vs-alternative guidance (e.g., vs document_read_around for context windows).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

document_read_aroundA
Read-onlyIdempotent

Verbatim text around a search match: the match's units plus before/after units (defaults: 1 page, 3 paragraphs, 15 lines). The match is always inside the excerpt; anything left out on either side is stated with a cursor. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
afterNo
beforeNo
encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
match_idYes
max_charsNoMost characters returned; a cut is always stated, with a cursor to continue.
input_pathYesAbsolute file path.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
matchNo
sourceNo
excerptNo
invariantNo
truncatedNoTrue if anything in the requested range was not returned.
next_cursorNoPass as `cursor` to document_read to continue exactly.
before_cursorNo
schema_versionYes
remaining_charsNoCharacters of the requested range after the excerpt.
requested_unitsNo
unreadable_unitsNoUnreadable units inside this excerpt.
operation_completedYes
omitted_before_charsNo
range_unreadable_unitsNoEvery unreadable unit of the whole requested range, returned or not.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive behavior, yet the description adds real value beyond them: the match is guaranteed to be inside the excerpt, any omitted text is reported with a cursor, and cuts from max_chars are always stated. That is substantive behavioral context not available from annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core behavior and cursor semantics; the bracketed schema version tag is minor noise but the rest earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return structure need not be explained, and the description covers the excerpt guarantee and truncation/cursor behavior. Remaining gap is the precise meaning of a 'unit' for the before/after parameters, which matters given a 6-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% and before/after have no schema descriptions, so the description carries the load by giving their unit-relative defaults (1 page, 3 paragraphs, 15 lines) and confirming the match is always included regardless. It does not explain what a 'unit' resolves to per document type, leaving some ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Verbatim text around a search match') and makes the scope precise by requiring a match_id, which implicitly distinguishes it from document_read/document_search. It does not name a sibling explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the required match_id and the 'around a search match' framing — it clearly follows a search — but there is no explicit when-to-use vs. document_read or document_search, and no when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docx_to_markdownA
Idempotent

Markdown (CommonMark + pipe tables + strikethrough) from a DOCX, via sandboxed Pandoc. The report checks text, numbers, identifiers, links, headings, lists, tables and bold/italic, and lists what Markdown cannot carry (headers/footers, comments, equations, images). [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
overwriteNo
input_pathYesAbsolute file path.
output_pathNo
report_pathNo
max_differencesNoDifferences kept in the full report; the rest are counted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint=false, destructiveHint=false, and idempotentHint=true, so the safety profile is partly covered. The description adds useful behavioral context beyond annotations: the conversion runs via sandboxed Pandoc, and the report explicitly checks specific content categories and lists unsupported DOCX features.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and front-loads the conversion direction and output format. The second sentence is dense but informative, though the trailing schema version tag is of limited value to an agent selecting the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema and annotations reduce the burden, and the description explains the report scope and sandboxing. However, for a 6-parameter tool with 50% schema description coverage, it still leaves important invocation details (overwrite, paths, report handling) unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%, and the description does not compensate by explaining parameters. It does not clarify overwrite, output_path, report_path, or how detail/max_differences affect the report beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact conversion direction and output dialect: Markdown (CommonMark + pipe tables + strikethrough) from a DOCX, via sandboxed Pandoc. This clearly distinguishes it from siblings like markdown_to_docx and pdf_to_markdown without needing to name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are provided. The description never says when to choose this tool over pdf_to_markdown, markdown_to_docx, or validate_conversion, so the agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_reportA
Read-onlyIdempotent

The full validation report behind a report_id (held for the last 200 operations of this server session), or one section of it: axes, structure_checks, differences, warnings, unsupported_checks. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
sectionNoall
report_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNo
sectionNo
report_idNo
schema_versionYes
operation_completedYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and openWorld=false, so the safety profile is covered. The description adds a genuinely useful non-annotated trait: the report is only retained for the last 200 operations of the session, warning the agent of expiry. It omits error behavior when a report_id is stale.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the core purpose first, then section options, then the retention caveat. The trailing schema-version tag is minor noise but the whole is tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the description covers purpose, section selection, and retention. The one gap is not stating which operation produces the report_id, leaving provenance implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the burden. It lists the section values (axes, structure_checks, differences, warnings, unsupported_checks), helping map the enum, but does not explain what each section contains and adds nothing about the report_id format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (the validation report behind a report_id) and explicitly frames the choice between the full report and one section. A verb is only implied ("get"), and no sibling tool is named, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent is expected to already hold a report_id from a prior validation. The retention note (last 200 operations) hints at when a report is still retrievable, but no explicit when-to-use, exclusions, or alternatives are offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

images_to_pdfA
Idempotent

One PDF page per image (JPEG, PNG, TIFF, BMP, GIF, WEBP). JPEGs are embedded byte-for-byte; EXIF orientation and the optional clockwise rotations are applied by page placement, never by re-encoding. The report checks bytes or pixels (colour and alpha separately) and orientation. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
inputsYes
overwriteNo
page_sizeNoimage
rotationsNoClockwise degrees, one per image.
output_pathYesAbsolute file path.
report_pathNo
max_differencesNoDifferences kept in the full report; the rest are counted.
apply_exif_orientationNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only weak annotations (readOnly=false, idempotent=true, destructive=false), the description carries most of the load and does so well: JPEGs are embedded byte-for-byte, EXIF orientation and rotations are applied at page-placement time rather than by re-encoding, and the report verifies bytes or pixels (colour and alpha separately) plus orientation. It still omits what happens when output_path already exists with overwrite=false and whether report_path is always written.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense, front-loaded sentences: the page-per-image rule and format list come first, followed by encoding guarantees and report checking. Only the trailing '[docbridge schema 0.2.4]' tag is dispensable noise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained; the description focuses usefully on conversion fidelity and report content. Remaining gaps are the page_size and overwrite semantics and the role of report_path, which an agent would need to inspect the schema for.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 44% across 9 parameters. The description meaningfully clarifies rotations (clockwise, applied by page placement, not re-encoding) and EXIF orientation behavior, but says nothing about page_size, overwrite, report_path, or how detail interacts with get_report, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource: build a PDF with one page per image, and enumerates the accepted input formats (JPEG, PNG, TIFF, BMP, GIF, WEBP). No sibling in the list does image-to-PDF conversion, so an agent can immediately distinguish this tool from pdf_merge, docx_to_markdown, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the operation itself (you have images and want a PDF). There is no explicit when-to-use vs. when-not guidance, no mention of prerequisites, and no pointer to related tools such as validate_conversion or document_compare for verifying the produced PDF.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

markdown_to_docxA
Idempotent

DOCX from Markdown (CommonMark + pipe tables + strikethrough) via sandboxed Pandoc: headings, paragraphs, bold/italic, lists, tables, links, code blocks. Images are never fetched (alt text is kept). The report checks text, numbers, identifiers and structure. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
overwriteNo
input_pathYesAbsolute file path.
output_pathNo
report_pathNo
max_differencesNoDifferences kept in the full report; the rest are counted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false), and the description adds genuinely useful context beyond them: processing is sandboxed Pandoc, images are never fetched and alt text is preserved, and a structural report is produced. These are behavioral facts not derivable from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and packs supported-feature detail efficiently. The clauses are dense and slightly run-on ('sandboxed Pandoc: ... links, code blocks'), but every sentence carries information relevant to invoking the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the description covers conversion semantics and image handling. It nonetheless omits how output_path/report_path/overwrite behave (destination and file-clobbering behavior) for a tool whose default overwrite is false, leaving a gap for an agent deciding where output lands.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 57%, with detail, encoding, and max_differences well documented in the schema itself. The description adds only an oblique reference ('The report checks...') and says nothing about output_path, report_path, or overwrite, so it does not compensate for the uncovered parameters. Baseline 3 fits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (produce DOCX from Markdown) and enumerates supported constructs (CommonMark, pipe tables, strikethrough, headings, lists, tables, code blocks). It clearly identifies the conversion direction, though it never names the sibling markdown-family tools to differentiate itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and description (you have Markdown, you want DOCX), and the sandboxed-Pandoc note hints at the processing context. However, there is no explicit when-to-use/when-not guidance and no mention of the adjacent converters (txt_to_docx, docx_to_markdown, validate_conversion).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pdf_mergeA
Idempotent

Merge PDFs, in the order given, into one PDF. Bookmarks, links, annotations and form fields are carried. The report checks page count, and that every output page is pixel-identical (rendered), text-identical and rotation-identical to the source page it came from. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
inputsYesPDF paths, in output order.
overwriteNo
raster_dpiNoRender resolution for the page-identity check.
output_pathYesAbsolute file path.
report_pathNo
max_differencesNoDifferences kept in the full report; the rest are counted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare this is a non-read-only but idempotent, non-destructive write, and the description adds real value beyond them: exactly what content is carried (bookmarks, links, annotations, form fields) and the built-in verification (page count, pixel/text/rotation identity). It stops short of stating failure behavior or what happens to the output path on error.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action, followed by carried content and verification. The trailing '[docbridge schema 0.2.4]' tag is minor noise but the rest earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't explain return values, and it covers action, carried content, and verification. Only error/failure handling and overwrite behavior are unstated, which is a minor gap for a 7-param write tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, so the schema documents most parameters, but the description usefully explains why raster_dpi, max_differences, and the report exist by describing the page-identity check. This adds meaning beyond the field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Merge PDFs ... into one PDF') and pins down the ordering semantics ('in the order given'), which is the key ambiguity for a merge operation. It is easily distinguished from siblings like pdf_split and images_to_pdf.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is strongly implied by the name and the description, but there is no explicit when/when-not guidance and no pointer to alternatives (e.g., pdf_split for the inverse, images_to_pdf for non-PDF sources). Adequate but leaves selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pdf_splitA
Idempotent

Split a PDF into one file per page range ('1-3', '5', '7-' to the end) or one file per page (each_page=true). Files are named _p-.pdf. The report checks each output's page count and that every page is identical to its source page. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
rangesNo1-based ranges; exactly one of ranges/each_page.
each_pageNo
overwriteNo
input_pathYesAbsolute file path.
output_dirNo
raster_dpiNoRender resolution for the page-identity check.
report_pathNo
max_differencesNoDifferences kept in the full report; the rest are counted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes beyond the annotations (which only cover safety: non-destructive, idempotent) by disclosing the output naming convention and the per-page identity/page-count verification the report performs. It does not explain overwrite-collision behavior, which matters for a write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the core behavior and mode switch. The trailing '[docbridge schema 0.2.4]' tag is noise that earns no place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the description covers modes, naming, and verification. For a 9-parameter mutation tool, however, it leaves overwrite semantics and several optional parameters unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 56%; the description usefully adds range format examples ('1-3', '5', '7-') that the schema lacks. But overwrite, output_dir, report_path, detail, and max_differences get no mention in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Split a PDF') with both operating modes named inline and concrete range syntax examples. An agent can immediately distinguish it from the sibling pdf_merge.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The two modes (ranges vs each_page) are explained, which implies how to invoke it, but there is no explicit when-to-use-this-vs-alternatives guidance and no exclusions. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pdf_to_markdownA
Idempotent

Markdown from a PDF's text layer, with page markers. Headings, lists and tables are INFERRED (reported NOT_CHECKED); the text, numbers and identifiers are validated against the PDF. Pages with no text layer are marked, never guessed (no OCR). [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
overwriteNo
input_pathYesAbsolute file path.
output_pathNo
report_pathNo
page_markersNo
detect_tablesNo
max_differencesNoDifferences kept in the full report; the rest are counted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (destructive=false, idempotent=true), and the description adds substantial context beyond them: headings/lists/tables are INFERRED and flagged NOT_CHECKED, text/numbers are validated, and no-OCR pages are marked rather than guessed. It stops short of describing write/overwrite behavior or output file handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core output and its reliability model come first, the no-OCR guarantee last. Every clause carries meaning; only the schema version tag in brackets is arguably filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the fidelity model is well covered. However, for a tool that writes Markdown to a path, the description never explains the output_path/overwrite/report_path contract, leaving real gaps given the low parameter coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 38% across 8 parameters. The description alludes to page markers and the report, but overwrite, output_path, report_path, and detect_tables are undocumented in both description and schema, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (Markdown from a PDF) and immediately scopes it to the text layer with page markers, which cleanly separates it from docx_to_markdown and the OCR-based alternatives. An agent knows exactly what artifact this produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the appropriate context (PDFs that have a text layer) and states the failure mode for pages without one, but never names an alternative tool or an explicit when-to-use/when-not-to-use rule against siblings like docx_to_markdown or validate_conversion. Usage is inferable, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

txt_to_docxA
Idempotent

DOCX from plain text: one paragraph per line, tabs kept, form feeds as page breaks. Characters DOCX cannot store are refused with their positions, never dropped. The report checks text and numbers. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
overwriteNo
input_pathYesAbsolute file path.
output_pathNo
report_pathNo
max_differencesNoDifferences kept in the full report; the rest are counted.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (idempotent, non-destructive); the description adds real behavioral substance: one paragraph per line, tabs preserved, form feeds become page breaks, and unencodable characters are refused with positions rather than dropped. That failure-mode disclosure is genuinely useful and not recoverable from the schema. It omits overwrite semantics (default false) and report behavior timing, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight clauses, conversion rules front-loaded, no filler sentences. The final sentence ('The report checks text and numbers') is somewhat cryptic and the '[docbridge schema 0.2.4]' tag is low-value, but overall density is high.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and this is a non-destructive, idempotent writer. However, for a file-writing tool the description never states what happens when output_path exists without overwrite, nor how encoding and the report parameters interact with the conversion summary it advertises. Adequate but with clear gaps for a 7-parameter mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 57%, and the schema already documents detail, encoding, and max_differences well. The description references the report ('The report checks text and numbers') without tying it to report_path, detail, or max_differences, and never mentions overwrite or output_path defaults. Marginal added meaning over the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource direction ('DOCX from plain text') and enumerates the exact text-to-DOCX mapping rules, which distinguishes it from markdown_to_docx and other conversion siblings. It never uses the tool name's phrasing, so an agent must infer that txt_to_docx is the txt→docx converter, but the transformation semantics make the intent unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the input format ('plain text', 'a .md/.txt input'). There is no explicit when-to-use versus markdown_to_docx or when-not guidance, and the version tag '[docbridge schema 0.2.4]' occupies space that routing guidance could fill. An agent can guess the right tool from the format description but gets no help choosing between the sibling converters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_conversionA
Idempotent

Given a source document and a converted one (made by anything), run every check that is possible for the format pair and return the validation report: text, numeric and identifier fidelity, structure, and every property NOT_CHECKED or UNSUPPORTED. Writes nothing unless report_path is set. [docbridge schema 0.2.4]

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNosummary: report_summary + report_id (get_report has the rest). full: the whole report inline.summary
overwriteNo
report_pathNo
source_pathYesAbsolute file path.
target_pathYesAbsolute file path.
max_differencesNoDifferences kept in the full report; the rest are counted.
source_encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.
target_encodingNoText encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses.

Output Schema

ParametersJSON Schema
NameRequiredDescription
toolYes
errorNo
reportNoThe full report. Over MCP only with detail='full'.
outputsNo
report_idNoPass to get_report for the full report (held for this server session).
disclaimerNo
report_pathNo
report_summaryNo
schema_versionYes
conversion_notesNo
operation_completedYesThe operation ran and wrote its outputs. Says NOTHING about fidelity: read the report status.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-destructive, idempotent, but NOT read-only. The description resolves that tension well by stating 'Writes nothing unless report_path is set', which explains why readOnlyHint is false for a mostly-read operation. It also discloses behavior beyond annotations by naming the report contents (fidelity categories plus every NOT_CHECKED or UNSUPPORTED property).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the trigger condition and the returned checks in two dense sentences with no filler. The trailing '[docbridge schema 0.2.4]' tag is metadata rather than agent-facing guidance, a minor deduction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a validation tool whose return shape is covered by an output schema, the description gives enough: it defines inputs, the optional side effect, and the report's scope. Safety and idempotency come from annotations. The only real gap is the absence of routing against document_compare.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already documents most parameters. The description adds meaning for report_path (the conditional write) and alludes to the report scope, but says nothing about detail, overwrite, max_differences, or the encoding parameters, so it only partially compensates for the uncovered fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (validate) and resource (a conversion between a source document and a converted one), and enumerates the check categories it returns: text, numeric and identifier fidelity, structure. It does not name or differentiate itself from the sibling document_compare, which is the most likely source of confusion for an agent, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Given a source document and a converted one (made by anything)' sets the precondition and clarifies it works with output from any converter, which is helpful. However, it never says when to prefer this over document_compare, or when a validation is unnecessary, leaving the when-to-use decision largely implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 16 tool updatesv0.2.4
    • First observeddocbridge_contract
    • First observeddocument_compare
    • First observeddocument_extract
    • First observeddocument_outline
    • First observeddocument_read
    • First observeddocument_read_around
    • First observeddocument_search
    • First observeddocx_to_markdown
    • First observedget_report
    • First observedimages_to_pdf
    • First observedmarkdown_to_docx
    • First observedpdf_merge
    • First observedpdf_split
    • First observedpdf_to_markdown
    • First observedtxt_to_docx
    • First observedvalidate_conversion

TDQS

A3.8/5.0

Scored across 16 tools

Disambiguation4/5

Most tools target clearly distinct operations: read, read_around, outline, search, extract each occupy a different slice of document inspection, and merge/split/convert tools are format-specific. Minor overlap exists between document_compare and validate_conversion (both diff documents) and between document_read and document_read_around, but descriptions clarify boundaries well.

Naming Consistency4/5

Names follow readable, domain-specific patterns: document_* for inspection, <format>_to_<format> for conversions, and verb_noun for get_report/validate_conversion. It is not one single verb_noun convention throughout, but each family is internally consistent and predictable.

Tool Count4/5

16 tools is slightly on the heavy side but justified by the broad domain (inspection, extraction, four conversion routes, merge/split, compare, validate). Each tool addresses a distinct format or operation, so few feel redundant.

Completeness4/5

The surface covers a full read-search-extract-convert-validate lifecycle across PDF/DOCX/MD/TXT plus merge/split and image input. Minor gaps remain: no direct pdf_to_docx or markdown_to_pdf route (only via markdown), but core workflows and validation are well covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers