visimark
Provides document integrity checks for Markdown files, evaluating embedded formulas and anchors to verify computed numbers, detect stale values, date errors, undefined names, cycles, and other inconsistencies.
VisiMark
A document integrity layer for Markdown: it checks that every number in a document still matches the formula that produced it, so an agent's correct formula can't ship with a wrong total.
VisiMark checks the numbers in Markdown documents, especially ones an agent wrote. Every computed value carries its formula, a machine proves the two still agree, and CI fails when they don't. The document stays plain Markdown.
This matters most when the Markdown is written or edited by an AI agent.
Agents are reliable at writing formulas and unreliable at the arithmetic those
formulas describe: Net = Price * Qty is something an agent gets right,
50.00 is something it guesses. Write the formula instead of the number and
the number stops being a claim and becomes a derivation — reviewable in a
diff, re-runnable, enforceable in CI. The agent writes, VisiMark verifies, Git
records the result.
A VisiMark document is an ordinary Markdown file. It renders correctly on GitHub, in VS Code preview, and through pandoc to both HTML and Word, today, with no plugin — verified, not assumed. What VisiMark adds is that every computed number in it carries the formula that produced it, that a machine can prove the two still agree, and that a change to either shows up as a small, readable diff.
A wrong invoice that renders clean
docs/example-invoice-drift.md is a real B2B
invoice after someone raised the on-call hours from 12 to 20 and updated
nothing that depends on it. It renders as a clean, plausible invoice on
GitHub, in a Markdown preview, anywhere — which is the entire argument for
this project. Every number in it carries the formula that produced it; the
syntax follows further down. Here is visimark check reading the same file a
reviewer just skimmed and approved:
$ visimark check docs/example-invoice-drift.md
docs/example-invoice-drift.md
STALE lines.Net · On-call support 3120.00 ≠ 5200.00 Qty * Rate
STALE lines.VAT · On-call support 717.60 ≠ 1196.00 Net * vat
STALE lines.Gross · On-call support 3837.60 ≠ 6396.00 Net + VAT
STALE lines.Gross · Discovery workshop 4428.50 ≠ 4428.00 Net + VAT
STALE lines.net_total 23300.00 ≠ 25380.00 SUM(Net)
STALE lines.vat_total 5359.00 ≠ 5837.40 SUM(VAT)
STALE lines.gross_total 28659.00 ≠ 31217.40 SUM(Gross)
STALE lines.gross_total 28659.00 ≠ 31217.40 SUM(Gross)
STALE schedule.Amount · Signature 8597.70 ≠ 9365.22 Share * lines.gross_total
STALE schedule.Amount · Delivery of backend 11463.60 ≠ 12486.96 Share * lines.gross_total
STALE schedule.Amount · Acceptance 8597.70 ≠ 9365.22 Share * lines.gross_total
STALE schedule.covered 28659.00 ≠ 31217.40 SUM(Amount)
STALE terms.early_pay_total 28085.82 ≠ 30593.05
STALE terms.early_pay_saved 573.18 ≠ 624.35
STALE 8 prose anchors bound to the values above
DATE schedule.Due · Delivery of backend "15.10.2026"
Dates must be ISO 8601 calendar dates: YYYY-MM-DD.
Unambiguous — `visimark fmt --fix-dates` rewrites it to 2026-10-15.
DATE schedule.Due · Acceptance "11/12/2026"
Dates must be ISO 8601 calendar dates: YYYY-MM-DD.
Ambiguous: 2026-12-11 or 2026-11-12, 29 days apart. Fix by hand.
NOTE schedule.Days · 2 rows not verified (upstream DATE errors)
UNDEF terms.eur_total unknown name `fx_rate`
did you mean `fx_eur`?
VECTOR recon.variance `schedule.Amount` is a column, not a value.
Wrap it in an aggregate: SUM(schedule.Amount)
CYCLE late_fees.base → late_fees.fee → late_fees.total → late_fees.base
27 problems (22 stale, 5 errors)
$ echo $?
1Twenty-seven problems: a payment date ambiguous by twenty-nine days, a cell someone nudged by hand to make a column look right, a circular reference — all invisible on the rendered page, all caught before a human had to notice.
Related MCP server: Claude Writer's Aid MCP
Why this matters now
More and more of the documents that carry numbers are written as text and kept in Git: reports and research notes, estimates and budgets, quotes and invoices, project plans, financial summaries. Text is excellent for collaboration, review and version control. But ordinary Markdown has no way to say this number came from these inputs, and it must still agree with them — so the moment an input changes, every figure downstream of it is a guess until a human re-checks it by hand.
VisiMark adds that missing layer. The working stack for collaborating with an agent is a text editor over lightly formatted artifacts; Markdown already covers the prose, and this covers the calculation. The arithmetic is not the value — the audit trail is.
What it looks like
| Item | Price | Qty | Net |
|-------|------:|----:|------:|
| pen | 5.00 | 10 | 50.00 |
| paper | 0.10 | 100 | 10.00 |
```vmark #order
Net = Price * Qty
total = SUM(Net)
```
Order total: **60.00**<!--vmark=order.total-->Four ideas, and that is the whole format:
A
vmarkblock declares formulas for the table above it, and names the sheet.Columns are uniform —
Net = Price * Qtyis one rule for every row, not a formula per cell. Columns with no rule are inputs, and are never overwritten.Aggregates are scalars, declared alongside. This is what replaces the totals row; tables stay rectangular.
An anchor materialises a scalar into a sentence. The HTML comment is invisible in every target renderer, so the prose reads normally while the number stays machine-checkable.
The tutorial
docs/tutorial.md is the end-to-end tutorial: twenty-eight
short chapters from a plain Markdown table to a checked document, a CI job and a
script reading values back out. It teaches the language in dependency order, every
finding check can report, and the one habit that keeps a green check meaningful.
Every transcript in it is real.
There is a side-by-side reader for it at
docs/tutorial.html,
which shows each block's Markdown source next to its rendering, in lockstep.
Worked examples
docs/example-invoice.md is a complete B2B invoice
that computes itself: line items, VAT, a payment schedule derived from the gross
total, early-payment terms, a currency conversion, and a reconciliation that
proves the instalments sum to the invoice. Its appendix explains each mechanism.
docs/example-invoice-drift.md is that same
invoice with the drift shown at the top of this README — the 27 problems
transcript above is check reading this exact file, and its appendix walks
through every one of the 27 problems.
docs/example-quote-plain.md is the other
direction: a quote with no VisiMark in it at all — no vmark block, no
anchor, just a table and a total written in prose, the way an agent hands one
over before anyone has wired it up. visimark infer reads it and proposes the
rules that reproduce every number already there:
$ visimark infer docs/example-quote-plain.md
docs/example-quote-plain.md table at line 10 — 4 rows, 7 columns
column rules
Revenue = Seats * Fee 4/4 rows
Materials = Revenue * 0.08 4/4 rows
Delivered = Revenue + Materials 4/4 rows
constants worth naming
0.08 also appears as "8%" in prose, line 18
scalars matching figures in prose
72 line 17 = SUM(Seats) seats_total
27600.00 line 18 = SUM(Revenue) revenue_total
2208.00 line 19 = SUM(Materials) materials_total
29808.00 line 19 = SUM(Delivered) delivered_total
530.00 line 20 = AVG(Fee) fee_avg
no rule found — treating as inputs
Module, Format, Seats, Fee
also fits, not proposed
Delivered = Revenue * 1.08 prefers a rule over materialised columns
Delivered = Materials * 13.5 prefers a rule over materialised columns
docs/example-quote-plain.md table at line 24 — 3 rows, 4 columns
column rules
Amount = Share * unnamed1.delivered_total 3/3 rows
scalars matching figures in prose
29808.00 line 30 = SUM(Amount) amount_total
no rule found — treating as inputs
Stage, Share, Due
also fits, not proposed
Amount = Share * 29808 prefers a rule over materialised columns
4 rules, 0 aliases, 6 scalars, 6 anchors.A rule is proposed only if it reproduces every row exactly, at that column's
own precision — never a best fit, never a threshold, and also fits, not proposed is listed rather than silently dropped, because a rule over
materialised columns beating one with a bare constant is a judgment call worth
seeing. --write inserts exactly the blocks and anchors above and rewrites
nothing else. A rule that fits every row but one is never written; it is
reported as a near-miss instead — the tool telling you the document already
has a wrong number in it, before anyone runs check on it. The document's own
appendix walks through every mechanism, including that near-miss case.
infer is the way in for the document check would otherwise have nothing to
say about: one with no formulas at all. Pairing the two closes the loop —
infer gets a plain table wired up, and check keeps it that way.
How it works
VisiMark parses the document, builds a dependency graph across every sheet, sorts it topologically, and evaluates in decimal arithmetic. Circular dependencies are reported with the full path through the cycle.
flowchart LR
parse[Parse Markdown] --> deps[Build dependency graph]
deps --> sort[Topological sort]
sort --> evaluate[Evaluate in decimal]
evaluate --> cycle{"Cycle?"}
cycle -->|yes| report[Report the CYCLE path]
cycle -->|no| values[Computed values]The CLI is the product. An agent must be able to verify a document without an editor; a VS Code extension is a later, thin wrapper.
The commands
Five of them. check is the one that matters; the rest exist to get a
document into a state check can be strict about, or to explain what it did.
flowchart LR
plain[Plain table] -->|"infer --write"| wired[Rules and anchors in the file]
wired -->|edit an input or a rule| stale[Stored values disagree]
stale -->|fmt| wired
stale -->|check| fail[Exit 1]
wired -->|check| pass[Exit 0]
wired -.-> evalCmd["eval / explain"]Command | What it does | Options | What it writes | Exit codes |
| Recomputes every formula and reports the numbers that no longer agree with it | — | nothing, ever |
|
| Repairs stale values in place, by splicing the bytes of each number it owns |
| computed cells and anchored values only — never inputs, prose or headings |
|
| Works out which rules reproduce the numbers a document already has, and proposes them |
| nothing, unless |
|
| Prints the computed values — all of them, or one by name |
| nothing |
|
| Prints each sheet's inputs, rules and evaluation order |
| nothing |
|
| Prints what a builtin function does — signature, parameters, errors, worked examples — or lists all sixteen |
| nothing |
|
Every option, every exit code and every finding check can report is
tabulated in docs/cli-reference.md.
check is read-only, so it is safe to point at anything. fmt repairs stale
values and only stale values: every other kind of problem is a question a
person has to answer, so it reports those and leaves them alone.
A green check has to mean something
A document with no formulas in it has nothing to disagree with, so a checker
that only compares numbers to rules would call it clean — the most misleading
answer it could give. check reports a table with no rules attached to it as a
problem in its own right:
$ visimark check quote.md
quote.md
COVERAGE a table with no `vmark` rules — nothing in this document is checked
run `visimark infer` to derive them, or mark it `<!--vmark:no-formulas-->`
1 problem (0 stale, 1 error)Two things keep that from being annoying. It needs a table to be present, so prose — a README, a changelog — is never asked for arithmetic it does not have. And it is counted across the whole document, so a reference table that really is all input passes as long as some other table carries a rule.
When a document genuinely has nothing to derive, say so in the document:
<!--vmark:no-formulas-->That marker is the only way out, and deliberately so. It lives in the file
rather than in a workflow flag, so it travels with the content, shows up in
review, and turns up in a grep. visimark infer --write writes it for you
when it finds nothing whatsoever to derive — and refuses to when it found a
near-miss or two rules it cannot choose between, because those mean the
document does have arithmetic and wants a person to look. The marker is
checked like anything else: add rules to a marked document later and check
tells you the marker is now wrong.
In CI
The whole point of check is that it runs somewhere other than a human's
judgment, so the CI story is one line:
npx visimark check **/*.mdThat exits non-zero on the first disagreement, which is all most CI systems
need. A GitHub Actions workflow can do the same with the composite action
this repo ships (action.yml) instead of hand-rolling the
npx line:
- uses: michal-niedzwiedzki/visimark@v0.1.5
with:
files: "docs/**/*.md"There is nothing to configure and no strictness dial to find: pointing it at a
glob is the whole setup. Any document under that glob with a table and no rules
is a failure, which is why the <!--vmark:no-formulas--> marker above belongs
in the file rather than in this workflow — the decision is about a document,
not about a CI run.
A project already on remark/remark-lint adds the same checks with
remark-lint-visimark
instead — see docs/ci.md chapter 24.
A project on markdownlint adds
them with
markdownlint-rule-visimark
— see docs/ci.md chapter 25.
An agent reaches the same engine over
MCP with
visimark-mcp, which serves
every command as a tool, the authoring discipline as resources, and writes
nothing unless an operator opens the write gate:
npm i -g visimark-mcp # or: bun add -g visimark-mcp
npx visimark-mcp # or: bunx visimark-mcp
claude mcp add visimark -- npx -y visimark-mcpThe full surface is docs/mcp.md, and chapter 29 of
docs/ci.md covers running it beside a CI
check.
Diffable by construction
An .xlsx is a zip of XML: change one cell and code review can tell you the
file changed, and essentially nothing more. VisiMark documents review like
source, and that is a design constraint rather than a side effect of being
text.
fmt never re-renders the Markdown. It locates each value it owns by position
and splices the original byte buffer, so a rewrite touches the characters of
that number and nothing else — no reflowed paragraphs, no renormalised emphasis
markers, no realigned table columns, none of the four-hundred-line diff a
round-trip through a Markdown printer would produce for a one-cell change. It
also writes only what it owns: computed cells and anchored values. Input
columns, prose and headings are human territory and are never touched.
Raising one input in the worked invoice — on-call hours from 12 to 20, the
very edit the drift example above leaves unpropagated — makes fmt update 6
cells and 9 anchors, and the result is a 13-line diff in a 127-line
document. Every changed line is a figure that genuinely depends on that
input, so the diff is the propagation: a reviewer sees the VAT, the three
milestone instalments, the early-payment terms and the EUR conversion all move
together, and can check that they moved for the right reason.
The other half is that the diff contains everything. The formula lives in the document, so a changed rule shows up as a changed rule. Nothing outside the file can alter a number — no plugins, no config, no clock. And because an aggregate takes a column rather than an expression, every intermediate is materialised on the page: a total is always the sum of numbers the reviewer can see.
What it refuses to do
Where a value could mean two things, VisiMark errors rather than guesses.
Dates are ISO 8601 only — YYYY-MM-DD, ten characters. 15.10.2026 is
rejected with an offered fix, because 15 cannot be a month. 11/12/2026 is
rejected outright, because it is 11 December or 12 November depending on where
its author lives, and no amount of care catches that by reading. Thousands
separators are rejected for the same reason.
A column may carry a currency symbol or a physical unit — $5.50, 12 N —
and VisiMark strips it to compute and puts it back when it writes. What it will
not do is let one column mean two things: a column holding both $5.00 and
€5.00 is an error, not a sum. The decoration is inert, never converted and
never propagated through a formula.
A name bound twice in one scope is an error rather than a silent overwrite.
There are no boolean literals. Comparisons produce booleans and IF() consumes
them, but a boolean is never written into a cell — a materialised value is a
number, a date, or a string, so the word true in a column stays the string it
looks like.
There is no plugin architecture, and there will not be one. A document's
numbers depend on its own text and the version of VisiMark reading it, and on
nothing else — no extension modules, no config file, no environment, no
network, no clock. A registry of host-supplied functions would produce
documents whose arithmetic cannot be checked from the document, which is the
one thing the format exists to prevent. When the built-in vocabulary is too
small the answer is a new primitive in the engine, readable by everyone and
runnable by everyone; when a value genuinely comes from outside, it belongs in
an input column where a human wrote it down. Requests to grow that vocabulary —
and proposals for any other language or tooling change — go through
docs/vocabulary-catalogue.md, which records
every one and the decision on it; the review process is
docs/issue-runbook.md.
This makes the format smaller, not merely stricter: there is no locale, no
configuration, and no rule for what a bare / means.
Status
All five commands are implemented, in TypeScript. Install the visimark
command with bun add -g visimark or npm i -g visimark — it runs under
whichever of Bun or Node is on your PATH — or run it without installing with
npx visimark. (On Windows the npx / npm i -g shims need sh on PATH,
which Git Bash or WSL provide.) All three worked examples pass as the
acceptance suite — check on the drift invoice reproduces the transcript above
byte-for-byte, fmt leaves the clean invoice untouched, and infer on
docs/example-quote-plain.md — a quote with no
formulas in it at all — reproduces the transcript in that document's own
appendix. The design is
written up in docs/visimark-design.md, including the
deferred work and the known tensions; the implementation plan is
docs/superpowers/plans/2026-09-03-visimark-cli.md.
The editor support is implemented too: one language server
(packages/visimark-lsp) wrapping the same engine, and a VS Code client
(editors/vscode) — live diagnostics, fmt behind the editor's own
format-on-save, quick fixes, inlay hints, CodeLens and hover.
The extension is not published to a marketplace yet; to build and install it from a clone (Bun, like the rest of the repo's tooling):
bun run vscode-install # build, package and install (also reinstalls)
bun run vscode-uninstall # remove it againReload the window afterwards, then open docs/example-invoice-drift.md. Both
targets need the code CLI on your PATH. For development, press F5
instead — that runs the extension straight from editors/vscode in a separate
Extension Development Host, so uninstall the packaged copy first or you will see
every diagnostic twice.
The Obsidian plugin is for people who keep notes in Obsidian, read them on
a phone and will never open a terminal. It marks every computed value in reading
mode and Live Preview, explains where a value came from, and sweeps a whole vault
for notes that disagree with themselves. It is a client of the engine rather than
of the language server, it is not published to npm, and it does nothing on a note
that has no vmark block. Install it from the
latest plugin release
(the ones tagged without a v, such as 0.2.1) with
BRAT or by copying its three
files into a vault —
editors/obsidian/README.md has the steps and
says what each feature does.
Releases are tag-driven: pushing a vX.Y.Z tag publishes the engine to npm and
the extension to both the VS Code Marketplace and Open VSX. The workflow needs
three repository secrets — NPM_TOKEN, VSCE_PAT and OVSX_PAT. The checklist
for cutting one is docs/releasing.md.
For agents
skills/visimark/SKILL.md is an agent skill for
authoring and verifying these documents. Copy it to ~/.claude/skills/visimark/
to install it. Its central warning is one worth stating here too: a green check
is evidence of agreement, not of derivation. Change an input and confirm the
checker starts complaining before believing a document is wired up. The
COVERAGE finding described above exists so that an agent cannot report a
green build on a document with no build in it, but the habit is still the
better safeguard.
There is also an MCP server, visimark-mcp, for an agent
working in a repository it has never seen: npx visimark-mcp or
bunx visimark-mcp, or claude mcp add visimark -- npx -y visimark-mcp. It
serves the skill above as a resource, so the discipline arrives with the
verifier rather than separately. It is read-only unless started with
--allow-write and given a host-declared root.
Editor support is specified in
docs/visimark-editor-plugins-design.md:
one language server — continuous check as diagnostics, fmt behind
the editor's own format-on-save, quick fixes, and inlay hints that show the
computed value without touching the bytes — with VS Code as the first client.
git clone … && cd visimark && bun install
bun test # the full suite, the three examples included
bunx visimark check docs/example-invoice-drift.mdbun install builds the engine and links the visimark command into
node_modules/.bin, so bunx visimark works in a fresh clone. To run the CLI
straight from source without a build, use bun packages/visimark/src/cli/main.ts check FILE.
The project began as a CSV-based idea and moved to Markdown so that several small sheets can live inside one master document, and so that the file renders as a document rather than as data. The name is a nod to VisiCalc — the first spreadsheet software, originally developed for the Apple II by VisiCorp and later ported to the IBM PC.
Out of scope
VisiMark is deliberately not a spreadsheet replacement. No grid, no presentation layer, no cell styling, no locale, no Excel file compatibility, and no attempt at Excel's function library. Use other tools for neat presentation — and a spreadsheet when what you want is a spreadsheet.
Use VisiMark when what you want is a document: plain text, readable without the tool, reviewable in an ordinary pull request, writable by a human or an agent — with numbers that can be checked on every commit.
Available Tools
8 toolsvisimark_checkCheck a VisiMark documentARead-only
Verify that every computed number in a Markdown document still agrees with the formula that produced it, and report what does not. Findings are a successful result, not an error. A document given as content has no directory, so imports and generated artifacts cannot be verified; skipped names the ones that were not checked.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Path to a Markdown document on disk. | |
| content | No | The document as text, for a draft that is not on disk. Give this or `path`, never both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safe read-only profile, but the description adds two traits they do not cover: that a report of discrepancies is the expected success outcome, and that a content-supplied document silently cannot verify imports or generated artifacts, surfaced via 'skipped'. That is meaningful operational context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action before the caveats, and each sentence carries distinct information (what it does, how to read findings, input-mode limitation). No filler, though the final clause packing 'skipped' and the import caveat is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must sketch the return shape, and it does so at a workable level: findings plus a 'skipped' list. It does not detail the structure of individual findings or the exact return format, but for a two-param, zero-required read tool this is close to sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value the schema does not: the mutual exclusivity of path and content ('Give this or path, never both') and the functional consequence of choosing content (no directory, so imports/artifacts are unverifiable). That is actual semantic enrichment of the parameter choice.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (verify) and a precise resource (computed numbers in a Markdown document vs. their producing formula), plus what the result reports. It is clearly a validation pass, which is inferable as distinct from siblings like visimark_eval or visimark_infer, but it never names or contrasts an alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real usage context: 'Findings are a successful result, not an error' steers the agent away from treating a non-empty result as failure, and it explains that content-based documents cannot verify imports. However, there is no explicit when-to-use-this-vs-alternative guidance against visimark_eval or visimark_ref, leaving sibling selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visimark_evalEvaluate a VisiMark documentARead-only
Evaluate a document and return its values, its assertions and its charts. Give get to select a single named value. Give scenarioPath or scenarioContent to substitute parameters and see what moves — a draft scenario against a draft document, with no temp file. A false assertion is a successful result reporting problems.
| Name | Required | Description | Default |
|---|---|---|---|
| get | No | One value's qualified name, e.g. `lines.gross_total`. | |
| path | No | Path to a Markdown document on disk. | |
| content | No | The document as text, for a draft that is not on disk. Give this or `path`, never both. | |
| scenarioPath | No | Path to a scenario JSON file. | |
| scenarioContent | No | A scenario as JSON text. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read-only, non-destructive profile, and the description adds genuinely non-obvious context beyond that: a false assertion is reported as a successful result, and scenario evaluation needs 'no temp file' (no side effects on disk). It still doesn't characterize the response shape or whether failures throw vs. report.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core verb and result, then parameter routing, then the critical caveat about false assertions. No filler or repetition of schema text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so (values, assertions, charts). It covers the main invocation modes; the only gap is the relationship/ordering between `path`/`content` and evaluation errors, which is mostly covered by 'never both'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so 3 is the baseline, but the description earns above it by explaining functional semantics the schema does not: what `get` selects, that scenarioPath/scenarioContent substitute parameters, and that scenario evaluation is a draft-against-draft, no-temp-file operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Evaluate a document') and enumerates the return surface (values, assertions, charts). It does not name any sibling (visimark_check, visimark_ref, visimark_infer), so an agent must infer why evaluation differs from checking or referencing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives conditional guidance for two parameters ('Give `get` to select a single named value', 'Give `scenarioPath` or `scenarioContent` to substitute parameters'), which implies usage. However it never states when to prefer this tool over visimark_check or the other siblings, and offers no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visimark_explainExplain a VisiMark document's structureARead-only
Describe what a document declares: its sheets, input columns, computed rules, anchored scalars, derived precision and imports. Give sheet to narrow it.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Path to a Markdown document on disk. | |
| sheet | No | A sheet id, without the leading `#`. | |
| content | No | The document as text, for a draft that is not on disk. Give this or `path`, never both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds that output is a structural summary and that omitting `sheet` yields the whole document, but says nothing about failure modes (missing file vs. malformed document) or the shape of the explanation. Adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler, and the enumeration of what gets explained is placed before the scoping tip. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on the burden of describing the return content by listing the declared elements. Safety is covered by annotations. The remaining gap is error/edge-case behavior, which is minor for a read-only inspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters (path, sheet, content) are already documented, including the path/content mutual exclusion. The description's only added semantic is that `sheet` narrows the output scope, which is marginal beyond the schema's own wording. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ("Describe") and resource ("what a document declares") and enumerates the concrete categories returned: sheets, input columns, computed rules, anchored scalars, derived precision, imports. This is far more than a restatement of the title. It does not, however, name or distinguish itself from siblings like visimark_ref or visimark_eval, which likely also surface document information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: inspecting a document's declared structure. The one directive offered ("Give `sheet` to narrow it") tells the agent how to scope output but not when this tool is preferable to visimark_ref, visimark_check, or visimark_eval. No exclusions or alternative-routing guidance is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visimark_fmtPlan a VisiMark repairARead-only
Plan the edits that would bring a document's stored numbers back into agreement with its formulas, and name the generated artifacts it would write. Nothing is written: pass the returned plan to visimark_fmt_apply to land it. It repairs stale computed values, anchors and generated artifacts, and nothing else — never prose, never a column with no rule, never any other finding class. A document given as content has no directory, so imports and generated artifacts cannot be verified; skipped names the ones that were not checked.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Path to a Markdown document on disk. | |
| content | No | The document as text, for a draft that is not on disk. Give this or `path`, never both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description reinforces and extends this: it lists exactly what is repaired (stale computed values, anchors, generated artifacts), what is never touched (prose, columns with no rule, other finding classes), and the `skipped` behavior for content-only input. That is meaningful behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each load-bearing: purpose first, then the no-write/apply contract, then the repair scope, then the content-mode caveat. No filler and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry return-value context, and it does name the key surfaces (the plan's generated artifacts and `skipped`). It stops short of describing the full shape of the returned plan, so a small gap remains for an agent that needs to inspect the plan programmatically.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the `path`/`content` contract is already documented; the description still adds semantic meaning by explaining the consequence of using `content` (no directory, so imports and generated artifacts cannot be verified) and what `skipped` then reports. That goes beyond the schema text, though it covers only one of the two parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (plan the edits) and resource (a document's stored numbers vs. its formulas), and adds that it names the artifacts it would write. It explicitly routes to the sibling `visimark_fmt_apply` for applying the plan, so an agent can distinguish this from the write-side tool without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use and a next-step alternative ('pass the returned plan to `visimark_fmt_apply`'), plus the destructive-path caveat ('Nothing is written'). It does not distinguish this plan step from the other read-side siblings such as `visimark_check` or `visimark_ref`, which is the remaining gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visimark_fmt_applyApply a VisiMark repairADestructive
Land the plan visimark_fmt returned: splice the document's computed cells and anchors, and write the generated artifacts it named. Refused if the document changed since the plan was computed, and refused entirely unless the operator started the server with --allow-write and the host declared a root.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | The document to write. There is nothing to apply to a string. | |
| plan | Yes | The result of `visimark_fmt`, unedited apart from dropping entries you rejected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so safety direction is covered. The description goes further, disclosing a staleness guard (refused if the document changed since the plan was computed), an operator precondition (--allow-write), and a host requirement (declared root) — meaningful behavioral context beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, action front-loaded, with preconditions trailing. No filler; each clause either defines the operation or a condition that gates it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema, the description supplies the operation, the plan-sourcing dependency, and both environmental preconditions an agent must satisfy before calling. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with two documented parameters, so the schema already carries the primary semantics. The description adds only indirect hints ('artifacts it named', dropping rejected entries), which is the expected baseline rather than extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names specific verbs (splice, write) and resources (computed cells, anchors, generated artifacts), and explicitly ties itself to the sibling visimark_fmt as the plan producer. An agent can distinguish this apply step from the fmt planning step without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the exact upstream relationship ('Land the plan visimark_fmt returned') and the two refusal conditions (document changed since plan, missing --allow-write or host root), making clear when the call will fail rather than succeed. This is explicit when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visimark_inferPropose VisiMark rules for a plain tableARead-only
Derive the formulas an existing Markdown table already obeys, and propose them as vmark rules. Run this before hand-authoring rules for a document that already has its numbers. Returns proposals only and never writes; visimark_infer_apply writes them.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Path to a Markdown document on disk. | |
| content | No | The document as text, for a draft that is not on disk. Give this or `path`, never both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is partially given. The description adds real value by confirming the output is 'proposals only and never writes' and that a separate tool performs the mutation, which clarifies the non-mutating contract beyond the hint. It stops short of describing proposal format or behavior on tables with no derivable formulas.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and benefit before the usage rule and the write-safety caveat. Every sentence carries distinct information; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, annotation-covered tool with no output schema, the description covers purpose, usage, and the non-writing contract. Only edge-case behavior (e.g., a table with no derivable formulas, or error handling when neither path nor content is supplied) is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself states the 'give this or `path`, never both' constraint, so the baseline is 3. The description adds no parameter-level syntax or format detail beyond what the schema documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('Derive the formulas an existing Markdown table already obeys') and the concrete artifact produced ('propose them as `vmark` rules'). It also distinguishes itself from the sibling visimark_infer_apply, so an agent can tell the propose step from the write step without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when: 'Run this before hand-authoring rules for a document that already has its numbers.' And it names the alternative with its selecting condition: proposals are returned here, while `visimark_infer_apply` writes them. The routing decision is fully determined by the text.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visimark_infer_applyApply proposed VisiMark rulesADestructive
Insert the vmark rules visimark_infer proposed. Pass the plan you were given, with any proposal you rejected removed. Refused if the document changed since the plan was computed, and refused entirely unless the operator started the server with --allow-write and the host declared a root.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | The document to write. There is nothing to apply to a string. | |
| plan | Yes | The result of `visimark_infer`, unedited apart from dropping entries you rejected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false and destructiveHint=true; the description adds real behavioral context: the write is gated on the operator starting the server with --allow-write and the host declaring a root, and it is refused as stale if the document changed since the plan was computed. That is exactly the extra information an agent needs before invoking a mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler: the action and its source come first, then the input-shaping rule, then the two refusal conditions. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive write with no output schema, the description covers provenance, input shaping, and both failure gates. It does not say what happens on success (e.g., counts of applied rules or partial-application behavior), a minor gap given no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, including the note that plan must be unedited apart from dropping rejected entries. The description restates the plan handling but adds no syntax, format, or constraint detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ("Insert the vmark rules") plus the exact resource and its origin ("the rules visimark_infer proposed"), which separates it from the sibling visimark_infer that only proposes them. An agent can tell it is the write-side step of the infer/apply pair without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the required workflow (use the plan returned by visimark_infer, drop rejected entries) and the two refusal conditions (document changed, missing --allow-write/root). It does not contrast with visimark_fmt_apply, so the agent must infer which apply tool fits which artifact.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visimark_refVisiMark function referenceARead-only
Look up VisiMark's builtin functions: signature, arity, precision, errors and examples. Give name to look it up, or omit name for the whole reference. Reads no file.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | A builtin's name, e.g. `SUM`. Omit for the whole reference. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds the non-obvious trait 'Reads no file,' signaling this is a pure in-memory reference lookup rather than a file-consuming operation like eval/check, and it discloses the shape of the returned content. Useful context beyond the safety annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the returned content and followed by the invocation modes. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing what a lookup returns (signature, arity, precision, errors, examples), and the readOnly annotations cover the safety profile. The only gap is sibling differentiation, which is not strictly required for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single `name` parameter is fully documented as an example (`SUM`) with the omit behavior. The description restates 'omit `name` for the whole reference' without adding new semantics, so the schema is doing the heavy lifting and the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up) and resource (builtin functions) and enumerates what the entry contains: signature, arity, precision, errors and examples. It does not, however, differentiate itself from siblings like visimark_explain or visimark_check, which plausibly overlap with 'explain a function'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the two invocation modes (pass `name` to look up one builtin, omit it for the whole reference), which is genuinely operative guidance. It never states when to prefer this tool over visimark_explain or visimark_check, so the routing decision is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
visimark_check - First observed
visimark_eval - First observed
visimark_explain - First observed
visimark_fmt - First observed
visimark_fmt_apply - First observed
visimark_infer - First observed
visimark_infer_apply - First observed
visimark_ref
TDQS
Scored across 8 tools
Each tool has a clearly distinct role: ref (lookup builtins), explain (describe declarations), eval (compute values), check (verify agreement), infer (propose rules), fmt (plan repairs), and their _apply counterparts. The plan-then-apply split is explicit and the read-only analyzers (check vs eval vs explain) are differentiated by their descriptions.
All tools use a consistent visimark_ snake_case prefix with verb-based names, and the _apply suffix reliably marks write actions. Minor deviation: 'ref' and 'fmt' are abbreviations while others are full verbs, but the pattern remains predictable.
Eight tools is well-scoped for a document-analysis/formula-verification server, with each tool earning its place. Read/plan/apply responsibilities are cleanly separated without redundancy.
The surface covers the full lifecycle: reference lookup, document description, evaluation, verification, inference, formatting, and the apply operations to land each plan, all with write-safety gating. No obvious gaps for the stated domain.
Related MCP Connectors
Markdown workspace for AI agents: read, write, organize, and share markdown documents.
Create, edit, review, and explicitly publish Live or Snapshot Markdown Documents in mdedit.ai.
Free mechanical checks for AI text: unnamed counts, dangling references, bad arithmetic, misquotes.
Deterministic signed verification of numeric & financial claims for AI agents & spreadsheets.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables intelligent Markdown document analysis including extracting table of contents with hierarchy, detecting numbering issues like duplicates and discontinuities, and generating formatted TOC content in multiple output formats.1Apache 2.0
- AlicenseBqualityFmaintenanceProvides intelligent manuscript analysis and writing assistance for markdown projects, including semantic search, quality checks, terminology consistency, link validation, progress tracking, and comprehensive writing statistics.35131 npm13MIT
- AlicenseNot gradedqualityBmaintenanceEnables sharing Markdown documents for collaborative review with inline annotations and structured change requests.7 npmMIT

openmarkdownofficial
FlicenseNot gradedqualityAmaintenanceLocal-first Markdown editor whose MCP server lets a coding agent and a human co-edit the same .md file — it opens and reveals files in the editor and reads or section-edits their contents. Tools: open_file, reveal, read_section, write_section, wait_for_change.63-