Skip to main content
Glama

VisiMark

A document integrity layer for Markdown: it checks that every number in a document still matches the formula that produced it, so an agent's correct formula can't ship with a wrong total.

VisiMark checks the numbers in Markdown documents, especially ones an agent wrote. Every computed value carries its formula, a machine proves the two still agree, and CI fails when they don't. The document stays plain Markdown.

This matters most when the Markdown is written or edited by an AI agent. Agents are reliable at writing formulas and unreliable at the arithmetic those formulas describe: Net = Price * Qty is something an agent gets right, 50.00 is something it guesses. Write the formula instead of the number and the number stops being a claim and becomes a derivation — reviewable in a diff, re-runnable, enforceable in CI. The agent writes, VisiMark verifies, Git records the result.

A VisiMark document is an ordinary Markdown file. It renders correctly on GitHub, in VS Code preview, and through pandoc to both HTML and Word, today, with no plugin — verified, not assumed. What VisiMark adds is that every computed number in it carries the formula that produced it, that a machine can prove the two still agree, and that a change to either shows up as a small, readable diff.

A wrong invoice that renders clean

docs/example-invoice-drift.md is a real B2B invoice after someone raised the on-call hours from 12 to 20 and updated nothing that depends on it. It renders as a clean, plausible invoice on GitHub, in a Markdown preview, anywhere — which is the entire argument for this project. Every number in it carries the formula that produced it; the syntax follows further down. Here is visimark check reading the same file a reviewer just skimmed and approved:

$ visimark check docs/example-invoice-drift.md
docs/example-invoice-drift.md

  STALE   lines.Net       · On-call support         3120.00 ≠ 5200.00    Qty * Rate
  STALE   lines.VAT       · On-call support          717.60 ≠ 1196.00    Net * vat
  STALE   lines.Gross     · On-call support         3837.60 ≠ 6396.00    Net + VAT
  STALE   lines.Gross     · Discovery workshop      4428.50 ≠ 4428.00    Net + VAT
  STALE   lines.net_total                          23300.00 ≠ 25380.00   SUM(Net)
  STALE   lines.vat_total                           5359.00 ≠ 5837.40    SUM(VAT)
  STALE   lines.gross_total                        28659.00 ≠ 31217.40   SUM(Gross)
  STALE   lines.gross_total                        28659.00 ≠ 31217.40   SUM(Gross)
  STALE   schedule.Amount · Signature               8597.70 ≠ 9365.22    Share * lines.gross_total
  STALE   schedule.Amount · Delivery of backend    11463.60 ≠ 12486.96   Share * lines.gross_total
  STALE   schedule.Amount · Acceptance              8597.70 ≠ 9365.22    Share * lines.gross_total
  STALE   schedule.covered                         28659.00 ≠ 31217.40   SUM(Amount)
  STALE   terms.early_pay_total                    28085.82 ≠ 30593.05
  STALE   terms.early_pay_saved                      573.18 ≠ 624.35
  STALE   8 prose anchors bound to the values above

  DATE    schedule.Due    · Delivery of backend   "15.10.2026"
          Dates must be ISO 8601 calendar dates: YYYY-MM-DD.
          Unambiguous — `visimark fmt --fix-dates` rewrites it to 2026-10-15.

  DATE    schedule.Due    · Acceptance            "11/12/2026"
          Dates must be ISO 8601 calendar dates: YYYY-MM-DD.
          Ambiguous: 2026-12-11 or 2026-11-12, 29 days apart. Fix by hand.

  NOTE    schedule.Days   · 2 rows not verified (upstream DATE errors)

  UNDEF   terms.eur_total   unknown name `fx_rate`
          did you mean `fx_eur`?

  VECTOR  recon.variance    `schedule.Amount` is a column, not a value.
          Wrap it in an aggregate: SUM(schedule.Amount)

  CYCLE   late_fees.base → late_fees.fee → late_fees.total → late_fees.base

  27 problems (22 stale, 5 errors)
$ echo $?
1

Twenty-seven problems: a payment date ambiguous by twenty-nine days, a cell someone nudged by hand to make a column look right, a circular reference — all invisible on the rendered page, all caught before a human had to notice.

Related MCP server: Claude Writer's Aid MCP

Why this matters now

More and more of the documents that carry numbers are written as text and kept in Git: reports and research notes, estimates and budgets, quotes and invoices, project plans, financial summaries. Text is excellent for collaboration, review and version control. But ordinary Markdown has no way to say this number came from these inputs, and it must still agree with them — so the moment an input changes, every figure downstream of it is a guess until a human re-checks it by hand.

VisiMark adds that missing layer. The working stack for collaborating with an agent is a text editor over lightly formatted artifacts; Markdown already covers the prose, and this covers the calculation. The arithmetic is not the value — the audit trail is.

What it looks like

| Item  | Price | Qty |   Net |
|-------|------:|----:|------:|
| pen   |  5.00 |  10 | 50.00 |
| paper |  0.10 | 100 | 10.00 |

```vmark #order
Net   = Price * Qty
total = SUM(Net)
```

Order total: **60.00**<!--vmark=order.total-->

Four ideas, and that is the whole format:

  • A vmark block declares formulas for the table above it, and names the sheet.

  • Columns are uniform — Net = Price * Qty is one rule for every row, not a formula per cell. Columns with no rule are inputs, and are never overwritten.

  • Aggregates are scalars, declared alongside. This is what replaces the totals row; tables stay rectangular.

  • An anchor materialises a scalar into a sentence. The HTML comment is invisible in every target renderer, so the prose reads normally while the number stays machine-checkable.

The tutorial

docs/tutorial.md is the end-to-end tutorial: twenty-eight short chapters from a plain Markdown table to a checked document, a CI job and a script reading values back out. It teaches the language in dependency order, every finding check can report, and the one habit that keeps a green check meaningful. Every transcript in it is real.

There is a side-by-side reader for it at docs/tutorial.html, which shows each block's Markdown source next to its rendering, in lockstep.

Worked examples

docs/example-invoice.md is a complete B2B invoice that computes itself: line items, VAT, a payment schedule derived from the gross total, early-payment terms, a currency conversion, and a reconciliation that proves the instalments sum to the invoice. Its appendix explains each mechanism.

docs/example-invoice-drift.md is that same invoice with the drift shown at the top of this README — the 27 problems transcript above is check reading this exact file, and its appendix walks through every one of the 27 problems.

docs/example-quote-plain.md is the other direction: a quote with no VisiMark in it at all — no vmark block, no anchor, just a table and a total written in prose, the way an agent hands one over before anyone has wired it up. visimark infer reads it and proposes the rules that reproduce every number already there:

$ visimark infer docs/example-quote-plain.md
docs/example-quote-plain.md  table at line 10 — 4 rows, 7 columns

  column rules
    Revenue    = Seats * Fee                  4/4 rows
    Materials  = Revenue * 0.08               4/4 rows
    Delivered  = Revenue + Materials          4/4 rows

  constants worth naming
    0.08   also appears as "8%" in prose, line 18

  scalars matching figures in prose
    72        line 17  = SUM(Seats)                seats_total
    27600.00  line 18  = SUM(Revenue)              revenue_total
    2208.00   line 19  = SUM(Materials)            materials_total
    29808.00  line 19  = SUM(Delivered)            delivered_total
    530.00    line 20  = AVG(Fee)                  fee_avg

  no rule found — treating as inputs
    Module, Format, Seats, Fee

  also fits, not proposed
    Delivered = Revenue * 1.08    prefers a rule over materialised columns
    Delivered = Materials * 13.5  prefers a rule over materialised columns

docs/example-quote-plain.md  table at line 24 — 3 rows, 4 columns

  column rules
    Amount  = Share * unnamed1.delivered_total  3/3 rows

  scalars matching figures in prose
    29808.00  line 30  = SUM(Amount)               amount_total

  no rule found — treating as inputs
    Stage, Share, Due

  also fits, not proposed
    Amount = Share * 29808    prefers a rule over materialised columns

4 rules, 0 aliases, 6 scalars, 6 anchors.

A rule is proposed only if it reproduces every row exactly, at that column's own precision — never a best fit, never a threshold, and also fits, not proposed is listed rather than silently dropped, because a rule over materialised columns beating one with a bare constant is a judgment call worth seeing. --write inserts exactly the blocks and anchors above and rewrites nothing else. A rule that fits every row but one is never written; it is reported as a near-miss instead — the tool telling you the document already has a wrong number in it, before anyone runs check on it. The document's own appendix walks through every mechanism, including that near-miss case.

infer is the way in for the document check would otherwise have nothing to say about: one with no formulas at all. Pairing the two closes the loop — infer gets a plain table wired up, and check keeps it that way.

How it works

VisiMark parses the document, builds a dependency graph across every sheet, sorts it topologically, and evaluates in decimal arithmetic. Circular dependencies are reported with the full path through the cycle.

flowchart LR
  parse[Parse Markdown] --> deps[Build dependency graph]
  deps --> sort[Topological sort]
  sort --> evaluate[Evaluate in decimal]
  evaluate --> cycle{"Cycle?"}
  cycle -->|yes| report[Report the CYCLE path]
  cycle -->|no| values[Computed values]

The CLI is the product. An agent must be able to verify a document without an editor; a VS Code extension is a later, thin wrapper.

The commands

Five of them. check is the one that matters; the rest exist to get a document into a state check can be strict about, or to explain what it did.

flowchart LR
  plain[Plain table] -->|"infer --write"| wired[Rules and anchors in the file]
  wired -->|edit an input or a rule| stale[Stored values disagree]
  stale -->|fmt| wired
  stale -->|check| fail[Exit 1]
  wired -->|check| pass[Exit 0]
  wired -.-> evalCmd["eval / explain"]

Command

What it does

Options

What it writes

Exit codes

visimark check FILE...

Recomputes every formula and reports the numbers that no longer agree with it

—

nothing, ever

0 clean · 1 findings · 2 bad usage or unreadable file

visimark fmt FILE...

Repairs stale values in place, by splicing the bytes of each number it owns

--fix-dates rewrites unambiguous non-ISO dates

computed cells and anchored values only — never inputs, prose or headings

0 clean · 1 problems it cannot fix remain · 2 bad usage or unreadable file

visimark infer FILE...

Works out which rules reproduce the numbers a document already has, and proposes them

--write inserts what it proposed

nothing, unless --write — and then it only ever inserts

0 whatever it finds, because it is advisory · 2 bad usage or unreadable file

visimark eval FILE

Prints the computed values — all of them, or one by name

--get NAME, --json

nothing

0 · 2 bad usage, unreadable file, or no such name

visimark explain FILE

Prints each sheet's inputs, rules and evaluation order

#sheet limits it to one sheet

nothing

0 · 2 bad usage, unreadable file, or no such sheet

visimark ref [NAME]

Prints what a builtin function does — signature, parameters, errors, worked examples — or lists all sixteen

--json

nothing

0 · 2 no such function

Every option, every exit code and every finding check can report is tabulated in docs/cli-reference.md.

check is read-only, so it is safe to point at anything. fmt repairs stale values and only stale values: every other kind of problem is a question a person has to answer, so it reports those and leaves them alone.

A green check has to mean something

A document with no formulas in it has nothing to disagree with, so a checker that only compares numbers to rules would call it clean — the most misleading answer it could give. check reports a table with no rules attached to it as a problem in its own right:

$ visimark check quote.md
quote.md

  COVERAGE a table with no `vmark` rules — nothing in this document is checked
           run `visimark infer` to derive them, or mark it `<!--vmark:no-formulas-->`

  1 problem (0 stale, 1 error)

Two things keep that from being annoying. It needs a table to be present, so prose — a README, a changelog — is never asked for arithmetic it does not have. And it is counted across the whole document, so a reference table that really is all input passes as long as some other table carries a rule.

When a document genuinely has nothing to derive, say so in the document:

<!--vmark:no-formulas-->

That marker is the only way out, and deliberately so. It lives in the file rather than in a workflow flag, so it travels with the content, shows up in review, and turns up in a grep. visimark infer --write writes it for you when it finds nothing whatsoever to derive — and refuses to when it found a near-miss or two rules it cannot choose between, because those mean the document does have arithmetic and wants a person to look. The marker is checked like anything else: add rules to a marked document later and check tells you the marker is now wrong.

In CI

The whole point of check is that it runs somewhere other than a human's judgment, so the CI story is one line:

npx visimark check **/*.md

That exits non-zero on the first disagreement, which is all most CI systems need. A GitHub Actions workflow can do the same with the composite action this repo ships (action.yml) instead of hand-rolling the npx line:

- uses: michal-niedzwiedzki/visimark@v0.1.5
  with:
    files: "docs/**/*.md"

There is nothing to configure and no strictness dial to find: pointing it at a glob is the whole setup. Any document under that glob with a table and no rules is a failure, which is why the <!--vmark:no-formulas--> marker above belongs in the file rather than in this workflow — the decision is about a document, not about a CI run.

A project already on remark/remark-lint adds the same checks with remark-lint-visimark instead — see docs/ci.md chapter 24.

A project on markdownlint adds them with markdownlint-rule-visimark — see docs/ci.md chapter 25.

An agent reaches the same engine over MCP with visimark-mcp, which serves every command as a tool, the authoring discipline as resources, and writes nothing unless an operator opens the write gate:

npm i -g visimark-mcp   # or: bun add -g visimark-mcp
npx visimark-mcp        # or: bunx visimark-mcp
claude mcp add visimark -- npx -y visimark-mcp

The full surface is docs/mcp.md, and chapter 29 of docs/ci.md covers running it beside a CI check.

Diffable by construction

An .xlsx is a zip of XML: change one cell and code review can tell you the file changed, and essentially nothing more. VisiMark documents review like source, and that is a design constraint rather than a side effect of being text.

fmt never re-renders the Markdown. It locates each value it owns by position and splices the original byte buffer, so a rewrite touches the characters of that number and nothing else — no reflowed paragraphs, no renormalised emphasis markers, no realigned table columns, none of the four-hundred-line diff a round-trip through a Markdown printer would produce for a one-cell change. It also writes only what it owns: computed cells and anchored values. Input columns, prose and headings are human territory and are never touched.

Raising one input in the worked invoice — on-call hours from 12 to 20, the very edit the drift example above leaves unpropagated — makes fmt update 6 cells and 9 anchors, and the result is a 13-line diff in a 127-line document. Every changed line is a figure that genuinely depends on that input, so the diff is the propagation: a reviewer sees the VAT, the three milestone instalments, the early-payment terms and the EUR conversion all move together, and can check that they moved for the right reason.

The other half is that the diff contains everything. The formula lives in the document, so a changed rule shows up as a changed rule. Nothing outside the file can alter a number — no plugins, no config, no clock. And because an aggregate takes a column rather than an expression, every intermediate is materialised on the page: a total is always the sum of numbers the reviewer can see.

What it refuses to do

Where a value could mean two things, VisiMark errors rather than guesses.

Dates are ISO 8601 only — YYYY-MM-DD, ten characters. 15.10.2026 is rejected with an offered fix, because 15 cannot be a month. 11/12/2026 is rejected outright, because it is 11 December or 12 November depending on where its author lives, and no amount of care catches that by reading. Thousands separators are rejected for the same reason.

A column may carry a currency symbol or a physical unit — $5.50, 12 N — and VisiMark strips it to compute and puts it back when it writes. What it will not do is let one column mean two things: a column holding both $5.00 and €5.00 is an error, not a sum. The decoration is inert, never converted and never propagated through a formula.

A name bound twice in one scope is an error rather than a silent overwrite.

There are no boolean literals. Comparisons produce booleans and IF() consumes them, but a boolean is never written into a cell — a materialised value is a number, a date, or a string, so the word true in a column stays the string it looks like.

There is no plugin architecture, and there will not be one. A document's numbers depend on its own text and the version of VisiMark reading it, and on nothing else — no extension modules, no config file, no environment, no network, no clock. A registry of host-supplied functions would produce documents whose arithmetic cannot be checked from the document, which is the one thing the format exists to prevent. When the built-in vocabulary is too small the answer is a new primitive in the engine, readable by everyone and runnable by everyone; when a value genuinely comes from outside, it belongs in an input column where a human wrote it down. Requests to grow that vocabulary — and proposals for any other language or tooling change — go through docs/vocabulary-catalogue.md, which records every one and the decision on it; the review process is docs/issue-runbook.md.

This makes the format smaller, not merely stricter: there is no locale, no configuration, and no rule for what a bare / means.

Status

All five commands are implemented, in TypeScript. Install the visimark command with bun add -g visimark or npm i -g visimark — it runs under whichever of Bun or Node is on your PATH — or run it without installing with npx visimark. (On Windows the npx / npm i -g shims need sh on PATH, which Git Bash or WSL provide.) All three worked examples pass as the acceptance suite — check on the drift invoice reproduces the transcript above byte-for-byte, fmt leaves the clean invoice untouched, and infer on docs/example-quote-plain.md — a quote with no formulas in it at all — reproduces the transcript in that document's own appendix. The design is written up in docs/visimark-design.md, including the deferred work and the known tensions; the implementation plan is docs/superpowers/plans/2026-09-03-visimark-cli.md.

The editor support is implemented too: one language server (packages/visimark-lsp) wrapping the same engine, and a VS Code client (editors/vscode) — live diagnostics, fmt behind the editor's own format-on-save, quick fixes, inlay hints, CodeLens and hover.

The extension is not published to a marketplace yet; to build and install it from a clone (Bun, like the rest of the repo's tooling):

bun run vscode-install     # build, package and install (also reinstalls)
bun run vscode-uninstall   # remove it again

Reload the window afterwards, then open docs/example-invoice-drift.md. Both targets need the code CLI on your PATH. For development, press F5 instead — that runs the extension straight from editors/vscode in a separate Extension Development Host, so uninstall the packaged copy first or you will see every diagnostic twice.

The Obsidian plugin is for people who keep notes in Obsidian, read them on a phone and will never open a terminal. It marks every computed value in reading mode and Live Preview, explains where a value came from, and sweeps a whole vault for notes that disagree with themselves. It is a client of the engine rather than of the language server, it is not published to npm, and it does nothing on a note that has no vmark block. Install it from the latest plugin release (the ones tagged without a v, such as 0.2.1) with BRAT or by copying its three files into a vault — editors/obsidian/README.md has the steps and says what each feature does.

Releases are tag-driven: pushing a vX.Y.Z tag publishes the engine to npm and the extension to both the VS Code Marketplace and Open VSX. The workflow needs three repository secrets — NPM_TOKEN, VSCE_PAT and OVSX_PAT. The checklist for cutting one is docs/releasing.md.

For agents

skills/visimark/SKILL.md is an agent skill for authoring and verifying these documents. Copy it to ~/.claude/skills/visimark/ to install it. Its central warning is one worth stating here too: a green check is evidence of agreement, not of derivation. Change an input and confirm the checker starts complaining before believing a document is wired up. The COVERAGE finding described above exists so that an agent cannot report a green build on a document with no build in it, but the habit is still the better safeguard.

There is also an MCP server, visimark-mcp, for an agent working in a repository it has never seen: npx visimark-mcp or bunx visimark-mcp, or claude mcp add visimark -- npx -y visimark-mcp. It serves the skill above as a resource, so the discipline arrives with the verifier rather than separately. It is read-only unless started with --allow-write and given a host-declared root.

Editor support is specified in docs/visimark-editor-plugins-design.md: one language server — continuous check as diagnostics, fmt behind the editor's own format-on-save, quick fixes, and inlay hints that show the computed value without touching the bytes — with VS Code as the first client.

git clone … && cd visimark && bun install
bun test                    # the full suite, the three examples included
bunx visimark check docs/example-invoice-drift.md

bun install builds the engine and links the visimark command into node_modules/.bin, so bunx visimark works in a fresh clone. To run the CLI straight from source without a build, use bun packages/visimark/src/cli/main.ts check FILE.

The project began as a CSV-based idea and moved to Markdown so that several small sheets can live inside one master document, and so that the file renders as a document rather than as data. The name is a nod to VisiCalc — the first spreadsheet software, originally developed for the Apple II by VisiCorp and later ported to the IBM PC.

Out of scope

VisiMark is deliberately not a spreadsheet replacement. No grid, no presentation layer, no cell styling, no locale, no Excel file compatibility, and no attempt at Excel's function library. Use other tools for neat presentation — and a spreadsheet when what you want is a spreadsheet.

Use VisiMark when what you want is a document: plain text, readable without the tool, reviewable in an ordinary pull request, writable by a human or an agent — with numbers that can be checked on every commit.

Available Tools

8 tools
visimark_checkCheck a VisiMark documentA
Read-only

Verify that every computed number in a Markdown document still agrees with the formula that produced it, and report what does not. Findings are a successful result, not an error. A document given as content has no directory, so imports and generated artifacts cannot be verified; skipped names the ones that were not checked.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoPath to a Markdown document on disk.
contentNoThe document as text, for a draft that is not on disk. Give this or `path`, never both.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safe read-only profile, but the description adds two traits they do not cover: that a report of discrepancies is the expected success outcome, and that a content-supplied document silently cannot verify imports or generated artifacts, surfaced via 'skipped'. That is meaningful operational context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action before the caveats, and each sentence carries distinct information (what it does, how to read findings, input-mode limitation). No filler, though the final clause packing 'skipped' and the import caveat is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must sketch the return shape, and it does so at a workable level: findings plus a 'skipped' list. It does not detail the structure of individual findings or the exact return format, but for a two-param, zero-required read tool this is close to sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds value the schema does not: the mutual exclusivity of path and content ('Give this or path, never both') and the functional consequence of choosing content (no directory, so imports/artifacts are unverifiable). That is actual semantic enrichment of the parameter choice.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (verify) and a precise resource (computed numbers in a Markdown document vs. their producing formula), plus what the result reports. It is clearly a validation pass, which is inferable as distinct from siblings like visimark_eval or visimark_infer, but it never names or contrasts an alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real usage context: 'Findings are a successful result, not an error' steers the agent away from treating a non-empty result as failure, and it explains that content-based documents cannot verify imports. However, there is no explicit when-to-use-this-vs-alternative guidance against visimark_eval or visimark_ref, leaving sibling selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visimark_evalEvaluate a VisiMark documentA
Read-only

Evaluate a document and return its values, its assertions and its charts. Give get to select a single named value. Give scenarioPath or scenarioContent to substitute parameters and see what moves — a draft scenario against a draft document, with no temp file. A false assertion is a successful result reporting problems.

ParametersJSON Schema
NameRequiredDescriptionDefault
getNoOne value's qualified name, e.g. `lines.gross_total`.
pathNoPath to a Markdown document on disk.
contentNoThe document as text, for a draft that is not on disk. Give this or `path`, never both.
scenarioPathNoPath to a scenario JSON file.
scenarioContentNoA scenario as JSON text.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a safe read-only, non-destructive profile, and the description adds genuinely non-obvious context beyond that: a false assertion is reported as a successful result, and scenario evaluation needs 'no temp file' (no side effects on disk). It still doesn't characterize the response shape or whether failures throw vs. report.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core verb and result, then parameter routing, then the critical caveat about false assertions. No filler or repetition of schema text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so (values, assertions, charts). It covers the main invocation modes; the only gap is the relationship/ordering between `path`/`content` and evaluation errors, which is mostly covered by 'never both'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so 3 is the baseline, but the description earns above it by explaining functional semantics the schema does not: what `get` selects, that scenarioPath/scenarioContent substitute parameters, and that scenario evaluation is a draft-against-draft, no-temp-file operation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Evaluate a document') and enumerates the return surface (values, assertions, charts). It does not name any sibling (visimark_check, visimark_ref, visimark_infer), so an agent must infer why evaluation differs from checking or referencing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives conditional guidance for two parameters ('Give `get` to select a single named value', 'Give `scenarioPath` or `scenarioContent` to substitute parameters'), which implies usage. However it never states when to prefer this tool over visimark_check or the other siblings, and offers no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visimark_explainExplain a VisiMark document's structureA
Read-only

Describe what a document declares: its sheets, input columns, computed rules, anchored scalars, derived precision and imports. Give sheet to narrow it.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoPath to a Markdown document on disk.
sheetNoA sheet id, without the leading `#`.
contentNoThe document as text, for a draft that is not on disk. Give this or `path`, never both.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds that output is a structural summary and that omitting `sheet` yields the whole document, but says nothing about failure modes (missing file vs. malformed document) or the shape of the explanation. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler, and the enumeration of what gets explained is placed before the scoping tip. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on the burden of describing the return content by listing the declared elements. Safety is covered by annotations. The remaining gap is error/edge-case behavior, which is minor for a read-only inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all three parameters (path, sheet, content) are already documented, including the path/content mutual exclusion. The description's only added semantic is that `sheet` narrows the output scope, which is marginal beyond the schema's own wording. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ("Describe") and resource ("what a document declares") and enumerates the concrete categories returned: sheets, input columns, computed rules, anchored scalars, derived precision, imports. This is far more than a restatement of the title. It does not, however, name or distinguish itself from siblings like visimark_ref or visimark_eval, which likely also surface document information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: inspecting a document's declared structure. The one directive offered ("Give `sheet` to narrow it") tells the agent how to scope output but not when this tool is preferable to visimark_ref, visimark_check, or visimark_eval. No exclusions or alternative-routing guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visimark_fmtPlan a VisiMark repairA
Read-only

Plan the edits that would bring a document's stored numbers back into agreement with its formulas, and name the generated artifacts it would write. Nothing is written: pass the returned plan to visimark_fmt_apply to land it. It repairs stale computed values, anchors and generated artifacts, and nothing else — never prose, never a column with no rule, never any other finding class. A document given as content has no directory, so imports and generated artifacts cannot be verified; skipped names the ones that were not checked.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoPath to a Markdown document on disk.
contentNoThe document as text, for a draft that is not on disk. Give this or `path`, never both.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description reinforces and extends this: it lists exactly what is repaired (stale computed values, anchors, generated artifacts), what is never touched (prose, columns with no rule, other finding classes), and the `skipped` behavior for content-only input. That is meaningful behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each load-bearing: purpose first, then the no-write/apply contract, then the repair scope, then the content-mode caveat. No filler and the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must carry return-value context, and it does name the key surfaces (the plan's generated artifacts and `skipped`). It stops short of describing the full shape of the returned plan, so a small gap remains for an agent that needs to inspect the plan programmatically.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the `path`/`content` contract is already documented; the description still adds semantic meaning by explaining the consequence of using `content` (no directory, so imports and generated artifacts cannot be verified) and what `skipped` then reports. That goes beyond the schema text, though it covers only one of the two parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb (plan the edits) and resource (a document's stored numbers vs. its formulas), and adds that it names the artifacts it would write. It explicitly routes to the sibling `visimark_fmt_apply` for applying the plan, so an agent can distinguish this from the write-side tool without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use and a next-step alternative ('pass the returned plan to `visimark_fmt_apply`'), plus the destructive-path caveat ('Nothing is written'). It does not distinguish this plan step from the other read-side siblings such as `visimark_check` or `visimark_ref`, which is the remaining gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visimark_fmt_applyApply a VisiMark repairA
Destructive

Land the plan visimark_fmt returned: splice the document's computed cells and anchors, and write the generated artifacts it named. Refused if the document changed since the plan was computed, and refused entirely unless the operator started the server with --allow-write and the host declared a root.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesThe document to write. There is nothing to apply to a string.
planYesThe result of `visimark_fmt`, unedited apart from dropping entries you rejected.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so safety direction is covered. The description goes further, disclosing a staleness guard (refused if the document changed since the plan was computed), an operator precondition (--allow-write), and a host requirement (declared root) — meaningful behavioral context beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, action front-loaded, with preconditions trailing. No filler; each clause either defines the operation or a condition that gates it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no output schema, the description supplies the operation, the plan-sourcing dependency, and both environmental preconditions an agent must satisfy before calling. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with two documented parameters, so the schema already carries the primary semantics. The description adds only indirect hints ('artifacts it named', dropping rejected entries), which is the expected baseline rather than extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names specific verbs (splice, write) and resources (computed cells, anchors, generated artifacts), and explicitly ties itself to the sibling visimark_fmt as the plan producer. An agent can distinguish this apply step from the fmt planning step without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the exact upstream relationship ('Land the plan visimark_fmt returned') and the two refusal conditions (document changed since plan, missing --allow-write or host root), making clear when the call will fail rather than succeed. This is explicit when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visimark_inferPropose VisiMark rules for a plain tableA
Read-only

Derive the formulas an existing Markdown table already obeys, and propose them as vmark rules. Run this before hand-authoring rules for a document that already has its numbers. Returns proposals only and never writes; visimark_infer_apply writes them.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoPath to a Markdown document on disk.
contentNoThe document as text, for a draft that is not on disk. Give this or `path`, never both.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is partially given. The description adds real value by confirming the output is 'proposals only and never writes' and that a separate tool performs the mutation, which clarifies the non-mutating contract beyond the hint. It stops short of describing proposal format or behavior on tables with no derivable formulas.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and benefit before the usage rule and the write-safety caveat. Every sentence carries distinct information; nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, annotation-covered tool with no output schema, the description covers purpose, usage, and the non-writing contract. Only edge-case behavior (e.g., a table with no derivable formulas, or error handling when neither path nor content is supplied) is left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself states the 'give this or `path`, never both' constraint, so the baseline is 3. The description adds no parameter-level syntax or format detail beyond what the schema documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource ('Derive the formulas an existing Markdown table already obeys') and the concrete artifact produced ('propose them as `vmark` rules'). It also distinguishes itself from the sibling visimark_infer_apply, so an agent can tell the propose step from the write step without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when: 'Run this before hand-authoring rules for a document that already has its numbers.' And it names the alternative with its selecting condition: proposals are returned here, while `visimark_infer_apply` writes them. The routing decision is fully determined by the text.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visimark_infer_applyApply proposed VisiMark rulesA
Destructive

Insert the vmark rules visimark_infer proposed. Pass the plan you were given, with any proposal you rejected removed. Refused if the document changed since the plan was computed, and refused entirely unless the operator started the server with --allow-write and the host declared a root.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesThe document to write. There is nothing to apply to a string.
planYesThe result of `visimark_infer`, unedited apart from dropping entries you rejected.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false and destructiveHint=true; the description adds real behavioral context: the write is gated on the operator starting the server with --allow-write and the host declaring a root, and it is refused as stale if the document changed since the plan was computed. That is exactly the extra information an agent needs before invoking a mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler: the action and its source come first, then the input-shaping rule, then the two refusal conditions. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter destructive write with no output schema, the description covers provenance, input shaping, and both failure gates. It does not say what happens on success (e.g., counts of applied rules or partial-application behavior), a minor gap given no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, including the note that plan must be unedited apart from dropping rejected entries. The description restates the plan handling but adds no syntax, format, or constraint detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ("Insert the vmark rules") plus the exact resource and its origin ("the rules visimark_infer proposed"), which separates it from the sibling visimark_infer that only proposes them. An agent can tell it is the write-side step of the infer/apply pair without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the required workflow (use the plan returned by visimark_infer, drop rejected entries) and the two refusal conditions (document changed, missing --allow-write/root). It does not contrast with visimark_fmt_apply, so the agent must infer which apply tool fits which artifact.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visimark_refVisiMark function referenceA
Read-only

Look up VisiMark's builtin functions: signature, arity, precision, errors and examples. Give name to look it up, or omit name for the whole reference. Reads no file.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoA builtin's name, e.g. `SUM`. Omit for the whole reference.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds the non-obvious trait 'Reads no file,' signaling this is a pure in-memory reference lookup rather than a file-consuming operation like eval/check, and it discloses the shape of the returned content. Useful context beyond the safety annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the returned content and followed by the invocation modes. No filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing what a lookup returns (signature, arity, precision, errors, examples), and the readOnly annotations cover the safety profile. The only gap is sibling differentiation, which is not strictly required for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single `name` parameter is fully documented as an example (`SUM`) with the omit behavior. The description restates 'omit `name` for the whole reference' without adding new semantics, so the schema is doing the heavy lifting and the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (look up) and resource (builtin functions) and enumerates what the entry contains: signature, arity, precision, errors and examples. It does not, however, differentiate itself from siblings like visimark_explain or visimark_check, which plausibly overlap with 'explain a function'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two invocation modes (pass `name` to look up one builtin, omit it for the whole reference), which is genuinely operative guidance. It never states when to prefer this tool over visimark_explain or visimark_check, so the routing decision is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedvisimark_check
    • First observedvisimark_eval
    • First observedvisimark_explain
    • First observedvisimark_fmt
    • First observedvisimark_fmt_apply
    • First observedvisimark_infer
    • First observedvisimark_infer_apply
    • First observedvisimark_ref

TDQS

A4.2/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a clearly distinct role: ref (lookup builtins), explain (describe declarations), eval (compute values), check (verify agreement), infer (propose rules), fmt (plan repairs), and their _apply counterparts. The plan-then-apply split is explicit and the read-only analyzers (check vs eval vs explain) are differentiated by their descriptions.

Naming Consistency4/5

All tools use a consistent visimark_ snake_case prefix with verb-based names, and the _apply suffix reliably marks write actions. Minor deviation: 'ref' and 'fmt' are abbreviations while others are full verbs, but the pattern remains predictable.

Tool Count5/5

Eight tools is well-scoped for a document-analysis/formula-verification server, with each tool earning its place. Read/plan/apply responsibilities are cleanly separated without redundancy.

Completeness5/5

The surface covers the full lifecycle: reference lookup, document description, evaluation, verification, inference, formatting, and the apply operations to land each plan, all with write-safety gating. No obvious gaps for the stated domain.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    A
    maintenance
    Local-first Markdown editor whose MCP server lets a coding agent and a human co-edit the same .md file — it opens and reveals files in the editor and reads or section-edits their contents. Tools: open_file, reveal, read_section, write_section, wait_for_change.
    63
    -