Skip to main content
Glama

diagonaldiagrams

Diagrams and charts your agent can draw, and check before it shows you. One command takes a short spec, Mermaid, or a CSV, lays it out, draws an SVG, and audits the drawing for labels on labels, arrows through boxes, and text that spills.

Try it in the browser: the real engine, running in your browser.

An AI agent platform with framed groups and icons, drawn and audited by diagonaldiagrams

Install

Claude Code, as a plugin. This adds the drawing skill:

/plugin marketplace add lucas19919/diagonaldiagrams
/plugin install diagonaldiagrams@diagonaldiagrams

Any agent, as a command. This puts diagonaldiagrams on PATH:

pip install git+https://github.com/lucas19919/diagonaldiagrams

As MCP tools (render, describe, validate, audit, export, infer):

claude mcp add --scope user diagonaldiagrams -- diagonaldiagrams mcp

Or clone it and run py -3 graph.py describe (Windows) or python3 graph.py describe. Agents can follow .agents/skills/diagonaldiagrams-install/SKILL.md to install, and .agents/skills/graph-engine/SKILL.md to draw.

Related MCP server: excalidraw-mcp

One command

diagonaldiagrams mermaid flow.mmd -o out/flow.svg
{"ok": true, "svg": "out/flow.svg", "html": "out/flow.html", "receipt": "out/flow.graph.json", "items": 9,
 "layout": ["row 1: Your agent describes the diagram", "row 2: Check the description", "row 3: Makes sense? (decision)", ...]}

It checks the input, draws, audits the drawing, and prints one line of JSON. ok: false comes with errors, each with a fix; change the input and run it again. layout describes the drawn structure, so an agent can confirm it without looking at a picture.

The audit reads the finished SVG: labels on labels, labels too wide for their box or off the canvas, boxes on boxes, arrows through a box or across a label, two arrows on one line, and nodes inside a frame they don't belong to. When it finds a problem it can fix, the layout repairs itself. It reads any SVG, so diagonaldiagrams audit also checks a figure an agent wrote by hand.

Input

Command

Input

Draws

mermaid

Mermaid flowchart, sequenceDiagram, or erDiagram

Picks flow, seq, or schema

flow

JSON

A decision flow, with frames (groups) and icons

arch

JSON

Services and the calls between them

seq

JSON

A sequence of messages

schema

JSON

Tables and their relations

bar, line, scatter

CSV

Comparisons, trends, relationships

geo

CSV or GeoJSON

Places on a built-in coastline

chart <type>

CSV or JSON

About 60 more: sankey, gantt, heatmap, treemap, box, radar, math, ...

describe <type> prints the contract for any type. Long names wrap inside their boxes. Loops are drawn back up the outside. Built-in icons (describe icons) cover users, databases, servers, queues, and more; in Mermaid, write fa:fa-database.

Output

File

Contents

OUT.svg

The drawing. Title and subtitle are on the figure; points and bars have hover tooltips.

OUT.html

The same figure, plus Copy SVG, Export SVG, and Export PNG.

OUT.graph.json

The receipt. export redraws from it.

export OUT --to png|pdf|svg|drawio|excalidraw writes other formats. draw.io and Excalidraw exports keep boxes, frames, and the routed arrows attached to their boxes, so a person can open the figure and drag things around. PNG and PDF use cairosvg if installed, else headless Edge or Chrome.

Measured

Five ways for an agent to draw the same flowchart, at 5, 11, 20 and 40 steps, two runs each, each run a fresh Claude Sonnet agent.

Tokens per flowchart for five approaches: diagonaldiagrams stays between 1.3k and 2.1k, the others climb to 8k to 28k at 40 steps

Tokens per flowchart, mean of two runs:

Steps

diagonaldiagrams

Hand-written SVG

Graphviz

Mermaid CLI

D2

5

1.3k

0.9k

2.5k

2.7k

2.5k

11

1.5k

2.0k

5.7k

5.6k

5.3k

20

1.5k

5.7k

4.4k

3.2k

9.5k

40

2.1k

8.1k

12.5k

20.5k

28.0k

All eight runs per approach:

Tokens

Calls

Screenshots

Time

diagonaldiagrams

12.7k

29

0

6 min

Hand-written SVG

33.4k

51

2

26 min

Graphviz

50.1k

91

4

28 min

Mermaid CLI

63.9k

115

42

31 min

D2

90.5k

170

23

38 min

What the runs show:

  • Every approach produced a figure without overlapping text; the difference is what it cost to get there. With another tool the agent renders a PNG and looks, or writes its own geometry checks, and that cost grows with the figure. With diagonaldiagrams it reads the verdict and the outline, and took no screenshots in any run.

  • Hand-written SVG is cheapest at 5 steps, because those agents mostly never looked at their output (see below).

  • Where others do better: at 20 and 40 steps diagonaldiagrams draws one tall column. The Mermaid agents spent much of their extra tokens rearranging 40 steps into side-by-side stages, which reads better.

How it was run:

  • Every run got the same task: a process described in prose, the number of steps, "clean, readable, no overlapping text", save figure.svg, and reply with the path and any failed commands. Only the first lines differ: which tool, and the command to run it.

  • Tools: mermaid-cli 12 using the system Edge; Graphviz 16.1 as the official WebAssembly build; D2 0.7 from its official npm package (dagre layout); diagonaldiagrams with its skill; and no tool at all.

  • Counted: the run's own tokens, meaning what the agent wrote, what came back to it, and screenshots at Claude's image rate. Left out: the harness (system prompt, tool definitions, reminders), which is the same whatever the task; the final report; and thinking, which transcripts hide. scripts/bench_tokens.py recounts any run from its transcript.

  • Caveats. Two runs per point, and single runs differ by up to 2×. The no-tool prompt said not to use any renderer, and most agents took that to mean they should not look either. Graphviz and D2, as installed, wrote SVG only and there was no image viewer, so several agents spent calls looking for one. Twenty agents ran at once and shared one browser pane. All agents were Claude Sonnet.

Every figure was drawn from a file in examples/ and passed its own audit. Rebuild them with py -3 scripts/build_site.py.

How a figure is made

Checkout platform with frames

Checkout request, a sequence diagram

Shop data model

Signups line chart

Tests

py -3 tests_smoke.py

Exit 0 means every check passed.

License

MIT

Available Tools

6 tools
auditC

Check any SVG for collisions, clipping, and arrows through boxes.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the three issue categories it detects but says nothing about whether the operation is read-only, whether it reports or mutates, what the output looks like, or any performance/limit characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no waste; the scope of checks is stated immediately. It is efficient, though it is arguably too terse given the surrounding gaps rather than genuinely concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description should explain what the audit returns (report, pass/fail, issue list) and how the path argument is supplied. Neither is covered, so an agent cannot fully predict invocation or result shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter 'path' is never mentioned in the description. 'Check any SVG' loosely implies the input is an SVG, but it does not clarify that the parameter is a filesystem path (string) rather than SVG markup, leaving the required argument's format ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check') and resource (SVG) with three concrete check categories (collisions, clipping, arrows through boxes), so an agent knows exactly what the tool inspects. However, it does not differentiate itself from the closely related sibling 'validate', leaving overlap unresolved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus the sibling 'validate', nor any prerequisites or exclusions. The agent must infer that 'audit' is broader/diagnostic and 'validate' is rule-based, with nothing in the text to support that inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describeC

The spec shape, flags, and budgets for one figure type. No type: the list of types.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the useful 'no type → list of types' branching behavior, but says nothing about read-only status, auth/permission needs, return structure, or failure modes for a tool whose whole purpose is introspection.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary content description, so nothing is padded. However, the brevity tips into under-specification rather than crispness — the fragments are terse to the point of ambiguity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 1-param introspection tool with no annotations and no output schema, the description should explain what the returned spec shape/flags/budgets actually look like, since nothing else documents the return value. It hints at return content but leaves the agent without enough to call or interpret confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It does add real meaning for the single optional 'type' param by defining the no-type fallback (returns the list of types), which the bare string schema does not convey. It still omits what valid type values look like and the format of the per-type output.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gestures at content ('spec shape, flags, and budgets') but never states a clear verb — it reads as a noun phrase, leaving the agent to infer that the tool 'describes' things. The resource ('one figure type') is identifiable and distinct from siblings like render/validate/audit, but the domain jargon ('spec shape') keeps the purpose somewhat opaque.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage signal is the conditional 'No type: the list of types,' which explains one calling mode but not when to choose describe over render, validate, audit, export, or infer. No context, prerequisites, or exclusions are offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exportC

Write png, pdf, or svg from a rendered figure's receipt.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNo
outNo
pathYes
layoutNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. "Write" implies a filesystem mutation, but the description never says whether existing files are overwritten, whether directories are created, what permissions or paths are required, or what the tool returns. Only the output format list is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the output formats front-loaded and no filler. It is efficiently sized, though its brevity reflects under-specification rather than disciplined editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations, no output schema, and zero parameter documentation, the description is too thin. An agent lacks the information needed to supply "path", "out", and "layout" correctly or to anticipate side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description must compensate and largely does not. It echoes the png/pdf/svg values already encoded in the schema's enum and says nothing about "path" (required), "out", or "layout".

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb ("Write") and the output formats (png, pdf, svg) sourced from a rendered figure's "receipt", so the general intent is graspable. However, "receipt" is unexplained domain jargon and the description does not clearly distinguish this tool from the sibling "render", which an agent might reasonably confuse with producing figure output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the siblings (render, describe, validate, audit, infer), nor any prerequisite such as needing a prior render step. The phrase "from a rendered figure's receipt" hints at an ordering dependency but never states it explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inferC

Propose a scatter spec from a CSV.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataNo
pathNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it discloses nothing about behavior: whether it is read-only, whether it mutates state, what it returns, or whether it needs a file path versus inline data. 'Propose' hints at a non-destructive suggestion but is never made explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words, which is structurally sound. However, it is under-specified rather than genuinely concise, so brevity here costs more than it saves.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter coverage, the description is the only information source and it is one sentence. It leaves the agent without enough to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters (data, path), so the description must compensate but does not. It only vaguely references a CSV input and never clarifies the distinction between supplying inline data versus a path, nor which is expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (propose) and resource (scatter spec) with the input source (CSV), so the agent can tell what it produces. It does not, however, differentiate itself from siblings like render or describe, leaving overlap ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives. With siblings such as render, describe, and validate, the agent has no signal about which tool to reach for in a given situation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

renderB

Validate, draw, and audit one figure. Writes OUT.svg, OUT.html, and a receipt. ok false: read errors, change the spec or options, call again.

ParametersJSON Schema
NameRequiredDescriptionDefault
outYesoutput path ending .html
dataNoCSV text, inline
pathNoor a spec/CSV file on disk
specNoJSON spec, inline
typeYesmermaid (data = Mermaid text), flow, arch, seq, schema, bar, line, scatter, geo, or any describe type
optionsNoCLI flags without dashes, e.g. {"x": "month", "y": "signups", "group": "product", "title": "...", "highlight": ["Atlas"]}. describe lists them.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses side effects (writes OUT.svg, OUT.html, and a receipt) and a failure mode (ok false indicates read errors), but says nothing about overwriting existing files, permissions, or the precedence between inline spec/data and path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences with no padding; outputs are named early. The telegraphic 'ok false' phrasing is terse to the point of being cryptic, costing a little clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param tool with nested objects and no output schema, the description covers the write side effects and a failure signal, which is the right scope. It still omits how the 'ok'/receipt is surfaced to the caller and how spec vs path precedence works, leaving gaps an agent would hit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters, including the 'type' enum values and options shape. The description adds no parameter meaning beyond mentioning 'spec or options' in its error note, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states concrete actions (validate, draw, audit) and names the concrete artifacts produced (OUT.svg, OUT.html, a receipt), so an agent knows this renders a figure to disk. However it gives no differentiation from the siblings validate and audit, which it claims to also perform, leaving overlap unresolved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is a hint of recovery guidance ('ok false: ... call again'), but nothing states when to use render versus describe, export, validate, or audit. An agent gets no routing signal for choosing this tool over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validateB

Check a spec or CSV without drawing.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataNoCSV text, inline
pathNoor a spec/CSV file on disk
specNoJSON spec, inline
typeYes
optionsNoCLI flags without dashes, e.g. {"x": "month", "y": "signups", "group": "product", "title": "...", "highlight": ["Atlas"]}. describe lists them.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses only the absence of rendering; it says nothing about what 'check' means (schema validity? CSV parse errors?), what the result looks like, or whether it can fail silently. For a tool with zero annotation coverage this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste. It is arguably over-compressed for a five-parameter tool, but the structure itself is efficient and earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five parameters (one nested options object), one required field, no annotations, and no output schema, the description should explain what validation means and what failure returns. A one-line sentence leaves an agent unable to predict the outcome of a call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is ~80%, so the schema already documents data, path, spec, and options. The description adds only 'spec or CSV,' which maps loosely to three of the five parameters and gives no format, precedence (data vs path vs spec), or meaning for the undocumented 'type' parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb (check/validate) and the resources it operates on (a spec or CSV), and the phrase 'without drawing' implicitly distinguishes it from the sibling render tool. The name 'validate' alone is generic, but the description supplies the missing object of the verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Without drawing' hints at when to prefer this over render (dry-run validation), but no explicit when-to-use or when-not-to-use guidance is given. The alternative (render) is only implied, never named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedaudit
    • First observeddescribe
    • First observedexport
    • First observedinfer
    • First observedrender
    • First observedvalidate

TDQS

B3.2/5.0

Scored across 6 tools

Disambiguation4/5

Each tool has a distinct role: describe (spec shape), render (draw), validate (dry-run check), audit (SVG QA), export (format conversion), infer (CSV to spec). The main overlap is between validate and render (which internally validates) and between audit and render (which internally audits), but descriptions make the boundaries reasonably clear.

Naming Consistency5/5

All six tools use a consistent single lowercase verb convention (describe, render, validate, audit, export, infer). No mixed casing, no noun_verb hybrids, no stylistic deviations.

Tool Count5/5

Six tools is well-scoped for a diagram spec/render/export pipeline. Each tool covers a distinct stage (introspect, author, check, QA, convert, generate) and none feels redundant or bolted on.

Completeness4/5

The surface covers a full lifecycle: discover spec shape, infer from CSV, validate, render, audit, and export to multiple formats. Minor gaps like batch rendering or listing prior outputs exist, but core workflows have no dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers