Skip to main content
Glama

Ferry is an MCP server that lets AI agents build animated, narrated presentations of codebase changes. The presentations open in your browser and update live while the agent builds them.

Ferry demo: an agent builds a deck live, then walks a sequence diagram and a stepped diff

▶ Watch the full demo (1:50): the agent builds the deck live, walks the behavior, architecture and code, then applies a review comment from the viewer.

Ferry carries reviewers from the old code to the new. It's built on a few ideas that make change explainers easy to follow:

  • Maximum stability. Code keeps its identity between steps. Only changed lines enter or leave. Lines that only change indentation glide sideways instead of being deleted and re-added.

  • Behavior before code. Sequence diagrams play the broken behavior, then replay the fix in the same slots. After that, a stepped diff shows the change that causes it.

  • Interruptible motion. Every channel is a closed-form spring. Retargeting keeps the current position and velocity, so pressing ← mid-transition reverses smoothly instead of jumping.

  • Plans, not pixels. Agents send declarative JSON. The server validates it, resolves git, highlights code with Shiki, and returns actionable warnings.

Install with your agent

Paste this into Claude Code, Codex, Cursor, or any agent that can run shell commands:

Install the Ferry MCP server for me. Clone https://github.com/danowicz/ferry-mcp into ~/.ferry-mcp (or pull if it's already there), then follow ~/.ferry-mcp/INSTALL.md: check Node ≥ 22.18, run npm install and npm test, and register Ferry with the MCP clients I use. Tell me when it's ready and what I need to restart.

INSTALL.md holds the steps the agent follows: prerequisites, verification, and config snippets for each client.

Related MCP server: MCP Walkthrough

Install by hand

git clone https://github.com/danowicz/ferry-mcp.git ~/.ferry-mcp
cd ~/.ferry-mcp && npm install   # Node ≥ 22.18

Claude Code

claude mcp add ferry --scope user -- node ~/.ferry-mcp/bin/ferry.js

Claude Desktop / Cursor / any MCP client (use the absolute path; clients don't expand ~)

{
  "mcpServers": {
    "ferry": { "command": "node", "args": ["/Users/you/.ferry-mcp/bin/ferry.js"] }
  }
}

Then ask your agent something like:

Explain the changes on this branch with a Ferry presentation.

The explain_changes prompt walks the agent through the whole flow. The viewer starts automatically at http://localhost:4747 (set FERRY_PORT to change it). Decks are stored in ~/.ferry/decks (set FERRY_HOME to change it).

Tools

Tool

What it does

authoring_guide

Slide types, JSON shapes, storytelling rules

inspect_changes

Changed files, churn, and numbered changes (blocks of added, removed, or re-indented lines)

create_deck / update_deck

Deck metadata and theme

add_slides / update_slide / remove_slides / reorder_slides

Edit the deck; warnings flag unmatched callouts, unknown ids, folded targets

get_deck / list_decks

Read back authoring JSON and outlines

delete_deck

Deletes a deck with its change plan and chat

draft_deck_from_git

Skeleton deck: title with stats, file map, one stepped diff per significant file

open_deck

Opens the live viewer

wait_for_feedback

Waits until you send a plan from the viewer, then returns it to implement (outline and slide JSON)

get_feedback

Returns what's pending without waiting

resolve_feedback

Closes a sent plan with a reply shown in the viewer

reply_feedback

Asks you a question in the viewer

export_deck

One self-contained HTML file (fonts and code inlined, works offline)

Review, plan, then implement

Ferry splits a review into three steps, all in the viewer:

  1. Review the change through the deck the agent built: behavior, architecture, then the code, step by step.

  2. Plan what should change next. Press C and switch the chat to Plan (or press Shift+Tab). Describe a change ("make the retry interval configurable", "add a test for the live-server guard") and Pin the code line it's about if you like. Claude reads the code and drafts the change as a plan deck: a second set of slides with the proposed code as hand-written diffs, the approach, and the tests to add. Nothing in the repository changes while you plan. Open the plan with the Plan button in the bar, keep chatting to revise it, and go back and forth until the slides say what you want.

  3. Implement: in the panel's Plan tab, press Send plan to agent.

    • If your agent (the one that built the deck) is waiting in wait_for_feedback, the panel says Agent is listening and the agent receives the plan: its outline and every slide's authoring JSON.

    • Otherwise Claude Code running in the viewer implements it. It edits the code in the deck's repository on the checked-out branch, never commits, and you follow along in the Chat tab. If it doesn't finish, the plan goes back so you can send it again.

    • The agent replies when it's done (Done or Declined). Reply under it to reopen. Review the result with git diff as usual.

Copy as prompt gives you a prompt for any other agent to implement the plan.

The chat

The chat is a conversation with a Claude Code agent dedicated to the deck. It runs headless (claude -p) in the deck's repository, using your existing Claude Code login.

  • Every message carries what you're looking at: the slide, the step, and anything you Pin (a code line, node, row…).

  • Ask mode answers questions and can edit the deck you're reviewing. Plan mode edits only the plan deck. Neither can touch repository files: only a sent plan does, and each mode can only write to its own deck.

  • Replies stream into the panel along with what the agent is doing ("Reading bind.ts", "Adding slides"), and slides update live.

  • The conversation continues across messages. New chat starts over, and Stop interrupts.

Set FERRY_CHAT_MODEL to pick a model (e.g. sonnet for faster replies), or FERRY_CLAUDE_BIN if claude isn't on your PATH. Chat logs are stored in ~/.ferry/chat/ and sent plans in ~/.ferry/feedback/. The viewer only accepts JSON requests from localhost pages, so other websites can't send instructions to your agent.

Slide types

title · section · points (cards, list, checklist) · files · diff (from git, before/after, hand-written lines, or plain code; with steps, focus, callouts, morph or review mode) · sequence (before/after phases) · flow (architecture with status changes and travelling packets) · metrics (rolling numbers) · compare (code, points, markdown, or a draggable image wipe for UI changes) · markdown.

Viewer keys

→ / Space: next step · ← back · ↑ ↓ slides · C chat & plan · P pin · O overview · N speaker notes · V voice narration (auto-advances) · T theme (midnight, tokyo, evergreen, paper) · S slow motion · R replay · F fullscreen · ? help

CLI

node bin/ferry.js serve [--open]        # standalone viewer
node bin/ferry.js list                  # saved decks
node bin/ferry.js delete <deck-id>…     # delete decks (or use the trash button on the home page)
node bin/ferry.js export <deck-id> [out.html]
npm run demo                            # builds the showcase deck through MCP
npm run record                          # re-records docs/demo.mp4 (needs ffmpeg and Playwright's Chromium)
npm run banner                          # re-renders docs/banner.png
npm test                                # end-to-end check of every tool

Layout

src/           MCP server (Node runs TypeScript directly)
  mcp.ts         tools, prompt, resource
  schema.ts      authoring input (zod) — what agents send
  compile.ts     authoring → render model: ids, git, highlighting, steps, warnings
  diff.ts        line diff with re-indent detection, change blocks, folding
  model.ts       render model shared with the viewer
  server.ts      viewer HTTP server + live updates (SSE) + feedback API
  feedback.ts    change requests, agent presence, change-plan formatting
  chat.ts        viewer chat: headless Claude Code per deck, streamed over SSE
viewer/src/    presentation app (bundled with esbuild on demand)
  motion.ts      closed-form interruptible springs
  feedback.ts    side panel: chat tab, change-plan tab, pinning
  markdown.ts    safe Markdown for chat replies
  slides/        one view per slide type
scripts/       demo, smoke test, screenshot walker, demo-video recorder, banner
docs/          banner, demo video and GIF

Available Tools

18 tools
add_slidesAdd slidesA

Append slides to a deck (or insert after a slide id). Returns the outline and any warnings (unmatched callouts, unknown ids). Changes appear live in an open viewer. See authoring_guide for slide types.

ParametersJSON Schema
NameRequiredDescriptionDefault
afterNoInsert after this slide id. Default: append.
slidesYes
deck_idYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavior: it returns the outline plus warnings (unmatched callouts, unknown ids) and that changes appear live in an open viewer. It does not address mutation semantics such as permission needs, idempotency, or whether appended slides can be undone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action and with no filler. The parenthetical warning examples add useful detail without bloating the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a large polymorphic schema, no annotations, and no output schema, the description covers the essentials: what it does, the return shape, the live-viewer side effect, and where to find slide-type detail. Permissions and mutation semantics are the remaining gap, but the tool is largely self-explanatory at the agent level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema description coverage is only 33%, with deck_id and slides undocumented at the top level, so the description must compensate. It clarifies the 'after' insert behavior and offloads the many slide types to authoring_guide, but deck_id and the overall slides contract get nothing beyond what the nested schema already shows.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Append slides to a deck') and covers the insert variant, which is enough to distinguish it from create_deck (new deck) and update_slide (single slide). It does not name any sibling explicitly, so it falls short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'append... or insert after a slide id', but there is no explicit when-to-use vs when-not guidance and no named alternative such as update_slide or reorder_slides. The pointer to authoring_guide is a helpful related-tool reference but not a selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

authoring_guideFerry authoring guideA
Read-only

Slide types, JSON shapes, and storytelling rules for explaining code changes. Read once before building a deck.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes the safety profile, so the description only needs to add context, which it does via the "read once" cue implying idempotent, cheap, cacheable content. It stops short of saying the output is prose documentation, but the gaps are minor given the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, and the content inventory is front-loaded before the usage instruction. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only reference tool this is essentially complete: what it contains, when to read it, and that it is safe. Only a hint about the return being prose/markdown is absent, which is a negligible omission here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate — the 0-param baseline applies. No syntax or format detail is required for an argument-free call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Slide types, JSON shapes, and storytelling rules for explaining code changes" names the exact content the tool delivers and is unmistakably distinct from the deck-manipulation siblings like create_deck and add_slides. The agent knows immediately this is reference material, not an action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Read once before building a deck" gives an explicit trigger and even a frequency constraint, telling the agent both when to call it and when not to call it again. No alternatives are needed since no sibling provides this documentation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_deckCreate deckA

Create an empty presentation deck. Returns its id and live viewer URL. Add content with add_slides.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoBranch, PR, or commit label, e.g. "PR #482".
repoNoRepository label shown in the footer, e.g. "acme/api".
themeNomidnight (warm dark, default), tokyo, evergreen, or paper (light).
titleYes
subtitleNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose two useful behaviors: the deck starts empty and the call returns an id plus a live viewer URL, which hints at sharing semantics. It says nothing about permissions, whether the viewer URL is public, or what happens if title collides, leaving notable gaps for a creation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then return values, then the follow-up tool. No filler; each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description adequately covers the return shape (id and live viewer URL). Combined with the 60% parameter documentation, an agent has enough to invoke it, though the absence of any permission or side-effect detail keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60% (ref, repo, theme documented; title/subtitle self-evident), and the description adds no parameter meaning whatsoever. The theme enum default is only explained in the schema. Baseline 3 is appropriate since coverage is not low enough to demand description-side compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create an empty presentation deck') and the 'empty' qualifier distinguishes it from content-bearing siblings like draft_deck_from_git. It does not explicitly differentiate itself from draft_deck_from_git, which also creates a deck, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Add content with add_slides' gives a clear next-step workflow cue, but there is no when-to-use/when-not guidance and no mention of the alternative draft_deck_from_git for repo-based deck creation. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_deckDelete deckA
Destructive

Permanently delete a deck, its change plan and its viewer chat. Only do this when the user asks: it cannot be undone.

ParametersJSON Schema
NameRequiredDescriptionDefault
deck_idYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the safety profile is covered. The description adds genuinely non-redundant context: the operation is permanent, irreversible ('cannot be undone'), and cascades to the change plan and viewer chat. It does not discuss permissions or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence carrying the action, the cascade, the usage gate, and the irreversibility warning. No filler, nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool, the description plus the destructiveHint annotation give an agent everything needed to call it safely. The only omission is any meaning for deck_id, which is trivial given the parameter name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single required parameter deck_id is never described in the text. However, with only one obviously-named identifier parameter there is little room for ambiguity, so this is adequate rather than deficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('delete') and resource ('deck'), and goes further by enumerating the cascade (change plan, viewer chat). This clearly distinguishes it from siblings like remove_slides or update_deck, which do not destroy the deck entity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit gating condition: 'Only do this when the user asks', which is strong usage guidance for an irreversible action. It does not name any alternative tool or describe a non-destructive path, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draft_deck_from_gitDraft deck from gitA

Create a skeleton deck from a git range: title with stats, a file map, and one stepped diff slide per significant file (one step per hunk). Then refine: add narration, callouts, behavior slides, and a review checklist with update_slide/add_slides.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNo
headNo
repoYesAbsolute path to the git repository.
themeNomidnight (warm dark, default), tokyo, evergreen, or paper (light).
titleNo
max_filesNoDiff slides to create, by churn (default 8).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantial work by disclosing exactly what gets generated, including the granularity rule 'one step per hunk'. It omits side-effect and return semantics (does it create a new deck, and what identifier is returned to feed update_slide/add_slides?), which matters for a tool with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the generation behavior and structured as generate-then-refine. Dense but every clause carries information; the parenthetical '(one step per hunk)' is a worthwhile precision detail rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with no annotations and no output schema, the description covers creation, artifact structure, and the follow-up workflow well. The notable gap is the return value/identifier needed to chain into update_slide and add_slides, which the absence of an output schema makes more important.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: base, head, and title are undocumented in the schema. The description compensates by framing the input as 'a git range' (implying base/head are range endpoints) and by defining significance via churn/significant files, which maps to max_files. It still leaves the title parameter's origin (auto vs. supplied) ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Create), resource (skeleton deck), and source (a git range), then enumerates the generated artifact: title with stats, file map, and stepped diff slides. This clearly differentiates it from the generic create_deck sibling by scoping it to git-derived content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly describes the two-phase workflow: first generate the skeleton, then 'refine: add narration, callouts, behavior slides, and a review checklist with update_slide/add_slides', naming the sibling tools to use next. It lacks an explicit when-not or prerequisite statement (e.g., must a deck already exist), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_deckExport deckB

Write the deck as one self-contained HTML file (viewer, fonts, and code inlined) to share or attach to a PR.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoAbsolute output path. Default ~/.ferry/exports/<id>.html
deck_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose meaningful output traits (fully inlined, self-contained, shareable), but says nothing about whether it overwrites an existing file, whether the deck must exist, or what permissions are needed for a write to disk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the action and the artifact, devoting its parenthetical to the only detail that matters for the output. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description is the only source of behavioral context, and it covers the nature of the produced file reasonably well but omits the default output location (left entirely to the schema), overwrite behavior, and any note that deck_id must reference an existing deck. Adequate but with noticeable gaps for a filesystem-writing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: 'path' is documented in the schema (including the default ~/.ferry/exports/<id>.html), but the required 'deck_id' has no description in either the schema or the tool description. The description adds no parameter meaning at all, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb (write) plus resource (the deck) and the artifact produced: one self-contained HTML file with viewer, fonts, and code inlined. That distinguishes it in practice from read-oriented siblings like get_deck or open_deck, though it never names an alternative to make the contrast explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The trailing phrase 'to share or attach to a PR' implies the scenario for exporting rather than viewing, which is useful implicit context. But there is no explicit when-to-use versus open_deck/get_deck, no prerequisites, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_deckGet deckB
Read-only

Return the deck outline and the authoring JSON of its slides (or one slide), e.g. to edit with update_slide.

ParametersJSON Schema
NameRequiredDescriptionDefault
deck_idYes
slide_idNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes this is a non-mutating read, and the description usefully discloses the return shape (outline plus authoring JSON). It says nothing about size limits, pagination, or behavior when slide_id is omitted versus supplied, so it adds only moderate context over the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the return payload and appends the use case, with no filler. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with no output schema, describing the returned outline and authoring JSON is the right core content. However, deck_id semantics and what happens when slide_id is absent or invalid remain unaddressed, leaving an agent to guess at invocation edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameter burden. The phrase "or one slide" hints that slide_id narrows the result to a single slide, but deck_id is never explained and the required/optional distinction is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (deck outline plus the authoring JSON of its slides or one slide), which clearly separates it from list_decks or inspect_changes. It stops short of naming a sibling tool explicitly for differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"e.g. to edit with update_slide" ties the read to the write sibling that consumes its output, which implies the read-before-edit workflow. There is no guidance on when not to use it (e.g. vs list_decks for discovery), so usage is only partially covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_feedbackGet feedbackA

Return the pending change plan left in the viewer (without waiting) and mark it as being worked on: code changes to make in the repository the deck explains. Use when the user says they left feedback.

ParametersJSON Schema
NameRequiredDescriptionDefault
deck_idYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the two most important traits: the call is non-blocking ('without waiting') and it has a side effect (marks the plan as being worked on). It omits what happens when no plan is pending and any permission requirements, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The trigger phrase is front-loaded and there is little filler, but the middle clause 'code changes to make in the repository the deck explains' is grammatically tangled and slows comprehension of the core meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-with-side-effect tool with no output schema and no annotations, the description covers the return concept and the mutation, but omits empty-result behavior and any auth or error conditions an agent would need for robust handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions deck_id or its format. The parameter name is self-explanatory and there is only one required param, so an agent can still call it correctly, but the description adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Return the pending change plan left in the viewer' plus the side effect of marking it 'as being worked on'. The parenthetical 'without waiting' implicitly separates it from the sibling wait_for_feedback, though it never names that sibling outright. Purpose is clear enough for correct selection among the feedback tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use when the user says they left feedback' gives a concrete trigger, and 'without waiting' signals the non-blocking distinction from wait_for_feedback. It stops short of naming resolve_feedback/reply_feedback as alternatives or stating when-not to use it, so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_changesInspect git changesA
Read-only

List changed files (status, churn) and their numbered changes (contiguous blocks of added/removed/re-indented lines) with previews. Change numbers match diff slides with a git source: assign them to steps with steps[].changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoBase ref (default HEAD). A branch name compares from its merge-base, like a PR.
headNoHead ref; omit to include uncommitted working-tree changes.
repoYesAbsolute path to the git repository.
pathsNoOnly show hunks for these files.
contextNoContext lines used for hunk numbering (default 3, same as diff slides).

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the safety profile is covered. The description goes further by disclosing the exact output shape (per-file status and churn, numbered changes with previews) and defining what a 'change' is, which is real behavioral context beyond the annotations. It omits pagination and any assumptions about repo state, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what it lists, and the second sentence covers the cross-tool referencing semantics. Dense but every clause earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns and does so adequately (files, status, churn, numbered changes, previews). It also explains how the change numbers interoperate with diff slides, which is the key integration detail. Minor gaps remain around edge cases and pagination.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (base, head, repo, paths, context) are already documented in the schema, including merge-base behavior and default context. The description adds no parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (changed files and their numbered changes), and unpacks the domain terms with parentheticals (status, churn, contiguous blocks of added/removed/re-indented lines). Clear on its own, but it never names or contrasts with any sibling such as draft_deck_from_git.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a genuine usage pointer for the output — change numbers match diff slides with a `git` source and can be assigned via steps[].changes — which implies this feeds the slide-authoring workflow. However, there is no explicit when-to-use or when-not-to-use versus sibling tools like draft_deck_from_git; the guidance is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_decksList decksA
Read-only

List saved decks, newest first.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered externally. The description adds one piece of behavioral context, the deterministic 'newest first' ordering, but says nothing about result size, pagination, or truncation behavior for a potentially unbounded list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource and the ordering front-loaded; no filler, no redundancy with the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description cannot be expected to enumerate return fields, but for a listing tool it omits whether results are paginated or bounded and what identifiers are returned for follow-up calls (get_deck, update_deck). Adequate for a trivial read tool but with visible gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is no parameter syntax the description could add or omit. No filtering or pagination parameters exist, so there is no gap for the description to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List saved decks') plus the sort order ('newest first'), which lets an agent distinguish it from get_deck (single retrieval) without opening a schema. It does not explicitly name any sibling, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'List saved decks' suggests an enumeration use case, but there is no explicit when-to-use, when-not, or routing to alternatives such as get_deck or inspect_changes. Nothing is misleading, but the agent must infer the selection context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_deckOpen deckB

Open the deck in the user's browser (live: later edits appear immediately). Returns the URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
slideNo1-based slide number to open at.
launchNoOpen a browser window (default true). False just returns the URL.
deck_idYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It usefully reveals that the view is live (later edits appear immediately) and that a URL is returned, but says nothing about whether launching a browser is a side effect with consequences, what happens if the deck doesn't exist, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with a compact parenthetical; the liveness note and return value are both front-loaded and no words are wasted. Not a 5 only because the parenthetical is doing a lot of unexplained work.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter tool with no output schema, the description covers purpose, the live-view semantic, and the return value. Remaining gaps (error behavior, deck existence requirements) are minor for a simple open-in-browser action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: slide and launch are documented in the schema and deck_id is not. The description's 'Returns the URL' ties loosely to the launch=false behavior already stated in the schema, so it adds only marginal semantic value — baseline 3 is right.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Open') and resource ('the deck') plus the destination ('in the user's browser'), which cleanly distinguishes it from read-only siblings like get_deck or list_decks. It does not explicitly name a sibling for differentiation, keeping it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Nothing states when to prefer this over get_deck (which presumably returns data without launching a browser) or over export_deck. The 'live' parenthetical hints at the use case but the choice between this and its closest siblings is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_slidesRemove slidesC

Remove slides by id.

ParametersJSON Schema
NameRequiredDescriptionDefault
deck_idYes
slide_idsYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. 'Remove' implies a destructive mutation, but it says nothing about irreversibility, required permissions, side effects, or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and free of filler, which is structurally clean. However, it is too sparse to be considered appropriately sized for a destructive tool with required parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and two required parameters with zero schema coverage. The description omits prerequisites, mutation implications, and return behavior, leaving substantial gaps for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only mentions 'by id' and does not clarify the roles of deck_id or slide_ids, their formats, or the minimum item requirement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Remove) and resource (slides), with scope limited to identifiers. It is clearly distinct from siblings like add_slides and delete_deck, though it doesn't explicitly name a sibling or mention the deck context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as delete_deck or update_slide. The phrase 'by id' hints at a prerequisite but provides no explicit usage conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reorder_slidesReorder slidesC

Set the slide order. Ids you omit keep their relative order after the listed ones.

ParametersJSON Schema
NameRequiredDescriptionDefault
orderYes
deck_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the useful partial-order semantic (omitted IDs keep relative order and go after the listed ones), but says nothing about permissions, whether unlisted slides are preserved, or failure behavior for invalid IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the core purpose front-loaded before the ordering rule. Nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and 0% schema description coverage, the description is thin: it omits deck_id semantics, error handling, and any return/confirmation behavior. It covers only the ordering edge case.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, so the description must compensate. It meaningfully clarifies the `order` array as a partial ordering where omitted IDs are appended, but `deck_id` receives no explanation at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Set the slide order'), which clearly identifies the operation. It does not explicitly differentiate itself from siblings like add_slides or remove_slides, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives or the contrast with add_slides/remove_slides. The agent must infer from the name alone that this is a reordering/all-mutation operation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reply_feedbackReply in feedback chatA

Post a message in the viewer chat without closing anything: ask a clarifying question about a request (by id), or say something about the whole deck (no id). The user answers in the viewer; wait_for_feedback returns their answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoFeedback id to answer; omit for a deck-wide message.
textYes
deck_idYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the message is posted in the viewer chat, that it does not close or resolve anything, and that the exchange is asynchronous ('the user answers in the viewer; wait_for_feedback returns their answer'). It omits auth/permission requirements and whether messages persist or can be edited.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and its non-destructive nature, then the two modes. Dense but every clause carries information; only slight density cost keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the two invocation modes and the downstream answer loop, which is what an agent needs to call it correctly. It lacks error/edge-case guidance (e.g., invalid id, mismatched deck_id), leaving minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (deck_id and text are undocumented), so the description is expected to compensate. It does clarify the id/no-id branching, but that behavior is already stated in the schema's id description, and it adds nothing for deck_id or text beyond their obvious meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Post a message in the viewer chat') and immediately delineates two distinct modes of use (with an id vs. without). The phrase 'without closing anything' implicitly contrasts it with resolve_feedback, so an agent can separate it from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly says when to include an id (asking a clarifying question about a request) and when to omit it (deck-wide message), and names the follow-up path via wait_for_feedback. It stops short of explicitly stating when to use resolve_feedback instead, so it lacks full alternative coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_feedbackResolve feedbackB

Close change requests after making the code changes (or declining them). Each reply appears in the viewer next to the request.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
deck_idYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses one useful behavioral trait — the reply is surfaced in the viewer next to the request — and that closing is a terminal action, but says nothing about permissions, reversibility, or what happens when only some items succeed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the visible side effect. Nothing is padded, though the second sentence is a side note rather than core invocation guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the intent and one side effect but omits the result shape, partial-failure behavior, and whether a batch must be all-or-nothing. Adequate but incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema description coverage is reported as 0%, so the description should compensate; it does partly by explaining that a reply records 'what changed, or why not' and that declines are a variant of closing. It adds no detail on deck_id or batching semantics beyond the schema's own nested field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb ('close') and resource ('change requests') plus the two possible dispositions (changes made, or declined). It is distinguishable from siblings like get_feedback and reply_feedback, though it never names those alternatives explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'after making the code changes (or declining them)' implies the moment to call it, i.e. once work on a request is finished. There is no explicit when-not guidance or reference to the sibling tools that read or reply to feedback.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_deckUpdate deckC

Change deck metadata or theme.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNo
repoNo
themeNomidnight (warm dark, default), tokyo, evergreen, or paper (light).
titleNo
deck_idYes
subtitleNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. "Change" implies mutation, but the description says nothing about permission requirements, whether changes are reversible, how omitted fields are handled, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is structurally clean. It is arguably too terse for a six-parameter mutation tool, so brevity here is under-specification rather than effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, no annotations, no output schema, and near-zero schema coverage, the description should compensate for the missing structured context but does not. An agent cannot determine field semantics or behavioral constraints from what is written.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only the theme enum is documented). The description loosely implies theme and "metadata," but adds no meaning for title, subtitle, ref, repo, or the required deck_id, leaving five of six parameters unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ("Change") and resource ("deck"), and hints at two attribute categories (metadata, theme). However, "metadata" is vague and does not tell the agent which fields (title, subtitle, ref, repo) are actually in scope, so the purpose is only partially specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use update_deck versus siblings like update_slide, create_deck, or get_deck, and no preconditions are stated. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_slideUpdate slideA

Replace one slide with a new definition (send the complete slide). Keeps its id and position.

ParametersJSON Schema
NameRequiredDescriptionDefault
slideYesThe complete slide, in the same shape as an add_slides item (see authoring_guide).
deck_idYes
slide_idYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose the key trait: this is a full replacement, not a patch, and the slide keeps its id and position. It says nothing about required permissions, failure modes when slide_id is unknown, or whether the deck must be unlocked, which matters for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no waste, with the replace semantics and the caller obligation front-loaded before the retention guarantee. Nothing here is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-param mutation with no annotations, no output schema and a free-form nested object, the description covers the destructive replace contract but omits permissions, error behavior, and the meaning of deck_id/slide_id. Adequate to invoke correctly, incomplete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: slide is documented as 'the complete slide, in the same shape as an add_slides item', while deck_id and slide_id are bare strings with no meaning given in either schema or description. The description reinforces the all-or-nothing semantics of slide, which is genuinely useful for a free-form nested object, but it does not compensate for the two undocumented identifiers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Replace one slide') plus the replacement scope ('with a new definition'), which lets an agent separate it from add_slides, remove_slides and reorder_slides without opening their schemas. It never names an alternative sibling explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(send the complete slide)' implies when this tool applies (you already have a slide and want to change it) and how it must be called, but there is no explicit when-not guidance or named alternative for partial edits. Usage is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_feedbackWait for feedbackA

Wait until the user sends a change plan from the viewer's feedback panel (C key), then return it: code changes they want in the repository the deck explains, each with the slide, step and pinned element (often a code line) it was written on, plus the slide's authoring JSON. Returns at once if requests are already pending. The viewer shows the user that you are listening. Make the changes in the code (not the slides), resolve_feedback, then call this again to keep reviewing together.

ParametersJSON Schema
NameRequiredDescriptionDefault
deck_idYes
timeout_secondsNoHow long to wait (default 300). On timeout, tell the user how to send feedback, or wait again.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the blocking wait semantics, the immediate-return case, the fact the viewer signals to the user that the agent is listening, and the shape of the returned feedback. Timeout behavior is delegated to the schema, and no auth or failure-mode detail is given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical behavior (blocking wait, immediate return) is front-loaded, and every sentence carries information. The middle sentence enumerating return fields is dense and slightly meandering, but nothing is truly wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description explains what is returned, and the workflow loop with resolve_feedback is spelled out. It is nearly complete for a 2-parameter wait tool, with only the undocumented deck_id leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: timeout_seconds is documented in the schema, but deck_id has no description anywhere. The prose mentions neither parameter, so it adds no meaning beyond the schema and only partially compensates for the undocumented deck_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific blocking verb (wait), the exact trigger (a change plan sent from the viewer's feedback panel), and what comes back (code changes with slide/step/pinned element plus authoring JSON). It is clearly distinguishable from the sibling read tools get_feedback and inspect_changes because it describes a long-poll that returns pending feedback.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operating context: call it to receive change plans, act on the code rather than the slides, then call resolve_feedback and loop back to this tool to keep reviewing together. It also notes it returns at once if requests are already pending. It stops short of stating when not to use it (e.g. versus get_feedback for a non-blocking poll).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 18 tool updatesv0.1.0
    • First observedadd_slides
    • First observedauthoring_guide
    • First observedcreate_deck
    • First observeddelete_deck
    • First observeddraft_deck_from_git
    • First observedexport_deck
    • First observedget_deck
    • First observedget_feedback
    • First observedinspect_changes
    • First observedlist_decks
    • First observedopen_deck
    • First observedremove_slides
    • First observedreorder_slides
    • First observedreply_feedback
    • First observedresolve_feedback
    • First observedupdate_deck
    • First observedupdate_slide
    • First observedwait_for_feedback

TDQS

A3.5/5.0

Scored across 18 tools

Disambiguation5/5

Each tool targets a distinct resource or action, from deck and slide CRUD to git inspection and viewer feedback. The only subtle overlap is between get_feedback and wait_for_feedback, but their descriptions clearly distinguish immediate retrieval from blocking wait, and resolve_feedback vs reply_feedback are similarly well separated.

Naming Consistency4/5

All operational tools follow a consistent snake_case verb_noun pattern (e.g., create_deck, add_slides, update_slide, wait_for_feedback). The lone deviation is authoring_guide, which is a noun phrase rather than an action, but the overall convention is predictable.

Tool Count4/5

18 tools is slightly above the ideal 3-15 range, but the server's domain is genuinely broad: deck lifecycle, slide editing, git diff import, viewer control, and feedback handling. Each tool earns its place, though a few could conceivably be consolidated.

Completeness5/5

The surface covers full CRUD for decks and slides, list/open/export operations, git change inspection and skeleton deck generation, plus a complete feedback loop (wait, get, resolve, reply). No obvious dead ends exist for the stated purpose of building and refining code-explanation decks.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to present interactive code walkthroughs with voice narration, opening files, highlighting code, and showing inline explanations with synchronized text-to-speech.
    5
    16 npm
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Turn your AI coding agent into a producer of interactive, narrated walkthroughs — code, whiteboard, and 3D casts, each a single self-contained HTML file that opens in any browser. Runs locally over npx.
    1
    -
  • A
    license
    A
    quality
    B
    maintenance
    Enables LLMs to create animated, narrated tours of GitHub pull requests, with a local browser viewer showing diffs and code highlights while speaking the narration aloud.
    7
    MIT