maple
OfficialThis server exposes Maple review-comment tools for an agent working on a branch.
List comments:
list_commentsreturns a branch's review comments newest first and can filter byopen,resolved,needs_reverify, ororphaned.Wait for comments:
wait_for_commentsblocks until a new comment arrives or the timeout elapses, with an optional cursor to only return newer comments.Resolve comments:
resolve_commentmarks a comment resolved with the commit SHA that addressed it and an optional note.Get context:
get_comment_contextreturns the anchor, viewport, surrounding markup, replay links, and mock state needed to act on a comment.Start solo mode:
start_solostarts a localhost bridge paired to one deployed preview; comments are stored locally under.maple/and never count toward the merge gate.
Provides a merge gate and reviewer identity on GitHub: reviewers sign in through GitHub Device Flow so every comment is written as its author (no secret in the preview), and a maple/visual-review GitHub check fails while any visual comment is unresolved—optionally until a reviewer approves—holding the pull request merge until reviews are resolved.
A reviewer points at something on a preview deployment and says what is wrong. Maple captures where they pointed, what they were looking at and who they are, hands it to a coding agent in a form it can act on, and holds the merge until every comment is resolved.
Status: 0.x. All eight packages are published, with provenance, through a trusted publisher. 0.x makes no compatibility promise: a public interface is broken when breaking it is the right shape, and the changeset says what broke.
How it works
Mount: One route in your application, and one script in the preview build.
Comment: A reviewer points to an issue on the app. Maple records all the context needed for the agent to pick it up.
Fix: Your agent monitors new comments via the MCP, implements a fix and marks it as resolved.
Gate: A CI check holds the merge until all comments are resolved, and all visual gates pass.
Code got fast. Planning, definitions of done and edge cases did not, so they get skipped and surface in testing. Maple moves that review to the preview, where the comment can still be acted on.
Related MCP server: loupe-mcp
What it does
Each package's README shows the rest: the CLI, the design lint, flags and roles in a mock and the connectors.
Why another one
Pincushion, Vercel Toolbar, Chromatic, BugHerd and Marker.io each do some of this. None combines all four of:
Open source, Apache-2.0, with no hosted service required.
Deployed previews. Maple runs on the preview URL your CI already builds, so anyone with the link can comment — a designer, a product manager, a client. Tools like Agentation run against
localhost, which means the only person who can leave a comment is the person running the build.A merge gate — CI blocks while a visual comment is unresolved, and optionally until somebody says they looked. See
docs/gate.md.An agent loop — the agent reads comments, fixes, and resolves them.
Pick a setup
The SDK route has no default store: without one, its comment endpoints answer 404. Smallest setup first:
Add it to an app
The quickest wiring, for trying Maple on your own machine. Sharing a preview needs a store per reviewer: see Pick a setup.
npm install @maple-kit/core @maple-kit/ui @babel/core@babel/core is an optional peer, and the tagger is what needs it: leave it out
and tagger: true has nothing to transform with.
Mount the route and the tagger from the build, then render the overlay:
// vite.config.ts
import { createCommentStore } from "@maple-kit/core";
import { githubStore } from "@maple-kit/core/connectors";
import { maple } from "@maple-kit/core/vite";
const preview = process.env.MAPLE_PREVIEW === "1";
export default defineConfig({
plugins: [
react(),
maple({
tagger: preview,
route: {
// Local trial only: every comment is written as this one token.
// The route has no default store; without one its comment endpoints 404.
store: createCommentStore(
githubStore({ owner: "acme", repo: "web", token: process.env.GITHUB_TOKEN! }),
),
},
}),
],
});// main.tsx
import { Maple } from "@maple-kit/ui/maple";
createRoot(root).render(
<>
<App />
<Maple branch={import.meta.env.VITE_MAPLE_BRANCH} />
</>,
);The shared GITHUB_TOKEN is for a local trial only. A preview other people
review builds the store per request from each reviewer's own token, through the
resolver in docs/github-auth.md, so a comment is
authored by whoever wrote it and no GitHub secret sits in the preview.
Next.js uses withMaple from @maple-kit/core/next and a catch-all route at
app/api/maple/[...maple]/route.ts; examples/next-app
has both. The setup-maple-org skill walks through the GitHub App a reviewer
signs in with.
Use with Claude Code
The Maple plugin bundles the skills that set Maple up and act on its comments, the MCP server that reads and resolves them, and the Stop hook that keeps an agent working while they are open:
/plugin marketplace add maple-kit/maple
/plugin install maple@maple-kitThe server and the hook read GITHUB_TOKEN, MAPLE_GITHUB_OWNER and
MAPLE_GITHUB_REPO from the environment Claude Code starts in. In a project
where the last two are unset, the hook lets every stop through.
To install only the skills, for Claude Code or any other agent the
skills CLI supports, run npx skills add maple-kit/maple.
Vendor-agnostic by construction
Maple stores nothing itself. A connector is one file implementing plain Promise-returning methods:
import type { StoreConnector } from "@maple-kit/core/connectors";
export function myStore(options: MyOptions): StoreConnector {
return {
name: "my-store",
async list(query) {
/* … */
},
async append(comment) {
/* … */
},
};
}There are six kinds — store, media, observability, identity, gate and
classifier — and a connector's capabilities are exactly the methods it defines.
Run maple connectors to print the matrix from the code, or see
docs/connectors.md.
Packages
Package | What it is |
Server SDK, overlay controller, connector contracts, build plugins. | |
The marks, the island and the composer: the overlay a reviewer uses. | |
Hooks over the controller, for an overlay in your own design system. | |
The MCP server an agent talks to, and a Stop hook. | |
The | |
Rewrites a page's API responses, flags and role into a named state. | |
Scores a comment as it is written, and plans a mock from a sentence. | |
Design-system rules read off the page the browser laid out. |
Documentation
maple review: the overlay on a running app, nothing wired inConfiguration: every environment variable, and which are secrets
Examples: Vite and Next: real applications Maple mounts into
The CLI:
review,solo,setup,connectors,mock plan
Solo mode: a guest's comments on their own machine
Drafts and publishing: why a comment is unsent until it is not
The JSX tagger: how a comment becomes
file:lineAnchoring a region: how a dragged box finds its content again
Screenshots: the picture taken at pick time
Replies: decided, not built
The agent loop: the MCP tools and the Stop hook
@maple-kit/mcp: the server and the hooksetup-maple-agent-loop: connect and verify an agentmaple-review: turn a pull request's comments into a worklist
The merge gate: what blocks a merge, what approving does, which App to pin
Maple Mock: a model picks the state, code writes every byte
@maple-kit/mock: flags, roles and the runtimeThe assist tier: what a score is, and what it is never allowed to be
@maple-kit/classifier: scoring and mock planningDesign lint: the rendered rules and the tiers around them
@maple-kit/lint: the rules as a package
@maple-kit/core: server SDK, overlay controller, connector contracts@maple-kit/ui: the overlay a reviewer uses@maple-kit/react: hooks for an overlay in your own design systemcontribute-connector: scaffold a connector and run the contract suite
Deploying Maple in an organisation: trust boundaries and a hardening checklist
GitHub authentication: Device Flow, the two Apps, the cookie, revocation
setup-maple-org: register and install the AppThe overlay and CSP: what Maple asks of your policy
What a comment carries: the page text a comment stores
SECURITY.md: reporting a vulnerability
Releasing: changesets, trusted publishing, a new package name
Evals, with cases for assist, doc drift and mock plans
Demo recorder: how the clip above is made
Ported helpers, network mocks and the AI tier notes
Security
Report a vulnerability privately, as SECURITY.md describes. To run
Maple in an organisation, read docs/security.md: who
trusts what, which secrets exist where, and a checklist to work through before a
preview is shared.
See it running
The project site is maple-kit.org. To run the Vite example yourself:
nvm use
pnpm install
pnpm --filter @maple-kit/example-vite dev # http://localhost:5173A real Vite application with three comments already on it: marks on the page,
the island in the corner, and all three picks working against an SDK route the
dev server mounts. examples/vite-app says what it does
and does not prove.
Development
nvm use # or fnm use, mise install — .nvmrc pins the version
pnpm install
pnpm hooks
pnpm lint && pnpm typecheck && pnpm testRequires Node 24, the active LTS, and pnpm 10. Switch before the install:
pnpm 10 and 11 load node:sqlite, which Node 23 does not have, so on the wrong
version pnpm crashes rather than telling you the version is wrong. See
CONTRIBUTING.md.
Licence
Apache-2.0. Contributions are accepted under the
Developer Certificate of Origin — sign off with git commit -s. There is
no CLA.
Available Tools
5 toolsget_comment_contextGet a comment's contextARead-only
Return everything needed to act on one comment: anchor, viewport, surrounding markup and any replay link. A comment written under a Maple Mock carries mock.recipe and mock.replay, a link to the page in that state.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The comment's id. | |
| branch | Yes | Head branch of the pull request the comments are on, such as `feat/login`. Not a pull request number. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description carries most of the burden and does add real domain context: the returned fields and the semantics of mock.recipe and mock.replay for comments written under a Maple Mock. It still omits error/not-found behavior, but the added payload detail goes meaningfully beyond the safety hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary outcome and followed by a single clarifying note about mock-authored comments. Nothing is padded, though the second sentence is somewhat niche.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on the job of describing the return payload, and it does list the key components (anchor, viewport, markup, replay link). It is adequate for a read-only two-parameter tool, with only minor gaps around error handling and return shape detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (id, branch) are documented in the schema, including the useful 'not a pull request number' clarification. The description adds nothing about parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (return) plus a concrete resource (one comment's context) and an enumeration of what that context contains: anchor, viewport, surrounding markup, replay link. It is clearly distinguishable from siblings like list_comments (bulk) and resolve_comment (mutation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'everything needed to act on one comment' implies the use case, but there is no explicit when-to-use guidance and no alternatives named. An agent must infer that this is the follow-up to list_comments rather than being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_commentsList review commentsBRead-only
Return the review comments on a branch, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| branch | Yes | Head branch of the pull request the comments are on, such as `feat/login`. Not a pull request number. | |
| statuses | No | Only these states. Omit for all of them. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds one useful behavioral trait — results are ordered newest first — but says nothing about pagination, result caps, or permissions, which matters for a listing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the resource, scope, and ordering, and zero filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with annotations covering safety and no output schema, the description covers purpose, scope, and ordering adequately. It stops short of describing result size, pagination, or how it relates to 'wait_for_comments', leaving a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself is strong (the branch description explicitly warns it is not a PR number, and statuses enumerates the filter values). The description adds nothing beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Return the review comments') plus the scoping unit ('on a branch') and ordering ('newest first'). It is clear on its own, but it does not explicitly distinguish itself from the sibling 'wait_for_comments', so an agent must infer the difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no mention of the obvious alternative 'wait_for_comments' or 'get_comment_context'. The agent gets no signal about whether to poll this or wait, or whether to use it for a single comment versus a list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_commentResolve a review commentB
Mark a comment resolved, recording the commit that addressed it.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The comment's id. | |
| sha | Yes | Commit you believe addresses it. | |
| note | No | What you changed, for the reviewer. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation only supplies readOnlyHint=false, which the description corroborates by framing this as a state-changing resolution. The description adds the meaningful behavioral detail that the addressing commit is recorded, but says nothing about idempotency (what happens if already resolved), reversibility, or whether 'note' is posted to the reviewer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action and its effect; no filler, no redundancy. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Parameters are fully covered by the schema and there is no output schema to explain, so the remaining burden is behavioral. For a mutation tool with essentially no annotation coverage beyond readOnlyHint, the description stops short of covering failure modes or post-resolution behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so 'id', 'sha', and 'note' are already documented in the schema. The description only restates the sha's role ('the commit that addressed it'), adding marginal value beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Mark resolved') and resource ('a comment'), plus the side effect of recording the addressing commit. It does not explicitly differentiate from siblings, but the siblings (list_comments, wait_for_comments, get_comment_context, start_solo) are clearly different operations, so the risk of confusion is low.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives, no prerequisites (e.g. required permissions or whether the comment must be open), and no exclusions. The agent must infer context entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_soloKeep a preview's comments on this machineA
Start a bridge on localhost and return the link that pairs a deployed preview with it. The reviewer opens the link in their browser and the comments they write there land in .maple/ for list_comments and wait_for_comments to read. They never count toward the merge gate.
| Name | Required | Description | Default |
|---|---|---|---|
| previewUrl | Yes | The deployed preview's URL, such as `https://feat-login.preview.example`. Only that origin is paired with the bridge. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, so the description carries most of the burden and delivers: it starts a localhost bridge, comments persist under .maple/, and they explicitly do not count toward the merge gate. That last point is a genuine behavioral trait not derivable from structured fields. It omits lifecycle details such as how the bridge is stopped or whether it persists across sessions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action and return value, followed by the workflow and the merge-gate caveat. Every sentence carries information an agent needs; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly tells the agent what is returned (a pairing link) and where comments end up. For a stateful, non-read-only tool it could say more about shutdown or multiplicity of bridges, but the core call-and-consequence picture is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single previewUrl parameter already documents the format and the origin-pairing constraint. The description only restates that the preview is 'paired' with the bridge, adding no syntax or edge-case detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb and resource — 'Start a bridge on localhost and return the link' — and immediately explains what the link does, which compensates for the opaque name 'start_solo'. It also names the sibling tools (list_comments, wait_for_comments) that consume the output, so an agent can place it in the workflow without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is clear: use this when you want a deployed preview's reviewer comments to land on this machine. It describes the reviewer-side workflow and the merge-gate behavior, but never states when NOT to use it or what alternative exists for other comment-routing modes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_commentsWait for review commentsARead-only
Block until a new comment arrives or the wait elapses. Returns status "timeout" rather than failing when nothing arrives.
| Name | Required | Description | Default |
|---|---|---|---|
| branch | Yes | Head branch of the pull request the comments are on, such as `feat/login`. Not a pull request number. | |
| cursor | No | Timestamp from a previous call. Only comments newer than this are returned; omit it to drain everything already there. | |
| timeoutMs | No | How long to wait. Clamped to 55 seconds, under every client's ceiling. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds a genuinely important non-obvious trait: a timeout returns status "timeout" rather than raising an error, which changes how an agent should handle the result. It does not restate safety, and the only missing trait (the 55s clamp) is already covered by the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the blocking condition comes first and the timeout return contract second, so the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does explain the timeout status, which is the critical case. It stops short of describing the shape of a successful return or how to continue waiting from a prior cursor, leaving a modest gap for a blocking tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema itself documents branch, cursor, and timeoutMs semantics thoroughly, including the 55-second clamp and the drain-everything behavior of an omitted cursor. The description adds nothing about parameters, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (block until a new comment arrives or the wait elapses) on a specific resource, which is more informative than the name alone because it establishes blocking semantics rather than a one-shot fetch. It implicitly contrasts with the polling-style sibling list_comments but never names it, so differentiation relies on inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The blocking behavior implies the scenario (waiting for new review activity instead of repeatedly polling), but there is no explicit when-to-use statement, no named alternative, and no guidance on cursor resumption across calls. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
get_comment_context - First observed
list_comments - First observed
resolve_comment - First observed
start_solo - First observed
wait_for_comments
TDQS
Scored across 5 tools
Each tool targets a distinct operation in the review-comment workflow: listing existing comments, blocking for new ones, resolving a comment, fetching detailed context for one comment, and starting a local bridge. There is no meaningful overlap between these purposes.
All tool names use snake_case and follow a verb-led pattern (list_comments, wait_for_comments, resolve_comment, get_comment_context, start_solo). The convention is consistent throughout.
Five tools is well-scoped for a specialized review-comment bridge server. Each tool serves a clear, non-redundant role in the workflow.
The surface covers reading, waiting for, resolving, and getting context on comments, plus starting the bridge. Minor gaps exist, such as no explicit unresolve or reply operation, but core workflows are fully supported.
Maintenance
Related MCP Connectors
Comment on AI-generated webpages; feedback flows back to your coding agent. Free, MIT, local-first.
Human feedback for AI agents: share HTML, get a live review link, read anchored notes as markdown.
A design agent in your coding agent's loop: briefs each screen, reviews UX and visual quality.
Human-in-the-loop for AI coding agents — ask questions, get approvals via Slack.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceHuman-to-AI code review bridge. Annotate UI elements in the browser with review comments, and AI agents read the feedback via MCP to fix code automatically — with full element context (CSS selector, styles, DOM path, accessibility info). 10 MCP tools, framework-agnostic Web Component, zero-config install via uvx.9-
- AlicenseAqualityAmaintenanceTurns product feedback pinned to a live UI into an actionable backlog for Claude Code — list comments, open one with its target element's HTML, computed styles and screenshot, and update its status.467,614,124 npm1MIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to capture pixel-accurate screenshots and DOM state from localhost apps, receive user-drawn instructions and reference images, and manage implementation review cycles.11 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables coding agents to pull structured UI feedback captured in the browser — including element selectors, bounding boxes, computed styles, screenshots, and annotations — and to mark issues as fixed.MIT