loupe-mcp
This MCP server lets an AI agent work through Loupe's product-feedback backlog: read comments, fetch full context, propose UI changes, and close the loop with status updates.
List comments (
list_comments): view the project's feedback backlog, optionally filtered by page path (url) orstatus(open,in_progress,done).Get full comment context (
get_comment): retrieve one comment byid, including the request, page, target element HTML, and computed styles.Submit a proposed change (
propose_change): post rewrittenhtml(plus optionalcssandnotes) that resolves the comment, storing it for dev review with a live preview.Update status (
update_status): move a comment toin_progresswhen starting work anddonewhen shipped, signalling the PM.Follow a suggested flow: list → get context → propose change → update status, with all tools executed directly (task support forbidden).
Provides a Composer package that embeds the Loupe widget into Laravel apps, stores comments in your database, and exposes the backlog to Claude over MCP with your authorization rules.
Loupe
Loupe is an open-source visual feedback tool. You pin comments to elements of a live web page, and an AI agent such as Claude Code picks them up over MCP (Model Context Protocol, the standard that lets an agent call external tools).

Contents
Related MCP server: anriss
What you can do
Pin feedback where it belongs. Comment on a single element, drag a region, or leave a free note on the page. Pins re-anchor after a redeploy: Loupe finds the same element again, even when its markup changes.
Capture what you see. Attach a screenshot or a short screen recording of a region. Elements marked with
data-loupe-redactare left out of every screenshot. Screen recordings capture the real pixels, redacted elements included.Discuss in threads. Reply under each comment,
@mentionteammates, and react with 👍 🎉 👀 🙏 ❤️ 🚀.Triage on a five-stage board. Move comments through Queue, To Do, In Progress, In Review and Resolved.
Hand the backlog to an agent. The
@loupekit/mcpserver gives Claude Code 19 tools, includinglist_comments,get_comment,propose_changeandupdate_status. The Laravel package ships its own MCP server with those 4 core tools when the optionallaravel/mcppackage is installed (php artisan mcp:start loupe).Run it inside Laravel. The
loupekit/laravelpackage stores comments in your database, uses your authentication and gates, and serves the board on your routes.Route tickets between apps. Loupe Hub, a separate server you host, sends a comment from one project to another in the same organization (a group of projects in Hub), and syncs status changes and replies both ways.
Choose how to install
Integration | Use it when | Guide |
Script tag | You want the widget on any page without a build step. | |
npm package | Your app uses a bundler: React, Vue, or any single-page app. | |
Laravel | Your app is built on Laravel 11, 12 or 13. | |
Browser extension | You want to comment on a site you cannot change. | |
MCP clients | You want Claude Code, or another MCP client, to work through the comments. | |
Local server and dashboard | You want the Node API and the Kanban board (columns that comment cards move across) on your machine. | |
Loupe Hub | You want to route tickets between several apps. |
Try it in five minutes
This quick start runs the local server, leaves one comment on the demo page, and shows it on the board.
Prerequisites
Node 24. The local server runs its TypeScript files directly with
node index.ts, which needs Node 24. Runnode --version. You should seev24followed by a minor version.npm, which comes with Node.
git, to clone the repository.
Steps
Clone the repository and move into it:
git clone https://github.com/mohamed-ashraf-elsaed/loupe.git cd loupeRun every later command from this folder.
Install the workspace dependencies:
npm installYou should see npm finish with an
added ... packagessummary.Build the packages:
npm run buildYou should see the command finish with no line that starts with
npm error.Create the demo project:
npm run seedYou should see
Seeded project: pk_demo_acme, followed by anadmin keyline. The admin key is the project secret. The dashboard asks for it in its address. Its default value issk_demo_acme_0f3b9c.Start the server:
npm startYou should see:
[loupe] API + static on http://localhost:8787 (dashboard: /dashboard/ · demo: /demo/)Leave this terminal running.
In your browser, open http://localhost:8787/demo/.
You should see the Q3 Performance Overview page with the Loupe panel open on the right and a short tour. Click Next through the tour, then click Done.
Leave a comment:
In the panel, on the Home tab, click ✛ Pin feedback on this page.
Click the Send invite button in the invite form.
Type a title and a description, then click Comment.
You should see a numbered pin, 1, on the Send invite button.
In a new browser tab, open the board:
http://localhost:8787/dashboard/?key=<ADMIN_KEY>Replace
<ADMIN_KEY>with the admin key thatnpm run seedprinted in step 4.You should see the board with five columns: Queue, To Do, In Progress, In Review and Resolved. Your comment is a card in the Queue column.
Troubleshooting
Symptom | Cause | Fix |
| Another program is using port 8787. | Stop the other program, or start on another port with |
| Your Node version is older than 24. | Install Node 24, check it with |
Clicking Comment on the demo page fails. | You ran | Unset |
The board is empty. | You have not left a comment yet. | Do step 7, then reload the board. |
Next steps
Follow the full walkthrough, which also connects Claude Code: Leave your first comment locally.
Add Loupe to your own app with one of the guides in Choose how to install.
Packages
Package | Registry | What it is |
npm | The embeddable widget: inspect, comment, capture, re-anchor. | |
npm | Shared types, board stages and helpers used by every package. | |
npm (bin | The MCP server over stdio (standard input and output, which Claude Code uses to start and talk to a local server), plus a local bridge on 127.0.0.1: a small HTTP server that lets the widget reach the agent, for example in its Chat tab. | |
Packagist | The Laravel package. Mirrored from | |
Not published; load unpacked (see the guide) | The Manifest V3 browser extension. A private workspace. | |
Not published; run from this repo | The Node API that also serves the dashboard and the demo. | |
Not published; run from this repo | The Kanban triage board. | |
Not published; run from this repo | Loupe Hub: organizations, projects and ticket routing. See How Loupe Hub works. |
Documentation
Contributing
Author
Loupe is created and maintained by Mohamed Ashraf Elsaed.
💼 LinkedIn: mohamedashrafelsaed
🐙 GitHub: @mohamed-ashraf-elsaed
✉️ Email: m.ashraf.saed@gmail.com
License
MIT © Mohamed Ashraf Elsaed. See LICENSE.
Available Tools
19 toolsadd_thread_messageA
Reply on a thread without changing its status — progress notes, questions, or the preview URL when it goes live. The status is left exactly as it was.
| Name | Required | Description | Default |
|---|---|---|---|
| message | Yes | Markdown. Say something useful; a person reads this. | |
| thread_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It does disclose the key non-side-effect (status unchanged), but says nothing about notifications recipients receive, message ordering/threading, editability, or permissions required. Adequate on the one trait it covers, silent on the rest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and short, but the two sentences are partially redundant — the second restates the non-mutation of status already asserted in the first. One clause of the text is not earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter append tool with no output schema and no annotations, the description covers the essential behavioral promise (non-destructive to status) and intended content. Missing details about downstream notification or ordering are minor at this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'message' has a schema description (Markdown) while 'thread_id' has none. The description adds useful content guidance for the message body but no format or identifier semantics for thread_id, so it only marginally compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (reply/add a message on a thread) and immediately scopes it against the status-changing siblings by declaring it does not alter status. The examples of message content (progress notes, questions, preview URL) make the intended action concrete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context via content examples and implicitly routes status changes to siblings (update_status, mark_thread_addressed) by stating the status is left untouched. It never names those alternatives explicitly, so the routing is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_pr_for_threadA
Open a pull request for a fix — or, more usually, add it to the one the repo already has. One working branch per repo accumulates every fix as its own commit, and the PR body keeps a table of them. If that PR was merged or closed, a fresh branch is started automatically. Use get_element_context first to find the file, then pass the full new contents of each file you changed.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | Owner/name, e.g. "acme/web". | |
| files | Yes | Every file you changed. They land as one atomic commit. | |
| thread_id | Yes | The thread this fix answers. | |
| base_branch | No | What to branch from. Defaults to "main". | |
| branch_name | No | Override the branch name. Rarely needed. | |
| description | Yes | One line for the PR table, e.g. "larger checkout button". | |
| revision_of | No | Set to a thread id when this is a revision — it gets its own revision-* branch rather than joining the accumulating one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the one-branch-per-repo accumulation model, the PR-body table, atomic commit landing, and automatic fresh branching when a PR is merged or closed. It omits auth/permission requirements and failure modes, but the mutation semantics are unusually well explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the tool's action and the dominant case, then the branch lifecycle, then the prerequisite. No filler, though the branch-lifecycle detail could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must carry behavior, and it does for the write lifecycle. What it does not cover is what the call returns (PR identifier/URL) or permission requirements, leaving a modest gap for a 7-parameter mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 7 parameters are already documented in the schema, including the full-contents-not-a-diff rule that the description reiterates. The description adds no parameter-level syntax or constraints beyond what the schema states, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource (open a PR for a fix) and immediately qualifies the common case (adding to the repo's existing PR). An agent can distinguish it from siblings like propose_change or mark_thread_addressed without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly orders the workflow: 'Use `get_element_context` first to find the file,' then pass full file contents. It also explains the merged/closed fallback condition. It stops short of saying when to prefer a sibling such as propose_change over opening a PR.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_source_for_selectionA
Just the ranked source files for the current selection or a stored thread — for when you only need to know where the component lives. Heuristic: verify the top candidate before editing it.
| Name | Required | Description | Default |
|---|---|---|---|
| thread_id | No | Work from a stored comment instead of a live selection. | |
| selection_id | No | A specific selection by correlation id (see get_selection_history). Omit for the most recent. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a useful trait — output is a ranked list and the top candidate is a heuristic that should be verified before editing — but says nothing about what happens when no selection exists, latency, or result count.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences: scope first, then the practical heuristic. No filler and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with no output schema and no annotations, the description conveys what is returned (ranked source files) and how to treat it. It is adequate, though a note on what happens with an unresolved selection would close the last gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both thread_id and selection_id are already documented in the schema, including the 'omit for most recent' behavior. The description's 'current selection or a stored thread' only restates that mapping without adding syntax or precedence rules, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('find' / 'ranked source files') and scopes it to a current selection or stored thread, with the qualifier 'Just ... for when you only need to know where the component lives' that implicitly separates it from richer context siblings like get_element_context. It never names that sibling explicitly, so the differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a use condition ('when you only need to know where the component lives'), which is more than nothing, but names no alternative and no exclusions. An agent must infer that get_element_context or get_selection_history are the richer alternatives for other needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_activity_summaryA
What this machine's agent sessions have been doing: sessions, tool-call counts, files touched, and whether a session is still live. Use it to understand what has already been tried before starting, or to answer "what have you been working on?".
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many sessions to summarize. Default 5. | |
| session_id | No | Just this session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It implicitly characterizes this as a read-only summary and scopes it to 'this machine', which is useful, but it says nothing about permissions, rate limits, result size, or pagination behavior despite exposing a limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: the first front-loads what is returned, the second immediately follows with when to use it. No filler, no redundancy, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey return values, and it does by listing sessions, tool-call counts, files touched, and liveness. Combined with complete schema coverage for the two optional parameters, an agent has what it needs, though scope limits (e.g., how far back the summary reaches) go unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (limit and session_id) are already documented with defaults and bounds. The description adds no parameter-level detail beyond generically referring to 'sessions', so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (agent sessions on this machine) and enumerates the specific content returned: session counts, tool-call counts, files touched, and liveness. It is concrete enough to distinguish it from a generic status tool, though it doesn't explicitly contrast itself with close siblings like get_files_touched or get_recent_events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives two explicit use cases: checking what has already been tried before starting work, and answering 'what have you been working on?'. That is clear context for invocation. It stops short of naming alternatives or exclusions (e.g., when to prefer get_recent_events over this), so it doesn't reach a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_commentA
Get the full context for one comment: the request, the page, the target element's HTML, and its computed styles — everything needed to make the change.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The comment id from list_comments. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description indicates a read operation returning context. Lacks explicit mention of safety or side effects, but for a read tool this is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence (20 words) that is front-loaded with the primary action and includes explanatory details without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Simple tool with one param, no output schema. Description sufficiently explains what the return value contains, making it complete for an agent to understand usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% coverage with description for 'id'. Tool description adds that id comes from list_comments, providing useful context beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'Get the full context for one comment' and lists specific components (request, page, target element HTML, computed styles). Distinct from sibling tools list_comments and update_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies use after list_comments to retrieve detailed context for a specific comment. No explicit when-not-to-use or alternative mention, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_companion_messagesA
Read what the person watching has said while you were working. Messages are ALSO delivered automatically on every other tool result, so you rarely need this — use it when you want to check before finishing a task, or when someone asked you to wait for their input.
| Name | Required | Description | Default |
|---|---|---|---|
| drain | No | Set false to read without consuming. Default true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and discloses the key non-obvious trait: messages arrive automatically on every other tool result, so this tool is normally redundant. It does not disclose that a call consumes/drains the queue by default (the drain parameter implies it), which is a notable side-effect omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the resource, then the critical caveat and usage conditions. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema and no annotations, the description is nearly complete, covering purpose and routing. The one gap is the consume-on-read default behavior, which it leaves entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single drain parameter is documented in the schema as "Set false to read without consuming. Default true." The description adds no detail about the drain/consume semantics, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Read") and resource ("what the person watching has said while you were working"), making clear this retrieves inbound companion messages. Easily distinguished from siblings like reply_to_companion (outbound) or get_activity_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when NOT to use it ("Messages are ALSO delivered automatically on every other tool result, so you rarely need this") and names the two conditions that justify a call: checking before finishing, or waiting for requested input. This is a textbook when/when-not pairing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dashboard_urlA
The local URL of the live activity dashboard — a page a person can open to watch what the agent is doing right now. Give them this when they ask to see progress.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that the URL is local and that the target is a human-viewable live page, but says nothing about prerequisites (does the dashboard server need to be running?), URL stability, or whether the page is shareable outside the machine. Modest added value for a side-effect-free getter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler. The core identity of the return value leads, and the usage instruction follows immediately, so both sentences earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter getter with no output schema, the description should convey the return shape, and 'the local URL' does so implicitly. It omits any note about whether the dashboard must be running or whether the URL is ephemeral, which is a small residual gap given there is no output schema to lean on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to explain; the baseline of 4 applies. The description adds that the returned value is a local URL rather than an identifier or status object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (the live activity dashboard URL) and immediately clarifies what that artifact is: 'a page a person can open to watch what the agent is doing right now.' No sibling tool overlaps with this role, so the agent can place it precisely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger: 'Give them this when they ask to see progress.' That is a clear use condition. However, it names no alternatives (e.g. get_activity_summary) and states no when-not-to-use case, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_element_contextA
Full context for one element: the element, the page, its key computed styles, the ranked source files that probably render it, and a ready-made edit prompt. Pass thread_id to work from a stored comment, or nothing to use the latest live selection. Call this before editing anything.
| Name | Required | Description | Default |
|---|---|---|---|
| thread_id | No | Work from a stored comment instead of a live selection. | |
| selection_id | No | A specific selection by correlation id (see get_selection_history). Omit for the most recent. | |
| include_prompt | No | Include the edit prompt. Defaults to true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully describes the return bundle (element, page, styles, ranked sources, edit prompt) and that it is a read step preceding edits, but it discloses nothing about auth requirements, cost/rate limits, or how expensive the 'ranked source files' resolution is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences: the first front-loads what the tool returns, the second gives the selection rule and the sequencing advice. No filler, no redundancy with structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly takes on the job of telling the agent what comes back, and it does so thoroughly. It is nearly complete; only the absence of any note on behavior when no selection exists or on the meaning of 'ranked' keeps it from 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (thread_id, selection_id, include_prompt) are already documented in the schema with meanings and defaults. The description's parameter handling restates the thread_id behavior rather than adding format or edge-case detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Full context for one element') and enumerates exactly what 'context' means: the element, the page, computed styles, ranked source files, and an edit prompt. This clearly separates it from siblings like find_source_for_selection (which returns source files only) and get_latest_selection (which returns a selection only).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete branching rule — pass thread_id to work from a stored comment, pass nothing to use the latest live selection — plus a sequencing directive ('Call this before editing anything'). It doesn't explicitly name a sibling as an alternative or state when not to use it, so it falls just short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_files_touchedA
Every file the agent has touched in the recorded sessions, in first-seen order. Use it to see the blast radius of the work so far before adding to it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose scope ('recorded sessions') and ordering, which is useful. It never states that this is a read-only, side-effect-free operation, nor whether results are paginated or bounded — gaps that matter more here because no annotations cover them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The resource and its ordering come first and the usage cue follows, so the reader gets the essential facts immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and low-risk, and the description covers scope and ordering. However, with no output schema and no annotations, it should say at least roughly what a returned entry looks like and whether the list spans one session or all sessions, neither of which is resolvable from the structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema has nothing to explain and the baseline of 4 applies. There is no argument syntax an agent could get wrong, and the description reasonably notes the scoping ('recorded sessions') that substitutes for a filter parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and retrieval verb — every file the agent has touched — plus the ordering ('first-seen order'), which is real, non-obvious information. It does not, however, distinguish itself from potentially overlapping siblings like get_activity_summary or get_recent_events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it to see the blast radius of the work so far before adding to it' gives a clear situational trigger, which is better than nothing. But it names no alternatives and sets no exclusions, so an agent weighing this against get_activity_summary or get_recent_events gets no help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_latest_selectionA
What the user last selected in the browser, with its element, computed styles and the source files most likely to render it. Use this when someone says "this element" or "what I'm looking at" and no thread exists yet.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the disclosure burden. It does describe the return shape (element, computed styles, sources), but omits behavior when no selection exists or whether the selection is live vs cached. The single usage hint is helpful but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what the tool yields and followed by the triggering condition. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the return payload and the scenario for calling it. It is largely complete for a zero-param read tool, with only the no-selection/edge-case behavior left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters with 100% schema coverage, so the baseline of 4 applies; there is no parameter semantics to document or compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (the user's last browser selection) and enumerates the returned payload: element, computed styles, and likely source files. It implicitly contrasts with get_selection_history via 'last' but never names siblings to differentiate itself explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete trigger ('this element' / 'what I'm looking at') and a condition for use ('no thread exists yet'), which effectively routes the agent away from the thread tools. It stops short of naming the alternative tools (e.g. get_thread_conversation, get_selection_history) directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recent_eventsA
The raw agent event stream — tool calls, prompts, session starts and stops — newest first. Use it when you need to know exactly what has run, rather than a summary.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter to one type, e.g. "tool_use" or "prompt_submit". | |
| limit | No | How many events. Default 30. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses content shape and newest-first ordering, but says nothing about pagination behavior beyond the schema's limit, permission requirements, or read-only nature — a modest but real gap for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste: the first states what the stream is and how it is ordered, the second states when to pick it over the alternative. Nothing is repeated and the key scoping detail is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with two optional params and no output schema, the description covers the resource, its contents, and its ordering well enough to call correctly. Minor remaining gap is that it doesn't characterize the payload fields of an event, which is not documented anywhere else.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (type filter and limit with a 200 maximum) are already documented in the schema. The description adds no filter syntax or default-limit context beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource precisely ('the raw agent event stream'), enumerates its contents (tool calls, prompts, session starts/stops) and states the ordering (newest first). It also implicitly distinguishes itself from the summary-style sibling, so an agent can tell what it returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit selection rule: use it when you need to know exactly what has run 'rather than a summary'. That routes the agent away from the summary tool, though it does not name the sibling explicitly or state any exclusion conditions (e.g. volume constraints).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_selection_historyA
The recent element selections, newest first — useful when the user clicked a few things and you need to pick the right one.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many to list. Defaults to 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses ordering ('newest first') and that multiple selections are returned, but says nothing about permissions, whether selections persist across sessions, or the shape of each entry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence that is front-loaded with what the tool returns and ordering, then the usage cue. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameter-light read tool with no output schema, the description covers what is returned, the ordering, and the motivating scenario. The main residual gap is that it doesn't hint at what constitutes one selection entry, but that is minor at this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — the single 'limit' parameter is already documented with its default (10) and maximum (50). The description adds no meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('recent element selections') plus ordering ('newest first'), so the agent knows it returns a list rather than a single item. It implicitly separates itself from the sibling get_latest_selection via 'history'/'recent' but never names that sibling explicitly, so differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete usage scenario: 'when the user clicked a few things and you need to pick the right one.' That is clear context for when to reach for this tool, but no exclusion or named alternative (e.g. get_latest_selection for the single most recent pick) is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_thread_conversationA
The whole conversation on a thread: the original request, then every reply with its author. Read it before answering so you are not repeating something already said.
| Name | Required | Description | Default |
|---|---|---|---|
| thread_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the shape of the result (original request first, then all replies with authors), which is useful, but says nothing about permissions, pagination, or whether resolved/closed threads are included.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the payload description front-loaded and the read-before-replying guidance placed after it. Nothing is wasted and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description usefully explains what comes back, which compensates for the missing output schema. However, the sole parameter is completely undocumented and no error or edge-case behavior is covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions thread_id. An agent gets no guidance on the identifier's format, where it is obtained, or what happens with an invalid or missing thread.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (the whole conversation on a thread) and delineates its scope: the original request plus every reply with author attribution. This implicitly separates it from siblings like get_comment and list_comments, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Read it before answering so you are not repeating something already said" gives a concrete situation for invoking the tool with a rationale. There is no statement of when not to use it or which sibling to prefer for partial reads, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
install_agent_hooksA
Install the Claude Code hooks that report tool use, prompts and sessions to Loupe. Idempotent (running it twice changes nothing), backed up before writing, and it never touches another tool's hook entries. Opt-in: nothing installs these on its own.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | The settings file to write. Defaults to ~/.claude/settings.json. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so well: it discloses idempotency ('running it twice changes nothing'), a pre-write backup, and non-interference with other tools' hook entries. These are exactly the mutation-safety traits an agent needs before writing to a settings file, and none are derivable from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action, then safety properties, then the opt-in constraint. Every clause earns its place with no filler or repetition of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no annotations and no output schema, the description covers the safety-relevant behavior (idempotency, backup, scoping) that matters most. It stops short of stating what the call returns or how to confirm/undo the install, a minor gap given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional 'path' parameter already documented in the schema. The description adds no format or path guidance beyond it, so the baseline 3 for high-coverage schemas is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Install the Claude Code hooks that report tool use, prompts and sessions to Loupe.' This is unmistakably distinct from every sibling (get_*, list_*, reply_*, update_status), all of which are read/query or messaging operations. An agent can route to it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Opt-in: nothing installs these on its own' line implies the manual trigger condition, but it never states explicitly when an agent should call this versus leaving it alone, nor any prerequisites. No alternatives exist among siblings, so the routing burden is low, leaving usage merely implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_commentsA
List Loupe product-feedback comments for the project as a task backlog. Each item carries its board stage, priority and change type, so you can start with the most urgent. Use this to see what a PM has flagged, then work through the items.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Filter to a single page path, e.g. /checkout. | |
| repo | No | Filter to one repository, e.g. "org/repo". | |
| branch | No | Filter to one branch, e.g. "main". | |
| status | No | Filter by stage. Omit for all. | |
| priority | No | Filter by priority. | |
| changeType | No | Filter by change type. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It usefully discloses what each returned item contains (board stage, priority, change type) and the intended ordering by urgency, but says nothing about read-only semantics, pagination, result limits, or sort order guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and result shape. The closing "then work through the items" is mildly redundant with "start with the most urgent," but the description stays tight overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does well to sketch the returned item fields and urgency framing. However, it leaves gaps around pagination, result volume, and sorting behavior for a six-filter listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six filters (url, repo, branch, status, priority, changeType) are already documented in the schema. The description adds no syntax, format, or combination guidance beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("List") and resource ("Loupe product-feedback comments") scoped to a project, and frames the return as a task backlog. It is distinguishable from the singular sibling get_comment, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this to see what a PM has flagged, then work through the items" gives a clear usage context and intent. It offers no explicit exclusions or named alternatives (e.g. get_comment for a single thread), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_thread_addressedA
Hand a thread back to a human: it moves to In Review and, optionally, posts your closing note. Use this when the change is ready. It CANNOT resolve a thread — only a person does that — which is why there is no status argument.
| Name | Required | Description | Default |
|---|---|---|---|
| message | No | A short note for the reviewer — what changed and where to look. Include a preview URL if there is one. | |
| thread_id | Yes | The thread you have addressed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the state transition to In Review, that the note is optional, and that resolution is impossible here. It omits permission requirements, idempotency, and what confirmation the caller receives, which keeps it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and effect, then the trigger, then the capability boundary. Every sentence earns its place and nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema, the description covers purpose, effect, and the key constraint. It does not describe the return/confirmation the agent should expect, which is the only remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented (message and thread_id). The description adds meaning beyond the schema by explaining the deliberate absence of a status argument, heading off an incorrect invocation; it adds no format or syntax detail, only rationale.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('hand a thread back to a human'), the resulting state change ('moves to In Review'), and an optional side effect ('posts your closing note'). This clearly separates it from siblings like add_thread_message, create_pr_for_thread, or get_thread_conversation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use this when the change is ready') and a clear capability boundary ('It CANNOT resolve a thread — only a person does that'), which functions as a when-not. It does not name a sibling tool to use instead for adjacent cases, so it stops short of the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_changeA
Submit the modified UI for a comment: the rewritten HTML (and optional CSS) that resolves the PM's request. This stores your proposal on the comment so the dev team can review the code and a live preview in the dashboard. Use get_comment first to see the original element, its computed styles, and the screenshot.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The comment id from list_comments. | |
| css | No | Accompanying CSS. Omit if the styling is inlined in the HTML. | |
| html | Yes | The modified element markup that implements the requested change. | |
| notes | No | A short explanation of what you changed and why. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses persistence ('stores your proposal on the comment') and downstream consumption ('dev team can review the code and a live preview'), which is useful. However, it omits whether a proposal overwrites prior proposals, permission/auth requirements, and whether the action is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and payload, followed by the storage/consumer effect and the prerequisite lookup. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no output schema or annotations, the description covers purpose, payload expectations, storage behavior, and the prerequisite read step. It stops short of describing the response or edge-case behavior (overwriting, validation failure), leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real context: it clarifies that html is the change-resolving markup and css is optional and should be omitted when styling is inlined, plus that id should come from a prior lookup. This meaningfully supplements the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Submit the modified UI for a comment') and immediately identifies the payload (rewritten HTML with optional CSS) that resolves the PM's request. It also explains where the result lands ('stores your proposal on the comment'), so an agent can distinguish this from read-only siblings like get_comment or list_comments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear operational sequence: call get_comment first to see the original element, computed styles, and screenshot, then submit the proposal. That establishes when to use this tool relative to a specific sibling, though it does not state exclusions (e.g. what to do if the proposal is rejected or how it differs from update_status).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_to_companionA
Answer the person watching, in the panel they are looking at. Use this when a companion message needs a response, when you need a decision before continuing, or when you finish something they asked about.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Your reply. Markdown is fine. | |
| inReplyTo | No | The id of the message you are answering. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It conveys that the reply surfaces in the user's panel, but says nothing about threading behavior with inReplyTo, whether the thread gets marked addressed, delivery/notification semantics, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, each earning its place: the first front-loads the action and destination, the second enumerates the triggers. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema or annotations, this covers purpose and timing but omits important routing context — notably how it relates to mark_thread_addressed and add_thread_message, and what happens to the thread after replying.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents 'body' (Markdown allowed) and 'inReplyTo' (id of the message being answered). The description adds no parameter-level detail beyond the schema, which is the expected baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and target (answer the person watching, in their panel), which is distinguishable from sibling reads like get_companion_messages and thread-oriented add_thread_message. It stops short of naming those siblings, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three concrete triggers: a companion message needs a response, a decision is needed before continuing, or something they asked about is finished. These are genuine when-to-use conditions, though no sibling alternative or exclusion is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_statusA
Move a comment along the board. Set In Progress when you start it, and In Review when the change is ready for a human — only a person resolves a comment, so never set Resolved yourself.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| status | Yes | Stage: queue / todo / in_progress / in_review / resolved. Legacy "open" and "done" are accepted too. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses a critical policy constraint beyond the schema: the Resolved state is human-only. It does not cover idempotency, permission requirements, or failure behavior, but the most consequential behavioral gotcha is surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the action and then the per-status guidance. Every clause carries operational weight, including the trailing prohibition, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter mutation with no output schema and no annotations, the description supplies the status semantics and the key restriction. It is slightly thin on what 'id' refers to and on the result of a successful move, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the 'id' parameter is undocumented in the schema, but 'Move a comment along the board' implies id is the comment identifier. The description adds real meaning to the status parameter by mapping values to workflow intent (In Progress at start, In Review when ready for human, never Resolved), which goes beyond the schema's list of accepted strings.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: moving a comment's status along the board, which clearly distinguishes it from siblings like mark_thread_addressed or reply_to_companion. It does not explicitly name a sibling alternative, so it stops short of a 5, but an agent knows exactly what this tool mutates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance tied to workflow stages (In Progress when you start, In Review when ready for a human) plus a hard when-not rule (never set Resolved yourself, because only a person resolves). This is exactly the routing context an agent needs before invoking, with nothing left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
v0.14.1- Added
add_thread_message - Added
create_pr_for_thread - Added
find_source_for_selection - Added
get_activity_summary - Added
get_companion_messages - Added
get_dashboard_url - Added
get_element_context - Added
get_files_touched - Added
get_latest_selection - Added
get_recent_events - Added
get_selection_history - Added
get_thread_conversation - Added
install_agent_hooks - Changed
list_comments6 fields changed- added
Input schema / properties / branchAdded value: +{ + "description": "Filter to one branch, e.g. \"main\".", + "type": "string" +} - added
Input schema / properties / changeTypeAdded value: +{ + "description": "Filter by change type.", + "type": "string" +} - added
Input schema / properties / priorityAdded value: +{ + "description": "Filter by priority.", + "type": "string" +} - added
Input schema / properties / repoAdded value: +{ + "description": "Filter to one repository, e.g. \"org/repo\".", + "type": "string" +} - changed
Input schema / properties / status / descriptionPrevious value: -"Filter by status. Omit for all."New value: +"Filter by stage. Omit for all." - removed
Input schema / properties / status / enumRemoved value: -[ - "open", - "in_progress", - "done" -]
- Added
mark_thread_addressed - Added
reply_to_companion - Changed
update_status2 fields changed- added
Input schema / properties / status / descriptionAdded value: +"Stage: queue / todo / in_progress / in_review / resolved. Legacy \"open\" and \"done\" are accepted too." - removed
Input schema / properties / status / enumRemoved value: -[ - "open", - "in_progress", - "done" -]
1 tool update
v0.8.0- Added
propose_change
3 tool updates
v0.5.2- First observed
get_comment - First observed
list_comments - First observed
update_status
TDQS
Scored across 19 tools
Most tools target distinct resources or actions, but there is some overlap among the context-gathering tools: get_element_context, get_comment, get_latest_selection, and find_source_for_selection can all return element or source context. The detailed descriptions help clarify when to use each, though an agent could still hesitate between get_comment and get_element_context with a thread_id.
All tool names use snake_case and begin with a verb (get_, list_, update_, add_, create_, etc.), following a predictable verb_noun pattern. Minor variations like prepositions or abbreviations (create_pr_for_thread) are still consistent and readable.
With 19 tools, the set is slightly heavy for a single server, but it covers several distinct domains: agent activity monitoring, companion messaging, feedback thread management, element selection, and PR creation. Each tool appears to earn its place, though the count is at the upper edge of what feels well-scoped.
The tools cover core workflows for activity tracking, companion chat, comment/thread interaction, element context retrieval, change proposals, and PR creation. A few gaps exist, such as creating or deleting comments directly and searching/filtering comments, but agents can work around these in most scenarios.
Maintenance
Related MCP Connectors
Read, reply to and resolve website feedback threads from your coding agent.
1Visual feedback from your website's visitors as tasks for your coding agent.
41Comment on AI-generated webpages; feedback flows back to your coding agent. Free, MIT, local-first.
Triage app feedback and store reviews, draft replies and release notes, tell reporters what shipped.
Related MCP Servers
- FlicenseAqualityCmaintenanceEnables visual browser feedback collection directly into Claude Code. Users can point at elements in their browser and send annotated feedback that Claude can act on immediately.121-
- FlicenseNot gradedqualityAmaintenanceEnables UI feedback loop by clicking elements, leaving comments, and letting AI coding agents (via MCP) resolve annotations interactively.2-
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to pull, triage, and resolve user feedback pinned directly on live web prototypes via MCP tools.26 npm3MIT
- AlicenseNot gradedqualityAmaintenanceEnables coding agents to pull structured UI feedback captured in the browser — including element selectors, bounding boxes, computed styles, screenshots, and annotations — and to mark issues as fixed.MIT