Skip to main content
Glama

Loupe

Loupe is an open-source visual feedback tool. You pin comments to elements of a live web page, and an AI agent such as Claude Code picks them up over MCP (Model Context Protocol, the standard that lets an agent call external tools).

Screenshot of a Loupe comment thread with a screenshot, a reply with an @mention, and a thumbs-up reaction

Contents

Related MCP server: anriss

What you can do

  • Pin feedback where it belongs. Comment on a single element, drag a region, or leave a free note on the page. Pins re-anchor after a redeploy: Loupe finds the same element again, even when its markup changes.

  • Capture what you see. Attach a screenshot or a short screen recording of a region. Elements marked with data-loupe-redact are left out of every screenshot. Screen recordings capture the real pixels, redacted elements included.

  • Discuss in threads. Reply under each comment, @mention teammates, and react with 👍 🎉 👀 🙏 ❤️ 🚀.

  • Triage on a five-stage board. Move comments through Queue, To Do, In Progress, In Review and Resolved.

  • Hand the backlog to an agent. The @loupekit/mcp server gives Claude Code 19 tools, including list_comments, get_comment, propose_change and update_status. The Laravel package ships its own MCP server with those 4 core tools when the optional laravel/mcp package is installed (php artisan mcp:start loupe).

  • Run it inside Laravel. The loupekit/laravel package stores comments in your database, uses your authentication and gates, and serves the board on your routes.

  • Route tickets between apps. Loupe Hub, a separate server you host, sends a comment from one project to another in the same organization (a group of projects in Hub), and syncs status changes and replies both ways.

Choose how to install

Integration

Use it when

Guide

Script tag

You want the widget on any page without a build step.

Embed Loupe with a script tag

npm package

Your app uses a bundler: React, Vue, or any single-page app.

Install Loupe from npm

Laravel

Your app is built on Laravel 11, 12 or 13.

Install Loupe in a Laravel app

Browser extension

You want to comment on a site you cannot change.

Use the browser extension

MCP clients

You want Claude Code, or another MCP client, to work through the comments.

Connect MCP clients

Local server and dashboard

You want the Node API and the Kanban board (columns that comment cards move across) on your machine.

Run the local server

Loupe Hub

You want to route tickets between several apps.

Self-host Loupe Hub

Try it in five minutes

This quick start runs the local server, leaves one comment on the demo page, and shows it on the board.

Prerequisites

  • Node 24. The local server runs its TypeScript files directly with node index.ts, which needs Node 24. Run node --version. You should see v24 followed by a minor version.

  • npm, which comes with Node.

  • git, to clone the repository.

Steps

  1. Clone the repository and move into it:

    git clone https://github.com/mohamed-ashraf-elsaed/loupe.git
    cd loupe

    Run every later command from this folder.

  2. Install the workspace dependencies:

    npm install

    You should see npm finish with an added ... packages summary.

  3. Build the packages:

    npm run build

    You should see the command finish with no line that starts with npm error.

  4. Create the demo project:

    npm run seed

    You should see Seeded project: pk_demo_acme, followed by an admin key line. The admin key is the project secret. The dashboard asks for it in its address. Its default value is sk_demo_acme_0f3b9c.

  5. Start the server:

    npm start

    You should see:

    [loupe] API + static on http://localhost:8787  (dashboard: /dashboard/ · demo: /demo/)

    Leave this terminal running.

  6. In your browser, open http://localhost:8787/demo/.

    You should see the Q3 Performance Overview page with the Loupe panel open on the right and a short tour. Click Next through the tour, then click Done.

  7. Leave a comment:

    1. In the panel, on the Home tab, click ✛ Pin feedback on this page.

    2. Click the Send invite button in the invite form.

    3. Type a title and a description, then click Comment.

    You should see a numbered pin, 1, on the Send invite button.

  8. In a new browser tab, open the board:

    http://localhost:8787/dashboard/?key=<ADMIN_KEY>

    Replace <ADMIN_KEY> with the admin key that npm run seed printed in step 4.

    You should see the board with five columns: Queue, To Do, In Progress, In Review and Resolved. Your comment is a card in the Queue column.

Troubleshooting

Symptom

Cause

Fix

npm start fails, or the server does not start on port 8787.

Another program is using port 8787.

Stop the other program, or start on another port with PORT=9000 npm start.

npm start fails on node index.ts.

Your Node version is older than 24.

Install Node 24, check it with node --version, then run npm start again.

Clicking Comment on the demo page fails.

You ran npm run seed with LOUPE_DEMO_SECRET set. The demo page signs requests with the default secret.

Unset LOUPE_DEMO_SECRET, run npm run seed again, then reload the demo page.

The board is empty.

You have not left a comment yet.

Do step 7, then reload the board.

Next steps

Packages

Package

Registry

What it is

@loupekit/sdk

npm

The embeddable widget: inspect, comment, capture, re-anchor.

@loupekit/shared

npm

Shared types, board stages and helpers used by every package.

@loupekit/mcp

npm (bin loupe-mcp)

The MCP server over stdio (standard input and output, which Claude Code uses to start and talk to a local server), plus a local bridge on 127.0.0.1: a small HTTP server that lets the widget reach the agent, for example in its Chat tab.

loupekit/laravel

Packagist

The Laravel package. Mirrored from packages/laravel to loupe-laravel.

Extension

Not published; load unpacked (see the guide)

The Manifest V3 browser extension. A private workspace.

Server

Not published; run from this repo

The Node API that also serves the dashboard and the demo.

Dashboard

Not published; run from this repo

The Kanban triage board.

Hub

Not published; run from this repo

Loupe Hub: organizations, projects and ticket routing. See How Loupe Hub works.

Documentation

Contributing

Author

Loupe is created and maintained by Mohamed Ashraf Elsaed.

License

MIT © Mohamed Ashraf Elsaed. See LICENSE.

Available Tools

19 tools
add_thread_messageA

Reply on a thread without changing its status — progress notes, questions, or the preview URL when it goes live. The status is left exactly as it was.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageYesMarkdown. Say something useful; a person reads this.
thread_idYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It does disclose the key non-side-effect (status unchanged), but says nothing about notifications recipients receive, message ordering/threading, editability, or permissions required. Adequate on the one trait it covers, silent on the rest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and short, but the two sentences are partially redundant — the second restates the non-mutation of status already asserted in the first. One clause of the text is not earning its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter append tool with no output schema and no annotations, the description covers the essential behavioral promise (non-destructive to status) and intended content. Missing details about downstream notification or ordering are minor at this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'message' has a schema description (Markdown) while 'thread_id' has none. The description adds useful content guidance for the message body but no format or identifier semantics for thread_id, so it only marginally compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (reply/add a message on a thread) and immediately scopes it against the status-changing siblings by declaring it does not alter status. The examples of message content (progress notes, questions, preview URL) make the intended action concrete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context via content examples and implicitly routes status changes to siblings (update_status, mark_thread_addressed) by stating the status is left untouched. It never names those alternatives explicitly, so the routing is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_pr_for_threadA

Open a pull request for a fix — or, more usually, add it to the one the repo already has. One working branch per repo accumulates every fix as its own commit, and the PR body keeps a table of them. If that PR was merged or closed, a fresh branch is started automatically. Use get_element_context first to find the file, then pass the full new contents of each file you changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoYesOwner/name, e.g. "acme/web".
filesYesEvery file you changed. They land as one atomic commit.
thread_idYesThe thread this fix answers.
base_branchNoWhat to branch from. Defaults to "main".
branch_nameNoOverride the branch name. Rarely needed.
descriptionYesOne line for the PR table, e.g. "larger checkout button".
revision_ofNoSet to a thread id when this is a revision — it gets its own revision-* branch rather than joining the accumulating one.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the one-branch-per-repo accumulation model, the PR-body table, atomic commit landing, and automatic fresh branching when a PR is merged or closed. It omits auth/permission requirements and failure modes, but the mutation semantics are unusually well explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the tool's action and the dominant case, then the branch lifecycle, then the prerequisite. No filler, though the branch-lifecycle detail could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, so the description must carry behavior, and it does for the write lifecycle. What it does not cover is what the call returns (PR identifier/URL) or permission requirements, leaving a modest gap for a 7-parameter mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 7 parameters are already documented in the schema, including the full-contents-not-a-diff rule that the description reiterates. The description adds no parameter-level syntax or constraints beyond what the schema states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource (open a PR for a fix) and immediately qualifies the common case (adding to the repo's existing PR). An agent can distinguish it from siblings like propose_change or mark_thread_addressed without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly orders the workflow: 'Use `get_element_context` first to find the file,' then pass full file contents. It also explains the merged/closed fallback condition. It stops short of saying when to prefer a sibling such as propose_change over opening a PR.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_source_for_selectionA

Just the ranked source files for the current selection or a stored thread — for when you only need to know where the component lives. Heuristic: verify the top candidate before editing it.

ParametersJSON Schema
NameRequiredDescriptionDefault
thread_idNoWork from a stored comment instead of a live selection.
selection_idNoA specific selection by correlation id (see get_selection_history). Omit for the most recent.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a useful trait — output is a ranked list and the top candidate is a heuristic that should be verified before editing — but says nothing about what happens when no selection exists, latency, or result count.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences: scope first, then the practical heuristic. No filler and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool with no output schema and no annotations, the description conveys what is returned (ranked source files) and how to treat it. It is adequate, though a note on what happens with an unresolved selection would close the last gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both thread_id and selection_id are already documented in the schema, including the 'omit for most recent' behavior. The description's 'current selection or a stored thread' only restates that mapping without adding syntax or precedence rules, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('find' / 'ranked source files') and scopes it to a current selection or stored thread, with the qualifier 'Just ... for when you only need to know where the component lives' that implicitly separates it from richer context siblings like get_element_context. It never names that sibling explicitly, so the differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a use condition ('when you only need to know where the component lives'), which is more than nothing, but names no alternative and no exclusions. An agent must infer that get_element_context or get_selection_history are the richer alternatives for other needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_activity_summaryA

What this machine's agent sessions have been doing: sessions, tool-call counts, files touched, and whether a session is still live. Use it to understand what has already been tried before starting, or to answer "what have you been working on?".

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many sessions to summarize. Default 5.
session_idNoJust this session.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It implicitly characterizes this as a read-only summary and scopes it to 'this machine', which is useful, but it says nothing about permissions, rate limits, result size, or pagination behavior despite exposing a limit parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: the first front-loads what is returned, the second immediately follows with when to use it. No filler, no redundancy, every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey return values, and it does by listing sessions, tool-call counts, files touched, and liveness. Combined with complete schema coverage for the two optional parameters, an agent has what it needs, though scope limits (e.g., how far back the summary reaches) go unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (limit and session_id) are already documented with defaults and bounds. The description adds no parameter-level detail beyond generically referring to 'sessions', so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (agent sessions on this machine) and enumerates the specific content returned: session counts, tool-call counts, files touched, and liveness. It is concrete enough to distinguish it from a generic status tool, though it doesn't explicitly contrast itself with close siblings like get_files_touched or get_recent_events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives two explicit use cases: checking what has already been tried before starting work, and answering 'what have you been working on?'. That is clear context for invocation. It stops short of naming alternatives or exclusions (e.g., when to prefer get_recent_events over this), so it doesn't reach a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_commentA

Get the full context for one comment: the request, the page, the target element's HTML, and its computed styles — everything needed to make the change.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe comment id from list_comments.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description indicates a read operation returning context. Lacks explicit mention of safety or side effects, but for a read tool this is acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence (20 words) that is front-loaded with the primary action and includes explanatory details without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Simple tool with one param, no output schema. Description sufficiently explains what the return value contains, making it complete for an agent to understand usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% coverage with description for 'id'. Tool description adds that id comes from list_comments, providing useful context beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Get the full context for one comment' and lists specific components (request, page, target element HTML, computed styles). Distinct from sibling tools list_comments and update_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies use after list_comments to retrieve detailed context for a specific comment. No explicit when-not-to-use or alternative mention, but context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_companion_messagesA

Read what the person watching has said while you were working. Messages are ALSO delivered automatically on every other tool result, so you rarely need this — use it when you want to check before finishing a task, or when someone asked you to wait for their input.

ParametersJSON Schema
NameRequiredDescriptionDefault
drainNoSet false to read without consuming. Default true.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and discloses the key non-obvious trait: messages arrive automatically on every other tool result, so this tool is normally redundant. It does not disclose that a call consumes/drains the queue by default (the drain parameter implies it), which is a notable side-effect omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the resource, then the critical caveat and usage conditions. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema and no annotations, the description is nearly complete, covering purpose and routing. The one gap is the consume-on-read default behavior, which it leaves entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single drain parameter is documented in the schema as "Set false to read without consuming. Default true." The description adds no detail about the drain/consume semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Read") and resource ("what the person watching has said while you were working"), making clear this retrieves inbound companion messages. Easily distinguished from siblings like reply_to_companion (outbound) or get_activity_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when NOT to use it ("Messages are ALSO delivered automatically on every other tool result, so you rarely need this") and names the two conditions that justify a call: checking before finishing, or waiting for requested input. This is a textbook when/when-not pairing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dashboard_urlA

The local URL of the live activity dashboard — a page a person can open to watch what the agent is doing right now. Give them this when they ask to see progress.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that the URL is local and that the target is a human-viewable live page, but says nothing about prerequisites (does the dashboard server need to be running?), URL stability, or whether the page is shareable outside the machine. Modest added value for a side-effect-free getter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler. The core identity of the return value leads, and the usage instruction follows immediately, so both sentences earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter getter with no output schema, the description should convey the return shape, and 'the local URL' does so implicitly. It omits any note about whether the dashboard must be running or whether the URL is ephemeral, which is a small residual gap given there is no output schema to lean on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to explain; the baseline of 4 applies. The description adds that the returned value is a local URL rather than an identifier or status object.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (the live activity dashboard URL) and immediately clarifies what that artifact is: 'a page a person can open to watch what the agent is doing right now.' No sibling tool overlaps with this role, so the agent can place it precisely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Give them this when they ask to see progress.' That is a clear use condition. However, it names no alternatives (e.g. get_activity_summary) and states no when-not-to-use case, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_element_contextA

Full context for one element: the element, the page, its key computed styles, the ranked source files that probably render it, and a ready-made edit prompt. Pass thread_id to work from a stored comment, or nothing to use the latest live selection. Call this before editing anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
thread_idNoWork from a stored comment instead of a live selection.
selection_idNoA specific selection by correlation id (see get_selection_history). Omit for the most recent.
include_promptNoInclude the edit prompt. Defaults to true.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully describes the return bundle (element, page, styles, ranked sources, edit prompt) and that it is a read step preceding edits, but it discloses nothing about auth requirements, cost/rate limits, or how expensive the 'ranked source files' resolution is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences: the first front-loads what the tool returns, the second gives the selection rule and the sequencing advice. No filler, no redundancy with structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on the job of telling the agent what comes back, and it does so thoroughly. It is nearly complete; only the absence of any note on behavior when no selection exists or on the meaning of 'ranked' keeps it from 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (thread_id, selection_id, include_prompt) are already documented in the schema with meanings and defaults. The description's parameter handling restates the thread_id behavior rather than adding format or edge-case detail, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Full context for one element') and enumerates exactly what 'context' means: the element, the page, computed styles, ranked source files, and an edit prompt. This clearly separates it from siblings like find_source_for_selection (which returns source files only) and get_latest_selection (which returns a selection only).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete branching rule — pass thread_id to work from a stored comment, pass nothing to use the latest live selection — plus a sequencing directive ('Call this before editing anything'). It doesn't explicitly name a sibling as an alternative or state when not to use it, so it falls just short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_files_touchedA

Every file the agent has touched in the recorded sessions, in first-seen order. Use it to see the blast radius of the work so far before adding to it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose scope ('recorded sessions') and ordering, which is useful. It never states that this is a read-only, side-effect-free operation, nor whether results are paginated or bounded — gaps that matter more here because no annotations cover them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The resource and its ordering come first and the usage cue follows, so the reader gets the essential facts immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and low-risk, and the description covers scope and ordering. However, with no output schema and no annotations, it should say at least roughly what a returned entry looks like and whether the list spans one session or all sessions, neither of which is resolvable from the structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema has nothing to explain and the baseline of 4 applies. There is no argument syntax an agent could get wrong, and the description reasonably notes the scoping ('recorded sessions') that substitutes for a filter parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and retrieval verb — every file the agent has touched — plus the ordering ('first-seen order'), which is real, non-obvious information. It does not, however, distinguish itself from potentially overlapping siblings like get_activity_summary or get_recent_events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it to see the blast radius of the work so far before adding to it' gives a clear situational trigger, which is better than nothing. But it names no alternatives and sets no exclusions, so an agent weighing this against get_activity_summary or get_recent_events gets no help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_latest_selectionA

What the user last selected in the browser, with its element, computed styles and the source files most likely to render it. Use this when someone says "this element" or "what I'm looking at" and no thread exists yet.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It does describe the return shape (element, computed styles, sources), but omits behavior when no selection exists or whether the selection is live vs cached. The single usage hint is helpful but not rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what the tool yields and followed by the triggering condition. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the return payload and the scenario for calling it. It is largely complete for a zero-param read tool, with only the no-selection/edge-case behavior left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters with 100% schema coverage, so the baseline of 4 applies; there is no parameter semantics to document or compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (the user's last browser selection) and enumerates the returned payload: element, computed styles, and likely source files. It implicitly contrasts with get_selection_history via 'last' but never names siblings to differentiate itself explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger ('this element' / 'what I'm looking at') and a condition for use ('no thread exists yet'), which effectively routes the agent away from the thread tools. It stops short of naming the alternative tools (e.g. get_thread_conversation, get_selection_history) directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_recent_eventsA

The raw agent event stream — tool calls, prompts, session starts and stops — newest first. Use it when you need to know exactly what has run, rather than a summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoFilter to one type, e.g. "tool_use" or "prompt_submit".
limitNoHow many events. Default 30.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses content shape and newest-first ordering, but says nothing about pagination behavior beyond the schema's limit, permission requirements, or read-only nature — a modest but real gap for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste: the first states what the stream is and how it is ordered, the second states when to pick it over the alternative. Nothing is repeated and the key scoping detail is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with two optional params and no output schema, the description covers the resource, its contents, and its ordering well enough to call correctly. Minor remaining gap is that it doesn't characterize the payload fields of an event, which is not documented anywhere else.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (type filter and limit with a 200 maximum) are already documented in the schema. The description adds no filter syntax or default-limit context beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource precisely ('the raw agent event stream'), enumerates its contents (tool calls, prompts, session starts/stops) and states the ordering (newest first). It also implicitly distinguishes itself from the summary-style sibling, so an agent can tell what it returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit selection rule: use it when you need to know exactly what has run 'rather than a summary'. That routes the agent away from the summary tool, though it does not name the sibling explicitly or state any exclusion conditions (e.g. volume constraints).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_selection_historyA

The recent element selections, newest first — useful when the user clicked a few things and you need to pick the right one.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many to list. Defaults to 10.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses ordering ('newest first') and that multiple selections are returned, but says nothing about permissions, whether selections persist across sessions, or the shape of each entry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence that is front-loaded with what the tool returns and ordering, then the usage cue. No filler; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameter-light read tool with no output schema, the description covers what is returned, the ordering, and the motivating scenario. The main residual gap is that it doesn't hint at what constitutes one selection entry, but that is minor at this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — the single 'limit' parameter is already documented with its default (10) and maximum (50). The description adds no meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('recent element selections') plus ordering ('newest first'), so the agent knows it returns a list rather than a single item. It implicitly separates itself from the sibling get_latest_selection via 'history'/'recent' but never names that sibling explicitly, so differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a concrete usage scenario: 'when the user clicked a few things and you need to pick the right one.' That is clear context for when to reach for this tool, but no exclusion or named alternative (e.g. get_latest_selection for the single most recent pick) is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_thread_conversationA

The whole conversation on a thread: the original request, then every reply with its author. Read it before answering so you are not repeating something already said.

ParametersJSON Schema
NameRequiredDescriptionDefault
thread_idYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the shape of the result (original request first, then all replies with authors), which is useful, but says nothing about permissions, pagination, or whether resolved/closed threads are included.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the payload description front-loaded and the read-before-replying guidance placed after it. Nothing is wasted and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the description usefully explains what comes back, which compensates for the missing output schema. However, the sole parameter is completely undocumented and no error or edge-case behavior is covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions thread_id. An agent gets no guidance on the identifier's format, where it is obtained, or what happens with an invalid or missing thread.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (the whole conversation on a thread) and delineates its scope: the original request plus every reply with author attribution. This implicitly separates it from siblings like get_comment and list_comments, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Read it before answering so you are not repeating something already said" gives a concrete situation for invoking the tool with a rationale. There is no statement of when not to use it or which sibling to prefer for partial reads, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_agent_hooksA

Install the Claude Code hooks that report tool use, prompts and sessions to Loupe. Idempotent (running it twice changes nothing), backed up before writing, and it never touches another tool's hook entries. Opt-in: nothing installs these on its own.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoThe settings file to write. Defaults to ~/.claude/settings.json.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so well: it discloses idempotency ('running it twice changes nothing'), a pre-write backup, and non-interference with other tools' hook entries. These are exactly the mutation-safety traits an agent needs before writing to a settings file, and none are derivable from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action, then safety properties, then the opt-in constraint. Every clause earns its place with no filler or repetition of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no annotations and no output schema, the description covers the safety-relevant behavior (idempotency, backup, scoping) that matters most. It stops short of stating what the call returns or how to confirm/undo the install, a minor gap given no output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single optional 'path' parameter already documented in the schema. The description adds no format or path guidance beyond it, so the baseline 3 for high-coverage schemas is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Install the Claude Code hooks that report tool use, prompts and sessions to Loupe.' This is unmistakably distinct from every sibling (get_*, list_*, reply_*, update_status), all of which are read/query or messaging operations. An agent can route to it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Opt-in: nothing installs these on its own' line implies the manual trigger condition, but it never states explicitly when an agent should call this versus leaving it alone, nor any prerequisites. No alternatives exist among siblings, so the routing burden is low, leaving usage merely implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_commentsA

List Loupe product-feedback comments for the project as a task backlog. Each item carries its board stage, priority and change type, so you can start with the most urgent. Use this to see what a PM has flagged, then work through the items.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoFilter to a single page path, e.g. /checkout.
repoNoFilter to one repository, e.g. "org/repo".
branchNoFilter to one branch, e.g. "main".
statusNoFilter by stage. Omit for all.
priorityNoFilter by priority.
changeTypeNoFilter by change type.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It usefully discloses what each returned item contains (board stage, priority, change type) and the intended ordering by urgency, but says nothing about read-only semantics, pagination, result limits, or sort order guarantees.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and result shape. The closing "then work through the items" is mildly redundant with "start with the most urgent," but the description stays tight overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description does well to sketch the returned item fields and urgency framing. However, it leaves gaps around pagination, result volume, and sorting behavior for a six-filter listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six filters (url, repo, branch, status, priority, changeType) are already documented in the schema. The description adds no syntax, format, or combination guidance beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("List") and resource ("Loupe product-feedback comments") scoped to a project, and frames the return as a task backlog. It is distinguishable from the singular sibling get_comment, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use this to see what a PM has flagged, then work through the items" gives a clear usage context and intent. It offers no explicit exclusions or named alternatives (e.g. get_comment for a single thread), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark_thread_addressedA

Hand a thread back to a human: it moves to In Review and, optionally, posts your closing note. Use this when the change is ready. It CANNOT resolve a thread — only a person does that — which is why there is no status argument.

ParametersJSON Schema
NameRequiredDescriptionDefault
messageNoA short note for the reviewer — what changed and where to look. Include a preview URL if there is one.
thread_idYesThe thread you have addressed.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the state transition to In Review, that the note is optional, and that resolution is impossible here. It omits permission requirements, idempotency, and what confirmation the caller receives, which keeps it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and effect, then the trigger, then the capability boundary. Every sentence earns its place and nothing is repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema, the description covers purpose, effect, and the key constraint. It does not describe the return/confirmation the agent should expect, which is the only remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented (message and thread_id). The description adds meaning beyond the schema by explaining the deliberate absence of a status argument, heading off an incorrect invocation; it adds no format or syntax detail, only rationale.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('hand a thread back to a human'), the resulting state change ('moves to In Review'), and an optional side effect ('posts your closing note'). This clearly separates it from siblings like add_thread_message, create_pr_for_thread, or get_thread_conversation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use this when the change is ready') and a clear capability boundary ('It CANNOT resolve a thread — only a person does that'), which functions as a when-not. It does not name a sibling tool to use instead for adjacent cases, so it stops short of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_changeA

Submit the modified UI for a comment: the rewritten HTML (and optional CSS) that resolves the PM's request. This stores your proposal on the comment so the dev team can review the code and a live preview in the dashboard. Use get_comment first to see the original element, its computed styles, and the screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe comment id from list_comments.
cssNoAccompanying CSS. Omit if the styling is inlined in the HTML.
htmlYesThe modified element markup that implements the requested change.
notesNoA short explanation of what you changed and why.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It discloses persistence ('stores your proposal on the comment') and downstream consumption ('dev team can review the code and a live preview'), which is useful. However, it omits whether a proposal overwrites prior proposals, permission/auth requirements, and whether the action is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and payload, followed by the storage/consumer effect and the prerequisite lookup. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no output schema or annotations, the description covers purpose, payload expectations, storage behavior, and the prerequisite read step. It stops short of describing the response or edge-case behavior (overwriting, validation failure), leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real context: it clarifies that html is the change-resolving markup and css is optional and should be omitted when styling is inlined, plus that id should come from a prior lookup. This meaningfully supplements the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Submit the modified UI for a comment') and immediately identifies the payload (rewritten HTML with optional CSS) that resolves the PM's request. It also explains where the result lands ('stores your proposal on the comment'), so an agent can distinguish this from read-only siblings like get_comment or list_comments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear operational sequence: call get_comment first to see the original element, computed styles, and screenshot, then submit the proposal. That establishes when to use this tool relative to a specific sibling, though it does not state exclusions (e.g. what to do if the proposal is rejected or how it differs from update_status).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reply_to_companionA

Answer the person watching, in the panel they are looking at. Use this when a companion message needs a response, when you need a decision before continuing, or when you finish something they asked about.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesYour reply. Markdown is fine.
inReplyToNoThe id of the message you are answering.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It conveys that the reply surfaces in the user's panel, but says nothing about threading behavior with inReplyTo, whether the thread gets marked addressed, delivery/notification semantics, or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, each earning its place: the first front-loads the action and destination, the second enumerates the triggers. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema or annotations, this covers purpose and timing but omits important routing context — notably how it relates to mark_thread_addressed and add_thread_message, and what happens to the thread after replying.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents 'body' (Markdown allowed) and 'inReplyTo' (id of the message being answered). The description adds no parameter-level detail beyond the schema, which is the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and target (answer the person watching, in their panel), which is distinguishable from sibling reads like get_companion_messages and thread-oriented add_thread_message. It stops short of naming those siblings, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives three concrete triggers: a companion message needs a response, a decision is needed before continuing, or something they asked about is finished. These are genuine when-to-use conditions, though no sibling alternative or exclusion is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_statusA

Move a comment along the board. Set In Progress when you start it, and In Review when the change is ready for a human — only a person resolves a comment, so never set Resolved yourself.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
statusYesStage: queue / todo / in_progress / in_review / resolved. Legacy "open" and "done" are accepted too.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses a critical policy constraint beyond the schema: the Resolved state is human-only. It does not cover idempotency, permission requirements, or failure behavior, but the most consequential behavioral gotcha is surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the action and then the per-status guidance. Every clause carries operational weight, including the trailing prohibition, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation with no output schema and no annotations, the description supplies the status semantics and the key restriction. It is slightly thin on what 'id' refers to and on the result of a successful move, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% and the 'id' parameter is undocumented in the schema, but 'Move a comment along the board' implies id is the comment identifier. The description adds real meaning to the status parameter by mapping values to workflow intent (In Progress at start, In Review when ready for human, never Resolved), which goes beyond the schema's list of accepted strings.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: moving a comment's status along the board, which clearly distinguishes it from siblings like mark_thread_addressed or reply_to_companion. It does not explicitly name a sibling alternative, so it stops short of a 5, but an agent knows exactly what this tool mutates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance tied to workflow stages (In Progress when you start, In Review when ready for a human) plus a hard when-not rule (never set Resolved yourself, because only a person resolves). This is exactly the routing context an agent needs before invoking, with nothing left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv0.14.1
    • Addedadd_thread_message
    • Addedcreate_pr_for_thread
    • Addedfind_source_for_selection
    • Addedget_activity_summary
    • Addedget_companion_messages
    • Addedget_dashboard_url
    • Addedget_element_context
    • Addedget_files_touched
    • Addedget_latest_selection
    • Addedget_recent_events
    • Addedget_selection_history
    • Addedget_thread_conversation
    • Addedinstall_agent_hooks
    • Changedlist_comments6 fields changed
      • addedInput schema / properties / branch
        Added value: +{
        +  "description": "Filter to one branch, e.g. \"main\".",
        +  "type": "string"
        +}
      • addedInput schema / properties / changeType
        Added value: +{
        +  "description": "Filter by change type.",
        +  "type": "string"
        +}
      • addedInput schema / properties / priority
        Added value: +{
        +  "description": "Filter by priority.",
        +  "type": "string"
        +}
      • addedInput schema / properties / repo
        Added value: +{
        +  "description": "Filter to one repository, e.g. \"org/repo\".",
        +  "type": "string"
        +}
      • changedInput schema / properties / status / description
        Previous value: -"Filter by status. Omit for all."New value: +"Filter by stage. Omit for all."
      • removedInput schema / properties / status / enum
        Removed value: -[
        -  "open",
        -  "in_progress",
        -  "done"
        -]
    • Addedmark_thread_addressed
    • Addedreply_to_companion
    • Changedupdate_status2 fields changed
      • addedInput schema / properties / status / description
        Added value: +"Stage: queue / todo / in_progress / in_review / resolved. Legacy \"open\" and \"done\" are accepted too."
      • removedInput schema / properties / status / enum
        Removed value: -[
        -  "open",
        -  "in_progress",
        -  "done"
        -]
  2. 1 tool updatev0.8.0
    • Addedpropose_change
  3. 3 tool updatesv0.5.2
    • First observedget_comment
    • First observedlist_comments
    • First observedupdate_status

TDQS

A3.9/5.0

Scored across 19 tools

Disambiguation4/5

Most tools target distinct resources or actions, but there is some overlap among the context-gathering tools: get_element_context, get_comment, get_latest_selection, and find_source_for_selection can all return element or source context. The detailed descriptions help clarify when to use each, though an agent could still hesitate between get_comment and get_element_context with a thread_id.

Naming Consistency5/5

All tool names use snake_case and begin with a verb (get_, list_, update_, add_, create_, etc.), following a predictable verb_noun pattern. Minor variations like prepositions or abbreviations (create_pr_for_thread) are still consistent and readable.

Tool Count4/5

With 19 tools, the set is slightly heavy for a single server, but it covers several distinct domains: agent activity monitoring, companion messaging, feedback thread management, element selection, and PR creation. Each tool appears to earn its place, though the count is at the upper edge of what feels well-scoped.

Completeness4/5

The tools cover core workflows for activity tracking, companion chat, comment/thread interaction, element context retrieval, change proposals, and PR creation. A few gaps exist, such as creating or deleting comments directly and searching/filtering comments, but agents can work around these in most scenarios.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    C
    maintenance
    Enables visual browser feedback collection directly into Claude Code. Users can point at elements in their browser and send annotated feedback that Claude can act on immediately.
    12
    1
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to pull structured UI feedback captured in the browser — including element selectors, bounding boxes, computed styles, screenshots, and annotations — and to mark issues as fixed.
    MIT