GitWarren
Summary: This server lets agents manage Git repositories and conduct code reviews entirely on the local machine, including creating and participating in review discussions, all without any cloud dependency.
Repository management: List, add, update, and remove tracked git repositories, with live git state (branch, existence).
Review creation & lifecycle: Create, list, update, and delete reviews comparing two refs (including uncommitted work), and open/close them.
Review comments: List all comment threads, add new line-level or review-level comments (with optional screenshots), reply within threads, resolve/reopen threads, and edit/delete one's own comments.
Agent identity: Set a session label to attribute comments consistently; tool name is fixed from the handshake.
Navigation links: Every review/comment result provides
guiUrl(and optionallywebUrlfor tailnet) to open in the GitWarren UI, which works even when the app isn't running.
Provides tools for reviewing local git repositories, including browsing diffs, comparing branches, and managing review conversations.
GitWarren
Code review for your own git repositories, on your own machines. Your machines, your agents, no one else's server — and no account.
Built for the moment a coding agent — Claude Code, Codex, or anything else that edits files on your disk — has just finished, and its work is sitting in your worktree uncommitted. Read that diff here, on your own machine, before it becomes a commit.
gitwarren.com has downloads for macOS, Windows and Linux.
What it does
Tell GitWarren which local git repositories you care about, then open reviews against them — a review is a comparison of two refs, presented the way a pull request is, with conversation, commits and files changed tabs.
Review work before it is a commit. If the branch you are reviewing is checked out in a worktree, GitWarren finds that worktree — wherever it is — and folds its staged, unstaged and untracked changes into the diff. That is exactly when review is most useful, and it is the part that makes this worth having.
Nothing is cached. Every branch name, commit and diff on screen is read from git at the moment it is shown.
Nothing leaves your machine. Reviews live in one SQLite file in your application-data directory. No account, no telemetry; the desktop app's one outbound request is the update check.
Agents are first-class. Local AI agents get the same capabilities through an MCP server over stdio — they open reviews, read your comments, reply in a thread and resolve what they fixed.
Desktop app or browser. It runs as an Electron app on macOS, Windows and Linux, or as a command that serves the same review UI into a browser tab. Same renderer either way; the shell is the only thing that differs.
Your other machines too. A machine with no screen at all — a VPS, a WSL distro, a box an agent works on — runs the headless half and is reviewed from somewhere else, over SSH, over
wsl.exe, or over your own tailnet. Reviews live on the machine the code is on and stay there; nothing is replicated or relayed anywhere but the computers you already own.
Comments go on the line, the way a pull request does them. The agent that wrote the code answers in the same thread, while you are still reading the diff:
Related MCP server: Batch Review
Install
The desktop app
Download it from gitwarren.com, or on macOS:
brew install --cask klarluft/tap/gitwarrenThe command line
Serves the same review UI in a browser instead of an Electron window — for a machine that will not have the app on it, or one with no screen at all:
brew install klarluft/tap/gitwarren-cli # macOS and Linux, brings its own Node
curl -fsSL https://gitwarren.com/install.sh | sh # macOS and Linux, no Homebrew needed
npx gitwarren serve # anywhere Node 22.14+ is, including WindowsThen gitwarren serve --open. See
The gitwarren command line.
Into a coding agent
If you arrive from a coding agent, start here instead. The plugin brings the MCP server and a note that teaches the agent when to open a review and how to answer your comments:
/plugin marketplace add klarluft/gitwarren-app # Claude Code, then:
/plugin install gitwarren@gitwarren
gemini extensions install https://github.com/klarluft/gitwarren-app
npx skills add klarluft/gitwarren-app # the note alone, for any agentCursor, Codex, VS Code and Kiro read the same repository from their plugin screens. See Installing it as a plugin.
Documentation
Guide | What is in it |
What a review is, how uncommitted work is found, reading and navigating a large diff, marking files reviewed, opening a file in your editor | |
The MCP tools, installing the plugin, how an agent links you back into the app, comment threads, images | |
| |
Reaching a repository over SSH, over | |
The stack, the shared service layer both surfaces call, and what is and is not stored | |
Running from source, the project layout, database migrations | |
Cutting a release, auto-update, code signing and notarization | |
The fields Glama's admin page actually reads, and why they say what they say | |
What GitWarren does not do, and why | |
The design plan behind the remote-host model |
Agents setting GitWarren up on a user's machine should read llms-install.md.
Contributing
Contributions are welcome. Anything larger than a bug fix starts as a discussion in Ideas, so the shape can be agreed before you spend time on it; once it is settled it becomes an issue. See CONTRIBUTING.md for the development workflow, the two design constraints that changes need to respect, and what a good pull request looks like here.
Before a first contribution can be merged you will be asked to sign the Contributor License Agreement. A bot handles it on your pull request; it takes about ten seconds and only happens once. The CLA keeps copyright in the codebase in one place, which is what makes it possible to offer GitWarren under a commercial licence alongside the GPL, or to change licence later, without having to track down every past contributor. You keep full ownership of your work and can use it elsewhere however you like.
Support and privacy
Questions, ideas and setups worth copying go to Discussions — Q&A if you are stuck on something, Ideas for a feature, and Show and tell for an agent or remote-machine arrangement other people should steal. Bugs you can describe — what GitWarren does, and when — go to issues. Anything you would rather not post publicly goes to contact@klarluft.com.
GitWarren keeps everything on your machine: reviews live in one SQLite file in your application-data directory, the diff is read from your git worktree, and there is no account and no telemetry. The desktop app's one outbound request is the auto-update check against this repository's GitHub Releases. The website's privacy policy covers gitwarren.com itself.
License
GitWarren is free software, licensed under the GNU General Public License, version 3 or (at your option) any later version. The full text is in LICENSE.
In short: you may use, study, modify and redistribute it, including commercially. If you distribute a modified version, or a program that incorporates this one, you must release that under the GPL as well and make the source available. That reciprocity is the point — it keeps GitWarren and anything built on it open.
The copyright is held by Klarluft B.V. (Rotterdam, The Netherlands · KVK 86875590), and every contribution is covered by the CLA. Because the copyright sits in one place rather than being spread across contributors, a licence other than the GPL — for embedding GitWarren in a closed-source product, for instance — can be granted on request: email contact@klarluft.com.
Michal Wrzosek (michal@wrzosek.pl) is the creator of GitWarren and currently its main maintainer.
Copyright © 2026 Klarluft B.V.
This program is free software: you can redistribute it and/or modify it under
the terms of the GNU General Public License as published by the Free Software
Foundation, either version 3 of the License, or (at your option) any later
version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY
WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A
PARTICULAR PURPOSE. See the GNU General Public License for more details.
You should have received a copy of the GNU General Public License along with
this program. If not, see <https://www.gnu.org/licenses/>.Available Tools
17 toolsadd_repositoryAdd repositoryA
Start tracking a git repository. path may be any directory inside the working tree - it is resolved to the repository root before being stored, so the same repository cannot be added twice under two different paths. name defaults to the folder name. Fails with NOT_A_GIT_REPOSITORY if the path is not inside a git repository.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| path | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnlyHint=false, etc.), so the description carries the burden. It discloses key behavioral traits: path resolution to repository root (preventing duplicate adds), name defaulting behavior, and failure condition (NOT_A_GIT_REPOSITORY). This goes beyond the annotations and is valuable for the agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (three sentences) and front-loaded with the core purpose. Each sentence adds essential detail: purpose, path behavior, and failure mode. No fluff or redundant phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has simple inputs and no output schema, but the description covers all necessary aspects: how path is used, name defaulting, and error conditions. There is no missing information that an agent would need to invoke it correctly, given the moderate complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description thoroughly explains both parameters: `path` can be any directory inside the working tree and is resolved to root; `name` defaults to folder name. This adds meaning that the schema (only type and constraints) does not provide, fully compensating for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (start tracking) and the resource (a git repository). It distinguishes itself from siblings like list_repositories and remove_repository by focusing on the 'add' action, and it provides specific details about path resolution that prevent ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use: when you want to start tracking a git repository. It does not explicitly mention alternatives, but the sibling names (e.g., update_repository, remove_repository) imply different purposes. It provides a key exclusion: fails if not a git repository, which helps the agent decide when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_review_commentAdd review commentA
Start a new discussion on a review. Omit filePath and line for a comment on the review as a whole; give both to attach it to a line, the way a pull-request review comment works. side picks which side of the diff line counts on - "head" (the default) for the code as it will be, "base" to remark on a line the change removed. Add startLine to comment on a block rather than a single line; line is then the last line of it, and the one the whole range follows if the code moves. The line is not required to be part of the diff. A line the patch does not print is looked up in the file itself at the head, so a remark on code this branch did not touch anchors and follows that code exactly as one on a changed line does - people read these in the review's Browse files tab. Only a line that is in neither the diff nor the file (past the end of it, or on the "base" side outside the patch) is kept without an anchor; the returned thread says which happened. Comments are attributed automatically from the MCP handshake - see agent_identity.
To include a screenshot, write the file to disk first, then reference it as an ordinary markdown image:
The dropdown renders behind the modal:
The file is copied into GitWarren and the path rewritten, so the image survives /tmp being cleaned. Always write alt text - it is what agents without vision see.
If the path contains a space - "Screen Shot 2026-09-01 at 10.32.14.png", as macOS names screenshots - wrap it in angle brackets, or markdown does not read it as an image at all and it is left as plain text:
PNG, JPEG, GIF and WebP are accepted, up to 10 MB. A path that does not resolve is left in the text as written rather than failing the call, so check the returned body if it matters.
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| line | No | ||
| side | No | head | |
| changes | No | all | |
| filePath | No | ||
| reviewId | Yes | ||
| startLine | No | ||
| agentLabel | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations. It details how anchors are resolved (a line not in the diff is looked up in the head file, only a line in neither is kept unanchored), that comments are attributed from the MCP handshake, that unattached paths are left as text rather than failing the call, image type/size limits, and the distinct guiUrl vs webUrl behaviors. It confirms the annotations (write, non-idempotent, non-destructive) without contradicting them - no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Core positioning semantics are front-loaded, which is good. But the description sprawls: three paragraphs are devoted to markdown image syntax, path-escaping with spaces, file formats, and size limits, and another two long paragraphs to tailnet/webUrl mechanics. Valuable detail, but much of it could be condensed; it does not earn every sentence's place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no output schema and zero schema descriptions, the description covers the intricate comment-placement logic and even partial return-value fields (thread anchoring result, guiUrl, webUrl). The notable gap is two undocumented parameters - 'changes' and 'agentLabel' - which no structured field explains. Nearly complete, but those omissions keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the explanatory load. It does this superbly for the hard parameters - filePath/line pairing, the 'side' head/base meaning, and startLine's block semantics. But it leaves several parameters unaddressed: the 'changes' enum (committed/all/uncommitted) is never explained, and 'agentLabel' is not mentioned at all. Strong compensation for the tricky location params, incomplete for the rest.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb-plus-resource: 'Start a new discussion on a review.' This cleanly stakes out its place against the sibling set (reply_to_review_comment, update_review_comment, delete_review_comment, resolve_review_comment) - it is the one that opens a fresh thread, not one that replies to, edits, or deletes an existing one. The title reinforces it without needing the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Start a new discussion' framing plus the description of line anchoring ('the way a pull-request review comment works') gives clear context for when this tool applies. However, it never explicitly names the sibling alternative or states an exclusion - e.g., it does not say 'for replies to an existing comment use reply_to_review_comment.' The when-to-use is strongly implied but the when-not-to-use alternative is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_identityAgent identityAIdempotent
Show how comments from this session will be attributed, and optionally set a label for the session. The tool name ("Claude Code", "Codex", ...) is taken from the MCP handshake and cannot be changed - it is what makes attribution consistent across sessions. label is a short handle for this session ("auth-refactor"), worth setting when more than one agent is working the same review; it is remembered until this server process exits.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as idempotent and non-destructive, and the description adds useful behavioral context: the tool name comes from the MCP handshake and cannot be changed, and the label persists until the server process exits. This is consistent with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with the primary action front-loaded, followed by essential constraints and one illustrative example. Every sentence adds necessary context without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description covers what the tool does, when to use it, and how the label behaves. A precise statement of the returned attribution fields would be marginally helpful, but the current description is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for the 'label' parameter (0% coverage), so the description must compensate. It does so by explaining that 'label' is a short session handle, giving an example ('auth-refactor'), and clarifying its persistence and use case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Show') and a distinct resource ('how comments from this session will be attributed'), and also covers the optional 'set a label' behavior. This clearly distinguishes it from the CRUD repository/review/comment siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when setting the label is worthwhile ('when more than one agent is working the same review') and notes the label's lifetime. It does not explicitly discuss when not to use the tool, but no sibling alternative exists, so the guidance is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_reviewCreate reviewA
Open a review comparing two refs in a tracked repository. baseRef is what the changes are measured against (usually the trunk) and headRef is the branch under review; the diff is taken from their merge base, like a pull request. Both refs must exist and share history, or the call fails with INVALID_INPUT. title defaults to " into ". If the head branch is checked out in a worktree, its uncommitted changes are part of the review as well - a review can be opened on work that has never been committed. Passing the same ref as both endpoints is allowed and does exactly that: the review then holds only the uncommitted work on that ref, and its title defaults to "Uncommitted work on ".
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| baseRef | Yes | ||
| headRef | Yes | ||
| description | No | ||
| repositoryId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate a non-read, non-idempotent, non-destructive operation. The description goes far beyond that: it specifies the INVALID_INPUT failure mode, inclusion of uncommitted changes, URL behavior across server states and restarts, and token/cookie exchange. This is exactly the kind of behavioral context an agent needs and is not in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but each sentence carries unique value. It is front-loaded with the core action and then layers ref semantics, output behavior, and URL usage. It could be trimmed slightly without losing meaning, but the structure is logical and not redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description explains the return values (guiUrl, webUrl) and their intended use, error handling, and edge cases (uncommitted work, same ref, server not running). This is complete for an agent to invoke the tool correctly and present results appropriately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter meaning. It thoroughly explains baseRef and headRef (including merge-base, uncommitted work, same-ref case) and the title default, and also covers description implicitly by allowing it. This fully compensates for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, specific verb+resource: 'Open a review comparing two refs in a tracked repository.' It distinguishes this creation tool from sibling review tools (list, get, update, remove) by explaining the diff-from-merge-base semantics, so an agent can select it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives practical conditions for use (both refs must exist, title default, same-ref allowed for uncommitted work) and explains when to offer guiUrl vs webUrl based on user context. It doesn't explicitly contrast with alternatives, but the purpose is unambiguous and the guidance is rich enough for correct invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_review_commentDelete a commentADestructiveIdempotent
Delete one message. If it was the only message in its thread, the thread goes with it. As with editing, an agent can only delete messages written by its own tool.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry destructiveHint=true and readOnlyHint=false, so the description adds value by disclosing the thread-cascade behavior and the self-authored-message restriction. These details are not available in the structured metadata and are essential for a destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with the primary action and the most consequential side effect front-loaded. There is no filler, repetition, or irrelevant detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool, the description covers the core action, a key side effect, and the permission restriction, which is enough to invoke it safely. It does not specify behavior when a message has replies, but that level of detail is not clearly necessary for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain what `id` refers to or how to obtain it; the agent must infer that it is the comment/message ID. Because the schema provides no parameter description, the description should have compensated but did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a clear verb-resource pair, 'Delete one message,' and the tool name/title specify that the target is a review comment. This distinguishes it from comment-related siblings like update_review_comment and resolve_review_comment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a relevant permission constraint: an agent can only delete messages written by its own tool-twitter. However, it does not explicitly state when deletion should be chosen over alternatives like resolve_review_comment or update_review_comment; usage is mostly implied by the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_repositoryGet repositoryARead-only
Fetch one tracked repository by id, with its live git state.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, so the description correctly aligns with that. It adds the behavioral context that the response includes 'live git state', which informs the agent about the nature of the returned data. This goes beyond the annotation and adds value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the verb and resource. Every word adds meaning, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple fetch operation with one parameter and no output schema, the description is complete: it states what is fetched and that the result includes live git state. There is no missing information that an agent would need to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter (id) with 0% description coverage. The description directly references 'by id', explaining that the id is the repository identifier. While it doesn't detail the type or constraints, the schema provides those, and the description clarifies the semantic purpose of the parameter, which compensates for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (fetch) and the resource (one tracked repository) with an id parameter. It also adds specificity by mentioning 'live git state', which distinguishes it from a simple metadata fetch. This differentiates it from sibling tools like list_repositories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for fetching a single repository when you have its id, which is clear from the phrase 'by id'. However, it does not explicitly mention when to use this over alternatives or when not to use it, so it lacks explicit exclusions but provides clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_reviewGet reviewARead-only
Fetch one review by id, with the repository it belongs to attached (including the repository path, so the changes can be inspected with git directly).
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations (readOnlyHint=true, openWorldHint=false) by detailing the behavior of the returned guiUrl and webUrl, including conditions under which each is present, the token exchange behavior, and the troubleshooting advice for when the link fails. This is significant behavioral context that an agent would need to correctly act on the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long (~200 words) but each paragraph earns its place: the first sentence is the core function, the next explains the critical guiUrl behavior, and the last clarifies the webUrl alternative. The structure is logical and front-loaded with the essential action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers the returned data (review plus repository), the two URLs and their use cases, the token behavior, and the troubleshooting path. There is nothing an agent needs to know to call the tool and present results correctly that is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema coverage at 0%, the description must compensate. It does clarify that the 'id' parameter is the review identifier ('Fetch one review by id'), adding meaning beyond the bare integer type. However, it does not elaborate on the nature of the id (e.g., where it comes from) or any format expectations, so the addition is minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Fetch') and resource ('one review by id'), and distinguishes itself from the sibling list_reviews by specifying it returns a single review with its attached repository. This gives an agent an unambiguous understanding of what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied: use this tool when you have a review id and need the details. However, the description does not explicitly mention alternatives or provide exclusions, such as 'use list_reviews to browse all reviews.' The extensive guidance on how to present the returned links is more about output handling than about when to invoke the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_repositoriesList repositoriesARead-only
List every git repository tracked in GitWarren, each with its live git state (whether the folder still exists, whether it is still a repository, and the current branch). Git state is read fresh on every call and is never cached.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, but the description adds significant behavioral context: git state is read fresh on every call and never cached. It also details the specific state fields returned, which is useful for an agent. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, and every clause adds value. The description is efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with readOnlyHint annotation, the description fully covers what the tool does, what each result contains, and the live/caching behavior. No output schema exists, but the description sufficiently communicates the return content. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description need not explain them. The baseline for zero parameters is 4, and the description adds no unnecessary parameter information, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('every git repository tracked in GitWarren'), and enumerates exactly what is included (folder existence, repo status, current branch). This clearly distinguishes it from sibling tools like get_repository or add_repository.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it is for retrieving all repositories, and the sibling names make the distinction obvious. However, it does not explicitly state when to prefer this over get_repository or mention any exclusion criteria. It provides clear context but no explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_review_commentsList review commentsARead-only
Every discussion on a review: review-level threads and line comments alike, each with its full message history and who wrote each message. Line threads also carry an anchor saying where they land in the code as it stands right now - "anchored" (the line is unchanged), "moved" (the code shifted and anchor.line is its current line) or "outdated" (the line is gone, so the comment may be about code that has since been rewritten). Threads about files the diff does not contain are resolved against those files at the head rather than against the patch, so a comment on unchanged code reports an honest anchor too. Check the anchor before acting on a line comment. An outdated one still carries anchorSnapshot, the code as it read when the comment was written, which is how to tell what was being objected to before it was rewritten.
A body may contain images, written as markdown pointing at a gitwarren://attachment/... URL - that is an internal token, not something to fetch. Each one is resolved in the comment's attachments array, where path is a real file on disk: read it with your own image tools. alt is the description whoever attached it wrote, and is worth reading first - it is often enough on its own, and it is all you get if you cannot see images.
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| reviewId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=false, so the safety profile is already known. The description adds substantial behavioral context: anchor semantics (anchored/moved/outdated), resolution of threads against head files rather than the patch, attachment URL behavior (internal token, not fetchable), guiUrl behavior across restarts and serve modes, and webUrl tailnet behavior. The only minor gap is that it doesn't explicitly state the return format or pagination, but the description is rich enough that this is a small omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the core purpose, but it is quite long and covers a lot of edge-case behavior (tailnet, tokens, cookies, restart behavior) that could be trimmed or moved to a more structured format. Every sentence does earn its place in terms of unique information, but the length makes it harder to parse quickly. It is not bloated with repetition, but it is at the upper limit of what an agent can efficiently consume.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one parameter and no output schema, the description is remarkably complete. It covers the full range of what an agent needs to know: what results contain, how to interpret anchors, how to handle attachments, how to present links, and how to troubleshoot link failures. The absence of an output schema makes the description's detailed explanation of result fields (anchor, anchorSnapshot, attachments, guiUrl, webUrl) essential, and it delivers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for the single parameter. The description doesn't explicitly explain reviewId, but the tool name and title ('List review comments') plus the opening sentence make it clear that reviewId identifies the review whose comments are listed. With only one parameter and a self-evident name, the description's implicit coverage is adequate, though it could have explicitly stated 'reviewId: the review to list comments for.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear statement of what the tool does: lists every discussion on a review, including review-level threads and line comments, with full message history and authorship. It distinguishes itself from sibling comment tools (add, reply, resolve, update, delete) by being the read/list operation, and from list_reviews by operating on comments within a review. The scope is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides extensive usage guidance: check the anchor before acting on a line comment, use anchorSnapshot for outdated comments, read the alt text before fetching attachments, show guiUrl to the user, offer webUrl when the user is not at this machine, and both when uncertain. It also explains when not to fetch attachment URLs and how to handle a link that won't open. This is explicit when-to-use and how-to-use guidance that goes far beyond a basic description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_reviewsList reviewsARead-only
List reviews, newest activity first. Filter with repositoryId and/or status ("open" or "closed"); omit both to list every review across all tracked repositories. A review records the two refs being compared, not the commits they resolved to - read the repository with git to see the actual changes.
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| repositoryId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the readOnlyHint annotation by disclosing ordering, the ref-vs-commit interpretation, guiUrl behavior regardless of whether GitWarren is running, token/cookie behavior after restarts, and the tailnet conditions for webUrl. This gives the agent concrete expectations for results and failure modes. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main action and filtering are front-loaded and paragraphs separate core behavior from link guidance. However, the description is quite long for a two-optional-parameter list tool; the token/cookie troubleshooting and tailnet scenarios are detailed enough that not every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description tells the agent what results contain (guiUrl, conditional webUrl), how to present them, and how the review object should be understood. It also covers edge cases like a refused connection and restart-token behavior, so the tool can be invoked and interpreted without additional assumptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter meaning. It explains that repositoryId and status can be combined, that status is 'open' or 'closed', and that omitting both yields all reviews. It does not say where repositoryId comes from, such as list_repositories, but the core filtering semantics are clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States exactly what the tool does: lists reviews with newest activity first, with explicit filtering by repositoryId and/or status. It also clarifies the scope ('list every review across all tracked repositories') and distinguishes the review object's semantics from commit-level details, so an agent can tell it apart from get_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear invocation context: use filters when targeting a subset, omit both to list everything, and use git when actual changed content is needed. It does not explicitly name sibling alternatives such as get_review for single-review lookups, so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_repositoryRemove repositoryADestructiveIdempotent
Stop tracking a repository. This only removes it from GitWarren - the working copy on disk is never touched.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but the description adds a crucial nuance: the working copy on disk is never touched. This clarifies the exact scope of destruction, which is not inferable from annotations alone. It aligns with the destructive hint but provides essential context that the operation does not affect local files.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence states the action, the second clarifies a key behavioral distinction. All information is front-loaded and essential, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema and annotations covering destructive/idempotent behavior, the description provides the essential scope (removal from GitWarren, not disk). It is largely complete, though it could optionally mention the effect on associated data (e.g., reviews) or return behavior, but these are not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no description for the required 'id' parameter (0% coverage), and the tool description does not mention the parameter at all. The meaning of 'id' is likely inferable from the context (repository ID), but the description fails to explicitly state what it refers to, leaving the agent to rely on assumptions. With zero schema coverage, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Stop tracking') and resource ('a repository'), and immediately distinguishes what it does not do ('the working copy on disk is never touched'). This is a clear, non-tautological statement that differentiates it from sibling tools like get_repository or add_repository.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage ('Stop tracking a repository') but does not explicitly contrast with alternatives or state when not to use it. There are sibling tools like remove_review for other resources, but no explicit routing guidance is provided. The intent is clear from context, but it lacks explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_reviewRemove reviewADestructiveIdempotent
Delete a review. This removes only the review record - no branch, commit or file in the repository is touched. To file a finished review away without deleting it, set its status to "closed" instead.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false. The description adds valuable context by specifying that only the review record is affected and that no repository content is touched. It also mentions the non-destructive alternative. While it doesn't cover errors or permissions, for a simple destructive operation with annotations already present, this is strong behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The primary action is front-loaded, followed by a precise scope clarification and a useful alternative. Every sentence earns its place, and the structure is easily scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with a single parameter, the description is complete. It explains what the tool does, what it doesn't do, and when to use an alternative. With annotations covering safety and no output schema required, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explicitly explain the 'id' parameter. While the parameter name and tool name imply it is the review ID, the description adds no additional semantics beyond what the schema (type, constraints) already states. For a low-coverage case, the description should compensate by clarifying the meaning of the ID, but it fails to do so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Delete a review.' It immediately disambiguates scope by stating it removes only the review record and nothing in the repository, and contrasts with the 'closed' status alternative. This clearly distinguishes it from sibling tools like delete_review_comment or update_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use this tool (when you want to delete) and when not to (when you want to file away a review, use 'closed' status instead). This is a direct, actionable usage guideline with an alternative provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_to_review_commentReply to a review commentA
Add a message to an existing thread. Use this rather than opening a new thread when responding to something someone already raised, so the discussion stays in one place. Thread ids come from list_review_comments.
To include a screenshot, write the file to disk first, then reference it as an ordinary markdown image:
The dropdown renders behind the modal:
The file is copied into GitWarren and the path rewritten, so the image survives /tmp being cleaned. Always write alt text - it is what agents without vision see.
If the path contains a space - "Screen Shot 2026-09-01 at 10.32.14.png", as macOS names screenshots - wrap it in angle brackets, or markdown does not read it as an image at all and it is left as plain text:
PNG, JPEG, GIF and WebP are accepted, up to 10 MB. A path that does not resolve is left in the text as written rather than failing the call, so check the returned body if it matters.
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| threadId | Yes | ||
| agentLabel | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate a write operation (readOnlyHint false) and the description reinforces that by 'Add a message.' It goes well beyond annotations by disclosing side effects: file copying and path rewriting, size limits, behavior when a path fails (left as text), and detailed link semantics (guiUrl vs webUrl, token behavior, when each applies). This is extensive behavioral disclosure that fully prepares the agent for side effects and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section serves a purpose: purpose, usage distinction, image formatting, and link semantics. It is front-loaded with the core purpose, then structured logically. While it could be trimmed slightly, the density of critical information justifies its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (image handling, link nuances) and lack of output schema, the description covers all necessary ground: how to construct body, what happens with paths, how to interpret guiUrl and webUrl, and what to do when links fail. An agent can confidently invoke this tool without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 0%, the description adds significant meaning: it explains that threadId comes from list_review_comments and gives a detailed explanation of how to format body for images, including markdown syntax, alt text, space handling, and size limits. This far exceeds what the schema provides, fully compensating for the lack of parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Add a message to an existing thread.' It also distinguishes from the sibling tool by saying 'Use this rather than opening a new thread when responding to something someone already raised,' which differentiates it from add_review_comment. The purpose is specific, unambiguous, and directly contrasts with an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance is given on when to use this tool: 'Use this rather than opening a new thread when responding to something someone already raised.' It also directs the agent to get thread ids from list_review_comments. The description further provides detailed usage for images and links, making the intended context crystal clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_review_commentResolve or reopen a threadAIdempotent
Mark a discussion settled, or reopen one. Set resolved to true once the point has been addressed, false to bring it back. Resolving records who did it; it never deletes the messages, which stay readable.
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| resolved | Yes | ||
| threadId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description adds meaningful context beyond these: it notes that resolving records who did it, that it never deletes messages (reinforcing non-destructive), and explains the behavior of returned links (guiUrl, webUrl) and token handling. This enriches the behavioral profile without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed and well-structured with paragraphs, but it is somewhat verbose. The front-loaded purpose is clear, yet the extensive explanation of guiUrl/webUrl and token behavior could be trimmed for brevity while preserving essential usage guidance. It is informative but not concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (mutation with side effects and link outputs) and the absence of an output schema, the description covers necessary aspects: it explains the mutation's recording behavior, non-destructive nature, and the two link types and when to offer each. This is sufficient for an agent to call the tool and handle results appropriately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, and it does for the `resolved` parameter: it explains that `true` marks the point addressed and `false` reopens it. `threadId` is not elaborated, but its purpose is self-evident from the parameter name and the tool's purpose. The description adds semantic value for the key parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Mark a discussion settled, or reopen one,' which is a specific verb+resource pair that clearly identifies the tool's action. It explicitly states the two possible outcomes (resolve or reopen) and ties them to the `resolved` parameter, making the purpose unmistakable and distinct from sibling tools like `update_review_comment` which handles content edits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives situational context by saying 'Set `resolved` to true once the point has been addressed, false to bring it back,' which implies when to use it. However, it does not explicitly compare against alternatives like `update_review_comment` or `delete_review_comment`, nor does it state when not to use this tool, so the guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_repositoryUpdate repositoryAIdempotent
Rename a tracked repository, or repoint it at a moved working copy. Provide at least one of name or path. A new path is validated and resolved exactly as it is when adding.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| name | No | ||
| path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=true, covering the safety profile. The description adds behavioral detail by explaining that a new path is 'validated and resolved exactly as it is when adding', which is useful beyond annotations. It does not contradict annotations, and it provides a specific validation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action and then a usage constraint. No wasted words; every sentence contributes to understanding the tool's function and parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three parameters (one required), no output schema, and no nested objects, the description covers the primary operations, parameter requirements, and validation behavior. It does not mention return values or edge cases, but these are not essential given the simplicity. The description provides enough for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must compensate. It explains the role of name and path by stating 'Provide at least one of name or path' and clarifies path validation behavior. The required id is not explicitly described, but its purpose is obvious from context and sibling tools. The description adds meaning beyond the bare schema types and optionality.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to rename a repository or repoint it to a new working copy path. It uses specific verbs ('rename', 'repoint') tied to the resource ('tracked repository') and distinguishes from siblings like add_repository, remove_repository, and list/get by focusing on updating existing repositories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates when to use the tool ('Rename... or repoint...') and provides a concrete usage constraint ('Provide at least one of name or path'). It also references that path validation behaves as in add_repository, giving context. However, it does not explicitly mention when not to use it or alternatives, though the sibling names imply the distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_reviewUpdate reviewAIdempotent
Change a review's title, description or endpoints, or set its status to "closed" or "open" again. Provide at least one field. New refs are validated exactly as they are on creation.
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| title | No | ||
| status | No | ||
| baseRef | No | ||
| headRef | No | ||
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=false). It explains validation behavior ('New refs are validated exactly as they are on creation'), the mandatory presence of at least one field, and, most notably, the detailed behavior of the returned guiUrl and webUrl links, including how to handle a user reporting a refused connection and the token/cookie mechanism. This is rich behavioral disclosure that helps the agent understand side effects and interpretation of results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured in two paragraphs, front-loaded with the core purpose and requirement, then diving into link behavior. While it is somewhat long, every sentence adds value—especially the link explanations, which are crucial for correct agent behavior. There is no redundancy, and the structure logically separates the action from the result handling. It is appropriately detailed without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters, no output schema, but informative annotations, the description covers the essential aspects: what can be updated, the requirement for at least one field, validation, and the detailed semantics of the returned links (guiUrl and webUrl), including how to present them and how to troubleshoot a broken link. This is sufficient for an agent to invoke the tool correctly and interpret its result. The absence of an output schema is mitigated by the thorough explanation of the link behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It names the modifiable fields ('title, description or endpoints') and status, and clarifies that at least one must be provided. It also mentions validation for refs. However, it does not individually explain parameters like id, baseRef, headRef, or description beyond their types in the schema. The description adds some meaning but does not fully document each parameter's semantics, leaving gaps for an agent trying to construct valid input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: 'Change a review's title, description or endpoints, or set its status to closed or open again.' It explicitly lists the mutable fields and the status option, distinguishing it from sibling tools like create_review, remove_review, and get_review. The verb 'change' and resource 'review' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear usage requirement: 'Provide at least one field.' It also explains when to offer guiUrl vs webUrl, which guides post-update behavior. However, it does not explicitly state when to use this tool versus alternatives (e.g., create_review for new reviews), though the context of updating an existing review is implied by the tool name and sibling list. The guidance is adequate but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_review_commentEdit a commentAIdempotent
Replace the text of one message. An agent can only edit messages written by its own tool - correcting yourself is expected, rewriting someone else's review is not.
Each result carries guiUrl, a link that opens this in GitWarren. Show it to the user - it exists to be clicked, and it is how they see this without going and finding it themselves. The link works whether or not GitWarren is running right now, so hand it over without checking anything. If the user says it will not open - the browser reports that the connection was refused - then GitWarren is not running on their machine, and starting it makes the same link work. When the page is served by gitwarren serve or by a plugin rather than by the app, the link also carries that launch's token, which the page exchanges for a cookie; a link minted before a restart says it needs a token, and the newest link is the one that works.
When this machine is reachable on the user's tailnet, results also carry webUrl. The two links are for two situations and neither replaces the other: guiUrl opens GitWarren on the machine the user is sitting at, which is the right one when that is this machine; webUrl opens the same review in a browser from any of their devices - a phone, a laptop across the room - because it names this machine on their tailnet. Offer webUrl when the user is not at this machine, and both when you do not know.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| body | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
This is unusually transparent. The description discloses ownership constraints, the guiUrl return field and its behavior even when GitWarren is not running, token/cookie behavior across restarts, and the conditional webUrl field. These details go well beyond what the annotations provide and are genuinely useful for invoking the tool correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized and front-loaded with the core purpose, but it is long and spends significant space on edge cases around guiUrl, tokens, cookies, and webUrl. While much of that detail is relevant, the overall length is high for a simple two-parameter update tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description thoroughly covers the return fields and their behavior, including guiUrl and webUrl conditions, plus the ownership permission that affects when the tool may be used. The information is sufficient for an agent to decide whether to call the tool and what to do with the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% property description coverage, so the description carries the burden of explaining the parameters, but it never names or explicitly maps `id` or `body`. The phrase 'Replace the text of one message' hints at what `body` does, and 'one message' implies `id`, but the agent is left to infer parameter semantics from names and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Replace the text of one message.' It also clarifies the intended scope by saying an agent can only edit messages written by its own tool. However, it does not explicitly distinguish this tool from siblings like add_review_comment or reply_to_review_comment, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use and when-not-to-use guidance: 'correcting yourself is expected, rewriting someone else's review is not.' It does not name alternative tools or explicitly route the agent to a sibling for other cases, so the guidance is strong but lacks explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
v0.1.17- First observed
add_repository - First observed
add_review_comment - First observed
agent_identity - First observed
create_review - First observed
delete_review_comment - First observed
get_repository - First observed
get_review - First observed
list_repositories - First observed
list_review_comments - First observed
list_reviews - First observed
remove_repository - First observed
remove_review - First observed
reply_to_review_comment - First observed
resolve_review_comment - First observed
update_repository - First observed
update_review - First observed
update_review_comment
TDQS
Scored across 17 tools
Every tool targets a distinct resource plus action: repositories, reviews, and review comments each have clear list/get/create/update/delete boundaries. The potential confusion between add_review_comment and reply_to_review_comment is fully resolved by their descriptions (new thread vs. reply to existing thread), and resolve/update/delete comments are clearly separate operations.
The set overwhelmingly follows a snake_case verb_noun pattern (list_repositories, create_review, update_review_comment). It loses one point for minor inconsistencies: agent_identity is a noun rather than a verb-led name, and delete_review_comment uses 'delete' while the other destructive tools use 'remove'.
At 17 tools the server is slightly above the ideal 3-15 range, but the count is justified by covering two resource domains (tracked repositories and reviews) plus full threaded comment operations and an identity/attribution tool. It feels a touch heavy rather than bloated or redundant.
Repositories have full lifecycle coverage (list/get/add/update/remove), reviews have list/get/create/update/remove plus open/closed status handling through update, and comments support listing, adding, replying, resolving, editing, and deleting. The identity tool fills the attribution gap that a multi-agent review workflow needs, leaving no obvious dead ends.
Maintenance
Related MCP Connectors
Comment on AI-generated webpages; feedback flows back to your coding agent. Free, MIT, local-first.
Versioned artifact review for people and AI agents, with contextual comments and human control.
Agentic code review, no signup to try: reality gates + frontier-model review, with veto.
- ParleyOAuthdev.weldra
Coordination hub for AI coding agents: message teammates, ask humans, audit every event.
Related MCP Servers
- AlicenseBqualityCmaintenanceAI-powered code review server that analyzes git diffs and PRs with context from project guidelines and task lists. Supports integration with Claude Code and Cursor via MCP.3MIT
- AlicenseNot gradedqualityAmaintenanceA collaborative code and markdown review tool that bridges human reviewers and AI agents, enabling both to browse files, inspect git diffs, leave structured comments, and save a final review report from the same UI in real time.2MIT
- AlicenseAqualityBmaintenanceA local MCP server that provides a safe, explicit set of Git operations for version control tasks like status, diff, branching, staging, committing, fetching, merging, and pushing.1314 npmMIT
- FlicenseNot gradedqualityBmaintenanceA review handoff tool for agent-driven coding sessions that captures worktree diffs, creates shareable review URLs, and streams reviewer feedback back to the agent.1-