Skip to main content
Glama

mergesafe-mcp

MergeSafe's findings for your own coding agent (Claude Code, Cursor, or anything else that speaks MCP over stdio). Your agent reads what MergeSafe found on a pull request, fixes it on your machine, replies on the thread, and asks for a new review.

Tool

What it does

get_findings(repo, pr_number)

The latest review's findings, each with a self-contained agent_prompt read from its inline comment

request_review(repo, pr_number, tier?, addons?)

Asks MergeSafe to review the PR's current head

reply(repo, pr_number, comment_id, body)

Replies on a finding's thread, as you

report_reproduction(repo, pr_number, comment_id, command, exit_code, output_tail)

Posts the result of reproducing a finding (the command, its exit code, the last lines of output) on its thread, as you

Install

uvx mergesafe-mcp                 # run without installing
uv tool install mergesafe-mcp     # or install it

From source: uv tool install git+https://github.com/mergesafe-ai/mergesafe-mcp

Related MCP server: PR Review MCP Server

Configure

Variable

MERGESAFE_API_KEY

An organization API key from the MergeSafe app (ms_live_…)

MERGESAFE_API_URL

Optional. Defaults to https://app.mergesafe.ai

GITHUB_TOKEN

Optional. Defaults to gh auth token. Used to read MergeSafe's comments and to post your replies

Claude Code

claude mcp add mergesafe --env MERGESAFE_API_KEY=ms_live_... -- mergesafe-mcp

Cursor (~/.cursor/mcp.json or .cursor/mcp.json)

{
  "mcpServers": {
    "mergesafe": {
      "command": "mergesafe-mcp",
      "env": { "MERGESAFE_API_KEY": "ms_live_..." }
    }
  }
}

Then ask your agent: "Fix MergeSafe's blocking findings on PR 42 in acme/widgets."

What leaves your machine

  • To MergeSafe: your API key, the repo name and the PR number. MergeSafe's API returns the review's index (severity, title, file, line) and no code.

  • To GitHub: your own token, to read comments you can already see on github.com and to post your replies.

  • Your agent writes the fixes locally, with your tokens. MergeSafe never generates or applies a patch.

Available Tools

4 tools
get_findingsA

List MergeSafe's findings on the latest review of a pull request. repo is "owner/name". Each finding has severity (P0 most severe .. P3), title, problem, file, line, status, blocking, comment_id, comment_url, agent_prompt (a self-contained instruction for fixing it, read from the inline comment on GitHub) and comment_status: ok; no_comment (the finding has no inline comment); no_github_token (set GITHUB_TOKEN or run gh auth login, then call again); unavailable (the comment could not be read: open comment_url or the PR on GitHub). Intended loop: for each finding with blocking=true, check the claim against the file as it stands now; if it holds, apply the fix the agent_prompt describes and add a regression test; run the tests; commit and push; then call reply with "Fixed in " (or why it is not a real problem) on its comment_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoYes
pr_numberYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so: it enumerates every returned field (severity P0–P3, title, problem, file, line, status, blocking, comment_id, comment_url, agent_prompt) and defines the comment_status enum values plus the auth prerequisite (GITHUB_TOKEN or gh auth login) and the recovery path for each failure state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by the return-field glossary, then the workflow. The workflow portion is dense and partly belongs to the reply tool, but every element is informative and nothing is empty padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must supply the return contract and safety context itself — and it does, including field meanings, severity ordering, error statuses, auth requirements, and the corrective loop. Nothing an agent needs to call and use it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, so the description must compensate. It explains the repo format ('owner/name') clearly; pr_number is self-evident as an integer, leaving only a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) plus the exact resource (MergeSafe findings on the latest review of a pull request), and the scope ('latest review') makes it clearly distinct from siblings like reply or request_review. An agent can identify what this returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description spells out the intended loop that follows the call (check blocking findings, apply agent_prompt fix, test, commit/push, then call reply with 'Fixed in <sha>' on comment_id), which effectively names the alternative tool and the trigger for it. It stops short of stating when not to use this tool or handling non-blocking findings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replyA

Reply on one MergeSafe inline comment thread, as you (with your own GitHub token, not MergeSafe's). repo is "owner/name"; comment_id is the finding's comment_id from get_findings. Use it to say "Fixed in " after pushing a fix, or to explain why a finding is not a real problem.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
repoYes
pr_numberYes
comment_idYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it usefully discloses the identity used ('your own GitHub token, not MergeSafe's'), which matters for a write action. However, it omits failure behavior, required permissions beyond the token, and any rate-limit or idempotency notes, so significant gaps remain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with zero filler; the core action and the auth caveat are front-loaded before the usage examples. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-required-parameter mutation tool with no annotations and no output schema, the description covers purpose, identity, parameter provenance and intended usage well. It stops short of describing the return value or error handling, which an agent would still have to discover empirically.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for most parameters: repo format ('owner/name'), comment_id provenance ('the finding's comment_id from get_findings'), and body content guidance. Only pr_number goes unexplained, though its meaning is self-evident.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Reply on one MergeSafe inline comment thread') and immediately scopes it to a single thread, which separates it from siblings like get_findings and request_review. An agent can tell exactly what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives two concrete usage scenarios ('Fixed in <sha>' after a fix, or explaining why a finding isn't real) tied to a clear trigger condition. It does not name when not to use it or point to an alternative tool, so it falls short of the top mark.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_reproductionA

Report the result of reproducing one MergeSafe finding, as a reply on its thread (with your own GitHub token). repo is "owner/name"; comment_id is the finding's comment_id from get_findings. Pass the exact command you ran, its exit code (non-zero when the defect showed) and the tail of its output; the output is shortened to its last lines; a command over 500 characters is refused rather than shortened. Run it against the code as it stands, before fixing, and report what happened rather than what you expected: the reply records it on the thread as evidence for whoever judges the finding.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoYes
commandYes
exit_codeYes
pr_numberYes
comment_idYes
output_tailYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the auth requirement ('with your own GitHub token'), the external side effect (the reply is 'recorded on the thread as evidence'), and two hard behavioral constraints (output shortened to last lines, commands over 500 characters are refused rather than shortened). These are exactly the non-obvious traits an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and side effect are front-loaded in the first clause, and every sentence conveys a distinct constraint (format, auth, inputs, truncation, workflow). It is dense and slightly run-on, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-param mutation tool with no annotations and no output schema, the description covers the essentials an agent needs: auth, side effect, input provenance, and validation limits. Minor omissions remain (behavior on an invalid comment_id, whether the call is retryable, what pr_number affords) but nothing blocking is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and all six params are required, so the description must compensate, and it largely does: it defines the format of repo ('owner/name'), the source of comment_id, the semantics of exit_code ('non-zero when the defect showed'), the truncation behavior of output_tail, and a validation rule for command. pr_number is left to inference, which is the only gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Report the result of reproducing one MergeSafe finding, as a reply on its thread.' This clearly distinguishes it from the sibling get_findings (which supplies the finding) and from a generic reply tool, since it is scoped to reporting a reproduction result as thread evidence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing context ('comment_id is the finding's comment_id from get_findings') and an explicit workflow rule ('Run it against the code as it stands, before fixing'). It does not explicitly say when to use the sibling reply tool instead, so it stops short of full alternative-routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_reviewA

Ask MergeSafe to review the current head of a pull request. repo is "owner/name" (or the bare name in the key's own organization). tier and addons are optional; leave them out to use the repository's own settings. Returns MergeSafe's answer unchanged. Call it after pushing fixes; a head that was already reviewed is refused, not charged twice.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoYes
tierNo
addonsNo
pr_numberYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does real work: it discloses idempotency/duplicate-refusal behavior, billing semantics ("not charged twice"), and that the answer is passed back unchanged. It does not cover permissions or failure modes beyond the duplicate case, so it is good but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then parameter notes, then return and timing behavior. Every sentence carries information, though the parameter remarks and the closing operational note make it slightly denser than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description supplies the return semantics (answer unchanged), the billing/duplicate behavior, and default resolution for optional inputs. It leaves auth requirements and the meaning of tier values unstated, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does for the non-obvious parameters: repo format is given as "owner/name" or a bare name resolved against the key's organization, and tier/addons are marked optional with their default resolution (the repository's own settings). It never enumerates valid tier values or addon names, leaving part of the surface undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: ask MergeSafe to review the current head of a pull request. That is plainly distinct from the siblings get_findings, reply, and report_reproduction, so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Call it after pushing fixes" gives a clear trigger, and "a head that was already reviewed is refused, not charged twice" states a when-not condition. It stops short of naming an alternative tool for the already-reviewed case, so it is strong context rather than explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedget_findings
    • First observedreply
    • First observedreport_reproduction
    • First observedrequest_review

TDQS

A4.2/5.0

Scored across 4 tools

Disambiguation4/5

get_findings and request_review are clearly distinct (read vs. trigger a review), but report_reproduction and reply both post a reply on a MergeSafe inline comment thread, differing only in payload and intent. The descriptions do explain when to use each, so an agent can pick correctly, but the overlap is real.

Naming Consistency4/5

Three tools follow a verb_noun pattern (get_findings, request_review, report_reproduction) and one is a bare verb (reply). Mostly consistent with a single minor deviation that is still readable.

Tool Count4/5

Four tools is a tight, well-scoped surface for a review-feedback loop (fetch findings, request review, respond, report reproduction). It is on the lean side but each tool earns its place.

Completeness4/5

The intended loop of fetch findings, fix, push, reply, and re-request review is fully covered. Minor gaps exist — no tool to resolve/dismiss a thread or inspect the PR diff/files — but agents can work around these via GitHub directly.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers