Skip to main content
Glama

report_reproduction

Record the result of reproducing a MergeSafe review finding on its GitHub thread. Reply with the exact command, exit code, and output tail as evidence for judges.

Instructions

Report the result of reproducing one MergeSafe finding, as a reply on its thread (with your own GitHub token). repo is "owner/name"; comment_id is the finding's comment_id from get_findings. Pass the exact command you ran, its exit code (non-zero when the defect showed) and the tail of its output; the output is shortened to its last lines; a command over 500 characters is refused rather than shortened. Run it against the code as it stands, before fixing, and report what happened rather than what you expected: the reply records it on the thread as evidence for whoever judges the finding.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
repoYes
commandYes
exit_codeYes
pr_numberYes
comment_idYes
output_tailYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the auth requirement ('with your own GitHub token'), the external side effect (the reply is 'recorded on the thread as evidence'), and two hard behavioral constraints (output shortened to last lines, commands over 500 characters are refused rather than shortened). These are exactly the non-obvious traits an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and side effect are front-loaded in the first clause, and every sentence conveys a distinct constraint (format, auth, inputs, truncation, workflow). It is dense and slightly run-on, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-param mutation tool with no annotations and no output schema, the description covers the essentials an agent needs: auth, side effect, input provenance, and validation limits. Minor omissions remain (behavior on an invalid comment_id, whether the call is retryable, what pr_number affords) but nothing blocking is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and all six params are required, so the description must compensate, and it largely does: it defines the format of repo ('owner/name'), the source of comment_id, the semantics of exit_code ('non-zero when the defect showed'), the truncation behavior of output_tail, and a validation rule for command. pr_number is left to inference, which is the only gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Report the result of reproducing one MergeSafe finding, as a reply on its thread.' This clearly distinguishes it from the sibling get_findings (which supplies the finding) and from a generic reply tool, since it is scoped to reporting a reproduction result as thread evidence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing context ('comment_id is the finding's comment_id from get_findings') and an explicit workflow rule ('Run it against the code as it stands, before fixing'). It does not explicitly say when to use the sibling reply tool instead, so it stops short of full alternative-routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools