Skip to main content
Glama
ahines99

portco-mcp

by ahines99

certify_run

Approve or reject a generated bundle and its metrics by binding the review to the subject hash, ensuring rejected metrics are not published.

Instructions

Certify the generated bundle and its metrics. Reviewer principals only.

    `subject_hash` is the `subject_hash` of the pending certification items: it binds the approval to the
    exact packet under review (bundle, tests, waivers, open findings). `metric_decisions` maps
    metric -> approve|reject; unknown metric names are rejected, unlisted metrics are approved. Rejected
    metrics (and metrics derived from them) are not published.
    

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
run_idYes
commentNo
subject_hashYes
bundle_decisionNoapprove
metric_decisionsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
gateYes
run_idYes
reviewerYes
decisionsYes
approval_idYes
next_actionYes
subject_hashYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses the default-approve semantics for unlisted metrics, the rejection of unknown metric names, and the consequence that rejected metrics and their derivatives are not published. It also flags the authorization requirement. These are exactly the behavioral facts an agent needs before invoking a non-idempotent mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads action and authorization in the first two sentences, then details the two parameters that carry the real semantics. Dense but every clause adds information; no filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an output schema (so return values need no explanation), the description covers the critical decision semantics and side effects. The remaining gap is the unexplained run_id, comment, and bundle_decision parameters, which leaves the argument surface partly opaque.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does well for two of five parameters: subject_hash (binds approval to the exact packet under review) and metric_decisions (map of metric -> approve|reject, with rejection cascade). run_id, comment, and bundle_decision (default 'approve') are never explained, so coverage remains partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('certify the generated bundle and its metrics') and adds an authorization scope ('Reviewer principals only'). An agent can distinguish it from most siblings, though it never explicitly contrasts with publish_run, which is the nearest neighbor semantically.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'reviewer principals only' clause is a real eligibility constraint, and the metric_decisions paragraph implies usage context (acting on pending certification items). However, there is no explicit when-to-use/when-not guidance and no routing to publish_run or list_pending_reviews, leaving the workflow position to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.