Skip to main content
Glama
IDEAManagement

idea-base-mcp-server

Official

verify_task

Grade a task's acceptance criteria against its linked PR diff, persist the AI completion score and notes, and satisfy the ai_review verification gate without repeated 403 errors.

Instructions

Run AI verification: grades this task's acceptance_criteria against the diff of its linked PR (a Haiku relevance pre-screen, then a Sonnet deep review) and PERSISTS the result to ai_completion_score / ai_completion_notes. This is exactly what update_task_status checks when the task's verification_mode is "ai_review" or "all" — call this to satisfy that gate yourself rather than repeatedly hitting 403 VERIFICATION_REQUIRED with no way through it.

COSTS MONEY — DO NOT LOOP THIS. Every call is metered via recordAIUsage and recorded to ai_generations (Sonnet, large PR-diff context) for audit. It runs FREE while your account is within its plan's included monthly verification cap; once over that cap each call is billed as AI credit overage — 2 credits per call on this account. Calling this repeatedly hoping for a higher score spends real credits for no guaranteed gain; if a genuine score lands below the 0.8 gate, that is a finding about the work, not a reason to retry.

GRADES THE DIFF, NOT THE RUNNING SYSTEM. A high score means the PR's code changes look like they satisfy the criteria on paper — it is NOT proof the deployed/running behavior actually works. Never report this score as end-to-end verification; that is a human's job or an independent task-verify sub-agent's, not this tool's.

Requires the task to already have non-empty acceptance_criteria (set via update_task) and a way to see the work: either the task's own github_pr_url (set via update_task) or a pr_url/work_description passed here. Fails with a clear message naming what is missing rather than grading an empty contract — refusal is free, no credits spent.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pr_urlNoGitHub PR URL to grade against, for this call only — does not overwrite the task's stored github_pr_url. Omit to use the task's own linked PR.
task_idYesThe ID of the task to verify. Must already have non-empty acceptance_criteria.
work_descriptionNoFree-text description of the completed work to grade against. Used ONLY as a fallback when no PR diff is available (no pr_url given here and the task has no github_pr_url); ignored whenever a PR diff can be fetched.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv2.2.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses real cost (metered via recordAIUsage, 2 credits/call over cap, audit-logged to ai_generations), persistence behavior, prerequisites, and graceful failure (refusal is free, no credits spent). It also bounds what the score means (diff on paper, not deployed behavior).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is longer than typical, but every paragraph earns its place — cost warning, scope caveat, prerequisites — and the core action is front-loaded in the first sentence. The caps-emphasis blocks aid scanning rather than pad the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-annotation, no-output-schema tool this is unusually complete: it covers cost, side effects, prerequisites, failure modes, and the meaning of the result. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: pr_url applies 'for this call only' and does not overwrite the stored github_pr_url, and work_description is a fallback used only when no PR diff exists. The task_id prerequisite (non-empty acceptance_criteria) is reinforced in prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (AI verification grading acceptance_criteria against a linked PR diff) and describes the mechanism (Haiku pre-screen, Sonnet deep review) plus the persisted side effect (ai_completion_score / ai_completion_notes). It is clearly distinguishable from siblings like update_task_status and record_verification_feedback.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the trigger condition (task verification_mode of 'ai_review' or 'all', the same gate update_task_status checks) and gives strong when-not-to guidance: do not loop it hoping for a higher score, and do not treat it as end-to-end verification. It even routes the agent to the correct alternative for running-system proof.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.