Skip to main content
Glama

Audit Task

audit_task

Verify whether completed work meets a task specification using AI. Get a pass/fail verdict plus actionable criteria failures to fix before resubmitting.

Instructions

Verify whether completed work meets a task specification using AI.

Call get_fees() first to get the current fee amount and accepted assets. Each fee_hash is single-use (anti-replay protection).

Response shape (always check criteria_failed before deciding how to proceed):

{
  "verdict":         "PASS" | "FAIL",
  "status":          "approved" | "rejected",
  "score":           0-100,
  "summary":         "one-sentence explanation",
  "details":         "full reasoning",
  "criteria_met":    ["criterion A passed", "criterion B passed"],
  "criteria_failed": ["criterion C not met — specific reason"],
  "model_used":      "gemini-2.5-flash"
}

On PASS: proceed to payment or accept the deliverable. On FAIL: read criteria_failed to understand exactly what was missing. Each entry is a specific, actionable failure — not a generic rejection. Use them to tell the worker precisely what to fix before resubmitting.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
taskYesThe task requirements or specification the worker must meet.
workYesThe work, output, or proof of completion to evaluate against the specification.
fee_hashYesTransaction hash of the fee payment. For XRP: 64-char hex of an XRPL Payment tx. For USDC on Base: 0x-prefixed 66-char EVM tx hash. Each hash is single-use.
task_categoryNoEvaluation rubric. One of: default, creative, code, data, data_analysis, bug_bounty, legal, supply_chain.default
require_consensusNoWhen True, two AI models must independently agree before returning PASS. Recommended for high-stakes tasks.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds critical behavioral details: fee_hash is single-use (anti-replay protection), require_consensus triggers two-model agreement, and the response shape with criteria_failed for actionable feedback. This exceeds the annotation coverage and gives the agent essential context for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured: a purpose sentence, a prerequisite call, a response schema in a code block, and post-conditions. It is longer than average but every sentence adds value, and the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters and an output schema, the description provides the necessary workflow, response interpretation, and the single-use caveat. It covers the consensus option and fee prerequisite, making it complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented. The description adds context like fee_hash formats (XRP vs. USDC) and reiterates the single-use nature, but most semantics are covered by the schema. Baseline of 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verify) and resource (completed work against task specification). The purpose is unambiguous, but it does not explicitly distinguish itself from siblings like evaluate_escrow_work or assess_counterparty_and_job, which could overlap in function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear workflow: call get_fees() first, then use audit_task, and interpret PASS/FAIL outcomes. However, it does not explicitly state when to use this tool instead of sibling tools, leaving some ambiguity about alternative selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.