Skip to main content
Glama

judge_testing_implementation

Validate test implementation quality, coverage, and execution results after code review approval. Requires raw test runner output and test file list to approve or request improvements.

Instructions

Judge Testing Implementation

Description

Validate test quality, coverage, and execution results after code review is approved. The input MUST include real test evidence (raw test runner output and list of test files). Called when workflow_guidance.next_tool == "judge_testing_implementation".

Critical Tool Warning

  • Skipping this tool causes severe token inefficiency and wasted iterations.

  • Always invoke this tool at the appropriate stage to avoid extreme token loss and redundant processing.

  • Do not rely on assistant memory for identifiers. Always pass the exact task_id and recover it via get_current_coding_task if missing.

Args

  • task_id: string — Task UUID (required)

  • test_summary: string — Summary of the implemented tests (required)

  • test_files: list[string] — Paths to created/modified test files (required)

  • test_execution_results: string — Raw test runner output (required). For example, pytest/jest/mocha/go test/JUnit logs including pass/fail counts.

  • test_coverage_report: string — Coverage details (optional)

  • test_types_implemented: list[string] — e.g., unit, integration, e2e (optional)

  • testing_framework: string — e.g., pytest, jest (optional)

  • performance_test_results: string — Performance results (optional)

  • manual_test_notes: string — Manual testing notes (optional)

Returns

  • Response JSON schema (JudgeResponse):

{
  "$defs": {
    "FileReview": {
      "properties": {
        "path": {
          "description": "File path reviewed",
          "title": "Path",
          "type": "string"
        },
        "feedback": {
          "description": "Per-file feedback summary",
          "title": "Feedback",
          "type": "string"
        },
        "approved": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Optional per-file approval or risk flag",
          "title": "Approved"
        }
      },
      "required": [
        "path",
        "feedback"
      ],
      "title": "FileReview",
      "type": "object"
    },
    "LibraryPlanItem": {
      "properties": {
        "purpose": {
          "description": "Non-domain concern or integration point this library addresses",
          "title": "Purpose",
          "type": "string"
        },
        "selection": {
          "description": "Chosen library or internal utility (name and optional version)",
          "title": "Selection",
          "type": "string"
        },
        "source": {
          "description": "Source of solution: 'internal' for repo utility, 'external' for well-known library, 'custom' for in-house code",
          "title": "Source",
          "type": "string"
        },
        "justification": {
          "default": "",
          "description": "One-line rationale for the selection and any trade-offs",
          "title": "Justification",
          "type": "string"
        }
      },
      "required": [
        "purpose",
        "selection",
        "source"
      ],
      "title": "LibraryPlanItem",
      "type": "object"
    },
    "PlanRequiredField": {
      "description": "Specification for a required field in judge_coding_plan.",
      "properties": {
        "name": {
          "description": "Field name in the judge_coding_plan tool",
          "title": "Name",
          "type": "string"
        },
        "type": {
          "description": "Expected data type (string, list[str], list[dict], etc.)",
          "title": "Type",
          "type": "string"
        },
        "description": {
          "description": "What this field should contain",
          "title": "Description",
          "type": "string"
        },
        "required": {
          "description": "Whether this field is required",
          "title": "Required",
          "type": "boolean"
        },
        "conditional_on": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Task metadata field this requirement depends on (e.g., 'design_patterns_enforcement')",
          "title": "Conditional On"
        },
        "example_value": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Example of what this field should contain",
          "title": "Example Value"
        }
      },
      "required": [
        "name",
        "type",
        "description",
        "required"
      ],
      "title": "PlanRequiredField",
      "type": "object"
    },
    "RequirementsVersion": {
      "description": "A version of user requirements with timestamp and source.",
      "properties": {
        "content": {
          "title": "Content",
          "type": "string"
        },
        "source": {
          "title": "Source",
          "type": "string"
        },
        "timestamp": {
          "title": "Timestamp",
          "type": "integer"
        }
      },
      "required": [
        "content",
        "source"
      ],
      "title": "RequirementsVersion",
      "type": "object"
    },
    "ResearchScope": {
      "description": "Research scope enum for workflow-driven research validation.\n\nDetermines the depth and requirements for research validation:\n- NONE: No research required for this task complexity\n- LIGHT: Light research required (1+ authoritative domain source)\n- DEEP: Deep research required (2+ authoritative domain sources)",
      "enum": [
        "none",
        "light",
        "deep"
      ],
      "title": "ResearchScope",
      "type": "string"
    },
    "ReuseComponent": {
      "properties": {
        "path": {
          "description": "Repository path to the reusable component",
          "title": "Path",
          "type": "string"
        },
        "purpose": {
          "default": "",
          "description": "What part of the task this component will support",
          "title": "Purpose",
          "type": "string"
        },
        "notes": {
          "default": "",
          "description": "Any integration notes or caveats",
          "title": "Notes",
          "type": "string"
        }
      },
      "required": [
        "path"
      ],
      "title": "ReuseComponent",
      "type": "object"
    },
    "TaskMetadata": {
      "description": "Lightweight metadata for coding tasks that flows with memory layer.\n\nThis model serves as the foundation for the enhanced workflow v3 system,\nreplacing session-based tracking with task-centric approach.",
      "properties": {
        "task_id": {
          "description": "IMMUTABLE: Auto-generated UUID, primary key for memory storage",
          "title": "Task Id",
          "type": "string"
        },
        "created_at": {
          "description": "IMMUTABLE: Task creation timestamp (epoch seconds)",
          "title": "Created At",
          "type": "integer"
        },
        "title": {
          "description": "Display title for coding task (updatable)",
          "title": "Title",
          "type": "string"
        },
        "description": {
          "description": "Detailed coding task description (updatable)",
          "title": "Description",
          "type": "string"
        },
        "user_requirements": {
          "default": "",
          "description": "Current coding requirements (updatable)",
          "title": "User Requirements",
          "type": "string"
        },
        "state": {
          "$ref": "#/$defs/TaskState",
          "default": "created",
          "description": "Current task state (updatable, follows TaskState transitions)"
        },
        "task_size": {
          "$ref": "#/$defs/TaskSize",
          "description": "Task size classification for workflow optimization (XS=simple fixes, S=minor features, M=standard, L=complex, XL=major changes)"
        },
        "user_requirements_history": {
          "description": "History of requirements changes",
          "items": {
            "$ref": "#/$defs/RequirementsVersion"
          },
          "title": "User Requirements History",
          "type": "array"
        },
        "accumulated_diff": {
          "additionalProperties": true,
          "description": "Code changes accumulated over time",
          "title": "Accumulated Diff",
          "type": "object"
        },
        "modified_files": {
          "description": "List of file paths that were created or modified during task implementation",
          "items": {
            "type": "string"
          },
          "title": "Modified Files",
          "type": "array"
        },
        "test_files": {
          "description": "List of test file paths that were created during testing phase",
          "items": {
            "type": "string"
          },
          "title": "Test Files",
          "type": "array"
        },
        "test_status": {
          "additionalProperties": {
            "type": "string"
          },
          "description": "Status of different test types (unit, integration, e2e, etc.)",
          "title": "Test Status",
          "type": "object"
        },
        "updated_at": {
          "description": "Last update timestamp (epoch seconds)",
          "title": "Updated At",
          "type": "integer"
        },
        "tags": {
          "description": "Coding-related tags",
          "items": {
            "type": "string"
          },
          "title": "Tags",
          "type": "array"
        },
        "problem_domain": {
          "default": "",
          "description": "Concise statement of the problem domain and scope for this task",
          "title": "Problem Domain",
          "type": "string"
        },
        "problem_non_goals": {
          "description": "Explicit non-goals/boundaries to prevent scope creep and re-solving commodity concerns",
          "items": {
            "type": "string"
          },
          "title": "Problem Non Goals",
          "type": "array"
        },
        "library_plan": {
          "description": "Planned libraries/utilities per purpose; prefer internal reuse and well-known libraries; custom code only with justification",
          "items": {
            "$ref": "#/$defs/LibraryPlanItem"
          },
          "title": "Library Plan",
          "type": "array"
        },
        "internal_reuse_components": {
          "description": "Existing repository components/utilities to reuse with paths and purposes",
          "items": {
            "$ref": "#/$defs/ReuseComponent"
          },
          "title": "Internal Reuse Components",
          "type": "array"
        },
        "research_required": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether research is required by workflow guidance (None=undetermined, True=required, False=optional)",
          "title": "Research Required"
        },
        "research_scope": {
          "$ref": "#/$defs/ResearchScope",
          "default": "none",
          "description": "Research scope determined by workflow: none|light|deep"
        },
        "research_completed": {
          "anyOf": [
            {
              "type": "integer"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Epoch seconds when research validation passed",
          "title": "Research Completed"
        },
        "research_rationale": {
          "default": "",
          "description": "Explanation of why research was required and how the scope was determined",
          "title": "Research Rationale",
          "type": "string"
        },
        "expected_url_count": {
          "anyOf": [
            {
              "type": "integer"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "LLM-determined expected number of research URLs based on task complexity",
          "title": "Expected Url Count"
        },
        "minimum_url_count": {
          "anyOf": [
            {
              "type": "integer"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "LLM-determined minimum acceptable URL count for adequate research",
          "title": "Minimum Url Count"
        },
        "url_requirement_reasoning": {
          "default": "",
          "description": "LLM-generated explanation of why specific URL count is needed for this task",
          "title": "Url Requirement Reasoning",
          "type": "string"
        },
        "research_complexity_analysis": {
          "anyOf": [
            {
              "additionalProperties": true,
              "type": "object"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Detailed complexity analysis factors from LLM (domain, tech maturity, integration scope, etc.)",
          "title": "Research Complexity Analysis"
        },
        "internal_research_required": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether internal codebase research is needed (None=undetermined, True=required, False=not needed)",
          "title": "Internal Research Required"
        },
        "related_code_snippets": {
          "description": "Related code snippets from the codebase that are relevant to this task",
          "items": {
            "type": "string"
          },
          "title": "Related Code Snippets",
          "type": "array"
        },
        "risk_assessment_required": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether risk assessment is needed (None=undetermined, True=required, False=not needed)",
          "title": "Risk Assessment Required"
        },
        "identified_risks": {
          "description": "Areas that could be harmed by the proposed changes",
          "items": {
            "type": "string"
          },
          "title": "Identified Risks",
          "type": "array"
        },
        "risk_mitigation_strategies": {
          "description": "Strategies to mitigate identified risks",
          "items": {
            "type": "string"
          },
          "title": "Risk Mitigation Strategies",
          "type": "array"
        },
        "design_patterns_enforcement": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether design patterns are required for this task (None=undetermined, True=required, False=not needed)",
          "title": "Design Patterns Enforcement"
        },
        "plan_approved_at": {
          "anyOf": [
            {
              "type": "integer"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Timestamp when plan was approved by judge_coding_plan (None=not approved)",
          "title": "Plan Approved At"
        },
        "plan_rejection_count": {
          "default": 0,
          "description": "Number of times the plan has been rejected (max 1 allowed)",
          "title": "Plan Rejection Count",
          "type": "integer"
        },
        "code_approved_files": {
          "additionalProperties": {
            "type": "integer"
          },
          "description": "Dictionary mapping file paths to approval timestamps from judge_code_change",
          "title": "Code Approved Files",
          "type": "object"
        },
        "testing_approved_at": {
          "anyOf": [
            {
              "type": "integer"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Timestamp when testing was approved by judge_testing_implementation (None=not approved)",
          "title": "Testing Approved At"
        },
        "all_approvals_validated": {
          "default": false,
          "description": "Whether all required approvals (plan, code, testing) have been validated",
          "title": "All Approvals Validated",
          "type": "boolean"
        }
      },
      "required": [
        "title",
        "description",
        "task_size"
      ],
      "title": "TaskMetadata",
      "type": "object"
    },
    "TaskSize": {
      "description": "Task size classification for workflow optimization.\n\nSizes are based on estimated complexity and time requirements:\n- XS: Extra Small - Simple fixes, typos, minor config changes (< 30 minutes)\n- S: Small - Minor features, simple refactoring (30 minutes - 2 hours)\n- M: Medium - Standard features, moderate complexity (2-8 hours) - DEFAULT\n- L: Large - Complex features, multiple components (1-3 days)\n- XL: Extra Large - Major system changes, architectural updates (3+ days)\n\nThis classification determines planning complexity and validation depth:\n- XS/S: Basic planning requirements, streamlined validation\n- M: Standard planning and validation\n- L/XL: Comprehensive planning with enhanced validation (library plans, risk assessment, design patterns)\n\nAll tasks follow the unified workflow: CREATED \u2192 PLANNING \u2192 PLAN_APPROVED \u2192 IMPLEMENTING \u2192 REVIEW_READY \u2192 TESTING \u2192 COMPLETED",
      "enum": [
        "xs",
        "s",
        "m",
        "l",
        "xl"
      ],
      "title": "TaskSize",
      "type": "string"
    },
    "TaskState": {
      "description": "Coding task state enum with well-documented transitions.\n\nState Transitions:\n- CREATED \u2192 PLANNING: Task created, ready for planning phase (XS/S may skip to IMPLEMENTING)\n- PLANNING \u2192 PLAN_PENDING_APPROVAL: Plan created, awaiting user approval\n- PLAN_PENDING_APPROVAL \u2192 PLANNING: User requests plan changes\n- PLAN_PENDING_APPROVAL \u2192 PLAN_APPROVED: User approves plan\n- PLAN_APPROVED \u2192 IMPLEMENTING: Implementation phase started\n- IMPLEMENTING \u2192 IMPLEMENTING: Multiple code changes during implementation\n- IMPLEMENTING \u2192 REVIEW_READY: Implementation complete, ready for code review\n- REVIEW_READY \u2192 TESTING: Code review approved, ready for testing validation\n- TESTING \u2192 TESTING: Multiple test iterations\n- TESTING \u2192 COMPLETED: All tests validated; task completed successfully\n- Any state \u2192 BLOCKED: Task blocked by external dependencies\n- Any state \u2192 CANCELLED: Task cancelled\n- BLOCKED \u2192 Previous state: Unblocked, return to previous state\n\nUsage:\n- CREATED: Default state for new tasks, all tasks proceed to planning (unified workflow)\n- PLANNING: Planning phase in progress (set when planning starts)\n- PLAN_PENDING_APPROVAL: Plan created, awaiting user approval and potential iteration\n- PLAN_APPROVED: Plan validated and approved (set by judge_coding_plan)\n- IMPLEMENTING: Implementation phase in progress (set when coding starts)\n- REVIEW_READY: Implementation complete and ready for code review\n- TESTING: Testing/validation phase after code review approval\n- COMPLETED: Task completed successfully (set by judge_coding_task_completion)\n- BLOCKED: Task blocked by external dependencies (manual override)\n- CANCELLED: Task cancelled (manual override)",
      "enum": [
        "created",
        "planning",
        "plan_pending_approval",
        "plan_approved",
        "implementing",
        "testing",
        "review_ready",
        "completed",
        "blocked",
        "cancelled"
      ],
      "title": "TaskState",
      "type": "string"
    },
    "WorkflowGuidance": {
      "description": "Canonical workflow guidance model used across the system.\n\nReturned by tools to provide consistent next steps and instructions for\nthe coding assistant. This is the single source of truth for the\nWorkflowGuidance schema.",
      "properties": {
        "next_tool": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Next tool to call, or None if workflow complete",
          "title": "Next Tool"
        },
        "reasoning": {
          "default": "",
          "description": "Clear explanation of why this tool should be used next",
          "title": "Reasoning",
          "type": "string"
        },
        "preparation_needed": {
          "description": "List of things that need to be prepared before calling the recommended tool",
          "items": {
            "type": "string"
          },
          "title": "Preparation Needed",
          "type": "array"
        },
        "guidance": {
          "default": "",
          "description": "Detailed step-by-step guidance for the AI assistant",
          "title": "Guidance",
          "type": "string"
        },
        "research_required": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether research is required for this task (only determined for new CREATED tasks)",
          "title": "Research Required"
        },
        "research_scope": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Research scope: 'none', 'light', or 'deep' (only determined for new CREATED tasks)",
          "title": "Research Scope"
        },
        "research_rationale": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Explanation of research requirements (only determined for new CREATED tasks)",
          "title": "Research Rationale"
        },
        "internal_research_required": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether internal codebase analysis is needed (only determined for new CREATED tasks)",
          "title": "Internal Research Required"
        },
        "risk_assessment_required": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether risk assessment is needed (only determined for new CREATED tasks)",
          "title": "Risk Assessment Required"
        },
        "design_patterns_enforcement": {
          "anyOf": [
            {
              "type": "boolean"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "description": "Whether design patterns are required (only determined for new CREATED tasks)",
          "title": "Design Patterns Enforcement"
        },
        "plan_required_fields": {
          "description": "Structured specification of required fields for judge_coding_plan tool",
          "items": {
            "$ref": "#/$defs/PlanRequiredField"
          },
          "title": "Plan Required Fields",
          "type": "array"
        }
      },
      "title": "WorkflowGuidance",
      "type": "object"
    }
  },
  "properties": {
    "approved": {
      "description": "Whether the validation passed",
      "title": "Approved",
      "type": "boolean"
    },
    "required_improvements": {
      "description": "List of required improvements if not approved",
      "items": {
        "type": "string"
      },
      "title": "Required Improvements",
      "type": "array"
    },
    "feedback": {
      "description": "Detailed feedback about the validation",
      "title": "Feedback",
      "type": "string"
    },
    "suggested_diff": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Unified Git diff patch with suggested changes (optional). Provide when rejecting with concrete fixes or when proposing minor refinements.",
      "title": "Suggested Diff"
    },
    "reviewed_files": {
      "description": "Per-file reviews. Must include an entry for every file changed in the diff.",
      "items": {
        "$ref": "#/$defs/FileReview"
      },
      "title": "Reviewed Files",
      "type": "array"
    },
    "current_task_metadata": {
      "$ref": "#/$defs/TaskMetadata",
      "description": "ALWAYS current state of task metadata after operation"
    },
    "workflow_guidance": {
      "anyOf": [
        {
          "$ref": "#/$defs/WorkflowGuidance"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "LLM-generated next steps and instructions from shared method"
    }
  },
  "required": [
    "approved",
    "feedback"
  ],
  "title": "JudgeResponse",
  "type": "object"
}

Notes

  • Use after judge_code_change is approved. Follow workflow_guidance.next_tool for the next step.

  • Always use the exact task_id; recover it via get_current_coding_task if missing.

  • If test_files is empty or test_execution_results does not look like raw runner output, this tool will return approved: false and request real evidence (copy/paste the test run output and list the test files).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
task_idYes
test_filesYes
test_summaryYes
manual_test_notesNo
testing_frameworkNo
test_coverage_reportNo
test_execution_resultsYes
test_types_implementedNo
performance_test_resultsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
approvedYesWhether the validation passed
feedbackYesDetailed feedback about the validation
reviewed_filesNoPer-file reviews. Must include an entry for every file changed in the diff.
suggested_diffNoUnified Git diff patch with suggested changes (optional). Provide when rejecting with concrete fixes or when proposing minor refinements.
workflow_guidanceNoLLM-generated next steps and instructions from shared method
current_task_metadataNoALWAYS current state of task metadata after operation
required_improvementsNoList of required improvements if not approved
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full transparency burden and does disclose important behaviors: it requires real test evidence, returns approved:false when evidence is missing, and warns about token inefficiency if skipped. It does not explicitly describe side effects on task metadata, though the embedded response schema implies current_task_metadata is updated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The prose is organized and front-loaded, but the description is severely bloated by embedding the entire JudgeResponse JSON schema in the Returns section even though an output schema is already provided. Critical warnings and task_id guidance are also repeated across sections, making the description longer than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with 9 parameters and a rich output schema, the description is complete: it covers when to call it, prerequisites, required evidence, parameter semantics, failure behavior, and task_id recovery. The embedded output schema is redundant but does not leave the agent without needed context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The bare input schema has 0% description coverage, but the description's Args section fully compensates by explaining each parameter's meaning and providing concrete examples, such as 'Raw test runner output (required). For example, pytest/jest/mocha/go test/JUnit logs including pass/fail counts.' This adds substantial value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Validate test quality, coverage, and execution results after code review is approved.' It also clearly distinguishes this from sibling judge tools by tying invocation to workflow_guidance.next_tool and to appearing after judge_code_change is approved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit trigger condition ('Called when workflow_guidance.next_tool == "judge_testing_implementation"'), states it must be used after judge_code_change is approved, and explains when the tool will reject input (empty test_files or non-raw test_execution_results). It also warns against skipping it and tells the agent to recover task_id via get_current_coding_task if missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/groxaxo/mcp-llm-router'

If you have feedback or need assistance with the MCP directory API, please join our Discord server