Skip to main content
Glama

Generate Playwright test from report

test_gen_from_report

Turn a classified bug report into a Playwright regression test and open a draft GitHub PR with the spec. Use to lock in a regression as an E2E test.

Instructions

Generate a Playwright regression test from a classified report using your project LLM key, then open a draft GitHub PR with the spec. Requires the inventory_v2 plan plus GitHub and LLM keys configured. Returns { qaStoryId, prUrl }. Write; consumes LLM budget; NOT idempotent — each call opens a new PR. Use to lock in a regression as an E2E test; use generate_tdd_from_story to build a test from a mapped user story instead.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
reportIdYesReport UUID to turn into a Playwright PR (`report_id` is accepted too.)
projectIdNoProject UUID — defaults to configured project (`project_id` is accepted too.)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.1.12
    • changedInput schema / properties / projectId / description
      Previous value: -"Project UUID — defaults to configured project"New value: +"Project UUID — defaults to configured project (`project_id` is accepted too.)"
    • changedInput schema / properties / reportId / description
      Previous value: -"Report UUID to turn into a Playwright PR"New value: +"Report UUID to turn into a Playwright PR (`report_id` is accepted too.)"
  2. Changed1 schema field changedv0.1.10
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
  3. First observedv0.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, and the description adds important context beyond that: it consumes LLM budget, is NOT idempotent, and each call opens a new PR. It also discloses the return shape ({ qaStoryId, prUrl }). This is meaningful behavioral disclosure that goes beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it states the action, the prerequisites, the return value, the side effects, and the alternative in three sentences. Every sentence earns its place, and the most important behavioral warnings (consumes LLM budget, not idempotent) are placed before the usage guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no output schema, the description covers the key facts an agent needs: what it does, what it requires, what it returns, and how it differs from the sibling. It doesn't describe the PR contents or failure modes, but the annotations plus description cover the critical operational context. A 4 is appropriate because it's complete enough for correct invocation, with minor gaps around error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds a small amount of value by noting that reportId is the input to turn into a PR and that projectId defaults to the configured project, but it doesn't add substantial meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate'), a resource ('Playwright regression test from a classified report'), and a concrete outcome ('open a draft GitHub PR with the spec'). It also distinguishes itself from the sibling generate_tdd_from_story by naming the alternative and its different input source, so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool ('Use to lock in a regression as an E2E test') and names the alternative ('use generate_tdd_from_story to build a test from a mapped user story instead'). It also states prerequisites (inventory_v2 plan, GitHub and LLM keys configured), which is clear context for when the tool is applicable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.