Skip to main content
Glama

bg3_test_run

Run one auto or scripted BG3 test case end to end: stage real combat, perform scripted cast, verify effects and costs, then clean up. For player-mode cases, use staging plus a hotbar cast instead.

Instructions

Run one auto/script case end to end: stage (real combat, host acts first), scripted cast, verify, cleanup. Effects come from recorded events; costs are checked against the spell's loaded UseCosts (a scripted cast never charges them). player-mode cases need bg3_test_stage + a hotbar cast instead.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
waitNo
layerYes
layersNo
case_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and delivers substantial behavioral detail: real combat with host acting first, scripted casting, verification, cleanup, effects sourced from recorded events, and costs checked against the spell's loaded UseCosts without charging them. It still omits permissions, side effects, and output behavior, but the execution semantics are unusually transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loads the main action and pipeline before adding caveats about effects, costs, and player-mode cases. It avoids repetition and wastes little space, though the dense semicolon structure slightly reduces readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a test-runner tool with no annotations, four undocumented parameters, and many siblings, the description covers the high-level process and one important alternative path. It does not explain what layer, layers, or wait mean, leaving a meaningful gap for correct invocation. Output schema exists, so return-value explanation is not required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description must compensate for layer, layers, wait, and case_id. It only obliquely implies case_id by referring to one auto/script case, and says nothing about layer, layers, or wait. Most parameter meaning remains undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: running one auto/script case end to end, with stage, scripted cast, verify, and cleanup. It also distinguishes this from player-mode cases, which need bg3_test_stage plus a hotbar cast instead. However, it does not explicitly contrast with sibling tools like bg3_test_script or bg3_test_run_level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for using this tool on auto/script cases and explicitly names an alternative path for player-mode cases: bg3_test_stage plus a hotbar cast. It does not cover when to choose this over bg3_test_script or bg3_test_run_level, so guidance is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.