Skip to main content
Glama

aisafety_suite

Generate a test suite outline from a policy summary to identify safety test cases. Use to structure evaluation tests based on policy requirements.

Instructions

Generate a test suite outline.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
policyYesPolicy summary

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

C2.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It only restates the action of generating an outline and reveals nothing about side effects, output format, or how the policy input is used. No behavioral context is provided beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no wasted words, making it concise in length. However, it is under-specified and lacks useful structural elements such as input, output, or usage context, so it reads more as an incomplete statement than a well-crafted tool description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having only one parameter and no output schema, the description leaves key questions unanswered: what format the outline takes, what sections it includes, and how the policy affects generation. It also does not address sibling tools or any return behavior, making it insufficient for fully guided invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes 'policy' as 'Policy summary' with 100% coverage, so the baseline is 3. The description adds no additional meaning about how the policy should be formulated or what role it plays in generating the outline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate') and resource ('test suite outline'), clearly indicating the tool's core function. It does not explicitly distinguish itself from siblings aisafety_status and aisafety_scan, but the resource noun implies a distinct deliverable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus aisafety_status or aisafety_scan. There is no mention of appropriate contexts, prerequisites, or exclusions, leaving the agent without decision support for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools