Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

create_agent_response_test_route

Create tests that validate an ElevenLabs agent's replies against success and failure examples, success conditions, and simulated user scenarios.

Instructions

Create Agent Response Test

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameNo
typeNo
environmentNo
chat_historyNo
evaluation_modelNo
failure_examplesNoNon-empty list of example responses that should be considered failures
parent_folder_idNo
success_examplesNoNon-empty list of example responses that should be considered successful
tool_mock_configNoSimulation/preview-side config: tools are identified by IDs, resolved to names at runtime.
dynamic_variablesNoDynamic variables to replace in the agent config during testing
success_conditionNo
success_conditionsNoList of prompts that evaluate whether the simulation was successful. If provided, all criteria are evaluated and merged into a final result. Capped at the maximum number of evaluation criteria.
simulation_scenarioNoDescription of the simulation scenario and user persona for simulation tests.
tool_mock_overridesNoTest-specific response mocks, keyed by tool ID. Applied ahead of the tool's shared mocks and only within this test. Only take effect for tools that are mocked (see tool_mock_config).
simulated_user_modelNo
simulation_max_turnsNoMaximum number of conversation turns for simulation tests.
tool_call_parametersNo
check_any_tool_matchesNo
simulation_environmentNo
from_conversation_metadataNo
conversation_initiation_sourceNoEnum representing the possible sources for conversation initiation.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false, but the description adds nothing beyond them. With 21 parameters covering simulation config, tool mocking, and evaluation criteria, the description should disclose side effects, whether tests are persisted, and what happens on repeated calls — none of which is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short, but this is under-specification rather than conciseness. A single restated phrase cannot front-load meaningful information for a tool this complex.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 21-parameter, nested-object, mutation tool with no output schema and no annotations covering its semantics needs a substantial description. The one-line title restatement leaves the agent with essentially no information needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43% across 21 parameters, so the description carries a heavy burden — and it provides zero parameter information. Fields like tool_mock_overrides, success_conditions, simulation_scenario, and from_conversation_metadata are complex and nested, yet the description explains none of them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Create Agent Response Test" is a verbatim restatement of the tool name/title, which is the definition of tautology. It does not specify what an 'agent response test' is, what it operates on, or how it differs from sibling tools like update_agent_response_test_route or create_agent_test_folder_route.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to use this tool versus the 100+ siblings, no prerequisites, no mention of the agent or workspace context required. The agent must infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools