Skip to main content
Glama

calls_agent_duel

Start an agent-vs-agent VOICE test call: two AI voice agents share one LiveKit room — a 'caller' persona agent pursues a task brief against the 'callee' business agent under test. Use to evaluate booking flows, latency, and conversation quality without a human caller.

The CALLER agent should have an EMPTY voice_greeting (it must stay silent until the callee greets) and voice_filler_enabled=false. Afterwards inspect both call_ids with agents.traces_list / calls.get_transcript. A subscribe-only listen token is returned for listening in live.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
taskYesThe caller's brief — objective, persona details (name, phone), and when to end the call. Woven into its prompt as call instructions.
in_workspaceNoRun this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.
max_duration_sNoHard cap on the call in seconds (30-900, default 300).
callee_agent_idYesAgent under test (answers and greets first). From agents.list.
caller_agent_idYesCustomer-persona agent that places the call. Must be a different agent, active, with empty voice_greeting.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / properties / in_workspace
      Added value: +{
      +  "description": "Run this one call in this workspace id instead of the session's. Nothing is stored; other sessions are not affected.",
      +  "type": "integer"
      +}
  2. Added
  3. Removed
  4. Added

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly=false, destructive=false, idempotent=false). The description adds real operational context the annotations don't carry: the caller must be a different, active agent with empty voice_greeting and voice_filler_enabled=false (it stays silent until greeted), plus a subscribe-only listen token for live monitoring.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then a second paragraph of prerequisites and follow-up. Dense but every sentence earns its place; minor tightening possible in the caller-config sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the full workflow: pre-call caller configuration, live monitoring via the returned listen token, and post-call inspection of both call_ids. With no output schema, the mention of returned tokens and call_ids is valuable; a bit more on what a completed duel produces would complete it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description reinforces the caller-agent constraints (empty greeting, filler disabled) but adds little syntactic meaning beyond the already-documented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Start an agent-vs-agent VOICE test call' with the two roles (caller persona vs callee business agent under test) and the LiveKit room mechanism spelled out. This clearly distinguishes it from single-agent siblings like calls_make and calls_dispatch_agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use to evaluate booking flows, latency, and conversation quality without a human caller' gives a clear when-to-use condition. It also names the follow-up tools (agents.traces_list / calls.get_transcript) but stops short of explicitly contrasting with the nearest siblings (calls_make, agents_simulate_inbound).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.