Skip to main content
Glama

open_video_call

One-call bootstrap for 'video-call me' — talk to the human out loud, face to face. Same as open_remote_control (mints a private trusted channel + your identity + the human's identity + PIN + the pre-formed listener/reply commands), but ALSO returns a call_url: a meet.apuchat.com/call link that opens a Google-Meet-style video-call UI where your replies are spoken aloud and the human talks back by voice. Use when the human says 'video-call me', 'let me talk to you', 'call me', 'I want to speak out loud', 'talk to you like a person', or similar. YOUR side is IDENTICAL to a phone remote: you join and receive/reply plain TEXT — the human's speech is transcribed to text in their browser, and your text replies are spoken aloud in their browser. No audio/video flows through you; it stays a text channel underneath (max 8192 chars/msg). After this call: (1) join with the returned channel_id + token + agent.identity_key + owner_password; (2) arm receive with receiver_command_template (+ monitor_command_template or waiter_command_template); (3) run selftest_command_template; (4) relay operator_handoff_video to the human VERBATIM (it leads with a QR-page link + the one-tap call_url + the PIN-protected call_url_protected + the PIN). On each wake fire a send with kind:'status' first (the call shows an 'agent is working…' pose), then reply with reply_command_template.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
session_tokenNoOptional. Pass an account's session_token to attach the new channel to that account (shows up in /account). Otherwise an anonymous account is minted and a recovery_token is returned.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses behavioral traits: the text-only nature, no audio/video flowing through the agent, max message length (8192 chars), the exact steps to follow after the call, and the returned call_url artifacts. This goes far beyond what annotations would provide and gives the agent a complete operational picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured, with a front-loaded purpose and a numbered post-call procedure. Each sentence adds necessary details, though the density of procedural steps could be condensed slightly. Overall, it earns its length but is not maximally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description thoroughly explains the return values (channel_id, token, call_url, templates) and provides actionable steps for using them. It covers the full lifecycle from initiation to sending replies, making it complete for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single parameter session_token, and its description is already sufficient. The tool description does not add any additional meaning about the parameter, so the baseline 3 applies since the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear, specific purpose: 'One-call bootstrap for video-call me' — creating a video-call channel and returning a call_url. It explicitly distinguishes itself from the sibling tool open_remote_control by adding the video-call UI, making the function's unique role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: 'Use when the human says video-call me, let me talk to you, call me, I want to speak out loud...' and contrasts with open_remote_control, clearly stating the condition where this tool is appropriate and how it differs from the alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.