Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
speakA

Speak a short message aloud to the user through their speakers.

Use this to deliver a spoken TL;DR alongside (not instead of) your written answer: 1-3 conversational sentences summarizing the outcome, a finding, or a question. Never read code, file paths, or long explanations aloud.

You send only the text: voice, speed and language belong to the daemon (the user controls them on the dashboard). To deliberately switch your voice, call change_voice.

Concurrent speech is serialized: by default a new call WAITS for the current utterance to finish (queued), and for the user to finish speaking. Set interrupt=True to cut the current utterance off and speak immediately — use it only when your previous words are now stale (e.g. the user corrected you mid-answer).

Args: text: What to say. Plain conversational prose. Mark the key words the listener must catch with markdown bold (like this) — they get vocal emphasis and show bold on the live dashboard. Also supports inline speech tags like [pause] or [laugh] and wrapping tags like text. interrupt: Cut off any utterance currently playing and speak now. speaker: ONLY for subagents. If you are a subagent (Task/Agent tool), pass your role name here (e.g. "researcher") — the dashboard shows the message under that name with its own portrait, and the daemon gives you a stable voice distinct from the main agent's. The main agent must leave this empty.

announceA

Speak a quick spoken update WITHOUT waiting for it to finish.

Fire-and-forget: use this to tell the user what you just did and keep working ("done with X, moving on") — it returns immediately and plays in the background, queued behind any current speech. Use speak instead when you are asking a question or otherwise waiting for the user's reply. Like speak, it carries only text — voice/speed/language live in the daemon.

change_voiceA

Deliberately switch this agent's speaking voice from now on.

Updates your character in the listener daemon: the dashboard shows the new voice and every later speak/announce uses it (it also persists across restarts). Use list_voices to see the options. Speak itself carries no voice information — this call is the only way to change how you sound, so use it consciously (e.g. when the user asks for it).

Args: voice_id: Which voice to switch to. speaker: Move a named SPEAKER's voice instead of your own — the personas you address with speak(speaker=...). A voice already held by someone else is refused rather than duplicated, so two speakers never become indistinguishable by ear.

list_voicesA

List the Grok TTS voices available for the speak tool.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.5/5.0

Scored across 4 tools

Disambiguation4/5

The set is mostly distinct: speak delivers a blocking utterance, announce is fire-and-forget, change_voice and list_voices have clearly separate roles. speak and announce both produce speech, so their overlap could cause an agent to misselect when the blocking behavior matters, but the descriptions offer strong guidance.

Naming Consistency4/5

Naming is simple and readable with all verbs as commands, but it mixes one-word verb names (speak, announce) with verb_noun patterns (change_voice, list_voices). This minor inconsistency is not confusing and the style remains predictable.

Tool Count5/5

Four tools is well-scoped for a voice/speech server: each tool covers a necessary function—speaking, quick updates, voice switching, and voice enumeration. There is no bloat or obvious missing core capability for the stated purpose.

Completeness4/5

The tool set covers the main lifecycle of spoken interaction: speak, gently tell what you're doing, switch voices persistently, and discover available voices. One possible gap is lack of a way to query the currently active voice, but this is a minor issue that does not block common workflows.

Maintenance

ActivityActive
ResponsivenessResponsive