Skip to main content
Glama

Text to speech

parserail_speak

Convert text to natural-sounding speech audio for voice notes, IVR lines, and narration. Generate embeddable audio files, with credit costs from your account wallet.

Instructions

Text → natural speech audio, ready to embed. Voice notes, IVR lines, narration. Costs credits from the account wallet.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
paceNo
textYes
voiceNoVoice name; omit for the default narrator.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.5

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a meaningful behavioral consequence beyond annotations: it costs credits from the account wallet, which is important for an agent deciding whether to invoke it. It also indicates the output is embeddable audio. Annotations are readOnlyHint=false, consistent with a side-effecting operation, and there is no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core transformation is stated first, followed by use cases and a key side effect. Every sentence adds value and there is no redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter TTS tool, the description covers the essential context: input text, output audio, example uses, and cost. There is no output schema, so a bit more detail about the return format (e.g., audio URL or file) could help, but the 'ready to embed' phrase partially covers this.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, with only the voice parameter described in the schema. The description does not explain pace or text semantics, and while 'text' is obvious, the tool description adds no guidance on how parameters affect output. The low coverage requires more compensation than this provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly defines a specific function: converting text into natural speech audio, with concrete example use cases (voice notes, IVR lines, narration). This distinguishes it from the sibling parserail tools, which focus on text processing, parsing, and analysis rather than audio generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: when spoken audio is needed from text. It gives example applications and notes the credit cost, which helps the agent decide whether this is the right tool. It does not explicitly name alternatives or exclusions, but no other sibling appears to offer TTS.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.