Skip to main content
Glama

Transcription

parserail_transcribe

Transcribe audio recordings into accurate text with speaker labels and timestamps for meetings, calls, and voice notes. Uses account wallet credits only on successful processing.

Instructions

An audio recording → accurate text with speakers and paragraph timestamps. Meetings, calls, voice notes. Costs credits from the account wallet.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
audioNo
diarizeNoLabel speakers.
audioUrlNo
languageNoBCP-47 hint, e.g. "en".

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.5

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

It discloses a meaningful side effect beyond the annotations: 'Costs credits from the account wallet,' which is important given readOnlyHint=false. It also reveals useful output characteristics (speakers, paragraph timestamps). No contradiction exists with the annotations, though auth requirements or failure behavior are not covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tightly written and front-loaded with the core transformation. Use cases and the cost side effect are conveyed in minimal words with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no required parameters and no output schema, the description partially covers what the tool produces, including speakers and timestamps. However, it does not clarify whether audio must be passed inline via base64 or can be a URL, nor what happens when both are provided, so an agent must infer the calling convention from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%, with audioUrl and mimeType lacking descriptions. The description's mention of 'audio recording' only weakly hints at the input options and does not explain the relationship between the audio object and audioUrl, the purpose of language, or the base64 format. It adds little beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a clear action: transcribing an audio recording into text, with specific output features like speakers and paragraph timestamps. It lists concrete use cases (Meetings, calls, voice notes), but does not explicitly differentiate from sibling tools such as parserail_minutes or parserail_summarize.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The listed use cases provide clear context for when to use this tool, conveying it is for audio content. However, it does not mention when not to use it or name alternative sibling tools, so it falls short of explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.