Skip to main content
Glama

Submit audio for transcription + diarization (member)

ic_transcribe_submit

Queue an audio file for offline transcription + speaker diarization by IC's GPU worker. Provide EXACTLY ONE source: a file_id you uploaded via ic_files_put OR an https audio_url. SIZE: ic_files_put caps at ~3.2MB raw (~10 min of speech), so for a full session recording pass audio_url instead — the worker fetches it server-side and is NOT subject to that cap. Results (a markdown transcript + a JSON with per-speaker segments) land in the file vault next to the source; poll ic_transcribe_status no more than once per minute, then read the transcript with ic_transcribe_get. Args: { file_id?, audio_url?, language? (BCP-47 hint, e.g. 'en'), num_speakers_hint? (1..10) }. Returns: { ok, id, status: 'queued', queue_position }. Rate: 5 submissions per token per UTC day. Required scope: transcribe:submit (ic-member+).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
file_idNoA vault file id (f_...) from ic_files_put. Mutually exclusive with audio_url.
languageNoOptional BCP-47 language hint (e.g. 'en', 'es'). Omit to auto-detect.
audio_urlNoAn https URL to the audio. Mutually exclusive with file_id. No private/loopback hosts.
num_speakers_hintNoOptional hint for how many speakers to diarize (1..10).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (which only indicate non-idempotent, non-destructive), the description discloses that the tool queues a job, returns a status of 'queued' with a queue_position, and mentions the rate limit (5 per token per UTC day) and required scope (transcribe:submit). It also states that results land in the file vault, providing clear behavioral expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is verbose and somewhat redundant. The 'Args:' line at the end duplicates the parameter information already present in the schema and earlier prose. The size constraint and source requirements are stated multiple times. While not too long, it could be tightened to eliminate repetition, making it more concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by explicitly listing the return shape: { ok, id, status: 'queued', queue_position }. It also details the full workflow (queue, poll status, get transcript) and mentions rate limits and scope. For a tool that initiates asynchronous work, it provides sufficient context for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers 100% of parameters with descriptions. The tool description adds crucial semantic context: mutual exclusivity of file_id and audio_url, the size cap (3.2MB) for file_id, the necessity of https for audio_url, and the language hint format. This goes beyond the schema and significantly clarifies parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: queue audio for offline transcription and diarization. It uses specific verbs like 'Queue' and 'Provide' and explicitly names the source options (file_id or audio_url). It distinguishes itself from sibling tools by referencing ic_transcribe_status and ic_transcribe_get for polling and retrieval, ensuring the agent knows when to use this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Guidelines are explicit: provide exactly one source, explains the size cap for file_id vs audio_url (use audio_url for larger files), and instructs to poll ic_transcribe_status no more than once per minute and retrieve with ic_transcribe_get. This clarifies when to use this tool versus alternatives, including the prerequisite of using ic_files_put to obtain a file_id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.7/5.0
Disambiguation4/5

Most tools are clearly scoped to distinct actions (e.g., ic_hack_apply vs. ic_hack_register, ic_rooms_create vs. ic_rooms_join). A few pairs could confuse an agent: floor10_submit_highlight vs. floorcast_push both submit HighlightStories but to different queues, and ic_directory_search / ic_agent_directory_lookup / ic_admin_list_members overlap in searching members. Overall, the long descriptions help disambiguate, but the volume requires careful reading.

Naming Consistency3/5

The dominant pattern is ic_<domain>_<verb>_<object> (e.g., ic_admin_list_pending_events, ic_headsets_checkout), but there are notable deviations: floor10_* and floorcast_* prefixes break the ic_ convention, and a few tools use noun-style names (ic_health, ic_capabilities, ic_donations_total). Verb placement also varies (get_* vs *_get, e.g., ic_get_my_membership vs. ic_membership_set_profile). Still, most names are readable and predictable.

Tool Count1/5

175 tools is an extreme count for a single MCP server, far beyond the 50+ threshold that indicates an unwieldy surface. While the platform covers many domains (events, files, hackathon, headsets, prints, rooms, etc.), bundling everything into one server makes discovery and selection difficult. This would be better split into several narrowly-scoped servers.

Completeness4/5

The tool set covers nearly every lifecycle for each domain: CRUD for files/folders, full hackathon admissions and judging, headset lending with waivers and incidents, print farm submission and handoffs, and room coordination. Minor gaps exist: no delete for files/folders, no cancel for events, and some actions (like revoking a Z.ai key or tearing down a room) are explicitly left to human console use. Overall, the surface is remarkably comprehensive for the stated scope.