Transkriba — Russian audio transcription
Server Details
Transcribe Russian audio and video with timestamps and optional speaker labels from files or links.
- Status
- Healthy
- Uptime
- 100.0% over 22 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 7 tools
Tools are mostly distinct: create_upload handles local file upload, transcribe_file and transcribe_url both start transcription but differ by input source (uploaded vs public URL), which is clear from descriptions. get_transcript and list_transcripts both retrieve transcripts but differ by job_id lookup vs listing history, and search_transcripts is clearly distinct. Minor overlap between get_transcript and list_transcripts could cause slight confusion, but descriptions clarify.
All tools use snake_case with verb_noun pattern: create_upload, get_balance, get_transcript, list_transcripts, search_transcripts, transcribe_file, transcribe_url. Consistent and predictable. Minor deviation: transcribe_file and transcribe_url use verb_noun but transcribe is not a simple CRUD verb, still consistent with domain.
7 tools is well-scoped for a transcription service: upload, two transcription triggers, status/result retrieval, listing, search, and balance check. Each tool serves a distinct workflow step without redundancy.
Core lifecycle is covered: upload → transcribe → poll/get result, plus list and search. Missing operations like cancel/delete a job or delete transcripts are minor gaps; the main transcription workflow is complete. Balance check is a nice addition.
Available Tools
7 toolscreate_uploadПолучить ссылку для загрузки файлаAInspect
Путь для локального файла, у которого нет публичной ссылки. Передайте file_name и точный size_bytes. Возвращает upload_url, s3_key и обязательный content_type. Загрузите файл запросом PUT, передав Content-Type и Content-Length в точности из ответа (например, curl -X PUT -H 'Content-Type: application/octet-stream' -H 'Content-Length: 1048576' --data-binary @file.ogg ), после чего передайте s3_key и file_name в transcribe_file. Ссылка живёт 30 минут; один файл может занимать не более 256 МБ.
| Name | Required | Description | Default |
|---|---|---|---|
| file_name | Yes | Имя файла с расширением, например voice.ogg | |
| size_bytes | Yes | Точный размер файла в байтах, от 1 до 268435456. | |
| content_type | No | MIME-тип файла. По умолчанию application/octet-stream. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden—and it delivers. It discloses the return values, the required PUT request with exact Content-Type and Content-Length headers, the 30-minute link expiry, the 256 MB size limit, and the subsequent transcribe_file call. This gives an agent everything needed to understand the operation's behavior and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, required inputs, return values, upload procedure with a concrete curl example, expiry, size limit, and follow-up step. Information is front-loaded and flows logically from preparation through upload to transcription.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description manually enumerates the return fields and gives enough operational detail (PUT method, headers, expiry, size cap, next step) for an agent to successfully invoke this tool and continue the workflow. Nothing essential is missing for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces that size_bytes must be exact and that file_name/size_bytes are the core inputs, but it adds little beyond the schema's own parameter descriptions. The mention of content_type is about the response rather than the input parameter, so no significant extra semantic value is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The title states 'Get a link for uploading a file' and the description confirms it is the route for local files without a public link, returning upload_url, s3_key, and content_type. This clearly distinguishes it from sibling tools like transcribe_url and transcribe_file, which operate on already-available URLs or transcripts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: for a local file that has no public link. It also outlines the next step (pass s3_key and file_name to transcribe_file), giving clear context. However, it does not explicitly name alternatives like transcribe_url for files that do have a public link, so it falls short of full when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_balanceОстаток минутAInspect
Показывает, сколько минут расшифровки осталось на балансе. Стоит спросить перед пачкой длинных записей, чтобы не упереться в нехватку на середине.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The verb 'Показывает' conveys a read-only, non-destructive query. It adds the practical caveat about checking before large batches, which is valuable behavioral context beyond the literal reading. Lacks explicit statements about permissions or exact return format, but for a simple balance check the risk is low.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences. The first states the core function and the second provides a practical usage hint. There is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only query with no output schema, the description sufficiently conveys what the agent needs to know: it returns the remaining balance in minutes and suggests when to call it. It does not explicitly describe the return payload, but the meaning is clear enough for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
This tool takes zero parameters, so the schema covers everything. The description does not need to elaborate on inputs; the baseline of 4 for a no-parameter tool is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool shows the remaining transcription minutes on balance ('Показывает, сколько минут расшифровки осталось на балансе'), a specific verb and resource. It is immediately distinguishable from sibling tools like get_transcript or transcribe_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to call it: before a batch of long recordings to avoid hitting the limit mid-process. It does not name alternatives explicitly, but sibling contexts make it clear this is the balance query.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptЗабрать расшифровкуAInspect
Возвращает состояние задачи по job_id, а когда она готова — текст с таймкодами и метками спикеров, если при запуске было передано diarize=true. Пока идёт расшифровка, отвечает status=processing: вызывайте повторно с паузой в несколько секунд. Время зависит от длины записи, очереди и запуска обработчика.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Формат ответа. По умолчанию text — реплики с таймкодами. | |
| job_id | Yes | Идентификатор из transcribe_url |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the operation is non-blocking and requires polling, and that the response format depends on the 'diarize' flag passed during creationfake. However, it does not disclose the structure of the response (e.g., fields like status, text, timestamps) beyond what is implied, and there is no mention of error behavior or rate limits. This is a moderate disclosure given the lack of annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long and extremely concise. The core function is front-loaded in the first sentence, and the polling instructions are in the second. Every clause carries information with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a simple polling endpoint with 2 parameters and no output schema, so the description covers the essential usage: how to poll, what to expect during processing, and that the result may include speaker labels. However, it lacks details on error handling or the exact response fields, but given the simplicity, it is adequately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already has 100% coverage, documenting job_id as 'Идентификатор из transcribe_url' and format with its default. The description adds context about how the 'diarize' flag affects the output, but does not elaborate on the format parameter's specific syntax beyond what the schema provides. This is a baseline 3 given the schema's thoroughness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to retrieve the state of a transcription task by job_idhare, and when ready, the transcript text with timestamps and speaker labels. However, it does not explicitly differentiate it from siblings like transcribe_url or transcribe_file, which are used to create tasks, so an agent might wonder about the relationship to those. The verb 'Возвращает' (returns) and resource 'состояние задачи' (task state) are specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: it is for polling the status of an async transcription task, and it explicitly advises to call repeatedly with a pause when status is 'processing'. It does not, however, name alternative tools for when to use instead, but given the sibling list (which includes creation tools), the intended use is fairly unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_transcriptsСписок последних расшифровокAInspect
Перечисляет последние расшифровки учётной записи: id, имя файла, длительность и состояние. Полезно, чтобы найти job_id вчерашней записи, не спрашивая человека.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Сколько записей вернуть, до 50 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the disclosure burden. It clearly indicates a non-mutating, account-scoped list operation and names the output fields, which is good. However, it omits ordering, pagination, and default-limit behavior, leaving some transparency gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence defines the operation and output fields, and the second provides a practical use case. Both sentences earn their place and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional parameter and no output schema, the description covers the essential return fields and gives a realistic invocation context. Minor details such as explicit sorting direction and pagination are not critical for a basic recent-items listing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: the only parameter, limit, is already described in the input schema ('Сколько записей вернуть, до 50'). The description adds no extra parameter meaning, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a concrete verb ('Перечисляет') and specifies the resource ('последние расшифровки учётной записи') plus the returned fields: id, file name, duration, and status. It is clear and actionable, though it does not explicitly differentiate itself from the sibling search_transcripts beyond the word 'последние'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete use case: finding the job_id of yesterday's recording without asking a human. This makes the intended context clear, but it does not state when to prefer alternatives like search_transcripts or get_transcript.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_transcriptsПоиск по расшифровкамAInspect
Ищет фразу по текстам последних 200 готовых расшифровок и возвращает совпадения с таймкодом и именем спикера. Так можно найти, где именно на трёхчасовом созвоне обсуждали сроки, не перечитывая всё.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Сколько совпадений вернуть, до 50 | |
| query | Yes | Что искать |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It clearly explains the search scope (last 200 ready transcripts) and the output highlights (timestamp and speaker name). It does not cover edge cases like exact matching or default limits, but the core behavior is transparent enough for a search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two compact sentences with no filler. The behavior is front-loaded, and the example adds practical value without bloating the definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter search tool, the description is nearly complete: it states the scope, the query behavior, and the key output fields. With no output schema, it could be more explicit about response structure and default limit behavior, but these are minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: both 'query' and 'limit' already have descriptions. The tool description adds some context about what is searched and returned, but it does not substantially enrich the meaning of the parameters beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Ищет фразу по текстам последних 200 готовых расшифровок') and identifies the resource and scope clearly. It is easily distinguished from siblings like list_transcripts (listing metadata) and get_transcript (fetching one transcript) because it focuses on content search across recent transcripts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete use case: finding where a topic was discussed in a long call without rereading everything. It implies when the tool should be used, though it does not explicitly mention when not to use it or name alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribe_fileРасшифровать загруженный файлAInspect
Переводит в текст файл, загруженный по ссылке из create_upload. Возвращает job_id сразу: результат забирается инструментом get_transcript. Списывает минуты с баланса по факту длительности записи.
| Name | Required | Description | Default |
|---|---|---|---|
| s3_key | Yes | Ключ из ответа create_upload | |
| diarize | No | Разделить запись по спикерам. По умолчанию false. | |
| language | No | Код языка, например ru или en. Без него определяется сам. | |
| file_name | Yes | Человеческое имя записи для списка расшифровок |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses the async pattern (returns job_id immediately), the follow-up mechanism (get_transcript), and a billing side-effect (minutes deducted by actual duration). It omits failure modes and error handling, but the non-obvious async and cost behaviors are valuable disclosures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the core operation, the async return behavior, and the billing implication. The purpose is front-loaded and there is zero filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description compensates by explaining that the return value is a job_id and directing the agent to get_transcript for the result. It covers the critical operational surface (input source, async pattern, cost) for a mutation tool; minor gaps remain around supported file formats and error conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (s3_key, file_name, diarize, language) are already documented in the schema. The description reinforces s3_key's origin in create_upload but adds no per-parameter detail beyond what the schema provides, landing at the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Переводит в текст') and resource ('файл, загруженный по ссылке из create_upload'), which pinpoints both the operation and its input source. The mention of the create_upload link distinguishes it from the sibling transcribe_url, which presumably handles URLs instead of uploaded files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly lays out a pipeline: create_upload supplies the file, transcribe_file launches the job, and get_transcript retrieves the result. It gives clear workflow context but does not explicitly state when to prefer this over transcribe_url or exclude alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribe_urlРасшифровать запись по ссылкеAInspect
Переводит аудио или видео по ссылке в текст с таймкодами. Передайте diarize=true, если нужно разделение по спикерам. Понимает публичные записи YouTube, VK Видео, Rutube, Дзен, OK, Castbox и SoundCloud. Возвращает job_id сразу: результат забирается инструментом get_transcript. Списывает минуты с баланса по факту длительности записи.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Ссылка на запись | |
| diarize | No | Разделить запись по спикерам. По умолчанию false. | |
| language | No | Код языка, например ru или en. Без него определяется сам. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden and does well: it reveals that the tool is asynchronous (returns job_id immediately), that the transcript is fetched via get_transcript, and that billing debits minutes based on recording duration. It does not mention failure modes or auth requirements, but the key behavioral traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each earning its place: the core function and output, the diarize option, supported platforms, and the async/billing behavior. The most important information is front-loaded, and there is no filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers the critical operational details: what it accepts, what it returns, how to retrieve the result, speaker separation behavior, supported sources, and cost implications. An agent has sufficient information to invoke it correctly and know what to expect next.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaningful context beyond the schema by enumerating supported URL sources (YouTube, VK, Rutube, etc.) and by explaining the practical effect of diarize=true for speaker separation. This is genuinely helpful for correctly choosing parameter values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action—converting audio/video from a URL into text with timestamps—and clearly distinguishes this from the sibling transcribe_file by emphasizing 'по ссылке' (by link). The name and title reinforce the resource type, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it explains when to set diarize=true, lists supported platforms as public recordings, and routes the user to get_transcript for retrieving the result. It does not explicitly state when not to use this tool versus transcribe_file, but the 'по ссылке' framing and platform list imply the distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
create_upload2 fields changed- added
Input schema / properties / size_bytesAdded value: +{ + "description": "Точный размер файла в байтах, от 1 до 268435456.", + "type": "number" +} - changed
Input schema / requiredPrevious value: -[ - "file_name" -]New value: +[ + "file_name", + "size_bytes" +]
- Changed
transcribe_file1 field changed- changed
Input schema / requiredPrevious value: -[ - "s3_key" -]New value: +[ + "s3_key", + "file_name" +]
7 tool updates
- First observed
create_upload - First observed
get_balance - First observed
get_transcript - First observed
list_transcripts - First observed
search_transcripts - First observed
transcribe_file - First observed
transcribe_url
Related MCP Connectors
Audio & video to text, Russian-first: diarization, timestamps, summary, action items, subtitles.
Transcribe audio & video: diarization, timed SRT/VTT, podcasts, paste-a-link, whole-feed batch.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Verbatim transcription of public video/audio URLs to clean text, SRT, and timestamped records.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceThis service provides fast and reliable transcriptions for audio/video files and voice memos. It allows LLMs to interact with the text content of audio/video file.8MIT
- AlicenseAqualityBmaintenance100% Free AI audio and video transcription with speaker diarization and YouTube support.511 npmMIT
- AlicenseNot gradedqualityCmaintenanceTranscribes YouTube videos or audio files to Markdown, plain-text, and Word documents.MIT
- AlicenseAqualityBmaintenanceEnables AI assistants to transcribe audio and video from URLs or local files with high accuracy, speaker diarization, 119 languages, and word-level timestamps, while also supporting transcription management and caption export in SRT, WebVTT, or plain text.14264 npm11MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.