justcaptions
Server Details
Audio transcription, caption correction, translation and 15 animated video caption styles.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- SpaceSheepBoy/justcaptions-skill
- GitHub Stars
- 0
- Server Listing
- justcaptions
TDQS
Scored across 8 tools
Most tools target a distinct operation: transcribe, translate, correct, emoji-pick are all clear caption transforms. Minor overlap exists among the cost-related trio (estimate_cost, get_pricing, get_usage) which could occasionally be confused, but their descriptions distinguish them.
All 8 tools follow a consistent snake_case verb_noun pattern (correct_captions, estimate_cost, get_pricing, list_styles, pick_emojis, transcribe_audio). No deviations or mixed conventions.
8 tools is well-scoped for a caption processing service, with each tool covering a distinct step (transcribe, translate, correct, emoji, styles, pricing, cost, usage). No redundant or filler tools.
The surface covers the full caption pipeline plus cost/usage/pricing introspection, which is strong. Rendering is delegated to a local MCP, so no export/render tool exists here, a minor but intentional gap agents can work around.
Available Tools
8 toolscorrect_captionsBInspect
Correct caption text, preserving caption count and order. Does not realign word timing. Billed by input characters.
| Name | Required | Description | Default |
|---|---|---|---|
| captions | Yes | ||
| glossary | No | ||
| language | No | ||
| request_id | No | Unique request ID; reuse only when retrying identical work to avoid repeating paid calls. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false and destructiveHint=false, so the safety profile is partly covered. The description adds genuinely non-structured context — billing is per input character and word timing is left untouched — but says nothing about auth needs, rate limits, or how much the text itself may change.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying a distinct fact (action, scope limit, cost model), with the core contract front-loaded and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core text contract and billing model are covered for a 4-parameter mutation-style tool with no output schema. However, with 3 of 4 parameters undocumented and no explanation of what 'correct' actually alters or what comes back, an agent can invoke it but not fully predict the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (request_id is the sole documented field), yet the description describes no parameters at all. glossary and language are left completely undefined in both schema and description, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Correct caption text') and constrains the contract ('preserving caption count and order'), which immediately separates it from translate_captions and transcribe_audio. It stops short of naming a sibling explicitly, so an agent must still infer routing from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no named alternative. 'Does not realign word timing' implies a boundary against timing/alignment workflows, but it is a scope constraint rather than a comparison with the sibling tools an agent must choose between.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
estimate_costARead-onlyIdempotentInspect
Estimate additional cloud charges and available allowance before a batch. Includes active budget holds; does not reserve credit.
| Name | Required | Description | Default |
|---|---|---|---|
| text_chars | No | ||
| audio_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds genuinely useful non-obvious behavior: it includes active budget holds but does not reserve credit, which is exactly the kind of context an agent needs before a batch run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and followed by the key behavioral caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Behaviorally it is complete enough (no output schema, annotations cover safety), but with two parameters at 0% schema coverage and no output schema, the description leaves the caller without any indication of what the inputs mean or what the estimate returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate, but it never mentions text_chars or audio_seconds, nor their units or maximum bounds. The param names are self-explanatory enough to avoid a 1, but the gap is real at zero coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states a specific verb and resource: estimate cloud charges plus available allowance. An agent can distinguish it from get_pricing/get_usage by the 'before a batch' estimation framing, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'before a batch' implies the usage context, but the description never states when to prefer this over get_pricing or get_usage, nor any exclusion conditions. Usage is inferable but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pricingARead-onlyIdempotentInspect
Current API prices and monthly free allowances, no key needed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description still adds genuine value beyond those structured fields by disclosing that no API key is required and that the values are 'current' rather than cached, which affects how the agent invokes it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the resource stated first and the auth caveat trailing; no filler or redundancy. It is appropriately sized for a trivial no-arg lookup, though it is almost too terse to route confidently against sibling tools.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully names what comes back (prices plus monthly free allowances) and removes the main invocation ambiguity with the no-key note. For a zero-parameter read tool this is nearly complete, with only explicit sibling routing missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the schema coverage is 100%, so there is nothing for the description to clarify. Baseline 4 applies for a no-argument tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb + resource (prices, monthly free allowances) so the agent knows this returns pricing information rather than usage or cost estimates. It does not explicitly name or differentiate itself from the similar siblings estimate_cost and get_usage, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the static price-list lookup versus the per-request estimate_cost or past-consumption get_usage, but the description never states when to prefer this tool or name an alternative. 'No key needed' hints that it can be called without credentials, which is a precondition rather than a when-to-use rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usageARead-onlyIdempotentInspect
Current usage, estimated bill and monthly spend cap for your configured API key.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful context by enumerating what the response contains (usage, estimated bill, monthly spend cap), which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every clause carries information about scope or return content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool with no output schema, the description adequately conveys both the resource scope and the shape of the returned data. It could be slightly stronger by noting the billing window or that values reflect the key's own configuration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to explain. The baseline of 4 applies; nothing in the description misleads about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (current usage, estimated bill, monthly spend cap) scoped to the configured API key, which clearly separates it from siblings like estimate_cost and get_pricing. It lacks an explicit verb ('get'), but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the closely related estimate_cost or get_pricing siblings, and no prerequisites or exclusions are given. The agent must infer the distinction purely from the noun phrase.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_stylesARead-onlyIdempotentInspect
List 15 local rendering presets, aliases, style overrides and platform safe areas. This remote server returns the catalog; render via the local MCP.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely non-obvious behavioral context — that the remote server returns only the catalog and that rendering must be performed through the local MCP — which prevents an agent from expecting this call to apply a style.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the resource inventory front-loaded and the remote/local split stated second. Every clause carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool with no output schema, the description still conveys the shape and volume of the result (15 presets plus aliases, overrides, safe areas) and clarifies the remote-vs-local division of labor. Nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a 0-param tool is 4. No syntax or format details are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("List") and enumerates the exact resource contents: 15 local rendering presets, aliases, style overrides, and platform safe areas. None of the siblings (caption/audio/pricing tools) overlap, so an agent can identify this as the sole style-catalog tool without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description draws a clear boundary between this call and the actual work: this remote server only returns the catalog, while rendering happens "via the local MCP." That is actionable routing guidance, though it stops short of explicitly stating when this tool should be preferred over other catalog or pricing calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pick_emojisBInspect
Pick one emoji for each caption. Billed by input characters.
| Name | Required | Description | Default |
|---|---|---|---|
| captions | Yes | ||
| language | No | ||
| request_id | No | Unique request ID; reuse only when retrying identical work to avoid repeating paid calls. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly=false, destructive=false, openWorld=true, idempotent=false), so the bar is lower. The description adds a genuinely useful cost trait, 'Billed by input characters,' plus the one-emoji-per-caption output mapping, but says nothing about rate limits, latency, or what happens on failed items.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action front-loaded and the billing constraint second. Nothing can be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and three parameters, the description leaves the return shape (e.g., an array aligned to captions) and the meaning of 'language' unstated, though the billing model and cardinality are covered. Adequate for a simple tool, but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (just request_id), so captions and language are undocumented in the schema. 'One emoji for each caption' does clarify the input-to-output cardinality of the captions array, but language is left entirely unexplained in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (pick an emoji) and even the output cardinality (one per caption), which cleanly separates it from siblings like translate_captions or correct_captions. It never names those siblings or an operational boundary, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not, or alternative named anywhere. The only guidance-adjacent sentence is a pricing note, not a routing condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribe_audioAInspect
Transcribe extracted audio into segments and word timestamps. Max 12 MB decoded. Never send video. Charged by audio duration past the free allowance.
| Name | Required | Description | Default |
|---|---|---|---|
| glossary | No | ||
| language | No | ||
| mime_type | Yes | ||
| request_id | No | Unique request ID; reuse only when retrying identical work to avoid repeating paid calls. | |
| audio_base64 | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations declaring readOnlyHint=false and idempotentHint=false, the description still adds genuine value: the 12 MB decoded input cap, the pricing model (charged by audio duration beyond a free allowance), and the video exclusion. It does not describe failure behavior on oversized input or what happens to partial results, which keeps it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the operation and output, then constraints, then cost. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully names the return structure (segments + word timestamps) and covers the key cost and size constraints for a paid, open-world tool. It falls short on parameter-level meaning and error handling for the low-coverage schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (just request_id), so the description must compensate for glossary, language, mime_type, and audio_base64 — and it does not mention any of them. Only the size limit loosely touches audio_base64; glossary and language semantics are entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (transcribe), the resource (extracted audio), and the return shape (segments and word timestamps). This clearly separates it from siblings like translate_captions and correct_captions, which operate on already-produced captions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the prerequisite that audio must already be extracted ("extracted audio", "Never send video") and a hard size ceiling, but it names no sibling alternative and gives no when-not-to-use guidance beyond the video exclusion. Usage is implied rather than explicitly routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
translate_captionsAInspect
Translate caption lines, preserving count and order. Target-language word timing is estimated by the local renderer. Billed by input characters.
| Name | Required | Description | Default |
|---|---|---|---|
| captions | Yes | ||
| glossary | No | ||
| request_id | No | Unique request ID; reuse only when retrying identical work to avoid repeating paid calls. | |
| target_language | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond the annotations (readOnlyHint=false, openWorldHint=true, idempotentHint=false) by disclosing that timing is estimated by the local renderer and that the call is billed by input characters, which is valuable pre-call context. It does not cover failure behavior or whether output timing is approximate per-line, but the operational disclosure is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, with the core behavior (translate, preserve count/order) front-loaded and secondary details (timing estimate, billing) following. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey return behavior; it states count/order preservation and estimated timing, which helps. However, for a non-idempotent, billable, open-world tool with a 4-param schema at 25% coverage, the undocumented glossary and unsettled target_language format leave meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only request_id is documented), so the description carries the burden. It adds value for the captions parameter (count/order preserved) and ties billing to input characters, but the glossary parameter is left completely undocumented in both schema and description, and target_language format is unspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Translate caption lines') with the scope constraint that count and order are preserved. It is clearly a translation operation, though it does not distinguish itself from siblings like correct_captions or explain why one would translate here rather than elsewhere.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated — you use it to translate captions — and the mention of billing by input characters hints at cost-related siblings (estimate_cost, get_pricing) without naming them. No explicit when-to-use or when-not-to-use guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
- First observed
correct_captions - First observed
estimate_cost - First observed
get_pricing - First observed
get_usage - First observed
list_styles - First observed
pick_emojis - First observed
transcribe_audio - First observed
translate_captions
Related MCP Connectors
Transcribe, subtitle and dub videos into 100+ languages, and translate text.
Transcribe podcasts, YouTube and audio, then make show notes, chapters and translations.
1AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceThis service provides fast and reliable transcriptions for audio/video files and voice memos. It allows LLMs to interact with the text content of audio/video file.8MIT
- AlicenseAqualityBmaintenance100% Free AI audio and video transcription with speaker diarization and YouTube support.511 npmMIT
- AlicenseAqualityCmaintenanceTranscribes videos from 1000+ platforms (YouTube, TikTok, Vimeo, etc.) and local video files using OpenAI's Whisper model, with support for 90+ languages and multiple output formats.836 npm7MIT
- AlicenseAqualityBmaintenanceEnables AI assistants to transcribe audio and video from URLs or local files with high accuracy, speaker diarization, 119 languages, and word-level timestamps, while also supporting transcription management and caption export in SRT, WebVTT, or plain text.1488 npm11MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.