Samplecut
Server Details
Cut audio from YouTube, TikTok or Instagram links to MP3/WAV, and fetch video transcripts.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 2 tools
cut_audio produces an audio clip while get_transcript fetches timestamped captions; the two purposes are orthogonal and the descriptions even explain how they compose (use transcript times to pick a cut range). There is no plausible way to confuse them.
Both names follow a clean verb_noun snake_case pattern (cut_audio, get_transcript) with no stylistic deviations.
Two tools is thin for the scope; a user might reasonably expect companion operations such as listing available media/metadata or validating a URL before cutting. Still, each tool earns its place and the pair covers the primary workflow without redundancy.
The core lifecycle (find the right time range via transcript, then produce a cut clip from a YouTube/TikTok/Instagram URL) is covered, including error handling for missing captions and expired links. Minor gaps exist: no format/media inspection, no batch or multi-segment cutting, and no way to re-fetch an expired link without re-invoking.
Available Tools
2 toolscut_audioCut audioAInspect
Cut audio from a YouTube, TikTok, or Instagram URL and return a short-lived download link. Lead with download_url for cut/download requests; offer editor_url when the user wants to tweak the selection. Links expire after about an hour — ask again in chat for a fresh cut if expired. Omit start_sec and end_sec for the full track. format defaults to mp3 (wav allowed).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube, TikTok, or Instagram URL to cut | |
| format | No | Download format; default mp3 | |
| end_sec | No | Range end in seconds | |
| start_sec | No | Range start in seconds | |
| duration_sec | No | Clip length in seconds from start_sec (default 0) when end_sec is omitted |
Output Schema
| Name | Required | Description |
|---|---|---|
| full | Yes | |
| title | Yes | |
| format | Yes | |
| end_sec | No | |
| start_sec | No | |
| editor_url | Yes | |
| expires_at | Yes | |
| download_url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds behavior beyond the annotations: links expire after about an hour and the user must re-request, format defaults to mp3 with wav as the alternative. These are meaningful operational facts an agent needs. It doesn't cover auth or rate limits, but annotations already declare openWorldHint=true and destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then response guidance, then expiry and parameter defaults. Every sentence carries information; it is slightly dense but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Return keys (download_url, editor_url) and link expiry are explained, and an output schema exists so return structure needn't be detailed. Covers the essentials for calling and acting on the result, with only auth/permission context absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still adds value by clarifying the omit-both-for-full-track interaction and the mp3 default/wav option, which reinforces the schema's conditional duration_sec semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('cut audio') plus the accepted source platforms and the primary return artifact (a short-lived download link). It is clearly distinct from get_transcript, but it never explicitly names or contrasts the sibling, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage direction: prefer download_url for cut/download requests, offer editor_url when the user wants to tweak the selection, and omit start_sec/end_sec for the full track. There is no explicit when-not or alternative-tool guidance (e.g., 'use get_transcript instead for text'), so it is clear context rather than full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptGet transcriptARead-onlyInspect
Fetch timestamped captions for a YouTube, TikTok, or Instagram URL when the platform provides them. Use segment start_sec/end_sec (or the flat text) to choose a cut range, then call cut_audio with the same URL and those times. When available is false there are no captions — do not invent lyrics or timestamps; tell the user to pick times another way. Optional lang is a BCP-47 code (default en). No job_id is returned.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | YouTube, TikTok, or Instagram URL to fetch captions for | |
| lang | No | Optional BCP-47 caption language (default en) |
Output Schema
| Name | Required | Description |
|---|---|---|
| lang | Yes | |
| text | Yes | |
| source | Yes | |
| segments | Yes | |
| available | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so safety is covered. The description adds genuinely new behavior beyond the annotations: caption availability is conditional on the platform, the available=false case must not be hallucinated around, and lang defaults to en. The 'No job_id is returned' note is useful but only marginally so given an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the chaining instruction, then the failure guard, then two small footnotes. No sentence is filler; each adds a distinct operational fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Beyond the schema and annotations, the description supplies the workflow handoff to cut_audio and explains how to consume the returned segments, which is exactly what an agent needs for a step in a multi-tool pipeline. An output schema exists, so return-value documentation is correctly not duplicated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both url and lang are already documented at the same level the description restates (BCP-47, default en). The start_sec/end_sec references describe output segments rather than input parameters, so they help context but not parameter semantics. Baseline 3 is correct when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (fetch timestamped captions) plus the scope of supported platforms (YouTube, TikTok, Instagram) and the conditional 'when the platform provides them'. It also positions itself in the workflow relative to the only sibling, cut_audio, so an agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use segment start_sec/end_sec to choose a cut range, then call cut_audio with the same URL and those times. It also defines the failure path — when available is false, do not invent lyrics or timestamps and tell the user to pick times another way. This is when-to-use, how-to-chain, and when-not behavior in one place.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- First observed
cut_audio - First observed
get_transcript
Related MCP Connectors
Get transcripts from YouTube, TikTok, X, Instagram and more - even when captions are off.
Transcribe public videos & audio (YouTube, TikTok, IG) into accurate, timestamped text via API.
Timestamped transcripts, chapters and clip suggestions from YouTube, Twitch, Kick or TikTok links.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Related MCP Servers
- AlicenseAqualityCmaintenanceTranscribes videos from 1000+ platforms (YouTube, TikTok, Vimeo, etc.) and local video files using OpenAI's Whisper model, with support for 90+ languages and multiple output formats.836 npm7MIT
- AlicenseNot gradedqualityCmaintenanceTranscribes YouTube videos or audio files to Markdown, plain-text, and Word documents.MIT
- AlicenseBqualityDmaintenanceEnables downloading videos from platforms like YouTube and converting them to text using OpenAI Whisper and ffmpeg. It supports multiple output formats including TXT, JSON, SRT, and VTT for transcriptions.29 npmISC
- FlicenseNot gradedqualityCmaintenanceEnables extracting timestamped transcripts from YouTube and TikTok URLs, including title, channel, and duration, via a single tool.-
Glama MCP Gateway
Add one secure layer between your agents and this server.