io.github.AudialAI/audial-mcp
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@io.github.AudialAI/audial-mcpsplit the vocals and drums from ~/Music/song.mp3"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
audial-mcp
Run Audial's hosted audio tools from any MCP client: split stems, analyze and segment audio, master tracks, build sample packs, convert audio to MIDI, generate music, turn a one-shot into an editable Audial Synth preset, and sing lyrics in a reference voice.
mcp-name: io.github.AudialAI/audial-mcp
Install
You need an Audial account (user id + API key from https://audialmusic.ai) and
uv.
Claude Code
claude mcp add audial \
-e AUDIAL_USER_ID=your-user-id -e AUDIAL_API_KEY=your-api-key -e AUDIAL_RESULTS_DIR=~/Audial \
-- uvx audial-mcpClaude Desktop / Cursor / any client with a JSON config
{
"mcpServers": {
"audial": {
"command": "uvx",
"args": ["audial-mcp"],
"env": {
"AUDIAL_USER_ID": "your-user-id",
"AUDIAL_API_KEY": "your-api-key",
"AUDIAL_RESULTS_DIR": "~/Audial"
}
}
}
}Related MCP server: media-gen-mcp
Tools
Tool | What it does |
| Split a track into vocals, drums, bass, other (optionally retime / rekey) |
| BPM, key and other characteristics |
| Sections and component analysis |
| Mastering, optionally matched to a reference |
| Sample pack from a track |
| Audio to MIDI |
| Text to music, covers, remixes, extraction, completion with the Audial music model |
| One-shot → editable Audial Synth (Vital) preset |
| Lyrics + reference voice → sung vocal and MIDI |
| Browse previous results in your results folder |
What leaves your machine
Audio and text you pass to a tool are uploaded to Audial's API (https://api.audialmusic.ai)
over HTTPS and processed on Audial's servers; results are downloaded into
AUDIAL_RESULTS_DIR (default ~/Audial). Nothing else is sent.
Two other files are read if they exist, because the Audial SDK looks for them:
~/.audial/.audial_config.json and a .env in the directory the server was started from.
Credentials from your MCP client's config always win: audial-mcp snapshots its environment
before the SDK loads either file, so neither can redirect your credentials or the API host.
Every tool except list_results calls the Audial API, which requires an active Audial subscription on the account; without one the tool returns the API's subscription message.
Configuration
Variable | Required | Default |
| yes | |
| yes | |
| no |
|
| no |
|
| no | production API |
License
MIT. Source: https://github.com/AudialAI/audial-mcp
Available Tools
10 toolsanalyzeAnalyze audioARead-only
Analyze a track: BPM, key, loudness and other characteristics. Results are in metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Full path to a local audio file (.wav, .mp3, .aif, .aiff, .flac, .m4a, .ogg, .aac). ~ is expanded. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description usefully adds that results land in `metadata`, but it says nothing about whether analysis is synchronous or requires a follow-up retrieval call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core deliverable and output location front-loaded. 'and other characteristics' is mild filler, but the description is otherwise tight and waste-free.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't detail return values, and it does point to `metadata`. For a single-parameter read-only tool this is nearly complete; only the sync/async behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter and schema description coverage is 100% (the schema already documents accepted file formats and ~ expansion). The description adds no parameter detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Analyze') and resource ('a track') and enumerates the extracted characteristics (BPM, key, loudness). It is clearly distinct from the generative/stem siblings, though it never names an alternative explicitly to reinforce the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Correct usage is implied — point it at an audio file to obtain its characteristics — but there is no explicit when-to-use, no when-not-to-use, and no routing to or away from siblings such as list_results for retrieving results.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_midiAudio to MIDIC
Transcribe an audio file to MIDI.
| Name | Required | Description | Default |
|---|---|---|---|
| bpm | No | Override the detected tempo for note quantisation. | |
| file_path | Yes | Full path to a local audio file (.wav, .mp3, .aif, .aiff, .flac, .m4a, .ogg, .aac). ~ is expanded. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds no behavioral context such as output file handling, overwrite behavior, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and free of waste, clearly stating the core action. It could be slightly more informative without becoming verbose, but the conciseness is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description plus full schema coverage, annotations, and an output schema make invocation possible. However, the lack of usage guidance relative to sibling tools leaves a gap in helping an agent select this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are well documented in the schema itself. The description adds no parameter meaning beyond what the schema already provides, fitting the baseline of 3 for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Transcribe) and resource (audio file to MIDI), making the purpose clear. However, it does not explicitly differentiate from siblings like generate_music or sound2vital, leaving boundaries to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, nor are any prerequisites or exclusions mentioned. The description only states what the tool does, not when it is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_musicGenerate musicC
Generate music with the Audial music model.
Text to music, covers, remixes, stem extraction, completion, or analysis (understand).
| Name | Required | Description | Default |
|---|---|---|---|
| bpm | No | Tempo hint in beats per minute. | |
| seed | No | Seed for reproducible output. | |
| lyrics | No | Lyrics with optional section tags like [Verse] and [Chorus]. Omit for instrumental. | |
| prompt | Yes | Style description: genre, mood, instruments, vocal character. Tempo and key words in the prompt steer the model more than the numeric bpm/key fields, which are hints, not constraints. | |
| key_scale | No | Key hint, e.g. 'G major'. | |
| task_type | No | text2music (default), cover, remix, extract, lego, complete, or understand. | text2music |
| batch_size | No | Number of variations, 1-8. | |
| track_name | No | Track to extract/replace for extract/lego: vocals, drums, bass, guitar, piano, strings, synth, other. | |
| source_file | No | Local audio file; required for remix, extract, lego, complete, understand. | |
| audio_format | No | mp3, wav, flac, opus or aac. | mp3 |
| instrumental | No | Generate without vocals. | |
| audio_duration | No | Length in seconds (10-600). | |
| reference_file | No | Local audio file; required for cover. | |
| repainting_end | No | Remix end time in seconds (-1 = end). | |
| time_signature | No | e.g. '4/4'. | |
| vocal_language | No | Language code for vocals, e.g. 'en'. | en |
| negative_prompt | No | What to avoid, comma-separated. | |
| repainting_start | No | Remix start time in seconds. | |
| audio_cover_strength | No | 0-1 fidelity to the original for cover/remix. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, idempotentHint=false, openWorldHint=true). The description adds essentially nothing beyond that: it does not mention generation cost, latency, remote/service dependency, or how outputs are persisted, and its mode list merely restates values already in the task_type schema field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short lines, front-loaded with the core action and model. It is appropriately sized, though the trailing mode list is partly redundant with the schema and could have been replaced by a routing hint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 19 parameters, an output schema, and annotations, much of the burden is absorbed by structured fields, so the description is not dangerously thin. However, for a multi-mode tool it omits any guidance on mode selection or which parameters pair with which task, forcing the agent to reconstruct the workflow from individual schema fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already explains cross-parameter rules (source_file required for remix/extract/lego/complete/understand, reference_file for cover). The description adds no additional parameter meaning, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a clear verb+resource (generate music with the Audial model). The second enumerates modes, but two of them ('stem extraction', 'analysis (understand)') overlap with the sibling tools stem_split and analyze, muddying rather than sharpening differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no routing between this tool and siblings like stem_split, analyze, segment, or master. The mode list implies that task_type selects the operation, but the agent is left to infer which sibling to call for overlapping tasks such as separation or analysis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_samplesGenerate a sample packB
Extract a sample pack (one-shots and loops) from a track.
| Name | Required | Description | Default |
|---|---|---|---|
| genre | No | Genre hint. | |
| job_type | No | Sample pack job type accepted by Audial (default engine choice when omitted). | |
| file_path | Yes | Full path to a local audio file (.wav, .mp3, .aif, .aiff, .flac, .m4a, .ogg, .aac). ~ is expanded. | |
| components | No | Which components to sample, e.g. ['drums', 'bass']. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent, open-world operation, but the description adds no behavioral context: it does not say that output files are produced, whether the call runs as an async job (implied by job_type and the list_results sibling), how to retrieve results, or how long it takes. For a mutating tool with annotations only sketching the safety profile, this leaves a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the deliverable is named immediately. Nothing in it is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and all four parameters are documented in the schema. However, for a non-idempotent job-style tool that feeds a list_results workflow, the description omits async/result-retrieval context, leaving the agent to discover it elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so genre, job_type, file_path, and components are already documented in the schema. The description's mention of 'one-shots and loops' loosely maps to components but adds no format, required/optional, or default information beyond what the schema states. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Extract), resource (sample pack), and expands the resource's contents (one-shots and loops) with a scope (from a track). An agent knows what comes out, but the description never distinguishes this from the closest sibling, stem_split, which also decomposes a track.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this over stem_split, analyze, or segment, and no prerequisites or exclusions. The only routing signal is the tool name itself, so the agent must infer intent from the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_resultsList previous resultsARead-only
List previous Audial results in the results folder, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | No | Filter by tool name, e.g. 'stem_split'. Omit for all tools. | |
| limit | No | Maximum number of jobs to return, newest first. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds the ordering guarantee ('newest first') and the storage location ('results folder'), which is useful context, but says nothing about pagination limits or what happens when no results exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, efficient sentence with the resource and ordering front-loaded and zero filler. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the schema fully covers both optional parameters. For a simple read-only list tool this is nearly complete; only minor behavioral detail (pagination beyond limit, empty-folder behavior) is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters ('tool' filter and 'limit') are fully documented in the schema. The description's 'newest first' echoes the limit parameter's own description, adding no new semantic detail. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('previous Audial results') plus the scope ('in the results folder') and ordering ('newest first'). This clearly distinguishes it from the sibling generation tools, though it does not name any of them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'previous results' — an agent can infer this is the tool for retrieving already-produced outputs rather than creating new ones. However, there is no explicit when-to-use/when-not guidance or mention of alternatives for finding results.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
masterMaster a trackB
Master a mix, optionally matching a reference track.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Full path to a local audio file (.wav, .mp3, .aif, .aiff, .flac, .m4a, .ogg, .aac). ~ is expanded. | |
| reference_file | No | Optional reference track whose loudness and tone the master should match. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and idempotentHint=false. The description adds no further behavioral context such as side effects, output file creation, or whether the original mix is modified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single front-loaded sentence with no wasted words. The purpose is stated immediately and clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 100% schema coverage, existing output schema, and annotations, the description is sufficient for correct invocation. It lacks expanded usage context, but the surrounding structured fields compensate for the missing return and parameter details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with detailed descriptions for file_path and reference_file. The description only restates optional reference matching and adds no syntax, format, or constraint detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Master a mix' and notes optional reference matching. It is clear but does not explicitly differentiate itself from sibling tools such as analyze or stem_split.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides no explicit when-to-use guidance, no exclusions, and no alternative tools. Usage is only implied by the tool name and the phrase 'Master a mix'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
segmentSegment audioB
Detect song sections (intro, verse, chorus...) and analyze components within them.
| Name | Required | Description | Default |
|---|---|---|---|
| genre | No | Genre hint that improves section detection. | |
| features | No | Features to compute per segment. | |
| file_path | Yes | Full path to a local audio file (.wav, .mp3, .aif, .aiff, .flac, .m4a, .ogg, .aac). ~ is expanded. | |
| components | No | Components to analyze, e.g. ['vocals', 'drums']. | |
| analysis_type | No | Analysis type accepted by Audial's segmentation. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and idempotentHint=false. The description adds no behavioral context such as whether output files are created, whether repeated calls are safe, or what side effects to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no redundant or filler text. Every word contributes to stating what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and full schema coverage, return values and parameters are handled elsewhere. However, for a tool with multiple sibling analysis/generation tools, the description omits when to use it versus alternatives and does not disclose enough behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters like genre, features, components, and analysis_type are fully documented in the schema. The description mentions components generally but adds no syntax or format details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: detect song sections (intro, verse, chorus) and analyze components within them. The purpose is clear, but it does not explicitly distinguish itself from the sibling tool 'analyze' or 'stem_split'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no conditions for choosing this over siblings like analyze or stem_split, and no prerequisites beyond the schema. The description only implies usage from its function.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sound2vitalResynthesize a one-shot into a synth presetA
Turn a one-shot sample into an editable Audial Synth preset.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | A short one-shot (<= 20 s) audio file. The result is an editable Audial Synth (.vital) preset that reproduces its timbre. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and idempotentHint=false, covering mutation, external access, and non-idempotency. The description adds the <=20 s input constraint and that the result is an editable preset reproducing timbre, but omits file output behavior and other side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is appropriately sized for a one-parameter conversion tool and every word contributes to the action and output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists and the annotations carry the safety profile, so the description need not explain return values. It provides the input length constraint and output type, making it complete enough for correct invocation, though usage guidance remains thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single file_path parameter is fully documented as a short one-shot audio file. The description adds no syntax, format, or constraints beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific conversion: a one-shot sample becomes an editable Audial Synth preset. The verb and resource are clear, though it does not explicitly distinguish itself from sibling audio tools like generate_samples or analyze.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use, exclusions, or alternatives are provided. Usage is only implied by the conversion description, which is minimal for choosing between this tool and siblings such as generate_samples or analyze.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stem_splitSplit stemsC
Split a track into separate instrument stems, optionally retimed and rekeyed.
| Name | Required | Description | Default |
|---|---|---|---|
| stems | No | Stems to extract from: vocals, drums, bass, other, full_song_without_vocals. Default: vocals, drums, bass, other. | |
| algorithm | No | Separation algorithm. Default 'primaudio'. | primaudio |
| file_path | Yes | Full path to a local audio file (.wav, .mp3, .aif, .aiff, .flac, .m4a, .ogg, .aac). ~ is expanded. | |
| target_bpm | No | Retime the stems to this tempo (beats per minute). | |
| target_key | No | Transpose the stems to this key, e.g. 'A minor' or 'F# major'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true, so the safety/mutation profile is covered. The description adds nothing beyond echoing the optional retime/rekey parameters; it omits what happens to the source file, whether output stems are written to disk, and where results surface. With annotations doing the heavy lifting, the description contributes minimal extra context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core action front-loaded and no filler. It is efficient, though it spends its one clause on optional modifiers rather than routing or prerequisite information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, for a non-idempotent, open-world, file-consuming tool with five parameters, the description says nothing about where stems are produced or how the agent retrieves them (e.g., via list_results), leaving a meaningful operational gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (stems list, algorithm, file_path, target_bpm, target_key) are already documented with defaults and formats in the schema. The description only paraphrases target_bpm/target_key and adds no syntax or constraint detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('split a track into separate instrument stems') with an added capability note ('retimed and rekeyed'). It is clear what the tool does, but it never differentiates itself from siblings like segment, analyze, or master, which could overlap in the audio-processing family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given: nothing tells the agent when to prefer stem_split over segment, analyze, or master, nor any prerequisites (source file requirements, expected runtime, whether results need list_results). The optional retime/rekey behavior is mentioned but not framed as a decision criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
text2voxSing lyrics in a reference voiceB
Synthesize a sung vocal (and MIDI) from lyrics, a melody and a reference voice.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed for reproducible output. | |
| lyrics | Yes | Lyrics to sing. | |
| midi_file | No | MIDI melody (.mid). Give this or melody_audio_file. | |
| nfe_steps | No | Synthesis steps; more is slower and cleaner. | |
| lyrics_mode | No | auto (default) or an explicit lyric alignment mode. | auto |
| pitch_shift | No | Semitones to shift the melody. | |
| cfg_strength | No | Guidance strength. | |
| strict_pitch | No | Force exact pitch to the melody. | |
| no_pitch_bends | No | Disable pitch bends. | |
| reference_file | Yes | Short clip of the reference voice (timbre). | |
| reference_text | No | Transcript of the reference clip, if known. | |
| bend_smoothing_ms | No | Pitch-bend smoothing window in ms. | |
| leading_silence_s | No | Silence before the first note, seconds. | |
| melody_audio_file | No | Audio whose melody is transcribed and followed, if no MIDI. | |
| word_timestamps_file | No | Optional JSON word timings. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tool | Yes | |
| files | Yes | |
| summary | Yes | |
| metadata | Yes | |
| output_dir | Yes | |
| execution_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false and openWorldHint=true, so a generation/mutation profile is known. The description adds that the output includes both a sung vocal and MIDI, which is useful context. It still says nothing about runtime cost, GPU/quality tradeoffs, or the fact that reference_file is a required timbre input.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. It is efficient, though for a 15-parameter open-world generation tool it is arguably too terse to be optimally sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 15 parameters, an open-world generation annotation, and no usage guidance, a one-sentence description leaves real gaps: which melody source is required, how to pick among the many tuning knobs, and how it relates to sibling generators. The output schema and 100% param coverage cover return values and arguments, but the routing/usage layer is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 15 parameters (seed, nfe_steps, cfg_strength, strict_pitch, etc.) are already documented in the schema. The description adds no syntax, format, or interaction detail beyond it, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (synthesize) and resource (a sung vocal and MIDI) plus the three required inputs (lyrics, melody, reference voice). This clearly separates it from siblings like generate_music or generate_midi. It stops short of explicitly naming or contrasting those alternatives, so it does not reach the top of the scale.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as generate_music or which melody source (midi_file vs melody_audio_file) to prefer. The agent gets no routing help beyond the bare purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.1- First observed
analyze - First observed
generate_midi - First observed
generate_music - First observed
generate_samples - First observed
list_results - First observed
master - First observed
segment - First observed
sound2vital - First observed
stem_split - First observed
text2vox
TDQS
Scored across 10 tools
Most tools cover distinct audio operations such as stem splitting, analysis, mastering, MIDI transcription, and vocal synthesis, so many choices are clear. However, generate_music advertises stem extraction, analysis, and completion, overlapping with stem_split, analyze, and segment, which could cause misselection.
Names follow mixed conventions: verb-only (analyze, segment, master), verb_noun (generate_samples, list_results), and concatenated forms (stem_split, sound2vital, text2vox). They remain readable, but there is no predictable naming pattern.
With 10 tools, the set sits in the recommended 3-15 range and each tool represents a distinct audio or music capability. The surface does not feel bloated or thin for this domain.
Core workflows are covered: splitting, analysis, segmentation, mastering, sample/MIDI generation, music/vocal generation, synth preset conversion, and listing results. A minor gap is the lack of get_result/status or deletion/management operations beyond list_results.
Maintenance
Related MCP Connectors
Hosted MCP tools for FFmpeg-style video and audio processing through FFMPEG API.
MCP server for Producer/Riffusion AI music generation
AI image, video, voice and music generation over MCP, routed to Veo 3.1, Seedance 2.5 and more.
Generate images, video, audio and short films with 140+ AI models from any MCP client.
1
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to programmatically edit, analyze, and export audio projects through MCP tools, including multi-track editing, effects, transcription, and semantic search.1-
- FlicenseNot gradedqualityBmaintenanceEnables AI clients like Claude and ChatGPT to generate images and videos, animate images, create lip-synced videos, list TTS voices, and manage media via remote MCP tools.-
- AlicenseAqualityCmaintenanceEnables AI agents to master audio tracks to target LUFS/True Peak levels, remove Suno/Udio AI fingerprints, and retrieve mastering passports via a hosted MCP server.1133 npm1MIT
- AlicenseBqualityBmaintenanceEnables MCP-compatible agents to queue, run, and track local AI music generation jobs with dry-run defaults, license recording, and pluggable pipeline adapters for YuE or compatible backends.31AGPL 3.0