Skip to main content
Glama
AudialAI

io.github.AudialAI/audial-mcp

Official
by AudialAI

Generate music

generate_music

Turn text prompts or audio files into music, covers, remixes, stems, completions, or analysis with the Audial model.

Instructions

Generate music with the Audial music model.

Text to music, covers, remixes, stem extraction, completion, or analysis (understand).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
bpmNoTempo hint in beats per minute.
seedNoSeed for reproducible output.
lyricsNoLyrics with optional section tags like [Verse] and [Chorus]. Omit for instrumental.
promptYesStyle description: genre, mood, instruments, vocal character. Tempo and key words in the prompt steer the model more than the numeric bpm/key fields, which are hints, not constraints.
key_scaleNoKey hint, e.g. 'G major'.
task_typeNotext2music (default), cover, remix, extract, lego, complete, or understand.text2music
batch_sizeNoNumber of variations, 1-8.
track_nameNoTrack to extract/replace for extract/lego: vocals, drums, bass, guitar, piano, strings, synth, other.
source_fileNoLocal audio file; required for remix, extract, lego, complete, understand.
audio_formatNomp3, wav, flac, opus or aac.mp3
instrumentalNoGenerate without vocals.
audio_durationNoLength in seconds (10-600).
reference_fileNoLocal audio file; required for cover.
repainting_endNoRemix end time in seconds (-1 = end).
time_signatureNoe.g. '4/4'.
vocal_languageNoLanguage code for vocals, e.g. 'en'.en
negative_promptNoWhat to avoid, comma-separated.
repainting_startNoRemix start time in seconds.
audio_cover_strengthNo0-1 fidelity to the original for cover/remix.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
toolYes
filesYes
summaryYes
metadataYes
output_dirYes
execution_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation profile (readOnlyHint=false, idempotentHint=false, openWorldHint=true). The description adds essentially nothing beyond that: it does not mention generation cost, latency, remote/service dependency, or how outputs are persisted, and its mode list merely restates values already in the task_type schema field.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short lines, front-loaded with the core action and model. It is appropriately sized, though the trailing mode list is partly redundant with the schema and could have been replaced by a routing hint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 19 parameters, an output schema, and annotations, much of the burden is absorbed by structured fields, so the description is not dangerously thin. However, for a multi-mode tool it omits any guidance on mode selection or which parameters pair with which task, forcing the agent to reconstruct the workflow from individual schema fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already explains cross-parameter rules (source_file required for remix/extract/lego/complete/understand, reference_file for cover). The description adds no additional parameter meaning, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a clear verb+resource (generate music with the Audial model). The second enumerates modes, but two of them ('stem extraction', 'analysis (understand)') overlap with the sibling tools stem_split and analyze, muddying rather than sharpening differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no routing between this tool and siblings like stem_split, analyze, segment, or master. The mode list implies that task_type selects the operation, but the agent is left to infer which sibling to call for overlapping tasks such as separation or analysis.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.