Audionaut MCP Server
Audionaut
A free, open-source multitrack audio editor that AI agents can drive over MCP.
audionaut.app · Discussions · Sponsor
Audionaut is a free, open-source desktop application for effortless audio editing and recording. Whether you're working on music, podcasts, or multitrack recordings, it gives you precise cutting, per-track playlists, flexible multi-channel support, and clean exports — without the weight and complexity of a full DAW. Audionaut is written in modern C++ on the JUCE framework and runs natively on Windows, macOS, and Linux.
Use with Claude
Audionaut speaks MCP, so Claude and other agents can edit your sessions. With the app and Node.js 18+ installed:
claude mcp add audionaut -- npx -y audionaut-mcpKeep the project open in Audionaut and each edit arrives as one undo step. On macOS, keep projects in your Music folder. Claude Desktop setup and details: manual, chapter 11.
Checkout
Make sure to clone with submodules:
git clone --recursive https://github.com/kvoltmer/Audionaut.gitor once cloned do:
git submodule update --init --recursiveEssentia (audio analysis)
The analysis features (BIC segmentation, onset detection, beat tracking) link against a static build of the Essentia submodule. Build it once after checking out the submodules:
./Audionaut/Builds/build_essentia.shThe script builds Essentia's 3rd-party dependencies statically (slow, skipped on re-runs; force with FORCE_3RDPARTY=1), then configures and builds Essentia itself with its waf build system into Submodules/essentia/build. It patches Essentia's Linux-oriented 3rd-party build scripts in the working tree as needed for macOS/Apple Silicon, so the essentia submodule will show as dirty afterwards — don't commit those changes.
Prerequisites:
Python ≤ 3.11 — Essentia's bundled waf needs
distutils, removed in Python 3.12. The script picks a suitable interpreter automatically; override withPYTHON=... ./Audionaut/Builds/build_essentia.sh.pkg-config (e.g.
brew install pkg-config)CMake 3.x recommended — some 3rd-party deps ship very old CMakeLists that CMake 4.x refuses (e.g.
pip install "cmake~=3.31.0").
Building Essentia is optional: without it, ESSENTIA_ENABLED auto-detects off and the app and tests compile with the analysis features disabled.
demucs.cpp (stem separation)
Stem separation runs demucs.cpp, a C++ port of Meta's Demucs, compiled straight from the Submodules/demucs.cpp submodule — no separate build step. It needs the submodule's vendored Eigen (Submodules/demucs.cpp/vendor/eigen, a nested submodule that --recursive fetches); the CMake build detects it and sets AUDIONAUT_ENABLE_DEMUCS accordingly. The model weights are not in the repository: the app downloads them on first use into its Models folder, or point the CLI at a copy with --model.
Build
Xcode — open the Xcode project located here:
Audionaut/Builds/MacOSX/Audionaut.xcodeprojVisual Studio 2026 — open the Visual Studio solution located here:
Audionaut/Builds/VisualStudio2026/Audionaut.slnLinux Makefile — install dependencies:
sudo apt install libasound2-dev libjack-jackd2-dev ladspa-sdk libcurl4-openssl-dev libfreetype-dev libfontconfig1-dev libx11-dev libxcomposite-dev libxcursor-dev libxext-dev libxinerama-dev libxrandr-dev libxrender-dev libwebkit2gtk-4.1-dev libglu1-mesa-dev mesa-common-dev libxi-dev libegl-devthen compile:
cd Audionaut/Builds/LinuxMakefile/
make CONFIG=Release -j8Tests
See the Catch2 tests README.
Command-line tool (audionaut-cli)
audionaut-cli gives scripts, CI and AI agents headless access to .audium
projects — no GUI, no audio device. It builds alongside the tests:
cmake -B build -S Audionaut/Catch2Tests
cmake --build build -j8 --target AudionautCli
./build/AudionautCli_artefacts/AudionautCli --helpEvery command takes --json to emit exactly one machine-readable result
envelope on stdout ({"ok": true, "result": ...} or {"ok": false, "error": ...}) with all logging on stderr, plus --quiet. Exit codes: 0 success,
1 operation failed, 2 usage error, 3 feature unavailable in this build
(e.g. analyze without Essentia). Option values may be given as --opt value
or --opt=value.
audionaut-cli create song.audium --channels 2
audionaut-cli import song.audium take1.wav take2.wav --position 4.5
audionaut-cli info song.audium --json
audionaut-cli analyze song.audium --types sbic,beat_degara
audionaut-cli auto-edit song.audium --track 0 --measures 4
audionaut-cli assemble song.audium --duration 60 --mode random
audionaut-cli split song.audium --at 23 # bar 23; --unit beats|seconds|clocks
audionaut-cli create-region song.audium --name chorus --start 17 --end 25
audionaut-cli set-region song.audium --region chorus --rename drop --length 8
audionaut-cli place-clip song.audium --region drop --at 33
audionaut-cli move-clip song.audium --region drop --to 41
audionaut-cli remove-clip song.audium --at 41 --track 0
audionaut-cli clip-gain song.audium --region drop --gain -6 --db
audionaut-cli clip-fades song.audium --region drop --fade-in 1 --fade-out 2 --unit beats
audionaut-cli separate song.audium --track 1 # Drums/Bass/Other/Vocals tracks; needs the Demucs model
audionaut-cli export song.audium -o mix.wav --sample-rate 48000 --bit-depth 24A typical agent flow: create → import → analyze → auto-edit/assemble
→ export, checking ok in each --json envelope. Projects written by the
CLI open in the GUI app and vice versa.
End-user documentation lives in the User Manual.
CLI invocations report anonymous usage statistics under the same strictly
opt-in consent as the app (one cli_command event: verb and exit code) —
nothing is sent unless consent was granted in the app's settings. Set
AUDIONAUT_DISABLE_ANALYTICS=1 to switch the CLI's reporting off regardless
(CI environments, scripts).
The main Audionaut app also accepts the same verbs (Projucer-style): run the app binary with a verb and it executes headlessly and quits with the command's exit code, even while a GUI instance is open — a file argument or no arguments launches the GUI as usual.
./Audionaut.app/Contents/MacOS/Audionaut export ~/Music/song.audium -o ~/Music/mix.wavNote: the macOS app is sandboxed, so its in-app CLI can only reach
entitlement-covered locations such as ~/Music; the standalone
audionaut-cli build has no such restriction. On Windows, the app is a GUI
program — the shell prompt returns immediately and output interleaves with
it; prefer audionaut-cli for scripting there.
License
Audionaut is dual-licensed under GPL3 (or later) and a commercial license — see LICENSE.md for details.
Contributing
Contributions are welcome — see CONTRIBUTING.md for how to get started and the required Contributor License Agreement (CLA.md).
Support
If Audionaut is useful to you, consider sponsoring its development — sponsorship directly funds development time and keeps the project sustainable.
Available Tools
22 toolsanalyzeAnalyze audioA
Runs Essentia audio analysis (segment boundaries, beats, BPM) on the project's audio files - or one standalone audio file - and caches the results next to the project for auto_edit/assemble to use. Fails with essentia_unavailable in builds without Essentia.
| Name | Required | Description | Default |
|---|---|---|---|
| types | No | Comma-separated analysis types, e.g. "sbic,beat_degara" (default: the merge set) | |
| target | Yes | A .audium project package or a single audio file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses the caching side effect (results written next to the project) and a concrete failure mode (essentia_unavailable in builds lacking Essentia). It stops short of stating cost, duration, or whether re-analysis overwrites the cache.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the core purpose plus target is front-loaded before the failure-mode caveat. Slightly dense but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter, no-output-schema tool, the definition covers target scope, produced data, caching behavior, and a build-dependent failure mode. Only the exact return shape and re-run semantics are unstated, which is minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both target and types are already documented in the schema, including the default merge set. The description adds only indirect meaning by naming the categories of analysis (segment boundaries, beats, BPM) that map to the types parameter. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific verb (analyze) and the concrete outputs (segment boundaries, beats, BPM) via a named engine, and it explicitly scopes the target to the project's audio files or a standalone file. That distinguishes it from downstream consumers like auto_edit/assemble and from import/export audio tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: caching results 'for auto_edit/assemble to use' hints that this precedes editing, but there is no explicit when-to-run/when-not guidance or named alternative. The agent must infer the workflow position.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assembleAssemble arrangementB
Builds an arrangement of the given duration from the project's regions, sequentially or at random (requires regions - e.g. from auto_edit).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Selection mode (default sequential) | |
| seed | No | Random seed for reproducible random mode | |
| track | No | Track id (default 0) | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| crossfades | No | Apply crossfades at joints (default true) | |
| duration_seconds | No | Target duration (default 60) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It mentions the regions prerequisite and mode options but does not disclose mutation side effects, permissions, reversibility, or what happens to existing arrangements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. The core action, source, and modes are stated immediately, and the parenthetical prerequisite is efficiently placed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the high-level purpose and a key dependency, but for a 6-parameter mutation tool with no annotations and no output schema, it omits enough behavioral context (e.g., side effects, interaction with existing arrangements) that an agent might need to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds a prerequisite about regions but no additional syntax or format details beyond what the schema provides, making baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Builds') and resource ('arrangement'), and clarifies the source material ('project's regions') and selection mode. It also names a sibling tool (auto_edit) as a source of required regions, but does not explicitly distinguish itself from other arrangement-related siblings like place_clip.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a key prerequisite: 'requires regions - e.g. from auto_edit', which implies usage context. However, it does not state when to use this tool versus alternatives, nor does it include exclusions or explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auto_editAuto-edit clipA
Segments a clip on a track using cached analysis results (run analyze first). Splits the clip into musically-aligned regions.
| Name | Required | Description | Default |
|---|---|---|---|
| clip | No | Playlist item id (default: the track's clip) | |
| track | No | Track id (default 0; imported audio lands on a new track) | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| measures | No | Segment length in measures | |
| segments | No | Number of segments | |
| crossfades | No | Apply crossfades at joints (default true) | |
| duration_seconds | No | Target duration |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the analyze prerequisite and the musical-region outcome, but says nothing about whether the edit is destructive, how existing regions are affected, permissions needed, or what happens to the original clip.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences that front-load the action and the analyze prerequisite. Every sentence earns its place, with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema is rich and fully documented, and the description supplies the key analyze prerequisite. However, as a mutation tool with no annotations and no output schema, it still omits side-effect and return-behavior context that an agent would need to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all seven parameters. The description adds no parameter syntax, defaults, or interaction details beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it segments/splits a clip on a track into musically-aligned regions. It is clearly not a tautology and adds the cached-analysis basis, though it does not explicitly distinguish itself from sibling tools like split or create_region.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear prerequisite: run analyze first and use cached analysis results. That establishes when the tool is applicable, but it does not name alternatives or exclusions, so it falls short of a full when/when-not routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cleanup_regionsDelete unused regionsA
Deletes every region that has no clip on the timeline (the GUI's Delete Unused Regions). The audio files stay in the package.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well on the most important point: it defines the destructive boundary by telling the agent that the underlying audio files stay in the package. It omits irreversibility/undo and any auth or confirmation requirement, so it stops short of full behavioral coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and its scope, followed immediately by the side-effect clarification. Nothing is padded or repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema tool, the description supplies the operation, its scope, and the key side effect an agent must know before calling it. The only real gap is the absence of any note on reversibility, which matters for an unannotated destructive operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and schema description coverage is 100%, so the schema already explains the .audium project path. The description adds no format or path guidance beyond that, matching the baseline for a fully documented single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Deletes') and a precisely bounded resource ('every region that has no clip on the timeline'), which cleanly separates it from sibling mutators like create_region, set_region, and remove_clip. The parenthetical GUI mapping pins the intent unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The condition under which regions are targeted is explicit ('no clip on the timeline'), but there is no guidance on when to run this versus leaving regions alone, nor any prerequisite or safety note. Usable but inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clip_fadesSet clip fadesA
Sets the addressed clip's fade lengths, ramp offsets (0 clears; offsets may be negative to reach outside the clip) and curve exponents (0.1-4, 0.5 = equal power). Values are clamped against each other within the clip. The address must match exactly one clip.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | Timeline position of the clip (exclusive with region) | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| track | No | Track id | |
| region | No | Region name of the clip | |
| fade_in | No | Fade-in length | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| fade_out | No | Fade-out length | |
| fade_out_end | No | Fade-out ramp end offset from the clip end | |
| fade_in_curve | No | Fade-in curve exponent (0.1-4, 0.5 = equal power) | |
| fade_in_start | No | Fade-in ramp start offset from the clip start | |
| fade_out_curve | No | Fade-out curve exponent (0.1-4, 0.5 = equal power) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose two useful traits beyond the schema: that values are clamped against each other within the clip, and that the address must resolve to exactly one clip. It says nothing about permissions, reversibility, or what a successful response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, no filler, with the core action stated first and the address precondition last. Every clause carries information the agent needs (ranges, clearing behavior, clamping, addressing).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutation tool with no annotations and no output schema, the description covers the essential behavior (what is set, clamping, addressing constraint) and the schema covers addressing fields and defaults. It stops short of explaining the addressing scheme ('the addressed clip' is never defined) or failure behavior on non-matching addresses, which would be needed for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning not in the schema: ramp offsets of 0 clear the offset and negative offsets reach outside the clip, and curve exponents are documented with the 0.5 = equal-power reference point. These clarify the semantics of the offset and curve parameters beyond their terse schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Sets the addressed clip's fade lengths') and enumerates the exact settings touched: fade lengths, ramp offsets, curve exponents. It is unambiguous against siblings like clip_gain and clip_speed, which operate on entirely different clip properties.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the operation itself; there is no statement of when to prefer this over related tools or what preconditions besides addressing apply. The one genuine guideline given is the precondition 'The address must match exactly one clip,' which does tell the agent the call will fail on ambiguous addresses.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clip_gainSet clip gainA
Sets the addressed clip's gain - linear by default, dB with db=true - on every destination channel, or one channel. The address (at or region) must match exactly one clip.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | Timeline position of the clip (exclusive with region) | |
| db | No | Interpret gain as decibels | |
| gain | Yes | Gain value (linear unless db=true; linear must be >= 0) | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| track | No | Track id | |
| region | No | Region name of the clip | |
| channel | No | One destination channel (default: all) | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses linear-vs-dB gain interpretation, channel targeting, and the exact-one-clip addressing constraint, but it does not state permissions, undoability, or failure behavior when zero or multiple clips match.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core operation and then the addressing constraint. Every sentence earns its place without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no annotations and no output schema, the description covers invocation-critical addressing and gain semantics. However, it omits error behavior, permission requirements, and side-effect details, leaving it adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds meaningful cross-parameter semantics beyond individual schema fields: the at/region address must resolve to exactly one clip, and it reinforces the all-channels-versus-one-channel scope. It still does not detail unit, track, or project semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Sets the addressed clip's gain') and includes the exact addressing scope ('at or region must match exactly one clip'). An agent can distinguish this from sibling operations such as clip_fades or move_clip without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies a prerequisite ('address must match exactly one clip') and describes channel scope, but it does not name alternatives or explain when a gain change is preferable to other clip operations. Usage is implied from the verb and addressing rule rather than explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clip_speedSet clip speedA
Sets a clip's playback speed (ratio 0.25-4, 2.0 = double speed, half as long). The mode picks how: repitch (default, varispeed - pitch and length change together) or stretch (pitch-preserving time-stretch). Mode can also be changed on its own. lockTempo ties the clip to the project tempo (speed = project tempo / clip tempo, following tempo changes); the clip tempo is detected by beat tracking or given with tempo, which implies the lock.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | Timeline position of the clip (in `unit`) | |
| mode | No | How the speed is realised: varispeed or pitch-preserving | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| ratio | No | Speed ratio (0.25-4) | |
| tempo | No | The clip's own tempo in BPM; sets it and locks the clip | |
| track | No | Track id to narrow the match | |
| length | No | Fit the clip to this timeline duration (in `unit`) | |
| region | No | Region name of the clip | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| lockTempo | No | Lock the clip to the project tempo (true) or release it (false) | |
| semitones | No | Pitch shift in semitones (ratio = 2^(n/12)) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does well: it discloses the ratio range and mapping, repitch vs stretch semantics, lockTempo's tempo-following formula, and clip-tempo detection via beat tracking or explicit tempo. It still omits permissions, undo behavior, and side effects of mutating a clip, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded: the first sentence defines the core action, the next explains mode, and the last explains tempo locking. Every sentence earns its place without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutation tool with full schema descriptions but no annotations or output schema, the description covers the essential operational semantics and parameter interactions. It could be more complete by mentioning side effects, required permissions, or undo implications.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema for key parameters, including ratio-to-length mapping, mode defaults and behavior, and the relationship between lockTempo, tempo, and detected clip tempo.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Sets a clip's playback speed.' It immediately distinguishes the tool from siblings like clip_gain and clip_fades by targeting playback speed, so an agent can identify its function without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains operational choices such as mode selection and lockTempo behavior, implying when to use those features. However, it does not explicitly state when to use clip_speed versus other editing tools, nor does it describe when not to use it or what alternatives exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectCreate projectA
Creates a new empty Audionaut project (.audium package) with one track. Fails if the path already exists.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | Target path for the new .audium package | |
| channels | No | Channels on the initial track (default 2) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden but adds some useful context: it creates an 'empty' project with one track and fails if the path already exists. However, it omits permission requirements, behavior on partial path existence, the effect of the channels parameter, and any return information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the core action and artifact, followed by a failure condition. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple project-creation tool with full schema coverage and no output schema, the description is largely complete: it covers the action, artifact type, initial track, and failure condition. Without annotations, additional behavioral details (e.g., permissions) could be stated, but the core information is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameter meanings are fully documented in the schema. The description mentions 'one track' but does not explain the channels parameter or its default, adding no semantic detail beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Creates) and resource (Audionaut project), identifies the artifact type (.audium package), and describes the initial state (one track). It implicitly distinguishes itself from sibling tools like create_region or place_clip by focusing on project creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives, nor does it mention prerequisites or typical scenarios. It only states a failure condition (path already exists), which is behavioral rather than usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_regionCreate regionA
Creates a named region from a timeline range on every track whose clip fully contains it (the GUI's Create Region command). The arrangement itself is unchanged.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | Range end (must be after start) | |
| name | Yes | Name for the new region | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| start | Yes | Range start, e.g. 23 for bar 23 | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add one valuable trait: 'The arrangement itself is unchanged,' clarifying this is a non-destructive labeling operation. It still omits key behavior – what happens when no track contains the range, whether a duplicate name overwrites, permission needs, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and scope, with the second sentence delivering a distinct and useful non-obvious fact. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description need not explain return values, and the schema covers all parameters. The core creation semantics and the non-mutation guarantee are present, though edge-case behavior (no containing clip, name collision) is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters including units, ordering constraint (end must be after start), and path semantics. The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Creates a named region'), plus the precise scope condition ('from a timeline range on every track whose clip fully contains it'). It anchors to the GUI's Create Region command, though it never explicitly distinguishes itself from the sibling set_region, so sibling differentiation is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a real applicability condition – only tracks whose clip fully contains the range are affected – which implies when the tool is and isn't useful. However, it names no alternative (e.g. set_region) and gives no prerequisites or when-not guidance, so the agent must infer the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_audioExport audioA
Renders the project offline to a WAV file (no audio device needed). Pass region to bounce a single region instead - always dry, without any clip's gains or fades.
| Name | Required | Description | Default |
|---|---|---|---|
| track | No | Track id, to disambiguate same-named regions | |
| output | Yes | Output .wav path | |
| region | No | Bounce this region instead of the whole project | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| channels | No | Output channel count (default 2) | |
| bit_depth | No | Bit depth (default 24) | |
| multi_mono | No | Write one mono file per channel instead | |
| sample_rate | No | Sample rate in Hz (default 44100) | |
| start_seconds | No | Export start position | |
| length_seconds | No | Export length (default: whole project) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and usefully discloses that rendering is offline and needs no audio device. It also reveals an important region-export behavior: the bounce is always dry and ignores clip gains and fades. It does not cover overwrite behavior or failure cases, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences front-load the core action and then explain the region alternative. Every sentence carries information relevant to correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter export tool with no output schema, the description covers the core operation and the most important special case. The schema fully covers parameters, but the description does not document what is returned or whether existing output files are overwritten.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters, making 3 the appropriate baseline. The description adds meaningful semantics for the region option, but does not add syntax or meaning for the remaining parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: it renders a project offline to a WAV file. It also clearly distinguishes the whole-project mode from the region-bounce mode, so an agent can tell exactly what the tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to pass region to bounce a single region instead of the whole project, which gives clear usage context for the main branching behavior. It does not state when not to use the tool or name alternatives, but no sibling export tool exists to compare against.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_project_infoGet project infoB
Summarizes an Audionaut project: tempo, tracks, clips with positions/durations and their audio files. Set raw=true for the full persistence JSON instead of the summary.
| Name | Required | Description | Default |
|---|---|---|---|
| raw | No | Dump the full persistence JSON instead of the summary | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses what the summary includes and that raw=true changes output, which is useful. However, it doesn't mention whether this is a read-only operation, performance implications, permissions needed, or error behavior. For a likely read operation, this is a reasonable start but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the main purpose, followed by the raw parameter behavior. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool is relatively simple (2 params, no output schema). Description covers the summary content and the raw option, but lacks usage context and behavioral details like read-only nature. For a project info retrieval tool, it's adequate but not fully complete without annotations or usage guidelines.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both parameters described in the schema (raw and project). The description adds that raw=true gives full persistence JSON, slightly beyond the schema's 'Dump the full persistence JSON instead of the summary'. For project, no additional context on path format beyond schema's note. Baseline 3 appropriate when schema is high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it summarizes a project, enumerating the specific contents (tempo, tracks, clips with positions/durations and audio files). Distinguishes from siblings like create_project or analyze, but doesn't explicitly contrast with them. Verb+resource is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when to use or when not to use guidance. It suggests reading project info, but doesn't say alternatives like analyze or which situations warrant this tool. The mention of raw=true gives some usage hint but lacks broader context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_audioImport audioB
Imports one or more audio files into the project at the given position and saves it. Note: importing creates a new track holding the files.
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | Audio file paths to import | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| position_seconds | No | Timeline position for the files (default 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two non-obvious traits: the import persists by saving the project, and it creates a new track to hold the files. However, it says nothing about permissions, failure modes, undo/reversibility, or what happens to an existing track at the same position.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, both front-loaded with the action and its scope, and the side-effect caveat is placed after the main clause where it belongs. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no annotations and no output schema, the description covers purpose and one key side effect but omits prerequisites, supported formats, and error behavior. Adequate to start, but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so files, project, and position_seconds are already documented in the schema. The description's phrase 'at the given position' merely restates the position parameter without adding format or default semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (imports) and resource (audio files) plus the target scope (into the project at the given position) and the persistence effect (saves it). It does not explicitly distinguish itself from similar siblings like place_clip or assemble, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose import_audio over alternatives such as place_clip, nor any prerequisites (e.g. whether the .audium project must already exist, or which file formats are accepted). The trailing note about track creation is behavioral, not usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_clipMove clipA
Moves one clip to a new timeline position and/or onto another track. Address it by position (at) or region name; the address must match exactly one clip. to_track takes a track id, or "new" to create a track for it; without to the clip keeps its position. Give at least one of to and to_track.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | Current position of the clip (exclusive with region) | |
| to | No | Target timeline position (default: unchanged) | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| track | No | Track id | |
| region | No | Region name of the clip | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| to_track | No | Target track id, or "new" to create a track |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose meaningful behavior: the "new" value for to_track creates a track (a side effect), omitting to preserves the current position, and a non-unique address is an error. It does not mention permissions, undo, or how overlapping clips are handled on the target track.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and then the addressing rules, with almost no filler. The closing sentence partially restates the earlier "without to the clip keeps its position" point, costing a little density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no annotations and no output schema, the description covers addressing, defaults, the create-track side effect, and the required-argument rule. It stops short of describing collision behavior on the destination track or the tool's result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds a genuine cross-parameter constraint not expressed in the schema: at least one of to and to_track must be supplied, and at/region are alternative addressing modes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Moves one clip") and clarifies the two dimensions of movement: timeline position and/or track. It does not differentiate itself from the sibling place_clip, so an agent must infer the split between moving an existing clip and placing one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives operational constraints ("the address must match exactly one clip", "Give at least one of to and to_track") but never states when to choose this over place_clip or remove_clip. The guidance is about how to call it, not when to prefer it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
place_clipPlace clipA
Creates a new clip from an existing named region at a timeline position, on the track that owns the region (track only disambiguates same-named regions).
| Name | Required | Description | Default |
|---|---|---|---|
| at | Yes | Timeline position for the new clip | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| track | No | Track id, to disambiguate same-named regions | |
| region | Yes | Name of the region to place | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It usefully clarifies the track semantics ('track only disambiguates same-named regions') and reveals the clip is created from an existing region, but says nothing about permissions, what happens on name/position collisions, or whether the project is written to disk. For a no-annotation mutation tool this is only partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the parenthetical about track disambiguation is the only elaboration and it earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and 3 required params, the description covers the core action but omits return value (e.g., a new clip id), failure modes, and side effects. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description's phrasing about the timeline position and track disambiguation largely restates what the schema already documents for the 'at' and 'track' parameters, adding little new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (creates), resource (clip), source (an existing named region), and placement (a timeline position on the region's owning track). This is readily distinguishable from siblings like move_clip (moves an existing clip) and create_region (creates a region, not a clip).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the tool applies (turning a named region into a clip), but it never states when to use it versus alternatives such as move_clip or create_region, nor any prerequisites. Usage is only inferable from the purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_channelRemove channelA
Removes one channel (0-based) from a track, like deleting a channel strip in the GUI. Channel mappings and per-channel clip gains shift down; the audio files stay in the package.
| Name | Required | Description | Default |
|---|---|---|---|
| track | Yes | Track id | |
| channel | Yes | Channel index within the track (0-based) | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real value: it discloses that channel mappings and per-channel clip gains shift down and that the underlying audio files are not deleted. It omits reversibility, required permissions, and error behavior for out-of-range indices, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence stating the action, then two short clauses on effects; no filler. The GUI analogy is compact and aids comprehension, though it is optional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter mutation with no output schema and no annotations, the description covers the action and its downstream effects adequately. It stops short of addressing failure modes or persistence, which keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so project/track/channel are already documented, including the 0-based indexing. The description reiterates the channel-index semantics but adds no syntax or format beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Removes one channel ... from a track') plus the index base, and the scope phrase 'one channel' implicitly separates it from the sibling remove_track. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never states when to use this versus remove_track, remove_clip, or cleanup_regions, and names no alternative or precondition. It is purely a behavioral description, leaving tool selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_clipRemove clipsA
Removes the clip at a timeline position (one clip; ambiguity across tracks needs track), or every placement of a named region. delete_region also drops the region itself unless other clips use it.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | Timeline position of the clip (exclusive with region) | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| track | No | Track id | |
| region | No | Region name; removes every placement | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| delete_region | No | Also delete the region from the pool |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that this is a destructive removal, that 'region' removes every placement, and that 'delete_region' additionally drops the region from the pool unless other clips use it. It does not cover permissions, reversibility, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences front-load the primary action and then pack in the mode distinction and the 'delete_region' caveat. Every clause conveys useful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter destructive tool with no annotations and no output schema, the description covers the essential selection logic and the key side effect of 'delete_region'. It does not address edge cases such as no matching clip, multiple matches without 'track', or whether the operation is undoable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the position-vs-region choice, the need for 'track' when there is ambiguity across tracks, and the conditional effect of 'delete_region' when other clips still use the region.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it removes a clip. It also distinguishes the two removal modes (timeline position vs. named region) and clarifies the one-clip scope, so an agent can tell exactly what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains the two usage modes: use 'at' to remove one clip, or 'region' to remove every placement. It also notes that ambiguity across tracks requires 'track'. However, it does not compare against sibling tools like move_clip or cleanup_regions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_trackRemove trackA
Removes a whole track from the project, including its channels, clips and regions. Track ids are positions in the track list, so the ids of the tracks below shift up by one. The audio files stay in the package.
| Name | Required | Description | Default |
|---|---|---|---|
| track | Yes | Track id | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses the cascading destruction (channels, clips, regions), the side effect on track ids, and that audio files are preserved. It omits error behavior and whether the project file is written to disk immediately, but the essential mutation semantics are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the destructive action, followed by the id-shift caveat and the reassuring file-preservation detail. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive tool with no annotations and no output schema, the description covers scope, side effects, and non-destructive aspects of the operation. It does not address failure modes or return values, but the critical information an agent needs before calling it is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning beyond the schema by explaining that track ids are positional indices in the track list and that ids below the removed track shift up by one. That directly informs how to interpret and reuse the 'track' parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Removes a whole track') and immediately scopes what that encompasses (channels, clips, regions), which cleanly distinguishes it from sibling tools like remove_clip and remove_channel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The destructive scope is implied, but there is no explicit when-to-use guidance or reference to alternatives such as remove_clip or remove_channel, nor any stated prerequisite (e.g., project must be open/valid). Usage must be inferred from the purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
report_bugReport a bugA
Reports a bug in Audionaut to the maintainer as a GitHub issue. Use it when a tool misbehaves: an error that should not happen, a wrong or inconsistent result, a crash, or a project left in a bad state. Quote the exact tool call and error text - the CLI's code and message - and say what you expected instead; that is what makes the report reproducible. Never include the audio itself, and keep file paths to what is needed. Tell the user you are sending it; include their name or e-mail only if they offered it. If nothing could be sent, the reply carries a prefilled issue link to hand to the user.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | No | How to reproduce it: the tool calls in order, with their arguments | |
| title | Yes | One-line summary of the problem | |
| context | No | Project shape (tracks, clips, sample rate), platform, anything else relevant | |
| expected | No | What should have happened instead | |
| reporter | No | The user's name or e-mail for follow-up, only if they offered it | |
| description | Yes | What happened, including the exact error text if there was one |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations this description carries the full burden, and it does disclose meaningful behavior: the report is sent to the maintainer as a GitHub issue, the user must be told, reporter identity is optional, and on failure a prefilled issue link is returned. It omits auth/rate-limit or duplicate-suppression behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Six sentences, each carrying a distinct instruction: purpose, trigger conditions, content requirements, privacy constraints, user notification, and failure fallback. Front-loaded with the primary action and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a reporting tool with no annotations and no output schema, the description covers what an agent needs: reason to call it, what to supply, privacy limits, and the handling of the failure path (prefilled link). Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters, giving a baseline of 3. The description adds value by emphasizing which content matters most for reproducibility ('Quote the exact tool call and error text ... and say what you expected instead'), but it does not map that guidance to specific field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Reports a bug in Audionaut to the maintainer as a GitHub issue'), and the sibling set makes the distinction clear versus request_feature (feature requests) and the mutation tools. An agent can immediately tell what this does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates the triggering conditions ('an error that should not happen, a wrong or inconsistent result, a crash, or a project left in a bad state') and gives negative guidance ('Never include the audio itself, and keep file paths to what is needed'). That is when-to-use plus when-not-to-use in one place.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_featureRequest a featureA
Sends a feature request to the Audionaut maintainer as a GitHub issue. Use it when a task needs something the other tools cannot do: a missing verb or option, a limit that got in the way, or a workflow that took a clumsy workaround. Say what the user was trying to achieve and why the current tools fall short - that context is what makes the request actionable. Tell the user you are sending it; include their name or e-mail only if they offered it. If nothing could be sent, the reply carries a prefilled issue link to hand to the user.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | One-line summary of the requested feature | |
| context | No | What you were trying to do, which tool fell short, and any workaround used | |
| reporter | No | The user's name or e-mail for follow-up, only if they offered it | |
| description | Yes | What is missing, why it matters, and how it should behave |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses the external side effect (GitHub issue), privacy handling for reporter data ('only if they offered it'), and failure behavior (a prefilled issue link is returned). It omits auth or rate-limit details, but those are minor for this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core action, then usage conditions, then submission and failure behavior. No wasted phrasing; each sentence adds a distinct piece of guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no annotations and no output schema, the description covers what gets created, what to include, the privacy constraint, and the fallback path when submission fails. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all four parameters. The description adds real meaning on top: it tells the agent what to put in the context field (what the user was trying to achieve and why tools fell short) and how to treat reporter (only if offered), which goes beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: sends a feature request to the maintainer as a GitHub issue. This is clearly distinguishable from most siblings, though it never explicitly differentiates itself from the sibling report_bug tool, leaving a small ambiguity between a feature request and a bug report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions: a missing verb or option, a limit that got in the way, or a workaround. It does not name the alternative report_bug or state an exclusion for bugs, so the routing guidance is clear but not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
separate_stemsSeparate stemsA
Splits a clip into Drums/Bass/Other/Vocals tracks with the Demucs (htdemucs) source separator, aligned with the source clip. Slow: minutes for a full song, CPU only. Needs the model weights, which the Audionaut app downloads once (Settings > Separation); fails with model_missing until then, and with demucs_unavailable in builds without Demucs.
| Name | Required | Description | Default |
|---|---|---|---|
| clip | No | Playlist item id (default: the track's first clip) | |
| track | No | Track id (default 0; imported audio lands on a new track) | |
| project | Yes | Path to the .audium project package (absolute paths recommended) | |
| threads | No | Parallel segments (default: physical cores) | |
| mute_source | No | Mute the source track's channels (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses performance characteristics (slow, CPU only), a prerequisite (downloaded model weights), and two precise failure modes (model_missing, demucs_unavailable) with their causes. It also notes output alignment with the source clip.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and output, then adds the cost and failure conditions in two dense sentences. No filler or repetition; every clause adds operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no annotations or output schema, the description covers the operation, its results, runtime cost, prerequisites, and error codes. An agent has everything needed to call it and to interpret likely failures.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds context about what the separation produces but no syntax or default details beyond what the schema already documents for clip, track, project, threads, and mute_source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (splits) and resource (clip), names the exact output tracks (Drums/Bass/Other/Vocals), and identifies the model (Demucs htdemucs). It is clearly distinct from siblings like split and analyze.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong context: it is slow (minutes for a full song, CPU only) and requires model weights downloaded via Settings > Separation. It does not explicitly compare against alternatives like split, but the performance and prerequisite cues are enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_regionEdit regionA
Renames a region and/or retrims its source-relative range (clamped to the source audio). Clips have no length of their own, so a retrim affects every clip using the region.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | New source-relative end (exclusive with length) | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| start | No | New source-relative start | |
| track | No | Track id, to disambiguate same-named regions | |
| length | No | New length, keeping the start | |
| region | Yes | Name of the region to edit | |
| rename | No | New name (must not collide with an existing region) | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses that the range is clamped to the source audio and, importantly, that clips have no length of their own so a retrim affects every clip using the region – a non-obvious side effect. It stops short of covering permissions or what the response returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the primary action plus its most important consequence (retrim affects every clip) are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the key behavioral risk (shared-clip side effect) and the clamping constraint. It could add rename-collision behavior and required-permission context, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters (including unit, start/end/length exclusivity, and rename collision rules). The description only adds the source-relative/clamping framing, which is marginal beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (renames, retrims) and the resource (a region's source-relative range). It clearly contrasts with sibling create_region/cleanup_regions by implying an edit of an existing region, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'renames a region and/or retrims' – the agent can infer this edits an existing region rather than creating one. However, there is no explicit when-to-use vs. alternatives guidance (e.g., versus create_region or cleanup_regions) and no stated prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
splitSplit clipsA
Splits the clip under the given timeline position into two on every track that has one there (the GUI's Split command). Fails with nothing_to_split when no clip spans the position.
| Name | Required | Description | Default |
|---|---|---|---|
| at | Yes | Timeline position to split at, e.g. 23 for bar 23 | |
| unit | No | Unit for positions (default bars; bars/beats are 1-based, seconds/clocks absolute) | |
| project | Yes | Path to the .audium project package (absolute paths recommended) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden: it discloses that the split applies across multiple tracks and names the specific error code returned on failure. It still omits whether the split is undoable or what permissions/project state are required, so it is not fully complete for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the operation and then the failure mode, with no redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, yet the description covers the action's effect, its cross-track scope, and its error behavior, which is sufficient for a simple mutation. Only minor gaps remain, such as undo semantics or required project state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so 'at', 'unit', and 'project' are already documented in the schema. The description's phrase 'given timeline position' only restates the schema, adding no syntax or format detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb (splits), the resource (the clip under a given timeline position), and the scope (on every track that has a clip there), which clearly separates it from siblings like move_clip or remove_clip.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Anchors the tool to a familiar equivalent ('the GUI's Split command') and tells the agent the failure condition (nothing_to_split when no clip spans the position), but gives no explicit guidance on when to prefer this over alternatives such as create_region or separate_stems.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
22 tool updates
v0.2.0- First observed
analyze - First observed
assemble - First observed
auto_edit - First observed
cleanup_regions - First observed
clip_fades - First observed
clip_gain - First observed
clip_speed - First observed
create_project - First observed
create_region - First observed
export_audio - First observed
get_project_info - First observed
import_audio - First observed
move_clip - First observed
place_clip - First observed
remove_channel - First observed
remove_clip - First observed
remove_track - First observed
report_bug - First observed
request_feature - First observed
separate_stems - First observed
set_region - First observed
split
TDQS
Scored across 22 tools
Most tools target clearly distinct resources and actions, from project creation to clip gain/fades and stem separation. The main overlaps are request_feature vs report_bug (same GitHub issue mechanism, different intent) and the referenced but missing delete_region, which complicates remove_clip/cleanup_regions boundaries.
The set mostly follows a readable verb_noun pattern: get_project_info, create_project, place_clip, remove_track, report_bug. Minor deviations exist for clip_gain, clip_speed, clip_fades, and one-word tools like analyze/split, but these remain understandable.
22 tools is on the high side but reasonable for a DAW-like audio project server covering project, track, clip, region, analysis, export, and support operations. A few operations could likely be consolidated, but most tools earn their place.
Core workflows are covered: create/import/export, analyze, edit clips and regions, and remove tracks/channels. However, notable gaps remain, including no explicit track/channel create or rename operations, no project tempo/metadata update tool, and a delete_region operation referenced in remove_clip but absent from the tool list.
Maintenance
Related MCP Connectors
MCP server for Producer/Riffusion AI music generation
MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Create, co-edit, analyze, publish, and export collaborative step-sequencer sessions through MCP.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceMCP server that enables AI agents to read and edit Ableton Live sets, including tracks, clips, MIDI notes, devices, and scenes.1-
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to programmatically edit, analyze, and export audio projects through MCP tools, including multi-track editing, effects, transcription, and semantic search.1-
- AlicenseBqualityCmaintenanceThis MCP server enables AI assistants to control a live REAPER DAW instance, including transport, tracks, FX, MIDI, media, markers, rendering, and project state, with an escape hatch for arbitrary ReaScript commands.40MIT
- FlicenseAqualityBmaintenanceMCP server for controlling REAPER digital audio workstation from natural language. It enables AI agents to create tracks, import audio, adjust mix parameters, and manage sessions through typed operations and a REAPER bridge.22-