sublevel
Sublevel is an MCP server that lets Claude steer a Live set's low end toward a genre/reference target by measuring, planning, applying, verifying, and undoing changes—plus bass design, reference management, monitoring, mastering checks, and project scaffolding.
Session & targeting:
session,set_target(genre/reference/artist/favorites/file),find_references,add_reference,finishline(acceptance bands from finished records).Measurement & analysis:
measure,compare,readiness(club-floor checks),mastering_check,detect_key,read_notes,find_kicks,find_bass_patches,samples,project_reference.Mix mapping & monitoring:
map_mix,mixmap,monitor(periodic read-only readouts).Planning & guarded writing:
plan,apply(measure/verify/keep-or-revert),iterate(measure-plan-apply loop),verify,automate.Bass design & placement:
design_bass(headless Vital or in-Live search),deliver_patch,write_bassline,place_sample.Project setup:
scaffold,prepare_track(device chains on roles/tracks).Safety & history:
undo,revert_all,ab(A/B bypass),rate,history(change log with verdicts).Job control:
job(poll long-running operations),cancel.
Provides a Max for Live device and panel that runs the same low-end measurement, targeting, fixing, A/B, undo, and change-log workflow inside Ableton Live, with every action going through the same guarded engine as the MCP tools.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@sublevelcompare my kick and bass to a Tech House reference and tell me what's missing"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sublevel
A reference-matching low-end fixer for Ableton Live: an MCP server Claude drives and a Max for Live device on the master, both over one engine, backed by a corpus of 798 analyzed electronic tracks.
Mixing low end without accurate monitoring is guessing. Everything below about 150Hz is inaudible on laptop speakers and misrepresented on consumer headphones. Sublevel replaces the ear with measurement where measurement works, and says how confident it is by placing every number against records you actually listen to.
What makes it trustworthy
There is not a single invented threshold in this repo. A mix is not "thin" because a diff crossed -3dB; it is thin because it sits below p10 of tracks in its own genre. Every claim carries its percentile so it can be checked.
sub vs body: you -3.42dB (p8), reference 3.92dB (p82), -7.34dB.
Kick has punch but no weight underneath it.The corpus is scoped by genre, because the genres genuinely differ - House carries about 2dB less sub relative to kick body than Tech House or Minimal, and that is not a rounding artifact.
Related MCP server: Ableton Cookbook MCP
Status
Built and verified on a real set: measurement core, 798-track corpus with
section-aware finish-line profiles, genre-scoped calibration, the low-end and
club-ready planners, the guarded engine (measure, apply, verify, keep or
revert, with undo and A/B), headless Vital patch design, bassline writing,
the MCP server (36 tools) and the Max for Live panel. ARCHITECTURE.md has
the phase table; 251 tests, including the architecture rules, pass.
Needs you, once, in Live: python -m sublevel install-remote-script, then
select AbletonMCP as a Control Surface (Preferences > Link, Tempo & MIDI;
Input and Output both None) and restart Live.
Setup
python3.13 -m venv .venv
.venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytestThe MCP server is registered for this project in .mcp.json (python -m sublevel mcp). See ARCHITECTURE.md for the v2 layout and the tool contracts.
The Max for Live device
The same loop from inside Live, without Claude: aim at a genre or a record, measure, fix the low end, A/B, undo, read the changes.
.venv/bin/python -m sublevel m4l build # writes Sublevel.amxd into Live's User Library
.venv/bin/python -m sublevel serve # the HTTP shim the panel talks to, on 127.0.0.1:8766Drop Sublevel.amxd on the master (it passes audio through). The device is
a web view over surfaces/m4l/ui.html; every button is a POST to
/api/<tool> that calls the same function the MCP tool does, so the panel
cannot do anything Claude cannot, and nothing Claude cannot undo. The panel
and the MCP server each hold their own session and share the one Remote
Script socket: run one at a time.
Rebuilding the calibration
python -m sublevel.corpus extract ~/Downloads/beatport_tracks_* --out calibration/corpus.json
python -m sublevel.calibration build calibration/corpus.json
python -m sublevel.calibration show calibration/thresholds.json --genre "Tech House"Extraction is resumable and keyed by content hash, so re-running skips what is already analyzed. 798 tracks takes about 14 minutes on 5 workers.
MCP tools
The agent-facing surface (python -m sublevel mcp). Long operations return
a job id; poll with job.
tool | what it does |
| the open set: tracks, chains, target, budget, bypass state, jobs |
| genre / reference / artist / favorites / file; acceptance bands from the drop of finished records |
| loop-aligned capture; distance and gaps to the target; the club floor |
| which track owns which band |
| pure plan, guarded write, the loop; goals |
| the undo stack, the A/B bypass, what it sounded like, the change log |
| read-only readout every N bars while you play |
| what the master's limiter and compressors are contributing |
| the sample library and preset bank against the target's kick and bass |
| design a bass toward the reference (Vital offline, or a rack in Live in place) and put it on the bass track |
| the key, a clip's notes, a bassline under the kick, clip automation |
| the corpus and the named library |
Sound design
.vital presets are plain JSON with human-named parameters, and pedalboard
renders Vital offline, so the measurement engine doubles as a fitness function:
write a patch, render a note, measure, adjust - about 200 candidates a minute,
entirely headless, with only the winner going into Live. Targeting a real
tech-house bass preset from Vital's init patch converges in ~25 seconds.
Corpus percentiles are deliberately not used here. They describe finished, mastered records; a solo'd bass patch judged against them would produce nonsense with a confident percentile attached.
Safety
Sublevel writes to Live through exactly one path, Engine.apply, and every
write is measured before and after. A move is kept only if the objective
improved without costing kick punch, crest, mono compatibility or headroom;
otherwise it is put back. Every kept and reverted move is appended to
calibration/changes.jsonl with device, parameter, old value, new value, the
reason, the proposal it came from and the verdict. undo, revert_all and
the ab bypass are always available.
Two notes on the Ableton bridge:
The Remote Script Live runs is the upstream ableton-mcp script, patched and vendored at
src/sublevel/live/remote_script/. Upstream binds to0.0.0.0; the vendored copy binds to127.0.0.1.The third-party
abletonMCP server in.mcp.jsonhas telemetry that records the originating prompt and an opt-in dataset upload. It is registered withDISABLE_TELEMETRY=1; consent is yours to give.
What it does not do
It does not judge groove, sound design, arrangement, or whether a record is good. It speaks with confidence about sub, weight, punch, width and dynamics - the parts monitoring problems hide - and reports midrange as context rather than instruction, because most midrange problems are masking and arrangement, where the fix is removing a layer and no analyzer can tell you that.
Notes on the corpus
All 798 tracks are 320kbps MP3. That is transparent in the bottom two octaves,
so sub, punch, dynamics and mono correlation are trustworthy. It is not
transparent at the extremes: codec reconstruction overshoots sample peaks
(every track in this corpus reads above 0dBTP), and everything is lowpassed
around 20kHz. true_peak_db and the high band are marked advisory
wherever they surface, and are never presented as calibrated fact.
Available Tools
36 toolsabC
A/B by ear: bypass=true writes back every kept move's old value (the set as it was); bypass=false puts sublevel's moves back in. The undo stack is untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| bypass | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does disclose that the undo stack is untouched and explains the side effects of both bypass values. However, the mechanics behind 'writes back every kept move's old value' and 'puts sublevel's moves back in' are unexplained, leaving the actual state modification ambiguous and incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences and contains no wasted filler, so it is concise in length. However, it is dense and not well front-loaded: 'A/B by ear' is unclear and the long first sentence mixes two mode definitions without separating the core purpose. Conciseness is present, but clarity of structure is lacking.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description leaves out essential context: what 'A/B by ear' means, what a 'kept move' is, what a 'sublevel' is, and when this tool should be used relative to siblings like undo, compare, or revert_all. An output schema exists, so return values need not be described, but the operational context is too incomplete for an agent to reliably invoke this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the schema only states that 'bypass' is a required boolean with title 'Bypass'. The description compensates by defining both bypass=true and bypass=false behaviors, which is exactly what an agent needs to choose the right value. The definitions are still phrased in unstated domain jargon, preventing a perfect score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives specific behaviors - 'writes back every kept move's old value' and 'puts sublevel's moves back in' - so it is more than a tautologyaint, but the high-level purpose of 'A/B by ear' is never defined. The terms 'kept move', 'old value', 'sublevel', and 'the set as it was' are domain jargon that an agent cannot decode into a clear operation. It does not clearly distinguish itself from sibling tools beyond noting that the undo stack is untouched.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to invoke this tool versus alternatives like undo, revert_all, or compare. It implies two modes of operation (bypass=true and bypass=false), but never states the conditions under which an agent should choose this tool. No alternatives, exclusions, or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_referenceA
Ingest a local audio file into the reference library (and optionally tag it as a favourite so set_target(kind="favorites") uses it).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| path | Yes | ||
| favorite | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It states the primary mutation (adding a local file to the library) and the optional favorite side effect, but it does not say whether the file is copied or moved, whether duplicates are allowed, whether the operation is reversible, or what happens on invalid paths or unsupported audio formats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence front-loads the main action and uses a parenthetical for the optional favorite behavior. There is no filler, and every phrase conveys useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter add operation, the description provides enough to attempt a basic call: a path and a favorite flag. But complete invocation still depends on knowing name semantics, file handling behavior, and possible failure modes, none of which are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for 'path' by calling it a local audio file and for 'favorite' by explaining the set_target(kind='favorites') implication, but it never clarifies the optional 'name' parameter, leaving a real semantic gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Ingest a local audio file into the reference library.' It also differentiates itself from siblings by clarifying that this tool adds content, while find_references and set_target operate on or consume the reference library.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by 'ingest a local audio file into the reference library,' and the parenthetical ties favorite tagging to set_target. However, it does not explicitly state when to prefer this tool over find_references, project_reference, or place_sample, nor does it mention exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
applyA
Write a proposal to Live. With verify (default) it measures before and after and keeps the moves only if the objective improved without costing punch, crest, mono or headroom; otherwise it puts them back. Runs as a job: poll job(job_id). dry_run only runs the policy filter.
| Name | Required | Description | Default |
|---|---|---|---|
| verify | No | ||
| dry_run | No | ||
| proposal_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the behavioral disclosure. It explains the verify mechanism (measures before/after, keeps moves only if objective improves without certain costs, otherwise reverts), the async job execution, and the dry_run behavior. This is comprehensive and goes beyond a simple mutation, covering rollback and polling requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. The main purpose is front-loaded, and the critical behavioral details (verify, dry_run, job polling) are packed efficiently. Every sentence contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the operation (mutation with verification, rollback, async job), the description is complete. It tells the agent it needs a proposal_id, explains the two boolean parameters, and tells it to poll job(job_id). An output schema exists, so return format is covered elsewhere. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains verify (default behavior) and dry_run (only runs policy filter), and the proposal_id is self-explanatory from context. It does not explicitly define proposal_id but its role is obvious from the description's first sentence, so most parameters are effectively covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Write' and the resource 'a proposal to Live', which is specific and distinct from most siblings. It does not explicitly name alternative tools to differentiate itself, but the action is unambiguous enough for an agent to understand what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool by explaining it applies a proposal and runs as a job, instructing the agent to poll the job. It also explains dry_run's purpose, but it does not explicitly state when not to use this tool versus alternatives like verify, job, or undo. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
automateB
Write automation for one device parameter into a session clip as steps: points = [{"time": beats, "value": v}, ...]. Needs the extended Remote Script.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | No | ||
| clear | No | ||
| track | Yes | ||
| device | Yes | ||
| points | Yes | ||
| parameter | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden of disclosing side effects. It explains the structure of points and the external dependency, but it does not state whether existing automation is cleared, whether the operation is destructive, or how it interacts with undo/revert behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence followed by a minimal code example and a short prerequisite clause. Every element earns its place, and the main verb and resource are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too sparse for an agent to reliably perform a complete invocation. It includes the points format and a prerequisite, but leaves most parameter semantics and the call context unspecified, making it difficult to know the intended values for track, device, parameter, slot, and clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by defining parameter meanings. It only explains the points format (time in beats, value), while track, device, parameter, slot, and clear remain undocumented. The partial explanation is insufficient for a 6-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (write automation) and target (one device parameter into a session clip), and it clarifies the data format with the points array. It does not explicitly differentiate from sibling tools, but the resource and action are concrete enough to distinguish the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage guidance is the prerequisite 'Needs the extended Remote Script,' which signals a necessary environmental condition. It does not state when to prefer this tool over alternatives or mention exclusions, so the guidance remains implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancelB
Stop a running job between steps.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does add a meaningful detail — cancellation happens between steps, not mid-step — but it does not mention side effects, reversibility, or what happens if the job is not running or not between steps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, minimal sentence with no filler. Every word contributes meaning, and the key scoping condition ('between steps') is included without bloating the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema, this is a minimally viable description. However, it omits guidance on alternatives, doesn't clarify what happens in invalid states, and relies entirely on the schema for parameter semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention 'job_id' at all. The schema only indicates a required string titled 'Job Id', so the description adds no additional meaning about how to identify the job or where the ID comes from.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action ('Stop') and the resource ('a running job'), and adds a distinguishing nuance ('between steps'). It does not explicitly differentiate from sibling tools like 'job' or 'undo', but the verb and object make the primary purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'running job between steps' implies the intended usage context: only apply this to a job that is currently running and at a step boundary. However, there is no explicit statement of when not to use it or which sibling tool should be used instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compareA
Compare a measurement (default: the latest) with the target: per-axis gaps to the acceptance band, the 20-320Hz curve error per band, and what each gap means in plain words.
| Name | Required | Description | Default |
|---|---|---|---|
| measurement_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does not state whether the operation is read-only, what happens if no measurement or target exists, or how 'latest' is determined. It only lists output content, omitting side effects, dependencies, and potential error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently conveys purpose and outputs. It is concise and readable, though the list of outputs could be more structured for even quicker scanning. Overall, it is appropriately sized with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavior and default case, and an output schema exists to document return structure. However, it lacks information on preconditions (e.g., a target must be set) and error handling. For a simple tool with one optional parameter, this is mostly complete, but the missing preconditions and failure modes are notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It explains that the optional measurement_id defaults to null and that null means the latest measurement, which is not evident from the schema alone. This adds meaningful context for the only parameter, covering it well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (compare a measurement with the target) and specifies the exact outputs: per-axis gaps to the acceptance band, 20-320Hz curve error per band, and plain-word explanations. It distinguishes itself from related sibling tools by focusing on comparing to a target rather than measuring or monitoring.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not mention when to use this tool versus alternatives like 'verify', 'monitor', or 'measure'. It explains the default behavior (latest measurement) but provides no guidance on selection criteria or exclusions, leaving the agent to infer appropriate usage from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deliver_patchA
Put a designed Vital candidate into Live: written as an AU preset, loaded
onto the "Sublevel Bass" track (or track) with a 4-beat clip at the root note.
| Name | Required | Description | Default |
|---|---|---|---|
| track | No | ||
| candidate_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose concrete behavioral effects: the candidate is written as an AU preset, loaded onto 'Sublevel Bass' or the supplied track, and a clip is created. However, it does not disclose potential overwrites of existing clips, preset file locations, undo behavior, or whether other session state is affected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with no wasted words. The main action is front-loaded and the key scoping constraints ('Sublevel Bass' track, clip length, root note) are included without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values do not need explanation. Still, there are no annotations, no prerequisites, and no statement about whether the action overwrites prior delivery work or can be undone. For a mutation tool in a larger workflow, this leaves some gaps the agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description does add meaning by identifying candidate_id as the designed Vital candidate and `track` as an override for the default 'Sublevel Bass' track. It does not explain how to obtain candidate_id, expected ID format, or track semantics beyond the schema titles, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific action verb ('Put... into Live') and names the exact resource and outcome format ('designed Vital candidate', 'AU preset', 'Sublevel Bass' track, 4-beat clip at root note). It clearly distinguishes itself from sibling design and placement tools without needing schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use deliver_patch versus siblings such as design_bass, apply, or place_sample. The description implies it is a delivery/finalization step but gives no explicit context, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_bassA
Design a bass patch toward the target's bass profile (or a few words such as "deep round tight"). backend "vital": offline search, then deliver_patch loads it into Live. backend "live": searches the instrument on the "Sublevel Bass" track in place (creating the track with a Drift bass if missing) by writing its macros, firing a clip and recording. Runs as a job.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| brief | No | ||
| steps | No | ||
| passes | No | ||
| backend | No | vital |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It openly discloses side effects: creating the Sublevel Bass track with a Drift bass if missing, writing macros, firing a clip, recording, and loading the patch into Live. It also states that the operation 'runs as a job,' which is an important asynchronous behavior. This is transparent enough for an agent to predict non-read-only effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it states the action first, then details backend behavior. Every sentence adds either functional or side-effect information, though the backend branches are packed into dense prose that could be split into bullets.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the five optional parameters, no annotations, and two backend execution paths, the description covers the high-level behavior and side effects well. The output schema can explain return values, but the missing parameter definitions and the lack of a note about job lifecycle make it incomplete for correctly invoking the tool with custom settings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Brings `backend` to life by explaining what `vital` and `live` each do, and ties `brief` to the 'few words' target profile. However, `note`, `steps`, and `passes` are left completely unexplained, and with 0% schema coverage the description only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the core action explicitly: 'Design a bass patch toward the target's bass profile' and adds the free-text brief option. It clearly separates from searching/loading by naming `deliver_patch`, but it does not explicitly differentiate itself from `find_bass_patches`. The two backend modes also clarify scope, though a direct sibling comparison would make it a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains actionable context for choosing between the `vital` and `live` backends and explains what each does. It does not, however, state when to use this tool instead of related siblings like `find_bass_patches` or `deliver_patch`, nor does it mention prerequisites. The usage guidance is implied by the purpose rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect_keyC
The key: from the notes of a clip (the bass track's, or track's) when
there are enough, else from the last capture's audio, else the reference's.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | No | ||
| track | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It does disclose meaningful fallback behavior: notes first, then last capture audio, then reference. However, it does not describe failure behavior, side effects, or what happens when none of the sources are available.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with no wasted words. The fallback chain is compact, though somewhat awkwardly phrased and could be clearer with better punctuation or bullets.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return shape, but the description still leaves important gaps: no explanation of `slot`, no usage guidance, and an undefined threshold for 'enough'. It is adequate for a simple detection tool but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only partially explains the `track` parameter ('the bass track's, or track's'). The `slot` parameter is never explained, and the meaning of 'enough' notes is left undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys that the tool determines a musical key and specifies its data sources, but it never uses an explicit verb like 'detect' or 'return'. It is more a statement of how the key is derived than a clear statement of what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to call this tool vs. alternatives among the sibling tools, and names no exclusions or preconditions. The only implied usage is from the tool name itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_bass_patchesB
Rank the Vital preset bank against the target's bass profile (sub/body, brightness, sustain). Seeds for design_bass.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions ranking and seeding but does not disclose whether this is a read-only analysis, whether it modifies the session, how many results are returned, or what 'seeds' means operationally. The behavior of ranking and seeding is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the core action ('Rank the Vital preset bank') and includes the key context (target profile dimensions and downstream tool). Every word earns its place; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema and only one optional parameter, so the surface area is small. However, the description omits critical context: what the output looks like, how 'n' behaves, whether the ranking requires a prior target setup, and what 'seeds for design_bass' means in practice. For a tool that feeds another tool, more guidance is needed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description does not explain the 'n' parameter at all, leaving the agent to infer that it likely controls the number of ranked results. The default of 5 is in the schema, but the description adds no semantic meaning beyond the schema's bare property definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Rank') and resource ('Vital preset bank') against a defined target profile, and names the downstream tool it seeds ('design_bass'). It is clear about what the tool does, though it does not explicitly distinguish itself from sibling tools like find_kicks or find_references beyond the bass-specific context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: it is used to rank bass patches against a target bass profile and seeds design_bass. However, it does not explicitly state when to use this tool versus alternatives, nor does it mention prerequisites like having a target profile set or a preset bank loaded.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_kicksA
Rank every kick one-shot on the machine against the target's kick profile
(fundamental, decay, click, sub/body); pack narrows to one library. Needs a
target with a profile (a reference, or a genre with the finish line built).
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| pack | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the behavioral burden. The verb 'rank' implies a non-mutating analysis, and the description explains the matching basis and the pack filter. However, it does not say whether this operation is purely read-only, whether it can be slow over 'every kick', or what happens if no target profile exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loads the core action, and every phrase adds information. The second sentence adds a necessary prerequisite. It is slightly dense, with jargon packed into the parenthetical, but still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description does not need to explain return values flavors. It covers purpose, the pack parameter, and the target-profile prerequisite. But it omits any clarification of `n` and does not mention failure or readiness conditions, so coverage is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only explains `pack` ('narrows to one library'). The `n` parameter is completely undocumented beyond its schema default of 10, leaving the agent to infer that it likely controls result count. This is a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Rank every kick one-shot') against a defined criterion (the target's kick profile), and lists the profile dimensions. This clearly distinguishes it from sibling tools like find_bass_patches or find_references.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides an explicit prerequisite: a target with a profile (a reference or a genre with the finish line built). It also explains how the pack parameter narrows scope. However, it does not explicitly contrast with alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_referencesD
Records to aim at. like: nearest neighbours of a named track. character: words such as "heavy sub, tight punch, pumping".
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| like | No | ||
| genre | No | ||
| character | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does not mention whether the tool is read-only, requires prerequisites, mutates state, or has any side effects. No limitations or requirements are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short but cryptic. It lacks front-loading of the main purpose, uses unclear phrasing, and fails to structure the information logically. Conciseness without clarity is not effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is incomplete for a 4-parameter tool with an output schema. It does not mention return format, usage scenarios, or how it integrates with the other reference tools. An agent would struggle to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains 'like' as 'nearest neighbours of a named track' and 'character' with examples, but leaves 'n' (likely the number of results) and 'genre' completely unexplained. Since schema coverage is 0%, the description must compensate but only partially covers two of four parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Records to aim at' is vague and lacks a clear verb or resource. It hints at finding similar records via 'like: nearest neighbours of a named track' but does not explicitly state it is a search/find tool. It does not distinguish itself from siblings like find_kicks or find_bass_patches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. No exclusions, conditions, or context for selecting this over sibling tools like project_reference or find_kicks are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
finishlineA
What a finished record measures in the drop, for a genre / family / tempo band (default: the current target's scope): acceptance bands per axis, kick and bass profiles, arrangement, pump. From the deep corpus.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose what data it draws from ('deep corpus') and the default scoping behavior, which is useful behavioral context. However, it does not explicitly reveal how it behaves on missing scope beyond 'default: the current target's scope' and does not state output format or side effects (though output schema exists, mitigating part of that gap).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, front-loading the core purpose and listing the concrete contents returned. The wording is slightly awkward as a noun phrase rather than an imperative sentence, but no words are wasted and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has low complexity: one optional parameter, no required parameters, and an output schema exists so return value details are unnecessary in the description. The description explains scope meaning and default, lists what is measured, and names the source corpus, which is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the name 'Scope' with 0% description coverage. The description compensates by explaining that scope maps to a genre/family/tempo band and defaults to the current target's scope. This adds the key meaning missing from the schema, justifying a score above baseline; it loses one point for not elaborating on acceptable scope string formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies the resource and content clearly: it conveys what a finished record measures in the drop, i.e., acceptance bands per axis, kick and bass profiles, arrangement, and pump, sourced from the deep corpus. It is not a tautology and gives enough detail to distinguish it from siblings like monitor or measure, though it lacks an explicit action verb such as 'returns' or 'retrieves'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: to get finish-line benchmark measurements for a genre/family/tempo band or the current target's scope. It does not name explicit alternatives or exclusion criteria, and it does not explicitly state when to choose this tool over the many sibling measurement/comparison tools such as compare or readiness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
historyB
The change log for this set: every write with old/new value, reason, proposal id and verdict (kept / reverted / undone / rated).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden for side-effect transparency. Calling it a change log strongly implies a read-only operation, and the listed verdict types give useful context, but it does not explicitly state that it has no side effects or describe how the limit parameter affects the returned log.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the resource type and then lists the distinguishing entry fields. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema and a single optional parameter, the description is mostly sufficient, covering what entries contain and their verdict types. It is incomplete regarding the effect of limit and does not explicitly state the read-only nature of the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage for the only parameter, limit, is 0%, and the description never mentions limit or its meaning. The description therefore fails to clarify whether limit controls the number of log entries and how results are ordered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as the change log for the toolset and enumerates the fields carried by each entry (old/new value, reason, proposal id, verdict), which tells an agent what it will get. It lacks an explicit verb such as 'list' or 'get', and while it is distinguishable from undo/revert_all/rate by being a log, it does not directly contrast with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for reviewing past writes and their verdicts, which is useful context. However, it never explicitly says when to call history instead of undo, revert_all, or rate, nor does it mention exclusions or ordering of the log.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
iterateB
The loop: measure, plan, apply, verify, repeat until the goal is met or the policy's round cap. goal: "lowend" or "club_ready". Runs as a job.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | lowend | |
| track | No | ||
| rounds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden. It mentions 'Runs as a job' (implying asynchronous execution) and the termination condition, but omits side effects, state changes, reversibility, error handling, or progress monitoring. This is insufficient for a loop tool that likely modifies project state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise—two sentences with no filler. The core loop concept is front-loaded, and the goal values are stated efficiently. It earns high marks for economy, though the brevity contributes to the lack of detail in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that orchestrates a multi-step loop with three parameters, the description is notably incomplete. It lacks explanations for 'track' and 'rounds', does not clarify the meaning of goal values, and provides no guidance on monitoring or interpreting results. The presence of an output schema mitigates return-value ambiguity but not the missing parameter semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain all parameters. It only mentions 'goal' with possible values, but does not explain what 'lowend' or 'club_ready' mean, nor does it address 'track' or 'rounds' at all. The description adds minimal value beyond the schema, leaving most parameters ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: it runs a loop of measure, plan, apply, verify until the goal is met or the round cap. It specifies valid goal values ('lowend' or 'club_ready'), which distinguishes it from sibling tools that perform individual steps. This is a specific verb-resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage as the high-level orchestration loop, contrasting with sibling tools like measure, plan, apply, verify. However, it does not explicitly state when to use this tool versus calling those steps individually, nor does it mention any prerequisites or exclusions. The guidance is implicit, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobA
Poll a job: status, progress lines after since, and the result when done.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | ||
| job_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses that it is a polling operation, returns status, progress, and result, and supports incremental updates via `since`. However, it does not mention error handling, potential server load from frequent polling, or whether the job is automatically cleaned up after completion. The behavioral traits are partially disclosed but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly packed sentence that front-loads the core action ('Poll a job') and then lists the key information returned. Every word contributes to understanding, with no filler or redundancy. It is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a single required parameter and an output schema that likely describes the result structure, the description covers the essentials: it tells the agent what it does and what it returns. However, it doesn't describe polling frequency recommendations or how to interpret status values. The output schema might cover return values, but the description remains incomplete for edge cases like job errors or completion behavior, which could matter for an agent deciding when to poll.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions `since` as a progress line filter, but does not explain the format of `job_id` or the semantics of `since` beyond 'progress lines after.' The default of 0 is in the schema, but the description doesn't clarify that it starts from the beginning. Overall, the description adds some insight but not enough to fully cover the parameters, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool polls a job and specifies what it returns: status, progress lines after `since`, and the result when done. This distinguishes it from sibling tools like `apply` or `cancel`, though it doesn't explicitly name alternatives. The resource (`job`) is clear and the verb `poll` is specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: it is for checking ongoing job status, with the `since` parameter for incremental progress. However, it does not state when not to use it or mention alternative tools (e.g., maybe `monitor` for real-time status). It gives clear context but lacks explicit exclusions or comparisons.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_mixA
Work out which track owns each band by pulling each one 12dB down in turn and re-measuring the master from the same bar. Needed before the planner can correct the low-end curve on the tracks that carry it. Live must be playing with the loop brace on. Runs as a job (one capture per track).
| Name | Required | Description | Default |
|---|---|---|---|
| seconds | No | ||
| only_tracks | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the gain-pull measurement method, the live-loop requirement, and the job-based execution with one capture per track. It does not state whether the 12dB pulls are permanent or reverted, which is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, informative sentences: method and purpose first, prerequisite second, execution model third. No filler and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema provides return shape, the description supplies the needed context: why to run it, what it does, what must be true before running, and how it executes. The missing parameter documentation is the only notable incompleteness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions 'seconds' or 'only_tracks'. The agent cannot infer valid values, units, or track-selection semantics from the tool text, so this is a significant gap even though both parameters are optional.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific diagnostic procedure: determine which track owns each band by pulling each track down 12dB and re-measuring the master. This is much more precise than a generic 'map mix' and differentiates it from measurement siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly says this is needed before the planner can correct the low-end curve, and it explicitly requires Live to be playing with the loop brace on. It does not name alternative tools or exclusions, but the invocation context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mastering_checkA
Bypass the master's limiter / compressors, measure, restore, and say what they contribute: loudness added, crest and punch lost, whether the level is earned by the mix or borrowed from the ceiling. Runs as a job.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states 'Runs as a job' and 'restore', but it does not clarify whether the operation is destructive, whether the bypass is temporary, or what side effects occur on the session. The term 'restore' implies reversion, but it is not explicit, and there is no mention of permissions or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core action and immediately lists the analytical outputs. Every clause adds value, with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core functionality and expected outputs, and the presence of an output schema helps. However, it lacks guidance on when to invoke this tool versus siblings, and it does not clarify the exact side effects or prerequisites (e.g., whether a master bus must exist). For a job-based tool with no parameters, this is a notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is nothing for the description to add beyond the schema. Per calibration, a baseline of 4 is appropriate when no parameters exist; the description does not attempt to explain nonexistent inputs and stays focused on behavior and outputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Bypass'), names the resource ('master's limiter / compressors'), and details the exact outputs ('loudness added, crest and punch lost, whether the level is earned...'). It clearly differentiates from generic tools like 'measure' or 'compare' by specifying its unique analysis and restoration behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case (analyzing master bus processing) but does not explicitly state when to use it versus alternatives like 'measure' or 'compare'. It mentions 'Runs as a job' which hints at asynchronous execution, but there are no explicit 'when not to use' or alternative routing instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
measureA
Record the mix from the top of the loop and measure it: the low-end axes, dynamics, stereo, resonances, and the club-readiness floor. Live must be playing with the loop brace on. Returns a measurement id for compare().
| Name | Required | Description | Default |
|---|---|---|---|
| bars | No | ||
| section | No | loop |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It states that the tool records and measures the mix, and returns a measurement ID for compare(). It also specifies a runtime requirement (Live playing with loop brace). However, it does not clarify whether the tool is read-only, whether it modifies the session, or what happens to previous measurements. For a measurement tool, this is partially transparent but lacks explicit side-effect disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long and efficiently packs the core action, the specific measurements, a precondition, and the return value into a compact form. It is front-loaded with the main purpose and avoids fluff. Slightly denser than necessary, but still well-structured and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so the return value is presumably documented elsewhere. The description covers the purpose and a key precondition, and mentions the return ID's use case. However, it omits parameter semantics and alternative usage contexts. For a tool with two optional parameters and a defined output, the description is adequate but not comprehensive—it leaves the agent guessing about parameter meaning and alternative workflows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description does not explain the 'bars' or 'section' parameters at all. The parameters are marked optional with defaults, but their meaning remains undocumented in both the schema and description. The description mentions 'from the top of the loop,' which could hint at the 'section' parameter, but it is not explicit. This is a significant gap for an agent trying to set appropriate parameter values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: recording and measuring the mix, enumerating specific aspects (low-end axes, dynamics, stereo, resonances, club-readiness floor). It also mentions the return of a measurement ID for compare(), distinguishing it from sibling tools like map_mix or monitor. The verb 'measure' plus the detailed scope leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes a critical precondition: 'Live must be playing with the loop brace on.' This tells the agent when the tool can be used. However, it does not provide explicit guidance on when to prefer this tool over alternatives (e.g., 'use compare() after measuring'), nor does it mention any exclusions or alternative scenarios. The mention of compare() implies a workflow, but the when-to-use guidance is mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mixmapB
The stored mix map for this set: per track, how much of each band it carries.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It conveys that this is existing stored state and describes the data granularity, implying a non-mutating retrieval. However, it does not explicitly state that the tool has no side effects, how freshness works, or how it relates to current session state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler, and the core object is front-loaded. It could be slightly improved by phrasing it as an explicit operation, such as 'Returns the stored mix map...', but it remains concise and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless getter with an output schema, the description covers the essential meaning of the data. It is not fully complete because it lacks an explicit action verb and any indication of when to prefer this tool over map_mix, leaving minor room for misinterpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the description cannot add parameter-level detail. The baseline for a parameterless tool is 4, and the description appropriately focuses on the meaning of the data being returned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies what the tool exposes: the stored mix map for the current set, including per-track band distribution. It is specific about the resource and data, but lacks an explicit verb like 'get' or 'returns,' so it does not fully differentiate the read behavior from sibling map_mix.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this tool versus alternatives such as map_mix or monitor. 'Stored' weakly suggests a read-only operation, but the description does not state when to use it or when to use a sibling instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
monitorA
A live readout while you play: every interval_bars, capture (without
seeking) and report distance and gaps against the target. Read-only.
action "start" returns a job id whose progress lines are the readout;
"stop" cancels it. windows=0 runs until stopped.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | start | |
| windows | No | ||
| interval_bars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does admirably: it declares read-only behavior, 'without seeking', the return of a job id, progress lines as readout, and cancel semantics. These go well beyond the sparse schema and make side effects predictable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences, with the core purpose front-loaded and no filler. Every clause adds useful information: read-only, no seeking, cadence, job lifecycle, and windows semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and the tool's moderate complexity, the description is complete: it covers the operational lifecycle, cadence, and termination condition. Nothing needed to invoke or stop monitoring correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three parameters. It explains action ('start'/'stop'), interval_bars (capture interval), and windows (=0 runs until stopped). Each parameter receives meaningful semantic context that is absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb-resource pair: it provides a 'live readout' of 'distance and gaps against the target' while playing. It also clarifies the read-only nature and the start/stop job model, which distinguishes it from one-shot siblings like 'measure' or 'compare'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: use it 'while you play', run continuously with windows=0, and stop via the 'stop' action. It does not explicitly name alternative tools or say when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
place_sampleB
Put an audio file into Live: a new audio track (or track) gets a clip
in slot with the sample. Logged like any other change.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| slot | No | ||
| track | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It does disclose that the action mutates Live by creating a track/clip and that the change is logged like other changes, which is useful context. However, it does not mention validation behavior, what happens when track is null vs specified, or any file-related constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence. It communicates the action, the target context, and the logging behavior without wasted words or redundant schema repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has an output schema, so return-value explanation is not needed. Still, the description omits preconditions, track-creation behavior, and slot bound expectations. It is adequate but not fully self-sufficient for an agent deciding how to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the roles of 'track' and 'slot' in broad terms, but leaves 'path' and the null/default semantics of track underspecified. The agent still has to infer key parameter meaning from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific action ('Put an audio file into Live') and clearly states the resulting state: a track gets a clip in the given slot. This distinguishes it from sibling tools like add_reference or find_kicks without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case but gives no explicit guidance on when to choose this tool versus alternatives such as add_reference or samples. There are no prerequisites, exclusions, or routing cues beyond the general action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
planA
Plan moves toward the target without touching Live. goal: "lowend" steers the low end into the target's band; "club_ready" fixes whatever the club floor failed (mono, headroom, resonances, loudness, plus the low-end axes). track: force the moves onto one track index instead of the mix map's owner. Returns a proposal id for apply(), the moves, anything refused, and the devices it would need added.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | lowend | |
| track | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and discharges it well: "without touching Live" discloses non-mutation, and mentioning "anything refused" and "devices it would need added" reveals refusal and side-condition behavior. It does not address persistence, auth, or failure modes, but the key safety trait is explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first clause, and the goal:/track: layout is structurally scannable with no filler. Slight deduction for jargon density that forces the agent to infer domain meanings.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (so return values needn't be re-explained), the description covers both parameters, both goal modes, non-mutation, and refusal/device-addition behavior. Remaining gaps are undefined domain terms ("club floor", "mix map") and error/failure behavior, which are minor against the strong param coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description fully compensates: it enumerates goal's valid values with their distinct effects and explains track's override-of-owner behavior. This is exactly the compensation needed when the schema is silent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Plan moves toward the target without touching Live" states a specific verb and resource, and the explicit reference to a "proposal id for apply()" differentiates it from the apply sibling. Minor deduction for domain jargon ("club floor", "mix map's owner") that assumes insider knowledge an agent won't have.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives meaningful context for goal selection (“lowend” vs “club_ready”) and the track override semantics, and it implies the plan→apply workflow via "returns a proposal id for apply()." However, it gives no explicit when/when-not guidance against other planning/verification siblings like iterate, verify, or monitor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_trackB
Give a track the chain finished projects put on its role (kick: Saturator, Drum Buss, EQ Eight; bass: Saturator, EQ Eight, Glue; master: Glue, Saturator, EQ Eight, Limiter, Utility). role is read from the track name unless given. Adds only what is missing, in order, at the end of the chain.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | ||
| track | Yes | ||
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses key behavior beyond annotations, specifically idempotence (adds only what is missing, in order, at the end) and role fallback from track name. However, dry_run defaults to true yet the description never mentions dry-run semantics, which is important for whether a call actually mutates the chain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences carry meaning without filler; the core behavior, role maps, role fallback, ordering, idempotence, and chain position are all included efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return-value explanation is unnecessary. The description is sufficient for a domain-expert agent to understand the core operation, but lacks clarity on what track numbers refer to and whether dry_run actually prevents mutation, leaving real invocation gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description helpfully explains role semantics (read from track name unless given, unless explicitly provided), but track is unexplained as an integer identifier and dry_run behavior is left entirely to inference from its name/default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a concrete behavior with resource (track chain) and specific role-to-chain mappings, and explains role fallback logic. It is not a tautology and likely distinguishes prepare_track from vague siblings, though it does not explicitly name siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: prepare a track by appending the standard role chain for kick, bass, master, with role inferred unless explicit. It gives no explicit when-to-use vs alternatives or when-not-to-use guidance, so usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_referenceB
What finished projects are built like: with no name, the profile across all indexed .als files (what sits on kicks, basses, buses and the master, tempo, length); with a name, that project's tracks and chains.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does describe what the tool returns (profile, tracks, chains) and the two modes, which is useful. However, it does not state whether the tool is read-only, if it has any side effects, requires permissions, or has performance implications. Since it appears to be a query tool, the lack of explicit read-only declaration is a minor gap, but the description adds some value by detailing the output content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence but packs in the two distinct behaviors and the details of the output. It is reasonably concise and front-loaded with the core purpose ('What finished projects are built like'), but the opening phrase is a bit verbose and could be streamlined to 'Returns the build profile of finished projects'. No extra fluff is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only one optional parameter and an output schema (which presumably documents the return structure), the description is fairly complete. It explains both modes and the key aspects of the output (what sits on kicks, basses, buses, master, tempo, length; tracks and chains). It does not mention any prerequisites (e.g., projects must be indexed) or limitations, but these are not critical given the tool's simplicity. Overall, it provides enough context for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It does this effectively by explaining that the optional 'name' parameter, when null/absent, returns the global profile, and when set, returns that project's tracks and chains. This adds clear semantic meaning beyond the bare schema definition, though it does not specify the exact format or source of the name (e.g., project ID vs. title).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: it returns a profile of finished projects, either across all indexed .als files (when no name is given) or for a specific project (when a name is given). The verb is implicit but the resource and the two distinct modes are specified, which distinguishes it from siblings like find_references. However, the phrasing 'What finished projects are built like' is indirect and could be more explicit about the action (e.g., 'get project profile').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the conditional behavior based on the name parameter but does not provide guidance on when to use this tool versus alternatives. It does not mention when not to use it, such as 'use find_references to locate projects first' or 'this is for finished projects only'. The context of when this tool is appropriate is only implied through the description of its output, not stated explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rateA
Record what a kept proposal sounded like: verdict "good" or "bad", with an optional note. Goes into the change log next to the numbers.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| verdict | Yes | ||
| proposal_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It goes beyond the schema by stating the result goes 'into the change log next to the numbers,' which discloses the persistence side effect. It does not detail overwrite behavior or prerequisites, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loads the core action, and omits filler. Every clause earns its place: verdict values, optional note, and the change-log destination.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter tool with an output schema, the description provides the necessary verdict semantics and side-effect context. The main remaining gap is defining what 'kept proposal' means precisely, but this is likely domain context an agent would already have.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the description compensates by clarifying the verdict parameter's allowed values ('good' or 'bad'), noting that the note is optional, and implicitly linking proposal_id to the kept proposal. This adds real meaning beyond the bare schema field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Record') and a clear resource ('a kept proposal'), and states exactly what is recorded: a verdict of 'good' or 'bad' plus an optional note. This clearly distinguishes it from sibling tools like history, compare, or verify.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear context for use: after a proposal has been kept)Skip, you record how it sounded. It does not explicitly name alternatives or exclusion criteria, but the intended scenario is concrete enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
readinessB
The club floor on a measurement (default: latest): seven pass/fail checks plus the DJ-mixer boost test, each with its consequence and the fix.
| Name | Required | Description | Default |
|---|---|---|---|
| measurement_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It does not indicate whether the tool is read-only, whether it has side effects, or any requirements. It only describes output content. This is insufficient for a tool that could plausibly be read-only but is not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the main purpose and details. It is concise without unnecessary words, though it packs many details into one sentence which slightly reduces readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional parameter) and the presence of an output schema, the description covers the core purpose and output elements. However, it lacks usage context and behavioral disclosure, leaving some gaps for an agent to fully understand when and how to use it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter (measurement_id) with 0% description coverage, but the description clarifies that the default is 'latest' and that the measurement is the subject of the checks. This adds meaning beyond the schema, though it doesn't explain the parameter format or further details. The default null in schema aligns with 'latest' via the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific purpose: evaluating readiness via seven pass/fail checks plus a DJ-mixer boost test, each with consequence and fix. This distinguishes it from sibling tools like 'mastering_check' or 'measure' by focusing on readiness assessment with a defined set of checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by describing what the tool returns, but it does not explicitly state when to use it versus alternatives. For example, it doesn't say 'Use this before mastering' or contrast with 'mastering_check'. The context is implicit but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_notesD
The notes of a MIDI clip.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | No | ||
| track | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose any behavioral traits such as whether the operation is read-only, what it returns (despite an output schema existing), potential side effects, or error conditions. The description is too sparse to inform an agent about the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short, but it is under-specified rather than concise. It lacks essential details and is not front-loaded with actionable information. Every sentence should earn its place; this single sentence provides almost no value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there are two parameters (one required) and an output schema, the description is grossly incomplete. It does not explain how to select the MIDI clip (via slot/track), what the output contains, or any assumptions. An agent cannot reliably call this tool based on the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not explain 'slot' or 'track' at all, nor their roles or how they affect the result. The description adds no meaning beyond the raw schema names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'The notes of a MIDI clip' is a noun phrase without a verb, so it doesn't clearly state the action. It adds the context of MIDI clip but does not say what the tool does (e.g., 'read', 'retrieve', 'get'). It is nearly a restatement of the tool name 'read_notes' and does not differentiate from any siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no mention of context, prerequisites, or when not to use it. The description provides no decision support for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revert_allC
Put every parameter sublevel changed in this session back.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the basic action of reverting changes but does not reveal whether the operation is destructive, irreversible, or what the final state looks like. The term 'back' is vague, and side effects are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no fluff. It is front-loaded with the action and scope, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is incomplete for a state-modifying tool. It does not clarify what 'sublevel' means, the extent of 'this session', or any potential side effects or prerequisites. Given the presence of an output schema, return values may be covered, but the behavioral context is lacking, leaving the agent uncertain about invocation safety.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, so the description is the only source of semantic information. It clearly explains that the tool reverts all parameter changes made in the current session, which is meaningful and sufficient given the empty schema. The 0-parameter baseline applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a verb ('put back') and a resource ('every parameter sublevel changed in this session'), but 'parameter sublevel' is ambiguous. It clearly indicates a revert operation, but the exact scope and meaning of 'sublevel' are not defined, making it only partially clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not mention when to use this tool versus alternatives like 'undo', 'session', or 'history'. It implies reverting all session changes, but this is not stated explicitly, and no comparison or exclusion criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
samplesA
What the sample index holds: one-shots per kind, and the packs the kicks
come from. Build or refresh it with python -m sublevel samples index.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the tool can build or refresh the index via a command, implying a state-changing action, but it doesn't detail side effects, whether it reads or writes, or what the output schema contains. The description adds some context but leaves behavioral details vague.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose and ending with the build command. Every word earns its place; no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description is mostly complete. It explains the resource and the build/refresh action. It could be more explicit about whether the tool returns the index contents or just builds it, but the output schema likely covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics burden. The description explains what the index holds and how to build/refresh it, which is sufficient context for a no-argument tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explains what the sample index holds ('one-shots per kind, and the packs the kicks come from') and how to build or refresh it. It is clear about the resource and the action, though it doesn't explicitly name a sibling tool to distinguish itself from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when you need to know what the sample index contains or need to build/refresh it. It doesn't explicitly state when not to use it or name alternatives, but the context is reasonably clear for a zero-parameter informational tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scaffoldA
Set an empty project up like a finished one: tempo, KICK / BASS / DRUMS / PERC / SYNTHS / VOCALS / FX tracks with the device chains finished projects put on each role, two returns, and the master chain (Glue, Saturator, EQ Eight, Limiter, Utility). dry_run lists it; dry_run=false builds it. Groups have to be made by hand afterwards - the Remote Script cannot create them.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and provides valuable details: dry_run lists what will be built, dry_run=false actually builds it, and groups must be created manually. This goes beyond the schema. It does not state what happens if the project is not empty, which would be useful, but the disclosed behavior is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler: the main action and full resource list, the dry_run behavior, and a critical limitation. Every clause earns its place and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-boolean tool with an output schema, this description is largely complete: it covers action, scope, dry-run behavior, and a known limitation. The main missing piece is explicit handling of non-empty projects or prerequisites, but overall an agent can invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only provides a title and default for dry_run. The description adds practical meaning by explaining that dry_run lists the setup and dry_run=false executes the build. This significantly helps an agent understand the boolean's effect, which is especially important with 0% schema description coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Set an empty project up') and enumerates the exact resources created (tracks, returns, master chain). This is distinct from sibling tools like find_kicks or prepare_track, so an agent can identify what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the intended context ('Set an empty project up') but does not explicitly name when to use it over alternatives or when not to use it. It gives enough context for an agent to infer it is for initial project scaffolding, but no direct routing to related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionA
What sublevel knows about the open Live set: tracks and their device chains, the current target, budget spent, whether sublevel's moves are bypassed, open proposals and running jobs. Call this first.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It transparently frames the tool as a passive read of sublevel's internal knowledge, listing state categories rather than actions, which implies no mutation. It does not disclose details like cost, rate limits, or whether any cached data may be stale, so it falls short of being fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one information-dense sentence with a clear front-loaded subject and a concise list of covered state. Every phrase earns its place, and the imperative 'Call this first' is placed at the end as a natural follow-on. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool takes no parameters and has an output schema, so the description only needs to explain what the call represents and when to call it. It lists the full scope of session knowledge and gives usage direction. An agent has everything needed to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and no input schema properties, so there is no parameter meaning to explain. Baseline for a no-parameter tool is 4, and the description adds no unnecessary parameter noise. Nothing is missing here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a query of sublevel's knowledge about the open Live set, enumerating specific content: tracks, device chains, target, budget, bypass state, proposals, and jobs. It does not use a strong imperative verb like 'get' or 'retrieve', but 'What sublevel knows...' is unambiguous. The detail helps distinguish it from state-related siblings like history, readiness, and monitor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs 'Call this first,' giving clear when-to-use guidance for establishing session context. It does not mention alternatives or when not to use it, but for a session-introspection tool at the start of a workflow this is sufficient context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_targetB
Choose what the mix is steered toward.
kind: "genre" (value = genre name, or blank for the session's best guess), "reference" (value = a record name from the library or corpus, e.g. "Eating Glue"), "artist" (value = artist name; median of their corpus tracks), "favorites" (median of references tagged favourite), "file" (value = path). Returns the target's acceptance ranges (p25-p75) per axis and any warning that the chosen record is an outlier for its genre.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | genre | |
| genre | No | ||
| value | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden. It does disclose meaningful behavior: value semantics per kind, median aggregation for artist/favorites, and returned acceptance ranges plus outlier warnings. It does not state whether the target is persisted or applied to the current session or how it interacts with session/undo, leaving some behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is compact, front-loads the core purpose, and uses a scannable bullet list. The only minor issue is that the schema parameter names could have been aligned more explicitly with the bullet items.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers target kinds, value semantics, and the return payload shape despite the output schema being present. It is incomplete regarding the `genre` parameter, side effects on the current mix/session, and how 'blank' is represented, which matters for an agent with no annotations and 3 free-form parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds substantial meaning for `kind` and `value`, explaining valid values and per-kind behavior; schema coverage is 0%, so this compensation is necessary. However, the `genre` parameter is never explained, and its relationship to `value` for genre targets is ambiguous, leaving an agent unsure which field to populate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence uses a specific verb ('Choose') and a clear resource ('what the mix is steered toward'), and the bullet list concretely defines the kinds of targets. It does not explicitly differentiate itself from sibling tools like project_reference or find_references, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: an agent should call this when it wants to set the mix target, and the kind list indicates when each target form applies. However, there are no explicit conditions for when to use this versus alternative tools, no prerequisites, and no mention of when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
undoB
Undo the last n kept proposals (most recent first).
| Name | Required | Description | Default |
|---|---|---|---|
| n | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. 'Undo' implies mutation, but there is no mention of whether the action is destructive, reversible, or scoped beyond 'kept proposals'. The 'most recent first' ordering is helpful but does not address side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, focused sentence that front-loads the action and includes the key scoping detail. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool, the description conveys the essential operation. However, because this is a mutating 'undo' tool with no annotations and no usage context, an agent might still be uncertain about destructive consequences or when this tool is the right choice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain the parameter. It does this by connecting 'n' to 'the last n kept proposals', which gives meaningful semantics beyond the bare integer and default value. No range or constraints are stated, but the core meaning is clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Undo'), a resource ('kept proposals'), and a count ('n'), so an agent can tell what the tool does. It doesn't explicitly name sibling tools, but the 'kept proposals' scope helps distinguish it from broader revert tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus alternatives such as revert_all or history. The usage is only implied by the action 'undo', with no mention of prerequisites, limitations, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verifyA
Re-measure now and compare with the baseline taken before the first apply.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden for behavioral disclosure. It does explain that the tool re-measures and compares against a stored baseline, which suggests a non-mutating verification action. It does not state explicitly whether any state is saved, whether the operation is read-only, or what happens if no baseline exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence communicates the action and the comparison target with no filler or redundancy. Every word contributes to the tool's meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter tool with an output schema, the description is nearly complete: it identifies when to call it and what it compares against. It is slightly thin on preconditions such as what would happen if no baseline has been recorded or if no apply has been run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the baseline score of 4 applies. The description adds a useful semantic context—measuring against a baseline—but no parameter-level documentation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete action—re-measure now and compare with a baseline—rather than merely repeating the name 'verify'. It is clear enough to be distinguished from siblings like 'measure' or 'compare', though it does not explicitly name them or state what verification confirms.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before the first apply' implies this is meant for use after an apply step, giving a clear temporal context. However, it never explicitly says 'use this after apply' or contrasts it with sibling tools like 'measure' or 'compare', so the agent must infer the exact selection condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_basslineA
A bassline that sits under the kick: root from the reference's bass (or the
key, or root as a MIDI pitch), kick grid from a drum clip or the last capture.
style: sub_hold, rolling_offbeat, follow_kick, two_note. dry_run shows the
piano roll; dry_run=false writes it to slot slot of the bass track (created
if missing) or track, replacing what is there.
| Name | Required | Description | Default |
|---|---|---|---|
| bars | No | ||
| root | No | ||
| slot | No | ||
| style | No | sub_hold | |
| track | No | ||
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does a good job: it discloses root selection priority, style options, dry_run behavior, track creation, and that existing content is replaced. The only notable omission is what the bars parameter controls, but the main write and replace semantics are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, and nearly every clause adds information. It is a single dense run-on rather than clearly separated sentences, but there is little waste and the key behavioral details are packed efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers side effects, destination selection, root resolution, and style options, which is strong for a tool without annotations. However, the bars parameter is unexplained, the term 'last capture' is not defined, and there is no guidance on style selection or interaction with sibling bass-design tools. The output schema presumably covers return values, so that part is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain the parameters, and it explains most of them: root, style, dry_run, slot, and track. The bars parameter is entirely undocumented, which is a real gap, but the description adds substantial meaning beyond the bare schema for the other five parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool produces: a bassline synced to the kick, with configurable root, style, and write target. It lacks an explicit verb at the start and does not differentiate itself from sibling tools like design_bass or find_bass_patches, but the resource and behavior are clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when this tool is used—when a bassline under the kick is needed—but gives no explicit guidance on when to choose it over alternatives. It also does not mention when to use dry_run versus a direct write, leaving that decision to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
36 tool updates
v0.1.0- First observed
ab - First observed
add_reference - First observed
apply - First observed
automate - First observed
cancel - First observed
compare - First observed
deliver_patch - First observed
design_bass - First observed
detect_key - First observed
find_bass_patches - First observed
find_kicks - First observed
find_references - First observed
finishline - First observed
history - First observed
iterate - First observed
job - First observed
map_mix - First observed
mastering_check - First observed
measure - First observed
mixmap - First observed
monitor - First observed
place_sample - First observed
plan - First observed
prepare_track - First observed
project_reference - First observed
rate - First observed
read_notes - First observed
readiness - First observed
revert_all - First observed
samples - First observed
scaffold - First observed
session - First observed
set_target - First observed
undo - First observed
verify - First observed
write_bassline
TDQS
Scored across 36 tools
Tools are grouped into clear functional clusters—reference, measurement, mixing moves, and job control—and most have distinct resource-action purposes. The measurement family (measure/compare/verify/monitor) and revert family (undo/revert_all/ab) could be confused, but their scopes are different enough that a careful agent should select correctly.
Most actions use an imperative verb_target pattern and information tools are named as plain nouns (session, history, samples, mixmap), which is a consistent semantic convention. A few bare verbs or abbreviations like ab, job, and rate break the pattern slightly, but the all-snake_case style keeps it predictable.
At 36 tools, this is well beyond the 25-tool 'too many' threshold, making the surface heavy even for a broad DAW-assistant domain. Several tools could be consolidated into grouped subcommands—especially the measurement, reference, and job-related tools—without losing clarity.
The set covers the full workflow: references and target selection, measurement and comparison, planning/apply/iterate, undo and history, plus creative helpers like bassline writing, patch design, and project scaffolding. Minor gaps like reference-library removal/editing exist, but core workflows are complete with no dead ends.
Maintenance
Related MCP Connectors
- mozonicOAuthcom.mozonic
AI mixing and mastering: analyze your mixes, run DSP autofix, render stems, and master tracks.
Audio features + harmonic set-building for tracks by name/ISRC. Spotify audio-features replacement.
AI music production assistant — audio profiling, AI mixing sessions, and service inquiries.
Audio mastering for AI agents: LUFS/True Peak targets, Suno/Udio AI-fingerprint removal.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceEnables AI-assisted electronic music production in Ableton Live, with built-in genre theory for indie dance, tech house, melodic techno, and house music.6-
- AlicenseNot gradedqualityCmaintenanceBridges AI assistants with Ableton Live, enabling real-time control, offline project analysis, version tracking, and rack/preset parsing for music production workflows.MIT
- AlicenseBqualityBmaintenanceEnables AI-powered music production in Ableton Live through natural language, with tools for composition, arrangement, mixing, and sound design.52MIT
- AlicenseAqualityAmaintenanceStyle-aware auto-mixing and mastering for Ableton Live, providing tools for analysis, preview rendering, and release checks.1044 PyPI3MIT