Skip to main content
Glama
synopsys0

PostFader V10 — FL Studio MCP Server

PostFader — FL Studio MCP server

The AI copilot for FL Studio

Your AI can finally work inside FL Studio.

PostFader connects Claude, Codex, Cursor, and other local MCP-compatible AI clients to the FL Studio project you already have open. Ask it to inspect your session, diagnose a mix, control loaded plug-ins, clean up routing, build MIDI parts, organize patterns and Playlist tracks, add section markers, or make supported changes from natural language.

All release assets · Setup guide · Explore what PostFader can do

V10: 134 tools · 8 live resources · Windows and macOS · Open source · No PostFader account

PostFader V10 (10.0.0) brings the complete 134-tool workflow to the Windows/macOS packages, Codex ZIPs, Claude Desktop MCPB, and Python distribution. See the V10 release notes for upgrade steps and evidence boundaries.

Starts read-only and never saves your project automatically. Native macOS plug-in loading, note inspection, and saved-project rendering are included; complete live qualification of those newer host paths remains pending.

What it can do · Workflows · Feature depth · Install · AI clients · Documentation

Unofficial community project; not made by or affiliated with Image-Line.

Related MCP server: flstudio-mcp

Not just another note sender

PostFader is a production layer for FL Studio—not only a way to send notes or change isolated controls. It gives your AI useful context from the project that is open now, plus tools to analyze exported audio, work with loaded plug-ins, compose musical parts, and carry separately requested supported changes back into the session.

🔎 Understand your project

🩺 Diagnose the mix

Inspect mixer routing, Channel Rack generators, loaded effects, patterns, Playlist tracks, transport, undo/redo history position, step sequences, and parameters exposed by loaded plug-ins.

Measure an exported bounce, compare it with a reference, examine vocal-versus-instrument masking, monitor peaks during playback, and surface evidence your AI can prioritize.

🎛️ Control the session

🎹 Create and transform music

Rename and color tracks, adjust levels and panning, manage sends and routing, control transport, organize channels and patterns, edit steps, and change supported loaded plug-in parameters.

Generate chords, melody, bass, and drums; export multi-track Type-1 MIDI; estimate tempo and key; transcribe monophonic audio; and prepare or transform Piano Roll material.

Workflows with PostFader

🔎 Understand the project already open

“Show me every instrument that is not routed to the mixer.”

“Which effects are loaded on my lead vocal?”

“Where are my drums routed, and which mixer inserts are peaking too high?”

PostFader reads FL Studio's current project state: mixer inserts and sends, loaded mixer effects, Channel Rack generators, patterns, Playlist track state, transport, undo/redo history bounds, the current step grid, and exposed plug-in parameters. Your AI can answer from the actual session instead of relying on a project description pasted into chat.

🩺 Diagnose an exported mix and decide what to improve

“What are the three highest-impact problems in this bounce?”

“Compare this mix with my reference and explain the biggest differences.”

“Check these vocal and instrumental exports for likely masking.”

Mix Doctor turns an exported bounce into producer-readable technical findings about level, dynamics, tonal balance, stereo behavior, and export readiness. Reference analysis compares loudness and tonal balance across aligned audio. Masking analysis uses synchronized vocal and instrumental renders to report possible spectral-overlap regions.

During a chosen observation window, a peak watch samples the mixer inserts included in the watch and remembers the highest level it observed for each. Use those results to build a gain-staging plan from a playback section or full pass that fits the window rather than a single instant. These are sampled observations, so they do not prove that every transient was captured. Your AI can use other reported findings to create a separate, reviewable mix plan.

🎛️ Control the session and the plug-ins you already use

“Rename insert 4 to Lead Vocal, color it purple, pan it 10% left, and confirm the changes.”

“Mute the backing-vocal tracks and lower the send to the reverb bus.”

“Find the feedback parameter on the delay that is already loaded and reduce it.”

PostFader can control mixer volume, pan, mute, solo, arm, selection, stereo separation, sends, and routing. It can change tempo, playback, loop mode, recording state, and song position; organize channels, patterns, and Playlist tracks; and edit step sequences.

For loaded effects and Channel Rack generators, PostFader asks FL Studio which parameters the plug-in exposes. It can inspect names and values, search or scan large parameter surfaces within explicit limits, set a known value, target the number a plug-in displays, or choose an exact named option. Bundled profiles for selected FL Studio stock effects add known parameter roles for supported workflows without pretending every plug-in has the same controls.

For those selected profiles, your AI can turn supported goals such as “tame harshness,” “control dynamics,” “limit peaks,” “shorten the reverb,” or “create a rhythmic echo” into matching parameter roles. Intent resolution is read-only; choosing values and applying a change remain separate steps.

On macOS, plugins_list_available reads FL's native Add menu and plugins_load adds an instrument or an effect on a specified mixer track, then identifies the new instance through FL's bridge. Removal and reordering remain unavailable, as does reliable effect-slot bypass or wet/dry control.

Plugin Atlas adds offline product knowledge for plug-ins whether or not an instance is currently loaded. Its bundled Image-Line catalog and selected third-party records describe purposes, techniques, limitations, and explicit stock alternatives. Atlas keeps that knowledge separate from runtime matching, control-adapter evidence, and the three honest availability states. See the Plugin Atlas guide or inspect the installed bundle with postfader-plugin-atlas.

🎚️ Choose a coherent sound palette

“Create a melodic bass track and choose all the sounds yourself.”

“Keep the lead in Drop B, but make the bass and texture feel bigger.”

Sound Selection turns that direction into a deterministic palette chosen from the generators and effects already loaded in the project. It can select a product and exact preset for each role, preserve core identity sounds, plan a section variation, map a drum kit's reported pads, and pass role targets into a Production Run. User preferences and exclusions always win; balanced planning uses bounded local recency only to distinguish similarly suitable choices.

Sound Selection does not use random preset roulette or pretend to hear FL Studio's output. It reads preset identity back after bounded navigation, keeps explicit local feedback and usage history separate from project state, and reports a concise blocker when the requested sound is not loaded. Load the instrument pool manually, then see the Sound Selection guide for examples and exact boundaries.

🎹 Compose, transform, and organize musical ideas

“Create an eight-bar D Dorian melody with a bassline and drum pattern.”

“Transpose this Piano Roll part up an octave and humanize the velocities.”

“Estimate the tempo and key of this sample.”

“Transcribe this monophonic melody into a reviewable note sequence.”

Generate deterministic chord progressions, melodies, bass parts, and drums from a musical brief; melody, bass, and drum generation also accept reproducible seeds. Export separate parts in one Type-1 MIDI file and verify the written file's structure and content digest. Audio tools can estimate tempo and a global major or minor key, while monophonic transcription creates a note sequence that can be reviewed and exported in a separate step.

PostFader can also find and prepare a pattern FL Studio reports as empty, add section markers, organize Playlist tracks, record a supported automation value, and inspect existing Piano Roll notes or prepare append, replace, quantize, transpose, humanize, duplicate, delete, or clear operations. Piano Roll application uses FL Studio's separate script workflow, so PostFader reports the evidence it actually has instead of claiming controller-side note readback.

From one request to a complete production workflow

TIP

You ask: “The vocal feels buried. Find the most likely cause, show me what you would change, and fix only the highest-confidence problem.”

PostFader workflow

  1. Reads the mixer, routing, and relevant loaded plug-ins.

  2. Analyzes the bounce or synchronized renders you provide.

  3. Reports possible level, tonal, dynamics, or synchronized-input masking findings.

  4. Your AI prioritizes the reported evidence and builds a reviewable plan from supported operations.

  5. Returns the proposed changes for a plan-only request, or uses the existing authorization when you asked it to fix the problem.

  6. Enables session writes once and applies the supported changes within that request's scope.

  7. Reports the observed result and any evidence limitation.

A narrow remote control stops at individual commands. PostFader connects those commands into a production workflow.

Production Runs: task-scoped autonomy

Production Runs let your connected AI turn one outcome-oriented request into a bounded, multi-stage plan. Ask it to finish a track, build around a loop, transform a genre while preserving a vocal, work only on one section, or mix without changing notes. The AI submits the structured plan; PostFader validates scope, resolves references, applies supported operations, and records truthful receipts.

Creation runs now begin with one silent readiness scorecard covering the live bridge, Piano Roll, generator pool, drum map, patterns, loaded processing, and known manual handoffs. A ready run keeps that bounded context through palette, composition, note application, processing, and finalization instead of rescanning the complete project before each change. Sound choices retain confidence and alternatives; generated notes can adapt to known articulation, envelope, register, and polyphony; supported loaded effects can be planned by semantic goal and applied through the existing verified setters.

Autonomy belongs to that request only—there is no permanent autonomous-mode toggle. A plan-only request never changes FL Studio. An authorized execution run enables the existing session write gate once, then continues until the submitted plan completes or reaches a real blocker. Earlier verified changes remain visible if a later operation fails; PostFader never claims rollback, replays an ambiguous mutation, or saves the project automatically.

The consolidated result reports technical execution, arrangement delivery, processing, manual handoff, and audible quality separately. Technical success never means PostFader heard or approved the song. See Creation pipeline, Production Runs, and Sound Selection.

See the Production Runs guide for chat examples, the supported MVP operation set, continuation and stop behavior, durable run lifetime, and FL Studio limitations.

Production Runs now survive MCP restarts: postfader_list_runs finds retained plans and receipts, and postfader_continue_run resumes remaining work. An interrupted operation with an unknown outcome is never replayed.

piano_roll_read_notes inspects existing note timing, pitch and expression through the Piano Roll script bridge without enabling musical edits. postfader_render_saved_project starts a separate WAV render from a saved .flp; inspect or cancel it through the render job tools. Unsaved live edits are not included. These new paths have synthetic coverage; live acceptance on FL Studio remains pending.

Genre requests now supply editable instrument-role defaults, with ten style profiles and eighteen instrument families. Explicit preferences stay in control.

Creation Review, Revision, and Delivery

After a Production Run creates a playable draft, export one bounce and ask the connected AI to review and improve it. PostFader validates and measures the selected file globally and by known song section, combines that evidence with your explicit feedback, protects accepted sounds or notes with independent locks, and compiles the smallest bounded revision into the existing Production Run executor. One revision pass uses one readiness preflight and one task-scoped write authorization.

Export the revised bounce with matching settings and PostFader can compare the two versions, report improvements and regressions separately, and prepare an exact Playlist, export, and delivery handoff. Technical measurements and section-energy proxies never substitute for your artistic approval. Sessions can persist locally without storing audio bytes or paths you chose not to retain. The review workflow does not save, create Playlist clips, insert plug-ins, or hear FL Studio's live output. See the Creation Review guide, which documents its 13 MCP tools and 9 corresponding Production Run operations.

Feature depth

Mix and finish

  • Run Mix Doctor on an exported bounce.

  • Compare loudness and tonal balance when candidate and reference inputs align and pass readiness checks.

  • Examine synchronized vocal and instrumental renders for possible spectral overlap.

  • Measure peaks, loudness, dynamics, tonal balance, and stereo behavior.

  • Watch sampled mixer peaks during a chosen observation window and build gain-staging plans.

  • Run a finish assessment and use your AI to turn selected recommendations into reviewable one-shot plans.

Control the session

  • Work with mixer inserts, sends, routing, and transport.

  • Inspect and organize Channel Rack generators, patterns, and Playlist tracks.

  • Read undo/redo history bounds and edit the current step sequence.

  • Discover and control supported parameters exposed by loaded effects and generators.

  • Plan and apply a coherent sound palette with exact preset verification, drum-pad mapping, continuity, and bounded novelty.

  • Read bundled Plugin Atlas product knowledge and compare it with observed loaded plug-ins without turning the catalog into a runtime allowlist.

Create music

  • Generate deterministic chord progressions; melody, bass, and drums also accept reproducible seeds.

  • Export multi-track Type-1 MIDI and verify the written content.

  • Estimate tempo and global major or minor key from an audio file.

  • Turn monophonic audio into a reviewable note sequence.

Edit and arrange

  • Append, replace, quantize, transpose, humanize, duplicate, delete, or clear Piano Roll material through the separate FL Studio script workflow.

  • Find and prepare patterns FL Studio reports as empty, add section markers, organize Playlist tracks, and record supported automation values.

Bring your own AI

  • Claude: Claude Desktop and Claude Code

  • Codex: CLI, IDE extension, and desktop Codex

  • Cursor: IDE and CLI

  • OpenCode

  • Grok Build

  • T3 Code through an MCP-capable provider

  • Other local hosts compatible with stdio MCP servers

More than a basic FL Studio MCP

To make the distinction concrete, the baseline below is deliberately defined as a local MCP with transport commands, individual controls, point-in-time reads, predefined parameter mappings, and note dispatch. It is not a survey of every other project.

Capability

Narrow baseline used here

PostFader V10

Play, stop, and change individual controls

Transport and individual controls

Yes, plus wider session workflows

Read the open project

Selected state only

Mixer, channels, loaded plug-ins, patterns, Playlist tracks, undo/redo history, steps, and transport

Diagnose exported audio

Not part of the baseline

Mix Doctor, peaks, loudness, tonal balance, stereo analysis, masking, and references

Monitor levels through playback

Point-in-time meter reads

Process-local per-insert peak watches sample across a chosen observation window

Move from diagnosis to a separate apply request

Not part of the baseline

Diagnose → propose → review → apply → report

Work with loaded plug-ins

Predefined controls

Runtime parameter discovery, exact controls, named options, and selected stock-effect profiles

Generate musical parts

Individual note dispatch

Chords, melody, bass, drums, transcription, and Type-1 MIDI export

Transform Piano Roll content

Not part of the baseline

Append, replace, quantize, transpose, humanize, duplicate, delete, and clear

Help organize an arrangement

Not part of the baseline

Pattern preparation, markers, Playlist track tools, and automation helpers

Install on Windows and macOS

Not part of the baseline

Guided platform packages for both

Work across AI clients

Single-host setup

Claude, Codex, Cursor, OpenCode, Grok Build, and other local MCP hosts

Check supported changes

Command dispatch only

Reads supported controls back from FL Studio after the change


Quick installation

  1. Download and extract the matching platform package.

  2. Run the guided installer and select your virtual MIDI endpoint.

  3. Complete the documented FL Studio MIDI Settings stage.

  4. Connect your local AI client.

Codex users can choose the dedicated Codex ZIP for guided codex mcp add registration. Claude Desktop users can add the .mcpb after completing the same platform setup. Advanced users can install the wheel or source archive.

NOTE

Codex ZIPs, the Claude Desktop MCPB, wheel, and source archive still use the Universal Bridge, a virtual MIDI endpoint, and the documented FL Studio setup. PostFader does not install virtual MIDI software.

Python 3.10–3.14 is required. Python 3.13/3.14 or Windows ARM64 may require a native compiler for python-rtmidi.

Supported AI clients

PostFader runs as a local stdio MCP server, so the AI host must be able to launch it on the same computer as FL Studio.

Client or host

V10 setup path

Claude Desktop

Use the Windows/macOS package and generated claude-json; the optional .mcpb is an additional Claude Desktop wrapper, not the platform setup.

Claude Code

Use the Windows/macOS package and adapt the generated claude-json server values to Claude Code's MCP configuration.

Codex CLI, IDE extension, and desktop Codex

Use a Codex ZIP, or run postfader setup --client codex-toml --register-codex from a Python/source install.

Cursor IDE and Cursor CLI

Put the resolved executable, arguments, and environment values in Cursor's mcp.json.

OpenCode

Adapt the resolved values to opencode.json or opencode.jsonc.

T3 Code

Configure PostFader in the MCP-capable provider T3 Code is using; no T3-specific package is shipped.

Grok Build

Configure the local stdio server in Grok Build's MCP settings; no Grok-specific package is shipped.

Other local MCP hosts

Adapt the generated server values to the host's local stdio schema.

Grok on the web and Grok Bot require a publicly reachable remote HTTP MCP server and cannot use PostFader's current local packages directly.

Supported systems

Component

V10 support

PostFader

V10 / 10.0.0

FL Studio

FL Studio 2026, version 26.1.3 build 5336 or newer; live evidence is limited to the tested systems below.

FL MIDI scripting API

Version 44 or newer

Python

3.10 through 3.14

macOS

Qualified on macOS 27.0 arm64 with FL Studio Producer Edition 26.1.3 build 5336 and the built-in IAC bus.

Windows

Qualified on Windows 11 x64 with FL Studio Producer Edition 26.1.4 build 5589.

The V10 release notes distinguish automated platform checks from the historical live qualification matrix and the remaining experimental host-adapter paths.

How it works

flowchart LR
    A["Your AI client<br/>(MCP)"] --> B["PostFader<br/>runs locally"]
    B -->|"Virtual MIDI"| C["Universal Bridge<br/>inside FL Studio"]
    C --> D["Your open project"]
    B -->|"Files you choose"| E["Exported audio<br/>analysis"]
    B --> F["Mix and creative<br/>workflows"]
    F -->|"Type-1 MIDI"| G["Generated MIDI files"]
    F -->|"Optional script"| H["FL Piano Roll"]

The AI client calls PostFader's named MCP tools. Live FL Studio communication travels over a local virtual MIDI endpoint to the Universal Bridge controller script. Audio tools analyze files you select because FL Studio's scripting API does not expose its live audio buffer.


Built for real projects without pretending FL Studio exposes more than it does

  • PostFader starts read-only; opening or reloading a project resets write access.

  • It never saves the project automatically.

  • Supported direct changes are read back from FL Studio after they are made.

  • Workflows with narrower evidence say so instead of reporting full verification.

  • PostFader runs locally, requires no PostFader account, and has no PostFader telemetry.

For the exact boundaries, read What “verified” means, the write response contract, the security policy, and FL Studio API limitations.

What “verified” means

For a supported direct change, PostFader checks the target and current session, asks FL Studio to make the change, waits for a later controller update, and reads the control again. verified: true means the requested control state was observed afterward. It does not prove the choice sounds good or guarantee an undo point or rollback.

Writes affect the open project immediately. Ask the AI client to enable write mode only when you want changes, use a blank or disposable project for the first write test, and disable write mode when you are done.

FL Studio MCP questions

What is PostFader?

PostFader is an open-source FL Studio MCP server. It connects a local AI client to FL Studio through the Universal Bridge and virtual MIDI, with additional file-based audio analysis and composition workflows.

Does it work on Windows and macOS?

Yes. The release includes standard and Codex setup ZIPs for both platforms, an optional Claude Desktop MCPB, and Python packages. FL Studio, matching bridge installation, and a bidirectional virtual MIDI endpoint are required.

Which features are released?

V10 (10.0.0) exposes 134 tools and 8 resources, including Plugin Atlas, Sound Selection, recoverable Production Runs, Creation Review, Piano Roll note inspection, macOS plug-in loading, and saved-project rendering. Read the release notes for experimental feature boundaries.

Does it upload my music or save my project?

PostFader has no hosted service or telemetry and never saves your project automatically. Audio analysis runs on files you select. Your AI client may send tool arguments and results to its model provider under that client's settings. Keep private project data and local run journals out of public reports.

Documentation

Guide

What it covers

Setup and troubleshooting

Full installation, virtual MIDI, bridge, client configuration, upgrades, and diagnostics

Tool contracts

All current tools and 8 resources, exact arguments, results, refusals, and evidence boundaries

Sound Selection

Producer direction, coherent palettes, exact preset verification, drum maps, local history, and Production Run references

Production Runs

Bounded execution, durable checkpoints, explicit resume, and recovery without replaying unknown writes

Creation Pipeline

Readiness, sound-aware composition, phase timing, and semantic processing

Creation Review

Bounce evaluation, explicit feedback and locks, one bounded revision, before/after comparison, persistence, and delivery handoffs

Plug-in support

Parameter discovery, option controls, scan limits, troubleshooting, and compatibility evidence

Plug-in matrix

Evidence definitions, validated reports, and the contributor target backlog

Plugin Atlas

Offline product knowledge, runtime/evidence boundaries, and Atlas CLI usage

FL Studio constraints

What FL Studio's scripting API allows and where PostFader stops

Distribution and listings

Published versions, development scope, and verified MCP directory status

Architecture

Components, transport, bridge behavior, resources, and trust boundaries

V10 release notes

Features, upgrades, platform support, and known live-validation gaps

Distribution and listings

Verified releases, canonical descriptions, and MCP directory status

Security

Threat model, local trust boundaries, privacy, and vulnerability reporting

Early-user activation

A privacy-safe first-session and return-session checklist

Contributing

Development workflow and contribution guidelines

GitHub Discussions

Setup help, workflow sharing, ideas, and plug-in compatibility conversations

Detailed limitations

PostFader does not currently:

  • guarantee rollback or an FL Studio undo point;

  • save an FL Studio project or render unsaved live state;

  • hear or capture FL Studio's live audio output;

  • insert plug-ins on Windows, or remove or reorder them on either platform;

  • reliably control an effect slot's bypass or wet/dry mix;

  • create, move, or delete Playlist clips through the public scripting API;

  • read section-marker times or recorded automation points back from FL Studio;

  • infer named competing project tracks from a full-mix masking measurement; or

  • turn technical measurements into objective artistic truth.

Batch application is bounded but not atomic: earlier changes are not rolled back if a later operation cannot be verified. A lost or ambiguous mutation response is never replayed automatically. The local virtual MIDI bus is shared and unauthenticated, so use PostFader on a trusted, single-user workstation.

Audio and mix tools analyze files you explicitly select or recent bounces found in bounded FL Studio folders. Results can include paths, hashes, and measurements, but never audio samples. Your AI client is separate software and may send tool arguments and results to its model provider; review that client's privacy policy before using sensitive projects.


Contributing

Read CONTRIBUTING.md for the development workflow. To help expand real-world plug-in evidence, review the definitions and target backlog in the plug-in matrix, then submit a privacy-safe plug-in validation report.

Community setup help, workflow notes, feature ideas, and compatibility results belong in GitHub Discussions.

Security reporting

IMPORTANT

Do not open a public issue for a suspected vulnerability. Follow the private reporting instructions inSECURITY.md.

License

PostFader is available under the Apache License 2.0. See NOTICE for attribution details.

Available Tools

134 tools
arrangement_add_section_markersA
Destructive

Add bar/beat section markers; name readback is available, marker-time readback is not.

ParametersJSON Schema
NameRequiredDescriptionDefault
markersYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
ppqYes
commandYes
verifiedNo
warningsNo
requestedYes
applied_atYes
project_savedNo
names_verifiedYes
pulses_per_barYes
schema_versionNo
times_verifiedNo
after_marker_namesYes
undo_point_createdNo
before_marker_namesYes
session_fingerprintYes
verification_statusYes
time_signature_numeratorYes
session_precondition_appliedNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation risk is known. The description adds a useful behavioral limitation not in the annotations: name readback works, marker-time readback does not. This goes beyond the structured data and helps set expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-structured sentence front-loads the core action and then provides a concise, useful caveat. Every phrase earns its place; no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema and annotations cover return values and safety profile, but the definition lacks parameter semantics and usage context. It is adequate for a relatively simple add operation, yet the 0% schema description coverage means more burden falls on the description, which it only partially meets.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain parameters, but it only mentions 'bar/beat' which vaguely maps to bar_number and beat_offset. It does not explain the marker object structure, constraints (max 32), or defaults, leaving the agent to infer semantics from field names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Add') and resource ('bar/beat section markers'), and the tool name disambiguates the arrangement context. It also adds a distinguishing behavioral note about readback capability, which helps separate it from any potential marker-related operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or alternative-routing guidance is provided. The readback caveat implies a condition ('use if you need name readback, not marker-time readback'), but there is no direct comparison to other tools or clear exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arrangement_prepare_patternB
Destructive

Find an FL-reported empty pattern, select it, name/color it, and set length.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
colorNo
length_beatsNo
start_pattern_numberNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
lengthNo
outcomeYes
identityNo
verifiedYes
warningsNo
selectionYes
applied_atYes
project_savedNo
pattern_numberYes
schema_versionNo
one_session_targetYes
rollback_attemptedNo
automatic_replay_attemptedNo
pattern_was_reported_emptyNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint: true, readOnlyHint: false, and openWorldHint: true. The description adds a useful operation sequence—find, select, name/color, set length—but does not disclose what happens if the FL-reported pattern is no longer empty, whether the selection change is significant, or how failures behave. With annotations carrying the safety profile, the description provides some value but not full behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It lists the workflow in execution order: find, select, name/color, set length. Every phrase contributes to the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive composite tool with four parameters and no parameter-level schema descriptions, this short description is not fully adequate. It omits start_pattern_number semantics, fail behavior when no empty pattern exists, explicit when-to-use guidance, and side-effect detail. The annotations and output schema offset some completeness, but the description alone leaves an agent uncertain about safe or correct use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description maps three of the four parameters to meaningful actions: name, color, and length. It does not mention start_pattern_number, and it does not explain the color integer format or the meaning of selecting a starting pattern. Since schema description coverage is 0%, the description partially compensates but leaves a real semantic gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a concrete multi-step operation: find an FL-reported empty pattern, select it, name/color it, and set its length. It is distinct from a tautology and conveys the resource and actions. However, it does not explicitly contrast itself with the closely related sibling tools like fl_find_empty_pattern, fl_select_pattern, fl_set_pattern_identity, or fl_set_pattern_length.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives the operational steps but provides no explicit guidance about when to invoke this composite tool versus calling the individual sibling tools. It does not mention any prerequisites, exclusions, or cases where this tool should not be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_analyze_fileA
Read-onlyIdempotent

Measure level, spectrum, dynamics, stereo, and optional pitch of one render.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to an existing audio file bounced from FL Studio.
max_secondsNoOptional shorter analysis bound in seconds; the default reads up to 600.
include_pitchNoAlso run the monophonic pitch tracker; useful for a lead vocal stem, unreliable for a full mix.

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileYes
pitchNo
stereoYes
dynamicsYes
loudnessYes
spectrumYes
confidenceYes
limitationsYes
measured_atYes
interpretationNo
schema_versionNo
analyzer_versionYes
analyzer_versionsYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior. The description adds context beyond those annotations by specifying the single-render scope and flagging optional pitch analysis as a capability; this is useful behavioral context for a safe analysis tool. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the main action and target, and no filler or repetition of the title. Every word adds signal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a complete input schema and an output schema present, the high-level description plus the schema is enough to invoke the tool correctly. It is slightly incomplete only in that it does not situate the tool among the many audio-analysis siblings, but that gap is split with usage guidelines.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, max_seconds, and include_pitch fully. The description's mention of 'optional pitch' maps to include_pitch, but it adds no semantics beyond what the parameter descriptions already provide, so the schema-coverage baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb ('Measure') and names the resource ('one render') plus the measured aspects (level, spectrum, dynamics, stereo, optional pitch), so an agent knows what the tool does. It does not explicitly contrast it with sibling analyzers like audio_compare_files or audio_analyze_masking, so sibling differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this tool over alternatives such as audio_compare_files or audio_analyze_masking, and no exclusion criteria are given. The only implied context is that it analyzes a single rendered file, which is not enough to route an agent confidently among the audio analysis siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_analyze_maskingA
Read-onlyIdempotent

Measure per-band spectral overlap and vocal-minus-instrument level margins.

ParametersJSON Schema
NameRequiredDescriptionDefault
vocal_pathYesAbsolute path to the vocal render.
max_secondsNoOptional shorter analysis bound in seconds; the default reads up to 600.
instrument_pathYesAbsolute path to the instrumental render of the same section, rendered sample-synchronously.

Output Schema

ParametersJSON Schema
NameRequiredDescription
vocalYes
balanceNo
maskingNo
parametersYes
limitationsYes
measured_atYes
instrumentalYes
context_readyYes
interpretationNo
schema_versionNo
vocal_activityNo
analyzer_versionYes
readiness_reasonsYes
masking_analyzer_versionYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description's 'Measure' verb is consistent with those. The description adds the computed metric details but no additional behavioral traits such as input constraints or processing side effects, so it stays at baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action verb, and every word contributes to defining the tool's function. No filler or redundant rephrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a 100%-covered input schema, clear safety annotations, an output schema present, and a succinct purpose statement, nothing needed to call the tool correctly is missing. Return-value details are handled by the output schema, so the description does not need to explain them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with all three parameters documented in the input schema. The description restates the vocal/instrument theme but adds no semantic detail beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Measure') and names concrete resources: per-band spectral overlap and vocal-minus-instrument level margins. This is distinct from generic siblings like audio_analyze_file and audio_compare_files because it identifies the exact masking-oriented metric computed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for diagnosing frequency masking between a vocal and an instrumental render, but it never states explicit when-to-use conditions or contrasts with related siblings such as mix_masking_recommendations or audio_compare_files. Context is clear, but exclusions and alternatives are left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_compare_filesB
Read-onlyIdempotent

Measure band deltas over the aligned, loudness-matched common overlap.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_secondsNoOptional shorter analysis bound in seconds; the default reads up to 600.
candidate_pathYesAbsolute path to the candidate render; reported as the target.
reference_pathYesAbsolute path to the reference render.

Output Schema

ParametersJSON Schema
NameRequiredDescription
targetYes
alignmentYes
referenceYes
confidenceYes
parametersYes
band_deltasYes
centroid_hzYes
limitationsYes
measured_atYes
interpretationNo
schema_versionNo
analyzer_versionYes
comparison_readyYes
analyzer_versionsYes
loudness_matchingYes
target_rate_conversionNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds context about alignment and loudness matching, which is useful, but does not disclose output format, interpretation of band deltas, or any edge cases. It meets the baseline but lacks deeper behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence that front-loads the core action and scope. There is no fluff or redundancy, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (comparison of two files with optional max_seconds) and the presence of an output schema, the description is minimal. It does not explain what 'band deltas' mean or how results should be interpreted, though the output schema may cover return structure. It is adequate but leaves room for more context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with clear descriptions for reference_path, candidate_path, and max_seconds. The tool description adds no additional parameter semantics beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (measure band deltas) and context (aligned, loudness-matched common overlap), which clearly implies a comparison of two audio files. It is not a tautology and distinguishes itself from generic analysis tools like audio_analyze_file or audio_analyze_masking, though it could be more explicit about the comparison nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as audio_analyze_masking or audio_analyze_file. The description only states what the tool does, leaving the agent to infer appropriate use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_estimate_tempo_and_keyA
Read-onlyIdempotent

Estimate periodic tempo and global major/minor key with ranked ambiguity.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to a decoded audio file.
max_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
keyYes
pathYes
tempoYes
sha256Yes
channelsYes
truncatedYes
analyzed_atYes
limitationsYes
sample_rate_hzYes
schema_versionNo
analyzer_versionNo
mutations_appliedNo
source_duration_secondsYes
analyzed_duration_secondsYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds behavioral context by noting 'ranked ambiguity,' which tells the agent the tool returns multiple ranked candidates rather than a single definitive answer. It does not disclose details like processing time, file size limits, or what happens with silent or percussive audio, but given the strong annotation coverage, the description adds adequate value without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 11 words that front-loads the core action ('Estimate periodic tempo and global major/minor key') and adds the key differentiator ('with ranked ambiguity') at the end. Every word earns its place; there is no filler, repetition of the title, or redundant schema information. This is an exemplary concise definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a simple 2-parameter schema, an output schema (which presumably documents the ranked results), and annotations covering safety and idempotency. The description's mention of 'ranked ambiguity' aligns with the output schema's likely structure. The only missing context is a brief note on what max_seconds does and how it affects results, but given the tool's simplicity and the presence of an output schema, the description is nearly complete for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: the 'path' parameter is documented in the schema ('Absolute path to a decoded audio file'), but 'max_seconds' has no description beyond its type, default, and range. The tool description does not explain what max_seconds controls (e.g., how much audio to analyze) or how it affects the estimate. With half the parameters undocumented in both schema and description, the description does not fully compensate, but the parameter names and schema constraints are reasonably self-explanatory, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Estimate') and resource ('periodic tempo and global major/minor key') with a distinctive qualifier ('with ranked ambiguity'). It clearly distinguishes this from sibling audio analysis tools like audio_analyze_file or audio_transcribe_melody, though it doesn't explicitly name them. The title 'Estimate tempo and musical key' reinforces the same purpose, so the description adds the ranked-ambiguity detail that separates it from a generic analysis tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context: it is for estimating tempo and key from a decoded audio file, and the 'ranked ambiguity' phrasing suggests it returns multiple candidate interpretations. However, it does not explicitly state when to prefer this over sibling tools like audio_analyze_file or audio_transcribe_melody, nor does it mention any exclusions or prerequisites beyond the path parameter. The context is clear enough for a straightforward analysis tool, but the lack of explicit alternatives or when-not-to-use guidance keeps it at a 3.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_find_recent_bouncesA
Read-only

List the newest audio files in FL Studio's Rendered, Audio, and Projects folders.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of files to return, newest first.

Output Schema

ParametersJSON Schema
NameRequiredDescription
filesYes
limitYes
rootsYes
limitationsYes
searched_atYes
scan_truncatedYes
schema_versionNo
matched_file_countYes
returned_file_countYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint and destructiveHint annotations already establish this as a safe read operation. The description adds valuable context beyond the annotations by specifying the exact folders searched and the recency ordering, which informs the agent of scope. It does not discuss return format, but the presence of an output schema covers that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence that front-loads the core purpose and scope. Every word earns its place with zero waste, listing all three target folders in one compact clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only listing tool with a full output schema and annotations covering the safety profile, the description is largely complete. The only minor gap is the lack of explicit statement about what happens when no files exist or how the three folders' results are combined, but these are edge cases an agent can handle from the schema and output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents the limit parameter including its default (20) and bounds (1-200). The description adds minimal semantic value by referencing 'newest first' ordering, but this is also in the schema. Baseline 3 is appropriate given the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List'), resource ('newest audio files'), and explicitly names the three folders searched (Rendered, Audio, Projects). This distinguishes it clearly from sibling audio tools like audio_compare_files, audio_analyze_file, and audio_estimate_tempo_and_key, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys when to use the tool (when you need recent audio exports/bounces) through its clear scope, but it does not state explicit exclusions or when not to use it relative to siblings. There are no competing 'list audio' siblings, so no routing guidance is strictly necessary, though a brief note on what it is not would improve it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audio_transcribe_melodyB
Read-onlyIdempotent

Extract a reviewable note sequence from monophonic audio; it does not mutate FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to one isolated pitched source.
fmax_hzNo
fmin_hzNo
tempo_bpmNo
max_secondsNo
quantize_grid_beatsNo
minimum_note_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
methodNo
fmax_hzYes
fmin_hzYes
sequenceYes
limitationsYes
source_pathYes
tempo_sourceYes
source_sha256Yes
schema_versionNo
tempo_bpm_usedYes
transcribed_atYes
frame_hop_secondsYes
mutations_appliedNo
voiced_frame_shareYes
median_pitch_confidenceYes
analyzed_duration_secondsYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description's 'does not mutate FL' mostly repeats structured data. It adds the monophonic-input constraint, but does not disclose timing, failure modes, or what happens with polyphonic input.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence: it states the core purpose first and then the safety constraint. No filler words; it is as concise as possible while being informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters and 14% schema description coverage, the tool needed a fuller context picture. The output schema helps, but the description gives no parameter guidance, no expected return style, and no differentiation from audio_analyze_file or compose_melody, leaving important gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14%, so the description must compensate for the 7 parameters. It mentions no parameter by name or provides any guidance on fmin_hz, fmax_hz, tempo_bpm, max_seconds, quantize_grid_beats, or minimum_note_seconds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the action ('Extract a reviewable note sequence') and resource ('monophonic audio'), and it adds a clear safety boundary ('does not mutate FL'). It does not explicitly distinguish itself from siblings like audio_analyze_file or compose_melody, so it is not a full 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'from monophonic audio' implies when the tool should be used, and 'does not mutate FL' implies it is an analysis tool rather than a DAW-mutating operation. However, it gives no explicit guidance on when not to use it or which sibling is a better alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

automation_record_valueB
Destructive

Dispatch one REC_MIDIController value while playback and recording are active.

ParametersJSON Schema
NameRequiredDescriptionDefault
propertyYesChannel targets support volume/pan; mixer also supports stereo separation.
target_kindYesAutomation target namespace.
allow_masterNoExplicitly permit mixer target 0.
target_indexYes
expected_beforeNo
value_normalizedYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
commandYes
event_idYes
propertyYes
verifiedNo
warningsNo
applied_atYes
target_kindYes
target_indexYes
project_savedNo
schema_versionNo
after_normalizedNo
controller_valueYes
before_normalizedNo
undo_point_createdNo
session_fingerprintYes
requested_normalizedYes
control_value_verifiedYes
capture_conditions_heldYes
expected_before_appliedYes
process_rec_event_resultNo
automation_event_recordedNo
song_position_after_ticksNo
song_position_before_ticksNo
session_precondition_appliedNo
automation_event_verificationNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the important safety signals: readOnlyHint=false, destructiveHint=true, and idempotentHint=false. The description adds a relevant precondition and one-shot scope with 'while playback and recording are active' and 'one'. However, it does not explain what is mutated, whether existing automation is overwritten, or what the expected_before parameter is protecting against. There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact sentence with an active verb, no filler, and the core action front-loaded. It is appropriately concise, though the brevity leaves important operational detail to the schema and annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, a destructive side-effect, and a concurrency-related field like expected_before, a single sentence is not enough context for an agent to invoke this reliably. The description does not mention critical invocation semantics such as mixer target 0 protection, the normalized value range, or what happens if playback/recording is not active. The output schema may help with return shape, but not with these invocation concerns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers only about 50% of the parameters with descriptions, and the description adds no parameter-level meaning. It does not explain target_index, expected_before, value_normalized, or the safety role of allow_master. Given the incomplete schema coverage, the tool description needed to compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete action — 'Dispatch one REC_MIDIController value' — and the context in which it operates ('while playback and recording are active'). The tool name and annotation title reinforce that this is an automation-recording operation. It loses the top score because 'REC_MIDIController' is unexplained jargon and the description does not explicitly say that this records automation as opposed to merely sending a MIDI event.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'while playback and recording are active' gives an implicit usage precondition, but the description never says when to prefer this tool over sibling parameter tools like fl_set_mixer_volume, fl_set_mixer_pan, or fl_set_channel_mix. It does not provide exclusions, alternatives, or failure conditions, so an agent gets only a loose sense of when this is the right call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_basslineA
Read-onlyIdempotent

Generate a bounded bass part from Roman harmony without changing FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoTonic note name.C
seedNoDeterministic variation seed.
styleNoroots
octaveNo
tempo_bpmNo
collectionNoBundled scale/mode/raga name, or custom.major
progressionYes
beats_per_chordNo
custom_intervalsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
seedNo
notesYes
warningsNo
generatorYes
tempo_bpmNo
note_countYes
duration_beatsYes
engine_versionNo
schema_versionNo
pitch_collectionNo
mutations_appliedNo
note_digest_sha256Yes
time_signature_numeratorNo
time_signature_denominatorNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds 'without changing FL' and 'bounded,' which are mild behavioral clarifications, but it does not disclose additional traits such as output placement, determinism details, or any constraints beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It communicates the action, the input source, the output type, and the side-effect guarantee efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and strong annotations, the description does not need to explain return values or safety. However, given the tool has nine parameters and only one is clarified, the description is not fully self-sufficient for an agent deciding how to configure optional parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, so the description needs to compensate. It adds meaning for one key parameter by clarifying that the progression is in Roman-numeral form, but the remaining parameters (root, style, seed, octave, tempo, collection, beats_per_chord, custom_intervals) receive no additional semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Generate a bounded bass part from Roman harmony.' It clearly identifies the output as a bassline, which distinguishes it from sibling tools like compose_melody and compose_drums. The phrase 'without changing FL' also clarifies the read-only nature of the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: when a bass part needs to be generated from a Roman-numeral harmonic progression. It does not explicitly name alternatives or exclusion conditions, so it stops short of full routing guidance, but the intended use case is evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_chord_progressionA
Read-onlyIdempotent

Generate voice-led triads/sevenths without changing FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNoTonic note name, for example C, F#, or Bb.C
octaveNo
voicingNoDeterministic voicing strategy.close
velocityNo
tempo_bpmNo
collectionNoBundled scale/mode/raga name, or custom.major
progressionYesRoman chords such as I, vi, IV, V, or V7.
beats_per_chordNo
custom_intervalsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
seedNo
notesYes
warningsNo
generatorYes
tempo_bpmNo
note_countYes
duration_beatsYes
engine_versionNo
schema_versionNo
pitch_collectionNo
mutations_appliedNo
note_digest_sha256Yes
time_signature_numeratorNo
time_signature_denominatorNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the non-mutating safety profile. The description adds the specific 'without changing FL' context and the voice-led behavior, but does not disclose additional behavioral details such as how generated chords are returned or how parameters interact. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with high signal-to-noise: the verb, target resource, key musical behavior, and non-mutating constraint are all front-loaded. There is no filler or redundant restating of the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 9-parameter complexity, annotation coverage, and presence of an output schema, the description is adequate but thin. It does not explain when to select this tool over compose siblings, nor does it clarify the meaning of the less obvious parameters. The output schema reduces the need to describe return values, but broader usage context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 44%, and the description adds no parameter-level meaning. Several parameters (octave, velocity, tempo_bpm, beats_per_chord, custom_intervals) lack schema descriptions, and the tool description does not compensate. While some names are self-explanatory, custom_intervals in particular remains underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Generate') with a clear resource ('voice-led triads/sevenths') and a distinguishing constraint ('without changing FL'). This clearly differentiates it from sibling composition tools like compose_melody, compose_bassline, and compose_drums, as well as FL-mutating tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without changing FL' gives useful context that this tool is non-mutating, but the description does not explicitly state when to prefer this tool over alternatives like compose_melody, compose_bassline, or piano_roll_write_notes. Usage context is implied but no explicit when/when-not guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_drumsA
Read-onlyIdempotent

Generate mapped kick/snare/hat patterns without changing FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
barsNo
seedNoDeterministic variation seed.
styleNohouse
swingNoDelay offbeat eighths in beats.
drum_mapNoSelected semantic drum map; omit for explicit General MIDI fallback.
tempo_bpmNo
beats_per_barNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
seedNo
notesYes
warningsNo
generatorYes
tempo_bpmNo
note_countYes
duration_beatsYes
engine_versionNo
schema_versionNo
pitch_collectionNo
mutations_appliedNo
note_digest_sha256Yes
time_signature_numeratorNo
time_signature_denominatorNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering safety and determinism. The description adds 'without changing FL' which reinforces the read-only behavior but contributes little beyond annotations. It does not describe how output is delivered (e.g., returned pattern object) or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core action and resource, then states a key behavioral constraint. No wasted words; brevity is appropriate for a simple generation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, a complex drum_map schema, and a rich environment of sibling tools, the one-sentence description is insufficient. It does not explain how the output is provided (despite an output schema), what role seed/style play, or how the drum_map fallback works—leaving agents to infer critical usage context from parameter names alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43%, and the tool description does not compensate. It mentions 'mapped' which hints at the drum_map parameter, but seed, style, swing, and other parameters are not elaborated. The description adds minimal meaning beyond what parameter names and types already convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates mapped drum patterns (kick/snare/hat) and explicitly notes it does not modify the FL project. This distinctively separates it from sibling composition tools like compose_melody or compose_bassline.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (when drum patterns are needed and no project changes are desired) but does not explicitly mention alternatives or exclusion criteria. No sibling differentiation is provided, so the usage guidance is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_melodyA
Read-onlyIdempotent

Generate a bounded scale-aware melody without changing FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
barsNo
rootNoTonic note name.C
seedNoDeterministic variation seed.
contourNobalanced
densityNo
tempo_bpmNo
collectionNoBundled scale/mode/raga name, or custom.major
register_lowNo
beats_per_barNo
register_highNo
custom_intervalsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
nameYes
seedNo
notesYes
warningsNo
generatorYes
tempo_bpmNo
note_countYes
duration_beatsYes
engine_versionNo
schema_versionNo
pitch_collectionNo
mutations_appliedNo
note_digest_sha256Yes
time_signature_numeratorNo
time_signature_denominatorNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false; the description reinforces this with 'without changing FL' and adds output constraints ('bounded scale-aware'). However, it doesn't disclose details like output format or performance implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence that front-loads the action and outcome, with no redundancy. Every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 11 parameters and only a one-sentence description, the definition is far from complete. It lacks usage guidance and parameter semantics, though output schema partially covers return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 27%, with just root, seed, and collection described. The description doesn't elaborate on bars, contour, density, register, or custom_intervals, so agents lack essential parameter meaning beyond names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Generate'), resource ('a bounded scale-aware melody'), and scope ('without changing FL'), distinguishing it from sibling tools like compose_chord_progression or fl_set_tempo.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use versus alternatives; however, the name and description make the use case for melody generation clear. Sibling tools (compose_bassline, compose_drums) imply this is the melody-specific tool, but no exclusions or contextual triggers are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

copilot_capture_readonly_inspectionB
Read-onlyIdempotent

Capture a compact project, mixer, and effect inspection report.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_usedNoApply the conservative used-track heuristic.
max_pluginsNoMaximum loaded plug-ins whose parameters are previewed.
parameter_limitNoMaximum parameter indices previewed per plug-in.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeNo
mixerYes
projectYes
warningsNo
observed_atYes
observation_idYes
schema_versionNo
observation_atomicNo
parameter_previewsYes
prohibited_operationsNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false), so the description's burden is lighter. It adds the 'compact' scope and confirms the report covers project/mixer/effects, which aligns with read-only behavior. No contradiction with annotations. It does not describe the report's structure, but the presence of an output schema lessens that need.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. The core action and scope ('compact project, mixer, and effect inspection report') appear immediately, and nothing else is stated that could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a low-complexity tool: zero required parameters, all optional with defaults, full schema coverage, and an output schema that documents the return structure. Given those structured assets, the one-line description is largely sufficient for a read-only snapshot tool. A minor gap is that it never hints the report is for quick orientation or that params tune its depth, but this is secondary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — all three parameters (only_used, max_plugins, parameter_limit) are documented with clear descriptions and bounds. The description adds no parameter-level information, but with full coverage the baseline of 3 is appropriate; the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Capture') and a clear resource ('a compact project, mixer, and effect inspection report'). It indicates the report spans three domains, which is more precise than a bare tautology. It does not explicitly differentiate from the many sibling inspection tools (fl_get_project_summary, plugins_scan_loaded_plugins, fl_inspect_mixer_track), but the 'compact, combined report' framing offers some distinctiveness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus the roughly 200 siblings. It does not name alternatives, state exclusion conditions, or explain where a compact overview fits in the workflow. An agent deciding between this and fl_get_project_summary, plugins_inspect_parameter_map, or mix_get_plan gets no routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_apply_verified_batchB
Destructive

Apply a closed-union batch after one session preflight, without replay.

ParametersJSON Schema
NameRequiredDescriptionDefault
operationsYesOrdered absolute writes. Operation IDs and written fields must be unique; every attempted item gets its own later-tick receipt.
stop_on_unverifiedNoSkip remaining items after the first unverified receipt.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultsYes
verifiedYes
warningsNo
completedYes
applied_atYes
project_savedNo
skipped_countYes
schema_versionNo
stopped_reasonNo
attempted_countYes
requested_countYes
rollback_attemptedNo
stop_on_unverifiedYes
session_fingerprintYes
automatic_replay_attemptedNo
one_session_preflight_completedNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as a non-read-only, destructive, non-idempotent call. The description adds 'without replay', but doesn't say what a closed-union application does to existing state, when writes are committed, or what happens on unverified items — especially important for a destructive tool whose schema includes stop_on_unverified and per-item receipts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler or repeated schema information. Every clause ('closed-union', 'after one session preflight', 'without replay') is attempting to convey a distinguishing property rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent batch writer with 23 operation discriminated union variants, this one-sentence description is thin. It doesn't connect the batch to the required preflight step by name, explain what 'verified' means, or indicate how failures are handled; the rich schema and output schema carry most of the burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the operations parameter's own description adds crucial semantics (ordered absolute writes, unique operation IDs and written fields, per-item receipts). The tool description itself contributes no parameter-level information, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete action ('Apply') and object ('a closed-union batch') and adds two meaningful qualifiers: it happens after a single session preflight and without replay. This is more than a restatement of the tool name, but the jargon terms 'closed-union' and 'preflight' are not decoded, and no sibling tool is named, so it stops short of full clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'After one session preflight' implies a precondition and suggests this is the bulk-write counterpart to a preflight read, so the usage context is weakly signaled. However, it gives no explicit when-not guidance and does not name the individual fl_set_* setters or apply_plan siblings it should be preferred over or combined with.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_find_empty_patternA
Read-onlyIdempotent

Find the first default-empty pattern without changing the current pattern.

ParametersJSON Schema
NameRequiredDescriptionDefault
start_pattern_numberNoFirst pattern number to inspect.

Output Schema

ParametersJSON Schema
NameRequiredDescription
warningsNo
observed_atYes
schema_versionNo
observation_atomicNo
project_dirty_flagNo
empty_pattern_numberNo
start_pattern_numberYes
scanned_pattern_countYes
current_pattern_unchangedYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering safety. The description adds a meaningful behavioral detail—'without changing the current pattern'—which is not captured by any annotation. This goes beyond the structured metadata and clarifies the tool's side-effect-free nature on selection.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the core action and a key constraint with zero filler. Every word earns its place, and it is immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one optional parameter) and read-only, with annotations covering safety. An output schema exists (not shown but indicated), so return format is covered. The only minor gap is the meaning of 'default-empty pattern,' which is domain-specific but likely understood in context. Overall, nothing essential is missing for a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%; the parameter 'start_pattern_number' is already documented as 'First pattern number to inspect.' The description does not add additional semantics beyond what the schema provides. Baseline of 3 is appropriate when the schema fully explains the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Find'), a clear resource ('first default-empty pattern'), and adds a behavioral qualifier ('without changing the current pattern') that distinguishes it from mutation tools like fl_select_pattern or fl_set_pattern_identity. It is immediately actionable and not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a usage scenario: you want to locate an empty pattern while preserving the current selection. However, it does not explicitly name alternatives or state when not to use it (e.g., vs. fl_list_patterns or fl_select_pattern). The context is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_get_capabilitiesA
Read-onlyIdempotent

Report direct, partial, unavailable, and unvalidated integration paths.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
connectionYes
capabilitiesYes
generated_atYes
schema_versionNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover read-only, idempotent, and non-destructive behavior, and the description's 'Report' aligns with these hints. However, the description adds no additional behavioral context such as the nature of the report (e.g., whether it includes version info, connection status, or error handling). With annotations present, the bar is lower, and the description is consistent but contributes minimal extra transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the action ('Report') and the content (integration paths). Every word adds value, with no fluff or repetition. It is an excellent example of conciseness without sacrificing essential meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter tool with an output schema, the description is mostly complete. It specifies what is reported (direct, partial, unavailable, unvalidated integration paths) but does not elaborate on how these are determined or what the output schema contains. Given the presence of an output schema, the description is adequate but could benefit from a note on typical use cases or interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is trivially 100%. The description tells what the tool reports, which is meaningful even without parameters. Since there are no params to document, the description fully satisfies the need for parameter semantics, warranting the baseline high score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Report' and the resource 'integration paths', specifying four categories (direct, partial, unavailable, unvalidated). It distinguishes from siblings by focusing on capability status rather than specific parameter scanning or project details. The annotation title 'Get verified FL capabilities' adds further clarity, though the description alone is somewhat cryptic without context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention prerequisites, typical scenarios, or exclusions. While the tool is likely for capability discovery, the lack of explicit usage context forces the agent to infer from the name and siblings, which is insufficient for making correct tool selection decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_get_plugin_preset_countA
Read-onlyIdempotent

Read FL's authoritative preset count for one loaded plug-in.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesExplicit mixer effect or global channel-generator target.

Output Schema

ParametersJSON Schema
NameRequiredDescription
pluginYes
warningsNo
observed_atYes
preset_countYes
schema_versionNo
project_dirty_flagNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds useful context beyond annotations by emphasizing 'authoritative' and requiring a 'loaded plug-in,' which clarifies data source and a precondition. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It communicates the action, resource, source authority, and scope efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with strong annotations, a full input schema, and an output schema, the description is nearly complete. It lacks only explicit usage-alternative routing and edge-case notes, which are not critical for correctness here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already fully documents the target discriminator and both target variants with descriptions. The tool description does not need to add parameter details, and it does not; it only restates the 'one loaded plug-in' scope already implied by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (read), resource (preset count), and scope (one loaded plug-in), which distinguishes it from preset list/select tools. However, it does not explicitly name or contrast sibling tools like plugins_list_presets or fl_select_plugin_preset, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. The description implies a read-only count use case but does not state exclusions, preconditions beyond 'loaded,' or scenarios where a sibling tool would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_get_project_historyA
Read-onlyIdempotent

Read undo/redo bounds, current history position, hint, and dirty state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
historyYes
warningsNo
observed_atYes
schema_versionNo
observation_atomicNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds value by specifying what is read (bounds, position, hint, dirty state), giving the agent a concrete expectation of the tool's behavior without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that names the operation and all key data fields with no filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read-only getter with an output schema and rich annotations, this description supplies all necessary context: what is being read and that it is non-destructive. No preconditions or side effects need explanation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters and 100% schema coverage, so there is nothing for the description to add about parameter meaning. This aligns with the baseline for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Read' with a clear resource ('project history') and enumerates the exact data returned: undo/redo bounds, current history position, hint, and dirty state. This sharply distinguishes it from mutation siblings like fl_undo and fl_redo.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The read-only phrasing establishes that this is for inspecting history state rather than changing it, and the sibling context makes fl_undo/fl_redo the obvious mutating alternatives. It does not explicitly state 'use before undo/redo' or name an alternative, but the usage context is clearly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_get_project_summaryA
Read-onlyIdempotent

Read project metadata, counts, dirty state, version, and transport state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
ppqNo
warningsNo
tempo_bpmNo
transportYes
connectionYes
dirty_flagNo
dirty_stateNo
observed_atYes
channel_countNo
pattern_countNo
project_genreNo
project_titleNo
project_authorNo
schema_versionNo
mixer_track_countNo
undo_history_countNo
playlist_track_countNo
undo_history_positionNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile: readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds the scope of data read (counts, dirty state, version, transport state), which is useful context but doesn't clarify what 'dirty state' entails or the nature of the return. No contradiction with annotations; with the annotations present, the bar is lower and this adds acceptable value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single 11-word sentence, front-loaded with the active verb and resource, then an efficient enumeration of content areas. Every word earns its place with zero fluff. This is model concise writing for a zero-param read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool with a present output schema (which covers return values) and strong annotations, the description is nearly complete. It enumerates the data domains covered. The only gap is not addressing the overlap with fl_get_transport_state and fl_get_project_history, which is a minor omission given the low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the rubric baseline is 4. There is nothing for the description to document — schema coverage is 100% vacuously. No parameter description is needed and none is provided, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and a clear resource scope ('project metadata, counts, dirty state, version, and transport state'). The name 'project summary' plus the enumerated content makes the purpose concrete. It does not explicitly differentiate from overlapping siblings like fl_get_transport_state, which also reads transport state, but the breadth of the summary is clear enough on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is provided. Critically, the description silently overlaps with fl_get_transport_state (transport state) and fl_get_project_history without noting when to prefer one over the other. Usage context is only implicit in the tool name itself; there are no explicit alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_get_selected_rangeC
Read-onlyIdempotent

Read raw endpoints and PPQ without claiming meter or rendering semantics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
warningsNo
end_ticksNo
observed_atYes
range_orderNo
start_ticksNo
raw_end_timeYes
timebase_ppqNo
raw_time_unitNo
duration_ticksNo
raw_start_timeYes
schema_versionNo
semantic_scopeNo
selection_stateNo
observation_atomicNo
project_dirty_flagNo
safe_for_renderingNo
selection_presenceNo
raw_end_display_hintNo
interpretation_statusNo
raw_start_display_hintNo
inactive_start_semanticsNo
repeated_read_consistentNo
render_endpoint_inclusivityNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds that it reads 'raw endpoints and PPQ' and does not 'claim meter or rendering semantics', which provides some context on the nature of data returned. However, it does not describe what the output looks like, any side effects (likely none given readOnly), or any limitations beyond the vague disclaimer. With annotations covering safety, the description offers limited additional behavioral insight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence with no filler. It front-loads the action ('Read') and a key detail (raw endpoints and PPQ). However, the phrasing is terse and uses jargon that may be unclear, but conciseness is positive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema (not shown in detail) that could clarify return values, but the description still lacks clarity on what 'endpoints' and 'PPQ' mean, and no guidance on when to use it. Given the complexities of FL Studio concepts, the description is too vague to fully guide an agent on what to expect or how to interpret the result. It doesn't mention any prerequisites or connection to other tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters)Skip; the description correctly adds no parameter semantics. Since there are no parameters to describe, the description does not need to compensate, but it also doesn't add value in this dimension. The baseline for 0 params is 4, but given that the description is vague about what the read returns, the parameter semantics dimension is moot; I'll rate 3 as neutral.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says 'Read raw endpoints and PPQ' which is specific about what is read, but it is not entirely clear what 'endpoints' and 'PPQ' refer to in the context of FL Studio. It mentions 'without claiming meter or rendering semantics', which hints at a distinction but the core purpose is somewhat ambiguous. It doesn't explicitly name the resource as a playlist selection, though the title is 'Get raw Playlist timeline selection', which is in annotations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not state when to use this tool or contrast it with alternatives. There is no mention of when to prefer this over other read tools like fl_get_project_summary or fl_get_transport_state. The vague phrasing 'without claiming meter or rendering semantics' implies it is for raw data, but it does not explicitly state the context in which this is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_get_step_sequenceB
Read-onlyIdempotent

Read an explicit current-pattern/channel grid and its conflict digest.

ParametersJSON Schema
NameRequiredDescriptionDefault
channel_indexYesGlobal channel index.
pattern_numberYesExplicit current pattern number.

Output Schema

ParametersJSON Schema
NameRequiredDescription
cellsYes
digestYes
warningsNo
step_countYes
index_scopeNo
observed_atYes
channel_indexYes
pattern_numberYes
schema_versionNo
grid_resolutionNo
digest_algorithmNo
observation_atomicNo
project_dirty_flagNo
current_pattern_numberYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety behavior. The description adds only the notion of an 'explicit' grid and conflict digest; it does not disclose additional behavioral traits such as failure behavior, cache effects, or output semantics beyond what structured annotations and output schema provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It states the action and resource immediately and avoids repeating any schema or annotation information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only, two-parameter tool with full schema coverage, idempotency annotations, and an output schema present, the description is sufficiently complete to invoke the tool correctly. The main gap is the absence of usage guidance relative to sibling read tools, but that is not critical for a focused getter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter has a basic description ('Global channel index', 'Explicit current pattern number'). The description's phrase 'explicit current-pattern/channel grid' loosely maps to the two required parameters but adds little semantic detail beyond the schema. Baseline 3 is appropriate because the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and a specific resource ('explicit current-pattern/channel grid') plus mentions 'conflict digest', clearly indicating a read operation on a particular structure. It is distinct from the setter sibling fl_set_step_sequenceanding, but the phrase 'current-pattern/channel grid' relies on domain jargon and could be slightly clearer for an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives like fl_list_patterns, piano_roll_read_notes, or fl_get_selected_range. The description implies a read use case but does not state exclusions, prerequisites, or when another read tool might be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_get_transport_stateA
Read-onlyIdempotent

Read playback, recording, loop mode, position, and song length.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
playingNo
loop_modeNo
recordingNo
tempo_bpmNo
song_length_msNo
precount_enabledNo
metronome_enabledNo
song_position_displayNo
song_position_normalizedNo
time_signature_numeratorNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety and side-effect behavior. The description adds value by specifying exactly which transport aspects are exposed (playback, recording, loop mode, position, song length), which is not captured in the annotations. It does not contradict any annotations and provides behavioral context beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the verb 'Read' and enumerates the fields concisely. Every word adds value, with no redundancy or filler. The structure is ideal for a simple getter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for a parameterless read with annotations covering safety and an output schema (not shown) that presumably defines the return structure. It lists all key transport aspects an agent would want to query. The only minor gap is not stating that it returns the current values, but that is implied by 'Read'. Overall, it is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing to document. The schema coverage is trivially 100% and the description needs no parameter details. Per the rubric, a baseline of 4 applies for zero parameters, and since the description is entirely sufficient for a parameterless read, a score of 5 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and the resource 'transport state', explicitly listing the attributes (playback, recording, loop mode, position, and song length). This distinguishes it from sibling setter tools like fl_set_playing or fl_set_song_position, which are inherently mutating. The tool's read-only nature is made obvious, so an agent can immediately understand its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage as the read counterpart to transport setters but does not explicitly state when to use it over other getters or mention that modifications require fl_set_* tools. While the verb 'Read' makes its purpose clear, there is no explicit guidance on when not to use it or how it contrasts with other read tools like fl_get_selected_range. The guidance is present but not detailed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_inspect_mixer_trackA
Read-onlyIdempotent

Read one track's state, effects, built-in EQ, and outgoing routes.

ParametersJSON Schema
NameRequiredDescriptionDefault
track_indexYesZero-based mixer index. Index 0 is always Master.

Output Schema

ParametersJSON Schema
NameRequiredDescription
trackYes
routesNo
warningsNo
builtin_eqNo
observed_atYes
schema_versionNo
observation_atomicNo
project_dirty_flagNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds the scope of what is read (state, effects, EQ, routes), which is modest context beyond the annotations but not rich behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. The core purpose ('Read one track's state') appears first, and every word carries meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only tool with well-documented schema, safety annotations, and an output schema, the description is nearly complete. The only gap is the absence of differentiation from list-type siblings, which is minor given the clear read scope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents track_index well, including the zero-based convention and that index 0 is Master. The description adds no parameter-specific meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Uses a specific verb ('Read') with a clear resource ('one track') and enumerates the exact content inspected (state, effects, built-in EQ, outgoing routes). This is easily distinguished from sibling fl_list_mixer_tracks, which lists tracks rather than inspecting one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as fl_list_mixer_tracks or fl_get_capabilities. The description does not mention any exclusion criteria, prerequisites, or selection conditions, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_list_channelsA
Read-onlyIdempotent

List globally addressed channels, mix state, routing, and generator identity.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
partialNo
channelsYes
warningsNo
observed_atYes
schema_versionNo
observation_atomicNo
project_dirty_flagNo
total_channel_countYes
scanned_channel_countYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering the safety profile. The description adds useful context about the scope ('globally addressed') and the content of results (mix state, routing, generator identity), which goes beyond the annotations. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the action and lists the key outputs. Every word contributes value, with no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, read-only, zero-parameter tool, the description is sufficiently complete. It states what is listed and what aspects are covered. An output schema exists, so the return format is handled elsewhere. No critical missing information for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description doesn't need to explain any. The schema is empty and fully covered. The baseline of 4 for zero-parameter tools applies, and the description provides no misleading info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action (list) and the resource (globally addressed channels), and specifies what information is returned (mix state, routing, generator identity). This distinguishes it from sibling tools like fl_list_mixer_tracks or fl_list_patterns, which target different resources. The verb and resource are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when channel information is needed, but it does not explicitly contrast with alternatives or state when not to use it. With many sibling list tools, an agent might need more guidance on selecting this over others. However, the purpose is clear enough for basic selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_list_mixer_tracksA
Read-onlyIdempotent

List mixer tracks, current levels, selection state, and loaded effects.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_usedNoApply a conservative used-track heuristic. False is the authoritative default.
max_tracksNoOptional early page limit for a large mixer scan.
include_peaksNoInclude instantaneous meter values; these are not audio analysis.

Output Schema

ParametersJSON Schema
NameRequiredDescription
tracksYes
partialNo
warningsNo
only_usedYes
observed_atYes
schema_versionNo
total_track_countYes
observation_atomicNo
project_dirty_flagNo
scanned_track_countYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds no extra behavioral traits (e.g., pagination, performance caveats, or that it queries live state). It is consistent with annotations, so no penalty, but also no added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that front-loads the verb and resource and then lists the useful data fields. There is zero filler and it immediately communicates the tool's core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read-only listing tool with full schema coverage, annotations, and an output schema, this description is basically enough. It could mention that it returns all tracks or hint at the optional filtering params, but the existing structured fields already cover those.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents each parameter (only_used, max_tracks, include_peaks) with descriptions. The tool description adds no extra meaning to the parameters, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('mixer tracks'), plus the kinds of data returned ('current levels, selection state, loaded effects'). This is clear and mostly distinguishes from the singular inspection sibling (fl_inspect_mixer_track), though it does not explicitly say 'all tracks' or name that alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a read-only overview of mixer tracks is needed, but it does not explicitly contrast with fl_inspect_mixer_track or other listing tools. There is no when-not-to-use guidance, though the tool name and sibling list make the common case understandable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_list_patternsA
Read-onlyIdempotent

List pattern identity, length, current state, and empty/default status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
patternsYes
warningsNo
observed_atYes
schema_versionNo
observation_atomicNo
project_dirty_flagNo
current_pattern_numberYes
maximum_pattern_numberYes
reported_pattern_countYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is well covered. The description aligns with these annotations and adds the specific returned fields, but it does not add operational caveats such as scope or ordering. With such strong annotation coverage, the description does not need to carry much additional behavioral burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. The action and the returned attribute categories are front-loaded, and every word contributes to the agent's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only listing tool with a declared output schema and fully covered safety annotations, the description is complete. It identifies the resource and the fields surfaced, and nothing else is needed for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

This tool has zero parameters and the schema coverage is complete, so there are no parameter semantics for the description to clarify. The description usefully indicates what the response concerns—pattern identity, length, state, and empty/default status—which is the only relevant semantic for a no-argument listing call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with an active verb ('List') and a specific resource ('patterns'), then enumerates the returned dimensions: identity, length, current state, and empty/default status. This is clear enough for an agent to tell it apart from mutating pattern tools like fl_set_pattern_identity and fl_set_pattern_length, though it does not explicitly name sibling alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus related siblings such as fl_find_empty_pattern, fl_select_pattern, or the fl_set_pattern_* tools. Usage context is only implicit in the verb 'List' rather than explicitly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_list_playlist_tracksA
Read-onlyIdempotent

List every one-based Playlist track and its controllable state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
tracksYes
warningsNo
observed_atYes
schema_versionNo
total_track_countYes
observation_atomicNo
project_dirty_flagNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful context beyond the annotations by specifying one-based indexing and that the result includes each track's controllable state. The annotations (readOnlyHint, openWorldHint, idempotentHint, destructiveHint) already cover safety, so the description's extra detail about indexing and content is sufficient and non-contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, efficient sentence with no wasted words. The action and resource are front-loaded, and the specific qualifiers (one-based, controllable state) are included without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only listing tool with an output schema available, the description fully covers what the tool does and what it returns. No additional context (e.g., pagination, limits) is necessary for an agent to call it correctly, and the annotations cover the safety profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is trivially 100%. The description does not need to explain parameter semantics, and the baseline for 0 parameters is 4. The description's mention of 'one-based' and 'controllable state' adds context about the output rather than parameters, which is not required here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('List'), the specific resource ('Playlist tracks'), and defines the scope ('every one-based' and 'its controllable state'). It distinguishes this from sibling list tools like fl_list_mixer_tracks and fl_list_patterns by naming the resource type explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (if you need playlist tracks, use this) but provides no explicit when-to-use guidance or alternatives. It does not mention cases where the agent should prefer fl_list_mixer_tracks or fl_list_patterns, nor any exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_redoA
Destructive

Move to the next absolute undo-history position and verify it.

ParametersJSON Schema
NameRequiredDescriptionDefault
expected_beforeNoOptional history position/count/dirty guard.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
directionYes
applied_atYes
project_savedNo
bridge_commandYes
schema_versionNo
requested_positionYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as destructive and non-read-only, so the description's main added value is the phrase 'and verify it', which hints that the tool actively confirms the resulting history state. It does not explain guard-failure behavior or what exactly gets verified, though the output schema may cover the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no wasted words, front-loading the core action ('Move...') before the secondary verification detail. It avoids repeating the title or restating schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-optional-parameter tool with a rich schema, output schema, and safety-related annotations, the description is nearly sufficient. The main gap is the missing explicit relationship to fl_undo and guidance about when the guards should be supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% coverage with detailed descriptions for expected_before and session_fingerprint, including their roles as guards. The description adds no parameter-specific semantics, which is acceptable because the schema already carries that burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete action ('Move to the next...') on a defined resource ('absolute undo-history position') and adds a verification step. It is less immediately transparent than 'redo the last undone change', but the title annotation and sibling fl_undo disambiguate the intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'next' qualifier implies this is the forward counterpart to fl_undo, so an agent can infer when to use it. However, the description never explicitly says 'use after fl_undo' or names alternatives, prerequisites, or conditions under which this tool should be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_route_channel_to_mixerB
Destructive

Set one global channel's absolute mixer destination and verify it.

ParametersJSON Schema
NameRequiredDescriptionDefault
channel_indexYesGlobal channel index.
expected_beforeNoOptional guarded channel fingerprint/destination.
mixer_destinationYesAbsolute mixer destination; -1 leaves it unassigned.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
index_scopeNo
channel_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
requested_mixer_destinationYes
session_precondition_appliedNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already carry destructiveHint=true and readOnlyHint=false, and the description does not contradict them. The added 'verify it' hints at a post-write check but does not explain what verification entails, what side effects occur, or what happens to the previous routing. A minimal but not misleading disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with zero filler. The core action and object are front-loaded, and every word contributes to understanding what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema provides the guardrails and parameter docs, and annotations cover safety, so the description is not dangerously incomplete. However, it does not explain operational semantics such as what verification means, when guards are needed, or what happens on a mixed route failure. For a 4-parameter destructive route tool, more operational context would be expected.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the parameters already document 'global channel index', 'absolute mixer destination', the expected_before guard, and session_fingerprint meaning. The description mostly paraphrases the schema rather than adding new parameter-level semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the action (Set), the resource (one global channel), and what is being changed (absolute mixer destination), plus a verification step. It does not explicitly distinguish from sibling tools like fl_set_channel_mix or fl_set_mixer_send, so it misses full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives, no mention of prerequisites, and no explanation of how the optional expected_before or session_fingerprint guards should guide invocation. The agent is left to infer usage context from the schema alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_select_channelB
Destructive

Select one global channel exclusively and verify the complete selection.

ParametersJSON Schema
NameRequiredDescriptionDefault
channel_indexYesGlobal channel index.
expected_beforeNoOptional exact selected-channel list guard.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
exclusiveNo
applied_atYes
index_scopeNo
channel_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo
after_selected_channel_indicesYes
before_selected_channel_indicesYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose destructiveHint=true and readOnlyHint=false, and the description adds 'exclusively' (implying replacing the current selection) and 'verify' (implying a confirmation step). It does not detail what state is destroyed or what the verification actually returns, but it does extend beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact sentence with no wasted words; it says what it does in the first clause and adds an extra trait in the second. It is efficient, though nothing is grouped or front-loaded in a way that improves digestibility because it is too short to need it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has only three parameters and no output schema, so the description carries more responsibilities. It does not objectively cover what 'verify' returns, what happens if the selection is not accepted, or how session_fingerprint and expected_before affect the result, although the schema already partially explains those parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already documented in the input schema. The description merely says 'global channel', echoing the schema's 'Global channel index,' and adds no extra meaning about expected_before or session_fingerprint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Select') and resource ('one global channel') plus a distinguishing quality ('exclusively' and 'verify the complete selection'). It is clear enough to separate it from sibling tools like fl_select_mixer_track or fl_select_pattern, though it never explicitly names an alternative to differentiate against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use the tool versus other selection tools, no conditions are mentioned, and no explicit 'when not to use' is given. The phrase 'exclusively' suggests the primary use, but an agent is left on its own to infer whether this tool is the right one rather than fl_select_mixer_track or a channel-toggle action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_select_mixer_trackB
Destructive

Make one mixer track active and verify the active-track getter.

ParametersJSON Schema
NameRequiredDescriptionDefault
track_indexYesZero-based mixer index to make active.
allow_masterNoRequired to select mixer track 0 (Master).
expected_beforeNoOptional expected active track index.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
track_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
targeted_masterNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
after_active_track_indexNo
before_active_track_indexNo
requested_active_track_indexYes
session_precondition_appliedNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a non-read-only, destructive, non-idempotent mutation, so the bar is lower. The description adds one useful behavioral detail: the operation verifies the active-track getter after selecting. It does not disclose side effects on the previously active track or failure behavior on guard mismatch, though some of that is encoded in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence and front-loads the core action. The trailing 'verify the active-track getter' phrase is slightly awkward but not verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a guarded write with 4 parameters, but the schema and annotations carry most of that context and an output schema exists. What is missing is any use-case framing or relationship to sibling mixer tools, so an agent gets no help deciding when this tool is the right one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 even without parameter details in the description. The description adds no parameter-level meaning; it merely names the action. The schema already explains track_index, allow_master, expected_before, and session_fingerprint adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('make one mixer track active') on a specific resource ('mixer track'), which separates it from mixer property setters and read-only inspectors. The appended 'verify the active-track getter' clause slightly muddies whether the tool is a selector or a verifier, but the primary verb/resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus fl_list_mixer_tracks, fl_inspect_mixer_track, or the many fl_set_mixer_* siblings. It neither states prerequisites nor excludes cases such as master selection or concurrency-guarded writes. The schema hints at these, but the description itself provides no usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_select_patternA
Destructive

Select one pattern and verify FL's current-pattern getter.

ParametersJSON Schema
NameRequiredDescriptionDefault
pattern_numberYesPattern number to make current.
expected_beforeNoOptional expected current pattern guard.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
project_savedNo
bridge_commandNo
schema_versionNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
after_pattern_numberNo
verification_summaryYes
before_pattern_numberNo
expected_before_appliedNo
requested_pattern_numberYes
session_precondition_appliedNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this as non-read-only and destructive; the description adds the post-selection verification of the getter, which is a useful behavioral detail beyond annotations. However, it does not clarify what state is affected or why the operation is flagged destructive, so the extra disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler, and both clauses carry information: the selection action and the verification behavior. It is appropriately sized for this tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema plus fully described parameters make basic invocation clear, and the description states the action and the verify behavior. For a mutation tool with a concurrency guard, it would benefit from a sentence on when to use it and what state it changes, but the structured fields fill most gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: pattern_number, expected_before, and session_fingerprint all have descriptive text, including the concurrency-guard semantics. The description adds no parameter-level information, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair ('Select one pattern') and adds a second, distinctive behavior (verifying FL's current-pattern getter), so an agent can tell this is the pattern-selection tool rather than fl_list_patterns or fl_set_pattern_identity. It does not explicitly name alternatives, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The action itself implies when to call it (when the agent wants to make a pattern current), but there are no explicit when-to-use/when-not-to-use conditions or alternative tool mentions. With dozens of pattern and mixer-selection siblings, the routing guidance is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_select_plugin_presetC
Destructive

Navigate to an exact preset and require later-idle-tick identity readback.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesExplicit mixer effect or global channel-generator target.
preset_nameNoExact reported preset name.
preset_indexNoExact reported preset index.
expected_currentNoOptional stale-read guard for the current preset.
settle_tick_limitNoLater idle ticks allowed for plug-in settling.
target_fingerprintNoObserved target-identity guard.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.
max_navigation_stepsNoBound on next/previous navigation.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
pluginYes
targetYes
outcomeYes
verifiedYes
warningsNo
applied_atYes
project_savedNo
bridge_commandNo
schema_versionNo
navigation_stepsYes
settle_tick_limitYes
target_fingerprintNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
max_navigation_stepsYes
navigation_directionYes
verification_summaryYes
requested_preset_nameNo
requested_preset_indexNo
expected_before_appliedNo
session_precondition_appliedNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds a behavioral detail beyond annotations by mentioning a post-navigation identity readback, but the phrase 'later-idle-tick identity readback' is unexplained and cryptic. There is no contradiction with the annotations, but the added context is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler words; it front-loads the core action. However, the second clause is dense and lacks explanation, so it earns a 4 rather than a 5 for ideal clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a complex schema (8 parameters, nested target discriminator, multiple concurrency and fingerprint guards), but the description only gives a two-part imperative. An agent cannot infer how to correctly set expected_current, target_fingerprint, session_fingerprint, settle_tick_limit, or max_navigation_steps from this description. The presence of an output schema does not compensate for missing operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents parameters like preset_name, preset_index, expected_current, and the fingerprint guards. The description adds no additional parameter meaning, and the baseline of 3 applies because the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a verb ('Navigate') and a resource ('exact preset'), which is more than a tautology)Skip the vague phrase 'require later-idle-tick identity readback' obscures the core operation. It does not explicitly state that this selects a plugin preset, and the readback jargon makes the purpose partly opaque.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool instead of siblings like plugins_list_presets, plugins_get_current_preset, or fl_get_plugin_preset_count. The description implies a selection operation but offers no context on prerequisites, alternatives, or conditions that would route an agent to this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_channel_identityA
Destructive

Set a channel's name and/or color with per-field readback proof.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoAbsolute channel name.
colorNoAbsolute FL 0x--BBGGRR color word. FL owns the high byte, so write verification compares the low 24 color bits.
channel_indexYesGlobal channel index.
expected_beforeNoOptional guarded channel fingerprint/name/color.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
index_scopeNo
channel_indexYes
name_verifiedNo
project_savedNo
bridge_commandNo
color_verifiedNo
requested_nameNo
schema_versionNo
requested_colorNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it destructive and non-read-only, lowering the bar. The description adds value by promising 'per-field readback proof', a behavioral guarantee of verification after the write. No contradiction observed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, only 15 words, front-loaded with the key action, resource, and distinct behavioral trait. Zero filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema documents all five parameters and output schema exists for return values. The description covers the primary purpose and the notable readback behavior, which is the main aspect not already exposed by the schema. It does not mention the optional guard parameters (expected_before/session_fingerprint), but these are thoroughly documented in the schema itself.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the description does not elaborate on parameter-specific semantics. Since the schema already carries descriptions for name, color, and expected_before, the description adds little to parameter understanding but does not need to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('set') with the exact resource ('channel's name and/or color') and a distinctive behavior ('per-field readback proof'). It clearly distinguishes this from sibling identity tools like fl_set_pattern_identity or fl_set_playlist_track_identity across different object types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool does not explicitly say when to use it instead of alternatives. The channel context is implied by the name and description, and 'per-field readback proof' could signal a use case for verifying writes, but no explicit when-to-use or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_channel_mixA
Destructive

Set channel volume, pan, and/or mute with per-field readback proof.

ParametersJSON Schema
NameRequiredDescriptionDefault
panNoAbsolute channel pan.
mutedNoAbsolute channel mute state.
channel_indexYesGlobal channel index.
expected_beforeNoOptional guarded channel fingerprint and/or mix fields.
volume_normalizedNoAbsolute channel volume.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
index_scopeNo
pan_verifiedNo
channel_indexYes
mute_verifiedNo
project_savedNo
requested_panNo
bridge_commandNo
schema_versionNo
requested_mutedNo
volume_verifiedNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
requested_volume_normalizedNo
session_precondition_appliedNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a non-read-only, non-idempotent, potentially destructive mutation. The description adds a useful, non-annotated behavior: 'per-field readback proof' shows the tool returns verification of what was applied. It does not disclose the concurrency guards or failure refusal semantics, but those are described in the schema properties such as session_fingerprint and expected_before, so the added value is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-formed sentence with the key verb and resource front-loaded. Every word contributes to understanding (set, channel, volume, pan, mute, readback proof), with no redundant explanation or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given rich annotations, a complete input schema, and an output schema, the description provides enough orientation for a correct call. The 'and/or' and 'per-field readback proof' cover the most useful non-schemaable context; missing guidance about concurrency guards is already place in schema parameter descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents each parameter's meaning. The description's 'and/or' suggests fields can be updated individually or together, but it adds little beyond the schema's explicit null defaults and descriptions. A baseline of 3 is appropriate when the schema carries the weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it sets channel volume, pan, and/or mute, and adds a distinctive 'readback proof' behavior. This distinguishes it from mixer-track siblings like fl_set_mixer_volume, fl_set_mixer_pan, and fl_set_mixer_mute by making the channel scope explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance about when to use this tool versus alternatives. It does not mention that it targets channel rack channels rather than mixer tracks, nor does it note the expected_before/session_fingerprint guards, leaving any usage decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_channel_pitchA
Destructive

Set normalized channel pitch and report normalized/semitone readback.

ParametersJSON Schema
NameRequiredDescriptionDefault
channel_indexYesGlobal channel index.
expected_beforeNoOptional fingerprint and/or pitch guard.
pitch_normalizedYesAbsolute FL channel pitch from -1.0 to 1.0.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
index_scopeNo
channel_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
requested_pitch_normalizedYes
session_precondition_appliedNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already communicate that this is mutating (readOnlyHint=false) and destructive (destructiveHint=true). The description adds the concrete mutation and the normalized/semitone readback, but it does not disclose any extra side effects, irreversibility, or how the optional guards affect execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with no filler: it states the action and the readback in a front-loaded way. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter setter with full schema coverage and an output schema, the core operation is clear. However, it lacks context about when to set pitch in a workflow, how it interacts with open-world state, or what the optional guards protect against beyond what the schema already says.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters including channel_index, pitch_normalized range, expected_before, and session_fingerprint. The description adds no parameter-level meaning beyond naming the pitch resource, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Set'), a precise resource ('normalized channel pitch'), and the additional readback behavior. This is enough to distinguish it from sibling setters like fl_set_channel_mix, fl_set_channel_solo, or fl_set_plugin_param.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus the many other fl_set_* tools, nor are preconditions stated such as the project being loaded or the channel index being valid. The only usage signal is the tool name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_channel_soloA
Destructive

Set a global channel solo state and verify it on a later FL tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
soloedYesAbsolute wanted solo state; never a toggle.
channel_indexYesGlobal channel index.
expected_beforeNoOptional fingerprint and/or solo-state guard.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
index_scopeNo
channel_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
requested_soloedYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, setting the baseline. The description adds valuable behavioral context: 'global' scope and the verification step on a later FL tick. This goes beyond the annotations by indicating the operation has an asynchronous verification aspect and affects the whole project. However, it does not mention potential side effects like exclusive solo behavior (unsoloing other channels) or the concurrency guard mechanics, which are partially in the schema. The description meaningfully supplements the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is concise and front-loaded with the core action. It avoids fluff and explicitly mentions verification. While it could be slightly more informative, it is appropriately sized for a straightforward set operation with verification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of a comprehensive input schema (100% parameter descriptions) and an output schema (declared present), the description need not repeat structured details. It covers the primary action, scope, and verification. The description is adequate for an agent to understand the tool's purpose and invoke it correctly, especially with the guards described in the schema. Missing details like why the guards exist are supplied by schema descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter having clear descriptions (e.g., 'Absolute wanted solo state; never a toggle', 'Global channel index', and explanatory notes for guards). The tool description itself adds no additional parameter-level detail beyond the schema, so the baseline of 3 is appropriate since the schema already carries the semantic weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the specific verb 'Set' with the resource 'global channel solo state' and adds the verification action 'verify it on a later FL tick'. It clearly distinguishes from sibling tools like fl_set_mixer_solo by specifying 'channel' rather than mixer, and the global scope is explicit. The action is unambiguous and easy to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for setting channel solo state but does not explicitly state when to use it versus alternatives like fl_set_mixer_solo or when not to use it. There is no mention of prerequisites or conditions under which this tool should be chosen over others. While the name and description make the primary use clear, the lack of explicit guidance on alternatives and exclusions leaves some gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_loop_modeA
Destructive

Set Pattern or Song loop mode without exposing FL's toggle-only API.

ParametersJSON Schema
NameRequiredDescriptionDefault
loop_modeYesAbsolute loop mode: 'pattern' or 'song'.
expected_beforeNoOptional expected current loop mode.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
project_savedNo
bridge_commandNo
schema_versionNo
after_loop_modeNo
before_loop_modeNo
undo_point_createdNo
verification_basisNo
requested_loop_modeYes
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the agent knows this is a mutating operation. The description adds value by explaining that FL's native API is toggle-only and this tool wraps it to provide an absolute set, which is a meaningful behavioral disclosure. The schema's session_fingerprint parameter also documents the concurrency guard, so the description doesn't need to repeat that. It doesn't mention side effects on playback, but the annotation covers the destructive nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and resource, then adds the crucial differentiator about the toggle-only API. Zero waste, no repetition of schema details, and the most decision-relevant information comes first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple setter with a full output schema and 100% parameter coverage, the description is nearly complete. The only gap is that it doesn't explicitly state the effect of expected_before (e.g., that the write will fail if the current state doesn't match), but the schema's parameter description covers that. The description plus schema plus annotations together give an agent everything needed to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters: loop_mode, expected_before, and session_fingerprint. The description adds the high-level intent ('Set Pattern or Song loop mode') but doesn't add meaning beyond the schema's own descriptions, which are already clear. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Set'), a specific resource ('Pattern or Song loop mode'), and a key differentiator: it avoids exposing FL's toggle-only API. This clearly distinguishes it from other transport/set tools like fl_set_playing or fl_set_write_mode, and tells the agent exactly what state it will produce.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: whenever the agent needs to set an absolute loop mode rather than toggle it. It doesn't explicitly name alternatives or exclusions, but the phrase 'without exposing FL's toggle-only API' signals that this is the safe, absolute setter among transport controls. Sibling names like fl_set_playing and fl_set_write_mode provide context, but the description itself could have been more explicit about when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_metronomeB
Destructive

Set the metronome absolutely and prove the later UI state.

ParametersJSON Schema
NameRequiredDescriptionDefault
enabledYesAbsolute metronome state.
expected_beforeNoOptional expected metronome state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
after_enabledNo
project_savedNo
before_enabledNo
bridge_commandNo
schema_versionNo
requested_enabledYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a write operation (readOnlyHint=false, destructiveHint=true), and the description adds a weak behavioral claim: 'prove the later UI state,' implying verification after the write. However, it does not explain what 'prove' means, what happens on mismatch, or what side effects occur beyond changing the metronome state. This goes somewhat beyond annotations but remains vague.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core verb, which is good, but the second clause is awkward and ambiguous ('prove the later UI state'). It is concise in length yet sacrifices clarity through imprecise wording, so it does not reach the level of a well-structured definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

As a simple setter with a complete input schema and an output schema, much of the needed information is present. However, the verification behavior implied by 'prove' is unexplained, and there is no guidance about when to supply expected_before or session_fingerprint, which are key to safe use. The tool is callable but not fully self-sufficient for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains 'Absolute metronome state', the optional expected state, and the session_fingerprint concurrency guard. The description adds no real parameter-level detail beyond the word 'absolutely', which loosely aligns with the enabled parameter. This meets the schema-heavy baseline without adding meaningful value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('set') on a clear resource ('metronome'), and 'absolutely' indicates an absolute enable/disable rather than a toggle. It does not name or distinguish itself from sibling read/write transport tools, so cross-tool differentiation is weak. The phrase 'prove the later UI state' is confusing, but the core purpose remains identifiable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as fl_get_transport_state or other fl_set_* tools. It does not mention when the concurrency guard or expected state should be supplied, nor when a simple write is insufficient. Usage context must be inferred entirely from the schema and sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_armA
Destructive

Set recording arm with one bounded toggle and later-tick readback.

ParametersJSON Schema
NameRequiredDescriptionDefault
armedYesThe absolute wanted recording-arm state.
track_indexYesZero-based mixer index. Index 0 is Master.
allow_masterNoRequired to target mixer track 0.
expected_beforeNoOptional expected current arm state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
after_armedNo
track_indexYes
before_armedNo
project_savedNo
bridge_commandNo
schema_versionNo
requested_armedYes
targeted_masterNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a mutating, destructive, non-idempotent operation. The description adds useful behavioral context beyond those annotations: the write is a single bounded toggle and the resulting state should be read back on a later tick. This gives the agent an important timing expectation that structured annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, front-loaded with the core action, and contains no filler. However, the phrasing 'bounded toggle' and 'later-tick readback' is packed with domain jargon and is slightly cryptic, so the conciseness comes at some cost to immediate readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that all parameters are fully described in the schema and an output schema exists, the description does not need to document return values or parameter syntax. It supplies the key operational caveat—readback on a later tick—and the binary nature of the write. The main missing piece is broader usage context, but that is already penalized under usage guidelines.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already well documented in the input schema. The description adds essentially no new parameter-level meaning; 'bounded toggle' loosely maps to the boolean armed property, but it does not enrich or explain any of the five parameters beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a clear operation and resource: 'Set recording arm'. The phrase 'one bounded toggle' signals a binary state change rather than a continuous adjustment, which helps distinguish it from sibling tools like fl_set_mixer_volume or fl_set_mixer_pan. It does not explicitly say 'mixer track', though the tool name and schema make that clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance about when to use this tool versus alternatives, nor any exclusionary context. The 'later-tick readback' hint implies the caller should verify the result after a processing tick, but it does not say which read tool to use or when the concurrency guards like expected_before or session_fingerprint should be supplied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_colorA
Destructive

Set a mixer color, accepting FL-owned differences in the high byte.

ParametersJSON Schema
NameRequiredDescriptionDefault
colorYesFL color word as unsigned 0xAABBGGRR integer.
track_indexYesZero-based mixer index. Index 0 is Master.
allow_masterNoRequired to target mixer track 0.
expected_beforeNoOptional expected current FL color word.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
after_colorNo
track_indexYes
before_colorNo
project_savedNo
bridge_commandNo
schema_versionNo
requested_colorYes
targeted_masterNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already communicate that this is a destructive, non-idempotent write operation, so the description's job is lighter. It adds a useful behavioral detail: the tool tolerates FL-owned differences in the high byte of the color word. This is non-obvious and valuable for correct use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes meaning, and the key behavior is stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema, output schema, and annotations, the short description is mostly sufficient. The main missing piece is explicit when/why to choose this tool over sibling mixer setters, but the schema compensates for most operational details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all five parameters with 100% coverage, so the baseline is 3. The description adds extra semantic meaning beyond the schema by clarifying that the high byte of the color value may differ from what FL reports, helping agents understand how strict the color value needs to be.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Set a mixer color') and adds a meaningful nuance about accepting FL-owned high-byte differences. It clearly distinguishes this tool from sibling mixer setters (volume, pan, mute, etc.) without needing to inspect the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is given about when to use this tool versus alternatives, and no prerequisites or exclusions are mentioned. The intended use is only implied by the name and description, which is insufficient given the large family of mixer-related setter tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_muteA
Destructive

Set one track's mute state and report the readback FL gave on a later tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
mutedYesThe wanted state: true mutes, false unmutes. This is stated, never toggled.
track_indexYesZero-based mixer index. Index 0 is Master and is refused unless allow_master is true.
allow_masterNoDeliberately target the master bus at index 0.
expected_beforeNoOptional expected current mute state; refuse if it changed.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
after_mutedNo
track_indexYes
before_mutedNo
project_savedNo
bridge_commandNo
schema_versionNo
requested_mutedYes
targeted_masterNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and idempotentHint=false, so the safety profile is known. The description adds valuable behavioral context by noting that the tool reports the readback FL gave on a later tick, implying a delayed/asynchronous read rather than an immediate confirmation. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with the core action front-loaded and the important readback behavior appended in the second clause. Every word earns its place, and there is no redundant restatement of the tool name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema, annotations, output schema, and the added readback behavior, the description provides most of the context an agent needs to invoke this tool correctly. It does not explicitly describe refusal cases such as master-track protection, expected_before mismatches, or session_fingerprint guards, but those are fully covered by the parameter schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains muted state semantics, zero-based track indexing, master refusal, expected_before, and session_fingerprint. The description adds no parameter-specific meaning beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific action: set one track's mute state, and adds the readback-reporting behavior. This clearly distinguishes it from sibling mixer setters like fl_set_mixer_volume, fl_set_mixer_solo, or fl_set_mixer_pan because the resource and mutation target are explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is used to set a track's mute state rather than toggle it, which gives the agent a usable context. However, it does not explicitly name alternatives or state when to prefer this over other mixer muting/state tools, so the usage guidance is mostly inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_nameA
Destructive

Rename one mixer track.

An empty name is not a blank label: FL puts the track's default back
("Insert 8"), and the result says so via `restored_default`.
ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesNew track name. Pass "" to restore FL's default.
track_indexYesZero-based mixer index. Index 0 is Master.
allow_masterNoRequired to target mixer track 0.
expected_beforeNoOptional expected current name; refuse if it changed.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
after_nameNo
applied_atYes
before_nameNo
track_indexYes
project_savedNo
bridge_commandNo
requested_nameYes
schema_versionNo
targeted_masterNo
restored_defaultNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare this as non-read-only and destructive, so the safety profile is known. The description adds genuinely useful behavioral context beyond that: an empty string is not a blank label, FL replaces it with the default ('Insert 8'), and the result reports this via `restored_default`. This is above the minimum and adds value without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler. The core action is first, followed only by the single caveat worth calling out. No information is repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-target rename tool, the description is complete: it states the action, the target scope, the empty-name edge case, and the result signal. Since the schema and annotations already cover parameters, destructive behavior, and output shape, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameters are already fully documented and the baseline is 3. The description adds meaning by explaining the non-obvious consequence of the empty-string case and naming the `restored_default` result marker, going slightly beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the specific verb+object 'Rename one mixer track', which states exactly what the tool does and distinguishes it from the many sibling fl_set_mixer_* tools. The scope is unambiguous even before reading the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use or when-not-to-use guidance. It never mentions alternatives or conditions such as when to use fl_set_mixer_color, fl_set_mixer_send, or another mixer setter. The only extra sentence covers empty-name behavior, not usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_panA
Destructive

Set one mixer pan and report the readback FL gave on a later tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
panYesPan position, -1.0 hard left through 0.0 centre to 1.0 hard right.
track_indexYesZero-based mixer index. Index 0 is Master and is refused unless allow_master is true.
allow_masterNoDeliberately target the master bus at index 0.
expected_beforeNoOptional expected current pan; refuse if it changed.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
after_panNo
applied_atYes
before_panNo
track_indexYes
project_savedNo
requested_panYes
bridge_commandNo
schema_versionNo
targeted_masterNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavioral context beyond the annotations: it reports the readback FL gave on a later tick, signaling that the write is verified asynchronously. The destructiveHint=true and idempotentHint=false annotations already convey the mutation profile, so the description does not need to repeat that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence conveys the action, scope, and observable result with no filler. The key purpose is front-loaded and every word contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema descriptions, output schema, and annotations, the description covers the essential behavior for a simple single-track setter. It could mention that concurrency guard parameters can cause refusal, but those are already explained in the parameter descriptions, so the overall package is complete enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the inputs are already fully documented. The description itself adds no additional parameter-level meaning beyond the schema, which is exactly the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Set'), a specific resource ('one mixer pan'), and adds the distinctive readback behavior. It clearly differentiates this from sibling mixer setters such as fl_set_mixer_volume or fl_set_mixer_mute.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied ('set one mixer pan'), but there is no explicit guidance about when not to use it or which sibling tool to prefer. With many similar mixer-setter tools available, some alternative routing would improve this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_sendA
Destructive

Route one mixer track to another, or stop routing it there.

A stated state, never a toggle. Sending *to* Master needs no flag; only
sending *from* Master does. Set the amount with `fl_set_mixer_send_level`
afterwards -- this call only decides whether the route exists.
ParametersJSON Schema
NameRequiredDescriptionDefault
enabledYesTrue to create the send, False to tear it down.
track_indexYesZero-based index of the sending track.
allow_masterNoRequired only to send FROM track 0.
expected_beforeNoOptional expected current route state; refuse if it changed.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.
destination_track_indexYesZero-based index of the receiving track.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
track_indexYes
after_enabledNo
project_savedNo
before_enabledNo
bridge_commandNo
schema_versionNo
targeted_masterNo
level_normalizedNo
requested_enabledYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
destination_track_indexYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag mutation and destructiveness; the description adds genuinely non-obvious runtime behavior: it is 'a stated state, never a toggle,' has an asymmetric Master-flag requirement, and scopes effects to route existence only. It doesn't add caveats about expected_before/session_fingerprint, but the schema covers those, so the added behavior context goes beyond annotations without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences front-load the core action and then add only the behavioral details an agent needs; every sentence earns its place, and the call-to-sibling for level setting is appended without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Combined with a rich schema that documents all six parameters and a clear output schema, the description supplies the missing operating semantics: route existence vs. level, absolute state, and the Master exception. No decision-relevant gap remains for selecting or calling the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is a 3, but the description adds semantic value for the boolean parameter: 'A stated state, never a toggle' tells the agent that enabled is absolute, not a flip. It also clarifies allow_master's asymmetry ('Sending to Master needs no flag; only sending from Master does'), complementing the schema's terse 'Required only to send FROM track 0.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names both actions ('Route one mixer track to another, or stop routing it there') with a specific verb and resource, making the tool's function immediately clear. It also distinguishes this from level-setting via a terse but precise scoping statement, so an agent won't confuse it with sibling mixer tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly directs the agent to use fl_set_mixer_send_level for the amount and clarifies that this call 'only decides whether the route exists,' which defines the boundary between the two tools. The master-flag guidance ('only sending from Master does') is a concrete when-to-include-flag rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_send_levelA
Destructive

Set how much of one track reaches another. 0.8 is unity, as on the fader.

The send must already exist; create it with `fl_set_mixer_send` first. FL
raises rather than reporting a level for a route that is not active, so
this is refused outright rather than written and reported unverified.
ParametersJSON Schema
NameRequiredDescriptionDefault
track_indexYesZero-based index of the sending track.
allow_masterNoRequired only to send FROM track 0.
expected_beforeNoOptional expected current send amount; refuse if it changed.
level_normalizedYesSend amount, 0..1. 0.8 is unity.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.
destination_track_indexYesZero-based index of the receiving track.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
send_activeNo
track_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
targeted_masterNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
after_level_normalizedNo
before_level_normalizedNo
destination_track_indexYes
expected_before_appliedNo
requested_level_normalizedYes
session_precondition_appliedNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as a destructive, non-idempotent write, and the description adds a non-obvious behavioral detail: FL raises rather than reporting a level for an inactive route, so the call is refused outright. This goes beyond the structured hints by explaining the validation/refusal model. It does not enumerate all side effects, but the destructive annotation carries that safety signal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short, focused paragraphs: the first defines the operation, the second gives the prerequisite and failure behavior. Every sentence earns its place, and there is no redundant restatement of annotations or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter write tool with a full input schema and an output schema, the description covers the core concept, the scale convention, the key prerequisite, and the failure mode. Nothing an agent needs in order to select and invoke the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline applies. The description's '0.8 is unity, as on the fader' adds a small calibration cue, but the schema already documents the same unity fact and describes each parameter's role. The description does not materially compensate for anything missing because nothing is missing from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action and object: 'Set how much of one track reaches another,' which clearly identifies this as a mixer send-level write. It also adds the key domain fact that 0.8 is unity, which distinguishes it from volume, pan, and send-creation operations. This is clearly distinct from sibling tools like fl_set_mixer_send and fl_set_mixer_volume.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the prerequisite: the send must already exist and must be created with `fl_set_mixer_send` first. It also explains when the operation should not proceed — for an inactive route — and that the tool refuses rather than writing an unverified value. This gives clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_soloB
Destructive

Set one mixer track's solo state and verify it on a later FL tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
soloedYesThe absolute wanted solo state; never a toggle.
track_indexYesZero-based mixer index. Index 0 is Master.
allow_masterNoRequired to target mixer track 0.
expected_beforeNoOptional expected current solo state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
track_indexYes
after_soloedNo
before_soloedNo
project_savedNo
bridge_commandNo
schema_versionNo
targeted_masterNo
requested_soloedYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint, non-idempotency, and open-world behavior. The description adds the useful detail that the sole state is verified on a later FL tick, which is not captured in the annotations, but it does not disclose failure semantics or what happens after a bridge reload or project load.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact, front-loaded sentence captures the core action and the unique verification behavior with no irrelevant content. It earns its place and avoids redundancy with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is straightforwardly clear for a simple read-write mixer tool, but this is a mutation with non-idempotent behavior, destructive annotation, and a concurrency guard. The description doesn't explain the significance of the session_fingerprint guard or verification failure, and the annotations/schema carry most of the burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters carry the full semantic load: absolute soloed state, zero-based track index, allow_master requirement, expected_before, and session_fingerprint concurrency guard. The description adds little beyond the schema, but the 'verify' phrase loosely hints at the expected_before mechanism.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('set'), resource ('mixer track'), and state ('solo state'), and clarifies it operates on a single mixer track. It clearly differentiates from the mute/volume/pan setters, but doesn't explicitly contrast with fl_set_channel_solo, so sibling differentiation is only implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this versus fl_set_channel_solo or fl_set_mixer_mute, no mention of prerequisites such as allow_master for track 0, and no exclusions. Usage is only implied by the tool name and schema, not actively explained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_stereo_separationA
Destructive

Set stereo separation and report FL's later-tick readback.

ParametersJSON Schema
NameRequiredDescriptionDefault
track_indexYesZero-based mixer index. Index 0 is Master.
allow_masterNoRequired to target mixer track 0.
expected_beforeNoOptional expected current stereo-separation value.
stereo_separationYesFL stereo-separation value from -1.0 to 1.0.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
track_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
targeted_masterNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
after_stereo_separationNo
expected_before_appliedNo
before_stereo_separationNo
requested_stereo_separationYes
session_precondition_appliedNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=false, destructiveHint=true, idempotentHint=false). The description adds a genuinely useful behavioral trait beyond annotations: the tool reports FL's 'later-tick readback,' implying the write may not be reflected immediately and the returned value can diverge from the requested value — a non-obvious timing nuance an agent should know before trusting the result. It doesn't contradict the write/destructive annotations, so no contradiction is flagged.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with two clauses, zero filler, and the essential action front-loaded before the readback caveat. Both clauses carry distinct useful meaning: the operation and the verification behavior. There is no wasted wording or redundancy with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a mutation with a concurrency guard and a reported readback, and it has an output schema to describe return values. The description mentions the readback but leaves ambiguity about why a later-tick read could differ and what the implications are. Given the openWorldHint and destructive annotations, slightly more behavioral context (e.g., what failure or divergence implies) would improve completeness, but the schema and annotations cover the mechanics, making this adequately minimal rather than deficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all five parameters, including the concurrency-guard semantics of session_fingerprint and the master-track gating of allow_master. Per the baseline rule, a 3 applies when the schema does the heavy lifting. The description adds no parameter-specific meaning, which is acceptable here since nothing is missing from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Set') with a specific resource ('stereo separation') and adds a non-obvious twist: 'report FL's later-tick readback,' signaling that the actual applied value may be verified asynchronously. The name alone distinguishes it from the many sibling mixer setters (volume, pan, mute, solo), and while the description doesn't explicitly contrast with siblings, it clearly states the operation. The readback clause adds definitional value beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance exists. The description offers no context on when to prefer this tool over fl_set_mixer_pan, fl_set_mixer_volume, or fl_set_mixer_send. There are no exclusions, alternatives, or conditions stated in the description. Some usage nuance is embedded in schema fields (allow_master is 'Required to target mixer track 0'), but the description itself leaves the agent to infer selection criteria, which is inadequate given the dense sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_volumeA
Destructive

Set one mixer fader and report the readback FL gave on a later tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
track_indexYesZero-based mixer index. Index 0 is Master and is refused unless allow_master is true.
allow_masterNoDeliberately target the master bus at index 0.
expected_beforeNoOptional expected current fader position; refuse if it changed.
volume_normalizedYesFader position, 0.0 silent to 1.0 maximum. 0.8 is FL Studio's 0 dB default.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
track_indexYes
project_savedNo
bridge_commandNo
schema_versionNo
after_volume_dbNo
targeted_masterNo
before_volume_dbNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
after_volume_normalizedNo
expected_before_appliedNo
before_volume_normalizedNo
requested_volume_normalizedYes
session_precondition_appliedNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a key behavioral trait beyond annotations: it reports the readback on a later tick, implying an asynchronous or delayed verification step. Annotations already mark it as destructive and non-read-only, so the description adds useful context about what happens after the write. The phrase 'later tick' is somewhat vague but still provides real behavioral insight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact sentence with no filler. The primary action ('Set one mixer fader') is front-loaded, and the additional reporting behavior is included without wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema, annotations, and output schema, the description completes the picture by adding the verification/readback behavior. It does not repeat schema details, which is appropriate. The main gap is the lack of explicit alternative routing to fl_set_mixer_volume_db, but that is a usage-guidelines issue rather than a completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter well-documented (e.g., 'Fader position, 0.0 silent to 1.0 maximum', master refusal, concurrency guard). The description itself adds no parameter meaning beyond what the schema already provides, so it stays at the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Set one mixer fader', and adds the distinctive behavior of reporting the readback. It clearly identifies what the tool does, but it does not explicitly distinguish itself from the closely related sibling fl_set_mixer_volume_db, which also sets mixer volume. The normalized vs. dB distinction is left to the parameter schema rather than the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description mentions 'one mixer fader' but does not explain when to use the normalized version over fl_set_mixer_volume_db, or any other mixer setter. No exclusions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_mixer_volume_dbC
Destructive

Search FL's fader curve and prove the requested dB value on a later tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
volume_dbYesTarget fader readback in dB.
track_indexYesZero-based mixer index; Master requires allow_master.
allow_masterNoDeliberately target Master at mixer index 0.
tolerance_dbNoMaximum accepted dB readback error.
expected_beforeNoOptional expected normalized and/or dB state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
track_indexYes
tolerance_dbYes
project_savedNo
bridge_commandNo
schema_versionNo
after_volume_dbNo
targeted_masterNo
before_volume_dbNo
search_iterationsYes
undo_point_createdNo
verification_basisNo
requested_volume_dbYes
session_fingerprintNo
verification_summaryYes
after_volume_normalizedNo
expected_before_appliedNo
before_volume_normalizedNo
session_precondition_appliedNo

TDQS

C2.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds a useful behavioral trait beyond the annotations: 'prove the requested dB value on a later tick' signals asynchronous or deferred verification, and 'Search FL's fader curve' hints at the dB-to-fader mapping logic. It does not contradict the destructiveHint or readOnlyHint annotations. The wording is terse but does disclose non-obvious timing behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, but it is not front-loaded with the actual purpose; the internal mechanism ('Search FL's fader curve') is placed before the operation. It sacrifices clarity for brevity and uses jargon ('prove', 'later tick') without definition. This is under-specification rather than effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even though the schema and output schema cover parameters and return values, the description does not state the core action or side effects of this destructive, non-idempotent write tool. It also fails to distinguish this tool from sibling mixer controls. For a 6-parameter mutating tool with async-style behavior, the description is materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents all six parameters with descriptions, so the baseline is 3 per high schema coverage. The description adds no parameter-specific meaning beyond the concept of a dB value. It neither clarifies the tolerance_db, expected_before, or session_fingerprint semantics, nor is it required to do so given the schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description leads with 'Search FL's fader curve,' which obscures the set operation implied by the tool name, and never explicitly says that the mixer fader is set to a target dB. 'Prove the requested dB value on a later tick' is an oblique way to describe the effect. An agent would struggle to confidently infer the primary action from this text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives such as fl_set_mixer_volume for normalized values. The description does not state prerequisites, exclusions, or a preferred context. Usage must be inferred entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_pattern_identityA
Destructive

Set pattern name and/or color with per-field later-tick proof.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoAbsolute pattern name.
colorNoAbsolute unsigned FL color word.
pattern_numberYesPattern number to edit.
expected_beforeNoOptional expected pattern name/color.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
name_verifiedNo
project_savedNo
bridge_commandNo
color_verifiedNo
pattern_numberYes
requested_nameNo
schema_versionNo
requested_colorNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=trueched and readOnlyHint=false, so the write nature is known. The description adds the cryptic 'per-field later-tick proof' concept but does not explain it; the session_fingerprint parameter description later clarifies the concurrency guard, partially compensating.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one clear, action-first sentence but sacrifices some clarity with the jargon-heavy 'per-field later-tick proof' phrase. Still, it is compact and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, rich parameter descriptions, and annotations covering safety. The only ambiguity — the meaning of 'later-tick proof' — is resolved in the session_fingerprint parameter description, making the tool effectively callable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter comprehensively explained in the schema. The description adds no extra parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Set'), the target ('pattern name and/or color'), and a distinctive behavior ('per-field later-tick proof'). This clearly differentiates it from sibling tools like fl_set_pattern_length or fl_select_pattern.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives that also modify pattern or channel identity. There is no mention of prerequisites, exclusions, or preferred scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_pattern_lengthB
Destructive

Set pattern length using Image-Line's API 39+ getter/setter pair.

ParametersJSON Schema
NameRequiredDescriptionDefault
length_beatsYesAbsolute pattern length in beats.
pattern_numberYesPattern number to edit.
expected_beforeNoOptional expected current length.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
project_savedNo
bridge_commandNo
pattern_numberYes
schema_versionNo
after_length_beatsNo
undo_point_createdNo
verification_basisNo
before_length_beatsNo
session_fingerprintNo
verification_summaryYes
requested_length_beatsYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as destructive and non-read-only, lowering the bar. The description adds the detail that it uses a getter/setter pair with API 39+, which provides some technical context, but it does not describe what happens to the previous pattern length or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the core operation. The clause 'using Image-Line's API 39+ getter/setter pair' is somewhat technical but still earns its place by providing useful implementation context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich schema, output schema, and annotations, the core operation is adequately specified. However, the description lacks explicit usage context and does not clarify the destructive effect beyond the annotation, leaving some gaps for an agent deciding whether this is the right tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are already individually documented in the schema. The description adds no additional parameter-level meaning, which matches the baseline of 3 for complete schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Set') and resource ('pattern length'), making the tool's purpose immediately understandable. However, it does not explicitly distinguish itself from sibling tools like fl_set_pattern_identity or fl_select_pattern, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives, and it names no sibling tools or exclusion criteria. Usage is only implied by the tool name and title.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_playingA
Destructive

Set playback to an absolute state and verify it on a later FL idle tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
playingYesAbsolute playing state; never a toggle.
expected_beforeNoOptional expected current playing state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
after_playingNo
project_savedNo
before_playingNo
bridge_commandNo
schema_versionNo
requested_playingYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate a destructive write (destructiveHint true, readOnlyHint false). The description adds unique behavioral context: that the change is verified on a later FL idle tick, implying asynchronous verification. This goes beyond what annotations convey, adding timing and verification semantics without contradicting the annotation safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly written sentence that conveys the core action, the absolute (non-toggle) nature, and the verification step. There is no fluff or repeated information from the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, full parameter descriptions in the schema, and annotations for safety/destructiveness, the description covers the essential purpose and timing. It does not elaborate on failure handling or the expected_before parameter, but those are documented in the schema. The description is sufficient for an agent to correctly invoke the tool in most scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so all three parameters (playing, expected_before, session_fingerprint) are already explained. The description does not add additional parameter-level details beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to set playback to an absolute state (not a toggle) and verify it later. This specific verb-resource pair distinguishes it from siblings like fl_stop (stop) and fl_set_song_position (position), and the 'absolute state' phrase clarifies it is for play/pause rather than a toggle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly name alternatives or state when to use this tool vs. others (e.g., fl_get_transport_state for reading). The phrase 'absolute state' implies it is for setting play/pause rather than toggling, but there is no direct guidance on when to prefer it over other playback-related tools. Usage is implied but not stated explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_playlist_track_identityB
Destructive

Set Playlist name and/or color with per-field later-tick proof.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoAbsolute Playlist track name.
colorNoAbsolute unsigned FL color word.
track_indexYesOne-based Playlist track index.
expected_beforeNoOptional expected track name/color.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
track_indexYes
name_verifiedNo
project_savedNo
bridge_commandNo
color_verifiedNo
requested_nameNo
schema_versionNo
requested_colorNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a destructive, non-idempotent write operation. The description adds a hint of concurrency behavior via 'per-field later-tick proof,' but it does not explain what the proof entails, when it can fail, or what side effects occur beyond changing name/color.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler and earns its place as a concise statement of the operation. However, the compressed jargon 'per-field later-tick proof' is not self-explanatory, so the conciseness comes at some cost to immediate clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the strong parameter schema, output schema, and annotations, the one-line description is enough for basic invocation. The main missing guidance—when to prefer this over state/identity siblings is already penalized under usage_guidelines, and the schema fills the remaining details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented. The phrase 'name and/or color' aligns with `name` and `color`, and 'per-field later-tick proof' loosely maps to `expected_before`, but the description does not add concrete semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Set') and resource ('Playlist name and/or color'), making it clear that the tool mutates Playlist track identity. It is differentiated enough from sibling identity tools by name and scope, though 'per-field later-tick proof' is cryptic and slightly obscures the core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool instead of related siblings like fl_set_playlist_track_state, fl_set_pattern_identity, or fl_set_channel_identity. Usage is only implied by the tool name and the 'Set' verb; no exclusions, prerequisites, or alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_playlist_track_stateB
Destructive

Set Playlist states; toggle-only selection is dispatched at most once.

ParametersJSON Schema
NameRequiredDescriptionDefault
mutedNoAbsolute mute state.
soloedNoAbsolute solo state.
selectedNoAbsolute selection state.
track_indexYesOne-based Playlist track index.
expected_beforeNoOptional expected mute/solo/selection state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
track_indexYes
mute_verifiedNo
project_savedNo
solo_verifiedNo
bridge_commandNo
schema_versionNo
requested_mutedNo
requested_soloedNo
requested_selectedNo
selection_verifiedNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the write/non-idempotent nature is carried by structured data. The description adds one behavioral nuance: 'toggle-only selection is dispatched at most once,' which is helpful but not elaborated (e.g., does this mean repeated selection toggles are coalesced?). It doesn't mention the concurrency guard or what happens to unspecified states, but annotations and schema partially cover the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, purpose first, with zero filler. The behavioral nuance is attached in a clear second clause. It is front-loaded and economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has six parameters with optional fields and a concurrency guard, but the description does not explain what null means for mute/solo/selected (presumably 'leave unchanged'), nor does it clarify the interaction between expected_before and session_fingerprint. Given the high schema coverage and annotations, an agent can probably call it correctly, but the absence of null semantics and a more explicit statement about the effect of a write is a measurable completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter has a descriptive line (e.g., 'Absolute mute state,' 'Optional bridge/project-session fingerprint'). The description adds value on the 'selected' behavior with the at most once note, but does not illuminate expected_before or session_fingerprint beyond schema text. Since the schema carries the parameter burden, the description plates at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Set Playlist states') and the schema confirms it covers mute, solo, and selected states on a playlist track. Though 'Playlist states' is a bit terse, the tool name and sibling context (e.g., fl_set_playlist_track_identity) distinguish it from identity-setting tools. It doesn't enumerate the exact states in the description, but the resource and verb are sufficiently clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no explicit guidance on when to prefer this tool over alternatives. It doesn't mention that it operates on Playlist tracks (vs mixer tracks), when to use the expected_before/session_fingerprint concurrency guards, or how it differs from identity manipulation. The single sentence is behavioral, not navigational. For a mutation tool with many siblings, this is a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_plugin_paramA
Destructive

Set one plug-in parameter; verified from FL's display string changing.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoExplicit mixer_effect or global channel_generator target. Mutually exclusive with legacy track_index/slot_index.
slot_indexNoLegacy zero-based effect slot 0 through 9. Supply it with track_index, or use target, never both.
track_indexNoLegacy zero-based mixer index. Supply it with slot_index, or use target, never both.
allow_masterNoDeliberately target the master bus at index 0.
expected_beforeNoOptional expected normalized value and/or exact display text; refuse if any supplied field changed.
parameter_indexYesParameter index as reported by plugins_inspect_parameter_map. Nothing here knows what the control does.
normalized_valueYesParameter value, normalized 0.0 to 1.0.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation already marks the tool as destructive (destructiveHint: true), so the description is not required to restate that. The description adds value by revealing that the operation is verified against FL's display string changing, implying an explicit post-write check. This is a behavioral trait not present in the annotations, so it earns credit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core action ('Set one plug-in parameter') and immediately adds a meaningful verification detail. No wasted words, and it is as concise as possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal, but the schema is rich and handles parameter semantics, target selection, and safety guards. Given the tool's complexity (8 parameters, two target addressing modes, concurrency guard), one might expect more usage context, but the schema compensates. Since an output schema exists, return value details are not required. Overall adequate but with some gaps in practical guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes all 8 parameters with full coverage, including detailed comments for target, legacy fields, expected_before, session_fingerprint, and normalized_value. The description adds no parameter-level semantics, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Set one plug-in parameter') on a resource, and adds a notable verification detail ('verified from FL's display string changing'). However, it does not explicitly differentiate itself from siblings like fl_set_plugin_param_display or fl_set_plugin_param_option, so it narrowly misses a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as fl_set_plugin_param_display or fl_set_plugin_param_option. The description lacks any condition or context that would help an agent decide between the set of similar parameter-setting tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_plugin_param_displayA
Destructive

Set a plug-in parameter using the units it displays, not a 0..1 guess.

Prefer this over `fl_set_plugin_param` for anything with real units.
Normalized 0..1 has no published mapping to ms, dB or Hz, and the curve
differs per control; this searches the control until its own readback
reports the number you asked for, so no curve is ever assumed.

Address the parameter by name where it has one ("Attack"), or by
what it displays where it does not ("Auto mode"). Run
`plugins_scan_parameters` first to see both.

Controls whose display is pure text -- "Chromatic", "Low Male" -- are
enumerations with no number to search on and are refused. Use
`fl_set_plugin_param_option`, which sets them by their option text and
also reports every option the control accepts.
ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoExplicit mixer_effect or global channel_generator target. Mutually exclusive with legacy track_index/slot_index.
parameterYesParameter index, or text matched against parameter names AND display strings (many real third-party controls have no name).
toleranceNoHow close counts. Defaults to 2% of the target, floor 0.01.
slot_indexNoLegacy zero-based effect slot 0 through 9. Supply it with track_index, or use target, never both.
target_unitNoOptional Hz, kHz, ms, seconds, dB, percent or ratio. Converts each display readback across unit prefixes. Omit for legacy first-number matching.
track_indexNoLegacy zero-based mixer index. Supply it with slot_index, or use target, never both.
allow_masterNoRequired to target mixer track 0.
target_valueYesThe number the plug-in displays: 20 for '20 ms', -18 for '-18.0 dB'. With target_unit='Hz', use 4000 for '4.0kHz'.
expected_beforeNoOptional expected normalized value and/or exact display text; refuse if any supplied field changed.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already communicate a write/destructive operation. The description adds valuable behavioral nuance: it searches the control via readback until the number matches, assumes no mapping curve, and refuses pure-text enumerations. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: purpose, preference, addressing strategy, and exclusions each get one clear block. No sentence is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-parameter write tool, the combination of description, annotations, and full schema coverage covers purpose, safety, parameter identification, and the sibling to use for text-only controls. An output schema exists, so the absence of return-format detail is not a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the baseline is met by the schema. The description adds meaningful extra guidance around how to address parameters—by name when available and by display string otherwise—and clarifies target values are display-unit numbers, not normalized guesses. It leaves tolerance and guard parameters to the schema, which is reasonable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Set a plug-in parameter using the units it displays, not a 0..1 guess.' It immediately differentiates this from fl_set_plugin_param and fl_set_plugin_param_option, so an agent can tell what this tool uniquely does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the agent to prefer this over fl_set_plugin_param for anything with real units, and directs pure-text enumerations to fl_set_plugin_param_option. It also instructs running plugins_scan_parameters first, giving a clear before-use workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_plugin_param_optionA
Destructive

Set a parameter that shows words rather than numbers: Key, Scale, Input Type.

Use this where `fl_set_plugin_param_display` refuses. That tool searches on
a number, and an enumeration has none.

**This moves the control while it looks.** FL cannot report a control's
options, so the only way to find them is to walk the parameter across its
range and read what it displays. The requested label must exactly match an
option, ignoring case. If it does not exist, the original value is restored
before the error, and the error lists every option that was found.

The result carries `options` -- the whole enumeration, in order -- so one
call is also how you discover what a control accepts.
ParametersJSON Schema
NameRequiredDescriptionDefault
optionYesThe exact option text to land on, e.g. 'A', 'Major', 'Low Male'.
targetNoExplicit mixer_effect or global channel_generator target. Mutually exclusive with legacy track_index/slot_index.
parameterYesParameter index, or text matched against names and displays.
slot_indexNoLegacy zero-based effect slot 0 through 9. Supply it with track_index, or use target, never both.
sweep_stepsNoSweep resolution. Raise it only if an option is being missed.
track_indexNoLegacy zero-based mixer index. Supply it with slot_index, or use target, never both.
allow_masterNoRequired to target mixer track 0.
expected_beforeNoOptional expected normalized value and/or exact display text; refuse if any supplied field changed.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a major side effect beyond the annotations: the tool moves the control while it looks, because FL cannot report options. It also explains exact-match case-insensitive matching, restoration of the original value on failure, and that errors list all found options.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loads the purpose, and uses bold to call out the key side effect. Every sentence contributes either selection guidance, behavioral disclosure, or discovery semantics with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool, the description covers the mechanism, failure mode, restoration guarantee, and options-discovery behavior. The rich schema and output schema handle the remaining parameter and return details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema carries most parameter meaning; the baseline is 3. The description adds value by clarifying that the option must exactly match a label ignoring case and that one call doubles as a discovery mechanism returning the whole options list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise action and resource: setting a parameter that displays words rather than numbers, with concrete examples (Key, Scale, Input Type). It also distinguishes the tool from fl_set_plugin_param_display by naming the sibling and stating the exact condition that selects this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs to use this tool where fl_set_plugin_param_display refuses, and explains why: that tool searches on a number while an enumeration has none. This gives an agent both a positive and a negative selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_precountA
Destructive

Set recording precount absolutely and verify it on a later FL tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
enabledYesAbsolute countdown-before-recording state.
expected_beforeNoOptional expected precount state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
after_enabledNo
project_savedNo
before_enabledNo
bridge_commandNo
schema_versionNo
requested_enabledYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint and non-idempotence, so the description adds value by revealing that the tool will later 'verify' the set value on a future FL tick, implying a follow-up check rather than a fire-and-forget write. This is useful behavioral insight beyond the schema. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single two-clause sentence perfectly communicates the core action and the verification behavior without any filler or unnecessary prose. It is front-loaded with 'Set...' and the rest is unnecessary but valuable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main action and verification but does not mention important preconditions or effects that are not already in the schema, such as failure behavior when session_fingerprint mismatches or what happens to existing precount. The output schema likely covers return values, so this is adequate but not exhaustive for a mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% coverage, each parameter having a clear description (e.g., 'Absolute countdown-before-recording state'). The description's 'absolutely' is redundant with the schema's 'Absolute'. No additional semantic meaning is given, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Set recording precount absolutely and verify it on a later FL transition' clearly identifies the specific verb (set), resource (recording precount), and the absolute (non-toggling) nature, plus a verification side-effect. This distinguishes it from siblings like fl_set_recording, which controls recording state, without confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus others, such as fl_set_recording or fl_set_metronome. It doesn't mention prerequisites, exclusions, or alternatives, leaving the agent to infer the appropriate call from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_recordingA
Destructive

Set recording absolutely; FL's toggle is dispatched at most once.

ParametersJSON Schema
NameRequiredDescriptionDefault
recordingYesAbsolute transport recording-arm state.
expected_beforeNoOptional expected recording state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
project_savedNo
bridge_commandNo
schema_versionNo
after_recordingNo
before_recordingNo
undo_point_createdNo
verification_basisNo
requested_recordingYes
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint and readOnlyHint, but the description adds behavioral nuance: 'FL's toggle is dispatched at most once' reveals that the tool may not act if the state is already as requested, and it's not a simple toggle. This goes beyond the annotations, adding useful context about how the operation is executed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly written sentence that conveys the core purpose and a key behavioral detail without fluff. It is front-loaded and every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write operation with concurrency guards, the description covers the main behavioral trait (absolute set) and the schema covers parameters and output. It does not explicitly mention failure modes (e.g., stale session fingerprint), but these are described in the schema. The presence of an output schema further reduces the need to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides full descriptions for all parameters (100% coverage), including the concurrency guard and expected state. The description adds no additional parameter information, so baseline 3 is appropriate since the schema carries the semantic burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool sets the recording state and uses 'absolutely' to indicate a deterministic set rather than a toggle. The resource is 'recording', which is distinct from sibling tools like fl_set_playing or fl_stop, so purpose is clear. However, it doesn't explicitly contrast with alternatives, so not a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It doesn't mention when to prefer this over a simple toggle or any exclusions. The description implies an absolute set but doesn't state conditions for use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_song_positionA
Destructive

Set a stopped transport's absolute normalized playhead position.

ParametersJSON Schema
NameRequiredDescriptionDefault
toleranceNoMaximum normalized readback error.
expected_beforeNoOptional expected current normalized position.
position_normalizedYesAbsolute normalized playhead position.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
toleranceYes
applied_atYes
project_savedNo
bridge_commandNo
schema_versionNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo
after_song_position_normalizedNo
before_song_position_normalizedNo
requested_song_position_normalizedYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds the critical precondition that the transport must be stopped and clarifies that the position is absolute and normalized. This goes beyond the annotations by specifying a behavioral constraint (stopped) and the exact positioning semantics, which is valuable for correct invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, front-loaded with the key action and constraints. There is no wasted verbiage, and every word contributes to the meaning. It is appropriately concise for the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (4 parameters, one with a nested object), the schema covers parameter semantics thoroughly, and annotations cover the destructive/non-idempotent nature. The description adds the key 'stopped' precondition and 'absolute normalized' semantics. However, it does not mention the session_fingerprint concurrency guard (though that is in the schema) or any failure modes, and it lacks explicit usage guidance. It is adequate but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter already having a detailed description (e.g., position_normalized as 'Absolute normalized playhead position', session_fingerprint with its concurrency guard semantics). The tool description adds no additional parameter information, so it does not improve on the schema. Per the rubric, the baseline of 3 applies when schema covers all parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Set'), a clear resource ('transport's playhead position'), and two qualifying attributes ('stopped' and 'absolute normalized'). This clearly distinguishes the tool from siblings like fl_set_playing or fl_stop, and conveys exactly what action it performs without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description only implies usage by mentioning 'stopped transport', but provides no explicit guidance on when to use this tool versus alternatives (e.g., fl_set_playing, fl_stop, fl_set_loop_mode). It does not state when not to use it or mention any exclusions, leaving the agent to infer from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_step_sequenceA
Destructive

Set absolute current-pattern cells only if the observed grid digest still matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
updatesYesUnique absolute cell states.
channel_indexYesGlobal channel index.
pattern_numberYesExplicit current pattern number.
expected_digestYesRequired digest from fl_get_step_sequence.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
index_scopeNo
channel_indexYes
project_savedNo
bridge_commandNo
cells_verifiedYes
pattern_numberYes
schema_versionNo
expected_digestYes
requested_updatesYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the description doesn't need to restate that. But the description adds the critical 'only if the observed grid digest still matches' - a concurrency guard that is a behavioral trait beyond annotations. However, it doesn't detail what happens on mismatch (does it refuse? partially apply? return error?) or other side effects, so a 3 is appropriate given annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One concise sentence that packs the key operational constraint (digest matching) up front. Every word is meaningful and the length is appropriate for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists (though not shown), the description doesn't need to explain return values. The description covers the purpose, the concurrency guard, and the context (absolute current-pattern cells) adequately for an agent to select and invoke the tool correctly, especially with annotations and schema providing additional details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (all parameters have descriptions), so the baseline is 3. The description doesn't add detailed semantics beyond what's already in the schema, but it does clarify the relationship between expected_digest and the concurrency guard. That is a slight addition, but not enough to push above 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific action (set absolute current-pattern cells) and a critical condition (only if observed grid digest matches), clearly distinguishing from siblings like fl_set_step_sequence's read counterpart and other pattern setters. The phrase 'set absolute current-pattern cells' precisely scopes the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly says when to use it: 'set absolute current-pattern cells' and when not: 'only if observed grid digest still matches'. It also implies the need to first read the digest via fl_get_step_sequence, and warns against use when digest is stale. This provides explicit context and exclusion (stale digest).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_tempoA
Destructive

Set project tempo while stopped and verify BPM on a later FL idle tick.

ParametersJSON Schema
NameRequiredDescriptionDefault
tempo_bpmYesAbsolute project tempo in BPM.
expected_beforeNoOptional expected current tempo in BPM.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
project_savedNo
bridge_commandNo
schema_versionNo
after_tempo_bpmNo
before_tempo_bpmNo
undo_point_createdNo
verification_basisNo
requested_tempo_bpmYes
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the write nature is known. The description adds the behavioral detail that the tool verifies the BPM on a later FL idle tick, which is beyond the annotations and useful for understanding postconditions. It does not contradict any annotation and provides meaningful context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that states the action, the precondition, and the verification step with no wasted words. The key information is front-loaded and every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and annotations cover the destructive nature, the description is largely complete. It includes the critical precondition (stopped) and the verification behavior. It does not discuss the concurrency guard or return values, but those are covered by schema and annotations, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter already has a clear description in the schema. The tool description adds nothing about parameters (e.g., no clarification of 'expected_before' or 'session_fingerprint'), so it stays at the baseline of 3 where the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Set') and resource ('project tempo') and adds a key precondition ('while stopped'). It also mentions a verification step. While it doesn't explicitly name a sibling tool to distinguish from, the purpose is clear and specific enough that an agent can tell what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for setting tempo and notes the precondition 'while stopped,' which gives some contextual guidance. However, it does not mention alternatives or when not to use it, nor does it explain why 'while stopped' is required or how to sequence with other tools like reading tempo.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_time_signature_numeratorA
Destructive

Set and prove beats per bar from FL's PPB/PPQ getter pair.

ParametersJSON Schema
NameRequiredDescriptionDefault
numeratorYesBeats per bar. FL exposes no denominator getter.
expected_beforeNoOptional expected current numerator.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
project_savedNo
bridge_commandNo
schema_versionNo
undo_point_createdNo
verification_basisNo
requested_numeratorYes
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false, so the write nature is known. The description adds that the tool 'proves' the set operation via FL's PPB/PPQ getter pair, and the schema explains the session_fingerprint is a concurrency guard that refuses after bridge reload or project load. This adds behavioral context beyond annotations, though it doesn't detail side effects or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that conveys the core action and verification mechanism without waste. It is front-loaded with the verb 'Set' and the resource, and the PPB/PPQ reference is meaningful to FL context. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and the schema covers all parameters, the description is mostly complete. It explains the core behavior (set and prove) and the schema covers the concurrency guard semantics. A minor gap is that it doesn't explicitly state the numerator range (1-32) or that it's a write operation, but those are in the schema/annotations. Overall adequate for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds the key semantic that the numerator is 'beats per bar' and that FL exposes no denominator getter, which clarifies why only the numerator is set. The expected_before and session_fingerprint parameters are well-described in the schema, so the description doesn't need to repeat them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Set and prove beats per bar from FL's PPB/PPQ getter pair' clearly identifies the action (set and prove), the resource (beats per bar / time signature numerator), and the mechanism (FL's PPB/PPQ getter pair). It distinguishes this from sibling tools like fl_set_tempo or fl_set_loop_mode, though it doesn't explicitly name a sibling alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context: it sets the numerator and verifies it via FL's PPB/PPQ getter pair. The schema adds that expected_before is an optional concurrency guard and session_fingerprint is a bridge/project-session guard. However, the description itself doesn't explicitly state when to use this vs alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_track_eqA
Destructive

Set gain and/or frequency on one built-in EQ band; at least one is required.

ParametersJSON Schema
NameRequiredDescriptionDefault
band_indexYesWhich band of the track's built-in three-band EQ: 0 low, 1 mid, 2 high.
track_indexYesZero-based mixer index. Index 0 is Master and is refused unless allow_master is true.
allow_masterNoDeliberately target the master bus at index 0.
expected_beforeNoOptional expected gain and/or frequency; refuse if any supplied field changed.
gain_normalizedNoBand gain, normalized 0.0 to 1.0; 0.5 is flat. Omit to leave the gain alone.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.
frequency_normalizedNoBand centre frequency, normalized 0.0 to 1.0. Omit to leave the frequency alone.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
applied_atYes
band_indexYes
track_indexYes
gain_verifiedNo
project_savedNo
bridge_commandNo
schema_versionNo
targeted_masterNo
frequency_verifiedNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
requested_gain_normalizedNo
session_precondition_appliedNo
requested_frequency_normalizedNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the write nature and side-effect risk are known. The description adds the at-least-one constraint but no additional behavioral context such as immediate application, undo behavior, or refusal conditions. With annotations covering the safety profile, this is adequate but not enriched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with the action front-loaded and the critical at-least-one condition included. No filler or redundant restatement of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The rich schema descriptions, output schema, and annotations cover parameter semantics, return values, and side-effect risk. The description supplies the one missing invocation rule (at least one of gain/frequency) and enough resource scoping to make a correct call. It does not over-explain details already present in structured data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a valuable cross-parameter constraint ('at least one is required') that the JSON schema does not enforce, and maps 'gain and/or frequency' to the two optional parameters. This goes beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Set'), names the exact resource ('one built-in EQ band'), and states what can be changed ('gain and/or frequency'). It clearly distinguishes this from plugin-param and other mixer setters by limiting scope to the built-in EQ, and it includes the at-least-one requirement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'built-in EQ band' qualifier gives context and implies this is not for plugin EQs, but there is no explicit when-to-use or alternative routing (e.g., 'use fl_set_plugin_param for plugin EQs'). The usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_set_write_modeA
DestructiveIdempotent

Enable or disable writes for this bridge session without restarting FL.

This changes no project value and never persists the setting. Enabling is
refused unless `confirm_user_present` is literally true. The result is
verified with a fresh bridge handshake before it reports success.
ParametersJSON Schema
NameRequiredDescriptionDefault
enabledYesAbsolute session write state. True exposes only the bounded verified write tools; false locks them again.
confirm_user_presentNoTrue asserts that the user requested project changes or write access in this task. No separate mode request is needed. Not required to disable.

Output Schema

ParametersJSON Schema
NameRequiredDescription
changedYes
verifiedNo
warningsNo
changed_atYes
bridge_modeYes
session_onlyNo
after_enabledYes
project_savedNo
before_enabledYes
bridge_commandNo
schema_versionNo
requested_enabledYes
write_mode_originYes
verification_basisNo
session_fingerprintYes
confirmation_appliedYes
confirmation_requiredYes
startup_default_enabledYes
runtime_write_mode_controlNo
session_precondition_appliedNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=false, destructiveHint=true, idempotentHint=true, but the description adds crucial context: it 'changes no project value and never persists the setting,' which clarifies that the tool is non-destructive itself and non-persistent. It also discloses the precondition for enabling (refused unless `confirm_user_present` is literally true) and that the result is verified with a fresh bridge handshake before success is reported. This goes beyond the annotations by explaining the verification and non-persistence aspects. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. The first sentence states the primary purpose immediately, and subsequent sentences provide essential clarifications: no project value change, no persistence, the confirmation requirement, and the verification step. Every sentence adds value without redundancy. It is front-loaded and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the key operational aspects: what the tool does, its side effects (none on project values), its persistence (none), the precondition for enabling, and the verification mechanism. It does not explicitly describe the return value, but an output schema exists, so that is covered. It also does not elaborate on the exact effect of disabling, but the schema mentions 'false locks them again,' which covers that. Overall, it is complete for a mode-toggle tool with these annotations and schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are well-documented in the schema. The description reinforces the meaning of `confirm_user_present` by stating that enabling is refused unless it is literally true, which adds practical context beyond the schema's description. It also clarifies the `enabled` parameter's effect via the phrase 'enable or disable writes.' Since the schema already covers the basics, the description adds a small but valuable semantic nuance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Enable or disable writes for this bridge session without restarting FL.' It specifies the resource (bridge session) and the action (enable/disable writes), and distinguishes itself from direct parameter-changing tools by noting it 'changes no project value.' This is a specific, unambiguous purpose that sets it apart from siblings like fl_set_mixer_volume.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly communicates when to use this tool: it is a session-level mode toggle, not a direct project modification. It states 'This changes no project value,' which tells the agent this is not for direct edits, and it clarifies that enabling requires `confirm_user_present` to be true. While it does not explicitly say 'use this before performing write operations,' the context and the contrast with other fl_set_* tools make the usage clear enough. Lacks an explicit exclusion or alternative mention, but the purpose is self-evident.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_stopA
Destructive

Stop playback, set normalized position to zero, and verify both fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
expected_beforeNoOptional expected playing and/or position state.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
verifiedYes
warningsNo
applied_atYes
after_playingNo
project_savedNo
before_playingNo
bridge_commandNo
schema_versionNo
playing_verifiedYes
position_verifiedYes
requested_playingNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo
after_song_position_normalizedNo
before_song_position_normalizedNo
requested_song_position_normalizedNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds a sequenced behavior beyond the annotations: it stops, rewinds to zero, and verifies both playing and normalized-position fields. Annotations already convey mutability and destructiveness, and the description's added verification step is valuable context rather than a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence states the action, the target state, and the verification behavior with no filler. Every word earns its place, making this a model of concise tool documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-optional-parameter tool with rich annotations and an output schema, the core behavior is fully captured. The main missing piece is a short note on expected_before's pre-condition role and the session_fingerprint concurrency guard, but those are already described in the schema, so the description remains adequate rather than incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents expected_before and session_fingerprint; the baseline stays at 3. The phrase 'verify both fields' maps loosely to playing and song_position_normalized, but it does not explain how expected_before acts as a pre-condition or how the session fingerprint gates the write, so the description adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies a concrete action sequence—stop playback, rewind to normalized position zero, verify the resulting state—so the tool's purpose is unambiguous. This combination clearly differentiates fl_stop from granular siblings like fl_set_playing and fl_set_song_position, even without naming them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to prefer fl_stop over related control tools, nor any prerequisites or exclusions. The agent must infer the use case from the name and sibling list, which is exactly what the description should have made explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_trigger_noteB

Audition a global channel with a bounded note-on/off dispatch receipt.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteYesMIDI note number.
velocityYesMIDI note-on velocity.
duration_msNoBounded audition duration in milliseconds.
midi_channelNoFL MIDI channel override; -1 uses the default.
channel_indexYesGlobal channel index.
expected_beforeNoOptional observation-scoped channel fingerprint guard.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteYes
velocityYes
warningsNo
dispatchedYes
duration_msYes
index_scopeNo
midi_channelYes
channel_indexYes
dispatched_atYes
note_off_sentYes
project_savedNo
bridge_commandNo
schema_versionNo
undo_point_createdNo
verification_basisNo
session_fingerprintNo
expected_before_appliedNo
session_precondition_appliedNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds the useful context that the note is bounded (note-on/off) and that a receipt is returned, complementing annotations that mark it non-readonly and non-idempotent. However, it does not disclose potential side effects such as audible playback, transport interaction, or whether any project state is changed beyond the audition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It is concise, though 'dispatch receipt' is slightly opaque and could have been clarified without much added length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters and an output schema, the short description is acceptable but relies heavily on the schema and title for context. It does not explain the optional concurrency guards (expected_before, session_fingerprint) or why an agent would choose this over related channel tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter already has a meaningful description, so the baseline of 3 applies. The tool description itself does not add parameter-level meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Audition') and identifies the resource ('a global channel') and the action ('bounded note-on/off dispatch receipt'). It is clear enough to distinguish from the many fl_set_channel_* siblings, though 'dispatch receipt' is jargon that could confuse an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives, nor any exclusions. It does not mention that this is for previewing a note without committing to pattern edits, which would be valuable among the large sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fl_undoA
Destructive

Move to the previous absolute undo-history position and verify it.

ParametersJSON Schema
NameRequiredDescriptionDefault
expected_beforeNoOptional history position/count/dirty guard.
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
verifiedYes
warningsNo
directionYes
applied_atYes
project_savedNo
bridge_commandYes
schema_versionNo
requested_positionYes
undo_point_createdNo
verification_basisNo
session_fingerprintNo
verification_summaryYes
expected_before_appliedNo
session_precondition_appliedNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and idempotentHint=false. The description adds useful context by noting the movement is to an absolute undo-history position and that verification occurs, but it does not disclose failure behavior or redo implications beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no waste. It states the action and the verification postcondition clearly, and every word contributes to understanding the tool's core behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Combined with the detailed parameter descriptions, annotations, and the presence of an output schema, the description is largely complete enough for correct invocation, including calling with no required parameters. It could be more explicit about prerequisites or alternatives, but the structured fields substantially reduce the burden on the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both optional parameters already have thorough descriptions, including the expected_before history guard and the session_fingerprint concurrency guard. The description itself adds no parameter info, but the schema carries the burden, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states what the tool does: it moves to the previous absolute undo-history position and verifies the result. The verb and resource are specific, and the word 'previous distinguishes it from the sibling fl_redo tool without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than explicit: the description conveys that you use this tool when you want to step backward in project history, but it does not name alternatives such as fl_redo or fl_get_project_history or state when not to use them. This is adequate but not strongly directive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

midi_export_type1B
Destructive

Write a standard Type-1 file, reopen it, parse it, and verify its digest/events.

ParametersJSON Schema
NameRequiredDescriptionDefault
ppqNo
pathYesAbsolute .mid/.midi output path whose parent already exists.
tracksYes
numeratorNo
overwriteNoExplicitly allow atomic replacement of an existing file.
tempo_bpmNo
denominatorNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
ppqYes
pathYes
sha256Yes
verifiedYes
tempo_bpmYes
byte_countYes
note_countYes
exported_atYes
midi_formatNo
track_countYes
schema_versionNo
header_verifiedYes
readback_sha256Yes
atomic_replace_usedNo
musical_track_countYes
note_events_verifiedYes
overwritten_existing_fileYes

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructiveHint=true and readOnlyHint=false, so a write is expected. The description adds valuable behavioral context: it reopens and parses the file after writing, implying a round-trip verification that annotations do not cover. This is a meaningful disclosure beyond the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It conveys the primary operation and the verification behavior efficiently, though it could benefit from more structure or detail while remaining concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 7 parameters, nested track specs, and an output schema, this description is insufficient. It lacks information about track structure, defaults, the meaning of digest/events, or any constraints beyond 'standard Type-1'. The output schema partially mitigates return-value concerns, but the input behavior is under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 29% (only path and overwrite have descriptions). The description does not compensate for this low coverage; it mentions no parameters at all. With 7 parameters including tracks, ppq, tempo, and time signature, the agent gets no additional semantic help from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (write), the resource (standard Type-1 file), and adds an extra verification step (reopen, parse, verify digest/events). It is specific and easily distinguishable from any hypothetical alternatives, even though no sibling export tools exist.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives, prerequisites (e.g., parent directory exists, but that's in the schema), or exclusions (e.g., not for Type-0 exports). The description only implies 'standard Type-1' but doesn't elaborate on use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_apply_planB
Destructive

Apply a plan once through the verified batch kernel.

ParametersJSON Schema
NameRequiredDescriptionDefault
plan_idYes
stop_on_unverifiedNoSkip remaining plan items after unverified proof.

Output Schema

ParametersJSON Schema
NameRequiredDescription
planYes
batchYes
applied_atYes
schema_versionNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, and idempotentHint=falseholo. The description adds the 'once' constraint and 'verified batch kernel' mechanism, which is some behavioral context. However, it does not explain what gets modified, what 'verified' entails, or failure behavior. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every word contributes meaning, and the core action is stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the large ecosystem of plan-related tools and the mutation nature of this operation, the description lacks workflow context: what 'verified' means, the relationship to mix_create_plan/mix_get_plan, and when to choose this over processing_apply_plan. Output schema and annotations cover safety and returns, but operational context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description's use of 'a plan' minimally maps to the plan_id parameter, but otherwise adds no detail beyond the schema. stop_on_unverified is already documented in the schema. With 50% schema description coverage, the description does not significantly compensate for the undocumented plan_id, but the parameter is inferable from context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies a clear verb ('Apply') and resource ('a plan'), with the qualifiers 'once' and 'verified batch kernel' adding operational specificity. It does not explicitly say 'mix plan' and could be confused with siblings like processing_apply_plan, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool vs alternatives. The description does not mention preconditions (e.g., plan must be reviewed/verified) or exclude cases, and it does not reference related siblings like mix_create_plan or processing_apply_plan.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_create_gain_stage_planA

Create, but do not apply, dB-fader changes from a peak watch.

ParametersJSON Schema
NameRequiredDescriptionDefault
watch_idYes
allow_masterNoExplicitly include Master in the proposed plan.
target_peak_dbfsNo
max_adjustment_dbNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
planNo
watchYes
warningsNo
generated_atYes
schema_versionNo
skipped_tracksNo
target_peak_dbfsYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds a key behavioral fact beyond the annotations: the operation creates a plan without applying fader changes. But it doesn't disclose what creating a plan entails (e.g., whether it replaces existing plans, persists, or returns a plan identifier), and openWorldHint=true leaves room for unspecified effects. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one front-loaded sentence that immediately communicates the most important distinction ('Create, but do not apply'). There is no filler, repetition of schema fields, or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the critical side-effect boundary (no application) and the output schema presumably handles return values. However, it omits workflow context such as watch_id being obtained from a prior peak-watch session and that the resulting plan is meant to be applied later. This is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, with only allow_master having a description. The tool description does not explain watch_id, target_peak_dbfs, or max_adjustment_db beyond their names; 'from a peak watch' and 'dB-fader changes' provide minimal context but do not compensate for the lack of parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Create'), a specific resource ('dB-fader changes from a peak watch'), and explicitly negates application with 'do not apply'. This makes the tool clearly distinguishable from apply/execute tools and from generic plan-creation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Do not apply' gives implicit usage guidance by signaling that this tool only builds a plan and does not execute it. However, it does not explicitly say when to prefer this tool over siblings like mix_create_plan or how it fits into the peak-watch workflow, and it names no alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_create_planC

Store a session-bound closed-union plan; no project value changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
rationaleNo
operationsYes
session_fingerprintNoOptional bridge/project-session fingerprint from a recent read. The write refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.

Output Schema

ParametersJSON Schema
NameRequiredDescription
titleYes
sourceYes
statusNo
plan_idYes
warningsNo
rationaleNo
created_atYes
operationsYes
schema_versionNo
session_fingerprintYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the broad safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds useful beyond-annotation context by stating the tool stores a plan and makes 'no project value changes.' However, it does not clarify side effects of repeated calls, whether an existing plan is overwritten, or what 'closed-union' means operationally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core action, and the 'no project value changes' clause is valuable. However, the term 'closed-union plan' is unexplained jargon, and the brevity borders on under-specification rather than effective conciseness for a tool of this complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a large schema, 23 operation variants, and many related plan/lifecycle tools, the description is too incomplete. It does not explain where the stored plan goes, how it is later retrieved or applied, what 'session-bound' implies for durability, or how this relates to mix_apply_plan and mix_get_plan. The presence of an output schema reduces the need to describe return values, but lifecycle and selection context are still missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, and the description provides no parameter-level guidance. It does not explain how title, rationale, operations, or session_fingerprint should be populated, nor does it help the agent understand the 23-way discriminated union inside 'operations.' The description must compensate for the sparse schema but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Store') and a specific resource ('session-bound closed-union plan'), and explicitly clarifies that no project values change. This is clear and helps distinguish plan creation from plan application, though it does not name or contrast siblings like mix_create_gain_stage_plan or mix_apply_plan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives such as mix_apply_plan, mix_get_plan, processing_plan, or mix_create_gain_stage_plan. The phrase 'session-bound' hints at a lifecycle, but no explicit when-to-use, prerequisites, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_doctorA
Read-onlyIdempotent

Diagnose a real bounce with explicit policy thresholds and no mutation.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoTechnical review target.balanced
vocal_pathNoOptional synchronized vocal stem.
max_secondsNo
candidate_pathYesAbsolute path to the candidate bounce.
reference_pathNoOptional absolute reference path.
instrumental_pathNoOptional synchronized instrumental stem.

Output Schema

ParametersJSON Schema
NameRequiredDescription
issuesYes
targetYes
analysisYes
warningsNo
diagnosed_atYes
warning_countYes
critical_countYes
policy_versionNo
schema_versionNo
masking_analysisNo
mutations_appliedNo
reference_comparisonNo
technical_export_readyYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the 'no mutation' constraint and 'explicit policy thresholds' context, but doesn't describe what those thresholds are or any other runtime behavior; no contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is one front-loaded sentence with every word carrying meaning: diagnose, real bounce, policy thresholds, no mutation. No filler or redundant background.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich annotations, high schema coverage, and presence of an output schema, the main invocation details are covered; the agent knows it needs a candidate_path and can optionally supply reference/instrumental/target. The only residual gap is the undefined meaning of 'policy thresholds,' which is likely resolved by the output schema or domain context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 83% schema description coverage, the schema already documents candidate_path, target, reference_path, vocal_path, and instrumental_path. The description adds no parameter-specific meaning and doesn't clarify the undocumented max_seconds or how 'policy thresholds' map to parameters, so it stays at the schema-heavy baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('diagnose') and identifies a concrete resource ('a real bounce'), and explicitly states it performs no mutation. This distinguishes it from mutation/planning siblings such as mix_apply_plan or mix_create_plan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for diagnosing a real bounce file with policy thresholds, but it never states when to prefer it over sibling diagnostics such as mix_masking_recommendations or audio_compare_files, nor lists exclusions. This leaves usage to inference rather than explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_finish_assessmentB
Read-onlyIdempotent

Run the end-to-end read-only finish assessment and stop at user export.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoTechnical review target.balanced
vocal_pathNo
max_secondsNo
candidate_pathYesAbsolute candidate bounce path.
reference_pathNo
instrumental_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
doctorYes
next_stepsYes
assessed_atYes
schema_versionNo
mutations_appliedNo
stopping_boundaryNo
plugin_compatibilityYes
render_available_through_fl_apiNo

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false. The description adds meaningful behavioral context by stating the run is 'end-to-end' and that it stops before user export, which implies it does not perform export side effects. This goes beyond what the annotations express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tightly constructed sentence with no filler. The core action and the stopping behavior are front-loaded. It is concise, though the brevity does sacrifice some clarifying detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite an output schema, the description leaves major gaps: it does not explain what the finish assessment consists of, what inputs are needed, or how this relates to sibling mix-assessment tools. For a tool with six parameters and a multi-step workflow, this is insufficient context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (one parameter described), and the tool description gives no parameter guidance at all. With low schema coverage, the description should compensate for the six parameters, but it does not explain candidate_path, target, or the optional reference/instrumental/vocal paths.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description names a specific verb ('Run') and resource ('end-to-end read-only finish assessment'), and adds a distinctive outcome ('stop at user export'). However, 'finish assessment' is not defined, and it does not differentiate from siblings like mix_doctor or mix_reference_recommendations, which could also be considered assessments.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, no scenarios, and no exclusions. Sibling tools with overlapping assessment purposes exist, but the description does not help an agent choose among them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_get_peak_watchA
Read-onlyIdempotent

Read cumulative sampled peaks without stopping the watch.

ParametersJSON Schema
NameRequiredDescriptionDefault
watch_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
statusYes
tracksYes
watch_idYes
only_usedYes
max_tracksYes
started_atYes
finished_atNo
frame_countYes
interval_msYes
limitationsNo
schema_versionNo
session_fingerprintNo
requested_duration_secondsYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds a meaningful behavioral detail beyond those annotations: it does not stop the running watch, and it returns cumulative sampled peaks. This clarifies both the side-effect profile and the data semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one tight sentence that front-loads the core action and the critical behavioral qualifier. Every word earns its place; there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single parameter, the existence of an output schema, and the read-only annotations, the description is nearly complete. It could also explicitly mention that a watch must already be running, since mix_start_peak_watch is a sibling, but the wording 'without stopping the watch' makes that context reasonably inferable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, watch_id, is not explained in the description, and schema description coverage is 0%. However, the schema provides a clear title and a 32-character hex pattern, and the parameter's meaning is evident from its name. The description adds no additional parameter-level value, but for a single obvious parameter this is acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Read cumulative sampled peaks.' It also includes the key qualifier 'without stopping the watch,' which clearly distinguishes this tool from mix_stop_peak_watch and mix_start_peak_watch. The purpose is immediately understandable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys a clear usage context: use this when you want to sample peaks while the watch continues running. It does not explicitly name alternatives or state when not to use it, but the 'without stopping the watch' phrasing effectively implies the relevant contrast with the stop tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_get_planB
Read-onlyIdempotent

Read one process-local reviewable plan.

ParametersJSON Schema
NameRequiredDescriptionDefault
plan_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
titleYes
sourceYes
statusNo
plan_idYes
warningsNo
rationaleNo
created_atYes
operationsYes
schema_versionNo
session_fingerprintYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds only 'process-local' and 'reviewable' context, which hints at ephemerality and review purpose but does not disclose limitations like missing-plan behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately short for a simple read operation, though slightly under-specified in content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one simple required parameterhol, an output schema present, and annotations covering read-only/idempotent behavior, minimal description is acceptable. However, the lack of usage guidance and the vague 'process-local' term leave gaps that an agent might need filled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not compensate: 'plan_id' is left entirely to the schema's title and pattern. The word 'one ... plan' weakly implies selection by ID, but no additional meaning is added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Read') and resource ('one process-local reviewable plan'), which is clear enough to distinguish the core action. It does not explicitly differentiate from siblings like mix_create_plan or processing_plan, but 'process-local reviewable plan' adds meaningful scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives such as mix_create_plan, mix_apply_plan, or postfader_review_get. The description implies a read operation but never states prerequisites, sequencing, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_inspect_plugin_compatibilityA
Read-onlyIdempotent

Report which loaded effects have known parameter-role adapters.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_usedNoFilter conservatively to used mixer tracks.

Output Schema

ParametersJSON Schema
NameRequiredDescription
matchesYes
warningsNo
observed_atYes
profiled_countYes
schema_versionNo
unprofiled_countYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, destructiveHint=false, so the safe, non-destructive nature is established. The description contributes scope: it reports only loaded effects and only those with known adapters, which is useful behavioral context. It does not describe the response format or the effect of only_used, but those are covered by the output schema and parameter schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with no filler; the core object (loaded effects) and the condition (known parameter-role adapters) come first. It earns its place, though 'parameter-role adapters' is domain jargon that is not expanded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with one optional well-documented parameter and an output schema, the description gives the essential semantic. It leaves minor gaps around when to use this versus sibling inspection/scanning tools and what an adapter/profile match means, but overall it is sufficient for a simple report tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, only_used, has a complete schema description plus a default value, so schema coverage is 100%. The description does not need to repeat or expand on it; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description has a specific verb ('Report'), a specific resource ('loaded effects'), and a precise filter ('known parameter-role adapters'), so an agent can understand this is an inspection/compatibility reporting tool. It is not a tautology, and it reads differently from parameter-scanning or loading siblings. However, it does not explicitly call out how it differs from nearby tools such as plugins_scan_loaded_plugins or mix_list_plugin_profiles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The read-only phrasing implies use when an agent needs to know which loaded effects already have parameter-role adapter coverage, likely before planning parameter assignments. It gives no explicit when-to-use/when-not-to-use context and does not name alternatives among the many inspection tools, so the agent must infer the selection from the tool name and wording.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_list_plugin_profilesB
Read-onlyIdempotent

List bundled parameter-role adapters and processing recipes.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryNoOptional exact profile category.

Output Schema

ParametersJSON Schema
NameRequiredDescription
profilesYes
warningsNo
profile_countYes
schema_versionNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the 'bundled' qualifier, indicating these are built-in profiles rather than user-created ones, which is a modest addition of context. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, efficient sentence with no filler. The description is front-loaded with the action verb and resource. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a single optional parameter, an output schema present, and annotations covering the safety profile, the definition is largely complete. The only gap is that the jargon 'parameter-role adapters' and 'processing recipes' is unexplained, which an agent might not fully parse without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single optional parameter 'category' is fully described in the schema ('Optional exact profile category'). The description adds nothing about the parameter beyond what the schema provides, so the baseline 3 for high coverage is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and names the resource ('bundled parameter-role adapters and processing recipes'). The terminology is somewhat jargony, but it clearly denotes a distinct resource type. It differentiates from siblings like plugins_list_available or plugins_list_presets through the 'parameter-role adapters' and 'processing recipes' framing, though it doesn't explicitly name those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many sibling list tools. No conditions, no exclusions, and no mention of what category values might be valid or when the optional filter should be applied. An agent must infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_masking_recommendationsA
Read-onlyIdempotent

Recommend bounded dynamic remediation from sample-synchronous stems.

ParametersJSON Schema
NameRequiredDescriptionDefault
vocal_pathYesAbsolute synchronized vocal stem path.
max_secondsNo
instrumental_pathYesAbsolute synchronized instrumental stem path.

Output Schema

ParametersJSON Schema
NameRequiredDescription
analysisYes
warningsNo
actionableYes
generated_atYes
remediationsYes
schema_versionNo
mutations_appliedNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the agent knows this is a safe, non-mutating operation. The description adds the behavioral context that the tool works on sample-synchronous stems and produces bounded dynamic remediation, which is useful. However, it does not disclose what the output looks like (though an output schema exists), whether it performs analysis only, or any limitations. With annotations covering the safety profile, a 3 is appropriate – the description adds some value but not rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core action ('Recommend') and key constraints ('bounded dynamic', 'sample-synchronous stems'). It is efficient with no wasted words. The title adds a redundant but harmless restatement. It could arguably be slightly more explicit about the output, but for its length it is well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema, the description need not explain return values. The annotations cover safety (read-only, idempotent, non-destructive). The description covers the core input requirement (sample-synchronous stems) and the nature of the output (bounded dynamic remediation). The main gap is the lack of guidance on max_seconds semantics and when to prefer this over audio_analyze_masking or mix_reference_recommendations, but for a read-only recommendation tool with an output schema, the description is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%: vocal_path and instrumental_path have descriptions ('Absolute synchronized vocal stem path' and 'Absolute synchronized instrumental stem path'), but max_seconds has no description beyond its schema type/default. The description's phrase 'sample-synchronous stems' reinforces the path parameters' meaning, adding a synchronization requirement not fully explicit in the schema. However, it doesn't explain max_seconds semantics (e.g., what happens when null, how it bounds the analysis). The description adds marginal value over the schema but doesn't fully compensate for the max_seconds gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Recommend bounded dynamic remediation from sample-synchronous stems' states a specific verb ('recommend'), a resource ('remediation'), and a key constraint ('bounded dynamic', 'sample-synchronous stems'). It distinguishes itself from siblings like mix_reference_recommendations and audio_analyze_masking by focusing on remediation recommendations from synchronized stems, though it doesn't explicitly name those siblings. The title 'Recommend masking remediation' reinforces the purpose, but the description alone is slightly jargon-heavy and could be clearer about what 'remediation' entails.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: when you have sample-synchronous vocal and instrumental stems and want bounded dynamic remediation recommendations. It does not explicitly state when NOT to use it or name alternatives like audio_analyze_masking or mix_reference_recommendations. The context of sibling tools suggests a mix-analysis workflow, but the description leaves the agent to infer the exact conditions. No exclusions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_reference_recommendationsA
Read-onlyIdempotent

Return bounded tonal review ranges only when alignment/readiness passes.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_secondsNo
candidate_pathYesAbsolute path to the candidate bounce.
reference_pathYesAbsolute path to the reference audio.

Output Schema

ParametersJSON Schema
NameRequiredDescription
warningsNo
actionableYes
comparisonYes
adjustmentsYes
generated_atYes
schema_versionNo
mutations_appliedNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context: it returns ranges only conditionally ('only when alignment/readiness passes'), implying it may return nothing or an empty result when the condition fails. This is useful beyond the annotations. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is compact and front-loads the core action ('Return bounded tonal review ranges') followed by the gating condition. No wasted words. The condition is placed at the end, which is acceptable given the sentence's brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are presumably documented there. The description covers the gating condition but leaves ambiguity about what 'alignment/readiness' refers to and what happens when the condition fails (empty result vs. error). Given the tool's moderate complexity and the presence of an output schema, the description is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%: candidate_path and reference_path have descriptions, but max_seconds has only a title and default. The description doesn't explain max_seconds semantics (e.g., how it bounds the analysis) or clarify the relationship between reference_path and candidate_path beyond their names. The description adds the concept of 'alignment/readiness' but doesn't map it to specific parameters. Baseline 3 is appropriate since the schema covers most parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('bounded tonal review ranges') with a condition ('only when alignment/readiness passes'). It distinguishes itself from sibling tools like mix_masking_recommendations and audio_compare_files by focusing on tonal review ranges gated by alignment/readiness. However, it doesn't explicitly name a sibling alternative, and the phrase 'bounded tonal review ranges' is somewhat jargon-heavy without elaboration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The condition 'only when alignment/readiness passes' implies a prerequisite: the tool should be used only after alignment/readiness checks succeed. This gives some usage context but doesn't explicitly state when NOT to use it or name alternatives. The sibling list includes postfader_creation_readiness and mix_doctor, which could be related, but the description doesn't route the agent to them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_resolve_processing_intentA
Read-onlyIdempotent

Map an intent to loaded profiled controls without applying settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
intentYesOutcome-level processing intent.
strengthNoReviewed artistic strength hint.
track_indexYesMixer track to inspect.

Output Schema

ParametersJSON Schema
NameRequiredDescription
readyYes
stepsYes
intentYes
strengthYes
warningsNo
resolved_atYes
track_indexYes
schema_versionNo
mutations_appliedNo
missing_categoriesYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds the useful prerequisite that the tool operates on 'loaded profiled controls', which is not in the annotations. It also reiterates non-applying behavior, providing context beyond the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no unnecessary words. It front-loads the core action and scope, making it immediately understandable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a read-only annotation, an output schema available, and a 100% parameter coverage schema, the description covers the essential purpose and prerequisite. It could optionally mention the output format, but since an output schema exists, that is not required. The description is adequate for an inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema. The tool description adds no additional parameter semantics beyond what is already present, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('map') and resource ('intent to loaded profiled controls') and explicitly notes it does not apply settings, clearly distinguishing it from applying tools like mix_apply_plan. It conveys the core functionality without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without applying settings' implies this tool is for inspection or preview, but it does not explicitly state when to prefer it over alternatives or when not to use it. No sibling alternatives are mentioned, making usage guidance only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_start_peak_watchB

Start a process-persistent sampled peak watch and return its first frame.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_usedNoRetain active/custom-named tracks plus Master.
max_tracksNoMaximum mixer indices scanned.
interval_msNoSampling interval.
duration_secondsNoWatch duration.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
statusYes
tracksYes
watch_idYes
only_usedYes
max_tracksYes
started_atYes
finished_atNo
frame_countYes
interval_msYes
limitationsNo
schema_versionNo
session_fingerprintNo
requested_duration_secondsYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=false and openWorldHint=true, so the description correctly implies a state-changing operation. The description adds the concepts of 'process-persistent' and 'sampled', and notes it returns the first frame, which goes beyond annotations. However, it does not disclose side effects such as whether starting multiple watches replaces existing ones, resource consumption, or how to stop it. Since annotations cover the basic safety profile (non-destructive), the description adds some value but lacks depth on behavioral nuances.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence that front-loads the primary action ('Start') and immediately conveys the resource and expected return. Every word earns its place; there is no fluff or repetition. It is appropriately concise for a tool that relies on schema for parameter details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool starts a persistent process, the description is incomplete. It does not explain how to retrieve subsequent frames (presumably via mix_get_peak_watch) or that the watch must be explicitly stopped (mix_stop_peak_watch). It also omits any mention of the watch scope (e.g., which mixer tracks are covered) beyond what parameters imply. An agent unfamiliar with the workflow may not know the sequence of calls. The output schema exists but is not shown; the description still should provide lifecycle context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All four parameters are fully documented in the schema with descriptions and defaults, and schema description coverage is 100%. The description itself adds no parameter-specific information beyond what the schema provides, so it does not compensate for any gaps. Baseline of 3 is appropriate because the schema carries the semantic load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('Start') and a specific resource ('process-persistent sampled peak watch') with an explicit outcome ('return its first frame'). It is distinguishable from siblings like mix_get_peak_watch (which presumably reads existing watches) and mix_stop_peak_watch, though it does not explicitly name them. The term 'mixer' is only in the title, not the description, but the tool name and title provide sufficient context. Slight gap: no explicit mention of what a 'peak watch' is for, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention that this should be called before mix_get_peak_watch, nor that mix_stop_peak_watch should be used to end the watch. There is no indication of prerequisites or typical workflow. An agent is left to infer usage from the tool name and siblings, which is insufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mix_stop_peak_watchA

Stop one process-local watch and return its final aggregate.

ParametersJSON Schema
NameRequiredDescriptionDefault
watch_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
statusYes
tracksYes
watch_idYes
only_usedYes
max_tracksYes
started_atYes
finished_atNo
frame_countYes
interval_msYes
limitationsNo
schema_versionNo
session_fingerprintNo
requested_duration_secondsYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry unsafe, non-idempotent, but not destructive hints. The description adds beyond these by scoping the watch as 'process-local' and by explicitly stating the stop returns a 'final aggregate', which helps the agent understand the outcome and scope of the operation. It does not, however, detail consequences like error behavior on nonexistent watches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tightly packed sentence: it states both the action and the return value with zero extraneous words. Structure is clear and fully front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an existing output schema, the description covers the essential operation and scoping. It omits edge-case behavior such as what happens when the watch was already stopped, but overall it is sufficiently informative for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. While 'one process-local watch' implies the watch_id parameter identifies the watch to stop, the description does not explicitly map the watch_id parameter to that identity, nor does it clarify format or constraints beyond the schema regex.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Stop') and resource ('process-local watch'), and clarifies the returned artifact ('final aggregate'). This clearly distinguishes it from sibling tools like mix_start_peak_watch and mix_get_peak_watch without requiring schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives. It does not mention that it pairs with mix_start_peak_watch or that it should be called only after a watch has been started, leaving the user to infer timing and selection context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

piano_roll_bridgeC

Manage the one-time-per-process arm for FL's separate Piano Roll script runtime.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoStatus only, write the bootstrap script, or confirm the user ran it once.status
confirm_user_ran_scriptNoRequired only for action='confirm', after the user manually ran Postfader Apply.

Output Schema

ParametersJSON Schema
NameRequiredDescription
platformYes
shortcutYes
observed_atYes
script_pathYes
script_existsYes
last_operationNo
schema_versionNo
last_request_idNo
scripts_directoryYes
setup_instructionYes
armed_this_sessionYes
last_requested_digestNo
prepared_this_sessionYes
last_requested_note_countNo
automatic_trigger_supportedYes
authoritative_fl_note_readback_availableNo

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds one meaningful behavioral trait beyond the annotations: the setup is 'one-time-per-process', which aligns with the non-idempotent annotation. However, it does not disclose what side effects occur (e.g., writing a bootstrap script) or what the 'arm' actually does at runtime. Annotations already indicate a mutating operation (readOnlyHint=false), so the description partially complements them but remains thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, and the 'one-time-per-process' qualifier is front-loaded. It is appropriately short for a tool whose parameter semantics live in the schema, though the cryptic wording ('arm') sacrifices some clarity for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even with a rich schema and output schema present, the description leaves major context gaps: it never explains what the bridge is, why a separate runtime exists, what 'arm' means, or when the tool must be invoked relative to other piano roll tools. An agent cannot determine the prerequisite ordering or the tool's role in a workflow from this description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both parameters have clear inline explanations, including the conditional need for confirm_user_ran_script. The tool description adds no parameter-level details, so the baseline of 3 applies because the schema carries the full semantic weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific resource (the one-time-per-process arm for FL's separate Piano Roll script runtime) but uses the vague verb 'manage' without naming the actual operations (status, prepare, confirm). It distinguishes itself from sibling note-reading/writing tools by focusing on the runtime bridge, but an agent would still need to open the schema to understand what the tool actually does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool or how it relates to alternatives like piano_roll_read_notes or piano_roll_write_notes. The 'one-time-per-process' phrase implies a setup scenario, but it never states whether this should be run before other piano roll tools or what the intended workflow is.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

piano_roll_read_notesA

Open a score and read notes without changing notes or enabling musical writes.

Requires the existing one-time piano_roll_bridge setup. Follow next_offset
to page; selected_only may return an empty page with a non-null next_offset.
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoRaw note indices per page.
offsetNoRaw score note offset.
channel_indexYesGlobal Channel Rack target index.
selected_onlyNoFilter selected notes within this raw page.
pattern_numberYesPattern to inspect.
session_fingerprintNoOptional expected bridge session.

Output Schema

ParametersJSON Schema
NameRequiredDescription
ppqNo
limitYes
notesNo
offsetYes
sourceNo
statusYes
targetYes
triggerYes
warningsNo
request_idYes
next_offsetNo
observed_atYes
channel_indexYes
project_savedNo
selected_onlyYes
notes_modifiedNo
pattern_numberYes
schema_versionNo
total_note_countNo

TDQS

A3.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds genuinely useful runtime details: bridge prerequisite, next_offset pagination, and the selected_only empty-page edge case. However, it claims the tool reads 'without changing notes or enabling musical writes' while the annotations declare readOnlyHint=false, which indicates the operation is not read-only. This is a direct annotation contradiction and forces a score of 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact, front-loaded sentences cover purpose, prerequisite, and pagination caveat without filler. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers setup, pagination, and a tricky selected_only edge case, while the output schema and input schema cover return and parameter details. The only real completeness issue is the unresolved contradiction between the description's read-only claim and readOnlyHint=false, which leaves the tool's behavior ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters. The description adds non-obvious behavioral semantics around next_offset and selected_only — specifically the possibility of an empty page with a non-null next_offset — which goes beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: 'Open a score and read notes' and explicitly excludes musical writes, distinguishing it from piano_roll_write_notes and piano_roll_transform. The title 'Inspect existing Piano Roll notes' reinforces this read-only intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: it is the tool to use for reading notes without musical writes, requires the one-time piano_roll_bridge setup, and explains pagination via next_offset. It does not explicitly name alternatives, but the read-only framing makes selection unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

piano_roll_transformB
Destructive

Quantize, transpose, humanize, duplicate, delete, or clear selected/all notes.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesClosed transform request read by FL's live score script.
auto_triggerNoAutomatically send the run-last-script shortcut.
channel_indexYesGlobal Channel Rack target index.
pattern_numberYesPattern to select before triggering the script.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeNo
statusYes
targetNo
triggerNo
warningsNo
operationYes
request_idYes
script_pathYes
requested_atYes
project_savedNo
script_sha256Yes
schema_versionNo
verification_targetNo
application_verifiedNo
requested_note_countNo
verification_triggerNo
requested_note_digestNo
script_runtime_evidenceNo
verification_script_sha256No
authoritative_note_readback_availableNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the delete/clear operations add no new information beyond the safety profile. The description does not disclose additional behavioral context such as how the tool interacts with the current piano roll, the auto_trigger mechanism, or whether it can be undone. It adds no value over annotations and the schema's parameter descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single clear sentence that front-loads the operation list and scope. Every word contributes for place zero filler. It is appropriately sized for a tool whose main purpose is simple though the underlying parameter space is complex.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-operation tool with a nested request schema, the description is too minimal. It does not explain which parameters apply to which operation, how the global channel_index/pattern_number context interacts with notes, or the role of the auto_trigger field. The output schema exists, so return values are covered, but an agent still lacks the operation–parameter mapping and scenario guidance to invoke the tool correctly without extensive manual schema exploration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage reported as 100%, the baseline is 3. The description lists operation names and the scope (selected/all) which mirrors the enum values already present in the schema, but does not explain how parameters like semitones, grid_beats, or repeats relate to each operation. No new semantic meaning is added beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names specific operations (quantize, transpose, humanize, duplicate, delete, clear) on a specific resource (notes) and the scope (selected/all). This clearly distinguishes it from siblings like piano_roll_read_notes and piano_roll_write_notes, so an agent can tell which tool to choose without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives, such as when to prefer it over piano_roll_write_notes or piano_roll_read_notes. The description provides no context about prerequisites, pattern/channel selection, or scenarios where this tool is the right choice, leaving the agent to infer usage from the operation list alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

piano_roll_write_notesC
Destructive

Generate a typed Piano Roll script, select its target, and report hotkey dispatch honestly.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoAppend to the score or clear all notes before adding these notes.append
notesYesBounded notes in quarter-note beat units.
auto_triggerNoSend FL's run-last-Piano-Roll-script shortcut after verified target selection.
channel_indexYesGlobal Channel Rack target index.
pattern_numberYesPattern to select before triggering the script.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modeNo
statusYes
targetNo
triggerNo
warningsNo
operationYes
request_idYes
script_pathYes
requested_atYes
project_savedNo
script_sha256Yes
schema_versionNo
verification_targetNo
application_verifiedNo
requested_note_countNo
verification_triggerNo
requested_note_digestNo
script_runtime_evidenceNo
verification_script_sha256No
authoritative_note_readback_availableNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state destructiveHint=true and readOnlyHint=false, so the agent knows it is a write operation. However, the description adds nothing else—no explanation of what is destroyed, side effects, or the meaning of 'report hotkey dispatch honestly.' It fails to provide behavioral context beyond what annotations already convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two cryptic sentences that do not front-load the core purpose. Phrases like 'report hotkey dispatch honestly' are ambiguous and add confusion rather than clarity. It is not concise in a helpful way.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters including a complex notes array, the description is severely incomplete. It does not explain the effect (writing notes), the meaning of mode/append/replace, the role of auto_trigger, or target selection. Schema describes parameters but the overall operation is unclear. The description fails to make the tool usable without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%—all parameters have descriptions in the schema. The tool description adds no parameter semantics. Baseline is 3 because the schema handles documentation, and the description offers no further value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description focuses on internal mechanics ('Generate a typed Piano Roll script, select its target, and report hotkey dispatch honestly') rather than the user-facing action of writing notes. The tool name implies writing notes, but the description does not clearly state that, and it could be mistaken for a scripting utility. It is not a tautology but is vague and does not specify the resource being modified.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus sibling tools like piano_roll_read_notes or piano_roll_transform. The description does not mention alternatives, prerequisites, or appropriate scenarios, leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_atlas_get_productB
Read-onlyIdempotent

Read one static Atlas product and its related descriptive records.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
vendorNo
productYes
adaptersNo
evidenceNo
schema_versionNo
registry_digestYes
stock_alternativesNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint, and openWorldHint, so the safety profile is covered. The description adds the 'static' qualifier and the fact that related descriptive records are returned, but it provides no additional behavioral detail beyond those facts, and it does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler and the key information front-loaded: what it reads, and that it is a single static product. Every word contributes to the tool's meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich annotations (read-only, idempotent, non-destructive, closed-world), a single required parameter, and the presence of an output schema, the description is largely sufficient. It only misses explicit usage guidance and identifier semantics, which are relatively minor for such a simple lookup operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the tool description needed to explain what product_id should contain and how it relates to the Atlas catalog. The description merely says 'Read one static Atlas product' and does not clarify the identifier format, source, or relationship to search results, leaving the agent to rely mostly on the parameter name and length constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Read'), a resource ('one static Atlas product'), and a scope ('related descriptive records'). It is clear, but it does not explicitly distinguish itself from sibling tools like plugins_atlas_search or plugins_atlas_inspect_loaded, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what the tool does but not when to use it over alternatives. There is no mention of 'use search when you lack an exact product ID' or 'use inspect_loaded for installed products,' so the agent has to infer the intended use case from the schema and sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_atlas_inspect_loadedC
Read-onlyIdempotent

Match the target-aware live Track B inventory to static Atlas knowledge.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
pluginsNo
warningsNo
observed_atYes
schema_versionNo
registry_digestYes

TDQS

C2.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds a bit of context by contrasting "live" live inventory against "static" Atlas knowledge. However, it does not explain behavioral details like how weak matches are handled, what matching entails, or what bounds the operation respects, so credit is limited.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler. It front-loads the core action 'Match' and is not bloated, though its brevity comes at the cost of clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the cryptic terminology, low parameter documentation, and no usage guidance, the description is not complete enough for correct selection and invocation. An output schema exists but does not compensate for the undefined parameters and domain-specific terms like 'Track B', 'weak', and 'target-aware.'

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has only one required 'request' object containing three parameters (only_used, match_limit, include_weak) with no individual descriptions, and schema description coverage is 0%. The description provides no meaning for these parameters, only the tautological label "Bounded live-inventory matching options." This is insufficient for an agent to know what values to supply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific-sounding but unexplained jargon: "target-aware live Track B inventory" and "static Atlas knowledge" never clearly identify the resource as loaded plug-ins. The annotation title clarifies that it matches loaded plug-ins to Plugin Atlas, but the description itself is too vague and does not distinguish this from sibling Atlas tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to use this tool versus alternatives like plugins_atlas_search, plugins_atlas_get_product, or plugins_scan_loaded_plugins. The phrase "live Track B inventory" implies inspecting loaded state, but it never explicitly states the use case, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_atlas_recommendB
Read-onlyIdempotent

Rank static plug-in choices or stock alternatives without changing FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
schema_versionNo
recommendationsNo
registry_digestYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is known. The description adds 'without changing FL', which reinforces the read-only nature but does not introduce new behavioral details like response format or rate limits. It adds minimal value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that front-loads the core function and a key constraint. No wasted words; it is appropriately sized for its purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complex schema (9 parameters) and the existence of an output schema, the description is too thin. It does not explain what 'static plug-in choices' or 'stock alternatives' mean, nor does it hint at the required inputs. An agent would need to inspect the schema to understand how to call this tool, making the description inadequate for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention any parameters, and schema description coverage is 0%. With 9 parameters in the nested request object, the description fails to explain their meanings or usage. The schema itself has a description, but that is not part of the tool description, so the agent gets no parameter guidance from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool ranks static plug-in choices or stock alternatives, using a specific verb ('Rank') and resource, and notes it does not change FL. This distinguishes it from siblings like plugins_atlas_search (searching) and plugins_atlas_get_product (product details).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for recommendations but does not explicitly say when to use this tool over alternatives such as plugins_atlas_search or plugins_atlas_get_product. The schema description mentions product_id for stock-alternative mode, but that is not in the tool description itself. No clear exclusions or alternative references are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_get_current_presetB
Read-onlyIdempotent

Read FL's current preset name and an index only when it is unique.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesExplicit mixer effect or global channel-generator target.

Output Schema

ParametersJSON Schema
NameRequiredDescription
pluginYes
warningsNo
observed_atYes
preset_countYes
schema_versionNo
observation_atomicNo
project_dirty_flagNo
current_preset_nameNo
session_fingerprintNo
current_preset_indexNo
current_preset_statusYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds one useful behavioral fact: the tool returns only when a uniqueness condition holds. However, it does not state what happens when that condition is not met—whether it returns an empty result, errors, or falls back. Since the annotations cover the core non-destructive nature, the description contributes partial context, earning a 3.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler: 'Read FL's current preset name and an index only when it is unique.' It is front-loaded with the action and resource. However, the phrase 'only when it is unique' introduces ambiguity that undermines clarity; it could have been clearer without extra length, but the structure itself is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a complete input schema (100% parameter coverage) and an output schema available, the description does not need to re-explain return types. However, it leaves open the critical case of what happens when the uniqueness condition fails, and it provides no guidance on when to select this tool relative to sibling preset tools. These gaps make the description incomplete for fully correct invocation, so a middle score is warranted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema fully defines the single 'target' parameter with its two variant object shapes and discriminators. The description does not need to repeat parameter mechanics and offers no semantic elaboration beyond the schema. By the rubric's baseline, the description adds no parameter-specific insight, so 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Read FL's current preset name and an index'. It conveys the tool's purpose and includes a distinguishing qualifier, 'only when it is unique'. However, the phrase 'only when it is unique' is ambiguous—whether it refers to the preset name, the index, or the preset itself—and does not explicitly differentiate this from siblings like plugins_get_plugin_preset_count or plugins_list_presets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives such as plugins_get_plugin_preset_count or plugins_list_presets. It only states an oblique condition ('when unique') without explaining what circumstances warrant this tool or when a different preset-related tool should be chosen. No exclusions or alternative selection criteria are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_inspect_pad_mapA
Read-onlyIdempotent

Read generic pad, MIDI-note, color, empty, and mute observations.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesExplicit mixer effect or global channel-generator target.

Output Schema

ParametersJSON Schema
NameRequiredDescription
padsYes
pluginYes
completeYes
warningsNo
pad_countYes
observed_atYes
schema_versionNo
observation_atomicNo
project_dirty_flagNo
session_fingerprintNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the read-only behavior is covered. The description adds value by specifying the kinds of observations returned (pad, MIDI-note, color, empty, mute), but it does not elaborate on what these mean or any quirks of the response. With annotations handling the safety profile, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence starting with the verb 'Read' and conveying the essential resource and content. Every word earns its place with no repetition of the title or schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema, the output schema, and the annotations that cover safety and idempotency, the description is sufficient for an agent to invoke the tool correctly. It is minimal, but nothing essential is missing for selection or basic usage; the schema explains the target parameter and the output schema explains the return value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single 'target' parameter is fully documented in the schema with a clear description and a discriminated union of target types. The tool description adds no parameter information, but the schema already carries the burden, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and names the exact resource ('pad map'), then enumerates the observation types covered: 'generic pad, MIDI-note, color, empty, and mute observations.' The title 'Inspect a plug-in pad map' reinforces this and distinguishes it from sibling plugins_inspect_parameter_map by the resource type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: an agent should call this when it wants pad-map observations rather than parameter-map or scan data. However, it provides no explicit when-to-use/when-not-to-use guidance or mention of alternatives, leaving the agent to infer selection from the tool name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_inspect_parameter_mapA
Read-onlyIdempotent

Read a bounded page of exposed parameters; never marks unknown controls safe.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNumber of parameter indices to scan in this page.
offsetNoFirst parameter index in this page.
targetNoExplicit mixer_effect or global channel_generator target. Mutually exclusive with legacy track_index/slot_index.
slot_indexNoLegacy zero-based effect slot (0 through 9). Supply it with track_index, or use target, never both.
name_filterNoOptional case-insensitive name substring.
track_indexNoLegacy zero-based mixer track index. Supply it together with slot_index, or use target, never both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior. The description adds meaningful context beyond those annotations by promising a conservative safety posture: unknown controls are never marked safe. This is a useful behavioral guarantee for an automation agent, though the concept of 'safe' is not elaborated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that front-loads the core action and follows with a key safety property. There is zero redundant language, and every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the rich input schema, output schema, and annotations, the description is nearly sufficient for correct invocation. The main gap is that it doesn't clarify how this relates to the sibling plugins_scan_parameters or what 'safe' means in output terms, but those are minor given the structured metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents limit, offset, target, track_index, slot_index, and name_filter. The description's phrase 'bounded page' loosely ties to limit/offset but adds no concrete parameter-level semantics beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and resource ('exposed parameters') and immediately distinguishes this tool from a full scan by noting it reads a 'bounded page.' The phrase 'never marks unknown controls safe' adds a distinct behavioral signature not present in sibling tool names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly implies this is for paged/partial inspection rather than whole-map scanning, so an agent gets context on when it fits. However, it never explicitly names alternatives like plugins_scan_parameters or states when not to use this tool, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_list_availableA

Read the native Add menu on macOS; opens/closes the menu and changes focus.

Reports exact loadable favorite names and instrument/effect kinds. This is menu availability, not proof of licensing or an exhaustive installed scan.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
sourceNo
entriesNo
platformYes
supportedYes
observed_atYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only carry limited hints, so the description adds useful behavioral context by stating that it opens/closes the native menu and changes focus. It also clarifies what the result represents and does not represent. No contradiction is present, though it could be more precise about focus restoration or missing menu handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences contain no filler: the action and side effect are front-loaded, output scope is stated, and the limitation is given. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With empty parameters, an output schema, and a clear description of the behavioral aim and limitations, the description provides everything an agent needs to select and invoke the tool correctly. Nothing essential to the caller is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so parameter spacing is already fully documented by the empty input schema. Per the zero-parameter baseline, the description does not need to compensate for anything.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly identifies the resource (native Add menu on macOS), the action (Read), and the exact output (loadable plugin names and instrument/effect kinds). It also distinguishes this from licensing proof and an exhaustive installed scan, differentiating it from siblings like plugins_scan_loaded_plugins.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear exclusions: it is menu availability, not proof of licensing or an exhaustive installed scan. However, it never names an alternative tool for those use cases, so the when-versus-alternatives guidance remains implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_list_presetsA
Read-onlyIdempotent

Read one deterministic preset page without changing the plug-in.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoBounded number of preset names in this page.
startNoFirst preset index to inspect.
targetYesExplicit mixer effect or global channel-generator target.
include_currentNoAlso report FL's current preset identity.
include_empty_namesNoRetain blank preset-name rows in the returned page.

Output Schema

ParametersJSON Schema
NameRequiredDescription
limitYes
startYes
pluginYes
partialNo
presetsYes
has_moreYes
warningsNo
truncatedNo
next_startNo
observed_atYes
preset_countYes
truncated_byNo
scanned_countYes
returned_countYes
schema_versionNo
duplicate_namesYes
blank_name_indicesYes
observation_atomicNo
project_dirty_flagNo
current_preset_nameNo
session_fingerprintNo
current_preset_indexNo
current_preset_statusYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false; the description reinforces this with 'Read', 'without changing', and 'deterministic'. The determinism language adds a slight extra guarantee beyond idempotence. No negative side effects or failures are mentioned, but those are not required given the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence containing an action, a resource, and a constraint. It is front-loaded and without vacuous filler, making the tool instantaneously easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema available and exhaustively described parameters, the description doesn't need to spell out return values. It could mention that targets include mixer effects or channel generators, but the schema already handles that. The tool is simple enough that an agent can invoke it correctly despite the terse summary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with every parameter (target, limit, start, include_current, include_empty_names) already meaningfully described in the input schema. The description's 'page' concept echoes start/limit but adds no deeper semantic, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the action ('Read') and the resource ('one deterministic preset page'), plus the critical constraint 'without changing the plug-in'. This clearly distinguishes it from mutating siblings like fl_select_plugin_preset or plugins_load, and from other plugin inspection tools like plugins_get_current_preset or plugins_scan_parameters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Without changing the plug-in' signals a read-only use case and implies that any tool modifying the plugin should be chosen instead, providing clear context. However, no explicit alternative tool names or 'use when/use instead' guidance is given, so it earns strong but not top marks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_loadA
Destructive

Load one macOS Add-menu plugin, then identify its new channel or effect slot.

Use plugins_list_available first. The task request authorizes the addition;
effect loading temporarily enables bridge writes only to select its track.
Unknown outcomes must be inspected before any new load attempt. Does not
save the project; Windows insertion is not implemented by this adapter.
ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesExact Add-menu name, kind and mixer destination for effects.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
requestYes
verifiedNo
warningsNo
menu_pathNo
dispatchedNo
observed_atYes
loaded_pluginNo
project_savedNo
undo_point_createdNo
session_fingerprintNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true, so the description doesn't need to repeat that, but it adds value by disclosing temporary write enablement, non-save behavior, and the need for post-load inspection. These details go beyond annotation signals, though a couple of specifics like permission requirements or revert behavior could be more explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense, front-loading the core action and outcome. Each sentence adds value: prerequisites, authorization context, side effects, and platform limitations. No unnecessary words or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema and detailed input schema, the description covers all essential context: when to use (after listing), authorization requirements, side effects (temporary write enablement), limitation (no Windows), and post-usage inspection. An agent has sufficient guidance to invoke it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a detailed plugin request schema that defines kind, name, track_index, allow_master, timeout_seconds, and session_fingerprint. The description mentions 'name, kind and mixer destination' but does not elaborate on each parameter beyond what the schema already provides, so the description adds minimal extra meaning, aligning with the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (load a plugin), the resource (macOS Add-menu plugin), and the expected outcome (identify new channel/effect slot). It distinguishes itself from the sibling plugins_list_available by explicitly instructing to use that first, making the purpose and scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage guidance: always call plugins_list_available first, only load when task request authorizes, and inspect unknown outcomes before further loads. It also notes platform limitations (no Windows insertion) and side effects (temporarily enables bridge writes), which helps the agent decide when to use this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_scan_loaded_pluginsA
Read-onlyIdempotent

Inventory loaded plug-ins without inserting or changing anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_usedNoApply the conservative used-track heuristic.
include_channel_generatorsNoAlso include Channel Rack generators with explicit channel_generator targets. False preserves the 0.11 mixer-effect-only response contract.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already set readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds 'without inserting or changing anything', which reinforces the read-only behavior but does not go beyond the annotations. It does not mention response format, pagination, or the effect of the include_channel_generators flag on the response contract. Since annotations carry most of the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that states the action and non-mutating property with zero extraneous words. It is perfectly front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only inventory tool with a clear output schema and comprehensive annotations, the description is largely sufficient. The only missing nuance is explicitly noting that the include_channel_generators flag affects the response contract, but that is already documented in the schema. The description covers the core purpose accurately, so it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters have detailed descriptions in the schema (e.g., 'Apply the conservative used-track heuristic.' and the note about preserving 0.11 response contract). The tool description adds no parameter-specific information, so it merely meets the baseline without enhancing clarity beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'inventory' and a clear resource 'loaded plug-ins', explicitly noting it does not mutate anything. This clearly distinguishes it from sibling tools like plugins_scan_parameters (which focuses on parameters) and plugins_load (which loads plugins). The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates this is for taking stock of loaded plugins, but it does not explicitly state when to prefer this over alternatives like plugins_scan_parameters or plugins_list_available. However, the read-only nature is implied and it is intuitive enough that an agent would use this to enumerate current plugins. No exclusions or conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plugins_scan_parametersA
Read-onlyIdempotent

De-pad a whole plug-in in one call; prefer this over paging a VST.

FL reports a padded maximum rather than a parameter count for VST plug-ins
-- often thousands of slots for a VST3 -- with the real controls sparse
inside it.
Paging that with `plugins_inspect_parameter_map` is about a thousand round
trips. This walks the range inside FL and returns only what is real, with
each control's display string, which is what actually identifies it.

Check `truncated` before treating the result as the whole plug-in.
ParametersJSON Schema
NameRequiredDescriptionDefault
endNoExclusive last index. Defaults to FL's reported count.
startNoFirst index to examine. Defaults to 0.
targetNoExplicit mixer_effect or global channel_generator target. Mutually exclusive with legacy track_index/slot_index.
slot_indexNoLegacy zero-based effect slot (0 through 9). Supply it with track_index, or use target, never both.
max_indicesNoStop after examining this many indices.
max_resultsNoStop after collecting this many real controls.
track_indexNoLegacy zero-based mixer track index. Supply it together with slot_index, or use target, never both.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, but the description adds valuable behavioral context: it explains the padding phenomenon, that the tool walks the range inside FL and returns only real controls with display strings, and that results may be truncated (via the `truncated` flag). This goes beyond the annotations and helps the agent understand edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, front-loading the main action and rationale. It uses three short paragraphs to explain the problem, the alternative, and the caution about truncation, with no wasted words. Each sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, output schema) and the rich schema and annotations, the description is complete. It explains the purpose, why it's needed, the alternative, and the critical `truncated` check. It doesn't repeat schema details but provides the contextual rationale an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter (target, start, end, max_indices, max_results, track_index, slot_index) having a clear description. The tool description doesn't add parameter-level detail, but since the schema already covers them, the baseline of 3 is appropriate. The description does not introduce ambiguity or omit necessary parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('De-pad') and resource ('a whole plug-in'), clearly distinguishing itself from the paging approach via plugins_inspect_parameter_map. It explains the problem (FL's padded maximum) and the solution (returns only real controls with display strings), making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs to 'prefer this over paging a VST' and names the alternative `plugins_inspect_parameter_map`, explaining the inefficiency of paging (thousands of round trips). It also warns to check `truncated` before treating the result as complete, providing clear guidance on when and how to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_continue_runA
Destructive

Continue or replace only a run's unexecuted remainder after a follow-up.

ParametersJSON Schema
NameRequiredDescriptionDefault
deltaYesUse mode=resume with no operations to continue the saved plan, append operations, or replace only the unexecuted remainder; an optional updated request may narrow scope or change task policy.
run_idYesProduction Run identifier retained in the local journal.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
summaryYes
blockersNo
receiptsNo
warningsNo
phase_planNo
run_contextNo
project_savedNo
timing_reportNo
attempted_countYes
completed_countYes
creation_outcomeNo
readiness_reportNo
total_operationsYes
generated_outputsNo
write_mode_activeNo
rollback_attemptedNo
session_fingerprintNo
project_state_digestNo
write_mode_enable_countNo
write_mode_disable_countNo
write_mode_shutdown_verifiedNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the mutation risk is known. The description adds the useful scoping detail that only the 'unexecuted remainder' is touched, implying completed portions are preserved, which is valuable behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single 13-word sentence with no filler. The core scoping constraint ('only... unexecuted remainder') is front-loaded, and every word contributes to meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with a rich output schema and descriptive parameter docs, the definition is minimally viable. However, it omits lifecycle context: whether the run must exist, what happens if it is already completed, and how this differs from postfader_execute_run. These gaps matter for a destructive tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies. The tool description adds no parameter information, but the schema's delta description meaningfully explains the mode=resume, append, and replace-remaining behaviors, so an agent can correctly populate both parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('continue or replace') and a precise resource ('a run's unexecuted remainder'), which is more informative than the bare title. It does not name sibling tools like postfader_execute_run or postfader_stop_run, so an agent must infer the differentiation from context rather than being told.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'After a follow-up' implies this tool is for resuming an existing run rather than starting one, but no alternatives or when-not-to-use conditions are stated. The delta parameter's schema description explains mode choices, but the tool description itself does not route the agent away from postfader_execute_run or other run lifecycle tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_creation_readinessB
Read-onlyIdempotent

Aggregate all detectable setup blockers without changing FL Studio.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYesClosed run plan whose complete setup needs are inspected.
requestYesTask-scoped creation objective and project constraints.

Output Schema

ParametersJSON Schema
NameRequiredDescription
scoreYes
blockersNo
warningsNo
dimensionsNo
limitationsNo
observed_atYes
overall_stateYes
manual_actionsNo
schema_versionNo
zero_mutationsNo
context_snapshotNo
mutations_performedNo
optional_enhancementsNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly states 'without changing FL Studio' which is a behavioral trait that matches the annotations (readOnlyHint=true, destructiveHint=false). The annotations already provide readOnly and non-destructive hints, so the description adds the concept of 'aggregate' which is not in annotations but is a small additional behavioral clarification. However, it doesn't disclose what happens if blockers are found or if it returns normally or throws. The description is consistent with annotations, so no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with exactly three word chunks: 'Aggregate all detectable setup blockers without changing FL Studio.' It is concise stubcount and front-loaded with the main verb and object database. It is appropriately sized for the complexity of the tool, and every word contributes to understanding the primary behavior and safety property. No extra words wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool with a large input schema and an output schema (which hedges on return details), the description is sufficient to understand the tool's core purpose. The annotations provide read-only and non-destructive hints, and the schema fully defines the inputs. The description adds the key element that it aggregates blockersaints without modifying state, which is complete for an agent to decide when to call it. The only missing piece is when to use it (e.g., before execution), but that is not essential for invocation

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema for the two parameters (request and plan) is fully described with descriptions in the schema, so the schema documentation coverage is 100%. However, the description of the tool adds no additional meaning beyond what the schema provides. The parameters are complex objects (ProductionRunRequest and ProductionRunPlan), and the schema has extensive definitions, so the agent can infer what they are. The description does not need to elaborate on the parameters since the schema covers them. Baseline is 3 for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Aggregate all detectable setup blockers without changing FL Studio' is specific in that it aggregates setup blockers and emphasizes it does not change anything dropped. However, the tool name includes 'creation_readiness', which is a narrow verb-resource pair, but the description does not mention the tool's inputs, only its output. It is also not differentiating enough from siblings like postfader_validate_run, which might also check blockers. The verb 'aggregate' is clear but the object 'setup blockers' is somewhat vague without knowing the context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage is to check for setup blockers without changing state, and is read-only, which is confirmed by annotations. It does not explicitly say when to use this tool vs alternatives like postfader_validate_run or postfader_execute_run. It also does not state that it should be called before execution, which would be a natural usage guideline. The description could be improved by saying 'use before postfader_execute_run to ensure no blockers'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_delivery_export_manifestA
Destructive

Create local delivery files without overwriting or saving the FL project.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesCreate-only JSON/Markdown delivery export options.

Output Schema

ParametersJSON Schema
NameRequiredDescription
digestYes
json_pathNo
delivery_idYes
json_sha256No
markdown_pathNo
project_savedNo
manifest_digestYes
markdown_sha256No
review_session_idYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description adds a meaningful guarantee: it creates files without overwriting or saving the FL project. This clarifies the destructiveHint and readOnlyHint signals by scoping them to local files rather than the project. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that conveys the core action and the key safety constraint. There is no filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for basic invocation, and the output schema plus annotations fill gaps. However, it does not explain what a review session is, what 'delivery files' contain, or how output_directory behaves, which an agent may need when choosing this tool among many postfader siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% because the request property is described as 'Create-only JSON/Markdown delivery export options.' The main description adds little beyond that, and nested property semantics (review_session_id, output_directory, formats) are left to the schema. Baseline 3 is appropriate given high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource ('Create local delivery files') and adds a distinguishing constraint ('without overwriting or saving the FL project'). It is reasonably distinct from render/save siblings, though it does not explicitly say 'delivery manifest' or reference the review session flow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage context: use this when you want local delivery files and do not want the FL project saved or overwritten. It does not name alternatives or give when-not-to-use exclusions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_delivery_manifestA
Read-onlyIdempotent

Build the final multi-dimensional delivery view without writing a file.

ParametersJSON Schema
NameRequiredDescriptionDefault
review_session_idYesReview Session whose current delivery view should be built.

Output Schema

ParametersJSON Schema
NameRequiredDescription
comparisonsNo
delivery_idYes
evaluationsNo
next_actionYes
final_run_idNo
generated_atNo
review_assetsNo
source_run_idYes
export_handoffNo
original_briefNo
manifest_digestNo
accepted_paletteNo
creation_outcomeNo
final_run_statusNo
playlist_handoffNo
accepted_sectionsNo
completion_targetNo
final_run_detailsNo
review_session_idYes
source_run_statusNo
technical_outcomeNo
pattern_placementsNo
processing_outcomeNo
source_run_detailsNo
arrangement_outcomeNo
final_revision_passNo
final_user_approvalNo
source_state_digestNo
final_revision_pass_idNo
unresolved_limitationsNo
audible_quality_outcomeNo
remaining_manual_actionsNo
accepted_role_assignmentsNo
accepted_generated_outputsNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description's 'without writing a file' adds a concrete, useful behavioral detail that reinforces the read-only nature in human terms, but it doesn't add richer context such as behavior with invalid/stale sessions or whether the view is regenerated fresh each call. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence containing zero filler. The verb comes first, the resource follows, and the distinguishing qualifier ('without writing a file') is placed at the end without bloating the sentence. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with 100% schema coverage, a rich annotation set, and an output schema, the description covers what matters: the purpose and the key non-persistence behavior. It stops just short of perfect because it does not explicitly name postfader_delivery_export_manifest as the file-writing counterpart, though the contrast is obvious from the wording and sibling list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter review_session_id is fully described ('Review Session whose current delivery view should be built'). The description's 'delivery view' vocabulary aligns with the parameter description, but it adds no parameter-specific semantics beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Build'), a specific resource ('the final multi-dimensional delivery view'), and adds the scope-limiting qualifier 'without writing a file,' which distinguishes it from the sibling postfader_delivery_export_manifest. An agent can tell at a glance what this tool produces and what it deliberately does not do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without writing a file' implies the usage boundary — use this when you need the delivery view in memory, not persisted — and the sibling name postfader_delivery_export_manifest makes the alternative inferable. However, the description never explicitly names the alternative or states when-to-use vs. when-not-to-use, leaving the routing to inference rather than direct guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_execute_runA
Destructive

Create and execute one task-scoped run until its plan completes or blocks.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYesClosed bounded plan to validate completely, then execute in order.
requestYesTask-scoped request. Mutating plans require authorized_to_modify=true because the present user explicitly asked to change the project.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
summaryYes
blockersNo
receiptsNo
warningsNo
phase_planNo
run_contextNo
project_savedNo
timing_reportNo
attempted_countYes
completed_countYes
creation_outcomeNo
readiness_reportNo
total_operationsYes
generated_outputsNo
write_mode_activeNo
rollback_attemptedNo
session_fingerprintNo
project_state_digestNo
write_mode_enable_countNo
write_mode_disable_countNo
write_mode_shutdown_verifiedNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the destructive/non-idempotent/read-only profile, so the description's job is to add context beyond them. It does: the run lifecycle ('until its plan completes or blocks'), the validate-completely-before-executing ordering, and the explicit mutation authorization gate. What it does not disclose is what 'blocks' concretely means, partial-execution behavior, or whether failures are reversible — but with destructiveHint=true already declared, this is solid supplementary disclosure rather than a gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One 13-word sentence that front-loads the action, names the resource, and states the termination condition. Every clause in the description and both parameter descriptions carries distinct information. No filler, no restatement of the title, and no bloat despite the enormous schema it accompanies.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's size and danger profile, the description is largely sufficient: it states the core lifecycle, the validate-then-execute ordering, and the auth precondition, while output schema and annotations cover the safety profile. It could be more complete by routing the agent toward postfader_validate_run for pre-flight checks or defining what conditions cause a run to block, but those are refinements, not essential omissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but both descriptions add genuine meaning beyond the bare refs. The plan description explains the consumption contract (validate completely, execute in order), and the request description explains the task-scoped nature and the authorized_to_modify=true requirement. These are not restatements of schema structure; they convey ordering and authorization semantics the schema alone cannot.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource ('Create and execute one task-scoped run') and adds a crisp termination condition ('until its plan completes or blocks'). This distinguishes it from sibling tools at a glance: validate_run only validates, continue_run resumes, stop_run stops, get_run/list_runs observe. No ambiguity about what this entry point does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description itself gives clear context for when this tool is the right choice (creating and executing a run end-to-end versus validating or continuing one). The parameter descriptions strengthen this: the plan description states it must be a 'Closed bounded plan to validate completely, then execute in order,' and the request description sets a critical precondition that mutating plans require authorized_to_modify=true. It stops short of explicitly naming alternatives or when-not-to-use, so it earns a 4 rather than a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_get_runA
Read-onlyIdempotent

Read a current or journaled run, its generated outputs and operation receipts.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesProduction Run identifier retained in the local journal.

Output Schema

ParametersJSON Schema
NameRequiredDescription
foundYes
stateNo
messageYes
process_localNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint, idempotentHint, and destructiveHint=false, and the description agrees with them by saying 'Read'. It adds useful behavioral context by defining the run states covered ('current or journaled') and what the call returns (outputs and operation receipts).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the action, scope, and return contents with no filler. Every word contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one fully documented parameter, an output schema, and annotation coverage for safety/idempotency, the description provides enough context to use the tool correctly. The mention of 'current or journaled' and the returned artifacts fills the remaining practical gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so run_id is already fully documented with a pattern and a meaningful description. The tool description adds no parameter-specific detail, but none is necessary given the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb, 'Read', and names the resource: a current or journaled run, along with its generated outputs and operation receipts. This clearly distinguishes it from run-list or run-execution siblings like postfader_list_runs and postfader_execute_run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'current or journaled run' gives clear context for when to call this tool: when you need the details, outputs, or receipts of a specific run. It does not explicitly name alternatives or exclusions, so it stops just short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_list_runsA
Read-onlyIdempotent

Find recent runs after an MCP restart without executing any operations.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum recent run summaries.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context: it specifically addresses the post-restart scenario and explicitly states that no operations are executed. This goes beyond what annotations provide by clarifying the tool's non-executing nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that front-loads the key purpose ('Find recent runs') and adds the critical context ('after an MCP restart', 'without executing any operations'). Every word earns its place; there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one optional parameter and a full output schema, the description is nearly complete. It covers the key scenario (post-restart) and the non-executing nature. The only minor gap is that it doesn't describe what 'recent runs' means in terms of ordering or time window, but the output schema likely covers the return structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single 'limit' parameter is fully documented in the schema. The description doesn't add parameter-specific details beyond what the schema provides, but it doesn't need to. The baseline of 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Find') and resource ('recent runs after an MCP restart'), which clearly identifies the tool's purpose. It doesn't explicitly differentiate from sibling tools like postfader_get_run or postfader_validate_run, but the 'after an MCP restart' context and 'without executing any operations' qualifier give it a distinct identity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: after an MCP restart, when you need to see recent runs without executing operations. It doesn't explicitly name alternatives or state when not to use it, but the context is reasonably clear. The sibling list includes postfader_get_run and postfader_validate_run, which could be alternatives, but the description doesn't mention them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_render_cancelB

Cancel monitoring and its owned process; on macOS the renderer may remain open.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesRender job ID from this MCP process.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
job_idYes
resultNo
statusYes
commandYes
platformYes
warningsNo
created_atYes
started_atNo
finished_atNo
output_pathYes
project_pathYes
launch_methodYes
process_localNo
cancel_requestedNo
output_directoryYes
process_return_codeNo
renderer_may_be_runningNo
includes_unsaved_changesNo
application_exit_observedNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate the operation is not read-only, not idempotent, and not destructive, which sets the basic safety profile. The description adds one useful behavioral note: on macOS the renderer may remain open. This is a specific caveat that helps manage expectations. However, it does not disclose other side effects, such as whether the job is marked canceled or if any cleanup occurs, leaving the agent with limited behavioral insight beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the main action and appends a platform-specific note. It is concise with no waste. However, it could be slightly clearer about the resource being canceled (render job vs monitoring), but structurally it is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists (not shown but indicated), the description does not need to explain return values. However, the description lacks clarity on what 'monitoring' refers to in the context of render jobs, and it does not mention the relationship to other postfader_render tools. The macOS caveat is useful, but overall the description leaves some ambiguity about the exact scope of the cancel operation, making it only adequately complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter job_id is fully described in the schema with 'Render job ID from this MCP process.' The tool description adds no additional meaning or constraints beyond the schema. Since schema coverage is 100%, a baseline of 3 is appropriate, and the description does not need to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Cancel') and resource ('monitoring and its owned process'), which strongly implies the intended action of canceling a render job. It distinguishes from siblings like postfader_stop_run by focusing on monitoring/process rather than runs. However, it does not explicitly say 'render job', relying on the annotation title for full clarity, which is a minor gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention that it is the cancel counterpart to postfader_render_saved_project or how it differs from postfader_stop_run. The agent is left to infer usage from the tool name and sibling list, which is insufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_render_get_jobB
Read-onlyIdempotent

Get render progress and decoded WAV evidence; completed also requires FL exit.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesRender job ID from this MCP process.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
job_idYes
resultNo
statusYes
commandYes
platformYes
warningsNo
created_atYes
started_atNo
finished_atNo
output_pathYes
project_pathYes
launch_methodYes
process_localNo
cancel_requestedNo
output_directoryYes
process_return_codeNo
renderer_may_be_runningNo
includes_unsaved_changesNo
application_exit_observedNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, idempotentHint, and destructiveHint=false, and the description does not contradict them. The 'completed also requires FL exit' phrase adds a behavioral nuance beyond those annotations, but it is ambiguous and does not clarify side effects or conditions. No annotation contradiction is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short two-clause sentence with no filler or repetition of the annotation data. The front-loaded 'Get render progress and decoded WAV evidence' is efficient, even though the trailing 'completed also requires FL exit' is syntactically opaque.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an inspect-by-id tool with one fully documented parameter and an available output schema, the description is mostly sufficient. However, it does not explain the saved-project render context, when to treat 'completed' as the terminal state, or what the 'FL exit' requirement means, so agents may still need to infer workflow timings from sibling tool names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The parameter schema already fully documents job_id, its type, pattern, and source ('Render job ID from this MCP process'). With 100% schema coverage, the description adds no additional meaning about how to obtain or format the job_id, which matches the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete action and resource: 'Get render progress and decoded WAV evidence.' It clearly identifies this as an inspection/polling tool rather than a render-starting or render-canceling tool, which distinguishes it from named siblings like postfader_render_saved_project and postfader_render_cancel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement about when to use this tool versus alternatives, nor any mention of polling cadence or the render-creation flow. The phrase 'completed also requires 2 exits' offers a thin timing clue, but fails to say whether to call this after postfadel_render_saved_project, before postfader_render_cancel, or how it relates to postfader_get_run.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_render_saved_projectA

Start FL's command-line WAV exporter in a separate process; saved state only.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesSaved .flp and parent output directory for a new WAV job.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
job_idYes
resultNo
statusYes
commandYes
platformYes
warningsNo
created_atYes
started_atNo
finished_atNo
output_pathYes
project_pathYes
launch_methodYes
process_localNo
cancel_requestedNo
output_directoryYes
process_return_codeNo
renderer_may_be_runningNo
includes_unsaved_changesNo
application_exit_observedNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is not read-only, but the description adds meaningful behavioral context: it launches a separate process, implying asynchronous execution, and uses only the saved project state. It does not cover details like file overwrite behavior or concurrent jobs, but there is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence front-loads the verb and resource, and the 'saved state only' qualifier earns its place as an important constraint. There is no redundant or filler phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and sibling tools handle job inspection and cancellation, the description covers the critical launch semantics: separate process and saved-state-only input. It does not explain timeout behavior or FL Studio path discovery, but those are represented in the schema defaults, so the remaining gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the nested request schema already documents project_path, output_directory, fl_studio_path, and timeout_seconds. The tool description adds only the high-level 'saved state only' constraint; the request schema already describes the inputs as a saved .flp and parent output directory, so the description contributes little beyond structured data.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: 'Start FL's command-line WAV exporter in a separate process.' The qualifier 'saved state only' plus the annotation title 'Render a saved FL Studio project' makes the tool's scope clear and distinguishes it from generic postfader execution or live-state operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'saved state only' implies the key precondition: the project must be saved before calling and unsaved changes will not be rendered. However, the description does not explicitly name alternatives or state when not to use this tool, so routing among sibling render/execute tools still depends mostly on the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_apply_revisionA
Destructive

Apply one bounded revision with one preflight and one write authorization.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesRecorded RevisionPlan and present task-scoped authorization.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
timingNo
blockersNo
warningsNo
started_atNo
finished_atNo
project_savedNo
source_run_idYes
affected_rolesNo
manual_handoffsNo
preflight_resultNo
retained_anchorsNo
revision_pass_idYes
revision_plan_idYes
affected_sectionsNo
generated_outputsNo
review_session_idNo
shutdown_verifiedNo
technical_outcomeNo
after_bounce_stateNo
operation_receiptsNo
processing_outcomeNo
rollback_attemptedNo
arrangement_outcomeNo
authorization_countNo
continuation_run_idNo
session_fingerprintNo
project_state_digestNo
source_evaluation_idYes
audible_quality_outcomeNo
write_mode_enable_countNo
source_production_run_idNo
write_mode_disable_countNo
readiness_preflight_countNo
automatic_replay_attemptedNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds specific operational behavior: it performs a preflight and requires one write authorization before modifying. This is genuinely useful behavioral context beyond the annotations, though it does not describe what gets destroyed or failure conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence with no filler. 'Apply,' 'bounded revision,' 'one preflight,' and 'one write authorization' each earn their place by communicating the core action and constraints efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested ReviewApplyRevisionRequest schema and destructive annotations, a one-sentence description is functional but thin. The output schema covers return details, and the annotations cover safety, but the description does not mention the need to coordinate with postfader_review_plan_revision or what happens when authorization is missing. It is minimally sufficient, not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter has its own description, 'Recorded RevisionPlan and present task-scoped authorization.' The prose adds little per-field meaning, but at 100% coverage the schema carries the burden, so the description is adequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Apply' and identifies the resource as a bounded revision, which clearly states the action. While it does not explicitly name the 'Creation Review' domain in the description or mention a particular sibling, 'Apply' clearly differentiates it from planning, evaluating, and comparing revisions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The wording 'with one preflight and one write authorization' implies the prerequisite that an authorization and a prior plan are needed, but there is no explicit guidance about when to choose this tool over postfader_review_plan_revision or other revision tools. The usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_attach_assetsA
Read-onlyIdempotent

Validate and attach explicit audio assets without changing FL Studio.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesExplicit caller-selected full mix, reference, stem, or section paths.

Output Schema

ParametersJSON Schema
NameRequiredDescription
assetsNo
statusNo
requestYes
blockersNo
feedbackNo
warningsNo
asset_setsNo
created_atNo
updated_atNo
comparisonsNo
evaluationsNo
section_mapNo
process_localNo
source_run_idYes
revision_plansNo
schema_versionNo
revision_passesNo
source_sectionsNo
source_snapshotNo
review_session_idYes
delivery_manifestsNo
current_next_actionNo
source_pattern_planNo
source_sound_paletteNo
source_note_sequencesNo
source_manual_handoffsNo
source_creation_outcomeNo
source_processing_receiptsNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description reinforces this with 'without changing FL Studio'. It also adds the validation behavior, which is not purely derivable from annotations. It does not describe failure modes or side effects on non-FL Studio state, but the annotations cover the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tightly structured sentence that front-loads the verb and key constraints. Every word adds value, and the key non-destructive scope is stated directly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the supporting annotations and a rich schema, the description is sufficient for understanding the tool's role in the review workflow. It lacks explicit usage guidance and details about validation criteria, but the schema and sibling context fill most gaps. The core behavior and non-destructive guarantee are clearly communicated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% at the top level, and the 'request' parameter has a description that clarifies asset categories. The nested object fields have titles and an enum but no descriptive text, and the tool description does not add parameter-level details. Baseline 3 is appropriate, as the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Validate and attach') and a specific resource ('explicit audio assets'), and adds a key scope constraint ('without changing FL Studio'). It helps distinguish from sibling tools that render, evaluate, or start reviews, though it does not explicitly name the 'creation review' context that appears in the title annotation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives. It implies use when assets need to be attached to a review, but lacks exclusions, alternative tool names, or workflow placement. This leaves the agent to infer usage from the name and siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_compareA
Read-onlyIdempotent

Compare before and after bounces without implying producer approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesDistinct aligned before/after assets and their revision objective.

Output Schema

ParametersJSON Schema
NameRequiredDescription
timingNo
warningsNo
after_assetYes
regressionsNo
stem_deltasNo
before_assetYes
improvementsNo
comparison_idYes
global_deltasNo
section_deltasNo
zero_mutationsNo
unknown_metricsNo
alignment_resultYes
mutations_appliedNo
unchanged_metricsNo
next_recommendationNo
user_approval_stateNo
technical_conclusionNo
expected_objective_resultsNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, destructiveHint, and idempotentHint, covering safety. The description adds a meaningful behavioral note—that invoking this comparison does not signal producer approval—which is not captured by the annotations and is valuable for correct usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that leads with the action and immediately states the key constraint. No unnecessary words; efficient and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety, an output schema, and a self-explanatory nested schema, the description is sufficient for an agent to invoke correctly. The only minor gap is not explicitly routing away from sibling tools, but the name and the approval nuance provide enough context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% coverage for the 'request' parameter with its description 'Distinct aligned before/after assets and their revision objective,' and the nested properties are self-explanatory. The description reinforces the alignment and distinctness of assets, adding slight nuance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Compare') and resource ('before and after bounces') and adds a distinctive behavioral nuance ('without implying producer approval'). This clearly distinguishes it from sibling tools like postfader_review_evaluate or audio_compare_files, which have different scopes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear functional constraint: this tool is for comparison and must not be used to imply approval. This is a form of when-to-use guidance. However, it does not explicitly name alternatives or state when not to use it, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_deleteA
Destructive

Delete one Review Session record without touching audio or the FL project.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmYesMust be true after an explicit request to delete review metadata.
review_session_idYesReview Session metadata to delete.

Output Schema

ParametersJSON Schema
NameRequiredDescription
deletedYes
messageYes
process_localNo
review_session_idYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and idempotentHint=false, and the description adds the key behavioral reassurance that audio and the FL project are untouched. It clarifies exactly what will be destroyed: review session metadata only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence conveys the action, resource, and destructive scope with zero wasted words. Important safety context (audio and project untouched) is included without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter delete operation with a confirm flag plus an output schema and destructive annotations, the description is nearly complete. It could have additionally stated that deletion is irreversible, but annotations and the confirm parameter already signal the destructive nature.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the parameter-level descriptions already explain review_session_id as the metadata to delete and confirm as a required explicit-deletion flag. The description adds no extra parameter meaning beyond this, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ("Delete one Review Session record") with a clear resource and an explicit scope boundary ("without touching audio or the FL project"). This distinguishes it from broader destructive audio/project tools and from review tools that do more than delete metadata.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the use case clear: the agent should use this when deleting a single review session record while preserving audio and project state. It does not explicitly name sibling alternatives, but the metadata-only scope is enough to route an agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_evaluateA
Read-onlyIdempotent

Measure one bounce globally and by known section; apply zero FL mutations.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesAttached asset set and optional authoritative section ranges.

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusNo
timingNo
findingsNo
warningsNo
evaluated_atNo
evaluation_idYes
source_run_idYes
top_prioritiesNo
zero_mutationsNo
analyzer_versionYes
asset_set_digestYes
goal_evaluationsNo
masking_analysisNo
mutations_appliedNo
review_session_idYes
stem_measurementsNo
section_map_digestYes
global_measurementsNo
unavailable_analysesNo
audible_quality_stateNo
reference_comparisonsNo
technical_audio_stateNo
analysis_policy_digestNo
arrangement_proxy_stateNo
processing_review_stateNo
energy_contrast_analysisNo
per_section_measurementsNo
generated_content_analysisNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The explicit 'apply zero FL mutations' goes beyond the annotations by clarifying that FL Studio state will not be touched, reinforcing the readOnlyHint=true, idempotentHint=true, and destructiveHint=false annotations. The 'by known section' phrase also discloses that the tool operates on caller-supplied section boundaries, which is useful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that states the operation, scope, and side-effect guarantee with zero filler. It front-loads the action ('Measure') and ends with the non-mutation note, making it easy to scan and parse for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and parameter semantics are fully documented, the description covers the essential behavioral contract: measure globally and by section without touching FL. It does not need to spell out return values or request construction, as the output schema and the required review_session_id field cover those. The only minor gap is not explicitly naming the review_session_id, but that is present in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the nested schemas already document the request fields, section ranges, and reference windows in detail. The description only adds that sections are 'known' (user-supplied), which maps to section_ranges but does not explain individual parameters. The schema does the heavy lifting, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Measure') and a specific resource ('one bounce') with scope ('globally and by known section'), which distinguishes it clearly from sibling tools like postfader_review_compare, postfader_review_get, and postfader_review_start. It is not a tautology and gives the agent a precise mental model of the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when a read-only measurement of a bounce is needed, over the whole file or per supplied section. It does not explicitly name an alternative or state when not to use it, but the 'apply zero FL mutations' line plus the sibling names (compare, get, start) make the intended use boundary reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_export_handoffB
Read-onlyIdempotent

Return one precise full-mix export request and only necessary stems.

ParametersJSON Schema
NameRequiredDescriptionDefault
review_session_idYesReview Session awaiting its next caller-exported bounce.

Output Schema

ParametersJSON Schema
NameRequiredDescription
handoff_idYes
next_actionYes
stem_reasonsNo
exact_end_barNo
include_tailsNo
exact_start_barNo
requested_stemsNo
exact_end_secondsNo
expected_locationNo
normalization_offNo
exact_start_secondsNo
recommended_filenameYes
bounded_discovery_rootNo
required_full_mix_exportNo
matching_settings_requiredNo
before_after_naming_conventionNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, idempotentHint, and destructiveHint. The description adds behavioral context by emphasizing 'precise' and 'only necessary stems,' suggesting a selective, non-mutating action. It does not contradict annotations and provides some extra detail beyond what annotations convey, though it does not explain the full behavior (e.g., whether it touches the review session or just generates a request).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that communicates the core purpose without excess. It is front-loaded with the key output ('full-mix export request') and the qualifier ('only necessary stems'). It is concise and effective, though it could benefit from a brief clause on when to use it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (1 parameter, output schema present, annotations cover safety), the description is mostly adequate. However, it does not explain what a 'handoff' means in the review workflow or how this tool fits with siblings like postfader_review_start or postfader_review_evaluate. Some workflow context is missing, but the presence of an output schema and safe annotations partially compensates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage for the single parameter 'review_session_id,' with a description explaining it awaits a bounce. The tool description does not add any extra parameter-specific meaning beyond this, so with high schema coverage the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'return' and identifies a precise output: 'one precise full-mix export request and only necessary stems.' This clearly states what the tool does and distinguishes it from sibling export-related tools like postfader_delivery_export_manifest by focusing on review handoff. However, it could be more specific about what a 'full-mix export request' entails (e.g., format or destination), so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives, such as postfader_review_get or postfader_review_apply_revision. There is no mention of prerequisites, workflow context, or exclusions. The name hints at usage, but the description itself does not provide explicit usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_getA
Read-onlyIdempotent

Read a Review Session, retained evidence, status, and exact next action.

ParametersJSON Schema
NameRequiredDescriptionDefault
review_session_idYesProcess-local or persisted Review Session identifier.

Output Schema

ParametersJSON Schema
NameRequiredDescription
foundYes
messageYes
sessionNo
process_localNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is established. The description adds value by specifying exactly what data is retrieved (retained evidence, status, exact next action), which goes beyond the annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero redundancy. The verb and core deliverables are stated immediately, making it highly efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one parameter and an existing output schema, the description covers the essential information: what is read and what is returned. It does not mention error cases or prerequisites, but those are likely handled by the schema and output structure. Minor gap: it could state that the session must already exist, but this is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole parameter review_session_id is already well-documented as 'Process-local or persisted Review Session identifier.' The description adds no further parameter information, so it meets the baseline 3 without exceeding it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read') and resource ('Review Session') and lists the returned content (evidence, status, next action). It clearly distinguishes from sibling review tools that start, evaluate, or modify sessions, though it does not explicitly name any alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. The description does not mention prerequisites, alternatives, or exclusions. It is only implied that this tool is for retrieving review session data, but an agent gets no direction on when to pick this over similar getters or state-changing operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_plan_revisionA
Read-onlyIdempotent

Compile and validate one bounded RevisionPlan before any project mutation.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesStrict revision request plus a closed traceable operation list.

Output Schema

ParametersJSON Schema
NameRequiredDescription
blockersNo
warningsNo
operationsNo
plan_digestNo
source_run_idYes
manual_actionsNo
revision_plan_idYes
mutations_appliedNo
review_session_idYes
targeted_findingsNo
protected_elementsNo
expected_objectivesNo
source_evaluation_idYes
subjective_objectivesNo
revision_request_digestNo
expected_measurable_movementsNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds value by stating the tool is a non-mutating pre-step and that the plan is 'bounded,' which is behavioral context beyond the annotations. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is tightly packed and front-loaded: verb, resource, scope, and timing all appear in order with zero filler. Every word contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a very complex schema, the description supplies the essential workflow context: this is the planning/validation step before mutation. Annotations cover safety and idempotency, and an output schema exists, so return-value detail is unnecessary. It could clarify what 'bounded' means or specify validation failure behavior, but the core call decision is fully supported.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter with 100% schema description coverage ('Strict revision request plus a closed traceable operation list'). The tool description adds no parameter-level meaning, so the schema carries the full burden. A 3 matches the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Compile and validate') with a clear resource ('one bounded RevisionPlan') and a decisive scoping phrase ('before any project mutation'). This distinguishes it from mutation tools like postfader_review_apply_revision without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'before any project mutation' gives clear temporal context for when to call this tool versus an apply/execute tool. It does not explicitly name alternatives or state when-not-to-use, but the placement cue is strong enough to route an agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_record_feedbackA

Record explicit feedback; silence and measurements never grant approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
feedbackYesExplicit structured producer feedback and independent locks.

Output Schema

ParametersJSON Schema
NameRequiredDescription
assetsNo
statusNo
requestYes
blockersNo
feedbackNo
warningsNo
asset_setsNo
created_atNo
updated_atNo
comparisonsNo
evaluationsNo
section_mapNo
process_localNo
source_run_idYes
revision_plansNo
schema_versionNo
revision_passesNo
source_sectionsNo
source_snapshotNo
review_session_idYes
delivery_manifestsNo
current_next_actionNo
source_pattern_planNo
source_sound_paletteNo
source_note_sequencesNo
source_manual_handoffsNo
source_creation_outcomeNo
source_processing_receiptsNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds a meaningful behavioral rule beyond the annotations: 'silence and measurements never grant approval.' This clarifies that the tool won't infer approval from indirect signals and that explicit verdicts are the only approval source. Combined with readOnlyHint=false, the write nature is understood, though no details about side effects or locking are given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense sentence that front-loads the core purpose and an essential constraint. Every word earns its place; there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a large and complex schema, the description is sparse but covers the most critical behavioral rule. It doesn't explain what the recorded feedback is used for or how it fits into the postfader review workflow, but the schema and sibling tool names supply much of that context. Adequate but not rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%—the single 'feedback' parameter and its nested types are fully described in the schema, including the phrase 'Explicit structured producer feedback and independent locks.' The tool description adds only the word 'explicit', which aligns with the schema but doesn't provide new semantic meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('Record') and resource ('explicit feedback'), which conveys the primary behavior. It includes a distinguishing constraint—'silence and measurements never grant approval'—but relies on the tool name to identify the 'creation review' scope, so it doesn't strongly differentiate from sibling tools like sound_selection_record_feedback.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly gives usage guidance: only explicit feedback should be recordedapsing, and silence/measurements should not be used as a basis for approval. This tells the agent when not to use the tool, though it doesn't explicitly name alternatives or describe the exact workflow context in which this should be invoked.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_startA
Read-onlyIdempotent

Start a bounded Review Session from one completed Production Run.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesTask-scoped review policy linked to a completed Production Run.

Output Schema

ParametersJSON Schema
NameRequiredDescription
assetsNo
statusNo
requestYes
blockersNo
feedbackNo
warningsNo
asset_setsNo
created_atNo
updated_atNo
comparisonsNo
evaluationsNo
section_mapNo
process_localNo
source_run_idYes
revision_plansNo
schema_versionNo
revision_passesNo
source_sectionsNo
source_snapshotNo
review_session_idYes
delivery_manifestsNo
current_next_actionNo
source_pattern_planNo
source_sound_paletteNo
source_note_sequencesNo
source_manual_handoffsNo
source_creation_outcomeNo
source_processing_receiptsNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, which already convey much of the safety profile. The description adds the bounded nature of the session and the requirement of a completed run, complementing the annotations. However, it doesn't explicitly clarify that this could still initiate analysis or planning work (despite readOnlyHint), nor does it discuss session persistence behavior; the schema covers defaults like persist_session but real-world effects like side effects on the run are not described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose and a key constraint. It avoids fluff and is appropriately sized for the complex schema provided separately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a very detailed schema and annotations, the description is functional but leaves a gap: it doesn't mention the interaction_policy, evaluation_policy, or user_feedback that the schema supports, nor does it clarify that this is the 'start' counterpart to postfader_review_stop. The output schema is not shown here, but the schema's richness compensates for much of the description's brevity. Still, the agent would benefit from knowing this tool is also the one that accepts user feedback for a bounded review.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Parameter coverage is 100% and the schema contains rich descriptions on nested objects (e.g., ReviewSessionRequest, CreationFeedback, ReviewPreserveRules). The description of the 'request' property matches the top-level description, so the text adds little new semantic meaning beyond the schema. The baseline 3 applies here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Start a bounded Review Session from one completed Production Run.' uses a specific verb ('Start'), a distinct resource ('bounded Review Session'), and a clear precondition ('completed Production Run'). It distinguishes this tool from siblings like postfader_continue_run and postfader_review_get by emphasizing session creation from a run, though it doesn't explicitly name those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description and the ReviewSessionRequest schema strongly imply this is the entry point for review sessions, especially given the required source_run_id and the annotation title 'Start a Creation Review.' While it doesn't explicitly state when not to use it or name alternatives, the context of sibling tools like postfader_continue_run and postfader_review_get makes the usage context reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_review_stopA

Stop future review work without undoing completed project changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
review_session_idYesReview Session whose future work should stop.

Output Schema

ParametersJSON Schema
NameRequiredDescription
assetsNo
statusNo
requestYes
blockersNo
feedbackNo
warningsNo
asset_setsNo
created_atNo
updated_atNo
comparisonsNo
evaluationsNo
section_mapNo
process_localNo
source_run_idYes
revision_plansNo
schema_versionNo
revision_passesNo
source_sectionsNo
source_snapshotNo
review_session_idYes
delivery_manifestsNo
current_next_actionNo
source_pattern_planNo
source_sound_paletteNo
source_note_sequencesNo
source_manual_handoffsNo
source_creation_outcomeNo
source_processing_receiptsNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description meaningfully augments the annotations by explaining that the operation is non-destructive to completed work and only halts future review activity. This is consistent with destructiveHint=false and readOnlyHint=false, adding useful behavioral nuance without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is one well-crafted sentence that front-loads the key behavior and immediately states the crucial non-destructive qualifier. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation tool with annotations and an output schema, the description is sufficient for invocation. It could be slightly more complete by contrasting with related siblings such as postfader_review_delete or postfader_stop_run, but that gap is minor given the clear action and parameter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter is fully documented in the schema with a clear description ('Review Session whose future work should stop.'). Since schema coverage is 100%, the tool description does not need to add parameter detail; it adds no new meaning beyond echoing the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('stop') and resource ('future review work'), and explicitly clarifies that completed project changes are preserved. This clearly differentiates it from potentially destructive siblings like postfader_review_delete or undo-like tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use this tool: when review work should stop but prior changes must remain intact. It does not explicitly name alternative tools or state exclusion criteria, but the context is strong enough for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_stop_runA

Stop future run operations without undoing completed project changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesProduction Run identifier retained in the local journal.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
summaryYes
blockersNo
receiptsNo
warningsNo
phase_planNo
run_contextNo
project_savedNo
timing_reportNo
attempted_countYes
completed_countYes
creation_outcomeNo
readiness_reportNo
total_operationsYes
generated_outputsNo
write_mode_activeNo
rollback_attemptedNo
session_fingerprintNo
project_state_digestNo
write_mode_enable_countNo
write_mode_disable_countNo
write_mode_shutdown_verifiedNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=false, destructiveHint=false, and idempotentHint=false, lowering the burden on the description. The description adds useful behavioral context by stating that completed project changes are not rolled back. However, it does not clarify whether an in-flight run is cancelled or whether the stop can be reversed by a later continue, leaving some behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the core action and the crucial negative scope. There is no filler, redundancy, or unnecessary detail, so every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter mutation with annotations, a complete schema, and an output schema present, the description is largely sufficient. The only minor gap is not stating whether a currently running operation is affected, but sibling tool names (postfader_continue_run, postfader_execute_run) partially fill this context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with run_id documented as 'Production Run identifier retained in the local journal' and a pattern constraint. The tool description adds no parameter-specific details, so the schema fully carries the semantic load; the baseline score applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a precise verb ('Stop') and a specific resource ('future run operations'), and immediately adds a key differentiator: 'without undoing completed project changes.' This clearly distinguishes it from sibling tools like fl_undo, postfader_execute_run, and postfader_continue_run even without inspecting their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes clear context: use this when you want to halt future run activity while preserving completed work. It does not explicitly name alternatives or exclusion criteria, but the 'without undoing' clause implicitly contrasts with rollback operations, providing sufficient situational guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

postfader_validate_runA
Read-onlyIdempotent

Validate a bounded Production Run and current capabilities without changing FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYesClosed ordered Production Run plan to validate without mutation.
requestYesTask-scoped objective, scope, preservation rules, allowed changes, completion target, and authorization inferred from the user's request.

Output Schema

ParametersJSON Schema
NameRequiredDescription
validYes
blockersNo
warningsNo
executableYes
plan_digestYes
validated_atYes
schema_versionNo
session_fingerprintNo
project_state_digestNo
required_capabilitiesYes
unsupported_operationsNo
resolved_operation_orderYes
expected_mutation_categoriesYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true. The description adds the specific guarantee 'without changing FL', reinforcing the non-destructive nature. It also mentions validating 'current capabilities', which is a useful behavioral detail not in the annotations. This adds value beyond what the annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no redundancy. It front-loads the key action ('Validate') and the critical constraint ('without changing FL'). Every word earns its place, and the structure is immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool takes two complex nested objects and has an output schema, so the return format is covered elsewhere. The description clearly states the non-mutating nature and the validation scope. It could mention what happens on validation failure (e.g., whether it returns blockers), but given the output schema, the description is adequate for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (request and plan) are well documented in the input schema. The description itself adds no additional parameter-level detail beyond what's in the schema. With high schema coverage, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Validate'), a specific resource ('a bounded Production Run and current capabilities'), and explicitly notes the non-mutating nature ('without changing FL'). This clearly distinguishes it from sibling tools like postfader_execute_run or postfader_continue_run, which are execution-oriented. An agent can immediately understand the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies a pre-execution validation step by emphasizing 'without changing FL' and 'current capabilities'. While it doesn't explicitly name alternatives like postfader_execute_run, the context strongly suggests this is a read-only pre-flight check. It could be improved by explicitly stating 'use before postfader_execute_run', but the intent is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

processing_apply_planA
Destructive

Apply a semantic plan through one task-scoped verified Production Run.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYesBounded semantic plan returned by processing_plan.
session_fingerprintYesRequired bridge/project-session fingerprint from a recent live read. The palette application refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.
authorized_to_modifyYesTrue only when the current user explicitly authorized these processing changes.

Output Schema

ParametersJSON Schema
NameRequiredDescription
run_idYes
statusYes
summaryYes
blockersNo
receiptsNo
warningsNo
phase_planNo
run_contextNo
project_savedNo
timing_reportNo
attempted_countYes
completed_countYes
creation_outcomeNo
readiness_reportNo
total_operationsYes
generated_outputsNo
write_mode_activeNo
rollback_attemptedNo
session_fingerprintNo
project_state_digestNo
write_mode_enable_countNo
write_mode_disable_countNo
write_mode_shutdown_verifiedNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the agent knows this mutates state. The description adds the crucial 'verified' and 'task-scoped' context, and the schema's session_fingerprint parameter description elaborates on the concurrency-guard/refusal behavior after bridge reload. The description itself is brief, but combined with the rich parameter descriptions (which are part of the schema, not annotations), the behavioral profile is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core purpose (apply a plan) and its key constraint (task-scoped, verified, one Production Run). It's appropriately sized for the tool's complexity. No waste, though it could have added a sentence on safety consequences of the destructive action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (3 required params, rich nested schema, output schema present, annotations covering destructive behavior), the description plus schema documentation fully covers what an agent needs: what to pass (plan from processing_plan, valid fingerprint, authorization flag), what happens semantically (verified run), and what safety class it falls into (destructive, not idempotent). The output schema exists, so return values need no description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining that the plan must come from processing_plan (in the plan parameter description) and by detailing the session_fingerprint's role as a concurrency guard that 'refuses after bridge reload or a reported project load.' It also clarifies that authorized_to_modify reflects explicit user authorization. This is meaningful semantic context beyond the raw field definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Apply') and resource ('a semantic plan through one task-scoped verified Production Run'). It clearly identifies the tool's function and differentiates it from sibling tools like processing_plan (which creates the plan) and mix_apply_plan (a different apply-plan tool). The annotation title confirms the purpose without replacing it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly establishes usage: you must already have a plan (from processing_plan, as referenced in the schema's plan parameter description). The 'verified Production Run' phrase signals this is the execution step after planning. It doesn't explicitly name alternatives or exclusions, but the workflow context is clear enough for an agent to select it over planning/read-only tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

processing_planA
Read-onlyIdempotent

Plan loaded-effect processing without enabling writes or mutating FL.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesRestrained processing goals resolved only against effects that are loaded, Atlas-matched, adapter-backed, and controllable.

Output Schema

ParametersJSON Schema
NameRequiredDescription
actionsNo
plan_idYes
warningsNo
candidatesNo
created_atNo
request_idYes
max_actionsNo
schema_versionNo
completion_targetYes
session_fingerprintNo
missing_capabilitiesNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety behavior is covered structurally. The description adds the 'loaded-effect' scope, which is useful, but it largely restates the non-mutating property already present in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence states the operation, the resource scope, and the safety constraint with no filler. Every word contributes to an agent's ability to understand and invoke the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a planning tool with only one parameter, a rich nested schema, and strong annotation coverage, the description covers the essential context. It could more explicitly route the agent to processing_apply_plan for the execution phase, but nothing needed to call it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the request parameter already has a detailed description: 'Restrained processing goals resolved only against effects that are loaded, Atlas-matched, adapter-backed, and controllable.' The tool description adds no parameter-level meaning beyond this, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Plan'), a specific resource ('loaded-effect processing'), and a clear constraint ('without enabling writes or mutating FL'). This distinguishes it from execution-oriented siblings like processing_apply_plan and establishes exactly what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without enabling writes or mutating FL' clearly signals this is the safe planning step and not the execution step. It does not name a specific alternative tool, such as processing_apply_plan, leaving the when-not-to-use guidance implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_applyB
Destructive

Apply exact presets in deterministic order and stop on unknown or unverified outcomes.

ParametersJSON Schema
NameRequiredDescriptionDefault
paletteYesA validated palette plan, section variation, or its process-local palette ID.
role_idsNoOptional bounded subset of palette roles.
persist_historyNoOverride this palette's task-scoped history policy.
settle_tick_limitNo
session_fingerprintYesRequired bridge/project-session fingerprint from a recent live read. The palette application refuses after bridge reload or a reported project load. This is a concurrency guard, not authentication or a durable project identity.
authorized_to_modifyYesTrue only when the current user explicitly authorized these project changes.
max_navigation_stepsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
stateYes
statusYes
blockersNo
receiptsNo
warningsNo
palette_idYes
schema_versionNo
verified_countNo
history_writtenNo
assignment_scopeNo
assignment_receiptsNo
session_fingerprintYes
failed_assignment_idNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already mark the tool as destructive and non-idempotent. The description adds value beyond that by guaranteeing deterministic order and a hard stop on unknown/unverified outcomes, which is useful context. It does not disclose what the stop means (e.g., no partial application), or authorization or rollback behavior, but it adds some meaningful context beyond the raw annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that immediately states the core action and the critical safety behavior. There is no fluff or redundancy, and the key caveat ('stop on unknown or unverified') is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with a required authorization flag, the description fails to convey that the tool should be invoked only after a planning step and with an explicit user authorization. It does not mention that the palette argument is a plan, nor the session-fingerprint concurrency requirement. The presence of an output schema reduces the need for return-value documentation, but the operational preconditions are left entirely to the structural schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 71%, the description adds no parameter-level semantics. It does not explain settle_tick_limit or max_navigation_steps, and it doesn't reinforce that the palette must be a validated plan or that session fingerprint is a concurrency guard rather than a durable identity. The schema provides some descriptions, but the description contributes no added value there.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action: 'apply exact presets in deterministic order' and adds a safety condition 'stop on unknown or unverified outcomes'. It clearly differentiates this from planning (e.g., sound_selection_plan) and feedback recording (sound_selection_record_feedback) by its execution focus, though it doesn't explicitly say 'palette'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance about when this tool should be used in a workflow, or when to prefer sibling tools like sound_selection_plan or sound_selection_get. It does not mention the need for a recent session fingerprint or explicit user authorization, which would help an agent know it is safe to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_create_variationC
Read-onlyIdempotent

Return a section delta instead of replacing the existing palette.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesSection-specific direction; anchors remain preserved by default.
sectionNoSection receiving the delta.
palette_idYes
replace_rolesNoRoles explicitly allowed to replace.

Output Schema

ParametersJSON Schema
NameRequiredDescription
sectionYes
blockersNo
warningsNo
conflictsNo
rationaleNo
assignmentsNo
plan_digestNo
variation_idYes
request_digestYes
schema_versionNo
base_palette_idYes
unchanged_role_idsNo
preserve_anchor_rolesNo
preset_discovery_coverageNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description's promise to 'return' a delta rather than modify the palette is consistent with that safety profile — no contradiction. The description adds the delta-vs-replace distinction beyond the annotations, but with the annotations already covering safety, the incremental behavioral context is modest; the description doesn't clarify how the returned delta is meant to be consumed or whether it persists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no wasted words, and it front-loads the key distinction. But it is terse to the point of under-specification: for a tool with a deeply nested request schema, one line of prose leaves significant context unstated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity — four parameters including a heavily nested SoundSelectionRequest with many sub-objects and defaults — a one-sentence description is insufficient. The output schema covers return values, but the description never explains what a 'variation' means in the workflow, how it relates to sound_selection_plan/sound_selection_apply, or why an agent would choose this over its siblings. An agent cannot confidently decide when this tool fits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with the schema already documenting the three meaningful parameters: request ('Section-specific direction; anchors remain preserved by default'), section ('Section receiving the delta'), and replace_roles ('Roles explicitly allowed to replace'). The description itself contributes essentially nothing about parameters, so it correctly relies on the schema to carry that burden — baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Return a section delta instead of replacing the existing palette' states a specific behavior and contrasts it against an implicit alternative (replacing the palette), which gives the tool some identity distinct from palette-replacing tools like sound_selection_apply. However, the phrase 'section delta' is jargon that isn't explained, and the description never mentions that the annotations' title identifies this as a planning tool ('Plan a Sound Palette variation'), leaving the actual function somewhat ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when you want a delta rather than a full replacement, but it never names a sibling tool or gives explicit when-to-use / when-not-to-use conditions. An agent comparing this against sound_selection_plan or sound_selection_apply gets no direct guidance on how to route between them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_getA
Read-onlyIdempotent

Look up one process-local palette without treating expiry as a server error.

ParametersJSON Schema
NameRequiredDescriptionDefault
palette_idYesProcess-local palette identifier.

Output Schema

ParametersJSON Schema
NameRequiredDescription
foundYes
stateNo
messageYes
process_localNo
schema_versionNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint false). The description adds genuine behavioral nuance beyond annotations: the 'without treating expiry as a server error' clause tells the agent how to handle expired palettes, which is valuable operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single 12-word sentence that front-loads the action and resource, then adds one critical behavioral clause. No filler, every phrase contributes to the agent's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter getter with a rich annotation set and an output schema, the description is complete. It conveys purpose, scope, and a subtle error-handling behavior, leaving little for an agent to guess. Return details are properly deferred to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — palette_id is already documented as 'Process-local palette identifier.' The description repeats the term 'process-local' but adds no additional parameter-level meaning, so it stays at the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('look up') and resource ('one process-local palette'), making the tool's core function immediately clear. The word 'one' distinguishes it from list-style siblings like sound_selection_inventory, and 'process-local' scopes it away from any global palette tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the usage context — retrieve a single palette by ID without surfacing expiry as an error — but never explicitly names alternatives or states when-not-to-use. With a large sibling family, some explicit routing guidance would help, but the basic use case is still inferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_history_resetA
DestructiveIdempotent

Explicitly remove bounded local selection history; project state is unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmYesMust be true after the user explicitly requested local history deletion.

Output Schema

ParametersJSON Schema
NameRequiredDescription
pathYes
existedYes
removedYes
warningsNo
recoverableNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true. The description adds clarity by stating that project state is unchanged, which is important given the destructive annotation. It also specifies removal of 'bounded local selection history,' providing context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the action and explicitly states the non-destructive nature to project state. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and the parameter is well-defined, the description is complete enough for an agent to call this tool correctly. It clearly states the effect and the requirement for confirmation. Minor gap: it does not mention what happens to the history after removal, but that is not critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the parameter 'confirm' is already well described. The description does not add additional detail about the parameter, but the schema is sufficient, so a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Explicitly remove bounded local selection history; project state is unchanged.' states a specific action (remove local selection history) and clarifies it does not affect project state, distinguishing it from potentially destructive operations. It is clear and precise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies that this is a cleanup operation for local history, and the parameter description reinforces the requirement for explicit user request. However, it does not explicitly state when to prefer this over alternatives like fl_undo or history-related tools, but the context makes it clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_history_statusA
Read-onlyIdempotent

Report the local history path, health, schema, and bounded record counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
pathYes
errorNo
existsYes
corruptYes
healthyYes
warningsNo
max_recordsYes
max_feedbackYes
record_countNo
feedback_countNo
schema_versionNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover non-destructive, idempotent, read-only behavior. The description adds genuine context beyond annotations: the tool reads a local history location and returns bounded record counts, which signals that output sizes are intentionally capped. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. Every phrase ('local history path,' 'health,' 'schema,' 'bounded record counts') adds distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-input inspection tool, the description plus the read-only/idempotent annotations and output schema cover almost everything an agent needs. The only minor gap is an explicit statement of when to choose this over sound_selection_inventory or sound_selection_get.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters and 100% schema coverage, there is no parameter burden for the description to carry. The tool takes no input, and the description focuses instead on what will be reported, which is the only semantic content relevant here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Report' and uniquely identifies the local Sound Selection history resource, listing four concrete output facets (path, health, schema, bounded record counts). It clearly distinguishes this status/inspection tool from mutating siblings like sound_selection_history_reset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use or when-not-to-use guidance and names no alternatives. Even though the read-only intent is inferable, the text itself leaves the agent to guess that this is a pre-operation inspection versus a debugging/support check.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_inventoryB
Read-onlyIdempotent

Read a compact loaded sound pool; Atlas-only products remain recommendations.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestNoOptional structured direction used to include the relevant target pool.
only_usedNoLimit mixer observations to used tracks; generators remain included.
preset_limitNoMaximum preset names per loaded target.
preset_startNoFirst preset index per target.
include_atlasNoEnrich loaded observations with local Plugin Atlas metadata.
include_currentNoRead current preset identities.
include_effectsNoInclude loaded effects; defaults from the request.
include_pad_mapsNoInspect generic generator pad maps.
include_empty_namesNoRetain blank preset names.

Output Schema

ParametersJSON Schema
NameRequiredDescription
warningsNo
observed_atNo
locked_rolesNo
loaded_effectsNo
schema_versionNo
loaded_generatorsNo
current_palette_idNo
session_fingerprintNo
known_unloaded_productsNo
preset_discovery_coverageNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide the core safety profile (readOnlyHint, openWorldHint, idempotentHint, no destructive hint). The description adds a useful behavioral nuance: Atlas-only products remain recommendations rather than being treated as loaded. This is extra context beyond the annotations, but it is fairly narrow and does not disclose other meaningful behavior such as pagination effects or how the request influences the pool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two terse clauses with no wasted words. The first clause states the action and object, and the second provides an important boundary condition. It earns its place and is front-loaded, though it is sparse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the schema, annotations, and output schema carry a lot of information, the description alone does not situate this tool among its sound_selection_* siblings. An agent can see it reads a loaded pool but has no explicit reason to choose it over sound_selection_plan or sound_selection_apply, nor does it state how its output should influence next steps. For a complex tool with a family of operations, this leaves a noticeable completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All top-level parameters have fully descriptive schema entries, so the parameter semantics are already well-covered. The description does not need to repeat them, and it adds a small contextual layer by hinting that the operation returns a "compact" pool (tying into preset_limit/preset_start). This matches the baseline expectation when schema coverage is 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a verb and resource: "Read a compact loaded sound pool." It gives a specific action and object, and the Atlas-only caveat adds further scoping. It does not explicitly name its sibling tools, so an agent has to infer how it differs from sound_selection_get or sound_selection_plan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. There is no mention of when to invoke it (e.g., before sound_selection_plan) or when not to use it. The only implication is that you use it to read the loaded sound pool, which is a faint and unsupported hint of context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_planA
Read-onlyIdempotent

Choose deterministic loaded-target assignments without changing FL or history.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesTask-scoped roles, direction, preferences, exclusions, continuity, and history policy.

Output Schema

ParametersJSON Schema
NameRequiredDescription
policyYes
blockersNo
drum_mapNo
warningsNo
conflictsNo
rationaleNo
palette_idYes
assignmentsNo
plan_digestNo
project_keyNo
anchor_rolesNo
flexible_rolesNo
request_digestYes
schema_versionNo
musical_directionNo
mutations_appliedNo
section_variationsNo
unused_candidate_idsNo
unused_candidate_targetsNo
preset_discovery_coverageNo
inventory_session_fingerprintNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true. The description adds beyond this by specifying 'deterministic' (same input → same output) and 'loaded-target assignments' (scope limited to already-loaded targets). This enriches the behavioral profile without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that leads with the core action and key constraint. There is no fluff, and it is front-loaded with the most important information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a planning tool with a rich output schema and extensive internal schema descriptions, the description is sufficient to convey that this tool creates a plan without side effects. It could benefit from naming the apply alternative for clarity, but given the annotations and schema, nothing critical is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema coverage is 100% and the `request` parameter already has a detailed description, so the tool description does not need to re-explain parameters. The description adds no parameter-specific meaning, but the schema fully covers it, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: 'Choose deterministic loaded-target assignments' – a clear verb and resource. It also clarifies a key constraint ('without changing FL or history'), which distinguishes it from mutating siblings like sound_selection_apply. The title 'Plan a coherent sound palette' reinforces the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a planning-only role by emphasizing 'without changing FL or history,' which hints that it is not for applying changes. However, it does not explicitly name alternative tools (e.g., sound_selection_apply) or provide conditions for selecting this tool over others. Usage guidance is adequate but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sound_selection_record_feedbackC

Update bounded local ranking feedback; silence is never inferred.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesExplicit accepted, rejected, or neutral palette feedback.

Output Schema

ParametersJSON Schema
NameRequiredDescription
historyYes
feedbackYes
warningsNo
persistedYes
schema_versionNo

TDQS

C2.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false, so the safety profile is known. The description adds a valuable non-obvious behavior: 'silence is never inferred' (i.e., absence of feedback is not treated as acceptance). This is a meaningful disclosure beyond the annotations and helps the agent understand the tool's semantics. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short sentence with no filler. It front-loads the action ('Update bounded local ranking feedback') and includes a critical behavioral note. However, the phrasing is somewhat opaque, which slightly reduces clarity, but it is concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and the request schema is fully documented, the description lacks essential context for correct use. It does not explain what 'bounded local ranking' means, when to record feedback, or how this tool fits into the broader sound_selection workflow. An agent would struggle to determine appropriate invocation conditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself provides detailed descriptions for the request object and its fields (e.g., 'Explicit accepted, rejected, or neutral palette feedback'). The description adds no new parameter-specific meaning beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update bounded local ranking feedback; silence is never inferred' is cryptic; it does not explicitly say 'record feedback' but the tool name and title make the purpose clear. It vaguely mentions 'bounded local ranking feedback' without explaining what that means, and it does not differentiate from sibling tools like sound_selection_apply or postfader_review_record_feedback.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus any of the many siblings. The description gives no context about when feedback should be recorded, what constitutes a valid use case, or any alternatives or exclusions. This leaves the agent without direction on tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 134 tool updatesv10.0.0
    • First observedarrangement_add_section_markers
    • First observedarrangement_prepare_pattern
    • First observedaudio_analyze_file
    • First observedaudio_analyze_masking
    • First observedaudio_compare_files
    • First observedaudio_estimate_tempo_and_key
    • First observedaudio_find_recent_bounces
    • First observedaudio_transcribe_melody
    • First observedautomation_record_value
    • First observedcompose_bassline
    • First observedcompose_chord_progression
    • First observedcompose_drums
    • First observedcompose_melody
    • First observedcopilot_capture_readonly_inspection
    • First observedfl_apply_verified_batch
    • First observedfl_find_empty_pattern
    • First observedfl_get_capabilities
    • First observedfl_get_plugin_preset_count
    • First observedfl_get_project_history
    • First observedfl_get_project_summary
    • First observedfl_get_selected_range
    • First observedfl_get_step_sequence
    • First observedfl_get_transport_state
    • First observedfl_inspect_mixer_track
    • First observedfl_list_channels
    • First observedfl_list_mixer_tracks
    • First observedfl_list_patterns
    • First observedfl_list_playlist_tracks
    • First observedfl_redo
    • First observedfl_route_channel_to_mixer
    • First observedfl_select_channel
    • First observedfl_select_mixer_track
    • First observedfl_select_pattern
    • First observedfl_select_plugin_preset
    • First observedfl_set_channel_identity
    • First observedfl_set_channel_mix
    • First observedfl_set_channel_pitch
    • First observedfl_set_channel_solo
    • First observedfl_set_loop_mode
    • First observedfl_set_metronome
    • First observedfl_set_mixer_arm
    • First observedfl_set_mixer_color
    • First observedfl_set_mixer_mute
    • First observedfl_set_mixer_name
    • First observedfl_set_mixer_pan
    • First observedfl_set_mixer_send
    • First observedfl_set_mixer_send_level
    • First observedfl_set_mixer_solo
    • First observedfl_set_mixer_stereo_separation
    • First observedfl_set_mixer_volume
    • First observedfl_set_mixer_volume_db
    • First observedfl_set_pattern_identity
    • First observedfl_set_pattern_length
    • First observedfl_set_playing
    • First observedfl_set_playlist_track_identity
    • First observedfl_set_playlist_track_state
    • First observedfl_set_plugin_param
    • First observedfl_set_plugin_param_display
    • First observedfl_set_plugin_param_option
    • First observedfl_set_precount
    • First observedfl_set_recording
    • First observedfl_set_song_position
    • First observedfl_set_step_sequence
    • First observedfl_set_tempo
    • First observedfl_set_time_signature_numerator
    • First observedfl_set_track_eq
    • First observedfl_set_write_mode
    • First observedfl_stop
    • First observedfl_trigger_note
    • First observedfl_undo
    • First observedmidi_export_type1
    • First observedmix_apply_plan
    • First observedmix_create_gain_stage_plan
    • First observedmix_create_plan
    • First observedmix_doctor
    • First observedmix_finish_assessment
    • First observedmix_get_peak_watch
    • First observedmix_get_plan
    • First observedmix_inspect_plugin_compatibility
    • First observedmix_list_plugin_profiles
    • First observedmix_masking_recommendations
    • First observedmix_reference_recommendations
    • First observedmix_resolve_processing_intent
    • First observedmix_start_peak_watch
    • First observedmix_stop_peak_watch
    • First observedpiano_roll_bridge
    • First observedpiano_roll_read_notes
    • First observedpiano_roll_transform
    • First observedpiano_roll_write_notes
    • First observedplugins_atlas_get_product
    • First observedplugins_atlas_inspect_loaded
    • First observedplugins_atlas_recommend
    • First observedplugins_atlas_search
    • First observedplugins_get_current_preset
    • First observedplugins_inspect_pad_map
    • First observedplugins_inspect_parameter_map
    • First observedplugins_list_available
    • First observedplugins_list_presets
    • First observedplugins_load
    • First observedplugins_scan_loaded_plugins
    • First observedplugins_scan_parameters
    • First observedpostfader_continue_run
    • First observedpostfader_creation_readiness
    • First observedpostfader_delivery_export_manifest
    • First observedpostfader_delivery_manifest
    • First observedpostfader_execute_run
    • First observedpostfader_get_run
    • First observedpostfader_list_runs
    • First observedpostfader_render_cancel
    • First observedpostfader_render_get_job
    • First observedpostfader_render_saved_project
    • First observedpostfader_review_apply_revision
    • First observedpostfader_review_attach_assets
    • First observedpostfader_review_compare
    • First observedpostfader_review_delete
    • First observedpostfader_review_evaluate
    • First observedpostfader_review_export_handoff
    • First observedpostfader_review_get
    • First observedpostfader_review_plan_revision
    • First observedpostfader_review_record_feedback
    • First observedpostfader_review_start
    • First observedpostfader_review_stop
    • First observedpostfader_stop_run
    • First observedpostfader_validate_run
    • First observedprocessing_apply_plan
    • First observedprocessing_plan
    • First observedsound_selection_apply
    • First observedsound_selection_create_variation
    • First observedsound_selection_get
    • First observedsound_selection_history_reset
    • First observedsound_selection_history_status
    • First observedsound_selection_inventory
    • First observedsound_selection_plan
    • First observedsound_selection_record_feedback

TDQS

B3.2/5.0

Scored across 134 tools

Disambiguation3/5

Many tools have clear distinct purposes, but there are overlapping tools such as fl_set_plugin_param, fl_set_plugin_param_display, fl_set_plugin_param_option, and fl_apply_verified_batch that may be confused. The postfader_* and mix_* groups also have many similar verbs (create, apply, plan) but are differentiated by domain prefixes.

Naming Consistency4/5

Most tools follow a consistent verb_noun or domain_verb_noun pattern (e.g., fl_set_mixer_volume, mix_create_plan). There are minor deviations like 'piano_roll_bridge' (noun_verb) and 'copilot_capture_readonly_inspection' (prefix_verb_adj_noun) but overall the pattern is predictable.

Tool Count2/5

The server has 134 tools, which is far beyond the typical well-scoped range. While the server covers multiple domains (FL Studio control, audio analysis, composition, rendering), the sheer number suggests insufficient consolidation and will likely overwhelm agents.

Completeness4/5

The server covers a wide range of workflows including project state, mixer control, plugin management, composition, audio analysis, and deliverable export. There are some potential gaps like missing explicit delete for certain entities (e.g., no fl_delete_channel or fl_delete_playlist_track), but the lifecycle coverage is generally strong given the domain breadth.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables AI assistants to control FL Studio on Windows through a local MCP server without cloud dependencies. Bridges MCP clients to FL Studio's Python scripting environment via file-based IPC for project management, transport control, and UI workflow automation.
    -
  • A
    license
    A
    quality
    C
    maintenance
    AI control for FL Studio via the Model Context Protocol — full in-DAW mixing (Mix Doctor diagnosis, gain staging, EQ/compression/reverb, reference matching), routing, and composition through Claude and any MCP client. 67 tools. Windows.
    67
    51
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Local MCP server for FL Studio on macOS, enabling control of projects, transport, mixer, plugins, automation, and piano roll operations.
    MIT