av-conductor
Provides tools for editing Max/MSP patches: adding and connecting objects, and reading documentation via the upstream maxmsp MCP server; the av-conductor can also send OSC scene data to Max receiver patches.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@av-conductorSet the scene to section intro with energy 0.2, then ramp energy to 1 over 20 seconds."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Claude_AV_MCP
Sandbox for audio-visual MCP experiments aimed at technical artists: Claude builds instruments and visuals and conducts live AV sets across Ableton Live, Max/MSP and a visual engine (TBD) — first on a MacBook Pro M5, later on a cheap Linux laptop with an all-open-source stack.
How it fits together
Server | Role | Source |
| Edit the Live set: tracks, clips, notes, devices, transport | ahujasid/ableton-mcp (upstream, unchanged) |
| Edit Max patches: add/connect objects, read docs | tiianhk/MaxMSP-MCP-Server (upstream, unchanged) |
| Shared scene (energy, section, palette, cues, ramps) → OSC to every program | this repo, |
Upstream servers author; the conductor directs; beat-accurate sync and audio reactivity stay in the realtime plane (Ableton Link, OSC between programs) where Claude never blocks them. Details: docs/ARCHITECTURE.md.
Related MCP server: Ableton MCP
Quick start (Mac)
# 1. Python env for the conductor (needs uv: https://docs.astral.sh/uv/)
uv sync --extra dev
uv run pytest
# 2. Upstream servers → vendor/ (or symlink your existing clones there)
scripts/fetch_upstream.sh
# 3. Watch the conductor's output with no apps running
cp config/av.example.toml config/av.toml # set [targets.monitor] enabled = true
uv run av-osc-monitor --port 7499Then open Claude Code in this folder — .mcp.json registers all three
servers. For Claude Desktop, merge
config/claude_desktop_config.example.json into its config (absolute paths).
In Max, open patches/max/av_receiver.maxpat; in Pd,
patches/pd/av_receiver.pd. Try:
"Set the scene to section intro with energy 0.2, then ramp energy to 1 over 20 seconds."
Layout
src/av_mcp/ av-conductor MCP server (scene, OSC targets, ramps, cues)
config/ av.example.toml (targets/ports/cues), Claude Desktop example
patches/max|pd/ OSC receiver patches → [r av-energy], [r av-section] …
visuals/ visual engine candidates + receiver sketches (TBD)
prompts/ prompt library
docs/ PLAN.md · ARCHITECTURE.md · LINUX_PORT.md
scripts/ fetch_upstream.sh
vendor/ upstream MCP servers (gitignored)Roadmap
docs/PLAN.md: Phase 0 plumbing → 1 audio/control loop → 2 pick the visual engine → 3 performance tooling → 4 Linux port (Pd, Ardour/SuperCollider, PipeWire, open-source visuals).
Privacy note
ableton-mcp sends anonymous telemetry by default. The configs here set
ABLETON_MCP_DISABLE_TELEMETRY=true and ABLETON_MCP_DISABLE_DATASET=true.
Available Tools
7 toolsget_stateC
Current scene, saved cue names, targets, and running ramps.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing beyond the returned content: no statement that it is a side-effect-free read, no auth requirements, no refresh/polling semantics. The 'get' name implies a safe read but this is never asserted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact clause with no filler, and the most important content (current scene) is front-loaded. It is efficient, though it is terse to the point of omitting the verb and any usage framing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain the return payload, and it sensibly summarizes it instead. However, with no annotations and no usage guidance, an agent still lacks context on when and why to invoke it relative to the six sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema has no properties, so there is nothing for the description to clarify. Baseline 4 applies for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase listing returned data ('Current scene, saved cue names, targets, and running ramps') rather than a verb+resource statement, so the retrieval action is only implied by the tool name. It does not contrast itself with siblings like recall_cue or resend_all, though the 'current state' semantics are distinguishable from the mutating siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus the sibling tools (save_cue, recall_cue, set_scene, ramp). No prerequisites, timing, or exclusions are stated. An agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rampB
Glide a numeric value over time (a build, a fade). Returns immediately.
param: "energy", "bpm", or "macro:" (e.g. "macro:brightness").
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| param | Yes | ||
| seconds | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden; it does disclose one important trait — the call returns immediately, so the glide continues in the background rather than blocking. It does not say what happens if a second ramp targets the same param, whether a ramp can be cancelled, or what side effects occur on the underlying value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very short and front-loaded with the core verb and effect. The trailing param line is slightly awkwardly formatted but adds real information rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not needed, and the fire-and-forget behavior is stated. But for a 3-required-param mutation tool with zero annotations and zero schema descriptions, the definition still leaves concurrency/interaction behavior and two parameter meanings unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does clarify the 'param' argument's accepted forms ('energy', 'bpm', 'macro:<name>') plus an example. However, 'to' and 'seconds' are left entirely unexplained beyond their names, so two of three parameters rely on inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and effect ('Glide a numeric value over time') with concrete examples ('a build, a fade'), which tells the agent exactly what this does. It is clearly distinct from siblings like set_scene or send_osc, though it never names or contrasts them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose ramp over alternatives such as set_scene or send_osc, nor any prerequisites (e.g. whether the param must already exist). The only contextual cue is 'Returns immediately', which hints at the async style but is not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recall_cueB
Jump to a saved cue (from config/av.toml or save_cue) and send /av/cue .
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the concrete side effect — sending /av/cue <name> — which is more than most, but it omits failure behavior for unknown names, whether the operation is idempotent, and whether it requires the cue to have been saved in the current session.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence covering action, name source, and emitted message. No filler, though the trailing OSC path detail is only marginally useful to an agent deciding whether to call it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the core action and parameter source. It is still thin on error conditions and preconditions for a tool that performs a state-changing dispatch.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It echoes the <name> placeholder and, helpfully, states where valid names originate (config/av.toml or save_cue), but it gives no format, casing, or naming constraints beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Jump to a saved cue" gives a specific action on a specific resource, and naming save_cue as the source makes the pairing with the sibling explicit. It is clear what the tool does, though "jump to" is slightly metaphorical for what is ultimately an OSC dispatch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentioning that cues come from config/av.toml or save_cue implies the recall-after-save workflow, which is useful context. However, there is no explicit when-to-use guidance or exclusion, and no statement about what happens if the name does not exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resend_allA
Re-broadcast the whole scene, e.g. after restarting Max or the visual engine.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that the operation is scene-wide (not selective), but says nothing about idempotency, cost, whether it must follow a state change, or side effects on a running engine. Output schema existing means return values need not be covered here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and scope, followed by a parenthetical use case. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter action with an output schema, the description covers what it does and when to reach for it, which is close to sufficient. Minor gaps remain around prerequisites (e.g. whether it must be called after set_scene) and what the broadcast actually covers, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is no parameter semantics to explain or get wrong. Schema coverage is 100% and nothing in the schema requires description-side compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (re-broadcast) and a clear resource scope (the whole scene), which distinguishes it from the targeted-send semantics implied by sibling send_osc. It does not explicitly name or contrast itself with any sibling tool, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete trigger scenario ('after restarting Max or the visual engine'), giving the agent a clear condition under which this tool is appropriate. It offers no exclusions or explicitly named alternatives (e.g. send_osc for single messages), so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_cueB
Snapshot the current scene under a name (kept until the server restarts).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one genuinely useful trait: the snapshot is volatile and lost on server restart. It omits other mutation behavior such as whether an existing name is overwritten and whether any prerequisite state is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero waste; the action comes first and the durability caveat is parenthetically attached rather than padded into its own sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and for a one-parameter tool the description covers the action and a key behavioral caveat. Only name-collision/overwrite behavior is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — the schema says only that 'name' is a string — so the description must compensate. 'Under a name' identifies it as a label but adds no uniqueness, overwrite, or format semantics, leaving the parameter effectively undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (snapshot) and resource (current scene) plus the naming parameter, so the operation is unambiguous. It never explicitly names the sibling recall_cue, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The pairing with recall_cue is strongly implied by the save/restore naming, and 'current scene' implies you call it while the desired state is live, but no alternative or when-not condition is ever stated. Usage is inferred rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_oscA
Escape hatch: send a raw OSC message, to all targets or the named ones.
Use for program-specific messages the scene model doesn't cover (e.g. "/visuals/layer/2/opacity" 0.5).
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | ||
| address | Yes | ||
| targets | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the broadcast default ("to all targets") and frames itself as an escape hatch, but says nothing about side effects, permissions, ordering, or whether omitted targets is safe. Output schema exists, so return values need not be described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the escape-hatch framing and the routing behavior, then the use case. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter raw-message tool with no annotations, the description covers purpose, usage, and target defaulting, but leaves most parameter semantics and side-effect behavior undocumented. Adequate but with clear gaps given zero schema description coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It implicitly conveys address and args via the example ("/visuals/layer/2/opacity" 0.5) and clarifies the targets parameter's default behavior ("all targets or the named ones"), but leaves arg typing and the full meaning of targets unspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("send a raw OSC message") and scopes it ("to all targets or the named ones"), which an agent can distinguish from scene-model siblings like set_scene. It is clear but does not explicitly name which sibling it replaces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for program-specific messages the scene model doesn't cover" gives a clear when-to-use condition with a concrete example address. It implies the alternative (the scene model / set_scene) without naming it explicitly, so no exclusion list is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_sceneA
Change the shared scene and broadcast it to every target over OSC.
energy and macros are 0..1. palette is a list of hex colors. Only fields you pass are changed. Note: bpm here is informational for the visuals; to change Live's tempo also call ableton-mcp's set_tempo (or use Ableton Link so everything follows Live).
| Name | Required | Description | Default |
|---|---|---|---|
| bpm | No | ||
| key | No | ||
| energy | No | ||
| macros | No | ||
| palette | No | ||
| section | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses the OSC broadcast to every target, the partial-update semantics (unspecified fields are left alone), and that bpm is not the real tempo. It omits reversibility, whether cues are needed to persist the scene, and error/rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, followed by compact parameter notes and a tempo caveat. Every sentence carries information, though the parameter notes are telegraphic enough to be slightly fragmented.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be explained, and the mutation semantics and OSC broadcast are covered. However, for a 6-parameter mutation tool with zero annotations and zero schema descriptions, the unexplained key and section parameters and the unstated meaning of a 'scene' leave real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must compensate. It defines ranges for energy and macros (0..1), the palette format (list of hex colors), and partially explains bpm, but key and section are left entirely undefined in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (change), resource (the shared scene), and a distinctive side effect (broadcast to every target over OSC), which separates it from the raw send_osc sibling. It does not name any sibling explicitly to contrast against, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit cross-tool routing for the tempo case: bpm here is informational, and changing Live's tempo requires ableton-mcp's set_tempo or Ableton Link. It also clarifies the partial-update contract ('only fields you pass are changed'). No guidance on when to use save_cue/recall_cue instead, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
get_state - First observed
ramp - First observed
recall_cue - First observed
resend_all - First observed
save_cue - First observed
send_osc - First observed
set_scene
TDQS
Scored across 7 tools
Each tool has a clearly distinct purpose: save/recall cues, raw OSC escape, full rebroadcast, state query, scene update, and time-based ramping. No two tools overlap in function or would cause misselection.
Names are consistently snake_case and mostly follow a verb_noun pattern (save_cue, recall_cue, send_osc, get_state, set_scene, resend_all). The lone 'ramp' breaks the pattern slightly, but the overall convention is clear.
Seven tools are well-scoped for an AV conductor, covering scene management, cue handling, OSC messaging, and ramping without bloat. Each tool earns its place.
Core lifecycle is present (create/read/update for scenes and cues), but notable gaps exist: no tool to delete or overwrite cues explicitly, and no way to cancel or stop an active ramp. These are common AV operations that agents cannot perform directly.
Maintenance
Related MCP Connectors
Deploy sims to any screen. Control your displays with Claude.
- SimSenseOAuthai.simsense
Deploy sims to any screen. Control your displays with Claude.
Create, co-edit, analyze, publish, and export collaborative step-sequencer sessions through MCP.
Real-time multi-track schedules for cooking, lab protocols, events and workouts. Validate and share.
Related MCP Servers
- AlicenseCqualityDmaintenanceConnects Claude AI to Ableton Live through the Model Context Protocol, enabling prompt-assisted music production with track creation, instrument loading, clip editing, and session control. Allows users to create complete musical arrangements by describing what they want in natural language.373MIT
- AlicenseNot gradedqualityDmaintenanceEnables natural language control over Ableton Live for generating musical patterns, melodies, and full song arrangements. It also provides tools for sample searching and mixing assistance through an OSC-based connection with Claude Desktop.1MIT
- AlicenseBqualityDmaintenanceEnables Claude to observe and compose in Ableton Live 12 via AbletonOSC, allowing pair-programming-style collaboration for musicians.13MIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude AI to control Ableton Live and Max for Live, allowing music production tasks like track management, MIDI editing, and pattern generation directly from conversation.-