gemini_music_generate
Generate music from a text prompt specifying mood, genre, instruments, or lyrics. Outputs an MP3 via Lyria models; single-turn, so include the full brief and a wait budget for long runs.
Instructions
Generate music from a text prompt (mood, genre, instruments, structure, or lyrics inline) via a Lyria model: lyria-3-clip-preview (30s instrumental clip, default, cheapest), lyria-3.5 (full-length song with vocals) or lyria-3-pro-preview (longer-form). Output is MP3, written to disk (or returned inline). Single-turn: a track cannot be refined by a follow-up call, so put the whole brief in the prompt. Runs long — give it a max_wait_ms budget (or async: true + gemini_get_result on a local install), or raise timeout_ms. Needs a funded account. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| model | No | Lyria model (default: lyria-3-clip-preview — 30s, $0.04). lyria-3.5 and lyria-3-pro-preview run minutes-long at $0.08. | |
| images | No | Optional reference image path(s) to condition the music | |
| inline | No | Return base64 audio inline instead of writing to disk. A 30s track is ~1.9MB of base64 — the default hands back a path instead | |
| prompt | Yes | Description of the music: mood, genre, instruments, tempo, structure, or lyrics | |
| filename | No | Base filename for the output audio (extension stripped; default: slugified prompt) | |
| background | No | Run the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Off by default — see gemini_video_generate | |
| images_url | No | Reference images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write audio to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| images_base64 | No | Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS clipboard as a reference | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Reference images as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. |