Skip to main content
Glama

transcribe_stem

Destructive

Transcribe a single-instrument stem into clean notes, removing harmonics and noise, with options for sustained and monophonic parts.

Instructions

Turn a recording of one instrument into notes, cleaned of what it is not.

Returns:
    Dictionary with how many notes were heard and how many survived cleaning, the
    pitch range, the span in beats, a sample of the notes, and the write result
    where one was asked for.

Note:
    **This transcribes pitch, and percussion has no pitch.** A drum kit returns cymbal
    and shell resonances, reporting success: measured on a four minute drum stem, 39
    notes for 488 beats, none of them a kick, and nothing across forty bars of steady
    playing. Never write drums from here. Held material fails the other way: a wavering
    sustained note reads as several, so a pad arrives at the right pitches with a rhythm
    nobody played. Measured 2026-09-19 over four stems in their busiest 20 seconds,
    written attacks ran 3.35 a second against 2.85 onsets on bass, 3.50 against 1.30 on
    a strummed guitar, and 5.35 and 5.30 against 0.10 and 0.55 on pad and keys.

    ``transcription_suspect`` flags that disagreement, ``kind`` says which way, and
    ``transcription_checked`` says whether the comparison ran at all, since it needs a
    macOS-only second decode: elsewhere no warning means nothing. Measured 2026-09-19
    against exported MIDI for five pitched stems, pitch content agreed 93.7 to 96.2
    percent: which notes, not when, the export's own timing having drifted. Harmonics
    are removed where a lower, louder note at the same time explains them;
    ``monophonic`` then leaves nothing sounding at the same time as anything else.
    ``scripts/eval_transcription.py`` recomputes these.

    Writing needs a clip: ``create_clip`` on the target slot first, long enough for the
    whole part, then this with ``confirm=True``. ``span_beats`` runs from beat 0 to the
    last note's end, so a clip of that length rounded up to a bar holds it. Writing
    replaces every note in the clip, with no undo here and no rollback in the LOM: a
    write that fails after the removal leaves nothing.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to the audio to transcribe. A separated stem, one instrument at a time. A whole mix transcribes as one tangle.
slotNoSession clip slot on that track.
tempoYesBeats per minute, to turn seconds into clip beats. Read the set's own tempo unless the recording has a different one. A recording whose tempo drifts will separate from one number over the length of a song.
trackNoTrack to write the notes into. -1 returns them instead of writing, which is the default and is how to look before committing.
confirmNoTrue writes into the clip, replacing every note already in it. False reports what would be written and changes nothing.
previewNoHow many notes to include in the answer.
sustainedNoTrue for material that holds its notes rather than striking them: a pad, a string section, an organ. It joins same-pitch fragments separated by a gap too small to play, which is how one held chord stops arriving as eight notes. Off by default because on struck material it costs agreement on pitch content, between one and ten points measured over five stems, and it does not repair a pad: it took one from 5.35 written attacks a second to 2.25, where the recording started 0.10. See transcription_suspect.
monophonicNoTrue for a part played one note at a time, such as a bass or a sung line. It reduces the result to the lowest note sounding, which is how a bass loses the harmonics a transcriber hears above it, and which throws away the upper voices of anything that plays chords.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark destructiveHint=true, and the description adds substantial context beyond that: writing replaces every note, has no undo and no LOM rollback, and a write that fails after removal leaves nothing. It also discloses nuanced failure modes around percussion and sustained pads, plus measurements behind transcription_suspect, which annotations alone could never convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but almost every sentence carries operational weight: measured failure rates, no-undo guarantees, and the macOS-only check caveat. It is front-loaded with purpose, then Returns, then behavioral notes. Some empirical figures could be trimmed without loss, but the structure is logical and dense rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, write-capable, destructive transcription tool, the description covers prerequisites, preview behavior, failure modes, output contents, and safety consequences. The output schema and annotations cover machine-readable details, so nothing an agent needs to call this tool correctly and safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all parameters in detail. The description adds meaningful operational context on top: 'Writing needs a clip' explains slot/confirm usage, and the span_beats/clip-length note gives practical meaning to tempo and output interpretation. A few parameters like path and preview get no new detail, but the schema fully carries them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line states a specific verb and resource: 'Turn a recording of one instrument into notes, cleaned of what it is not.' This clearly distinguishes it from mix analysis and from clip note reading/writing siblings, especially combined with the path schema contrast between a separated stem and 'a whole mix transcribes as one tangle.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit operational guidance: use separated stems, never write drums ('Never write drums from here'), and before writing create a clip and pass confirm=True. It also clarifies that track=-1 previews without writing. It does not explicitly name an alternative tool, but the when-to and when-not-to conditions are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.