Audio
OfficialProvides a check/normalize target for the Netflix loudness delivery specification, measuring audio against each rule of the spec and reporting which ones pass.
Supports running the audio bridge's agent chat against a local Ollama server (e.g. ANTHROPIC_BASE_URL=http://localhost:11434 with token ollama, or ollama launch pi --config) so audio can be measured, edited and played by locally hosted models.
audio

Audio playback, editing and analysis
Any Format — fast wasm codecs, no ffmpeg.
Non-destructive — virtual edits, infinite undo, instant clone.
Stream-first — playback/encode during decode, realtime editing.
Paged — no 2Gb memory limit, open 10Gb+ files.
Analysis — loudness, spectrum, beats, pitch, chords, key.
Modular – pluggable ops, tree-shakable.
CLI — playback, batch processing, scripting, unix pipes, tab completion.
Cross-platform — browsers, node, deno, bun.
Start Recipes API CLI FAQ Plugins Architecture Comparison
Start
Node
npm i audio
import audio from 'audio'
audio('voice.mp3').trim().normalize('podcast').fade(0.3, 0.5).save('clean.mp3')Browser
<script type="module">
import audio from 'https://esm.sh/audio'
audio('./song.mp3').trim().normalize().fade(0.5, 2).clip({ at: 60, duration: 30 }).play()
</script>CLI
npm i -g audio # or: npx audio …
audio voice.wav trim normalize podcast fade 0.3s -0.5s save clean.mp3Skill
npx skills add audiojs/audioMCP
npx add-mcp "npx -y audio --mcp" # asks which of your agents: Claude Code, Codex, Cursor, Gemini CLI, Kimi Code, Pi, OpenCode, Zed…Prompt: make ~/Desktop/interview.m4a podcast-ready and tell me the loudness before and after
One agent at a time: claude mcp add audio -- npx -y audio --mcp, qwen mcp add audio npx -y audio --mcp, droid mcp add audio "npx -y audio --mcp".
Editor and agents
npx audio --bridge # audio bridge on http://127.0.0.1:7777 key K agents claude, codex, pi, kimiConnect the editor to it with the key (once: the bridge keeps it), and its chat runs an agent of yours, which measures, looks at, edits and plays the sound open there; each tab keeps its conversations, each with the agent picked under the message. The bridge finds Claude Code, Codex, Pi, Gemini CLI, Qwen Code, Kimi Code, OpenCode, Kilo Code, Cline, Goose, Factory Droid, Cursor, Augment, Kiro and Mistral Vibe on PATH; any other that speaks ACP runs by its command line, --agent "my-agent --acp".
Any MCP agent gets the editor's tools (state, measure, look, edit, select, play, check, …), the running bridge found by itself: npx add-mcp "npx -y audio --mcp --editor".
An agent thinks with the model its own settings name: Pi takes local ones for good from ollama launch pi --config; Claude Code takes any Anthropic-compatible endpoint from the bridge's environment:
ANTHROPIC_BASE_URL=http://localhost:11434 ANTHROPIC_AUTH_TOKEN=ollama ANTHROPIC_API_KEY= ANTHROPIC_MODEL=qwen3-coder npx audio --bridgeModels |
|
Ollama |
|
Z.ai GLM |
|
Kimi |
|
Qwen |
|
DeepSeek |
|
Related MCP server: studiosphere-pulse-mcp
Recipes
Clean up
// master a raw take
let a = audio('raw-take.wav')
a.trim(-30).normalize('podcast').fade(0.3, 0.5)
await a.save('clean.wav')
// full restoration chain via ecosystem plugins (see API › Plugins)
a.gate(-45).dehum().deesser().compressor({ threshold: -18 }).limiter({ ceiling: -1 })
// cut 2:00–2:15, smooth the splice
a.remove({ at: 120, duration: 15 }).fade(0.1, { at: 120 })
// find clipped blocks
let clips = await a.stat('clipping')
// does it pass? each rule of the spec, measured
let { pass, rules } = await a.check('podcast') // acx, podcast, streaming, broadcast, netflixMaster & deliver
// master a song to a reference track: its tone (mid and side), width and loudness, under -1 dBTP
audio('mix.wav').master(await audio('reference.wav')).save('master.wav')
// audiobook chapter for ACX: RMS -23..-18 dB, peaks under -3 dB, floor under -60 dB, room tone at each end, 44.1 kHz
let ch = audio('chapter-01.wav')
.highpass(80).omlsa({ gMin: -12 }).compressor({ threshold: -24, ratio: 2.5 })
.normalize(-20, 'rms', { ceiling: -3.5 })
.trim().pad(1.5, 2).roomtone() // room tone, not digital silence
.resample(44100)
console.log(await ch.check('acx')) // passed 10 of 10 real narrations (.work/pro.md)
await ch.save('chapter-01.mp3', { bitrate: 192 })
// tighten pauses in a talking-head video, then hand the cuts to the video editor
let talk = audio('talk.mp4').shrink(0.3)
await talk.save('talk.m4a') // the sound, cut
await talk.save('talk.edl') // the same cuts for the picture (Premiere, Resolve)Compose
// podcast montage
let ep = audio([intro, interview.trim().normalize('podcast'), outro], { crossfade: 0.5 })
await ep.save('episode.mp3')
// voiceover over music
music.gain(-12).mix(voice, { at: 2 })
// ringtone: the chorus + fades
audio('song.mp3').crop({ at: 45, duration: 30 }).fade(0.5, 2).normalize().save('ringtone.mp3')
// split an audiobook into chapters
let [ch1, ch2, ch3] = audio('audiobook.mp3').split(1800, 3600)
// glitch: stutter + reverse
let v = a.clip({ at: 1, duration: 0.25 })
audio([v, v, v, v]).reverse({ at: 0.25, duration: 0.25 })Analyze
// waveform bars — and progressively, as it decodes
let [mins, peaks] = await a.stat(['min', 'max'], { bins: canvas.width })
a.on('data', ({ delta }) => appendBars(delta.max[0], delta.min[0]))
// features for ML
let mfcc = await a.stat('cepstrum', { bins: 13 })
let [loud, rms] = await a.stat(['loudness', 'rms'])
// notes, chords, key
let notes = await a.stat('notes') // [{time, duration, freq, midi, note, clarity}]
let chords = await a.stat('chords') // [{time, duration, label, root, quality, confidence}]
let key = await a.stat('key') // {tonic, mode, label, confidence}Record & generate
// mic take
let a = audio()
a.record()
// …later
a.stop()
a.trim().normalize()
// tone — any t => sample function
let tone = audio.from(t => Math.sin(440 * Math.PI * 2 * t), { duration: 2 })
// sonify data
let s = audio.from(t => Math.sin((200 + data[t / 0.2 | 0]) * Math.PI * 2 * t) * 0.5, { duration: data.length * 0.2 })Automate
a.gain(t => -12 * (0.5 + 0.5 * Math.cos(t * Math.PI * 4))) // 2Hz tremolo in dB
a.lowpass(t => 400 + 4000 * t) // filter sweep
a.pan({ t: [0, 2, 4], v: [-1, 1, -1] }) // curve L→R→L over 4s, serializable
music.ducker({ key: voice }) // sidechain (plugin)Stream & persist
// stream to network — encode/playback during decode
for await (let chunk of audio('2hour-mix.flac').highpass(40)) socket.send(chunk[0].buffer)
// serialize edits, restore later
let json = JSON.stringify(a) // { source, edits, ... }
let b = audio(JSON.parse(json)) // re-decode + replay editsAPI
Create
Method | Description |
| decode from file, URL, bytes, or a byte stream. Returns instantly — decodes in background; streams decode as they arrive. |
| wrap existing PCM, AudioBuffer, silence, or function. Sync, no I/O. |
let a = audio('voice.mp3') // file path
let b = audio('https://cdn.ex/track.mp3') // URL
let c = audio(inputEl.files[0]) // Blob, File, Response, ArrayBuffer
let d = audio() // empty, ready for .push() or .record()
let e = audio([intro, body, outro]) // concat (virtual, no copy)
let f = audio([a, b, c], { crossfade: 2 }) // concat with 2s crossfade
let g = audio(process.stdin) // byte stream: pipe, socket, fetch body – decodes as it arrives
// opts: { sampleRate, channels, crossfade, curve, storage: 'memory' | 'persistent' | 'auto' }
await a // await for decode — if you need .duration, full stats etc
let a = audio.from([left, right]) // Float32Array[] channels
let b = audio.from(3, { channels: 2 }) // 3s silence
let c = audio.from(t => Math.sin(440*TAU*t), { duration: 2 }) // generator
let d = audio.from(audioBuffer) // Web Audio AudioBuffer
let e = audio.from(int16arr, { format: 'int16' }) // typed array + formatProperties
Property | Description |
| total seconds, after edits. |
| channel count. |
| sample rate. |
| samples per channel. |
| what the speakers play now, in seconds: their latency compensated, smooth; at pause it holds where playback resumes. |
| true during playback. |
| true when paused. |
| 0..1 linear. Settable. |
| mute, independent of volume. Settable. |
| settable, mid-playback too: the span repeats, each seam a 10 ms equal-power crossfade. |
| 0.0625..16, settable mid-playback, click-free. The pitch stays (WSOLA, as a browser's media element plays at a speed) unless |
| true: at a rate other than 1, the pitch kept. false: varispeed, the pitch with the speed. Settable mid-playback. |
| true when playback reached the end, not after |
| true during a seek. |
| promise, resolves when playback sounds. |
| true during mic recording. |
| promise, resolves when fully decoded. |
| original source. |
| stored sample depth of the source: 16, 24, 32 (float); null for lossy or generated audio. Lossless |
|
|
| per-block stats (peak, rms, etc.). |
| edit list. |
| increments on each edit. |
Structure
Method | Description |
| strip leading/trailing silence (dB, default auto). On a live stream a given threshold streams, holding a silent tail until sound resumes; the automatic one reads the whole input, so it waits for the end. |
| shorten silent pauses to |
| keep range, discard rest. |
| delete range, close gap. |
| insert audio (default: at end), or a number of seconds of silence; |
| copy range (default: all) to this instance's clipboard. |
| copy, then remove. |
| insert the clipboard (default: at end). |
| slide a range to |
| zero-copy excerpt as a new instance. |
| zero-copy excerpts between timestamps. |
| silence at edges (seconds). |
| repeat n times. |
| reverse audio or range. |
| changes pitch and duration together. |
| changes duration, keeps pitch (phase-locked vocoder). A |
| move moments in time: |
| changes pitch, keeps duration. Semitones may be a curve |
| a voice's rises and falls wider or flatter about its median pitch: |
| moves the formants (the spectral envelope: a voice's vowels, the size of its head), keeps the pitch. A number, a curve |
| channel count (down per ITU-R BS.775: 7.1 → 5.1, stereo, mono), or a map: |
Every op takes a trailing {at, duration, channel} range, except channel-changing remix and crossover. Times are seconds or strings ('1:30', '2m'); negative counts from the end. FFmpeg's short names work wherever the long ones do: d for duration, xfade for crossfade (a.remove({ at: 1, d: 0.5, xfade: 0.01 })).
a.trim(-30) // strip silence below -30dB
a.remove({ at: '2m', duration: 15 }) // delete 2:00–2:15, close gap
a.remove(12.3, 0.4, '10ms') // cut a breath, crossfaded: no click
a.insert(intro, { at: 0 }) // prepend; .insert(3) appends 3s silence
a.copy(60, 30).paste(120) // duplicate the chorus at 2:00
a.cut(2, 1).paste(5) // move 2s–3s to 5s of the shortened timeline
a.move({ at: 2, duration: 1, to: 5 }) // slide 2s–3s over 5s–6s, silence left at 2s–3s
let [pt1, pt2] = a.split('30m') // zero-copy parts
let hook = a.clip({ at: 60, duration: 30 }) // zero-copy excerpt
a.stretch(1.1) // 10% longer, same pitch
a.warp([[1, 1], [2, 2.4], [3, 3]]) // the hit at 2s lands at 2.4s; 1s–3s keeps its length
a.pitch(-2) // 2 semitones down, same tempo
a.pitch({ t: [1, 1.2, 2, 2.2], v: [0, 3, 3, 0] }, { voice: true }) // a note drawn 3 semitones up, formants kept
a.intonation(1.5, { at: 2, d: 3 }) // 2s–5s: every rise and fall half as wide again
a.formant(-2) // a larger, darker voice on the same notes
a.remix([0, 0]) // L→both; .remix(1) for monoProcess
Method | Description |
|
|
| curves |
| remove DC, normalize. Loudness targets hold a true-peak ceiling, -1 dBTP by default: a lookahead limiter, then the loudness it took made back up. Presets per Apple Podcasts, Spotify, EBU R 128 (ITU-R BS.1770-4): |
| fill digital silence (≥ 10 ms under -90 dBFS: edited-out pauses, |
| overlay at |
| append with overlap, default 0.5s. |
| −1 left, 0 center, 1 right. |
| overwrite from |
| inline |
a.gain(-3) // reduce 3dB
a.gain(6, { at: 10, duration: 5 }) // boost range
a.gain(t => -12 * Math.cos(t * TAU)) // automate over time
a.fade(0.5, -2, 'exp') // 0.5s in, 2s exp fade-out
a.normalize('podcast') // -16 LUFS, -1 dBTP
a.normalize(-27, 'lufs', { ceiling: -2 }) // Netflix
a.mix(voice, { at: 2 }) // overlay at 2s
a.mix(bed, 0, -18) // music bed, 18 dB under
a.crossfade(next, 2) // 2s crossfade into next
a.crossfade(song2, 4, 'equal') // equal-power, for unrelated tracks
a.pan(-0.3, { at: 10, duration: 5 }) // pan left for rangeFilter
Method | Description |
| Butterworth pass filter; even integer order ≥ 2: 2 (12 dB/oct, default), 4 (24), 6, 8, … Other orders are rejected. |
| band-pass / notch. |
| phase shift, unity magnitude. |
| shelf EQ. |
| parametric EQ. |
| by type name, or a custom filter function. |
All biquads.
a.highpass(80).lowshelf(200, -3) // rumble + mud
a.eq(3000, 2, 1.5).highshelf(8000, 3) // presence + air
a.notch(50) // remove hum
a.allpass(1000) // phase shift at 1kHz
a.filter(customFn, { cutoff: 2000 }) // custom filter functionEffect
Method | Description |
| mid/side: |
| TPDF, default 16-bit. |
| headphone crossfeed, default 700 Hz, 0.3.≡ SoX |
| upsampling defaults to linear, downsampling to anti-aliased windowed sinc, its taps widening with the ratio. |
| N split frequencies → N+1 bands × channels, band-major. Linkwitz-Riley 4th order; bands sum back flat.≡ FFmpeg |
| match EQ: up to 8 parametric bands fit to the reference/source spectrum ratio. Tone only; loudness stays with |
| master to a reference track: |
| gain on a time × frequency region, |
| rebuild a damaged range (dropout, beep, click burst) from its surroundings. |
| remove a noise that holds still (hiss, hum and buzz, a fan, room tone, tape), learned where it plays alone: |
| neural speech denoising through the optional |
a.vocals() // isolate center-panned vocals
a.vocals('remove') // remove vocals (karaoke)
a.vocals({ model: 'umxhq' }) // vocals by a separation model
a.dither(16) // TPDF dither to 16-bit
a.dither(16, { shape: true }) // noise-shaped
a.crossfeed() // headphone crossfeed
a.resample(48000) // resample to 48kHz (linear)
a.resample(96000, { type: 'sinc' }) // high-quality windowed-sinc
a.match(reference, 0.7) // 70% of the way to its tone
a.spectral([1000, 4000], -30, { at: 2.1, duration: 0.3 }) // a cough
a.repair({ at: 1.2, duration: 0.05 }) // a dropout
a.repair({ at: 42, duration: 1 }) // a lost second of music: the passage that fits
a.denoise({ noise: { at: 1.2, duration: 0.5 } }) // hiss learned from a pause, 12 dB down everywhere
a.deepfilter() // speech out of noise: the noise 45 dB under the voice, room tone kept
a.rnnoise() // the same, streamingI/O
Method | Description |
| rendered PCM. |
| encode + write, format from extension. Lossless keeps the source depth; |
| encode to |
| the edits as a cut list for a video editor: |
| independent edits, shared pages. |
| feed PCM into a pushable instance; |
let pcm = await a.read() // Float32Array[]
let raw = await a.read({ format: 'int16', channel: 0 })
for await (let block of a) send(block) // async-iterable over blocks
await a.save('out.mp3') // format from extension
await a.save('book.mp3', { bitrate: 192 }) // ACX: 192 kbps CBR
await a.save('master.wav', { bitDepth: 24 }) // 24-bit
await a.save('talk.mp4') // video in, video out
let bytes = await a.encode('flac') // Uint8Array
let b = a.clone() // independent copy, shared pages
let src = audio() // pushable source
src.push(buf, 'int16') // feed PCM
src.stop() // finalizePlayback / Recording
Method | Description |
|
|
| take over |
| each ramps over 5 ms, none clicks; |
| mic. |
| the page's one AudioContext, which playback uses: made on first use, resumed by the first gesture; set your own before playing. |
Playback renders up to 2 s ahead into an AudioWorklet on audio.context (Node: @audio/speaker), so a busy main thread doesn't stop it, and sounds within milliseconds of play() (the device's own latency aside). An edit to the playing instance is heard ~50 ms later where it happens: the audio rendered ahead gives way, crossfaded. A source still arriving (decoding, pushed) plays what has come and goes on as more comes. Any channel count plays as it is; the device downmixes.
a.play({ at: 30, duration: 10 }) // play 30s–40s
await a.played // wait for sound
a.volume = 0.5; a.loop = true // live adjustments
a.muted = true // mute without changing volume
a.playbackRate = 1.5 // faster, the pitch kept
a.preservesPitch = false // the pitch follows the speed, as a tape's
a.pause(); a.seek(60); a.resume() // jump to 1:00
a.highpass(80) // an edit while playing: heard where it happens
b.play({ from: a }) // b takes over at the same place, crossfaded
b.stop() // end playback or recording
await audio.context.audioWorklet.addModule('./scrub.js') // your own nodes, on the same context
let scrub = new AudioWorkletNode(audio.context, 'scrub')
let mic = audio()
mic.record({ sampleRate: 16000, channels: 1 })
mic.stop()Metering
Method | Description |
| live per-block stats of what plays, delivered as it is heard: |
Option | Description |
| stat name, array of names, or omit for all block stats. |
|
|
| one-pole EMA time constant τ, in seconds. |
| peak-hold decay τ, in seconds. |
| spectrum resolution and range (when |
a.meter('rms', v => draw(v)) // scalar avg across channels
a.meter(['rms', 'peak'], v => draw(v)) // { rms, peak }
a.meter({ type: 'rms', channel: [0, 1] }, v => draw(v)) // [L, R]
a.meter({ type: 'spectrum', bins: 64, smoothing: 0.15 }, drawFFT) // Float32Array of mel bins
a.meter({}, ({ delta, offset }) => draw(delta)) // no type → all block stats
let m = a.meter({ type: 'rms' }) // pull form
requestAnimationFrame(function tick() { draw(m.value); requestAnimationFrame(tick) })
m.stop() // releaseAnalysis
Method | Description |
| one value; with |
|
|
| pass or fail against a delivery spec: |
Stat | Description |
| peak amplitude in dBFS. |
| RMS amplitude, linear (the CLI prints dBFS). |
| RMS of the quietest 0.4 s, dB: the room between words (ACX Check's measure, sample-exact). |
| the noise print of a range, as |
|
|
| integrated LUFS (ITU-R BS.1770-4; surround channels weighted, LFE excluded). |
| maximum 400 ms / 3 s loudness, LUFS (EBU Tech 3341). |
| loudness of the speech only, LUFS: speech found automatically (AES TD1008 dialog loudness). |
| DC offset. |
| clipped samples, at 16-bit full scale (±32767/32768) or beyond (scalar: timestamps, binned: counts). |
| silent ranges as |
| peak/RMS in dB. Sine ≈ 3dB, square ≈ 0dB. |
| spectral centroid in Hz (brightness). |
| spectral flatness: 0 tonal, 1 noise. |
| L/R phase correlation, −1 to +1. Mono returns 1. |
| peak envelope per bin, for waveforms. |
| mel spectrum in dB (A-weighted); of several channels, their mean power. |
| MFCCs. |
| tempo. |
| timestamps as |
| where the level jumps, up (a strike) or down (a stop), as |
|
|
|
|
|
|
Opts: bpm, beats, onsets take { minBpm, maxBpm, delta, frameSize, hopSize }; notes takes { minFreq, maxFreq, frameSize, hopSize, minDuration }; chords, key take { frameSize, hopSize, tuning } (frames of 16384 samples at 44.1 kHz, 0.34 to 0.51 s at other rates, every eighth of a frame; concert A read from the audio unless tuning in Hz is given); chords also boostN (no-chord bias, 0.1); key also method: 'nnls' | 'pcp'. chords needs @audio/mir-nnls-chroma and @audio/mir-chordino, key needs @audio/mir-nnls-chroma: GPL-2.0-or-later translations of the reference plugins, installed by choice (npm i @audio/mir-nnls-chroma @audio/mir-chordino); key with method: 'pcp' needs only the MIT @audio/mir-chroma and @audio/mir-key, installed with audio unless optional dependencies are skipped. notes with robust: true needs @audio/neural-pitch (weights inside): a network's pitch candidates in place of YIN's keep the notes where YIN loses them (Vocadito onsets F 0.76 against 0.53 at 0 dB SNR) and trail it slightly on clean audio, so YIN stays the default. notes with poly: true takes { minFreq, maxFreq, minDuration, onsetThreshold, frameThreshold } and needs @audio/neural-transcribe, whose model downloads on first use; bends are cents from the note's pitch per 11.6 ms frame, in 33.3-cent steps (in-tune notes read 0).
let loud = await a.stat('loudness') // LUFS
let [db, clips] = await a.stat(['db', 'clipping']) // multiple at once
let spec = await a.stat('spectrum', { bins: 128 }) // frequency bins
let [min, max] = await a.stat(['min', 'max'], { bins: 800 }) // peak envelope for canvas rendering
await a.stat('rms', { channel: 0 }) // left only → number
await a.stat('rms', { channel: [0, 1] }) // per-channel → [n, n]
let gaps = await a.stat('silence', { threshold: -40 }) // [{at, duration}, ...]
let bpm = await a.stat('bpm') // 120.5
let beats = await a.stat('beats') // Float64Array [0, 0.5, 1, ...]
let { bpm, confidence, beats, onsets } = await a.detect() // full pipeline, one pass
let notes = await a.stat('notes') // [{time, duration, freq, midi, note: 'A4', clarity}]
let chords = await a.stat('chords') // [{time, duration, label: 'Am', confidence}]
let k = await a.stat('key') // {label: 'C', mode: 'major', confidence}Meta
Property | Description |
| tags: |
| cover art |
|
|
| a marker at |
|
|
Parsed on decode, written on save; round-trips WAV, MP3, FLAC.
let a = await audio('song.mp3')
a.meta.title // 'Track Name'
a.meta.artist = 'Me' // mutate
img.src = a.meta.pictures[0].url // lazy Blob URL
a.crop({ at: 10, duration: 30 })
a.markers // re-projected — outside markers dropped, inside shifted
await a.save('edited.mp3') // tags + pictures preserved
await a.save('stripped.wav', { meta: false }) // opt outUtility
Method | Description |
| subscribe / unsubscribe. |
| returns the undone edit, for redo via |
| apply |
| release resources. Supports |
Event | Description |
| pages decoded/pushed. Payload: |
| any edit or undo. |
| stream header decoded. Payload: |
| playback position, as heard (~50 times a second). Payload: |
| playback started or resumed. |
| playback paused. |
| volume or muted changed. |
| playback ended: at its end, by |
| during save/encode. Payload: |
a.on('data', ({ delta }) => draw(delta)) // decode progress
a.on('timeupdate', t => ui.update(t)) // playback position
a.run(
['gain', { value: -3, at: 10, duration: 5 }],
['crop', { at: 1, duration: 2 }],
['fade', { in: 1, curve: 'exp' }],
['insert', { source: ref, at: 2 }],
)
a.undo() // undo last edit
b.run(...a.edits) // replay onto another file
JSON.stringify(a); audio(json) // serialize / restorePlugins
Method | Description |
| register an @audio contract factory, a stat |
| register an op: a |
| one descriptor, or all ops. |
| register a stat: |
Plugins also run without the engine: audio/batch over a whole signal, audio/stream over live chunks. Plugin tutorial.
import { compressor } from '@audio/dynamics-compressor/audio'
audio.use(compressor) // bring-your-own factory
a.freeverb({ room: 0.8 }) // registry plugin: loads on first render; tail composes
music.ducker({ key: voice }) // sidechain via the key option
await a.stat('truepeak') // stat plugins land on a.stat()
audio.op('crush', { params: ['bits'], process: (input, output, ctx) => {
let steps = 2 ** (ctx.bits ?? 8)
for (let c = 0; c < input.length; c++)
for (let i = 0; i < input[c].length; i++)
output[c][i] = Math.round(input[c][i] * steps) / steps
}})
a.crush(4) // custom op, chainable like built-insWorker
Call | Description |
| same API, engine in a Worker; the main thread keeps a few-KB facade. |
| same, once |
| your own worker entry: codecs, plugins, your own code and messages, then |
| hand an instance your worker made to the page as a facade. |
| the page's AudioContext, the same as |
Across the boundary clip(), split(), clone() return promises; op errors emit 'error'; functions don't cross, use {t, v} curves. play() renders in the worker straight into the page's AudioWorklet: the main thread can stall for seconds without a dropout. Architecture.
import audioWorker from 'audio/worker'
let a = audioWorker('track.mp3') // decode/edits/stats/encode in a Worker
a.gain(-3).fade(0.5)
let [mins, maxs] = await a.stat(['min','max'], { bins: 640 }) // transferred, zero-copy
a.play() // rendered in the worker, played by an AudioWorklet (Node: @audio/speaker)
// your own worker: its messages stay its own, and what it makes plays on the page
import audio from 'audio' // worker.js
import { expose } from 'audio/worker'
self.onmessage = ({ data }) => self.postMessage({ out: expose(audio(data.file).gain(-3)) })
let worker = new Worker('./worker.js', { type: 'module' }) // page
worker.onmessage = ({ data }) => audioWorker.adopt(data.out, { worker }).play()CLI
npm i -g audio, or without installing: npx audio …
audio [source] [transforms...] [sink] [options]A pipeline: a source produces audio, transforms reshape it, a sink consumes it. The default sink is stat — printing an overview.
# sources
FILE path, URL, or glob ('*.wav' for batch)
- stdin (or omit when piping)
record capture from microphone
# transforms (chained left-to-right)
gain fade trim normalize crop
clip remove reverse repeat pad
speed stretch pitch insert mix
crossfade remix pan split resample
highpass lowpass eq lowshelf highshelf
notch bandpass allpass vocals dither
crossfeed shrink crossover match spectral
repair copy cut paste
# sinks (terminate the chain — at most one)
stat [NAMES...] print analysis (default)
play [loop] open player UI
save PATH encode and write (or `-` for stdout); `192k` bitrate, `24bit` depth
# options
-f --force overwrite existing output
--format FMT override output format
--macro FILE apply edits from JSON
--cue FILE split at cue-sheet tracks (with split)
--verbose show progress
--help, -h help (or per-op: `audio gain --help`)
--mcp serve the CLI to AI agents as an MCP tool (stdio)
# named options, after an op or sink: name:value
normalize -27 lufs ceiling:-2 ducker key:voice.wav save out.m4a codec:alac
# compatibility shortcuts
-p ⇔ play -l ⇔ play loop -o PATH ⇔ save PATHPlayback
␣ pause · ←/→ seek ±10s · ⇧←/⇧→ seek ±60s · ↑/↓ volume · l loop · s save as · q quit
# play full song
audio song.mp3 play
# play fragment
audio song.mp3 10s..15s play
# play and loop a hook
audio song.mp3 30s..45s play loop
# play with effects applied live (streamable ops)
audio song.mp3 normalize broadcast highpass 80hz playEdit
# clean up
audio raw-take.wav trim -30db normalize podcast fade 0.3s -0.5s save clean.wav
# scope a range (applies to whole chain)
audio in.wav 1s..10s gain -3db save out.wav
# range on a single op
audio in.wav gain -3db 1s..10s save out.wav
# filter chain
audio in.mp3 highpass 80hz lowshelf 200hz -3db save out.wav
# concat
audio intro.mp3 + content.wav + outro.mp3 trim normalize fade 0.5s -2s save ep.mp3
# crossfade into next
audio track1.mp3 crossfade track2.mp3 2s save mixed.wav
# voiceover
audio bg.mp3 gain -12db mix narration.wav 2s save mixed.wav
# music bed ducked under the voice (sidechain)
audio bed.mp3 ducker key:voice.wav mix voice.wav save episode.wav
# loudness to any target, true peak held; delivery settings
audio book.wav normalize -20 lufs save book.mp3 192k
# fix a video's sound, keep the picture
audio talk.mp4 highpass 80hz 4 normalize podcast save talk.clean.mp4
# master to a reference track (tone, width, loudness)
audio mix.wav master reference.wav save master.wav
# shorten pauses; the same cuts as an EDL for the video editor (.fcpxml, .otio too)
audio talk.mp4 shrink 0.3 save talk.edl
# split
audio audiobook.mp3 split 30m 60m save 'chapter-{i}.mp3'
audio album.wav split --cue album.cue save '{i} - {title}.mp3' # cue-sheet tracks, tagged
# record
audio record 30s save voice.wavAnalysis
# overview (default sink)
audio speech.wav
# range overview — `audio FILE 0..10s` ⇔ `audio FILE stat 0..10s`
audio speech.wav 0..10s
# specific stats
audio speech.wav stat loudness rms
# tempo / beat grid / onsets
audio track.mp3 stat bpm
audio track.mp3 stat beats onsets
# loudness to spec: integrated, max momentary / short-term, speech only, true peak
audio mix.wav stat loudness momentary shortterm dialog truepeak
# pass or fail against a delivery spec (exit 1 on a fail); --json for scripts
audio episode.wav normalize podcast check podcast
audio chapter.wav check acx --json
# pitch / chords / key
audio song.mp3 stat notes
audio song.mp3 stat chords
audio song.mp3 stat key
# spectrum / cepstrum with bin count
audio speech.wav stat spectrum 128
audio speech.wav stat cepstrum 13
# stat after transforms (transforms apply, then stat)
audio speech.wav gain -3db stat dbBatch
audio '*.wav' trim normalize podcast save '{name}.clean.{ext}'
audio '*.wav' gain -3db save '{name}.out.{ext}'
audio 'chapters/*.mp3' check acx # a whole audiobook: one line per chapterStdin/stdout
cat in.wav | audio gain -3db save - > out.wav
curl -s https://ex.com/speech.mp3 | audio normalize save clean.wavTab completion
eval "$(audio --completions zsh)" # add to ~/.zshrc
eval "$(audio --completions bash)" # add to ~/.bashrc
audio --completions fish | source # fishFAQ
Built with
decode – codec decoding (13+ formats)
encode – codec encoding
filter – filters (weighting, EQ, auditory)
speaker – audio output
mic – audio input
pitch – pitch, chord, key analysis
audio-type – format detection
pcm-convert – PCM format conversion
This server cannot be deployed
Maintenance
Related MCP Connectors
Privacy-first audio intelligence: BPM, key, waveform. Audio never stored. Pay per second.
Transcript-based audio editing: transcribe audio, edit by word ID, export edited audio.
Trim, convert, resize, compress, and remix audio and video.
Audio features + harmonic set-building for tracks by name/ISRC. Spotify audio-features replacement.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables AI models to analyze audio files through numerical fingerprints, pitch tracking, and visual spectrograms without requiring direct audio playback. It provides tools for comparing audio iterations and detecting patterns using token-efficient analysis operations.14-
- AlicenseNot gradedqualityCmaintenancePrivacy-first audio intelligence: BPM, key, waveform. Audio never stored. Pay per second.MIT
- FlicenseNot gradedqualityBmaintenanceAnalyzes audio files to extract exact, reproducible measurements like loudness, tempo, key, spectral balance, and clipping for LLM-based DAW control.-
- AlicenseAqualityAmaintenanceRender, analyze, and verify audio (WAV or FLAC) through a fully offline, deterministic engine, exposed as MCP tools for AI agents.1214Apache 2.0