skills/showtime/references/audio.mdAudio tools: compose, effects, library, beats, mix, master
Read this when a video needs music, sound effects, a voice-over mix or a loudness check, and you
need the exact showtime audio commands and the audio/mix.json format. For choosing music by
tone, read music.md. For how loud, how many and where sounds go, read sound-design.md. For
music written in code inside a canvas film, read synth-score.md.
Everything runs locally. Every command has --help with examples, and most accept --json.
1. The five-minute path#
showtime audio music pick --for launch --dur 45 # a produced track (fetched when first mixed)
showtime audio lib fetch # optional: the whole ~249 MB library (~10-15 min); the starter part comes by itself
showtime audio compose --style underscore --dur 45 \
--sections 0:intro,6:verse,24:break,40:outro -o audio/bed.wav # which style: music.md
showtime audio sfx whoosh -o audio/whoosh.wav # prints JSON with "hit"
showtime audio mix audio/mix.json -o audio/mix.wav # mastered to -14 LUFS / -1 dBTP + report
showtime audio meter audio/mix.wav # verify, never guess
When a project's showtime.json has "audio": "audio/mix.json", showtime render runs the mix
itself. Run audio mix on its own while iterating: it takes seconds, and the report shows what
changed.
2. Music#
| Need | Command |
|---|---|
| A produced track for a launch, trailer or emotional piece | audio music pick --for launch --dur D (or search, info <id>), then "catalog": "<id>" in the mix |
| Music cut exactly to the edit, with section changes on picture changes | audio compose --style S --dur D --sections "t:name,..." |
| A real recording from the library | audio lib search --kind music --mood calm --min-dur 60 |
| A library track at exactly the video's length | audio fit track.opus --dur 42.5 -o bed.wav (bar-aligned loop or musical ending; --ending song keeps the track's own ending) |
| A short film cut to a produced track: the excerpt and scene changes on its phrases | audio cuts --for launch --dur 30 --scenes 5 (--apply <project> retimes the scenes and sets the track) |
| The beat grid of any track | audio beats track.mp3 (writes track.beats.json) |
| Temp music from a text prompt (non-commercial) | audio musicgen "prompt" --dur 20 -o draft.wav (optional tier) |
The produced-music catalog (lib/st/audio/music_catalog.json, 150+ tracks: Scott Buckley and
Kevin MacLeod under CC BY 4.0, CC0 and public-domain recordings; every track was downloaded, measured
and passed a quality gate: no muffled old transfers, hiss, crackle, clipping, mono, dead gaps or wild
loudness swings, thresholds in music.py):
audio music search [words] --for USE --shelf S --mood M --energy 0.2-0.5 --dur D --vocals noneranks tracks;pickreturns the best one (--n 1the next);presetslists the uses;statscounts. Shelves: cinematic, inspiring, ambient, corporate-tech, upbeat, documentary, tension, playful, lofi, piano, orchestral.- Nothing is bundled. A mix (or
audio music fetch <id>, orsetup --full) downloads the file from the creator's site on first use, announces it with its size, checks the pinned sha256 and caches it in~/.showtime/music($SHOWTIME_MUSIC_CACHE). Cached tracks work withSHOWTIME_OFFLINE=1;--seed DIRcopies from a folder of earlier downloads. - A catalog track skips its near-silent lead-in by default. For a short cut of a long track that builds
slowly,
"offset": "highlight"starts at its loudest sustained stretch (on a downbeat;music infoshows where), or give seconds. audio music veto <id>hides a track frompick(--undo,--import vetoes.json).audio music openverse <words>is a live CC BY / CC0 search of Freesound and Wikimedia Commons beyond the catalog (about 20 searches a minute, 200 a day);--fetch openverse:<id> -o <dir>saves one with its.license.json. Jamendo is left out unless you pass--source jamendo: its own terms add conditions to commercial use and its licensing program can claim videos, so check the artist's terms first.
audio compose details:
--styleis one of 18 styles (audio styles; defaultunderscore).--bpmand--keydefault to the style's own.--sections "0:intro,8:build,16:drop,26:outro". Names: intro, verse, build, drop, chorus, break, bridge, outro. Every marker lands exactly on a downbeat. The tempo is nudged per section (usually under 2 %, up to about 6-7 % for very short sections) and a section may get a 2-beat bar to make that happen. Above a 3 % bend, or with a 2-beat bar, compose warns and prints a--bpmthat puts every marker on whole bars (when one exists near the asked tempo), or the same markers moved to whole bars at the asked tempo. A marker inside the final hit's ring-out (for example27:outroin a 30 s bed whose hit lands at 27.14 s) merges into the ending, with a one-line note.- The piece ends with a final tonic hit, snapped to the grid inside a short ring-out. The file is
exactly
--durseconds (sample-exact).end_hitin the output is the time of that hit: put the logo there. - Transitions come with it: a riser that lands on each drop, reverse cymbals into energetic
sections, an impact on drops for cinematic styles, and a downlifter into outros (
--no-sfxremoves them). - Outputs:
bed.wav(mastered, -14 LUFS),bed.mid,bed.beats.json(beats, downbeats, bars, sections, events,end_hit, thebackendandsoundfontfile actually used, the plannotes),bed.stems/{drums,bass,pad,arp,lead,keys,sfx}.wav. --backend auto|sf|synth|hybrid: SoundFont instruments (GeneralUser GS), the numpy synth, or synth drums/bass with SoundFont keys and strings. Iftinysoundfontor the SoundFont is missing, it falls back to the synth automatically. It still works, it just sounds more electronic.--seed Ngives a different melody and variation with the same structure. Try 2 or 3 seeds.
3. Sound effects#
audio sfx-typeslists 56 procedural types in 7 categories, with their hit kind:onset: the sound hits at its start (impacts, clicks, bells).peak: the loudest moment (whooshes, sparkles).end: the build-up lands on the last sample (risers, reverse cymbals, swells, sparkle-up).
audio sfx <type> [--dur --key --intensity 0..1 --seed N] -o file.wavwritesfile.wavandfile.sfx.json, and prints the JSON:hit(seconds),hit_frame_30fps,category, and the level.audio mixreads the sidecar, so"align": "hit"uses the true hit automatically.- Tonal types (riser, swell, notification, success, bell, logo-sting, drone, and others) take
--key. Give them the music's key so effects sit in the harmony. - Output is loudness-matched per category (
sound-design.md§2), sogain_db: 0in a mix is a sensible level. --variants 4renders four seeds (x-1.wav...x-4.wav). Use a different variant for each repeat, so repeated clicks don't sound identical.- Library effects (Kenney, OpenGameArt and generated variants) carry
hit,categoryandharshnessin the catalog.
4. Library#
showtime audio lib fetch [--tier core|extended] [--no-generated] # idempotent; resumes
showtime audio lib search whoosh --kind sfx --distinct # words match ids, titles, tags, moods
showtime audio lib search --kind music --mood uplifting --bpm 100-130 --min-dur 60 --license cc0
showtime audio lib search --kind music --key Am --dur 45 # compatible keys rank higher
showtime audio lib info incompetech-voxel-revolution # details + the exact credit line
showtime audio lib credits <id> <id> -o CREDITS.txt
showtime audio lib stats | sources | generate
showtime audio lib index ~/Sounds/MyPack --name mypack --license "vendor-license" # your own pack, in place
- Library parts: the first command that reads the library fetches the starter part (~41 MB, about 1.5
min: common effects, three ambiences, six stingers, one bed per broad mood). A search that finds
fewer than 3 results fetches the category part it points at (
sfx,ambience,music-upbeat,music-calm,music-epic), once, announced with its size.showtime audio lib fetch --list-partsshows them;--part NAMEfetches one now. A library you pointSHOWTIME_LIBRARYat, or index yourself, is never filled automatically. - Tiers:
core(~249 MB) has every Kenney CC0 audio pack, OpenGameArt CC0 packs and ambiences, about 45 Kevin MacLeod tracks and stings (CC-BY 4.0, stored as Opus), and CC0 music from Komiku, Loyalty Freak Music and Musopen (Chopin).generatedis rendered locally and needs no download: about 160 effect variants and 28 composed beds.extendedadds about 200 MB more music and effects. Core sources are pinned (URL, size, sha256); most extended sources are not pinned yet, so their sha256 is recorded instate.jsonon the first download and every re-download must match it (trust on first use). Maintainers pin them withshowtime audio lib pin --tier extended.
- Licenses: CC0 and generated items need no credit. CC-BY items must be credited.
audio mixwritesCREDITS.txtnext to its output and lists the credits in the report;showtime rendercopies them intocredits.txtbeside the final, which is the file that ships. Paste them into the video description.--license cc0avoids the question entirely.- A library file copied into a project and used as
"file"keeps its license: the mix matches it to the catalog by content (size, then sha256). A file with a<file>.license.jsonsidecar uses that. A musicfilewith neither gets a warning (use"lib": "<id>", or add the sidecar).
- A library file copied into a project and used as
- Kinds and categories:
- Kinds are
music,stinger(short musical hits and jingles),sfx,ambienceandvoice(announcer words). - SFX categories are
ui,transition,impact,foley,fx,musicalandambience. --kind ambiencealso lists sfx items that can serve as a bed (at least 3 s, loopable or named or tagged as a loop, such as a 7 s water loop in a foley pack), ranked below real ambience items and marked "an sfx loop (catalogued as sfx)".
- Kinds are
- Moving the library: set
SHOWTIME_LIBRARY=/pathto keep it elsewhere, for example on an external disk. - Extra sound packs (
audio packs list): about 25 more CC0 packs (paper, keyboards, a typewriter, cloth, coins, UI, impacts, sci-fi, whooshes, city and crowd ambiences) that are not in the core tier. A mix naming one of their items ("lib": "oga-keyboard-typing/...") fetches the pack first;audio lib searchsays when an uninstalled pack matches;audio packs fetch <id> | --category C | --allinstalls ahead of time. Licenses and sources:lib/st/audio/sfx_packs.json.
5. The mix spec (audio/mix.json)#
{ "duration": 20.0,
"tracks": [
{"id": "bed", "kind": "music", "file": "audio/bed.wav", "gain_db": -2,
"fade_out": 1.2, "duck": {"under": "voice", "depth_db": 12, "carve": 0.4}},
{"kind": "voice", "file": "voice/vo.wav", "start": 0.6},
{"kind": "sfx", "file": "audio/whoosh.wav", "at": 3.0, "align": "hit"},
{"kind": "sfx", "synth": {"type": "impact", "key": "D", "intensity": 0.8, "seed": 3}, "at": 17.5},
{"kind": "sfx", "lib": "kenney-interface-sounds/click-001", "at": 5.2, "gain_db": -4, "pan": 0.3},
{"kind": "ambience", "lib": {"search": {"query": "rain", "kind": "ambience"}}, "loop": true, "gain_db": -3}
],
"sections": {"intro": 0, "demo": 4, "cta": 16},
"master": {"lufs": -14, "true_peak": -1}
}
- Sources: each track has exactly one.
file: a path.lib: a library id, or{"search": {...}, "pick": 0}.catalog: a produced-music id ("buckley-with-these-hands"), or{"use": "launch", "mood": "hopeful", "pick": 0}. Fetched on first use and credited; pair it with"fit": true.synth: anaudio sfxspec.compose: anaudio composespec. The duration defaults to the rest of the mix, and the result is cached by content hash.
- Levels:
level: "auto"is the default. Each track is first brought to its kind's reference: voice -16 LUFS, music -20 LUFS, ambience -32 LUFS, sfx by category.gain_dbis then relative to that. Use"level": "raw"to keep a file's own level. - Placement:
start(timeline seconds),offset(skip into the source),endordur.at+align(hit|start|end|peak) puts that point of the sound exactly onat.hitin the track overrides the measured hit (seconds from the file start).
- Length:
loop: trueorfit: truefills the window. Music loops on bar lines when a beat grid is known (a.beats.jsonsidecar, a composed track, or a library item), and ends musically. - Shape:
fade_inandfade_out, pluspanfrom -1 to 1. Every clip edge gets at least 3–5 ms of fade, so there are no clicks. - Ducking:
duck: {"under": "voice" | ["voice", "sfx"] | [track ids], "depth_db": 12, "attack": 0.08, "release": 0.5, "hold": 0.3, "lookahead": 0.12, "carve": 0..1}.- The duck follows the activity of the key tracks: speech for voice, and any other track too (a list of sfx ids ducks the bed under the key clicks). It starts 120 ms early, holds across short pauses and releases slowly.
carvealso cuts the bed only in the voice's own bands, biased to 1–4 kHz, and only while the voice speaks. 0.3–0.5 is transparent; above 0.7 it starts to sound hollow.
- Typing in a recorded demo:
{"kind": "sfx", "keystrokes": "media/events.json", "start": 10.8, "rec_offset": 7.85, "rate": 1.75}clicks once per typed character and key combo of ademo record, in sync with the recording placed atstart(rec_offset= the recording time shown there,rate= itsdata-rate), so no click times are converted by hand. - Typing on a page
typewriter:{"kind": "sfx", "typewriter": {"text": "npm create showtime", "cps": 18}, "start": 2.5}clicks once per typed character (and backspace) on the component's own timeline: give it the sametextorscript,cadence,cps,fitandseedas the page, andstart= the video time the typing starts (its clip's start plusdata-at). - Designed layers and masking: the report warns when an effect is under the rest of the mix. Tracks
of one family (same file stem or explicit id prefix such as
key-01..key-38, the same synth type) never count as "the rest" for each other; a riser whose hit is its end is judged before its end;"layer": "<id>"marks a track stacked on another (a braam on an impact);"texture": trueskips the check for a track meant to sit under (soft UI clicks under a voice). Masked tracks of one family are reported as one line. - Gain automation:
gain_points: [[t, db], ...]on any track (timeline seconds, linear in dB between points, end values hold), e.g.[[0, 4], [2.6, 4], [3.2, 0]]lifts a quiet hook by 4 dB.section_gainon a music bed sets a level per section with 0.25 s ramps:{"intro": 3}(the composed bed's section names) or{"0": 2.5, "5": 0}(seconds). This is the fix for "the intro is too quiet": a short sfx cannot lift a whole section, and renaming a section changes the arrangement. sections: optional names for the report. They default to the composed track's sections, or 5 s windows.master:lufs(default -14) andtrue_peak(default -1).engineisst(exact, the default),loudnorm(ffmpeg two-pass) ornone.- Output formats are
.wav(24-bit),.flac,.m4a,.mp3and.opus.
- Paths resolve against the mix.json folder, then the project root (the folder with
showtime.json), then the current directory.--rootoverrides.
mix.report.json contains:
- integrated LUFS, true peak and LRA, plus the pre-master levels
- per section: LUFS, RMS and the RMS of each bus
voice_to_music_db: how far the voice sits above the music and ambience while speaking (aim for 10–20)- per track: gain, level reference and alignment, plus duck and carve statistics; for a music bed
its landmarks on the output timeline (
end_hit,music_sections,downbeats) and, for a composed bed, thebackend_usedandsoundfont_file(the credit line for the SoundFont): line up the logo withend_hitfrom here, not from the cache; for each sfxabove_bed_db, its loudest 25 ms over everything else (under 0 dB it is probably masked, and the report warns) - the library and catalog items used, their credits (
credits) and full credit items (credit_items: title, artist, license, source, where the credit must go, Content ID status); a CC BY sound without credit text stops the mix with an error - warnings (voice too close to the music, a first section more than 6 LU under the loudest, masked effects, clipping, heavy limiting, non-commercial sources)
- paths relative to the report's folder, so a report can be copied or published
Read the report before you listen. Then listen.
Credits#
showtime render turns the report's credit items into credits.txt beside the video (the exact lines
each license asks for, courtesy lines for CC0, an optional end-card line) and a block in share.txt
between --- Credits (keep in the video description) --- and --- end credits --- (replaced on each
render; the rest of the file is kept). For Scott Buckley tracks it prints a Content ID note: his
Smart Content ID claims YouTube videos whose description lacks the credit. By hand:
showtime audio credits --report <mix.report.json> --out-dir <job>, or
showtime audio credits <id> ... [--end-card] for a quick look.
6. Analysis and delivery#
audio beats filewritesbpm,bpm_confidence,beats,downbeats,onsets(each tagged kick/snare/hat), anenergycurve (0.5 s),sections(energy phases VOID/LOW/MEDIUM/HIGH),moments(SURGE/DROP) andkey.- Trust
rhythmic/pacing: withbeat_cut, hard cuts may sit on beats. Withphrase_flow(ambient or rubato music), the grid is fictional: cut on phrases and energy changes instead. - For composed music the grid is exact by construction (
source: "composed"). - The tempo usually matches metadata within 1–3 %. On-beat vs off-beat phase can still be wrong on dense dance music with offbeat bass, so check the first downbeats on a waveform view when a hard sync matters.
- Trust
audio meter file [--windows 1] [--ffmpeg]gives integrated, short-term max, momentary max, LRA, true peak (8x oversampled), sample peak, RMS, clipped runs and DC.--ffmpegcross-checks with ffmpeg's ebur128.audio master in -o out [--target youtube|podcast|broadcast|...] [--preset mix|music|voice|none]normalises to within 0.05 LU with true peak at or below the ceiling. Master to WAV and encode once at the end: AAC and MP3 add up to about 0.5 dB of overshoot.
7. Platform notes#
- Everything portable: all DSP is numpy/scipy. ffmpeg comes from the resolver, for decoding, Opus encoding and the loudnorm engine, and paths go to it as plain arguments, so Windows drive letters are fine.
- SoundFonts: rendering uses
tinysoundfont. It has ready-made wheels for macOS arm64, Windows and Linux x64. On Intel macOS and Linux arm64 it is compiled at setup; if that fails, compose uses the numpy synth for every part (styles still work). - MusicGen lives in its own venv (
showtime setup --with musicgen). On CPU it runs roughly 10x slower than real time. Its weights are CC-BY-NC, so never use its output in a commercial video. The mix report warns when it is used.