Docs /
skills/showtime/references/sound-design.md Sound design: density, families, alignment, ducking and loudness targets
Read this when you place sound effects, balance music against a voice, or check whether a mix is
ready. It gives numbers to aim for and how to verify them with showtime audio meter and
mix.report.json, instead of guessing.
1. How many sounds#
- One designed sound per visual event that matters. Not every element that moves gets a sound.
- Density by tone (on average):
| Tone |
Effects |
| deadpan / minimal |
1–2 cues in the whole video |
| polished product |
a cue every 2–3 s, all subtle |
| cinematic |
2–4 big moments (riser into impact), little else |
| playful / chaotic |
up to 1 per second, short and light |
| explainer, data story, report, math |
1–3 in the whole video: tick, thock or a soft chime on the payoff |
- Whooshes only for moves that travel far or change scenes.
- Text pops: tick or pop on 20–40 % of them (the first, the last, the important one), not on
every word.
- Staggered lists: accent the first and the last item.
- One family per video: the same whoosh (different seeds or variants), the same click set, and
tonal effects in the music's key. A mixed bag of sources sounds cheap.
- Game-UI sounds read as cheap in serious work:
pop, ding, success, sparkle-up, coin,
power-up belong to playful pieces. A proof, a chart or a report gets tick, thock, a low
chime a few dB under the bed, or nothing.
- Silence is a sound. Let the beat before the reveal breathe (0.25–0.5 s near-silence).
- Room tone: put a quiet bed (
synth: room-tone or a library ambience at about -32 LUFS) under
talking-head cuts, so the gaps don't sound like dropouts.
2. Levels (what level: auto does, and why)#
| Element |
Reference level |
Notes |
| Voice-over |
-16 LUFS integrated |
then the master brings the whole mix to -14 |
| Music alone (intro, outro, no voice) |
-20 LUFS integrated |
gain_db +2..+6 for a music-driven video |
| Music under voice |
≈ -32 LUFS short-term |
the reference minus a 10–14 dB duck (default 12) (plus an optional carve) |
| Ambience |
-32 LUFS |
should be felt, not heard |
| UI (click, tick, pop, toggle) |
-24 LUFS momentary max, peaks ≤ -8 dBFS |
12–18 dB under the voice |
| Foley |
-22, peaks ≤ -6 dBFS |
|
| FX (glitch, laser, zap) |
-20 |
|
| Transitions (whoosh, riser) |
-18 |
|
| Musical (stinger, logo sting) |
-16 |
|
| Impacts |
-14 |
the loudest effect: one per video moment, not per cut |
- Procedural effects come out at these levels. Library effects are brought there when they enter
a mix.
gain_db then says "a bit more" or "a bit less" (-6 to +4 is normal).
- A very short click is quieter on a loudness meter than it sounds. Its peak cap keeps it from
poking out.
- Harsh sounds (
harshness: high in the catalog: lots of energy above 5 kHz) are tiring when
repeated. Prefer low/medium for anything that repeats.
- Soft UI clicks 12-16 dB under a voice are "likely masked" by design: mark them
"texture": true in
the mix so the report stops warning. Key clicks under a page typewriter come from a
{"typewriter": {...}} mix track (references/audio.md), not hand-placed times.
3. Alignment#
- The hit is the moment that matters, not the file start:
- an impact's transient
- a whoosh's loudest pass
- where a riser or reverse cymbal lands
align: "hit" with at: T puts the hit exactly on T. Procedural effects and library items carry
their hit. For other files, give "hit": seconds or use "align": "peak" for whooshes.
- Frame math: at 30 fps, frame N starts at N / 30 s. Put hits on the frame where the visual
event is first fully visible.
- An entry pop can lead the visual by 0–1 frame.
- Never let sound trail the picture by more than 2 frames (66 ms): it reads as out of sync.
- Sound 1–2 frames before a cut reads as anticipation, which is fine for whooshes into a cut.
- Riser → impact:
- A riser (
hit = its end) lands on the reveal downbeat.
- The impact starts on the same downbeat.
- A reverse cymbal or reverse-hit sucks into it.
- Make the riser 1–2 bars long: at 120 bpm a bar is 2 s.
- Transitions: a transition sound starts at the transition start. A success chime plays when
the result is fully visible, not when it starts to appear.
4. Ducking and carving#
| Setting |
Default |
Range |
| depth |
12 dB |
10–14 dB (up to 16 for dense music) |
| attack |
80 ms, with a 120 ms look-ahead (the duck is complete when the first word starts) |
|
| hold |
300 ms across pauses inside a sentence |
|
| release |
0.5 s |
0.4–0.8 s |
- Longer releases (up to 2 s) make the bed swell back only between sentences, not between words.
- Carve (
carve: 0.3–0.5) cuts the bed only where the voice lives (chosen bands, weighted
toward 1–4 kHz) and only while it speaks. The music keeps its low end and air, so it sounds
fuller than a deep flat duck.
- Combine a moderate duck (6–8 dB) with a carve.
- A hollow, phasey bed means the carve is too strong.
- Check it:
voice_to_music_db in mix.report.json should be 10–20 dB.
- Below 8, the report warns: words get masked on phone speakers.
- Above 25, the music sounds switched off. Raise the bed's
gain_db.
- Real footage with its own sound: duck the music under the footage's dialogue the same way.
Use the footage audio as a
voice track.
5. Mastering targets#
| Destination |
Integrated |
True peak |
Notes |
| Web, YouTube, X, LinkedIn, Reels, TikTok, Shorts |
-14 LUFS |
-1 dBTP |
house default; platforms turn louder masters down, and quieter ones just play quiet |
| Podcast / audio-first |
-16 LUFS |
-1 dBTP |
--target podcast |
| Broadcast / event screens |
-23 LUFS (EBU) or -24 LKFS (US) |
-1 / -2 dBTP |
only when asked |
| Music bed delivered separately |
-18 LUFS |
-1 dBTP |
--target music-bed |
- Loudness range (LRA) up to about 8 LU plays well on phones. Above 12, the quiet parts vanish on
small speakers.
- The limiter should do a little (0–4 dB). The report's
limiting_db above 6 means something
spiky (usually an impact) is too hot: lower that track instead.
- Master to WAV and encode once. AAC and MP3 add up to 0.5 dB of overshoot.
showtime deliver exports re-normalises every platform version.
6. Measure instead of guessing#
An agent cannot listen. These are the checks that stand in for ears:
showtime audio meter final.mp4 --ffmpeg
- Pass: integrated -14 ± 0.5 LUFS and true peak ≤ -1 dBTP, from our meter and ffmpeg's.
- Any clipped runs mean fail.
mix.report.json:
voice_to_music_db is 10–20.
- No warnings.
- The per-section
lufs follows the story: the quiet intro is quieter and the drop is louder.
A flat line means no dynamics. Jumps over 6 LU between neighbouring sections feel like
mistakes.
- Per-bus RMS per section (in the report): is the SFX bus louder than the voice anywhere? Then an
effect is covering words. Move it or lower it.
audio meter file --windows 1 shows RMS per second. Look for unexpected silences
(below -50 dBFS) and spikes.
- For effects: check the peak (
peak_db), the harshness and hit in the catalog or the .sfx.json.
- Then a human listens on laptop speakers and headphones before publishing. Say so when you
hand over the video: the numbers are necessary but not sufficient.
View or edit this page on GitHub