showtime
Docs / skills/showtime/references/footage-tools.md

Footage tools

Read this when you need one specific tool for real footage outside (or inside) an EDL edit: shot detection and contact sheets, face-tracked reframing, speech denoise, stabilisation, colour correction and looks, timeline views, media facts. The editing workflow itself is in editing.md; captions are in captions.md.

Every command has --help with examples and --json for machine-readable results, prints the paths it wrote, and never overwrites an earlier output (a -2 suffix is added). Default outputs never go beside the user's media: footage scenes writes to <job>/work/scenes/, and reframe, grade (and its --compare still), denoise, stabilize, luts --preview and view <media> to <job>/work/footage/. The job is --job where offered, else the job the file or the current folder is in, else the newest job; with no job at all, a fresh showtime-out/<tool>-<time>/. A source already inside a job keeps its outputs beside it; -o picks any path. transcribe writes to <job>/edit/transcripts/ the same way (with no job it makes a <name>-edit job rather than writing beside the footage); older <media folder>/edit/transcripts/ files are still found by edit cut and footage view. --edit-dir picks any folder, e.g. next to the footage when the user asks.

Inventory: probe and scenes#

showtime footage probe raw/take1.mp4           # size, display size, rotation, fps, VFR, HDR, audio tracks
showtime footage scenes raw/broll.mp4          # shots -> <job>/work/scenes/broll.scenes.json + labelled sheet PNG
showtime footage scenes raw/take1.mp4 --every 5 --job talk-edit   # one frame every 5 s (quick look at any clip)

For a VP9/VP8 webm with alpha (an overlay rendered with --alpha), probe reports pix_fmt: yuva420p with pix_fmt_stream: yuv420p and alpha: side channel: the alpha plane lives beside the colour planes, so ffmpeg's native decoder (and a plain snap) sees only the colour; libvpx decodes the alpha. ProRes 4444 reports its real format (yuva444p12le).

Shot detection uses PySceneDetect's adaptive detector (hard cuts, and fades reported at their end) and falls back to ffmpeg's scdet. --min-len (0.6 s) merges flicker; --threshold tunes it. Read the contact sheet PNG to see what is in the footage before planning B-roll or cutaways. A periodic sheet holds at most 60 frames: --every 0.25 --from 12 --to 16 looks closely at a range (the command says when it had to spread frames out); -o sheet.png names the file.

Trim and proxies#

showtime footage trim interview.mp4 --from 61 --to 142 -o <job>/work/excerpt.mp4   # frame-accurate excerpt
showtime footage trim demo.mp4 --width 1280 -o project/media/demo.mp4             # seek-friendly proxy for a page
showtime footage trim demo.mp4 --webm --no-audio -o project/media/demo.webm       # VP9 for Chromium without H.264

Re-encodes (H.264 + AAC, +faststart, a keyframe every --gop 0.5 s) unless --copy (instant, starts on a keyframe). Use it instead of a one-range EDL or bare ffmpeg.

Talking-head footage for a "cut the ums" demo: resource reels and press soundbites are already edited (no fillers left to cut); raw interview recordings and live shots are where fillers are.

Timeline views (self-review)#

showtime footage view take1.mp4 --from 12 --to 20            # filmstrip + waveform + words + pauses
showtime footage view take1.mp4 --from 12 --to 20 --mark 15.4
showtime edit view edit/edl.json                              # +-1.5 s around every cut of the render

Pauses of 0.4 s or more are shaded (with their length when 0.6 s or more), fillers are drawn in red, audio events in amber, cut points as red lines. The transcript is found automatically.

Reframe (16:9 -> 9:16, 1:1, 4:5)#

showtime footage reframe talk.mp4 --aspect 9:16               # face-tracked crop
showtime footage reframe talk.mp4 --aspect 1:1 --zoom 1.15    # plus a punch-in
showtime footage reframe talk.mp4 --aspect 9:16 --focus 0.3,0.5   # fixed crop centre, no tracking
showtime footage reframe screen.mp4 --aspect 9:16 --fit blur  # keep the whole frame on a blurred copy

Tracking: YuNet face detection at 8 fps, the largest confident face (preferring the one nearest the previous pick), gaps filled, a dead zone of 3.5 % of the width so small moves do not pan, a 1 s centred average, and hard shot changes respected (no pan across a cut). The face sits slightly above centre. No face found -> centre crop (or focus). Inside an EDL the same happens per range with "fit": "reframe" (or auto), and zoom > 1 punches in around the face.

Use blur (or contain) for screen recordings and slides: cropping cuts off UI. A crop that enlarges the source more than 1.5x (a 1280x720 clip cropped to 9:16 and scaled to 1080x1920 is 2.67x) looks soft: the command warns and suggests a smaller standard size (720x1280, else 540x960) or --fit blur.

Denoise speech#

showtime footage denoise interview.mov -o interview.clean.mov        # auto: best available
showtime footage denoise vo.wav -o vo.clean.wav --strength 0.7

auto picks DeepFilterNet (installed with showtime setup --with deepfilter) if present, else RNNoise (arnndn, models cb general / sh stronger / std), else ffmpeg afftdn. A 70 Hz high-pass runs first. strength below 1 keeps some room tone (sounds more natural): for DeepFilterNet it caps the reduction (0.5 = at most 21 dB), for RNNoise it is the wet mix, for afftdn the reduction amount. The report line names that, the noise floor before/after and the level of what was removed (the floor alone barely moves on audio with little steady noise, so two strengths can look alike there; the removed level differs). The output keeps the source's channel count (stereo stays stereo) and its exact length in samples (DeepFilterNet's ~30 ms frame delay is padded back), so a cleaned track still lines up with the picture. In an EDL use "audio": {"denoise": "auto"}: each source is cleaned once and cached.

Do not denoise clean studio audio; it can only lose detail. Denoise before loudness mastering (the renderer does this order for you).

Stabilise#

showtime footage stabilize walk.mp4 --strength 0.7 -o walk.stable.mp4

Two-pass vid.stab (motion analysis, then smoothing with automatic border-hiding zoom and a light sharpen). Builds without vid.stab (some ffmpeg builds, for example on Apple Silicon) fall back to deshake; the output says which was used. In an EDL: "stabilize": true on a range.

There is no --compare here on purpose: a before/after still cannot show shake. Judge motion by watching both, or compare the framing (the stabiliser zooms in to hide moving borders) with showtime snap walk.stable.mp4 --at 2,5 --compare walk.mp4 (writes a before|after compare.jpg).

Colour: correction and looks#

showtime footage grade take1.mp4 --analyze                     # measured stats + the auto correction
showtime footage grade take1.mp4 --auto --look teal-orange --strength 0.5 -o take1.graded.mp4
showtime footage grade take1.mp4 --preset punch --compare --at 4     # before/after PNG (-o must be .png/.jpg)
showtime footage luts                                          # list looks
showtime footage luts --preview take1.mp4 --at 5               # every look on one frame (PNG)
  • Order: correct first (auto: exposure via gamma, contrast when flat or foggy, saturation when dull; each change bounded), then a look at 40-70 %, then the BT.709 output conversion.
  • Presets (eq/curves): subtle punch warm cool film mono webcam lowlight.
  • Looks are original 33-point .cube LUTs generated by showtime (no third-party files): teal-orange warm-film clean-punch cool-tech mono-contrast bleach golden-hour matte night. The strength is baked into the LUT. Your own .cube works too (--look my.cube --strength 0.6).
  • Grade real footage only. UI captures and motion graphics keep their design colours.
  • In an EDL: "grade": "auto", "grade": "warm-film", or {"auto": true, "preset": "subtle", "lut": "teal-orange", "strength": 0.5}; per range too.

Screen recordings: auto zoom#

showtime footage autozoom ... forwards to the capture module's click-driven zoom (see tutorial-recording.md).

Platform notes#

  • All tools run on macOS (Apple Silicon and Intel), Windows and Linux. ffmpeg features differ by build; every optional filter has a fallback (vid.stab -> deshake, arnndn -> afftdn, zscale missing -> HDR shown without tone mapping plus a warning, lut3d missing -> look skipped with a warning). showtime doctor lists what your ffmpeg has.
  • Speech recognition, diarization, events and face tracking run on the CPU through CTranslate2 / ONNX Runtime; no GPU or account is needed.
  • Heavy steps (transcription, renders) use all cores: run them one after another, not in parallel with browser captures.

View or edit this page on GitHub