skills/showtime/references/workflows/footage-edit.mdWorkflow: edit real footage (talking head, interview, podcast, vlog)
Read this when the user hands you recorded video or audio and wants it edited: cut the ums and long
pauses, tighten a talking head, pull clips from an interview or podcast, pick the best takes, add
captions, reframe to vertical, clean the audio, fix the colour. Everything is driven by a word-level
transcript; you read text and images, never the raw video. The full reference is editing.md.
Inputs#
- One or more media files or a folder of takes. No path given: list the media files in the working
folder (
*.mp4 *.mov *.mkv *.webm *.m4a *.wav). Exactly one candidate: use it and say so. Ask only when there are none or several. - The goal: length, platform/aspect, what must stay, what must go.
- Helpful: speaker count, names and jargon (for the transcription prompt), a music bed, a brand kit.
Defaults#
Keep the source aspect, frame rate and framing (no punch-ins on cuts unless asked), except for Reels/TikTok/Shorts: --aspect 9:16 --captions bold-pop (face-tracked reframe). Remove fillers, shorten pauses over 0.5 s to 0.3 s; captions
bold-pop for vertical, clean for 16:9; --grade auto only when the footage looks off; denoise only
when noise is audible in the measurements; -14 LUFS / -1 dBTP for every edit (the same as everything else showtime delivers, podcasts and tutorials included; another level only when the user asks: --lufs N, or --keep-loudness to leave the source level). Ask at most:
target length/platform, and anything that must stay or go, and only when the request leaves it open.
Steps#
- Job.
showtime job init <name>-edit --goal "..."(add--platform reels|tiktok|shorts|youtubewhen named; qa then checks aspect, length cap and loudness for it). The printed folder is<job>; the edit files live in<job>/edit/, so<job>works as the argument ofedit renderandedit view. - Inventory.
showtime footage probe <file>per file (rotation, fps, HDR, audio tracks), andshowtime footage scenes <file> --every 5 --job <job>for a visual sheet in<job>/work/scenes/(never beside the user's footage); look at it. Done when: you know durations, orientation, audio tracks and what is on screen. - Transcribe.
showtime transcribe <folder or files> --edit-dir <job>/edit(without--edit-dirthe transcripts go to the job the media or the current folder is in, else a new<name>-editjob, never next to the user's footage; add--speakers 2for interviews,--prompt "names, jargon",--language esfor non-English,--audio-track 1for a mic on track 2). Transcripts are cached by content, so a re-run is instant. The first run loads the model (it printsloading model ...; about 30 s on a mid-range laptop, a fixed cost per run), then progress over the audio. Tell the user the time first for long files. Readguards.warningsin each transcript. - Read the material.
showtime pack <job>/editand readtakes_packed.mdend to end: the hook line, retakes (keep the last clean one), the payoff, dead air, events worth keeping. - Plan in 3-6 lines (keep/cut logic, order, aspect, captions, music, look). Mechanical work (fillers, pauses, captions, reframing, loudness) is stated as an assumption and done. Wait for the user only when the plan drops content: picks between takes, removes sentences, reorders. Done when: the plan is in SHOWTIME.md, confirmed when it drops content.
- EDL.
showtime edit cut <job>/edit/transcripts/*.json -o <job>/edit/edl.jsonwith the options that match the plan (--max-pause 0.5,--remove w40-w52,--keep-time 12.5-40,--aspect 9:16,--captions bold-pop,--grade auto,--music bed.wav,--denoise auto). Hand-edit ranges only with times copied fromtakes_packed.md.showtime edit check <job>(the job's latest EDL, or pass the EDL path): read the plan and the output time of every segment. Re-cutting never overwrites: the new EDL isedl-2.json(printed; from then on<job>inedit render/edit viewmeans it);--overwritereplacesedl.jsonin place. - Draft.
showtime edit render <job> --preview, thenshowtime edit view <job>and read every PNG (words, waveform and frames around each cut). Both print the file they used;edit viewalways shows the newest render (a re-render writespreview-2.mp4;--overwritereusespreview.mp4). Read the render report (edit renderprintsreport <path>and up to six warnings):frames_oktrue, no warnings. A source enlarged more than 1.5x (a 720p clip reframed to 1080x1920) is warned with a fix: a smaller EDL output size such as 720x1280, or--fit blur. Fix and re-render (unchanged segments are reused); at most three passes, then ask. - Final.
showtime edit render <job> -o <job>/final.mp4(this-owins over the EDL folder default inediting.md). Captions are burned last; a matching.srtwith the video's name is written beside it. - Poster.
showtime deliver poster <job> --at <t> --bake(the job's latest final) writesfinal.poster.mp4with the poster in frame 0 (andfinal.poster.png) and records both as the job's latest final and poster; a second bake of the same frame is skipped. The baked file is the deliverable: qa and exports follow it. Reels, TikTok and Shorts let the user pick a cover in the app, so the bake is optional there. - Verify.
showtime qa <job>(the latest final; it prints which, uses the job's platform and checks the job's captionsfinal.srt). Look at the sheet. For anything published, transcribe the output (showtime transcribe <the checked file> --edit-dir <job>/work/qa-transcript) and confirm no fillers or clipped words remain. Say what you could not check (you did not listen). - Deliver. Exports (
showtime deliver exports <job> --targets ...: the latest final), share copy, the delivery card.
Recipes#
| Goal | Additions |
|---|---|
| Tighten a talking head | edit cut --max-pause 0.5, no punch-ins: jump cuts in a tightened talking head are expected and read as clean; a zoom in and back out at every cut reads as a gimmick (edit check flags it). Reframe only on request, one scale held per sentence or section (editing.md section 6) |
| Vertical from landscape | --aspect 9:16 (face-tracked reframe), --captions bold-pop; or showtime footage reframe <file> --aspect 9:16 for a whole clip |
| Podcast clip | transcribe --speakers 2, pick ranges by speaker lines, captions clean (loudness stays at the -14 default unless asked). Long episode: transcribe <file> --from 21:30 --to 24:00 transcribes only that window (times stay on the episode's clock). Published transcript: align it (showtime voice align) and take speakers from its labels rather than diarization (captions.md, section 6) |
| Noisy room | showtime footage denoise <file> --strength 0.7 -o clean.mp4 first, or "audio": {"denoise": "auto"} in the EDL |
| Colour | showtime footage grade <file> --analyze, then --auto, or a look (showtime footage luts); --compare writes a before/after still |
| Shaky phone clip | showtime footage stabilize <file> or "stabilize": true on a range |
| Music under speech | "audio": {"music": "../music/bed.wav"} (ducked automatically); compose one with showtime audio compose --style underscore --dur <len> (calm; music.md) |
| B-roll or logo | overlays[] in the EDL with output times from edit check |
Pitfalls#
- Declaring the edit done because the render exited 0: read the report and the cut views first.
- Writing ranges from memory: copy times from
takes_packed.mdor letedit cutcompute edges. - Cutting exactly at word boundaries: keep the padding (50 ms before, 80 ms after).
- Computing caption or overlay times from range sums: use the output times from
edit check. - Using an English-only model on other languages: pass
--language. - Moving word times by hand in a transcript: fix spelling only.
Read next#
references/editing.md, references/captions.md, references/footage-tools.md,
references/platforms.md, references/qa.md.