skills/showtime/references/tutorial-recording.mdTutorial and product-demo recordings
Read this when the video walks through a real app: a feature tour, a tutorial episode, a launch
demo. It covers writing a demo script, pacing clicks and typing, showtime demo record, the
events file, auto zoom with a synthetic cursor and keycaps, and how to structure 60-120 s episodes.
For inspecting the repo and serving the app read references/capture.md; for an animated
walkthrough drawn entirely in code see templates/tutorial, and for a series of episodes that share
one look, sound motif and product UI (drawn in code, targeted by named rects) read
references/series.md (showtime new series).
1. The pipeline#
showtime demo init walkthrough.mjs # starter script with every helper
showtime demo record walkthrough.mjs --url http://localhost:3000 # or --serve ./dist
showtime autozoom showtime-out/walkthrough-demo-<time>/ # zoom + cursor + ripples + keycaps
The default output folder is showtime-out/<script>-demo-<time>/ (printed at the end); in a job,
pass <job>/work/demo as the outdir. --serve uses a plain static server; never record through
showtime server or preview (the command exits 1 with "wrong page").
demo record writes frames/ (clean frames, no cursor), events.json and demo.mp4.
autozoom (also showtime footage autozoom) writes autozoom.mp4 plus autozoom.camera.json.
Use the MP4 directly, or place the frames/video in an HTML composition and drive the camera, the
runtime cursor and keystroke components from events.json (section 5).
2. Writing a demo script#
A script is an ES module; every helper advances a virtual clock and records frames, so a script produces the same video on a fast or a slow machine.
export const options = { size: '1440x900', fps: 30 }; // CSS viewport; frames at --dpr (default 2)
export default async function (demo) {
await demo.goto('/'); // relative to --url / --serve
await demo.wait(1.2); // establish the screen
await demo.chapter('Create a project');
await demo.click('text=New project'); // glide (0.35-0.9 s by distance), press, release
await demo.type('#name', 'Launch video', { cps: 14 });
await demo.press('Meta+Enter'); // shown as keycaps by autozoom
await demo.waitFor('.toast'); // real-time wait, NOT recorded
await demo.wait(1.5); // read the result
await demo.scroll('#pricing', { duration: 1.2 });
await demo.hover('.plan.pro');
await demo.focus('.plan.pro', { zoom: 1.6, duration: 2 }); // explicit camera hint
await demo.wait(1.0);
}
| Helper | Does |
|---|---|
goto(path) |
navigate (not recorded); dismisses cookie banners; stops on bot walls |
wait(s) |
record s seconds of the current screen |
waitFor(selector | fn | seconds) |
wait in real time without recording: loading never shows |
moveTo(target, {duration}) |
move the cursor on an eased, slightly curved path |
click(target, {hold, move}), dblclick |
move, press (2 frames), release, hold 0.35 s |
type(target, text, {cps, clear, jitter}) |
click the field, then type at cps characters/s with a deterministic human rhythm |
type(text, {cps}) or type(null, text, {cps}) |
type into the field that already has focus (no click, the caret stays put); the focused element's box is logged so autozoom still frames it |
press('Meta+K') |
key or combo (Playwright names: Meta, Control, Shift, Alt, Enter, ArrowDown...) |
hover, select(target, value), drag(from, to) |
as named |
scroll(target | 'top' | 'bottom' | {y} | dy, {duration}) |
smooth, eased scroll |
focus(target, {zoom, duration}) |
tell autozoom to frame this element |
chapter(title), note(text) |
markers for titles, chapters, narration |
page |
the Playwright page (escape hatch) |
Targets are Playwright selectors (#id, .class, text=Save, role=button[name="Save"]) or
points {x, y} in CSS pixels. Off-screen targets are scrolled into view first.
Every pointer action first glides the cursor to its target: 0.35-0.9 s depending on the distance
(click(t, {move: 0.2}) sets it), and drag(from, to) glides to from before pressing. The press or
the drag start lands that much after the moment you call it, so to hit a narration word at time w,
start the action about w - 0.4 s (or pass a short move).
The page sees the recording machine's OS unless told otherwise: an app that shows "Ctrl K" on Linux
and "⌘K" on a Mac changes its labels and modifier keys with it. Record with --platform mac (or
export const options = { platform: 'mac' } in the script) when the narration names Mac shortcuts,
so a take on Windows or Linux matches. The app's own clock (Date.now(), "Just now" labels) runs in
real time, and a take can last several times its video length: keep time-of-day text out of shots,
or freeze it in the app's test mode.
Timing: CSS transitions and Web Animations are stepped frame by frame (--animations step), so
hover effects and toasts animate at true speed in the recording. JavaScript timers still run in real
time: gate on app state with waitFor, never on fixed sleeps.
3. Pacing rules#
- Capture at 2x (default
--dpr 2) so zooms up to 2x stay sharp. Use a 1280x800 to 1440x900 CSS viewport; bump the app's own zoom (125-150%) if text would be small at 1080p. - One action per sentence of narration. Hold 1.5-2.5 s on anything the viewer must read.
- Typing: 12-16 characters/s for short fields; paste or cut long text (never type a paragraph). After the first line of a field, omit the target so each line does not click (and move the caret).
- Park the pointer near what you type into or talk about, never on a label the voice names: the camera follows the pointer, and a parked pointer covers the thing on screen.
- Loading: skip it with
waitFor(the result appears instantly); never show a real wait over 1 s. - Dead air over 0.7 s gets cut or covered by a camera move.
- Clean the stage first: seeded demo data, no personal data, no notification badges, no dev overlays. Use fictional names and say so.
- Episodes of 60-120 s: a 3-5 s hook showing the end result, 3-5 steps of 10-25 s each, a 5 s
recap/CTA. Over 2 minutes, add chapters (
demo.chapter) with title cards of about 1.5 s or description timestamps (first at 0:00, at least three, each at least 10 s long).
4. Auto zoom#
showtime autozoom <demo-folder> # framed look, 1920x1080
showtime autozoom <demo-folder> --look plain --size 1080x1920 --fit cover # vertical follow-cam
showtime autozoom <demo-folder> --keys all --max-zoom 2.2 --bg "#0b1020,#312e81"
showtime autozoom screen.mp4 --cursor-log cursor.csv --cursor-scale 2 # your own recording
showtime autozoom screen.mp4 # no telemetry: follows screen changes
showtime autozoom <demo-folder> --preview # half size, fast
Zoom rules built in:
- Clicks zoom to 1.8x, typing 1.6x, drags 1.4x, scrolls stay wide; an element's box limits the
zoom so it fills about 70% of the view. Hard cap 2.8x (
--max-zoom, default 2.0). - The camera starts moving before the action: 0.15 s (click), 0.25 s (scroll/drag), 0.4 s (typing).
- Actions closer than 1.5 s and within 20% of the screen share one shot; shots less than 1.8 s
apart are bridged (the camera pans instead of zooming out and back in), but never across a
demo.chapter()marker, anddemo.wide()ends any shot at that moment (a wide beat); each shot holds at least about 1.6 s after its last click (1.2 s after typing), then zooms out. - A target in a window corner may use the framed look's padding, so it is not pinned to the picture's edge.
- The camera follows the cursor with a dead zone (central 75% of the view, 60% while typing).
- Motion is a damped spring (zoom in log space, response 0.55 s; pan 0.42 s).
- The synthetic cursor is 1.6x a normal pointer, dips on click with a ring ripple (0.35 s), and fades
after 2 s idle. Keycaps sit bottom-centre at 88% height, fade in 0.15 s, stay 1.5 s after the last
key, fade out 0.3 s.
--keys allalso shows typed text. --look framed(default) puts the window on a gradient with padding, rounded corners and a soft shadow; zooming in fills the frame.--keys-pos bottom|top|bottom-left|bottom-right|top-left|top-rightmoves the keycaps;--key-glyphsdraws ⌘ ⌥ ⇧ ⌃ ↵ instead of words. Every run also writes<name>.keys.json(keycaps and typed text with times and glyph labels) so a composition can draw its own (--keys off).- Output names never overwrite: a second run writes
autozoom-2.mp4andautozoom-2.camera.json, and says so first. Copy the file it names. - The text summary lists every shot (
start-end x zoom at x%,y%);--plan --jsonhasshot_list.
Changing the camera without a new take:
--hints hints.json: a list of events merged at plan time:{"type":"focus","t":12,"end":14, "x":640,"y":300,"zoom":1.4}adds a shot (CSS px),{"type":"wide","t":20}goes wide,{"type":"drop","t":8,"end":9.5}removes the actions in that span. To name a key the way the voice does, drop it and re-add it with a label:{"type":"key","t":45.47,"keys":".","label":". full stop"}(a lone punctuation key is labelled with its name by default, e.g.. Period).--cursor-offset 22.6-24:36,12moves the recorded pointer 36,12 CSS px inside 22.6-24 s, eased in and out (a parked pointer that covers a label); repeatable.
Your own screen recordings: record with the system cursor hidden if you can, and log the pointer
alongside as JSON [[t, x, y, down], ...] or CSV t,x,y,down in video pixels (--cursor-scale 2
for macOS point coordinates on a Retina capture). Clicks come from down changes and dwells
(the pointer resting 0.45-2.6 s) become softer zoom targets. Without a log, zooms follow regions of
the screen that change, and no cursor is drawn.
5. events.json (for compositions)#
{ "fps": 30, "frames": 412, "duration": 13.73, "viewport": {"width": 1440, "height": 900}, "dpr": 2,
"frame_size": {"width": 2880, "height": 1800}, "frame_pattern": "frames/%05d.jpg", "video": "demo.mp4",
"cursor": [[0.0, 892.8, 648.0, 0], ...],
"events": [{"t": 1.87, "type": "click", "x": 308, "y": 279, "bbox": [48, 252, 520, 53], "target": "#repo"},
{"t": 2.04, "end": 2.54, "type": "type", "text": "acme/web", "bbox": [...]},
{"t": 3.66, "type": "key", "keys": "Meta+K"},
{"t": 0.5, "end": 1.29, "type": "scroll", "from": 0, "to": 2492},
{"t": 0.5, "type": "chapter", "title": "Create a preview"}],
"chapters": [...],
"overlays": {"cursor_path": [{"at": 1.87, "x": 32.1, "y": 51.7, "click": true}],
"keystrokes": [{"at": 2.04, "text": "acme/web"}, {"at": 3.66, "keys": "Meta K"}]} }
Coordinates are CSS pixels of the viewport (multiply by dpr for frame pixels); cursor has one
row per frame. overlays.cursor_path and overlays.keystrokes are already in the format of the
runtime cursor and keystrokes components (x/y in % of the frame), so a composition can show the
clean frames as a video layer and draw its own cursor and keycaps in the page style. The camera
from autozoom --plan (*.camera.json: per frame [t, zoom, cx, cy], centre as a fraction of
the source frame) can drive a CSS transform on that layer:
transform: scale(zoom) translate((0.5 - cx) * 100%, (0.5 - cy) * 100%) with
transform-origin: 50% 50%.
Composing a recording into an HTML project#
When the tutorial needs step cards, a result-first hook or captions over the recording, use the
recording as a <video data-st> layer instead of an EDL:
showtime footage trim rec/demo.mp4 --width 1920 -o <job>/project/media/demo.mp4(a seek-friendly proxy), or useautozoom.mp4when its camera is what you want.- Place it:
<video data-st src="media/demo.mp4" data-offset="S">inside a scene;data-offsetis the recording time shown at the scene start. Chapter times fromevents.jsonminus the offset give each step card'sdata-start. - Voice lines go in
audio/mix.jsonat chapter time plus a short lead (0.3-0.6 s); the recording sets the timing here, soretime --from-voicedoes not apply. Build the caption words at video time (voice word times plus each line'sstart) forcaption-karaoke. - In a vertical short, frame the desktop recording as a card and push in per beat with a small
ST.onSeekcamera (scale/translate from waypoints in recording pixels), or withbrowser-framezoomsteps (x/y in % of the view).
6. Checklist before rendering the episode#
- The script runs start to finish twice with the same frame count.
- No personal data, tokens or real customer names in any frame (check the contact sheet or
showtime footage view). - Every step has narration or a caption; every click lands on something visible.
- Zoom never cuts off the thing being explained; re-run with
--max-zoom 1.6if it does. - Final render passes
showtime qa.