showtime.
Benchmark · v0.3.0 · round 4 · 30 Sep 2026 · claude-opus-5-5

showtime benchmark, v0.3.0: eight video requests, round 4

Earlier version: v0.2.0 (rounds 1 to 3)

The blind vote preferred showtime 0.3.0 (lean) over 0.2.0 on 4 of 6 tasks; the AI judge preferred 0.2.0 on 5

What was compared

  • showtime 0.3.0, lean review mode: new runs. At run time lean was showtime's default; it has since become an explicit opt-in.
  • showtime 0.2.0, default mode: the round-3 videos, reused, made from the same prompts. They are the videos voted on in round 3.
  • plain Opus 5.5: the same model in Claude Code with no video plugin. Reused from round 1 on the six older tasks; new runs on the two new tasks.
  • brag-slim: a launch-video-only skill, on the launch task (t1) only. Reused from round 1.

The 0.3.0 that ships uses the quality review mode by default: a separate critic round on every finished video. That default was not benchmarked in a blind round. The only quality-mode number here is one paid proof run on the data story (t2), at $3.03 API-equivalent; it was not on the blind boards and not ranked by the judge. Every table and chart below names the mode.

The reviewer's blind vote, pair by pair: 0.3.0 lean beat 0.2.0 default on t1, t4, t6 and t8 and lost t2 and t3. It beat plain Opus 5.5 on 7 of 8 tasks and lost the "make it like this reference" task (t10). On the launch video it beat brag-slim, the first time showtime has won that pair in a blind vote. The verified AI judge saw it differently: it ranked 0.2.0 above 0.3.0 lean on five of the six shared tasks, for polish reasons such as still holds, a blurred crossfade and a missing year label.

That disagreement is the main result, and it is why the shipped default is the quality mode and not lean. The voter is one person, showtime's author, and the vote is only partly blind.

4 of 6
blind-vote pairs won by 0.3.0 lean against 0.2.0 default (t1, t4, t6, t8; lost t2, t3)
7 of 8
blind-vote pairs won by 0.3.0 lean against plain Opus 5.5 (lost t10)
1 of 6
tasks where the AI judge ranked 0.3.0 lean above 0.2.0 default (t4, the filler cut)
1st time
showtime beat brag-slim on the launch video in a blind vote, after three losses
Small sample.8 tasks, 1 run per cell (2 for 0.3.0 on three tasks), and 1 human voter, who is showtime's author and had seen the older videos. Treat each task as one anecdote.

At a glance

The vote went to 0.3.0 lean on five tasks, to 0.2.0 on two, to plain Opus on one

First pick and the "which would you ship?" pairs are from the reviewer's blind boards (shuffled letters, no tool names). The judge's mean rank is over three rotated judgments per task; 1 is best, and the number of versions on the task is given with each rank.

Round 4 per task: blind vote and judge
TaskBlind voteAI judge, mean rank (1 = best)
First pick0.3.0 lean vs 0.2.00.3.0 lean vs plain0.3.0 lean0.2.0 defaultplain Opus
t1 launch video0.3.0 leanWONWON4.00 of 51.335.00
t2 data story0.2.0 defaultLOSTWON3.00 of 41.334.00
t3 vertical short0.2.0 defaultLOSTWON2.67 of 31.671.67
t4 filler cut0.3.0 leanWONWON1.33 of 31.673.00
t6 HTML report0.3.0 leanWONWON2.00 of 31.003.00
t8 math explainer0.3.0 leanWONWON3.00 of 41.674.00
t9 repo explainer (new)0.3.0 leannot runWON1.00 of 2not run2.00
t10 like a reference (new)plain Opus 5.5not runLOST2.00 of 2not run1.00

t9 (explain a repo) and t10 (make it like a reference) are new in round 4 and were run with 0.3.0 lean and plain Opus 5.5 only. The second 0.3.0 lean run on t1, t2 and t8 was judged but kept off the boards, to keep the vote short.

Vote and judge

The reviewer and the judge disagree about 0.3.0 lean against 0.2.0

Against plain Opus 5.5 the two instruments mostly agree: the vote went to 0.3.0 lean on 7 of 8 tasks and the judge ranked it higher on 6 of 8 (not on t3 and t10). Against 0.2.0 default they point in opposite directions: 4 of 6 for 0.3.0 lean in the vote, 1 of 6 with the judge.

Comparisons won by 0.3.0 lean, by instrument

Blind vote: "which would you ship?" pairs. AI judge: tasks where its mean rank was better. Drawn to scale.

0% 50% 100% Blind vote, vs 0.2.0 default: 4 of 6 pairs won Blind votevs 0.2.0 default 4 of 6 AI judge, vs 0.2.0 default: 1 of 6 tasks ranked above AI judgevs 0.2.0 default 1 of 6 Blind vote, vs plain Opus 5.5: 7 of 8 pairs won Blind votevs plain Opus 5.5 7 of 8 AI judge, vs plain Opus 5.5: 6 of 8 tasks ranked above AI judgevs plain Opus 5.5 6 of 8

Reading this. Each bar is a count of tasks, one pair or one ranking per task. A task can flip on a single run. The vote is one person's. The judge ranks still frames, a transcript and measurements; it cannot watch motion or hear sound.

On the five tasks where both were rated, the reviewer's 0-5 ratings averaged 4.2 for 0.3.0 lean and 4.4 for 0.2.0 default. The ratings tied on t4 and t6, where the pair went to 0.3.0 lean. So the ratings lean slightly to 0.2.0 while the pairs lean to 0.3.0.

What the judge held against 0.3.0 lean

These are polish defects, the kind a separate review of the finished video exists to catch. The lean mode in round 4 skipped that separate critic round to save cost. That is why the shipped default became the quality mode, which runs a critic round on every finished video. Whether it removes these defects has been checked on one proof run only, not in a blind round.

AI judge: mean rank per task, every version

Three judgments per task; further left is better; the grey track spans the versions on that task.

1st 2nd 3rd 4th 5th best t1, mean rank of 5: plain Opus 5.5 5.00; brag-slim 2.00; 0.2.0 default 1.33; 0.3.0 lean 4.00; 0.3.0 lean, 2nd run 2.67 t1of 5 t2, mean rank of 4: plain Opus 5.5 4.00; 0.2.0 default 1.33; 0.3.0 lean 3.00; 0.3.0 lean, 2nd run 1.67 t2of 4 t3, mean rank of 3: plain Opus 5.5 1.67; 0.2.0 default 1.67; 0.3.0 lean 2.67 t3of 3 t4, mean rank of 3: plain Opus 5.5 3.00; 0.2.0 default 1.67; 0.3.0 lean 1.33 t4of 3 t6, mean rank of 3: plain Opus 5.5 3.00; 0.2.0 default 1.00; 0.3.0 lean 2.00 t6of 3 t8, mean rank of 4: plain Opus 5.5 4.00; 0.2.0 default 1.67; 0.3.0 lean 3.00; 0.3.0 lean, 2nd run 1.33 t8of 4 t9, mean rank of 2: plain Opus 5.5 2.00; 0.3.0 lean 1.00 t9of 2 t10, mean rank of 2: plain Opus 5.5 1.00; 0.3.0 lean 2.00 t10of 2

Two runs of the same build differ a lot. 0.3.0 lean was run twice on t1, t2 and t8. On the math explainer the second run ranked first (mean 1.33), ahead of 0.2.0 (1.67), while the first run ranked 3.00. On the launch video the second run ranked 2.67 against the first run's 4.00. On all three tasks the two runs of the same build were 1.33 to 1.67 places apart; on t8 that was more than the gap between 0.3.0 lean and 0.2.0. Read every single-task result here as one draw.

The vote only saw the first run of each.

The launch video (t1)

showtime beat brag-slim in a blind vote for the first time; the judge put it fourth of five

In rounds 1, 2 and 3 the reviewer picked brag-slim on the launch video every time. In round 4 the reviewer picked 0.3.0 lean, and preferred it to brag-slim, to 0.2.0 default and to plain Opus 5.5 in all three pairs asked.

Two things temper that. First, the voter's view of the unchanged videos moved too: the same brag-slim and 0.2.0 files were on the round-3 board, where brag-slim won that pair; this time 0.2.0 won it, and plain Opus 5.5's silent video also beat brag-slim. Second, the judge ranked 0.3.0 lean 4th of 5 in all three judgments, above plain Opus 5.5 only, citing still holds and fewer 2.0 features. It ranked 0.2.0 first (mean 1.33) and brag-slim second (2.00). The board for this task has pairs and a pick but no 0-5 ratings.

Results by task

Eight tasks, one run each

Blind vote: the reviewer's 0-5 rating, first pick, and pairs won on the board. Judge rank: mean of three rotated judgments, with each judgment's place. Cost: the API-equivalent USD computed from the run's stream; the runs used a subscription and were not billed per token. Wall: the whole run. Reused cells keep the numbers of the round they were made in, on other days and under other load. qa: showtime qa from this round's checkout, the same thresholds for every file. Flagged claims: claims the fact-check judge marked unsupported or contradicted by the task's own files. Reviewer notes are quoted with spelling corrected.

t1 · launch video from a repo

Launch video for quillsort 2.0

Make a 30-second launch video for quillsort 2.0 from this repo that I can post on X.

Votepick 0.3.0 lean; won all 3 of its pairs
Judge1st 0.2.0 default · 0.3.0 lean 4th of 5
plain Opus 5.5: 'New in quillsort 2.0' feature list
plain Opus 5.5
frame at 24 s
brag-slim: natural sort demo, 'file2 before file10. Finally.'
brag-slim
frame at 12 s
showtime 0.2.0: terminal with plain sort output and 'file2 before file10'
0.2.0 default
frame at 6 s
showtime 0.3.0 lean: a terminal showing duplicates removed, 'Duplicates removed.'
0.3.0 lean
frame at 18 s
plain Opus 5.5reused, round 1brag-slimreused, round 10.2.0 defaultreused, round 30.3.0 lean0.3.0 lean, 2nd run
Blind votepairs won: 1 of 2no ratings on this boardpairs won: 0 of 3no ratings on this boardpairs won: 1 of 2no ratings on this boardpickpairs won: 3 of 3no ratings on this boardnot on the board
Judge rankmean 5.00of 5; judgments: 5th, 5th, 5thmean 2.00of 5; judgments: 1st, 2nd, 3rdmean 1.33of 5; judgments: 2nd, 1st, 1stmean 4.00of 5; judgments: 4th, 4th, 4thmean 2.67of 5; judgments: 3rd, 3rd, 2nd
Cost$0.88$2.09$3.73$4.01$3.81
Wall time4.4 min10.0 min14.4 min14.1 min12.5 min
Turns19361379
qaFAILsilent audio trackWARNcolour tags missingPASSWARNthree still holds of 3.6-4.6 sWARNstill holds of up to 4.9 s
Loudnessno audio−14.1 LUFS−14.0 LUFS−14.0 LUFS−14.0 LUFS
Flagged claims01demo output shown unsorted000

Review notes

  • No note and no ratings on this board. The reviewer preferred 0.3.0 lean in each pair, 0.2.0 over brag-slim, and plain Opus 5.5 over brag-slim.

Judge's reasoning, in short

All three judgments put 0.3.0 lean 4th: a section on --unique not shown as new, an empty terminal beat, holds of about 4 s, fewer new features, and a caption area left empty for several seconds. 0.2.0 was first twice for the tightest pacing and correct terminal output. brag-slim had the strongest hook but shows a dedupe result out of order. plain Opus 5.5 is silent, which the judge called a hard fail for a post on X.

t2 · data story from a CSV

45-second temperature data story

Turn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube.

Votepick 0.2.0 default; 0.3.0 lean lost to it, beat plain Opus
Judge1st 0.2.0 default · 0.3.0 lean 3rd of 4
plain Opus 5.5: dark bar chart, '2024 was the hottest year in 176 years of records'
plain Opus 5.5
frame at 36 s
showtime 0.2.0: decade averages bar chart, 'Since the 1970s, every decade beat the last'
0.2.0 default
frame at 27 s
showtime 0.3.0 lean: line chart, 'Since 1977, every year has been warmer than average'
0.3.0 lean
frame at 18 s
showtime 0.3.0 lean second run: dark bar chart with a year counter at 1897
0.3.0 lean, 2nd run (not voted)
frame at 9 s
plain Opus 5.5reused, round 10.2.0 defaultreused, round 30.3.0 lean0.3.0 lean, 2nd run
Blind vote★★★★★pairs won: 0 of 2★★★★★pickpairs won: 2 of 2★★★★★pairs won: 1 of 2not on the board
Judge rankmean 4.00of 4; judgments: 4th, 4th, 4thmean 1.33of 4; judgments: 2nd, 1st, 1stmean 3.00of 4; judgments: 3rd, 3rd, 3rdmean 1.67of 4; judgments: 1st, 2nd, 2nd
Cost$1.33$5.63$1.42$1.71
Wall time8.5 min21.3 min8.9 min9.1 min
Turns30204137
qaWARNno audio track; flat first frame; holds up to 5.3 sPASSPASSPASS
Loudnessno audio−14.0 LUFS−14.0 LUFS−14.1 LUFS
Flagged claims01credits NOAA, which the CSV does not name00

Review notes

Neither was the best, but C was cleaner; in B and C both, numbers sometimes disappear and then reappear when a new item was added.The reviewer. C was 0.2.0 default, B was 0.3.0 lean: the vanishing numbers were in both showtime videos.

Judge's reasoning, in short

0.2.0 told the fullest story in four beats: a yearly line, decade averages, the top-11 ranking and an end card with its source. 0.3.0 lean was clean and accurate but thinner, with the ranked bar chart held for about 15 s and odd 22-year axis ticks. The second 0.3.0 lean run, a bold dark design with a running year counter, was first once and second twice. plain Opus 5.5 has no audio track, which a YouTube video needs.

t3 · vertical short with voice-over and captions

30-second Reels short for quillsort 2.0

Make a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions.

Votepick 0.2.0 default; 0.3.0 lean lost to it, beat plain Opus
Judge0.2.0 default and plain Opus 5.5 tied first · 0.3.0 lean 3rd of 3
plain Opus 5.5: --natural card with a terminal and a caption
plain Opus 5.5
frame at 6 s
showtime 0.2.0: 'file2 before file10.' on notebook paper
0.2.0 default
frame at 0.6 s
showtime 0.3.0 lean: -i edits in place card with a terminal and a caption
0.3.0 lean
frame at 12 s
plain Opus 5.5reused, round 10.2.0 defaultreused, round 30.3.0 lean
Blind vote★★★★★pairs won: 0 of 2★★★★★pickpairs won: 2 of 2★★★★★pairs won: 1 of 2
Judge rankmean 1.67of 3; judgments: 1st, 1st, 3rdmean 1.67of 3; judgments: 2nd, 2nd, 1stmean 2.67of 3; judgments: 3rd, 3rd, 2nd
Cost$1.81$2.41$2.66
Wall time11.0 min11.3 min12.3 min
Turns464456
qaWARNlong caption line; flat first frame; last 2 s silentPASSPASS
Loudness−14.0 LUFS−14.0 LUFS−14.0 LUFS
Flagged claims002both false flags: the README, a listed source, states Python 3.8+ and pip install quillsort

Review notes

  • No note. The reviewer rated 0.2.0 5 of 5, 0.3.0 lean 4 and plain Opus 5.5 2.

Judge's reasoning, in short

Two judgments put 0.3.0 lean last: a left-heavy layout that leaves the right side and bottom half empty, a caption pressed against a card, and an in-place edit demo whose backup file shows only two lines, which looks broken. One judgment called it the strongest graphic design and put it second. plain Opus 5.5 was first twice for its clear one-card-per-change structure; the judge cannot hear the voice, which the reviewer marked down.

t4 · footage edit

Cut the fillers and long pauses from an interview

Cut the filler words and long pauses out of interview.mp4 and give me the tightened video.

Votepick 0.3.0 lean; won both of its pairs
Judge1st 0.3.0 lean (mean 1.33) · 0.2.0 default 1.67
the interviewee on camera, the same picture in every cut
0.3.0 lean, frame at 9.9 s. The other cuts show the same camera picture.
plain Opus 5.5reused, round 10.2.0 defaultreused, round 30.3.0 lean
Blind vote★★★★★pairs won: 0 of 2★★★★★pairs won: 1 of 2★★★★★pickpairs won: 2 of 2
Judge rankmean 3.00of 3; judgments: 3rd, 3rd, 3rdmean 1.67of 3; judgments: 2nd, 1st, 2ndmean 1.33of 3; judgments: 1st, 2nd, 1st
Cost$1.06$1.72$1.57
Wall time7.2 min6.3 min7.1 min
Turns234849
qaFAILkept a 5.2 s silence and frozen pictureWARN2 LU quiet (−16 LUFS); a 2.2 s silence inside; silent tailWARNa 2.2 s silence inside; 2.3 s silent tail
Loudness−16.0 LUFS−16.0 LUFS−14.0 LUFS

Review notes

I think they were all great. It was hard to pick a winner, but they were all very good.The reviewer (shortened). Rated 0.3.0 lean and 0.2.0 5 of 5 each, plain Opus 5.5 4.

Judge's reasoning, in short

Both showtime cuts removed every audible "uh" and the stutter, came in at about 49.5 s and left a 2.2 s silence and a silent tail. 0.3.0 lean was first twice because it is normalised to −14 LUFS (0.2.0 stayed at the source's −16). One judgment preferred 0.2.0 because 0.3.0 lean cut an "and" with an "uh". plain Opus 5.5 kept the fillers and a 5.2 s silence.

t6 · single-file HTML video report

Shareable HTML report on the Mauna Loa CO2 record

Make a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team.

Votepick 0.3.0 lean; won both of its pairs
Judge1st 0.2.0 default in all three · 0.3.0 lean 2nd
plain Opus 5.5: a dashboard page with a CO2 line chart and a play button
plain Opus 5.5
page on load
showtime 0.2.0: player poster, CO2 rose from 315.98 to 427.35 ppm, callout '2015: first year above 400'
0.2.0 default
page on load
showtime 0.3.0 lean: player poster, callout 'first year above 400 ppm' without a year
0.3.0 lean
page on load
plain Opus 5.5reused, round 10.2.0 defaultreused, round 30.3.0 lean
Blind vote★★★★★pairs won: 0 of 2★★★★★pairs won: 1 of 2★★★★★pickpairs won: 2 of 2
Judge rankmean 3.00of 3; judgments: 3rd, 3rd, 3rdmean 1.00of 3; judgments: 1st, 1st, 1stmean 2.00of 3; judgments: 2nd, 2nd, 2nd
Cost$1.40$0.98$0.91
Wall time6.5 min4.6 min6.2 min
Turns292329
HTML checksPASSsingle file, offline, plays, phone widthPASSsingle file, offline, plays, phone widthPASSsingle file, offline, plays, phone width
Flagged claims3a NOAA credit, a 1970 value off by 0.04 ppm, a 'longest record' line00

Review notes

B and C are almost the same. They were both a little bit slow … not very dynamic, but I like them both.The reviewer (shortened). B was 0.3.0 lean, C was 0.2.0 default; both rated 4 of 5.

Judge's reasoning, in short

The two showtime pages share one editorial design and the same correct figures. 0.2.0 was first in all three: its "2015: first year above 400" label is exact and it runs a tight 25 s. 0.3.0 lean adds a helpful 1959 baseline line, but its 400 ppm callout gives no year and one frame caught a heavy blurred crossfade. plain Opus 5.5's page has the most data but looks like a generic dashboard, with small text inside the player and no audio.

t8 · math explainer

Why the first n odd numbers add up to n²

Make a 45-second video that shows visually why the sum of the first n odd numbers is n squared.

Votepick 0.3.0 lean; won both of its pairs
Judge1st 0.3.0 lean, 2nd run · 0.2.0 default 1.67 · 0.3.0 lean (voted run) 3.00
plain Opus 5.5: a 7 by 7 grid of coloured layers with the sum
plain Opus 5.5
frame at 44.1 s
showtime 0.2.0: coloured L-shaped layers forming an n by n square, n squared boxed
0.2.0 default
frame at 44.2 s
showtime 0.3.0 lean: an n by n square of layers next to 1+3+5+...+(2n-1) = n squared
0.3.0 lean
frame at 43.9 s
showtime 0.3.0 lean second run: 'n odd numbers fill an n by n square'
0.3.0 lean, 2nd run (not voted)
frame at 44.1 s
plain Opus 5.5reused, round 10.2.0 defaultreused, round 30.3.0 lean0.3.0 lean, 2nd run
Blind vote★★★★★pairs won: 0 of 2★★★★★pairs won: 1 of 2★★★★★pickpairs won: 2 of 2not on the board
Judge rankmean 4.00of 4; judgments: 4th, 4th, 4thmean 1.67of 4; judgments: 2nd, 2nd, 1stmean 3.00of 4; judgments: 3rd, 3rd, 3rdmean 1.33of 4; judgments: 1st, 1st, 2nd
Cost$0.88$2.35$1.45$1.70
Wall time4.5 min14.3 min6.9 min8.6 min
Turns18463138
qaFAILblack first frame; no audio track; still holdsWARNone caption cue under 0.4 sPASSPASS
Loudnessno audio−14.0 LUFS−14.0 LUFS−14.0 LUFS
Flagged claims0000

Review notes

B was the best … the animation was different, but A was a little bit more clear. One weird thing about A is that the agent says "1 1 + 3", like that, it's kinda weird, but it was well explained in the end.The reviewer (shortened). B was 0.3.0 lean, A was 0.2.0 default. The remark about "1 1 + 3" is about the 0.2.0 video. The judges' speech check found no voice-over in it; it came with a captions file whose first cue reads "One.", so the remark may be about its captions.

Judge's reasoning, in short

All four give a correct L-shaped proof. The second 0.3.0 lean run was first twice for step titles that guide the viewer. 0.3.0 lean (the voted run) was third in all three: no guiding headings, and equations that overlap into a jumble during a transition around 35 s. plain Opus 5.5 has no audio track and starts and ends on black.

t9 · explain a repo · new in round 4

A 60-90 second explainer of an unfamiliar repo

Make a 60-90 second video that explains what this repo does and how it works, for developers seeing it for the first time.

Votepick 0.3.0 lean (5 of 5 against 4)
Judge1st 0.3.0 lean in all three
plain Opus 5.5: a pipeline diagram of the repo's modules
plain Opus 5.5
frame at 51 s
showtime 0.3.0 lean: 'Four rules, each returns findings' with the rules code
0.3.0 lean
frame at 51.8 s
plain Opus 5.50.3.0 lean
Blind vote★★★★★pairs won: 0 of 1★★★★★pickpairs won: 1 of 1
Judge rankmean 2.00of 2; judgments: 2nd, 2nd, 2ndmean 1.00of 2; judgments: 1st, 1st, 1st
Cost$1.44$3.42
Wall time11.2 min23.0 min
Turns3061
qaFAILno audio track; a 6.9 s frozen picturePASS
Loudnessno audio−14.0 LUFS
Flagged claims1output of a diff against a file that is not in the repo21 real: it says extra keys are warnings unless strict; extra is always a warning. The other (a __main__.py) exists in the repo

Review notes

  • No note. Two versions only, one pair.

Judge's reasoning, in short

plain Opus 5.5 explains a little more (exit codes, the diff and init commands, a pre-commit hook) but has no audio track and long frozen holds, one of 6.9 s. 0.3.0 lean walks the parse, compare and report steps with a module tracker, has music and captions, and ends on where to start reading; some code text is small for a phone. Neither has narration. The fact-check found one real imprecision in 0.3.0 lean (see the table).

t10 · make it like a reference · new in round 4

A 20-second announcement in the style of a reference video

Make a 20-second video announcing our spring plant swap (details in swap.md) in the same style as reference.mp4.

Votepick plain Opus 5.5; 0.3.0 lean lost its only pair
Judge1st plain Opus 5.5 in all three
the reference: red 'NIGHT RUN CLUB' on a near-black end card
The reference (not a candidate), drawn for the benchmark
reference, frame at 17 s
plain Opus 5.5: red 'HOLLIS STREET PLANT SWAP' on a near-black end card
plain Opus 5.5
frame at 19.6 s
showtime 0.3.0 lean: green 'HOLLIS STREET PLANT SWAP' on a dark green end card
0.3.0 lean
frame at 19.6 s
plain Opus 5.50.3.0 lean
Blind votepickpairs won: 1 of 1no ratings on this boardpairs won: 0 of 1no ratings on this board
Judge rankmean 1.00of 2; judgments: 1st, 1st, 1stmean 2.00of 2; judgments: 2nd, 2nd, 2nd
Cost$1.62$0.98
Wall time5.6 min4.0 min
Turns4429
qaFAIL4.8 LU too quiet (−18.8 LUFS)PASS
Loudness−18.8 LUFS−14.1 LUFS
Flagged claims00

Review notes

B was much closer to the reference and well explained. It went a little silent at the end, but it was good; it kept the colours and the idea.The reviewer (spelling corrected). B was plain Opus 5.5, the winner; the silence is in its video. No ratings on this board.

Judge's reasoning, in short

Both kept the reference's structure: a word-by-word opener, numbered cards with a progress bar, a wipe and a dark end card. plain Opus 5.5 kept the reference's cream, red and black palette and its very large type. 0.3.0 lean changed the accent to green to suit the plants and added small sentence-case lines the reference never uses; three of its six sampled stills are full-screen green wipes. plain Opus 5.5 was quiet (−18.8 LUFS, a qa FAIL); 0.3.0 lean was on target.

Release checks

Six of eight release checks passed; cost and the strict quality rule failed

Before the round, 0.3.0 set itself eight checks. They were computed by the benchmark's own report script from the round's files, on the 0.3.0 lean runs.

Round 4 release checks
CheckResultDetail
Cost: median at or under plain Opus 5.5FAIL0.3.0 lean $1.51 against plain Opus 5.5 $1.37, on the same 8 tasks (API-equivalent, from each run's stream)
Quality: judge and vote both at or above 0.2.0 on 5 of 6 tasksFAIL1 of 6. Only t4 passed both. The vote alone was 4 of 6 (t1, t4, t6, t8); the judge alone 1 of 6. The rule needs both to agree on each task.
Images read into the main context: median at most 12 per jobPASSmedian 4.5 (t4 highest at 11)
Full renders per job: median 1PASSmedian 1 (from each job's receipt; t4 rendered 4 times, t1 3, t8 0 full renders)
qa FAIL findings: 0 on every new 0.3.0 runPASS10 video runs checked, none failed; the HTML report has no qa verdict
Launch video (t1): not lastPASSjudge mean rank 4.0 of 5 (plain Opus 5.5 last at 5.0); blind pairs won 3 of 3
Every job writes a receiptPASS11 of 11 new 0.3.0 runs, all with a cost
Isolation: no run touched another showtime folderPASS13 new runs checked. The first pass flagged four showtime runs; that was a checker bug (it read encoded command text and counted the shared tools linked into the arm's own home as foreign), and the fixed checker found none.

Round 3 had a real isolation failure: two of its runs used another working copy of showtime. The isolation check was added for round 4 to catch that; it found nothing once its own bug was fixed.

Process and cost

Lean is the cheaper path; quality mode spends more for a critic round

Cost here is API-equivalent USD computed from each run's stream of model calls. The runs used a subscription, so nobody was billed these amounts. Mode matters: lean skips the separate review round, quality, the shipped default, adds a critic round on every finished video and costs more for it. On the one task measured all three ways, the data story (t2), 0.3.0 lean cost $1.42, 0.3.0 quality $3.03 (one proof run) and 0.2.0 default $5.63.

Cost per run, by task

API-equivalent USD; one run per bar; drawn to scale.

$0 $2 $4 $6 t1 t1, plain Opus 5.5: $0.88 $0.88 t1, 0.2.0 default: $3.73 $3.73 t1, 0.3.0 lean: $4.01 $4.01 t2 t2, plain Opus 5.5: $1.33 $1.33 t2, 0.2.0 default: $5.63 $5.63 t2, 0.3.0 lean: $1.42 $1.42 t2, 0.3.0 quality mode (one proof run, not blind-voted): $3.03 $3.03 quality mode, proof run t3 t3, plain Opus 5.5: $1.81 $1.81 t3, 0.2.0 default: $2.41 $2.41 t3, 0.3.0 lean: $2.66 $2.66 t4 t4, plain Opus 5.5: $1.06 $1.06 t4, 0.2.0 default: $1.72 $1.72 t4, 0.3.0 lean: $1.57 $1.57 t6 t6, plain Opus 5.5: $1.40 $1.40 t6, 0.2.0 default: $0.98 $0.98 t6, 0.3.0 lean: $0.91 $0.91 t8 t8, plain Opus 5.5: $0.88 $0.88 t8, 0.2.0 default: $2.35 $2.35 t8, 0.3.0 lean: $1.45 $1.45 t9 t9, plain Opus 5.5: $1.44 $1.44 t9, 0.3.0 lean: $3.42 $3.42 t10 t10, plain Opus 5.5: $1.62 $1.62 t10, 0.3.0 lean: $0.98 $0.98
Median cost and wall time
VersionTasksMedian costMedian wall
0.3.0 leanall 8$1.518.0 min
plain Opus 5.5all 8$1.376.9 min
0.3.0 lean6 older tasks$1.518.0 min
0.2.0 default6 older tasks$2.3812.8 min
plain Opus 5.56 older tasks$1.206.9 min
0.3.0 quality the shipped defaultt2 only, one proof run$3.03—

0.3.0 lean cost less than plain Opus 5.5 on t6 and t10, within $0.10 on t2, and more on the other five, most on the launch video ($4.01) and the repo explainer ($3.42), its two longest runs. Both 0.3.0 lean launch runs used helper agents, as the 0.2.0 launch run did. Wall times of reused cells come from other days and other machine load, so compare them loosely.

Other process numbers from the stream, for 0.3.0 lean: a median of 4.5 images read into the main agent's context per job and a median of one full render. The two runs of the same task were close in cost ($4.01 and $3.81 on t1, $1.42 and $1.71 on t2, $1.45 and $1.70 on t8).

Method

How round 4 was run

The same harness as rounds 1 to 3 (described in the v0.2.0 report), driven from a manifest that lists which cells are new and which are reused. The harness, tasks and scorers are in the showtime repository under benchmarks/.

The contenders

showtime 0.3.0, lean

The 0.3.0 release candidate on 30 September, before the quality default was merged. Its skill named "quick, lean" as the default mode: work without a separate critic round, check before one full render. Frozen snapshot of the checkout without benchmarks/, its own runtime folder.

All eight tasks; twice on t1, t2, t8

showtime 0.2.0, default

The round-3 cells of record: the videos voted on in round 3 (the rerun videos for t1, t2, t3 and t8, the original round-3 videos for t4 and t6). Reused, not rerun.

The six older tasks

plain Opus 5.5

Claude Code 2.1.283 with claude-opus-5-5 and no video skill or plugin; Claude Code's bundled skills stayed installed. Reused from round 1 on the six older tasks, new runs on t9 and t10.

All eight tasks

brag-slim

The one-file launch-video skill from the brag project (MIT). It won the launch task in the blind vote in rounds 1, 2 and 3. Reused from round 1.

Launch video (t1) only

showtime 0.3.0, quality

The shipped default: a critic round on every finished video. Not in this round. One paid proof run on t2 ($3.03), not voted and not ranked.

Not benchmarked

How each new run worked

  1. A fresh workspace and a fresh Claude Code setup per run, with its own config folder and home, on one Linux cloud machine, three runs at a time.
  2. The same everything else. Model claude-opus-5-5 at high effort, Claude Code 2.1.283, a 25-minute cap, a $15 budget cap, the same tools on the PATH and the same standard reply to any question. No run asked a question and no run hit a cap.
  3. Isolation checked after the fact. Every command and file read of every new run was scanned for another showtime folder (round 3's failure). See Release checks.

How the outputs were scored

LayerWhat it doesBlind?
Automatic metricsFormat checks, showtime qa on a copy of every file (reused files re-checked with this round's qa), a headless-browser probe for the HTML report, and process numbers from each run's stream.not needed
Fact-checkA judge lists every factual claim and marks it supported, unsupported or contradicted by the task's own files. New cells only; reused cells keep their round-3 or round-1 fact-check.yes
Ranking judgeThree judgments per task, each seeing every version as video-1..n in rotated order (frames, transcripts, measurements), scoring seven criteria and ranking them. Each judgment had to prove it opened every frame (tool-call log) and read a random number stamped on each: 591 of 591 images opened, 589 of 591 numbers right, 24 judgments, none redone.yes, positions rotated
Blind human voteOne reviewer, who is showtime's author. One board per task: the versions re-encoded alike under shuffled letters, "which would you ship?" pairs, 0-5 ratings, notes and a first pick. The t10 board showed the reference as "the style to match (not a candidate)". The letter key stayed apart until the votes were imported. The reviewer gave no ratings on the t1 and t10 boards.partly (see Limits)

The two new tasks

Both inputs were written for the benchmark. t9 uses a small working command-line tool (a README, a changelog, four modules and tests) that catches drift between .env and .env.example. t10 uses a 20-second reference drawn with ffmpeg only (flat colour, one open-licence typeface, a synthesised beat) and an event description. showtime 0.3.0 has a repo-explainer workflow and a reference workflow that plain Opus 5.5 does not; that is what these tasks test.

Changed after the round

What changed because of round 4, not retested

These changes were made after the vote, from the reviewer's notes and the losses. None has been through a blind vote or a judged round.

Limits

Read these before quoting a number