The blind vote preferred showtime 0.3.0 (lean) over 0.2.0 on 4 of 6 tasks; the AI judge preferred 0.2.0 on 5
What was compared
showtime 0.3.0, lean review mode: new runs. At run time lean was showtime's default; it has since become an explicit opt-in.
showtime 0.2.0, default mode: the round-3 videos, reused, made from the same prompts. They are the videos voted on in round 3.
plain Opus 5.5: the same model in Claude Code with no video plugin. Reused from round 1 on the six older tasks; new runs on the two new tasks.
brag-slim: a launch-video-only skill, on the launch task (t1) only. Reused from round 1.
The 0.3.0 that ships uses the quality review mode by default: a separate critic round on every finished video. That default was not benchmarked in a blind round. The only quality-mode number here is one paid proof run on the data story (t2), at $3.03 API-equivalent; it was not on the blind boards and not ranked by the judge. Every table and chart below names the mode.
The reviewer's blind vote, pair by pair: 0.3.0 lean beat 0.2.0 default on t1, t4, t6 and t8 and lost t2 and t3. It beat plain Opus 5.5 on 7 of 8 tasks and lost the "make it like this reference" task (t10). On the launch video it beat brag-slim, the first time showtime has won that pair in a blind vote. The verified AI judge saw it differently: it ranked 0.2.0 above 0.3.0 lean on five of the six shared tasks, for polish reasons such as still holds, a blurred crossfade and a missing year label.
That disagreement is the main result, and it is why the shipped default is the quality mode and not lean. The voter is one person, showtime's author, and the vote is only partly blind.
4 of 6
blind-vote pairs won by 0.3.0 lean against 0.2.0 default (t1, t4, t6, t8; lost t2, t3)
7 of 8
blind-vote pairs won by 0.3.0 lean against plain Opus 5.5 (lost t10)
1 of 6
tasks where the AI judge ranked 0.3.0 lean above 0.2.0 default (t4, the filler cut)
1st time
showtime beat brag-slim on the launch video in a blind vote, after three losses
Small sample.8 tasks, 1 run per cell (2 for 0.3.0 on three tasks), and 1 human voter, who is showtime's author and had seen the older videos. Treat each task as one anecdote.
At a glance
The vote went to 0.3.0 lean on five tasks, to 0.2.0 on two, to plain Opus on one
First pick and the "which would you ship?" pairs are from the reviewer's blind boards (shuffled letters, no tool names). The judge's mean rank is over three rotated judgments per task; 1 is best, and the number of versions on the task is given with each rank.
Round 4 per task: blind vote and judge
Task
Blind vote
AI judge, mean rank (1 = best)
First pick
0.3.0 lean vs 0.2.0
0.3.0 lean vs plain
0.3.0 lean
0.2.0 default
plain Opus
t1 launch video
0.3.0 lean
WON
WON
4.00 of 5
1.33
5.00
t2 data story
0.2.0 default
LOST
WON
3.00 of 4
1.33
4.00
t3 vertical short
0.2.0 default
LOST
WON
2.67 of 3
1.67
1.67
t4 filler cut
0.3.0 lean
WON
WON
1.33 of 3
1.67
3.00
t6 HTML report
0.3.0 lean
WON
WON
2.00 of 3
1.00
3.00
t8 math explainer
0.3.0 lean
WON
WON
3.00 of 4
1.67
4.00
t9 repo explainer (new)
0.3.0 lean
not run
WON
1.00 of 2
not run
2.00
t10 like a reference (new)
plain Opus 5.5
not run
LOST
2.00 of 2
not run
1.00
t9 (explain a repo) and t10 (make it like a reference) are new in round 4 and were run with 0.3.0 lean and plain Opus 5.5 only. The second 0.3.0 lean run on t1, t2 and t8 was judged but kept off the boards, to keep the vote short.
Vote and judge
The reviewer and the judge disagree about 0.3.0 lean against 0.2.0
Against plain Opus 5.5 the two instruments mostly agree: the vote went to 0.3.0 lean on 7 of 8 tasks and the judge ranked it higher on 6 of 8 (not on t3 and t10). Against 0.2.0 default they point in opposite directions: 4 of 6 for 0.3.0 lean in the vote, 1 of 6 with the judge.
Comparisons won by 0.3.0 lean, by instrument
Blind vote: "which would you ship?" pairs. AI judge: tasks where its mean rank was better. Drawn to scale.
Reading this. Each bar is a count of tasks, one pair or one ranking per task. A task can flip on a single run. The vote is one person's. The judge ranks still frames, a transcript and measurements; it cannot watch motion or hear sound.
On the five tasks where both were rated, the reviewer's 0-5 ratings averaged 4.2 for 0.3.0 lean and 4.4 for 0.2.0 default. The ratings tied on t4 and t6, where the pair went to 0.3.0 lean. So the ratings lean slightly to 0.2.0 while the pairs lean to 0.3.0.
What the judge held against 0.3.0 lean
Still holds. On the launch video, three stretches of 3.6 to 4.6 s with no change on screen (qa flagged them too) and an empty terminal beat. On the data story, the ranked bar chart stays on screen for about 15 seconds.
A blurred crossfade. On the HTML report, one sampled frame caught two scenes blurred over each other mid-transition.
A missing label. On the same report, the "first year above 400 ppm" callout does not say which year. The 0.2.0 page says 2015.
Crowded moments. On the math explainer, equations overlap into a jumble during a transition around 35 s. On the vertical short, a left-heavy layout leaves the right side and bottom half empty.
Less content. On the launch video, fewer of the 2.0 features; on the data story, a thinner story than 0.2.0's four beats.
These are polish defects, the kind a separate review of the finished video exists to catch. The lean mode in round 4 skipped that separate critic round to save cost. That is why the shipped default became the quality mode, which runs a critic round on every finished video. Whether it removes these defects has been checked on one proof run only, not in a blind round.
AI judge: mean rank per task, every version
Three judgments per task; further left is better; the grey track spans the versions on that task.
0.3.0 lean0.3.0 lean, 2nd run0.2.0 defaultplain Opus 5.5brag-slim
Two runs of the same build differ a lot. 0.3.0 lean was run twice on t1, t2 and t8. On the math explainer the second run ranked first (mean 1.33), ahead of 0.2.0 (1.67), while the first run ranked 3.00. On the launch video the second run ranked 2.67 against the first run's 4.00. On all three tasks the two runs of the same build were 1.33 to 1.67 places apart; on t8 that was more than the gap between 0.3.0 lean and 0.2.0. Read every single-task result here as one draw.
The vote only saw the first run of each.
The launch video (t1)
showtime beat brag-slim in a blind vote for the first time; the judge put it fourth of five
In rounds 1, 2 and 3 the reviewer picked brag-slim on the launch video every time. In round 4 the reviewer picked 0.3.0 lean, and preferred it to brag-slim, to 0.2.0 default and to plain Opus 5.5 in all three pairs asked.
Two things temper that. First, the voter's view of the unchanged videos moved too: the same brag-slim and 0.2.0 files were on the round-3 board, where brag-slim won that pair; this time 0.2.0 won it, and plain Opus 5.5's silent video also beat brag-slim. Second, the judge ranked 0.3.0 lean 4th of 5 in all three judgments, above plain Opus 5.5 only, citing still holds and fewer 2.0 features. It ranked 0.2.0 first (mean 1.33) and brag-slim second (2.00). The board for this task has pairs and a pick but no 0-5 ratings.
Results by task
Eight tasks, one run each
Blind vote: the reviewer's 0-5 rating, first pick, and pairs won on the board. Judge rank: mean of three rotated judgments, with each judgment's place. Cost: the API-equivalent USD computed from the run's stream; the runs used a subscription and were not billed per token. Wall: the whole run. Reused cells keep the numbers of the round they were made in, on other days and under other load. qa: showtime qa from this round's checkout, the same thresholds for every file. Flagged claims: claims the fact-check judge marked unsupported or contradicted by the task's own files. Reviewer notes are quoted with spelling corrected.
t1 · launch video from a repo
Launch video for quillsort 2.0
Make a 30-second launch video for quillsort 2.0 from this repo that I can post on X.
Votepick 0.3.0 lean; won all 3 of its pairs
Judge1st 0.2.0 default · 0.3.0 lean 4th of 5
plain Opus 5.5 frame at 24 s
brag-slim frame at 12 s
0.2.0 default frame at 6 s
0.3.0 lean frame at 18 s
plain Opus 5.5reused, round 1
brag-slimreused, round 1
0.2.0 defaultreused, round 3
0.3.0 lean
0.3.0 lean, 2nd run
Blind vote
pairs won: 1 of 2no ratings on this board
pairs won: 0 of 3no ratings on this board
pairs won: 1 of 2no ratings on this board
pickpairs won: 3 of 3no ratings on this board
not on the board
Judge rank
mean 5.00of 5; judgments: 5th, 5th, 5th
mean 2.00of 5; judgments: 1st, 2nd, 3rd
mean 1.33of 5; judgments: 2nd, 1st, 1st
mean 4.00of 5; judgments: 4th, 4th, 4th
mean 2.67of 5; judgments: 3rd, 3rd, 2nd
Cost
$0.88
$2.09
$3.73
$4.01
$3.81
Wall time
4.4 min
10.0 min
14.4 min
14.1 min
12.5 min
Turns
19
36
13
7
9
qa
FAILsilent audio track
WARNcolour tags missing
PASS
WARNthree still holds of 3.6-4.6 s
WARNstill holds of up to 4.9 s
Loudness
no audio
−14.1 LUFS
−14.0 LUFS
−14.0 LUFS
−14.0 LUFS
Flagged claims
0
1demo output shown unsorted
0
0
0
Review notes
No note and no ratings on this board. The reviewer preferred 0.3.0 lean in each pair, 0.2.0 over brag-slim, and plain Opus 5.5 over brag-slim.
Judge's reasoning, in short
All three judgments put 0.3.0 lean 4th: a section on --unique not shown as new, an empty terminal beat, holds of about 4 s, fewer new features, and a caption area left empty for several seconds. 0.2.0 was first twice for the tightest pacing and correct terminal output. brag-slim had the strongest hook but shows a dedupe result out of order. plain Opus 5.5 is silent, which the judge called a hard fail for a post on X.
t2 · data story from a CSV
45-second temperature data story
Turn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube.
Votepick 0.2.0 default; 0.3.0 lean lost to it, beat plain Opus
Judge1st 0.2.0 default · 0.3.0 lean 3rd of 4
plain Opus 5.5 frame at 36 s
0.2.0 default frame at 27 s
0.3.0 lean frame at 18 s
0.3.0 lean, 2nd run (not voted) frame at 9 s
plain Opus 5.5reused, round 1
0.2.0 defaultreused, round 3
0.3.0 lean
0.3.0 lean, 2nd run
Blind vote
★★★★★pairs won: 0 of 2
★★★★★pickpairs won: 2 of 2
★★★★★pairs won: 1 of 2
not on the board
Judge rank
mean 4.00of 4; judgments: 4th, 4th, 4th
mean 1.33of 4; judgments: 2nd, 1st, 1st
mean 3.00of 4; judgments: 3rd, 3rd, 3rd
mean 1.67of 4; judgments: 1st, 2nd, 2nd
Cost
$1.33
$5.63
$1.42
$1.71
Wall time
8.5 min
21.3 min
8.9 min
9.1 min
Turns
30
20
41
37
qa
WARNno audio track; flat first frame; holds up to 5.3 s
PASS
PASS
PASS
Loudness
no audio
−14.0 LUFS
−14.0 LUFS
−14.1 LUFS
Flagged claims
0
1credits NOAA, which the CSV does not name
0
0
Review notes
Neither was the best, but C was cleaner; in B and C both, numbers sometimes disappear and then reappear when a new item was added.The reviewer. C was 0.2.0 default, B was 0.3.0 lean: the vanishing numbers were in both showtime videos.
Judge's reasoning, in short
0.2.0 told the fullest story in four beats: a yearly line, decade averages, the top-11 ranking and an end card with its source. 0.3.0 lean was clean and accurate but thinner, with the ranked bar chart held for about 15 s and odd 22-year axis ticks. The second 0.3.0 lean run, a bold dark design with a running year counter, was first once and second twice. plain Opus 5.5 has no audio track, which a YouTube video needs.
t3 · vertical short with voice-over and captions
30-second Reels short for quillsort 2.0
Make a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions.
Votepick 0.2.0 default; 0.3.0 lean lost to it, beat plain Opus
Judge0.2.0 default and plain Opus 5.5 tied first · 0.3.0 lean 3rd of 3
plain Opus 5.5 frame at 6 s
0.2.0 default frame at 0.6 s
0.3.0 lean frame at 12 s
plain Opus 5.5reused, round 1
0.2.0 defaultreused, round 3
0.3.0 lean
Blind vote
★★★★★pairs won: 0 of 2
★★★★★pickpairs won: 2 of 2
★★★★★pairs won: 1 of 2
Judge rank
mean 1.67of 3; judgments: 1st, 1st, 3rd
mean 1.67of 3; judgments: 2nd, 2nd, 1st
mean 2.67of 3; judgments: 3rd, 3rd, 2nd
Cost
$1.81
$2.41
$2.66
Wall time
11.0 min
11.3 min
12.3 min
Turns
46
44
56
qa
WARNlong caption line; flat first frame; last 2 s silent
PASS
PASS
Loudness
−14.0 LUFS
−14.0 LUFS
−14.0 LUFS
Flagged claims
0
0
2both false flags: the README, a listed source, states Python 3.8+ and pip install quillsort
Review notes
No note. The reviewer rated 0.2.0 5 of 5, 0.3.0 lean 4 and plain Opus 5.5 2.
Judge's reasoning, in short
Two judgments put 0.3.0 lean last: a left-heavy layout that leaves the right side and bottom half empty, a caption pressed against a card, and an in-place edit demo whose backup file shows only two lines, which looks broken. One judgment called it the strongest graphic design and put it second. plain Opus 5.5 was first twice for its clear one-card-per-change structure; the judge cannot hear the voice, which the reviewer marked down.
t4 · footage edit
Cut the fillers and long pauses from an interview
Cut the filler words and long pauses out of interview.mp4 and give me the tightened video.
0.3.0 lean, frame at 9.9 s. The other cuts show the same camera picture.
plain Opus 5.5reused, round 1
0.2.0 defaultreused, round 3
0.3.0 lean
Blind vote
★★★★★pairs won: 0 of 2
★★★★★pairs won: 1 of 2
★★★★★pickpairs won: 2 of 2
Judge rank
mean 3.00of 3; judgments: 3rd, 3rd, 3rd
mean 1.67of 3; judgments: 2nd, 1st, 2nd
mean 1.33of 3; judgments: 1st, 2nd, 1st
Cost
$1.06
$1.72
$1.57
Wall time
7.2 min
6.3 min
7.1 min
Turns
23
48
49
qa
FAILkept a 5.2 s silence and frozen picture
WARN2 LU quiet (−16 LUFS); a 2.2 s silence inside; silent tail
WARNa 2.2 s silence inside; 2.3 s silent tail
Loudness
−16.0 LUFS
−16.0 LUFS
−14.0 LUFS
Review notes
I think they were all great. It was hard to pick a winner, but they were all very good.The reviewer (shortened). Rated 0.3.0 lean and 0.2.0 5 of 5 each, plain Opus 5.5 4.
Judge's reasoning, in short
Both showtime cuts removed every audible "uh" and the stutter, came in at about 49.5 s and left a 2.2 s silence and a silent tail. 0.3.0 lean was first twice because it is normalised to −14 LUFS (0.2.0 stayed at the source's −16). One judgment preferred 0.2.0 because 0.3.0 lean cut an "and" with an "uh". plain Opus 5.5 kept the fillers and a 5.2 s silence.
t6 · single-file HTML video report
Shareable HTML report on the Mauna Loa CO2 record
Make a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team.
Votepick 0.3.0 lean; won both of its pairs
Judge1st 0.2.0 default in all three · 0.3.0 lean 2nd
plain Opus 5.5 page on load
0.2.0 default page on load
0.3.0 lean page on load
plain Opus 5.5reused, round 1
0.2.0 defaultreused, round 3
0.3.0 lean
Blind vote
★★★★★pairs won: 0 of 2
★★★★★pairs won: 1 of 2
★★★★★pickpairs won: 2 of 2
Judge rank
mean 3.00of 3; judgments: 3rd, 3rd, 3rd
mean 1.00of 3; judgments: 1st, 1st, 1st
mean 2.00of 3; judgments: 2nd, 2nd, 2nd
Cost
$1.40
$0.98
$0.91
Wall time
6.5 min
4.6 min
6.2 min
Turns
29
23
29
HTML checks
PASSsingle file, offline, plays, phone width
PASSsingle file, offline, plays, phone width
PASSsingle file, offline, plays, phone width
Flagged claims
3a NOAA credit, a 1970 value off by 0.04 ppm, a 'longest record' line
0
0
Review notes
B and C are almost the same. They were both a little bit slow … not very dynamic, but I like them both.The reviewer (shortened). B was 0.3.0 lean, C was 0.2.0 default; both rated 4 of 5.
Judge's reasoning, in short
The two showtime pages share one editorial design and the same correct figures. 0.2.0 was first in all three: its "2015: first year above 400" label is exact and it runs a tight 25 s. 0.3.0 lean adds a helpful 1959 baseline line, but its 400 ppm callout gives no year and one frame caught a heavy blurred crossfade. plain Opus 5.5's page has the most data but looks like a generic dashboard, with small text inside the player and no audio.
t8 · math explainer
Why the first n odd numbers add up to n²
Make a 45-second video that shows visually why the sum of the first n odd numbers is n squared.
FAILblack first frame; no audio track; still holds
WARNone caption cue under 0.4 s
PASS
PASS
Loudness
no audio
−14.0 LUFS
−14.0 LUFS
−14.0 LUFS
Flagged claims
0
0
0
0
Review notes
B was the best … the animation was different, but A was a little bit more clear. One weird thing about A is that the agent says "1 1 + 3", like that, it's kinda weird, but it was well explained in the end.The reviewer (shortened). B was 0.3.0 lean, A was 0.2.0 default. The remark about "1 1 + 3" is about the 0.2.0 video. The judges' speech check found no voice-over in it; it came with a captions file whose first cue reads "One.", so the remark may be about its captions.
Judge's reasoning, in short
All four give a correct L-shaped proof. The second 0.3.0 lean run was first twice for step titles that guide the viewer. 0.3.0 lean (the voted run) was third in all three: no guiding headings, and equations that overlap into a jumble during a transition around 35 s. plain Opus 5.5 has no audio track and starts and ends on black.
t9 · explain a repo · new in round 4
A 60-90 second explainer of an unfamiliar repo
Make a 60-90 second video that explains what this repo does and how it works, for developers seeing it for the first time.
Votepick 0.3.0 lean (5 of 5 against 4)
Judge1st 0.3.0 lean in all three
plain Opus 5.5 frame at 51 s
0.3.0 lean frame at 51.8 s
plain Opus 5.5
0.3.0 lean
Blind vote
★★★★★pairs won: 0 of 1
★★★★★pickpairs won: 1 of 1
Judge rank
mean 2.00of 2; judgments: 2nd, 2nd, 2nd
mean 1.00of 2; judgments: 1st, 1st, 1st
Cost
$1.44
$3.42
Wall time
11.2 min
23.0 min
Turns
30
61
qa
FAILno audio track; a 6.9 s frozen picture
PASS
Loudness
no audio
−14.0 LUFS
Flagged claims
1output of a diff against a file that is not in the repo
21 real: it says extra keys are warnings unless strict; extra is always a warning. The other (a __main__.py) exists in the repo
Review notes
No note. Two versions only, one pair.
Judge's reasoning, in short
plain Opus 5.5 explains a little more (exit codes, the diff and init commands, a pre-commit hook) but has no audio track and long frozen holds, one of 6.9 s. 0.3.0 lean walks the parse, compare and report steps with a module tracker, has music and captions, and ends on where to start reading; some code text is small for a phone. Neither has narration. The fact-check found one real imprecision in 0.3.0 lean (see the table).
t10 · make it like a reference · new in round 4
A 20-second announcement in the style of a reference video
Make a 20-second video announcing our spring plant swap (details in swap.md) in the same style as reference.mp4.
Votepick plain Opus 5.5; 0.3.0 lean lost its only pair
Judge1st plain Opus 5.5 in all three
The reference (not a candidate), drawn for the benchmark reference, frame at 17 s
plain Opus 5.5 frame at 19.6 s
0.3.0 lean frame at 19.6 s
plain Opus 5.5
0.3.0 lean
Blind vote
pickpairs won: 1 of 1no ratings on this board
pairs won: 0 of 1no ratings on this board
Judge rank
mean 1.00of 2; judgments: 1st, 1st, 1st
mean 2.00of 2; judgments: 2nd, 2nd, 2nd
Cost
$1.62
$0.98
Wall time
5.6 min
4.0 min
Turns
44
29
qa
FAIL4.8 LU too quiet (−18.8 LUFS)
PASS
Loudness
−18.8 LUFS
−14.1 LUFS
Flagged claims
0
0
Review notes
B was much closer to the reference and well explained. It went a little silent at the end, but it was good; it kept the colours and the idea.The reviewer (spelling corrected). B was plain Opus 5.5, the winner; the silence is in its video. No ratings on this board.
Judge's reasoning, in short
Both kept the reference's structure: a word-by-word opener, numbered cards with a progress bar, a wipe and a dark end card. plain Opus 5.5 kept the reference's cream, red and black palette and its very large type. 0.3.0 lean changed the accent to green to suit the plants and added small sentence-case lines the reference never uses; three of its six sampled stills are full-screen green wipes. plain Opus 5.5 was quiet (−18.8 LUFS, a qa FAIL); 0.3.0 lean was on target.
Release checks
Six of eight release checks passed; cost and the strict quality rule failed
Before the round, 0.3.0 set itself eight checks. They were computed by the benchmark's own report script from the round's files, on the 0.3.0 lean runs.
Round 4 release checks
Check
Result
Detail
Cost: median at or under plain Opus 5.5
FAIL
0.3.0 lean $1.51 against plain Opus 5.5 $1.37, on the same 8 tasks (API-equivalent, from each run's stream)
Quality: judge and vote both at or above 0.2.0 on 5 of 6 tasks
FAIL
1 of 6. Only t4 passed both. The vote alone was 4 of 6 (t1, t4, t6, t8); the judge alone 1 of 6. The rule needs both to agree on each task.
Images read into the main context: median at most 12 per job
PASS
median 4.5 (t4 highest at 11)
Full renders per job: median 1
PASS
median 1 (from each job's receipt; t4 rendered 4 times, t1 3, t8 0 full renders)
qa FAIL findings: 0 on every new 0.3.0 run
PASS
10 video runs checked, none failed; the HTML report has no qa verdict
Launch video (t1): not last
PASS
judge mean rank 4.0 of 5 (plain Opus 5.5 last at 5.0); blind pairs won 3 of 3
Every job writes a receipt
PASS
11 of 11 new 0.3.0 runs, all with a cost
Isolation: no run touched another showtime folder
PASS
13 new runs checked. The first pass flagged four showtime runs; that was a checker bug (it read encoded command text and counted the shared tools linked into the arm's own home as foreign), and the fixed checker found none.
Round 3 had a real isolation failure: two of its runs used another working copy of showtime. The isolation check was added for round 4 to catch that; it found nothing once its own bug was fixed.
Process and cost
Lean is the cheaper path; quality mode spends more for a critic round
Cost here is API-equivalent USD computed from each run's stream of model calls. The runs used a subscription, so nobody was billed these amounts. Mode matters: lean skips the separate review round, quality, the shipped default, adds a critic round on every finished video and costs more for it. On the one task measured all three ways, the data story (t2), 0.3.0 lean cost $1.42, 0.3.0 quality $3.03 (one proof run) and 0.2.0 default $5.63.
Cost per run, by task
API-equivalent USD; one run per bar; drawn to scale.
plain Opus 5.50.2.0 default0.3.0 lean0.3.0 quality (t2 proof run)
Median cost and wall time
Version
Tasks
Median cost
Median wall
0.3.0 lean
all 8
$1.51
8.0 min
plain Opus 5.5
all 8
$1.37
6.9 min
0.3.0 lean
6 older tasks
$1.51
8.0 min
0.2.0 default
6 older tasks
$2.38
12.8 min
plain Opus 5.5
6 older tasks
$1.20
6.9 min
0.3.0 qualitythe shipped default
t2 only, one proof run
$3.03
—
0.3.0 lean cost less than plain Opus 5.5 on t6 and t10, within $0.10 on t2, and more on the other five, most on the launch video ($4.01) and the repo explainer ($3.42), its two longest runs. Both 0.3.0 lean launch runs used helper agents, as the 0.2.0 launch run did. Wall times of reused cells come from other days and other machine load, so compare them loosely.
Other process numbers from the stream, for 0.3.0 lean: a median of 4.5 images read into the main agent's context per job and a median of one full render. The two runs of the same task were close in cost ($4.01 and $3.81 on t1, $1.42 and $1.71 on t2, $1.45 and $1.70 on t8).
Method
How round 4 was run
The same harness as rounds 1 to 3 (described in the v0.2.0 report), driven from a manifest that lists which cells are new and which are reused. The harness, tasks and scorers are in the showtime repository under benchmarks/.
The contenders
showtime 0.3.0, lean
The 0.3.0 release candidate on 30 September, before the quality default was merged. Its skill named "quick, lean" as the default mode: work without a separate critic round, check before one full render. Frozen snapshot of the checkout without benchmarks/, its own runtime folder.
The round-3 cells of record: the videos voted on in round 3 (the rerun videos for t1, t2, t3 and t8, the original round-3 videos for t4 and t6). Reused, not rerun.
Claude Code 2.1.283 with claude-opus-5-5 and no video skill or plugin; Claude Code's bundled skills stayed installed. Reused from round 1 on the six older tasks, new runs on t9 and t10.
The shipped default: a critic round on every finished video. Not in this round. One paid proof run on t2 ($3.03), not voted and not ranked.
Not benchmarked
How each new run worked
A fresh workspace and a fresh Claude Code setup per run, with its own config folder and home, on one Linux cloud machine, three runs at a time.
The same everything else. Model claude-opus-5-5 at high effort, Claude Code 2.1.283, a 25-minute cap, a $15 budget cap, the same tools on the PATH and the same standard reply to any question. No run asked a question and no run hit a cap.
Isolation checked after the fact. Every command and file read of every new run was scanned for another showtime folder (round 3's failure). See Release checks.
How the outputs were scored
Layer
What it does
Blind?
Automatic metrics
Format checks, showtime qa on a copy of every file (reused files re-checked with this round's qa), a headless-browser probe for the HTML report, and process numbers from each run's stream.
not needed
Fact-check
A judge lists every factual claim and marks it supported, unsupported or contradicted by the task's own files. New cells only; reused cells keep their round-3 or round-1 fact-check.
yes
Ranking judge
Three judgments per task, each seeing every version as video-1..n in rotated order (frames, transcripts, measurements), scoring seven criteria and ranking them. Each judgment had to prove it opened every frame (tool-call log) and read a random number stamped on each: 591 of 591 images opened, 589 of 591 numbers right, 24 judgments, none redone.
yes, positions rotated
Blind human vote
One reviewer, who is showtime's author. One board per task: the versions re-encoded alike under shuffled letters, "which would you ship?" pairs, 0-5 ratings, notes and a first pick. The t10 board showed the reference as "the style to match (not a candidate)". The letter key stayed apart until the votes were imported. The reviewer gave no ratings on the t1 and t10 boards.
partly (see Limits)
The two new tasks
Both inputs were written for the benchmark. t9 uses a small working command-line tool (a README, a changelog, four modules and tests) that catches drift between .env and .env.example. t10 uses a 20-second reference drawn with ffmpeg only (flat colour, one open-licence typeface, a synthesised beat) and an event description. showtime 0.3.0 has a repo-explainer workflow and a reference workflow that plain Opus 5.5 does not; that is what these tasks test.
Changed after the round
What changed because of round 4, not retested
These changes were made after the vote, from the reviewer's notes and the losses. None has been through a blind vote or a judged round.
Quality is the default review mode. A critic round on every finished video; lean is an explicit opt-in. Measured once, on t2, at $3.03; not voted.
Data charts and HTML reports. From the t2 and t6 notes: numbers that vanish and reappear as items are added, and pacing that felt slow and not dynamic. The fixes also cover the judge's points: no blur over text, labels that carry their year or value, no 4-5 s holds.
The "like a reference" route (t10). The reference route is now required for such requests: a measured spec of the reference (cuts, timing, palette, text sizes, sound hits), a keep-or-change list, and a frame-by-frame comparison with the reference.
Limits
Read these before quoting a number
Lean, not the shipped default. Round 4 measured 0.3.0 in lean mode. The shipped default, quality mode, has one proof run on one task and no vote.
One run per cell. Two for 0.3.0 lean on t1, t2 and t8, and those two runs were 1.33 to 1.67 places apart in the judge's mean rank (t8: 3.00 against 1.33). Any single task can flip on another draw.
One human voter, who is showtime's author, and only partly blind. Letters were shuffled and nothing named a tool, but the reviewer had seen the reused 0.2.0, plain Opus 5.5 and brag-slim videos in earlier rounds and could recognise them. On t1 the vote on those unchanged videos moved from round 3. A claim that matters needs voters who built none of the tools.
Reused cells. 0.2.0 default comes from round 3; plain Opus 5.5 (six tasks) and brag-slim from round 1, same prompts, made on other days under other load. Only showtime 0.3.0 and the two new plain Opus 5.5 cells were run in round 4.
The judge's limits. It sees still frames, a transcript and measurements; it cannot watch motion or hear voice and music. Its three judgments per task are one model looking at rotated copies of one packet. The proof shows it opened and read the images, not that it judged well. The judge model is the same family as the agents that made the videos.
The fact-check judge makes mistakes. Both claims it flagged in 0.3.0 lean's vertical short are stated in the README, a listed source; one of the two in the repo explainer is real.
t9 and t10 have fewer comparisons. Two versions each and one pair each; no 0.2.0 and no outside tool. The t1 and t10 boards have no ratings.
Costs are API-equivalent estimates from the stream, at API prices. The runs used a subscription. Reused cells keep the cost and time of their own round.
The same author wrote the tasks, the harness, the scorers and the release checks, and showtime qa is the tested tool's own checker, applied the same way to every file. Nobody independent has reviewed the method.
Bundled skills and the network. Claude Code's own skills stayed installed for every contender, no one had a cloud key, and network access was not blocked.
Earlier rounds are in the v0.2.0 report, with their own limits.