showtime benchmark, rounds 1 to 3
showtime.
Benchmark · rounds 1-3 · 28-29 Sep 2026 · claude-opus-5-5

showtime benchmark, rounds 1 to 3: six video requests, five ways to answer them

Round 3, mixed: the judge put 0.2.0 below 0.1.0 on four tasks, the author's vote picked it on three

Round 3 (showtime 0.2.0): two of six tasks came out better than showtime's earlier video, four worse, no level. showtime 0.2.0 was run once on the six tasks with the original prompts and ranked, by a blind judge that had to prove it opened every frame, against its own round-1 or round-2 video and the other contenders' existing outputs. It ranked first on the filler cut (all five fillers removed, was two) and moved the launch video from 3rd to 2nd of 4; it ranked below its earlier video on t2, t3, t6, t8. Cost and time per run went up. The blind boards were then voted on by one person, showtime's author, in a partly blind vote (he had seen the earlier videos): showtime was the first pick on five of six tasks, 0.2.0 on three of them, the 0.1.0 video on two, and brag-slim won the launch video again. Head to head the 0.2.0 video beat the 0.1.0 video in three of six pairs. On the filler cut the vote went against the scorer and the judges. The launch default was changed after the vote and has not been re-voted.

2 of 6
round-3 tasks where the blind judge ranked showtime 0.2.0 above showtime's own earlier video (t1, t4)
5 of 5 fillers
removed from the interview by showtime 0.2.0, judged 1st by all three judgments (0.1.0: 2 of 5)
528 of 528
frame images the judges provably opened (tool-call log), 526 of 528 reading-check numbers read correctly
$2.55 vs $1.61
median API-equivalent cost per run, showtime 0.2.0 vs the earlier showtime runs (11.8 vs 7.6 min)
3 of 6
round 3: tasks where the author's partly blind vote picked showtime 0.2.0 first (t3, t6, t8); the 0.1.0 video was picked on t2, t4
3 to 3
round 3 vote, showtime 0.2.0 against the 0.1.0 video, pair by pair (ratings 4.3 against 4.2 of 5)
6 of 6
round 3 vote: pairs the 0.2.0 video won against plain Opus 5.5; it lost the launch video to brag-slim, the third vote in a row
1 voter
the round-3 vote is showtime's author, who had seen the round-1 and round-2 videos: not a panel, and only partly blind

Round 1 and the rematch: six one-sentence video requests were typed into Claude Code and run once each with plain Opus 5.5, once with showtime, and once with the open-source tool built for that kind of job. A blind AI judge ranked showtime first on 4 tasks and second on 2.

showtime lost the launch video to brag-slim and the filler cut to plain Opus 5.5. After fixes, those two cells were rerun as a rematch: the reviewer preferred showtime on the filler cut (a partly blind vote) and brag-slim on the launch again. The judge put showtime second in both.

Small sample.6 tasks, 3 contenders each, 1 run per cell, and 1 human reviewer, who is showtime's author. Treat each task as one anecdote. The totals show where to look, not what is true in general.
4 of 6
round 1: tasks where the blind reviewer's first pick was showtime
4× 1st 2× 2nd
round 1: showtime's place with the blind ranking judge (frames not always seen)
0 of 5
showtime video files that failed the loudness, black-frame and silence checks (plain Opus 5.5: 3 of 5)
$1.66 vs $1.20
median API-equivalent cost per run, showtime vs plain Opus 5.5

At a glance

Who won each task

The reviewer's pick is the video approved on a blind board with shuffled letters and no tool names. The judge is one AI ranking per task that saw every output side by side, in shuffled order, without names.

Launch video from a repo

Reviewer's pickbrag-slim
Judge's 1stbrag-slim
showtimereviewer 3rd · judge 2nd
Rematch
Reviewer's pickbrag-slim again
showtimejudge 2nd, flagged claims 4 → 0

Data story from a CSV

Reviewer's pickshowtime
Judge's 1stshowtime
Againsthyperframes, plain Opus 5.5

Vertical short with voice-over

Reviewer's pickshowtime
Judge's 1stplain Opus 5.5
showtimereviewer 1st · judge 2nd

Cut fillers from a talking clip

Reviewer's pickplain Opus 5.5
Judge's 1stshowtime
showtimereviewer 3rd · judge 1st
Rematch
Reviewer's pickshowtime (partly blind)
Judge's 1stvideo-use; showtime 2nd

Single-file HTML video report

Reviewer's pickshowtime
Judge's 1stshowtime
Againsthyperframes, plain Opus 5.5

Math explainer

Reviewer's pickshowtime
Judge's 1stshowtime
Againstvideo-use*, plain Opus 5.5

* The video-use run never called the video-use skill on this task, so in practice it was a second plain run.

Round 1 by contender

Reviewer's first picks

Share of the tasks each contender entered where the blind reviewer picked it

0% 50% 100% plain Opus 5.5: picked in 1 of 6 tasks plain Opus 5.5 1 of 6 showtime: picked in 4 of 6 tasks showtime 4 of 6 brag-slim: picked in 1 of 2 tasks brag-slim 1 of 2 hyperframes: picked in 0 of 2 tasks hyperframes 0 of 2 video-use: picked in 0 of 2 tasks video-use 0 of 2

Reviewer's head-to-head wins

Share of “which would you ship?” pairs won

0% 50% 100% plain Opus 5.5: won 5 of 12 pairs plain Opus 5.5 5 of 12 showtime: won 8 of 12 pairs showtime 8 of 12 brag-slim: won 3 of 4 pairs brag-slim 3 of 4 hyperframes: won 0 of 4 pairs hyperframes 0 of 4 video-use: won 2 of 4 pairs video-use 2 of 4

Judge's mean rank

Blind ranking judge, one judgment per task; further right is better

3rd 2nd 1st plain Opus 5.5: mean rank 2.50 over 6 tasks plain Opus 5.5 2.50 showtime: mean rank 1.33 over 6 tasks showtime 1.33 brag-slim: mean rank 2.00 over 2 tasks brag-slim 2.00 hyperframes: mean rank 2.50 over 2 tasks hyperframes 2.50 video-use: mean rank 2.00 over 2 tasks video-use 2.00

Median cost per run

API-equivalent USD reported by Claude Code (subscription runs)

$0 $4 $8 $12 plain Opus 5.5: median $1.20 over 6 runs plain Opus 5.5 $1.20 showtime: median $1.66 over 6 runs showtime $1.66 brag-slim: median $2.16 over 2 runs brag-slim $2.16 hyperframes: median $10.66 over 2 runs hyperframes $10.66 video-use: median $1.42 over 2 runs video-use $1.42

Median wall time per run

Minutes from start to finish, including any reply to a question

0 min 4 min 8 min 12 min 16 min plain Opus 5.5: median 6.9 min over 6 runs plain Opus 5.5 6.9 min showtime: median 7.7 min over 6 runs showtime 7.7 min brag-slim: median 13.3 min over 2 runs brag-slim 13.3 min hyperframes: median 15.5 min over 2 runs hyperframes 15.5 min video-use: median 9.2 min over 2 runs video-use 9.2 min

Reading these. Every mark is drawn to scale. showtime and plain Opus 5.5 entered all six tasks. Each other tool entered only the two tasks it was paired with, so its numbers come from a different, smaller set of tasks; compare it within those tasks below, not across this chart.

The rematch is left out here: it reran only showtime, on two tasks.

Round 1 totals per contender
ContenderTasksReviewer's 1st picksReviewer's head-to-head winsMean reviewer ratingJudge mean rankMedian costMedian wall timeqa FAIL
plain Opus 5.5all 61 of 65 of 123.02.50$1.206.9 min3 of 5
showtimeall 64 of 68 of 123.71.33$1.667.7 min0 of 5
brag-slimt1, t31 of 23 of 44.02.00$2.1613.3 min0 of 2
hyperframest2, t60 of 20 of 42.52.50$10.6615.5 min1 of 1
video-uset4, t80 of 22 of 43.02.00$1.429.2 min2 of 2

qa FAIL counts video files only; the HTML task has no qa verdict. Ratings are the reviewer's 0-5 stars on the blind boards. "Head-to-head wins" are the reviewer's answers to "which would you ship?" for every pair on a board.

Method

How the comparison was run

Every contender got the same one-sentence request, the same input files, the same model and the same limits. Only the video skill or plugin differed. The harness, tasks and scorers are in the showtime repository under benchmarks/.

The contenders

plain Opus 5.5

Claude Code 2.1.283 with the model claude-opus-5-5 and no video skill or plugin. Claude Code's own bundled skills were still present; it used the bundled dataviz skill on the data story and the HTML report.

All six tasks

showtime

The project under test: a Claude Code plugin for making video locally (HTML and canvas motion graphics, voice-over, music and sound, footage edits, captions). Tested as a frozen v0.1.0 pre-release snapshot loaded with --plugin-dir, without the benchmarks/ folder. The rematch used a new snapshot with the fixes.

All six tasks

brag-slim

The one-file skill from the brag project (MIT). It turns a project folder or a website into a short launch video with music, motion and share copy, built by the model with the tools already on the machine. Installed as shipped.

Launch video (t1), vertical short (t3)

hyperframes

HeyGen's open-source HTML-to-video framework for agents (Apache-2.0). Installed with 19 of its agent skills, its CLI (hyperframes@0.8.78 from npm) and its own headless browser.

Data story (t2), HTML report (t6)

video-use

Browser Use's open-source skill for editing footage by chatting with Claude Code (MIT). Installed as the whole repository plus its Python environment. Its filler detection expects an ElevenLabs transcript; no contender got any cloud key.

Filler cut (t4), math explainer (t8)

Each task got plain Opus 5.5, showtime, and the one outside tool whose home turf it is closest to. video-use is a footage editor, so the math task (no footage) is outside what it was built for; in that run it did not invoke its skill at all.

The six tasks, word for word

TaskPrompt typed into Claude CodeInputs in the workspaceContenders
t1 Launch videoMake a 30-second launch video for quillsort 2.0 from this repo that I can post on X.a small fictional Python repo: README, CHANGELOG, quillsort.py, landing pageplain, showtime, brag-slim
t2 Data storyTurn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube.NOAA global temperature departures, 1850-2025 (public domain)plain, showtime, hyperframes
t3 Vertical shortMake a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions.the same quillsort repoplain, showtime, brag-slim
t4 Filler cutCut the filler words and long pauses out of interview.mp4 and give me the tightened video.63.7 s of a NASA interview with STS-1 pilot Robert Crippen (public domain)plain, showtime, video-use
t6 HTML reportMake a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team.NOAA GML Mauna Loa annual CO2, 1959-2025plain, showtime, hyperframes
t8 Math explainerMake a 45-second video that shows visually why the sum of the first n odd numbers is n squared.noneplain, showtime, video-use

The repository defines eight tasks. This round was cut to these six to keep it to 18 runs; the narrated explainer (t5) and the logo sting (t7) were not run.

How each run worked

  1. A fresh workspace per run. A new git repository holding only the task's inputs, on a Linux cloud machine (32 vCPU), in a folder outside any home directory so no stray CLAUDE.md or skills folder could load.
  2. An isolated Claude Code setup per contender. Its own config folder and home, with opaque folder names so an agent could not tell which contender it was. An audit session confirmed each one loaded exactly its own skills: none for plain Opus 5.5, showtime:showtime plus its MCP server, brag-slim, 19 hyperframes skills, video-use.
  3. A frozen showtime snapshot of the files git ships, minus benchmarks/, so the showtime agent could not read the task rubrics.
  4. The same everything else. Model claude-opus-5-5 at high effort, Claude Code 2.1.283, the same ffmpeg, node, python and uv on the PATH, a 25-minute cap, a $15 budget cap, the same denied commands, and subscription authentication (a token from claude setup-token).
  5. One standard reply to questions. If an agent stopped to ask something before delivering, it got the same answer: I'm not available to answer questions right now. Use your best judgment, state your assumptions, and finish the video. The question still counts against it.
  6. Headless and timed. claude -p with streamed output; the harness logged every message, polled the workspace for the first output file, and killed leftovers at the end. No run hit the cap. A check after each run confirmed the operator's settings were untouched.

How the outputs were scored

LayerWhat it doesBlind?
Automatic metricsDelivered or not; format spec (length, aspect, audio, voice, captions); showtime qa on a copy of the file (loudness, true peak, clipping, silences, black and frozen stretches, frame 0, caption timing); local speech recognition for voice and footage tasks; filler, word and pause counts for the filler cut; a headless-Chrome probe with the network blocked for the HTML report.not needed
Fact-checkA judge lists every factual claim (spoken, from the transcript; on screen, from frames) and marks it supported, unsupported or contradicted by the task's own files. True facts from outside those files count as unsupported.yes
Ranking judgeOne Opus judgment per task. It sees every output as video-1..3 in shuffled order (frames, transcripts, measurements), scores seven criteria and ranks them.yes, position shuffled
Blind human reviewOne reviewer, who is showtime's author. One board per task: the three videos re-encoded with identical settings under shuffled letters, "which would you ship?" for each pair, 0-5 ratings, notes, and a final pick. HTML reports were screen-recorded by one script. The key mapping letters to tools stayed off the board until the vote was in.yes in round 1; partly in the rematch

A harness bug found and fixed during round 1

hyperframes' HTML report run ended its turn on a question (Do you already have an idea of how this should look or what it should say?). The harness's question detector did not recognise that phrasing, so the run never got the standard reply that every other contender would have received. The detector was widened, that one cell was rerun from scratch, and the whole round was rescored. The result shown for that cell is the rerun.

Results by task

Six tasks, one run each

First file: seconds until the first file of the deliverable's type appeared in the workspace. It is often a draft or an intermediate render, so it is not "time to a watchable video". Wall: whole run, including any reply to a question. Cost: the API-equivalent USD that Claude Code reports; the runs used a subscription and were not billed per token. qa: showtime qa, same thresholds for everyone. LUFS: integrated loudness (−14 is the usual target for social video). Flagged claims: claims the fact-check judge marked unsupported or contradicted.

t1 · launch video from a repo · rerun in the rematch

Launch video for quillsort 2.0

Make a 30-second launch video for quillsort 2.0 from this repo that I can post on X.

Reviewer's pickbrag-slim(both rounds)
Judgebrag-slim 1st · showtime 2nd · plain 3rd (both rounds)
plain Opus 5.5 launch video: 'New in quillsort 2.0' feature list
plain Opus 5.5
frame at 24 s
showtime round 1 launch video: 'Sorted.' next to a terminal listing
showtime, round 1
frame at 0.6 s
showtime rematch launch video: 'Sorted. Technically.' next to a terminal
showtime, rematch
frame at 0.6 s
brag-slim launch video: natural sort demo, 'file2 before file10'
brag-slim
frame at 12 s
plain Opus 5.5showtime r1showtime rematchbrag-slim
Reviewer★★★★★round 1: 2nd★★★★★round 1: 3rdnot pickedrematch, pick only★★★★★pickpicked in both rounds
Judge rank3rdrematch: 3rd2nd2nd1strematch: 1st
First file242 s420 s444 s577 s
Wall time265 s451 s474 s598 s
Cost$0.88$2.11$1.89$2.09
Turns · questions19 · 038 · 040 · 036 · 0
qaFAILsilent audio track; five still holds of 2.5-4.2 sPASSPASSWARNcolour tags missing
Loudnesssilent−14.0 LUFS−14.0 LUFS−14.1 LUFS
Flagged claims043 are true of the repo's code (the .bak file name, "62 lines", the source shown) but not in the listed sources; 1 frame shows unsorted output01demo output shown unsorted; a re-check during the rematch found 2

Review notes

  • The reviewer rated brag-slim and showtime equally (4 of 5) but preferred brag-slim in the head-to-head, and preferred plain Opus 5.5 over showtime.
  • plain Opus 5.5's video was marked down for being silent.
  • On-screen elements in brag-slim's video looked slightly unsteady to the reviewer.

Judge's reasoning, in short

brag-slim had the clearest hook (the "file10 before file2?" problem), large readable terminal text and no defects. showtime was close, with a witty opening and a good mix, but its "One Python file" code screen was too small to read on a phone. plain Opus 5.5 was silent and held still for long stretches.

In the rematch the judge said much the same: showtime covered the most content but its changelog and terminal text were small at phone size.

Rematch

showtime's rerun used the new fact-discipline and music defaults. It went from 4 flagged claims to 0 and still passed qa, but it did not win the reviewer over or move up with the judge. brag-slim won this task twice.

Caveat. The new fact-discipline rule in showtime's story guide originally used a quillsort example (backup name: notes.txt.bak (ran quillsort -i …)), taken from this very task, and the rematch agent read that guide. The drop to 0 flagged claims may partly reflect that, not only the general rule. The example has since been replaced with a generic one in the repository.

t2 · data story from a CSV

45-second temperature data story

Turn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube.

Reviewer's pickshowtime
Judgeshowtime 1st · hyperframes 2nd · plain 3rd
plain Opus 5.5 data story: bar chart of yearly departures, '2024 was the hottest year'
plain Opus 5.5
frame at 36 s
showtime data story: '+1.25 °C, 2024 the warmest year since 1850' over warming stripes
showtime
frame at 0.9 s
hyperframes data story: orange ranking of the ten warmest years
hyperframes
frame at 36 s
plain Opus 5.5showtimehyperframes
Reviewer★★★★★2nd★★★★★pick★★★★★3rd
Judge rank3rd1st2nd
First file238 s470 s1,180 s
Wall time511 s708 s1,208 s
Cost$1.33$1.97$11.72
Turns · questions30 · 010 · 021 · 17 sub-agents
qaWARNno audio track; seven still holds up to 5.3 sPASSFAIL4.8 LU too quiet; two black flashes; flat first frame
Loudnessno audio−14.0 LUFS−18.8 LUFS
Flagged claims01a running line-end label lagged the plotted value mid-draw1credits NOAA, which the CSV itself does not name

Review notes

  • showtime's charts were preferred, but the reviewer found the number labels crowded in places.
  • plain Opus 5.5's temperature chart was well liked; the video was marked down as mostly silent (its file has no audio track).
  • hyperframes' storytelling and graphics were praised, but the finish felt cheap to the reviewer.

Judge's reasoning, in short

showtime was the technically cleanest: on-target loudness, no gaps, no qa findings. hyperframes was rougher (too quiet, black flashes, holds). plain Opus 5.5 had no audio and long still stretches.

Caveat. On this task the judge reported that it could not open the frame images. It ranked from the measurements and transcripts only, at low confidence (2 of 5).

t3 · vertical short with voice-over and captions

30-second Reels short for quillsort 2.0

Make a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions.

Reviewer's pickshowtime
Judgeplain 1st · showtime 2nd · brag-slim 3rd
plain Opus 5.5 vertical short: --natural card on cream
plain Opus 5.5
frame at 6 s
showtime vertical short: 'Edit in place' card on dark teal with captions
showtime
frame at 12 s
brag-slim vertical short: --strip-blank terminal demo
brag-slim
frame at 18 s
plain Opus 5.5showtimebrag-slim
Reviewer★★★★★3rd★★★★★pick★★★★★2nd
Judge rank1st2nd3rd
First file646 s384 s889 s
Wall time661 s438 s1,002 s
Cost$1.81$1.49$2.23
Turns · questions46 · 031 · 052 · 0
qaWARNa 37-character caption line; flat first frame; last 2 s silentPASSWARNone-frame flash at the start; colour tags missing
Loudness−14.0 LUFS−14.0 LUFS−14.6 LUFS
Flagged claims6 flagged, 0 realall six (Python 3.8+, pip install, MIT, no dependencies, tagline) are in the README, a listed source; the judge missed them1a made-up quillsort list command1demo output shown unsorted

Review notes

  • showtime's reel was judged solid overall, but its subtitles were the weakest part.
  • brag-slim's reel was also rated highly; its voice sounded somewhat robotic.
  • plain Opus 5.5's voice sounded strongly robotic and its visuals unremarkable, which put it last.

Judge's reasoning, in short

plain Opus 5.5 covered all five changes with a clear card per feature and a strong hierarchy. showtime also covered all five with polished word-by-word captions and clean audio, but left the bottom half of each frame empty, used small card text and chopped the captions into short chunks. brag-slim had the strongest hook but its voice-over skipped the Python 3.7 drop.

The judge sees frames and a transcript, not the voice itself. The human reviewer ranked plain Opus 5.5 last, mainly for its voice.

t4 · footage edit · rerun in the rematch

Cut the fillers and long pauses from an interview

Cut the filler words and long pauses out of interview.mp4 and give me the tightened video.

Round 1reviewer's pick plain Opus 5.5; judge: showtime 1st · video-use 2nd · plain 3rd
Rematchreviewer's pick showtime (partly blind); judge: video-use 1st · showtime 2nd · plain 3rd
plain Opus 5.5 cut: the interviewee on camera
plain Opus 5.5
frame at 11.3 s
showtime round 1 cut: the interviewee on camera
showtime, round 1
frame at 10.2 s
showtime rematch cut: the interviewee on camera
showtime, rematch
frame at 10.6 s
video-use cut: the interviewee on camera
video-use
frame at 9.5 s
plain Opus 5.5showtime r1showtime rematchvideo-use
Reviewer★★★★★pick r1rematch: close second★★★★★round 1: 3rdpickrematch, partly blind, pick only★★★★★round 1: 2nd
Judge rank3rdrematch: 3rd1st2nd2ndrematch: 1st
First file224 s280 s222 s604 s
Wall time433 s489 s335 s843 s
Cost$1.06$1.60$1.26$1.94
Turns · questions23 · 035 · 041 · 035 · 1asked for an ElevenLabs key, then transcribed locally with Whisper
qaFAILkept a 5.2 s silence; 9 s still picture; 22 kHz audioWARN1.2 s gap; still holdsFAIL9.7 s still picture (the unedited source fails the same check)FAIL8.6 s still picture
Loudness−16.0 LUFS−14.0 LUFS−14.0 LUFS−15.9 LUFS
Length56.6 s51.2 s53.0 s47.6 s
Fillers removed3 of 52 of 52 of 54 of 5
Words kept in order98.8%100%100%99.4%
Longest pause left5.2 s1.2 s1.4 s2.5 s

The reference cut, which removes exactly the five known "uh"s and the two silences, scores 5 of 5 fillers, 99.4% of words, 51.5 s. A filler counts as removed when the words around it got at least 60% of its length shorter. Both showtime runs kept every word, and the rematch run reported that its two speech models heard almost no fillers, which is where it fell short on this measure. The 5-second gaps are a silent mission-patch card in the source; every contender kept it in some form.

Review notes

  • Round 1: the reviewer found all three cuts acceptable, but flagged a zoom in and back out at the cuts in showtime's version and a choppy feel in video-use's.
  • Rematch: the reviewer found all three good and showtime's and plain Opus 5.5's cuts nearly identical, and picked showtime.

Judge's reasoning, in short

Round 1: showtime's transcript was clean, its loudness on target and its longest inner silence 1.2 s. video-use was the tightest but left a 2.2 s gap and was quieter. plain Opus 5.5 kept the 5.2 s silence and a stutter. Rematch: video-use first for the tightest cut, showtime second for trimming less.

Caveat. In both rounds the judge reported that it could not find the frame images, so it ranked from measurements and transcripts only. It never saw the punch-in zooms the reviewer flagged. Round 1's showtime run added a 1.12× zoom on every other section to hide jump cuts; the rematch run had none.

Rematch

With punch-ins off by default, showtime's rerun went from the reviewer's 3rd (2 of 5) to the reviewer's pick. The vote was only partly blind: the other two videos were the same files the reviewer had already rated in round 1, so the new one was the only unfamiliar video on the board. The judge moved showtime from 1st to 2nd while ranking the same two competitor files again, which shows how noisy a single judgment is.

t6 · single-file HTML video report · includes the rerun cell

Shareable HTML report on the Mauna Loa CO2 record

Make a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team.

Reviewer's pickshowtime
Judgeshowtime 1st · plain 2nd · hyperframes 3rd
plain Opus 5.5 report: CO2 line chart with a play button
plain Opus 5.5
page on load
showtime report: player poster, 'Mauna Loa CO2, 1959-2025' over a chart
showtime
page on load
hyperframes report: embedded video player above headline numbers and key findings
hyperframes (rerun)
page at 3 s
plain Opus 5.5showtimehyperframes
Reviewer★★★★★2nd★★★★★pick★★★★★3rd
Judge rank2nd1st3rd
First file214 sa template page22 sa project draft50 sa composition page
Wall time389 s349 s651 s
Cost$1.40$1.05$9.60
Turns · questions29 · 028 · 039 · 16 sub-agents
Offline probePASSsingle file, 0 requests, plays, fits a phonePASSsamePASSsame
What playsa 37 s canvas animation, silenta 32 s player with 4 chapters and musican embedded 48 s MP4, silent
File size33 KB1.2 MB3.7 MB
Flagged claims54 true background facts not in the CSV (the observatory, NOAA, "longest record", the download feature); 1 wrong value: 1970 shown as 325.72 ppm, the CSV says 325.6801"rising faster every decade", but the 1990s rose slower than the 1980s

Review notes

  • showtime's report was the clear favourite and the only one with sound, but the reviewer found its music cheap-sounding and some chapter transitions abrupt.
  • The other two were seen as short videos with little story. The reviewer's written notes preferred hyperframes' clearer story and theme, but the head-to-head answer went to plain Opus 5.5.

Judge's reasoning, in short

showtime: a self-contained player with chapters, a correct headline (315.98 → 427.35 ppm), a chart that draws itself and a phone-sized layout; one flaw, axis labels behind the title on the poster. plain Opus 5.5: clean and accurate but silent and generic. hyperframes: a good written summary, but the headline contradicts its own notes.

Caveat. Every frame the judge saw of hyperframes' page showed only the video's poster, so its video content may not have reached the judge. The human board used a separate screen recording of each page.

t8 · math explainer

Why the first n odd numbers add up to n²

Make a 45-second video that shows visually why the sum of the first n odd numbers is n squared.

Reviewer's pickshowtime
Judgeshowtime 1st · video-use 2nd · plain 3rd
plain Opus 5.5 math video: 7 by 7 grid of L-shaped layers
plain Opus 5.5
frame at 36 s
showtime math video: coloured tile layers next to 1+3+5+7+9 = 5 squared
showtime
frame at 27 s
video-use arm math video: coloured square grid with running sums
video-use*
frame at 27 s
plain Opus 5.5showtimevideo-use*
Reviewer★★★★★3rd★★★★★pick★★★★★2nd
Judge rank3rd1st2nd
First file196 s132 sa draft scene render232 s
Wall time267 s478 s259 s
Cost$0.88$1.72$0.91
Turns · questions18 · 043 · 017 · 0
qaFAILblack first frame; 3.6 s black gap; no audioWARN0.47 s black blip; still holdsFAIL2.3 s black gap; no audio
Loudnessno audio−14.0 LUFSno audio
Math errors000

* This run had video-use installed but never invoked it, so it is effectively a second plain Opus 5.5 run. It says nothing about video-use itself.

Caveat. showtime ships a Manim example on this exact proof ("Odd numbers build squares"), and its Manim reference, which this run read, uses "Why n squared?" as a sample. The run wrote its own scenes and did not open the example files, but this task is not independent of showtime's own materials.

Review notes

  • The reviewer found all three acceptable. showtime's felt the most polished and best explained, but its music choice sounded cheap.
  • The video-use arm's version was well explained, but its animations looked childish to the reviewer; plain Opus 5.5's was less polished and less clear.
  • Only showtime's video had sound.

Judge's reasoning, in short

All three used the same correct proof (nested L-shaped layers). showtime had the tightest explanation, colour-matched equations and labelled braces, and was the only one with sound. The other two were silent, had black gaps, and let the grid run ahead of the equations.

Round 3 · showtime 0.2.0

The fixes that were never retested, and a judge that proves it looked

Rounds 1 and 2 changed showtime after each review, but only the launch video and the filler cut were rerun. Round 3 runs showtime 0.2.0 once on all six original requests, with the unchanged prompts, inputs, model, limits and reply-to-questions rule, and ranks the new video next to the round-1 or round-2 showtime video and the other contenders' existing outputs (they were not rerun). Every ranking counts only if the judge proves it opened the frames.

Blind boards for a human vote were built (one per task, letters shuffled, key kept apart) and have been voted on; the result is in the Round 3 vote section below, after Round 3b. The tables in this section are the judge's. New runs: showtime snapshot of git commit f9c3de9 for t2, t3, t4, t6 and t8, and git commit c8a5502 for t1, which was run after the launch-film fixes were merged. 0.2.0 changed the default transcription to a model that hears fillers, added a catalog of produced music tracks, and reworked the launch-video defaults, alongside the round-1 fixes that had not been retested (music taste, chart labels, captions, HTML transitions).

How the judge now proves it looked at the pictures

Earlier rounds disclosed three rankings made without the images. Now every judgment (ranking, pairwise and fact-check judges alike) has to pass two checks or it is discarded and asked again in a fresh session with a fresh packet: (1) the session's own tool-call log must show that every frame image of its packet was opened with the Read tool; (2) every image carries a random 5-digit number in a black strip at the bottom, stamped on every image of every contender alike and known only to the scorer, and the judge must report the number of each image it opened (at most 10 percent wrong). A number can only be known from the pixels. A unit test suite covers the rule, including a stand-in judge that opens nothing.

Result: 18 ranking judgments, 528 of 528 images opened, 526 of 528 numbers read correctly, 0 redone. The fact-check judges made 21 attempts and one failed the check (4 of 5 numbers) and was redone. A live test with the real judge model on synthetic images read 10 of 10 numbers. This proves the images were opened and read; it does not prove they were judged well.

TaskRanking judgmentsImages opened (tool-call log)Reading-check numbers rightRedone for failed proof
t1384 of 8484 of 840
t2384 of 8484 of 840
t3384 of 8484 of 840
t4384 of 8484 of 840
t63108 of 108106 of 1080
t8384 of 8484 of 840
all rankings18528 of 528526 of 5280

Round 3 at a glance

Each row: three judgments of one packet of four videos (the new showtime video, the round-1 or 2 showtime video, plain Opus 5.5, and the outside tool paired with the task), positions rotated. qa is showtime qa version 0.2.0 run on every file again, so the old files were re-checked with the new checker.

Taskshowtime 0.1.0 (judge)showtime 0.2.0 (judge)0.2.0 vs 0.1.0best other contenderqa 0.1.0 → 0.2.0costwall
t1 showtime 0.1.0 3rdshowtime 0.2.0 mean 1.67BETTER than 0.1.0brag-slim 1stPASS → PASS$1.89 → $3.86474 → 1158 s
t2 showtime 0.1.0 1stshowtime 0.2.0 2ndWORSE than 0.1.0plain Opus 5.5 3rdPASS → PASS$1.97 → $3.79708 → 1240 s
t3 showtime 0.1.0 1stshowtime 0.2.0 3rdWORSE than 0.1.0plain Opus 5.5 2ndPASS → PASS$1.49 → $2.51438 → 710 s
t4 showtime 0.1.0 3rdshowtime 0.2.0 1stBETTER than 0.1.0video-use 2ndWARN → WARN$1.26 → $1.72335 → 375 s
t6 showtime 0.1.0 1stshowtime 0.2.0 2ndWORSE than 0.1.0plain Opus 5.5 3rdHTML probe PASS → HTML probe PASS$1.05 → $0.98349 → 278 s
t8 showtime 0.1.0 1stshowtime 0.2.0 3rdWORSE than 0.1.0plain Opus 5.5 3rdWARN → WARN$1.72 → $2.59478 → 703 s

Reading this. "Better", "worse" and "level" compare the mean judge rank of the 0.2.0 video with the mean rank of the 0.1.0 video from the same packet. Each is one run against one run, with three judgments that are rotations of one packet, so a difference between two runs of the same tool may be run-to-run variation, not a change in the tool. Outputs of the other contenders are the ones from rounds 1 and 2; those tools were not rerun.

t1 · launch video from a repo · round 3

Launch video for quillsort 2.0

Make a 30-second launch video for quillsort 2.0 from this repo that I can post on X.

Judge (verified, 3 judgments)1st brag-slim · 2nd showtime 0.2.0 · 3rd showtime 0.1.0 · 4th plain Opus 5.5
plain Opus 5.5 output for this task
plain Opus 5.5
frame at 24 s
brag-slim output for this task
brag-slim
frame at 12 s
showtime 0.1.0 output for this task
showtime 0.1.0
frame at 0.6 s
showtime 0.2.0 output for this task
showtime 0.2.0
frame at 12 s
plain Opus 5.5brag-slimshowtime 0.1.0round 2 outputshowtime 0.2.0
Judge rank (of 4)4th3 judgments: 4th, 4th, 4thmean 1.333 judgments: 2nd, 1st, 1st3rd3 judgments: 3rd, 3rd, 3rdmean 1.673 judgments: 1st, 2nd, 2nd
Wall time265 s598 s474 s1,158 s
Cost$0.88$2.09$1.89$3.86
Turns · questions19 · 036 · 040 · 023 · 0
qa (0.2.0 checker)FAILthe audio track is silent; colour tags are missing ; untagged or non-BT.709 video shifts colour…; +5 moreWARNcolour tags are missing ; untagged or non-BT.709 video shifts colour…PASSPASS
Loudnessno audio-14.1 LUFS-14.0 LUFS-14.0 LUFS
Length30.0 s30.0 s30.0 s30.0 s
Flagged claims01demo output shown unsorted: "Grace" before "ada" (as in rounds 1 and 2)1how plain sort orders file names (locale-dependent; round 2 had counted 0 for this same video)1the module list a real grep printed on screen; the file's source is not among the listed sources

What is on screen

  • showtime 0.2.0: the hook "file2 comes before file10." on a cream ground, a terminal that shows plain sort getting it wrong and --natural --unique fixing it, an in-place edit that keeps a .bak, a grep that proves "no dependencies", an end card with the install command and the music credit. A terminal window carries the middle of the film. Music at −14 LUFS (a catalog track, credited on screen).
  • A real defect: at about 3.7 s the camera push over the hook crops the headline ("file2 comes befor"). Two of the three judgments name it. The launch fix that keeps text inside the frame during moves did not prevent it here.
  • showtime 0.1.0 (round 2): more variety (before/after file cards, a changelog checklist, a one-file graphic) but the judges caught mid-animation gaps in terminal output, misaligned changelog columns and tiny text. brag-slim: the strongest hook for two judgments, but its dedupe demo shows an unsorted result. plain Opus 5.5: silent (fails a post on X) with 3-4 s frozen holds.

Judges' reasoning, in short

Two judgments put brag-slim first and showtime 0.2.0 second; one put 0.2.0 first. All three put showtime 0.1.0 third and plain Opus 5.5 last. Mean rank: brag-slim 1.33, showtime 0.2.0 1.67, showtime 0.1.0 3.0, plain Opus 5.5 4.0. Rounds 1 and 2 also had brag-slim first (the reviewer picked it both times) and showtime second with a single judgment.

The agent ran the launch workflow's independent review round (one sub-agent), which took 1,158 s and $3.86, more than twice the earlier showtime run (474 s, $1.89). A second run of the same build, made while preparing this round, ran into the 25-minute cap for the same reason and is not used.

Fixes retested here: the new launch-video defaults

After the round-2 loss, 0.2.0 reworked how launch films are built: one story in five scenes, hand-offs by camera moves rather than 0.4 s pushes, a real produced music track, terminals that fit their window, text kept inside the frame. Result: the judges rank this video above the round-2 one in all three judgments (mean 1.67 against 3.0) and below brag-slim by a narrow margin, with a cropped headline as the visible flaw. In the round-3 human vote (one voter, partly blind) the reviewer picked brag-slim for a third time, and rated this video 3 of 5 (4 for the round-2 video, 5 for brag-slim), with the comment that the zooms and camera movements did not make much sense; see Round 3 vote. JUDGES: IMPROVED · VOTE: LOST AGAIN

t2 · data story from a CSV · round 3

45-second temperature data story

Turn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube.

Judge (verified, 3 judgments)1st showtime 0.1.0 · 2nd showtime 0.2.0 · 3rd plain Opus 5.5 · 4th hyperframes
plain Opus 5.5 output for this task
plain Opus 5.5
frame at 36 s
hyperframes output for this task
hyperframes
frame at 36 s
showtime 0.1.0 output for this task
showtime 0.1.0
frame at 0.9 s
showtime 0.2.0 output for this task
showtime 0.2.0
frame at 27 s
plain Opus 5.5hyperframesshowtime 0.1.0round 1 outputshowtime 0.2.0
Judge rank (of 4)mean 3.333 judgments: 4th, 3rd, 3rdmean 3.673 judgments: 3rd, 4th, 4th1st3 judgments: 1st, 1st, 1st2nd3 judgments: 2nd, 2nd, 2nd
Wall time511 s1,208 s708 s1,240 s
Cost$1.33$11.72$1.97$3.79
Turns · questions30 · 021 · 110 · 013 · 0
qa (0.2.0 checker)WARNcolour tags are None/None/bt709 ; untagged or non-BT.709 video shifts…; the video has no audio track; +8 moreFAILintegrated loudness -18.8 LUFS is -4.8 LU off the -14 LUFS target; frame 0 is a flat colour (mean luma 17); +4 morePASSPASS
Loudnessno audio-18.8 LUFS-14.0 LUFS-14.0 LUFS
Length45.0 s45.0 s45.0 s45.0 s
Flagged claims0100

What is on screen

  • showtime 0.1.0 (round 1): warming stripes, an animated line chart with 1976 marked, decade bars, a ranked chart of the eleven warmest years and a closing "49 years in a row" card.
  • showtime 0.2.0: the same editorial look in fewer scenes: a title card, the line chart, the ranked chart with clean end labels, and a closing "+1.01 °C" card. No stripes and no decade chart, so the run has less variety; the ranked chart and the closing card hold on screen for a long time.
  • plain Opus 5.5 has no audio track and ends on black; hyperframes shows "••••" where the ranked chart's year labels should be, black gaps and audio 4.8 LU too quiet (unchanged from round 1).

Judges' reasoning, in short

All three judgments put showtime 0.1.0 first and 0.2.0 second, for the reason they all gave: 0.2.0 has "less variety" and long holds; its style, numbers (correct) and audio (−14 LUFS, a captions file) were praised. All three judgments noted that the round-1 video's decade-bar labels crowd or overlap slightly; none noted crowding in the 0.2.0 video.

Cost and time went up: 1,240 s and $3.79 against 708 s and $1.97, with only 13 turns (the machine also carried two other runs and other jobs at the time).

Fix retested here: crowded chart labels

Round 1's reviewer found the number labels crowded. The 0.2.0 video has no crowded labels, but it also has no decade chart, the chart that had the problem, so this run does not show the collision fix working. In the round-3 vote the reviewer picked the 0.1.0 video and noted that its numbers are crowded in the bar charts (the same complaint as in round 1); the 0.2.0 video got no such note and was rated 4 of 5, the same as the 0.1.0 video. NOT CONCLUSIVE

t3 · vertical short with voice-over and captions · round 3

30-second Reels short for quillsort 2.0

Make a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions.

Judge (verified, 3 judgments)1st showtime 0.1.0 · 2nd plain Opus 5.5 · 3rd showtime 0.2.0 · 4th brag-slim
plain Opus 5.5 output for this task
plain Opus 5.5
frame at 6 s
brag-slim output for this task
brag-slim
frame at 18 s
showtime 0.1.0 output for this task
showtime 0.1.0
frame at 12 s
showtime 0.2.0 output for this task
showtime 0.2.0
frame at 6 s
plain Opus 5.5brag-slimshowtime 0.1.0round 1 outputshowtime 0.2.0
Judge rank (of 4)2nd3 judgments: 2nd, 2nd, 2nd4th3 judgments: 4th, 4th, 4th1st3 judgments: 1st, 1st, 1st3rd3 judgments: 3rd, 3rd, 3rd
Wall time661 s1,002 s438 s710 s
Cost$1.81$2.23$1.49$2.51
Turns · questions46 · 052 · 031 · 044 · 0
qa (0.2.0 checker)WARNcolour tags are missing ; untagged or non-BT.709 video shifts colour…; frame 0 is a flat colour (mean luma 248); +4 moreWARNcolour tags are missing ; untagged or non-BT.709 video shifts colour…; frame 0 differs sharply from frame 1 (mean diff 61/255): a one-frame…PASSPASS
Loudness-14.0 LUFS-14.6 LUFS-14.0 LUFS-14.0 LUFS
Length30.0 s30.0 s30.0 s30.0 s
Flagged claims01demo output shown unsorted (as in round 1)2"3.8 or newer" and "pip install quillsort", both in the README, a listed source: the fact-check judge was wrong0

What is on screen

  • showtime 0.1.0 (round 1): a dark theme, large high-contrast captions, a "Heads up: 2 changes" card for the breaking changes, all five changelog items, a pip-install end card.
  • showtime 0.2.0: a light cream theme with terminal cards, word-by-word captions with an outline, all five changelog items, qa PASS. The judges found the layout top-heavy with the lower 40 percent empty, some caption words dimmed to grey and the outlined captions muddy on the cream ground; one judgment also saw stretches with no caption.
  • plain Opus 5.5 was ranked second by all three judgments: clean, centred, a 1/5 to 5/5 counter, but a flat first frame, a silent last 2 s and two captions that go by fast. brag-slim was last: its voice-over skips the Python 3.7 drop.

Judges' reasoning, in short

All three judgments ranked showtime 0.1.0 first, plain Opus 5.5 second, showtime 0.2.0 third and brag-slim fourth. In round 1 the single judge had put showtime second behind plain Opus 5.5; with the frames verified it puts the round-1 video first.

Fix retested here: captions

Round 1's reviewer called the subtitles the weakest part, and the new clean-pop caption default was meant to fix that. The judges did not see an improvement: their notes on the 0.2.0 video (muddy outlines, greyed words, empty lower half) are at least as critical as their notes on the round-1 video. The two videos also differ in look for reasons beyond the caption style, since the agent chose a different theme each time. This round-3 run ran stale code; after the rerun (caption shadow fixed, Round 3b) the judges ranked it 2nd, behind the 0.1.0 video, and in the round-3 vote (one voter, partly blind) the rerun video was the first pick, 5 of 5 against 4 of 5 for the 0.1.0 video. Captions were not rated separately. MIXED: JUDGES NO IMPROVEMENT · VOTE PICKED IT

t4 · footage edit · round 3

Cut the fillers and long pauses from an interview

Cut the filler words and long pauses out of interview.mp4 and give me the tightened video.

Judge (verified, 3 judgments)1st showtime 0.2.0 · 2nd video-use · 3rd showtime 0.1.0 · 4th plain Opus 5.5
plain Opus 5.5 output for this task
plain Opus 5.5
frame at 11.3 s
video-use output for this task
video-use
frame at 9.5 s
showtime 0.1.0 output for this task
showtime 0.1.0
frame at 10.6 s
showtime 0.2.0 output for this task
showtime 0.2.0
frame at 9.9 s
plain Opus 5.5video-useshowtime 0.1.0round 2 outputshowtime 0.2.0
Judge rank (of 4)4th3 judgments: 4th, 4th, 4th2nd3 judgments: 2nd, 2nd, 2nd3rd3 judgments: 3rd, 3rd, 3rd1st3 judgments: 1st, 1st, 1st
Wall time433 s843 s335 s375 s
Cost$1.06$1.94$1.26$1.72
Turns · questions23 · 035 · 141 · 048 · 0
qa (0.2.0 checker)FAILsilence from 19.25s to 24.45s (5.2s) inside the video; integrated loudness -16.0 LUFS is -2.0 LU off the -14 LUFS target; +3 moreWARNintegrated loudness -15.9 LUFS is -1.9 LU off the -14 LUFS target; silence from 18.55s to 20.80s (2.2s) inside the videoWARNsilence from 21.75s to 23.15s (1.4s) inside the video; the last 2.5s are silent (sound ends at 50.50s)WARNintegrated loudness -16.0 LUFS is -2.0 LU off the -14 LUFS target; silence from 20.15s to 22.40s (2.2s) inside the video; +1 more
Loudness-16.0 LUFS-15.9 LUFS-14.0 LUFS-16.0 LUFS
Length56.6 s47.6 s53.0 s49.7 s
Flagged claimsnot checkedthe speaker's own wordsnot checkedthe speaker's own wordsnot checkedthe speaker's own wordsnot checkedthe speaker's own words
Fillers removed (rounds 1-2 scorer)3 of 54 of 52 of 55 of 5
Words kept in order98.8%99.4%100%100%
"uh"/"um" still audible (verbatim transcript)4290
Longest pause left5.2 s2.5 s1.4 s2.3 s

Fillers removed uses exactly the round-1 and round-2 scorer (Whisper turbo transcribes the source and each output; a filler counts as removed when the words around it shrank by at least 60% of its length), so the numbers are comparable with the earlier rounds: 0.1.0 scored 2 of 5 in both, video-use 4, plain Opus 5.5 3. The verbatim count is a different measure: showtime 0.2.0's new default speech model writes "uh" and "um" as words, and the same model was used on every file, but it is also the model showtime now ships, so treat it as the tested tool's own instrument. The reference cut (exactly the five known "uh" and the two silences) scores 5 of 5, 99.4%, 51.5 s. qa verdicts on old files changed for reasons unrelated to the videos: since round 2 the checker no longer fails a talking head that simply holds still.

What is on screen

  • showtime 0.2.0 removed every "uh" and the doubled words ("the the", "to to"): 49.7 s from 63.7 s, every content word kept, in order.
  • What is left: the silent mission-patch card in the middle (2.2 s) and a silent 2.3 s tail, and the loudness is −16.0 LUFS, the source's level (the 0.1.0 cut was lifted to −14.0). qa WARN.
  • showtime 0.1.0 (round 2) kept about nine "uh"s and several doubled words but was normalised to −14.0 LUFS; plain Opus 5.5 kept a 5.2 s silent still card; video-use is the tightest (47.6 s) with two "uh" left.

Judges' reasoning, in short

All three judgments ranked showtime 0.2.0 first: "the only version that removes all the filler", with the sentences intact. video-use second, showtime 0.1.0 third (barely changed the speech), plain Opus 5.5 last (a 5.2 s frozen silent card, a clipped word). Visually all four are the same letterboxed webcam picture, so the ranking rests on the transcript and the measurements.

In rounds 1 and 2 the same kind of ranking put showtime 1st, then 2nd, without seeing the frames; this one was verified.

Fix retested here: filler detection

0.2.0 replaced the default speech model and added a filler scan after it. This is the retest: 5 of 5 known fillers removed (was 2 of 5) with all words kept, and first place with every judgment. The reviewer's round-2 vote (partly blind) had picked the 0.1.0 video, whose fillers were mostly still there, and the round-3 vote did so again: the 0.1.0 video was the first pick (5 of 5) and this cut was rated 4 of 5, with no reason recorded. SCORER AND JUDGES: IMPROVED · VOTE: PREFERRED 0.1.0

t6 · single-file HTML video report · round 3

Shareable HTML report on the Mauna Loa CO2 record

Make a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team.

Judge (verified, 3 judgments)1st showtime 0.1.0 · 2nd showtime 0.2.0 · 3rd plain Opus 5.5 · 4th hyperframes
plain Opus 5.5 output for this task
plain Opus 5.5
page on load
hyperframes output for this task
hyperframes
page at 3 s
showtime 0.1.0 output for this task
showtime 0.1.0
page on load
showtime 0.2.0 output for this task
showtime 0.2.0
page on load
plain Opus 5.5hyperframesshowtime 0.1.0round 1 outputshowtime 0.2.0
Judge rank (of 4)3rd3 judgments: 3rd, 3rd, 3rd4th3 judgments: 4th, 4th, 4th1st3 judgments: 1st, 1st, 1st2nd3 judgments: 2nd, 2nd, 2nd
Wall time389 s651 s349 s278 s
Cost$1.40$9.60$1.05$0.98
Turns · questions29 · 039 · 128 · 023 · 0
Offline probePASSsingle file, 0 requests, plays, fits a phonePASSsingle file, 0 requests, plays, fits a phonePASSsingle file, 0 requests, plays, fits a phonePASSsingle file, 0 requests, plays, fits a phone
Flagged claims32 true background facts not in the CSV; 1 value wrong (1970: 325.72 ppm, CSV 325.68)1"rising faster every decade": the 1990s rose slower than the 1980s00

What is on screen

  • showtime 0.1.0 (round 1): a single-file player with four chapters (headline, the record, the rate, where we are), music, large editorial type, a chart that draws itself, decade-rate bars, a 32 s runtime.
  • showtime 0.2.0: the same player and look, four chapters, 25 s. The judges call it "tighter" and carrying "less context"; one judgment saw the closing count-up show a wrong number for a moment ("411.98 ppm") mid-animation. Fact-check: 0 flagged.
  • plain Opus 5.5 is a thorough silent dashboard; hyperframes is a silent embedded video whose headline the chart contradicts.

Judges' reasoning, in short

All three judgments ranked showtime 0.1.0 first and 0.2.0 second, then plain Opus 5.5 and hyperframes. All three said the two showtime pages look almost alike, and that 0.1.0 tells the fuller story.

The judge sees frames from a screen recording of each page playing, not the sound or the smoothness of a transition. It reports "an audio element but no speech", and cannot say whether the music or the transitions got better.

Fixes retested here: music taste and chapter transitions

Both are about how the page feels in motion and sound (round 1's reviewer found the music cheap-sounding and some transitions abrupt). The judges cannot hear or feel either, so this round says nothing about them. The recorded 0.2.0 page has an audio track (captured, peak −2.7 dB) and the human board played it. In the round-3 vote (one voter, partly blind) the 0.2.0 page was the first pick, 5 of 5 and marked as liked, over the 0.1.0 page at 4 of 5, while the judges ranked it second; music and transitions were not rated separately, so the vote does not isolate them. JUDGES: NOT JUDGEABLE · VOTE: PICKED, NOT ISOLATED

t8 · math explainer · round 3

Why the first n odd numbers add up to n²

Make a 45-second video that shows visually why the sum of the first n odd numbers is n squared.

Judge (verified, 3 judgments)1st showtime 0.1.0 · 2nd plain Opus 5.5 = showtime 0.2.0 = video-use (tied on mean rank)
plain Opus 5.5 output for this task
plain Opus 5.5
frame at 36 s
video-use output for this task
video-use*
frame at 27 s
showtime 0.1.0 output for this task
showtime 0.1.0
frame at 27 s
showtime 0.2.0 output for this task
showtime 0.2.0
frame at 27 s
plain Opus 5.5video-use*showtime 0.1.0round 1 outputshowtime 0.2.0
Judge rank (of 4)3rd3 judgments: 4th, 2nd, 3rd3rd3 judgments: 3rd, 4th, 2nd1st3 judgments: 1st, 1st, 1st3rd3 judgments: 2nd, 3rd, 4th
Wall time267 s259 s478 s703 s
Cost$0.88$0.91$1.72$2.59
Turns · questions18 · 017 · 043 · 047 · 0
qa (0.2.0 checker)FAILframe 0 is black (black for the first 0.30s): feeds, chat apps and…; black from 3.47s to 7.10s (3.63s); +6 moreFAILblack from 3.80s to 6.13s (2.33s); colour tags are missing ; untagged or non-BT.709 video shifts colour…; +6 moreWARNthe picture does not change from 0.00s to 3.13s (3.1s; small changes…; the picture does not change from 3.13s to 6.53s (3.4s; small changes…; +3 moreWARNblack from 9.03s to 9.47s (0.43s); the picture does not change from 10.40s to 14.23s (3.8s; small…
Loudnessno audiono audio-14.0 LUFS-14.0 LUFS
Length45.0 s45.0 s45.0 s45.0 s
Flagged claims0002two equations shown next to the wrong-sized grid mid-transition: "1+3+5 = 2²" (2×2 grid) and "1+3+5+7+9 = 4²" at about 24 s

What is on screen

  • showtime 0.1.0 (round 1): "Odd numbers as tiles", each layer wrapped around a growing square, the n + (n − 1) = 2n − 1 step, a closing "n layers make an n × n square"; music bed.
  • showtime 0.2.0: coloured L-shaped layers with colour-matched equations, the general (2n − 1) case with braces, music bed. But the text changes before the tiles: at 0:24 the contact sheet shows 1+3+5+7+9 = 4² beside a 4×4 grid, and both the judges and the fact-check catch a wrong equation on screen for a moment. Mean rank 3.0 (2nd, 3rd, 4th), level with plain Opus 5.5 and the video-use arm.
  • plain Opus 5.5 and the video-use arm (which never used its skill here) are silent, open on or hold black, and were each ranked between 2nd and 4th.

Judges' reasoning, in short

showtime 0.1.0 came first in all three judgments: "the equations always match the tiles". 0.2.0, plain Opus 5.5 and the video-use arm then traded places (a mean rank of 3.0 each), so the judges could not separate them. All three judgments call out the mismatched equations in the 0.2.0 video as an accuracy failure in a proof.

Caveat. showtime ships a Manim example on this same proof, written before the task existed, and its reference uses "Why n squared?" as a sample; round 1's run read that reference. That still holds in round 3 and favours showtime on this task.

Fix retested here: music

0.2.0 picks music from a catalog of produced tracks and prefers subtler or none for explainers. The judge cannot hear it, so its taste is untested here; the video keeps a music bed at −14.0 LUFS. A new defect appeared instead: equations that lead their tiles, flagged by the judges and the fact-check. NEW DEFECT

Round 3b · regression rerun of showtime 0.2.0

Two round-3 videos ran the wrong code; four tasks rerun after the fixes

Round 3's t2 (data story) and t3 (vertical short) did not measure showtime 0.2.0. The showtime arm's runtime folder was a link to a runtime folder shared with other work on the same machine. Its stable showtime command runs whichever skill folder was recorded there last, and that was another, unfinished working copy of showtime (24 commits older than the benchmarked commit, with uncommitted edits). showtime doctor tells agents that this command "always works", so the t2 agent sent all 18 of its showtime commands through it and the t3 agent all 12; the t1 agent noticed ("points at a different (older) install") and switched back. t4, t6 and t8 used the skill's own command and were not affected.

Causes found by comparing the 0.1.0 and 0.2.0 transcripts with the judges' reasons:

  • Stale code (t2, t3): with two skill folders of the same version, the command kept the one recorded first. Fixed: a version tie now goes to the skill that is running, and every benchmark arm gets its own command and record.
  • t1 headline cropped at ~3.7 s ("file2 comes befor"): the crop came from the fly-through scene transition, which the earlier keep-text-in-frame fix (camera moves inside a scene) did not cover, and showtime check never sampled inside that transition. Fixed: the fly-through fades every word outside the portal's own word before the frame edge reaches it; check now samples inside it and reports cut text as an error.
  • t3 "muddy, low-contrast" captions: every caption style drew a black shadow even with dark words on a light ground (a defect since 0.1.0, also in the shipped paper theme). Fixed: shadows use the outline colour.
  • t8 equations ahead of their tiles ("1+3+5 = 2²" beside a 2×2 grid): the agent added the next term before the tiles and the result moved; nothing in 0.2.0 caused it, but no rule forbade it. Added: every frame of a maths explainer must be a true statement that matches the picture.
  • t1 cost and time: 0.2.0's launch workflow made a vertical cut the default for X and LinkedIn; now only for vertical-first platforms. In the first rerun a new problem appeared (the terminal history line showed only the last line of the previous command's output, which read as broken output); the template's guidance was fixed and t1 was run once more.

Rerun: the fixed showtime 0.2.0, once per task, same prompts, inputs, model, limits and harness, each arm with its own command. Ranked with the same verified judge (3 judgments, frames opened and numbers read) against the same three other candidates as in round 3. t4 and t6 were not rerun (no 0.2.0 cause found; t6 lost first place because its video was shorter and one frame caught a count-up mid-animation).

Task0.2.0 mean rank, round 30.2.0 mean rank, rerun0.1.0 in the rerun judgingbest other contender (rerun judging)costwall
t1 launch video1.672.671.67brag-slim 1.67$3.86 → $3.731158 → 864 s
t2 data story2.001.671.33plain Opus 5.5 3.33$3.79 → $5.631240 → 1278 s
t3 vertical short3.002.001.33plain Opus 5.5 2.67$2.51 → $2.41710 → 676 s
t8 math explainer3.001.002.00video-use 3.00$2.59 → $2.35703 → 857 s

t1 was run twice: the first rerun (before the history-line fix) ranked 3.0 at $1.09 and 437 s, with no cropped headline but a terminal that read as broken output; the row above is the second rerun, after that fix. Its judges found no crop and no broken output, but ranked it behind brag-slim and the 0.1.0 video because it shows three features instead of four and keeps one look throughout. t2, t3 and t8 are from the first and only rerun. One run per task: a single rerun can move a rank by one place on variance alone, and the rerun judging is a new set of three judgments of a new packet, so the 0.1.0 column can differ from round 3's. Costs include any helper agents the run started (t2's rerun started a fact-check researcher and a critic, which found a real type problem).

Round 3 · human vote

The reviewer's vote on the round-3 boards: showtime was the pick on five of six tasks, 0.2.0 on three, and the launch video lost again

The blind boards built for round 3 (one per task, four unnamed videos with shuffled letters, "which would you ship?" pairs and a 0-5 rating for every video) were voted on 29 September. The six "Copy for your agent" digests were imported with the benchmark's own parser (human_board.py, parse_digest, the same way rounds 1 and 2 were imported) and tallied through the letter-to-tool key, which was kept apart until after the vote. Every task has one first pick, five pairwise answers (of the six possible pairs) and four ratings.

Which videos were on the boards. showtime 0.2.0 is the rerun video for t1 (the second rerun), t2, t3 and t8 (Round 3b, after the fixes), and the original round-3 video for t4 and t6. The showtime 0.1.0 video is the round-1 video for t2, t3, t6 and t8 and the round-2 video for t1 and t4. The other tools' videos are the round-1 outputs, as in round 3. So on four tasks the vote is about different showtime 0.2.0 videos from the ones in the Round 3 tables above.

How much this vote can carry

It is one person, and that person is showtime's author. The reviewer voted once on each of six tasks: 30 pairwise answers and 24 ratings in all. It is the author's own vote, not a panel, and it cannot say how anyone else would vote.

It is only partly blind. The letters were shuffled and no tool was named, but the reviewer had already seen the earlier videos (the round-1 and round-2 outputs, the outside tools' among them), so an old video can be recognised on sight. The direction of that bias is not known.

Nothing on the boards records the reason for a rating, except where a note or comment was written (quoted below).

Per task: first pick and ratings

Round 3 human vote per task
TaskFirst pickshowtime 0.2.00-5 ratingshowtime 0.1.00-5 ratingplain Opus 5.50-5 ratingoutside tool0-5 rating
t1 launch video from a repobrag-slim3/54/52/55/5first pick
t2 data story from a CSVshowtime 0.1.04/54/5first pick3/54/5
t3 vertical short with voice-over and captionsshowtime 0.2.05/5first pick4/52/52/5
t4 footage editshowtime 0.1.04/55/5first pick3/53/5
t6 single-file HTML video reportshowtime 0.2.05/5first pick4/52/52/5
t8 math explainershowtime 0.2.05/5first pick4/53/53/5

Comments and notes, in the reviewer's own words: t1, “The zooms here and camera movements didn't make much sense, the content was good though” (on the showtime 0.2.0 video, rated 3 of 5). t2, “B is the best but the numbers are crowded in the bar charts” ("B" was the showtime 0.1.0 video, the first pick; the reviewer rated the 0.2.0 video 4 of 5, the same as the 0.1.0 video and as hyperframes). t4, “All were good, but I think D nailed it” ("D" was the showtime 0.1.0 video). t6: the pick, the 0.2.0 video, was marked as liked. No other comments. The outside tool is brag-slim on t1 and t3, hyperframes on t2 and t6, video-use on t4 and t8. * The video-use run on t8 never called its skill.

Pairwise answers

Round 3 pairwise answers
Taskshowtime 0.2.0 in the pairs askedOther pairs asked
t1
  • WON showtime 0.2.0 vs plain Opus 5.5
  • LOST showtime 0.2.0 vs showtime 0.1.0
  • LOST showtime 0.2.0 vs brag-slim
  • brag-slim over showtime 0.1.0
  • showtime 0.1.0 over plain Opus 5.5
t2
  • WON showtime 0.2.0 vs plain Opus 5.5
  • WON showtime 0.2.0 vs hyperframes
  • LOST showtime 0.2.0 vs showtime 0.1.0
  • hyperframes over plain Opus 5.5
  • showtime 0.1.0 over hyperframes
t3
  • WON showtime 0.2.0 vs plain Opus 5.5
  • WON showtime 0.2.0 vs showtime 0.1.0
  • WON showtime 0.2.0 vs brag-slim
  • showtime 0.1.0 over brag-slim
  • showtime 0.1.0 over plain Opus 5.5
t4
  • WON showtime 0.2.0 vs plain Opus 5.5
  • WON showtime 0.2.0 vs video-use
  • LOST showtime 0.2.0 vs showtime 0.1.0
  • video-use over plain Opus 5.5
  • showtime 0.1.0 over video-use
t6
  • WON showtime 0.2.0 vs plain Opus 5.5
  • WON showtime 0.2.0 vs showtime 0.1.0
  • WON showtime 0.2.0 vs hyperframes
  • showtime 0.1.0 over hyperframes
  • showtime 0.1.0 over plain Opus 5.5
t8
  • WON showtime 0.2.0 vs plain Opus 5.5
  • WON showtime 0.2.0 vs showtime 0.1.0
  • WON showtime 0.2.0 vs video-use
  • showtime 0.1.0 over plain Opus 5.5
  • showtime 0.1.0 over video-use

Totals

Round 3 human vote totals
ContenderTasksFirst picksPairs wonMean 0-5 ratingand each rating
showtime 0.2.0all 63 of 614 of 184.3 3, 4, 5, 4, 5, 5
showtime 0.1.0all 62 of 612 of 164.2 4, 4, 4, 5, 4, 4
plain Opus 5.5all 60 of 60 of 122.5 2, 3, 2, 3, 2, 3
brag-slim2 of 61 of 22 of 43.5 5, 2
hyperframes2 of 60 of 21 of 53.0 4, 2
video-use2 of 60 of 21 of 53.0 3, 3

Reading this. showtime, either version, was the first pick on five of six tasks: 0.2.0 on three (t3, t6, t8), the 0.1.0 video on two (t2, t4) and brag-slim on the launch video (t1). Head to head, the 0.2.0 video won three of six pairs against the 0.1.0 video and lost three: it lost t1, t2 and t4 and won t3, t6 and t8. Its ratings were higher than the 0.1.0 video's on three tasks (t3, t6, t8), the same on one (t2) and lower on two (t1, t4); mean 4.3 against 4.2, a difference of 0.2 on a 0-5 scale from one voter. The 0.2.0 video beat plain Opus 5.5 in all six pairs and beat the outside tool in five of six; the loss was the launch video to brag-slim, for the third time in three votes on that task. Each outside tool has only two tasks, so its totals are two tasks and no more.

The vote and the judge on the same videos

For comparison, the verified judge's mean ranks for the same videos the vote used: the rerun judging for t1, t2, t3 and t8, the round-3 judging for t4 and t6. The two agree on the first place in two tasks, tie in one, and differ in three.

Human vote and judge on the same videos
TaskVote: 0-5 rating (1st pick first)Judge: mean rank on the same video (lower is better)1st place
t1brag-slim 5 · showtime 0.1.0 4 · showtime 0.2.0 3 · plain Opus 5.5 2showtime 0.1.0 1.67 · brag-slim 1.67 · showtime 0.2.0 2.67 · plain Opus 5.5 4.00Round 3b, second rerunTIED 1ST
t2showtime 0.1.0 4 · showtime 0.2.0 4 · hyperframes 4 · plain Opus 5.5 3showtime 0.1.0 1.33 · showtime 0.2.0 1.67 · plain Opus 5.5 3.33 · hyperframes 3.67Round 3b rerunSAME 1ST
t3showtime 0.2.0 5 · showtime 0.1.0 4 · plain Opus 5.5 2 · brag-slim 2showtime 0.1.0 1.33 · showtime 0.2.0 2.00 · plain Opus 5.5 2.67 · brag-slim 4.00Round 3b rerunDIFFERENT 1ST
t4showtime 0.1.0 5 · showtime 0.2.0 4 · video-use 3 · plain Opus 5.5 3showtime 0.2.0 1.00 · video-use 2.00 · showtime 0.1.0 3.00 · plain Opus 5.5 4.00round 3DIFFERENT 1ST
t6showtime 0.2.0 5 · showtime 0.1.0 4 · plain Opus 5.5 2 · hyperframes 2showtime 0.1.0 1.00 · showtime 0.2.0 2.00 · plain Opus 5.5 3.00 · hyperframes 4.00round 3DIFFERENT 1ST
t8showtime 0.2.0 5 · showtime 0.1.0 4 · video-use 3 · plain Opus 5.5 3showtime 0.2.0 1.00 · showtime 0.1.0 2.00 · video-use 3.00 · plain Opus 5.5 4.00Round 3b rerunSAME 1ST

Where they differ. The judge ranked the 0.1.0 video above the 0.2.0 video on t1, t2, t3 and t6; the vote agreed on t1 and t2 and put the 0.2.0 video first on t3 and t6. On the filler cut the two disagree the other way: the judge put 0.2.0 first in all three judgments (all five fillers removed, against about nine "uh"s left in the 0.1.0 cut), and the vote picked the 0.1.0 video, rating 0.2.0 4 and 0.1.0 5, with the note that all four were good. The vote gives no reason. The measured differences between the two cuts are a silent 2.2 s card in the middle of the 0.2.0 cut and its loudness of −16 LUFS against −14 LUFS for the 0.1.0 cut; whether either mattered is not known. On this task the human result goes against the scorer, and it stays in the tally.

After the vote: the launch video

The reviewer rated the 0.2.0 launch video 3 of 5 and wrote that the zooms and camera movements did not make much sense while the content was good. That video was rated third of four: it lost its pairs to the 0.1.0 (round-2) video and to brag-slim and beat only plain Opus 5.5. The judges' ranking of the same video was also third (mean 2.67, behind brag-slim and the 0.1.0 video at 1.67 each).

What was done about it, in order:

  • The launch default was changed to a calmer, motivated camera: a still camera by default, at most one fly-through from the hook into the product, match cuts and a dissolve between scenes.
  • A fresh t1 run of that default was ranked by the same verified judge: mean rank 3.0 (the voted video's was 2.67). The judges penalised the stillness.
  • The reviewer watched that run and called it much better. That is an unblinded impression of a video the reviewer knew to be the new one, not a vote, and it points the other way from the judges.
  • The fly-through was smoothed after that. No judged run and no vote on this page covers the smoothed version.

No blind re-vote has been done on the new default, so it is not known whether the change fixed what was voted down. The two instruments disagree about it: the judges scored the still version lower, and the reviewer liked it more. Each is one run and one opinion, and the launch task has now cost showtime the first place in every human vote on it.

What changed, and what has been tested

What changed, and what round 3 could tell

Each problem the reviews found in a showtime video became a change to showtime's defaults or checks. After round 1 only the launch video and the filler cut had been rerun. Round 3 reran every task with 0.2.0, but an AI judge sees frames and a transcript only: it can rank craft, structure and accuracy, and it cannot hear music or feel a transition. One person, showtime's author, has since voted on the round-3 boards, partly blind; the cards below say what that vote does and does not tell about each fix.

Music taste

Flagged on the HTML report (t6) and the math explainer (t8): the music sounded cheap.

A restrained underscore is the default, with taste rules, subtler or no music for explainers and 12 dB ducking under a voice. 0.2.0 adds a catalog of produced tracks, credited automatically.

VOTE: PICKED, NOT ISOLATEDBoth tasks were rerun with the new music; the judge cannot hear it. In the round-3 vote the 0.2.0 videos on both tasks were the first pick (5 of 5 each, t6 marked as liked), above the 0.1.0 videos (4 of 5). Music was not rated on its own and a video is more than its music, so this fits better music without showing it.

Punch-in zooms on cuts

Flagged on the filler cut (t4): a zoom in and back out at the cuts.

Zooms that hide jump cuts are no longer added by default, and a check flags a zoom that snaps in and back out.

TESTEDRematch and round 3: no zooms; the judges describe the 0.2.0 cut as the same untouched picture.

Crowded chart labels

Flagged on the data story (t2): number labels crowded in places.

Chart labels avoid collisions, and a labels_crowded check warns before render.

NOT CONCLUSIVEThe 0.2.0 data story has no crowded labels, but it also dropped the decade chart that had the problem, and the judges ranked it below the round-1 video. In the round-3 vote the round-1 video was the first pick, with the note that its bar-chart numbers are crowded; the 0.2.0 video was rated the same, 4 of 5.

Captions

Flagged on the vertical short (t3): the subtitles were the weakest part.

A cleaner caption style is the new default.

MIXEDThe judges called the round-3 video's captions muddy and low-contrast in places and ranked it 3rd of 4, below the round-1 video and plain Opus 5.5; that run used stale code, and after the rerun (shadow fixed) they ranked it 2nd. The round-3 vote picked the rerun video first (5 of 5 against 4 of 5 for the round-1 video). Captions were not rated separately.

HTML report transitions

Flagged on the HTML report (t6): some chapter transitions felt abrupt.

The data template's transitions between chapters were reworked.

JUDGES: NOT JUDGEABLE · VOTE: PICKEDThe judge ranked the 0.2.0 page 2nd, behind the round-1 page, which it described as almost identical but with more context. In the round-3 vote the 0.2.0 page was the first pick (5 of 5, marked as liked, the round-1 page 4 of 5); transitions were not rated separately.

Fact discipline

Flagged by the fact-check on the launch video (t1): 4 details not in the listed sources.

The launch workflow requires every on-screen fact to trace to a source file or to output the agent ran.

HELDRound 3: 0 flagged claims on t2, t3 and t6 (0.2.0). The two flags on t8 are equations shown ahead of their tiles for a moment, a timing defect, not invented facts. On t1 the fact-check flagged 1 claim in 0.2.0 (the module list a real grep printed; the file's source is not among the listed sources) and 1 in the 0.1.0 video, which round 2 had counted as 0: one more sign that the fact-check judge is not stable.

Filler detection (0.2.0)

Flagged in rounds 1 and 2: showtime removed only 2 of 5 known fillers, video-use 4.

A speech model that keeps "um" and "uh" as words, a scan for voiced pauses no word covers, shorter pauses and 20 ms crossfades at the cuts.

SCORER AND JUDGES: IMPROVED · VOTE: PREFERRED 0.1.05 of 5 removed with all words kept (the round-1 scorer); all three judgments ranked it 1st. Left: a silent 2.2 s card, and the loudness stayed at the source's −16 LUFS. The round-3 vote picked the 0.1.0 video (5 of 5) and rated this cut 4 of 5, with no reason recorded.

Launch-video defaults (0.2.0)

Flagged in rounds 1 and 2: the reviewer picked brag-slim on the launch video twice.

Five scenes in one story, camera-move hand-offs, a produced music track, terminals that fit, text kept inside the frame.

JUDGES: IMPROVED · VOTE: LOST AGAINJudged above showtime's round-2 video in all three judgments, below brag-slim by a narrow margin; the camera cropped the headline at 3.7 s (fixed in Round 3b). In the round-3 vote (rerun video) the reviewer picked brag-slim a third time and rated 0.2.0 3 of 5, below the round-2 video, with the comment that the camera moves did not make sense. The default was changed after that and has not been re-voted (see Round 3 vote).

Limits

Read these before quoting a number