showtime benchmark, rounds 1 to 3: six video requests, five ways to answer them
Round 3 (showtime 0.2.0): two of six tasks came out better than showtime's earlier video, four worse, no level. showtime 0.2.0 was run once on the six tasks with the original prompts and ranked, by a blind judge that had to prove it opened every frame, against its own round-1 or round-2 video and the other contenders' existing outputs. It ranked first on the filler cut (all five fillers removed, was two) and moved the launch video from 3rd to 2nd of 4; it ranked below its earlier video on t2, t3, t6, t8. Cost and time per run went up. The blind boards were then voted on by one person, showtime's author, in a partly blind vote (he had seen the earlier videos): showtime was the first pick on five of six tasks, 0.2.0 on three of them, the 0.1.0 video on two, and brag-slim won the launch video again. Head to head the 0.2.0 video beat the 0.1.0 video in three of six pairs. On the filler cut the vote went against the scorer and the judges. The launch default was changed after the vote and has not been re-voted.
Round 1 and the rematch: six one-sentence video requests were typed into Claude Code and run once each with plain Opus 5.5, once with showtime, and once with the open-source tool built for that kind of job. A blind AI judge ranked showtime first on 4 tasks and second on 2.
showtime lost the launch video to brag-slim and the filler cut to plain Opus 5.5. After fixes, those two cells were rerun as a rematch: the reviewer preferred showtime on the filler cut (a partly blind vote) and brag-slim on the launch again. The judge put showtime second in both.
At a glance
The reviewer's pick is the video approved on a blind board with shuffled letters and no tool names. The judge is one AI ranking per task that saw every output side by side, in shuffled order, without names.
* The video-use run never called the video-use skill on this task, so in practice it was a second plain run.
Share of the tasks each contender entered where the blind reviewer picked it
Share of “which would you ship?” pairs won
Blind ranking judge, one judgment per task; further right is better
API-equivalent USD reported by Claude Code (subscription runs)
Minutes from start to finish, including any reply to a question
Reading these. Every mark is drawn to scale. showtime and plain Opus 5.5 entered all six tasks. Each other tool entered only the two tasks it was paired with, so its numbers come from a different, smaller set of tasks; compare it within those tasks below, not across this chart.
The rematch is left out here: it reran only showtime, on two tasks.
| Contender | Tasks | Reviewer's 1st picks | Reviewer's head-to-head wins | Mean reviewer rating | Judge mean rank | Median cost | Median wall time | qa FAIL |
|---|---|---|---|---|---|---|---|---|
| plain Opus 5.5 | all 6 | 1 of 6 | 5 of 12 | 3.0 | 2.50 | $1.20 | 6.9 min | 3 of 5 |
| showtime | all 6 | 4 of 6 | 8 of 12 | 3.7 | 1.33 | $1.66 | 7.7 min | 0 of 5 |
| brag-slim | t1, t3 | 1 of 2 | 3 of 4 | 4.0 | 2.00 | $2.16 | 13.3 min | 0 of 2 |
| hyperframes | t2, t6 | 0 of 2 | 0 of 4 | 2.5 | 2.50 | $10.66 | 15.5 min | 1 of 1 |
| video-use | t4, t8 | 0 of 2 | 2 of 4 | 3.0 | 2.00 | $1.42 | 9.2 min | 2 of 2 |
qa FAIL counts video files only; the HTML task has no qa verdict. Ratings are the reviewer's 0-5 stars on the blind boards. "Head-to-head wins" are the reviewer's answers to "which would you ship?" for every pair on a board.
Method
Every contender got the same one-sentence request, the same input files, the same model and the same limits. Only the video skill or plugin differed. The harness, tasks and scorers are in the showtime repository under benchmarks/.
Claude Code 2.1.283 with the model claude-opus-5-5 and no video skill or plugin. Claude Code's own bundled skills were still present; it used the bundled dataviz skill on the data story and the HTML report.
github.com/anthropics/claude-code
All six tasks
The project under test: a Claude Code plugin for making video locally (HTML and canvas motion graphics, voice-over, music and sound, footage edits, captions). Tested as a frozen v0.1.0 pre-release snapshot loaded with --plugin-dir, without the benchmarks/ folder. The rematch used a new snapshot with the fixes.
github.com/FavioVazquez/showtime
All six tasks
The one-file skill from the brag project (MIT). It turns a project folder or a website into a short launch video with music, motion and share copy, built by the model with the tools already on the machine. Installed as shipped.
Launch video (t1), vertical short (t3)
HeyGen's open-source HTML-to-video framework for agents (Apache-2.0). Installed with 19 of its agent skills, its CLI (hyperframes@0.8.78 from npm) and its own headless browser.
github.com/heygen-com/hyperframes
Data story (t2), HTML report (t6)
Browser Use's open-source skill for editing footage by chatting with Claude Code (MIT). Installed as the whole repository plus its Python environment. Its filler detection expects an ElevenLabs transcript; no contender got any cloud key.
github.com/browser-use/video-use
Filler cut (t4), math explainer (t8)
Each task got plain Opus 5.5, showtime, and the one outside tool whose home turf it is closest to. video-use is a footage editor, so the math task (no footage) is outside what it was built for; in that run it did not invoke its skill at all.
| Task | Prompt typed into Claude Code | Inputs in the workspace | Contenders |
|---|---|---|---|
| t1 Launch video | Make a 30-second launch video for quillsort 2.0 from this repo that I can post on X. | a small fictional Python repo: README, CHANGELOG, quillsort.py, landing page | plain, showtime, brag-slim |
| t2 Data story | Turn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube. | NOAA global temperature departures, 1850-2025 (public domain) | plain, showtime, hyperframes |
| t3 Vertical short | Make a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions. | the same quillsort repo | plain, showtime, brag-slim |
| t4 Filler cut | Cut the filler words and long pauses out of interview.mp4 and give me the tightened video. | 63.7 s of a NASA interview with STS-1 pilot Robert Crippen (public domain) | plain, showtime, video-use |
| t6 HTML report | Make a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team. | NOAA GML Mauna Loa annual CO2, 1959-2025 | plain, showtime, hyperframes |
| t8 Math explainer | Make a 45-second video that shows visually why the sum of the first n odd numbers is n squared. | none | plain, showtime, video-use |
The repository defines eight tasks. This round was cut to these six to keep it to 18 runs; the narrated explainer (t5) and the logo sting (t7) were not run.
CLAUDE.md or skills folder could load.showtime:showtime plus its MCP server, brag-slim, 19 hyperframes skills, video-use.benchmarks/, so the showtime agent could not read the task rubrics.claude-opus-5-5 at high effort, Claude Code 2.1.283, the same ffmpeg, node, python and uv on the PATH, a 25-minute cap, a $15 budget cap, the same denied commands, and subscription authentication (a token from claude setup-token).I'm not available to answer questions right now. Use your best judgment, state your assumptions, and finish the video.The question still counts against it.
claude -p with streamed output; the harness logged every message, polled the workspace for the first output file, and killed leftovers at the end. No run hit the cap. A check after each run confirmed the operator's settings were untouched.| Layer | What it does | Blind? |
|---|---|---|
| Automatic metrics | Delivered or not; format spec (length, aspect, audio, voice, captions); showtime qa on a copy of the file (loudness, true peak, clipping, silences, black and frozen stretches, frame 0, caption timing); local speech recognition for voice and footage tasks; filler, word and pause counts for the filler cut; a headless-Chrome probe with the network blocked for the HTML report. | not needed |
| Fact-check | A judge lists every factual claim (spoken, from the transcript; on screen, from frames) and marks it supported, unsupported or contradicted by the task's own files. True facts from outside those files count as unsupported. | yes |
| Ranking judge | One Opus judgment per task. It sees every output as video-1..3 in shuffled order (frames, transcripts, measurements), scores seven criteria and ranks them. | yes, position shuffled |
| Blind human review | One reviewer, who is showtime's author. One board per task: the three videos re-encoded with identical settings under shuffled letters, "which would you ship?" for each pair, 0-5 ratings, notes, and a final pick. HTML reports were screen-recorded by one script. The key mapping letters to tools stayed off the board until the vote was in. | yes in round 1; partly in the rematch |
hyperframes' HTML report run ended its turn on a question (Do you already have an idea of how this should look or what it should say?
). The harness's question detector did not recognise that phrasing, so the run never got the standard reply that every other contender would have received. The detector was widened, that one cell was rerun from scratch, and the whole round was rescored. The result shown for that cell is the rerun.
Results by task
First file: seconds until the first file of the deliverable's type appeared in the workspace. It is often a draft or an intermediate render, so it is not "time to a watchable video". Wall: whole run, including any reply to a question. Cost: the API-equivalent USD that Claude Code reports; the runs used a subscription and were not billed per token. qa: showtime qa, same thresholds for everyone. LUFS: integrated loudness (−14 is the usual target for social video). Flagged claims: claims the fact-check judge marked unsupported or contradicted.
Make a 30-second launch video for quillsort 2.0 from this repo that I can post on X.
| plain Opus 5.5 | showtime r1 | showtime rematch | brag-slim | |
|---|---|---|---|---|
| Reviewer | ★★★★★round 1: 2nd | ★★★★★round 1: 3rd | not pickedrematch, pick only | ★★★★★pickpicked in both rounds |
| Judge rank | 3rdrematch: 3rd | 2nd | 2nd | 1strematch: 1st |
| First file | 242 s | 420 s | 444 s | 577 s |
| Wall time | 265 s | 451 s | 474 s | 598 s |
| Cost | $0.88 | $2.11 | $1.89 | $2.09 |
| Turns · questions | 19 · 0 | 38 · 0 | 40 · 0 | 36 · 0 |
| qa | FAILsilent audio track; five still holds of 2.5-4.2 s | PASS | PASS | WARNcolour tags missing |
| Loudness | silent | −14.0 LUFS | −14.0 LUFS | −14.1 LUFS |
| Flagged claims | 0 | 43 are true of the repo's code (the .bak file name, "62 lines", the source shown) but not in the listed sources; 1 frame shows unsorted output | 0 | 1demo output shown unsorted; a re-check during the rematch found 2 |
brag-slim had the clearest hook (the "file10 before file2?" problem), large readable terminal text and no defects. showtime was close, with a witty opening and a good mix, but its "One Python file" code screen was too small to read on a phone. plain Opus 5.5 was silent and held still for long stretches.
In the rematch the judge said much the same: showtime covered the most content but its changelog and terminal text were small at phone size.
showtime's rerun used the new fact-discipline and music defaults. It went from 4 flagged claims to 0 and still passed qa, but it did not win the reviewer over or move up with the judge. brag-slim won this task twice.
Caveat. The new fact-discipline rule in showtime's story guide originally used a quillsort example (backup name: notes.txt.bak (ran quillsort -i …)
), taken from this very task, and the rematch agent read that guide. The drop to 0 flagged claims may partly reflect that, not only the general rule. The example has since been replaced with a generic one in the repository.
Turn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube.
| plain Opus 5.5 | showtime | hyperframes | |
|---|---|---|---|
| Reviewer | ★★★★★2nd | ★★★★★pick | ★★★★★3rd |
| Judge rank | 3rd | 1st | 2nd |
| First file | 238 s | 470 s | 1,180 s |
| Wall time | 511 s | 708 s | 1,208 s |
| Cost | $1.33 | $1.97 | $11.72 |
| Turns · questions | 30 · 0 | 10 · 0 | 21 · 17 sub-agents |
| qa | WARNno audio track; seven still holds up to 5.3 s | PASS | FAIL4.8 LU too quiet; two black flashes; flat first frame |
| Loudness | no audio | −14.0 LUFS | −18.8 LUFS |
| Flagged claims | 0 | 1a running line-end label lagged the plotted value mid-draw | 1credits NOAA, which the CSV itself does not name |
showtime was the technically cleanest: on-target loudness, no gaps, no qa findings. hyperframes was rougher (too quiet, black flashes, holds). plain Opus 5.5 had no audio and long still stretches.
Caveat. On this task the judge reported that it could not open the frame images. It ranked from the measurements and transcripts only, at low confidence (2 of 5).
Make a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions.
| plain Opus 5.5 | showtime | brag-slim | |
|---|---|---|---|
| Reviewer | ★★★★★3rd | ★★★★★pick | ★★★★★2nd |
| Judge rank | 1st | 2nd | 3rd |
| First file | 646 s | 384 s | 889 s |
| Wall time | 661 s | 438 s | 1,002 s |
| Cost | $1.81 | $1.49 | $2.23 |
| Turns · questions | 46 · 0 | 31 · 0 | 52 · 0 |
| qa | WARNa 37-character caption line; flat first frame; last 2 s silent | PASS | WARNone-frame flash at the start; colour tags missing |
| Loudness | −14.0 LUFS | −14.0 LUFS | −14.6 LUFS |
| Flagged claims | 6 flagged, 0 realall six (Python 3.8+, pip install, MIT, no dependencies, tagline) are in the README, a listed source; the judge missed them | 1a made-up quillsort list command | 1demo output shown unsorted |
plain Opus 5.5 covered all five changes with a clear card per feature and a strong hierarchy. showtime also covered all five with polished word-by-word captions and clean audio, but left the bottom half of each frame empty, used small card text and chopped the captions into short chunks. brag-slim had the strongest hook but its voice-over skipped the Python 3.7 drop.
The judge sees frames and a transcript, not the voice itself. The human reviewer ranked plain Opus 5.5 last, mainly for its voice.
Cut the filler words and long pauses out of interview.mp4 and give me the tightened video.
| plain Opus 5.5 | showtime r1 | showtime rematch | video-use | |
|---|---|---|---|---|
| Reviewer | ★★★★★pick r1rematch: close second | ★★★★★round 1: 3rd | pickrematch, partly blind, pick only | ★★★★★round 1: 2nd |
| Judge rank | 3rdrematch: 3rd | 1st | 2nd | 2ndrematch: 1st |
| First file | 224 s | 280 s | 222 s | 604 s |
| Wall time | 433 s | 489 s | 335 s | 843 s |
| Cost | $1.06 | $1.60 | $1.26 | $1.94 |
| Turns · questions | 23 · 0 | 35 · 0 | 41 · 0 | 35 · 1asked for an ElevenLabs key, then transcribed locally with Whisper |
| qa | FAILkept a 5.2 s silence; 9 s still picture; 22 kHz audio | WARN1.2 s gap; still holds | FAIL9.7 s still picture (the unedited source fails the same check) | FAIL8.6 s still picture |
| Loudness | −16.0 LUFS | −14.0 LUFS | −14.0 LUFS | −15.9 LUFS |
| Length | 56.6 s | 51.2 s | 53.0 s | 47.6 s |
| Fillers removed | 3 of 5 | 2 of 5 | 2 of 5 | 4 of 5 |
| Words kept in order | 98.8% | 100% | 100% | 99.4% |
| Longest pause left | 5.2 s | 1.2 s | 1.4 s | 2.5 s |
The reference cut, which removes exactly the five known "uh"s and the two silences, scores 5 of 5 fillers, 99.4% of words, 51.5 s. A filler counts as removed when the words around it got at least 60% of its length shorter. Both showtime runs kept every word, and the rematch run reported that its two speech models heard almost no fillers, which is where it fell short on this measure. The 5-second gaps are a silent mission-patch card in the source; every contender kept it in some form.
Round 1: showtime's transcript was clean, its loudness on target and its longest inner silence 1.2 s. video-use was the tightest but left a 2.2 s gap and was quieter. plain Opus 5.5 kept the 5.2 s silence and a stutter. Rematch: video-use first for the tightest cut, showtime second for trimming less.
Caveat. In both rounds the judge reported that it could not find the frame images, so it ranked from measurements and transcripts only. It never saw the punch-in zooms the reviewer flagged. Round 1's showtime run added a 1.12× zoom on every other section to hide jump cuts; the rematch run had none.
With punch-ins off by default, showtime's rerun went from the reviewer's 3rd (2 of 5) to the reviewer's pick. The vote was only partly blind: the other two videos were the same files the reviewer had already rated in round 1, so the new one was the only unfamiliar video on the board. The judge moved showtime from 1st to 2nd while ranking the same two competitor files again, which shows how noisy a single judgment is.
Make a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team.
| plain Opus 5.5 | showtime | hyperframes | |
|---|---|---|---|
| Reviewer | ★★★★★2nd | ★★★★★pick | ★★★★★3rd |
| Judge rank | 2nd | 1st | 3rd |
| First file | 214 sa template page | 22 sa project draft | 50 sa composition page |
| Wall time | 389 s | 349 s | 651 s |
| Cost | $1.40 | $1.05 | $9.60 |
| Turns · questions | 29 · 0 | 28 · 0 | 39 · 16 sub-agents |
| Offline probe | PASSsingle file, 0 requests, plays, fits a phone | PASSsame | PASSsame |
| What plays | a 37 s canvas animation, silent | a 32 s player with 4 chapters and music | an embedded 48 s MP4, silent |
| File size | 33 KB | 1.2 MB | 3.7 MB |
| Flagged claims | 54 true background facts not in the CSV (the observatory, NOAA, "longest record", the download feature); 1 wrong value: 1970 shown as 325.72 ppm, the CSV says 325.68 | 0 | 1"rising faster every decade", but the 1990s rose slower than the 1980s |
showtime: a self-contained player with chapters, a correct headline (315.98 → 427.35 ppm), a chart that draws itself and a phone-sized layout; one flaw, axis labels behind the title on the poster. plain Opus 5.5: clean and accurate but silent and generic. hyperframes: a good written summary, but the headline contradicts its own notes.
Caveat. Every frame the judge saw of hyperframes' page showed only the video's poster, so its video content may not have reached the judge. The human board used a separate screen recording of each page.
Make a 45-second video that shows visually why the sum of the first n odd numbers is n squared.
| plain Opus 5.5 | showtime | video-use* | |
|---|---|---|---|
| Reviewer | ★★★★★3rd | ★★★★★pick | ★★★★★2nd |
| Judge rank | 3rd | 1st | 2nd |
| First file | 196 s | 132 sa draft scene render | 232 s |
| Wall time | 267 s | 478 s | 259 s |
| Cost | $0.88 | $1.72 | $0.91 |
| Turns · questions | 18 · 0 | 43 · 0 | 17 · 0 |
| qa | FAILblack first frame; 3.6 s black gap; no audio | WARN0.47 s black blip; still holds | FAIL2.3 s black gap; no audio |
| Loudness | no audio | −14.0 LUFS | no audio |
| Math errors | 0 | 0 | 0 |
* This run had video-use installed but never invoked it, so it is effectively a second plain Opus 5.5 run. It says nothing about video-use itself.
Caveat. showtime ships a Manim example on this exact proof ("Odd numbers build squares"), and its Manim reference, which this run read, uses "Why n squared?" as a sample. The run wrote its own scenes and did not open the example files, but this task is not independent of showtime's own materials.
All three used the same correct proof (nested L-shaped layers). showtime had the tightest explanation, colour-matched equations and labelled braces, and was the only one with sound. The other two were silent, had black gaps, and let the grid run ahead of the equations.
Round 3 · showtime 0.2.0
Rounds 1 and 2 changed showtime after each review, but only the launch video and the filler cut were rerun. Round 3 runs showtime 0.2.0 once on all six original requests, with the unchanged prompts, inputs, model, limits and reply-to-questions rule, and ranks the new video next to the round-1 or round-2 showtime video and the other contenders' existing outputs (they were not rerun). Every ranking counts only if the judge proves it opened the frames.
Blind boards for a human vote were built (one per task, letters shuffled, key kept apart) and have been voted on; the result is in the Round 3 vote section below, after Round 3b. The tables in this section are the judge's. New runs: showtime snapshot of git commit f9c3de9 for t2, t3, t4, t6 and t8, and git commit c8a5502 for t1, which was run after the launch-film fixes were merged. 0.2.0 changed the default transcription to a model that hears fillers, added a catalog of produced music tracks, and reworked the launch-video defaults, alongside the round-1 fixes that had not been retested (music taste, chart labels, captions, HTML transitions).
Earlier rounds disclosed three rankings made without the images. Now every judgment (ranking, pairwise and fact-check judges alike) has to pass two checks or it is discarded and asked again in a fresh session with a fresh packet: (1) the session's own tool-call log must show that every frame image of its packet was opened with the Read tool; (2) every image carries a random 5-digit number in a black strip at the bottom, stamped on every image of every contender alike and known only to the scorer, and the judge must report the number of each image it opened (at most 10 percent wrong). A number can only be known from the pixels. A unit test suite covers the rule, including a stand-in judge that opens nothing.
Result: 18 ranking judgments, 528 of 528 images opened, 526 of 528 numbers read correctly, 0 redone. The fact-check judges made 21 attempts and one failed the check (4 of 5 numbers) and was redone. A live test with the real judge model on synthetic images read 10 of 10 numbers. This proves the images were opened and read; it does not prove they were judged well.
| Task | Ranking judgments | Images opened (tool-call log) | Reading-check numbers right | Redone for failed proof |
|---|---|---|---|---|
| t1 | 3 | 84 of 84 | 84 of 84 | 0 |
| t2 | 3 | 84 of 84 | 84 of 84 | 0 |
| t3 | 3 | 84 of 84 | 84 of 84 | 0 |
| t4 | 3 | 84 of 84 | 84 of 84 | 0 |
| t6 | 3 | 108 of 108 | 106 of 108 | 0 |
| t8 | 3 | 84 of 84 | 84 of 84 | 0 |
| all rankings | 18 | 528 of 528 | 526 of 528 | 0 |
Each row: three judgments of one packet of four videos (the new showtime video, the round-1 or 2 showtime video, plain Opus 5.5, and the outside tool paired with the task), positions rotated. qa is showtime qa version 0.2.0 run on every file again, so the old files were re-checked with the new checker.
| Task | showtime 0.1.0 (judge) | showtime 0.2.0 (judge) | 0.2.0 vs 0.1.0 | best other contender | qa 0.1.0 → 0.2.0 | cost | wall |
|---|---|---|---|---|---|---|---|
| t1 | showtime 0.1.0 3rd | showtime 0.2.0 mean 1.67 | BETTER than 0.1.0 | brag-slim 1st | PASS → PASS | $1.89 → $3.86 | 474 → 1158 s |
| t2 | showtime 0.1.0 1st | showtime 0.2.0 2nd | WORSE than 0.1.0 | plain Opus 5.5 3rd | PASS → PASS | $1.97 → $3.79 | 708 → 1240 s |
| t3 | showtime 0.1.0 1st | showtime 0.2.0 3rd | WORSE than 0.1.0 | plain Opus 5.5 2nd | PASS → PASS | $1.49 → $2.51 | 438 → 710 s |
| t4 | showtime 0.1.0 3rd | showtime 0.2.0 1st | BETTER than 0.1.0 | video-use 2nd | WARN → WARN | $1.26 → $1.72 | 335 → 375 s |
| t6 | showtime 0.1.0 1st | showtime 0.2.0 2nd | WORSE than 0.1.0 | plain Opus 5.5 3rd | HTML probe PASS → HTML probe PASS | $1.05 → $0.98 | 349 → 278 s |
| t8 | showtime 0.1.0 1st | showtime 0.2.0 3rd | WORSE than 0.1.0 | plain Opus 5.5 3rd | WARN → WARN | $1.72 → $2.59 | 478 → 703 s |
Reading this. "Better", "worse" and "level" compare the mean judge rank of the 0.2.0 video with the mean rank of the 0.1.0 video from the same packet. Each is one run against one run, with three judgments that are rotations of one packet, so a difference between two runs of the same tool may be run-to-run variation, not a change in the tool. Outputs of the other contenders are the ones from rounds 1 and 2; those tools were not rerun.
Make a 30-second launch video for quillsort 2.0 from this repo that I can post on X.
| plain Opus 5.5 | brag-slim | showtime 0.1.0round 2 output | showtime 0.2.0 | |
|---|---|---|---|---|
| Judge rank (of 4) | 4th3 judgments: 4th, 4th, 4th | mean 1.333 judgments: 2nd, 1st, 1st | 3rd3 judgments: 3rd, 3rd, 3rd | mean 1.673 judgments: 1st, 2nd, 2nd |
| Wall time | 265 s | 598 s | 474 s | 1,158 s |
| Cost | $0.88 | $2.09 | $1.89 | $3.86 |
| Turns · questions | 19 · 0 | 36 · 0 | 40 · 0 | 23 · 0 |
| qa (0.2.0 checker) | FAILthe audio track is silent; colour tags are missing ; untagged or non-BT.709 video shifts colour…; +5 more | WARNcolour tags are missing ; untagged or non-BT.709 video shifts colour… | PASS | PASS |
| Loudness | no audio | -14.1 LUFS | -14.0 LUFS | -14.0 LUFS |
| Length | 30.0 s | 30.0 s | 30.0 s | 30.0 s |
| Flagged claims | 0 | 1demo output shown unsorted: "Grace" before "ada" (as in rounds 1 and 2) | 1how plain sort orders file names (locale-dependent; round 2 had counted 0 for this same video) | 1the module list a real grep printed on screen; the file's source is not among the listed sources |
--natural --unique fixing it, an in-place edit that keeps a .bak, a grep that proves "no dependencies", an end card with the install command and the music credit. A terminal window carries the middle of the film. Music at −14 LUFS (a catalog track, credited on screen).Two judgments put brag-slim first and showtime 0.2.0 second; one put 0.2.0 first. All three put showtime 0.1.0 third and plain Opus 5.5 last. Mean rank: brag-slim 1.33, showtime 0.2.0 1.67, showtime 0.1.0 3.0, plain Opus 5.5 4.0. Rounds 1 and 2 also had brag-slim first (the reviewer picked it both times) and showtime second with a single judgment.
The agent ran the launch workflow's independent review round (one sub-agent), which took 1,158 s and $3.86, more than twice the earlier showtime run (474 s, $1.89). A second run of the same build, made while preparing this round, ran into the 25-minute cap for the same reason and is not used.
After the round-2 loss, 0.2.0 reworked how launch films are built: one story in five scenes, hand-offs by camera moves rather than 0.4 s pushes, a real produced music track, terminals that fit their window, text kept inside the frame. Result: the judges rank this video above the round-2 one in all three judgments (mean 1.67 against 3.0) and below brag-slim by a narrow margin, with a cropped headline as the visible flaw. In the round-3 human vote (one voter, partly blind) the reviewer picked brag-slim for a third time, and rated this video 3 of 5 (4 for the round-2 video, 5 for brag-slim), with the comment that the zooms and camera movements did not make much sense; see Round 3 vote. JUDGES: IMPROVED · VOTE: LOST AGAIN
Turn global-temperature-anomaly.csv into a 45-second animated data story video for YouTube.
| plain Opus 5.5 | hyperframes | showtime 0.1.0round 1 output | showtime 0.2.0 | |
|---|---|---|---|---|
| Judge rank (of 4) | mean 3.333 judgments: 4th, 3rd, 3rd | mean 3.673 judgments: 3rd, 4th, 4th | 1st3 judgments: 1st, 1st, 1st | 2nd3 judgments: 2nd, 2nd, 2nd |
| Wall time | 511 s | 1,208 s | 708 s | 1,240 s |
| Cost | $1.33 | $11.72 | $1.97 | $3.79 |
| Turns · questions | 30 · 0 | 21 · 1 | 10 · 0 | 13 · 0 |
| qa (0.2.0 checker) | WARNcolour tags are None/None/bt709 ; untagged or non-BT.709 video shifts…; the video has no audio track; +8 more | FAILintegrated loudness -18.8 LUFS is -4.8 LU off the -14 LUFS target; frame 0 is a flat colour (mean luma 17); +4 more | PASS | PASS |
| Loudness | no audio | -18.8 LUFS | -14.0 LUFS | -14.0 LUFS |
| Length | 45.0 s | 45.0 s | 45.0 s | 45.0 s |
| Flagged claims | 0 | 1 | 0 | 0 |
All three judgments put showtime 0.1.0 first and 0.2.0 second, for the reason they all gave: 0.2.0 has "less variety" and long holds; its style, numbers (correct) and audio (−14 LUFS, a captions file) were praised. All three judgments noted that the round-1 video's decade-bar labels crowd or overlap slightly; none noted crowding in the 0.2.0 video.
Cost and time went up: 1,240 s and $3.79 against 708 s and $1.97, with only 13 turns (the machine also carried two other runs and other jobs at the time).
Round 1's reviewer found the number labels crowded. The 0.2.0 video has no crowded labels, but it also has no decade chart, the chart that had the problem, so this run does not show the collision fix working. In the round-3 vote the reviewer picked the 0.1.0 video and noted that its numbers are crowded in the bar charts (the same complaint as in round 1); the 0.2.0 video got no such note and was rated 4 of 5, the same as the 0.1.0 video. NOT CONCLUSIVE
Make a 30-second vertical short for Instagram Reels announcing what's new in quillsort 2.0 (see CHANGELOG.md), with a voiceover and captions.
| plain Opus 5.5 | brag-slim | showtime 0.1.0round 1 output | showtime 0.2.0 | |
|---|---|---|---|---|
| Judge rank (of 4) | 2nd3 judgments: 2nd, 2nd, 2nd | 4th3 judgments: 4th, 4th, 4th | 1st3 judgments: 1st, 1st, 1st | 3rd3 judgments: 3rd, 3rd, 3rd |
| Wall time | 661 s | 1,002 s | 438 s | 710 s |
| Cost | $1.81 | $2.23 | $1.49 | $2.51 |
| Turns · questions | 46 · 0 | 52 · 0 | 31 · 0 | 44 · 0 |
| qa (0.2.0 checker) | WARNcolour tags are missing ; untagged or non-BT.709 video shifts colour…; frame 0 is a flat colour (mean luma 248); +4 more | WARNcolour tags are missing ; untagged or non-BT.709 video shifts colour…; frame 0 differs sharply from frame 1 (mean diff 61/255): a one-frame… | PASS | PASS |
| Loudness | -14.0 LUFS | -14.6 LUFS | -14.0 LUFS | -14.0 LUFS |
| Length | 30.0 s | 30.0 s | 30.0 s | 30.0 s |
| Flagged claims | 0 | 1demo output shown unsorted (as in round 1) | 2"3.8 or newer" and "pip install quillsort", both in the README, a listed source: the fact-check judge was wrong | 0 |
All three judgments ranked showtime 0.1.0 first, plain Opus 5.5 second, showtime 0.2.0 third and brag-slim fourth. In round 1 the single judge had put showtime second behind plain Opus 5.5; with the frames verified it puts the round-1 video first.
Round 1's reviewer called the subtitles the weakest part, and the new clean-pop caption default was meant to fix that. The judges did not see an improvement: their notes on the 0.2.0 video (muddy outlines, greyed words, empty lower half) are at least as critical as their notes on the round-1 video. The two videos also differ in look for reasons beyond the caption style, since the agent chose a different theme each time. This round-3 run ran stale code; after the rerun (caption shadow fixed, Round 3b) the judges ranked it 2nd, behind the 0.1.0 video, and in the round-3 vote (one voter, partly blind) the rerun video was the first pick, 5 of 5 against 4 of 5 for the 0.1.0 video. Captions were not rated separately. MIXED: JUDGES NO IMPROVEMENT · VOTE PICKED IT
Cut the filler words and long pauses out of interview.mp4 and give me the tightened video.
| plain Opus 5.5 | video-use | showtime 0.1.0round 2 output | showtime 0.2.0 | |
|---|---|---|---|---|
| Judge rank (of 4) | 4th3 judgments: 4th, 4th, 4th | 2nd3 judgments: 2nd, 2nd, 2nd | 3rd3 judgments: 3rd, 3rd, 3rd | 1st3 judgments: 1st, 1st, 1st |
| Wall time | 433 s | 843 s | 335 s | 375 s |
| Cost | $1.06 | $1.94 | $1.26 | $1.72 |
| Turns · questions | 23 · 0 | 35 · 1 | 41 · 0 | 48 · 0 |
| qa (0.2.0 checker) | FAILsilence from 19.25s to 24.45s (5.2s) inside the video; integrated loudness -16.0 LUFS is -2.0 LU off the -14 LUFS target; +3 more | WARNintegrated loudness -15.9 LUFS is -1.9 LU off the -14 LUFS target; silence from 18.55s to 20.80s (2.2s) inside the video | WARNsilence from 21.75s to 23.15s (1.4s) inside the video; the last 2.5s are silent (sound ends at 50.50s) | WARNintegrated loudness -16.0 LUFS is -2.0 LU off the -14 LUFS target; silence from 20.15s to 22.40s (2.2s) inside the video; +1 more |
| Loudness | -16.0 LUFS | -15.9 LUFS | -14.0 LUFS | -16.0 LUFS |
| Length | 56.6 s | 47.6 s | 53.0 s | 49.7 s |
| Flagged claims | not checkedthe speaker's own words | not checkedthe speaker's own words | not checkedthe speaker's own words | not checkedthe speaker's own words |
| Fillers removed (rounds 1-2 scorer) | 3 of 5 | 4 of 5 | 2 of 5 | 5 of 5 |
| Words kept in order | 98.8% | 99.4% | 100% | 100% |
| "uh"/"um" still audible (verbatim transcript) | 4 | 2 | 9 | 0 |
| Longest pause left | 5.2 s | 2.5 s | 1.4 s | 2.3 s |
Fillers removed uses exactly the round-1 and round-2 scorer (Whisper turbo transcribes the source and each output; a filler counts as removed when the words around it shrank by at least 60% of its length), so the numbers are comparable with the earlier rounds: 0.1.0 scored 2 of 5 in both, video-use 4, plain Opus 5.5 3. The verbatim count is a different measure: showtime 0.2.0's new default speech model writes "uh" and "um" as words, and the same model was used on every file, but it is also the model showtime now ships, so treat it as the tested tool's own instrument. The reference cut (exactly the five known "uh" and the two silences) scores 5 of 5, 99.4%, 51.5 s. qa verdicts on old files changed for reasons unrelated to the videos: since round 2 the checker no longer fails a talking head that simply holds still.
All three judgments ranked showtime 0.2.0 first: "the only version that removes all the filler", with the sentences intact. video-use second, showtime 0.1.0 third (barely changed the speech), plain Opus 5.5 last (a 5.2 s frozen silent card, a clipped word). Visually all four are the same letterboxed webcam picture, so the ranking rests on the transcript and the measurements.
In rounds 1 and 2 the same kind of ranking put showtime 1st, then 2nd, without seeing the frames; this one was verified.
0.2.0 replaced the default speech model and added a filler scan after it. This is the retest: 5 of 5 known fillers removed (was 2 of 5) with all words kept, and first place with every judgment. The reviewer's round-2 vote (partly blind) had picked the 0.1.0 video, whose fillers were mostly still there, and the round-3 vote did so again: the 0.1.0 video was the first pick (5 of 5) and this cut was rated 4 of 5, with no reason recorded. SCORER AND JUDGES: IMPROVED · VOTE: PREFERRED 0.1.0
Make a shareable single-file HTML video report on the Mauna Loa CO2 record in mauna-loa-co2-annual.csv that I can send to my team.
| plain Opus 5.5 | hyperframes | showtime 0.1.0round 1 output | showtime 0.2.0 | |
|---|---|---|---|---|
| Judge rank (of 4) | 3rd3 judgments: 3rd, 3rd, 3rd | 4th3 judgments: 4th, 4th, 4th | 1st3 judgments: 1st, 1st, 1st | 2nd3 judgments: 2nd, 2nd, 2nd |
| Wall time | 389 s | 651 s | 349 s | 278 s |
| Cost | $1.40 | $9.60 | $1.05 | $0.98 |
| Turns · questions | 29 · 0 | 39 · 1 | 28 · 0 | 23 · 0 |
| Offline probe | PASSsingle file, 0 requests, plays, fits a phone | PASSsingle file, 0 requests, plays, fits a phone | PASSsingle file, 0 requests, plays, fits a phone | PASSsingle file, 0 requests, plays, fits a phone |
| Flagged claims | 32 true background facts not in the CSV; 1 value wrong (1970: 325.72 ppm, CSV 325.68) | 1"rising faster every decade": the 1990s rose slower than the 1980s | 0 | 0 |
All three judgments ranked showtime 0.1.0 first and 0.2.0 second, then plain Opus 5.5 and hyperframes. All three said the two showtime pages look almost alike, and that 0.1.0 tells the fuller story.
The judge sees frames from a screen recording of each page playing, not the sound or the smoothness of a transition. It reports "an audio element but no speech", and cannot say whether the music or the transitions got better.
Both are about how the page feels in motion and sound (round 1's reviewer found the music cheap-sounding and some transitions abrupt). The judges cannot hear or feel either, so this round says nothing about them. The recorded 0.2.0 page has an audio track (captured, peak −2.7 dB) and the human board played it. In the round-3 vote (one voter, partly blind) the 0.2.0 page was the first pick, 5 of 5 and marked as liked, over the 0.1.0 page at 4 of 5, while the judges ranked it second; music and transitions were not rated separately, so the vote does not isolate them. JUDGES: NOT JUDGEABLE · VOTE: PICKED, NOT ISOLATED
Make a 45-second video that shows visually why the sum of the first n odd numbers is n squared.
| plain Opus 5.5 | video-use* | showtime 0.1.0round 1 output | showtime 0.2.0 | |
|---|---|---|---|---|
| Judge rank (of 4) | 3rd3 judgments: 4th, 2nd, 3rd | 3rd3 judgments: 3rd, 4th, 2nd | 1st3 judgments: 1st, 1st, 1st | 3rd3 judgments: 2nd, 3rd, 4th |
| Wall time | 267 s | 259 s | 478 s | 703 s |
| Cost | $0.88 | $0.91 | $1.72 | $2.59 |
| Turns · questions | 18 · 0 | 17 · 0 | 43 · 0 | 47 · 0 |
| qa (0.2.0 checker) | FAILframe 0 is black (black for the first 0.30s): feeds, chat apps and…; black from 3.47s to 7.10s (3.63s); +6 more | FAILblack from 3.80s to 6.13s (2.33s); colour tags are missing ; untagged or non-BT.709 video shifts colour…; +6 more | WARNthe picture does not change from 0.00s to 3.13s (3.1s; small changes…; the picture does not change from 3.13s to 6.53s (3.4s; small changes…; +3 more | WARNblack from 9.03s to 9.47s (0.43s); the picture does not change from 10.40s to 14.23s (3.8s; small… |
| Loudness | no audio | no audio | -14.0 LUFS | -14.0 LUFS |
| Length | 45.0 s | 45.0 s | 45.0 s | 45.0 s |
| Flagged claims | 0 | 0 | 0 | 2two equations shown next to the wrong-sized grid mid-transition: "1+3+5 = 2²" (2×2 grid) and "1+3+5+7+9 = 4²" at about 24 s |
showtime 0.1.0 came first in all three judgments: "the equations always match the tiles". 0.2.0, plain Opus 5.5 and the video-use arm then traded places (a mean rank of 3.0 each), so the judges could not separate them. All three judgments call out the mismatched equations in the 0.2.0 video as an accuracy failure in a proof.
Caveat. showtime ships a Manim example on this same proof, written before the task existed, and its reference uses "Why n squared?" as a sample; round 1's run read that reference. That still holds in round 3 and favours showtime on this task.
0.2.0 picks music from a catalog of produced tracks and prefers subtler or none for explainers. The judge cannot hear it, so its taste is untested here; the video keeps a music bed at −14.0 LUFS. A new defect appeared instead: equations that lead their tiles, flagged by the judges and the fact-check. NEW DEFECT
Round 3b · regression rerun of showtime 0.2.0
Round 3's t2 (data story) and t3 (vertical short) did not measure showtime 0.2.0. The showtime arm's runtime folder was a link to a runtime folder shared with other work on the same machine. Its stable showtime command runs whichever skill folder was recorded there last, and that was another, unfinished working copy of showtime (24 commits older than the benchmarked commit, with uncommitted edits). showtime doctor tells agents that this command "always works", so the t2 agent sent all 18 of its showtime commands through it and the t3 agent all 12; the t1 agent noticed ("points at a different (older) install") and switched back. t4, t6 and t8 used the skill's own command and were not affected.
Causes found by comparing the 0.1.0 and 0.2.0 transcripts with the judges' reasons:
showtime check never sampled inside that transition. Fixed: the fly-through fades every word outside the portal's own word before the frame edge reaches it; check now samples inside it and reports cut text as an error.Rerun: the fixed showtime 0.2.0, once per task, same prompts, inputs, model, limits and harness, each arm with its own command. Ranked with the same verified judge (3 judgments, frames opened and numbers read) against the same three other candidates as in round 3. t4 and t6 were not rerun (no 0.2.0 cause found; t6 lost first place because its video was shorter and one frame caught a count-up mid-animation).
| Task | 0.2.0 mean rank, round 3 | 0.2.0 mean rank, rerun | 0.1.0 in the rerun judging | best other contender (rerun judging) | cost | wall |
|---|---|---|---|---|---|---|
| t1 launch video | 1.67 | 2.67 | 1.67 | brag-slim 1.67 | $3.86 → $3.73 | 1158 → 864 s |
| t2 data story | 2.00 | 1.67 | 1.33 | plain Opus 5.5 3.33 | $3.79 → $5.63 | 1240 → 1278 s |
| t3 vertical short | 3.00 | 2.00 | 1.33 | plain Opus 5.5 2.67 | $2.51 → $2.41 | 710 → 676 s |
| t8 math explainer | 3.00 | 1.00 | 2.00 | video-use 3.00 | $2.59 → $2.35 | 703 → 857 s |
t1 was run twice: the first rerun (before the history-line fix) ranked 3.0 at $1.09 and 437 s, with no cropped headline but a terminal that read as broken output; the row above is the second rerun, after that fix. Its judges found no crop and no broken output, but ranked it behind brag-slim and the 0.1.0 video because it shows three features instead of four and keeps one look throughout. t2, t3 and t8 are from the first and only rerun. One run per task: a single rerun can move a rank by one place on variance alone, and the rerun judging is a new set of three judgments of a new packet, so the 0.1.0 column can differ from round 3's. Costs include any helper agents the run started (t2's rerun started a fact-check researcher and a critic, which found a real type problem).
Round 3 · human vote
The blind boards built for round 3 (one per task, four unnamed videos with shuffled letters, "which would you ship?" pairs and a 0-5 rating for every video) were voted on 29 September. The six "Copy for your agent" digests were imported with the benchmark's own parser (human_board.py, parse_digest, the same way rounds 1 and 2 were imported) and tallied through the letter-to-tool key, which was kept apart until after the vote. Every task has one first pick, five pairwise answers (of the six possible pairs) and four ratings.
Which videos were on the boards. showtime 0.2.0 is the rerun video for t1 (the second rerun), t2, t3 and t8 (Round 3b, after the fixes), and the original round-3 video for t4 and t6. The showtime 0.1.0 video is the round-1 video for t2, t3, t6 and t8 and the round-2 video for t1 and t4. The other tools' videos are the round-1 outputs, as in round 3. So on four tasks the vote is about different showtime 0.2.0 videos from the ones in the Round 3 tables above.
It is one person, and that person is showtime's author. The reviewer voted once on each of six tasks: 30 pairwise answers and 24 ratings in all. It is the author's own vote, not a panel, and it cannot say how anyone else would vote.
It is only partly blind. The letters were shuffled and no tool was named, but the reviewer had already seen the earlier videos (the round-1 and round-2 outputs, the outside tools' among them), so an old video can be recognised on sight. The direction of that bias is not known.
Nothing on the boards records the reason for a rating, except where a note or comment was written (quoted below).
| Task | First pick | showtime 0.2.00-5 rating | showtime 0.1.00-5 rating | plain Opus 5.50-5 rating | outside tool0-5 rating |
|---|---|---|---|---|---|
| t1 launch video from a repo | brag-slim | 3/5 | 4/5 | 2/5 | 5/5first pick |
| t2 data story from a CSV | showtime 0.1.0 | 4/5 | 4/5first pick | 3/5 | 4/5 |
| t3 vertical short with voice-over and captions | showtime 0.2.0 | 5/5first pick | 4/5 | 2/5 | 2/5 |
| t4 footage edit | showtime 0.1.0 | 4/5 | 5/5first pick | 3/5 | 3/5 |
| t6 single-file HTML video report | showtime 0.2.0 | 5/5first pick | 4/5 | 2/5 | 2/5 |
| t8 math explainer | showtime 0.2.0 | 5/5first pick | 4/5 | 3/5 | 3/5 |
Comments and notes, in the reviewer's own words: t1, “The zooms here and camera movements didn't make much sense, the content was good though” (on the showtime 0.2.0 video, rated 3 of 5). t2, “B is the best but the numbers are crowded in the bar charts” ("B" was the showtime 0.1.0 video, the first pick; the reviewer rated the 0.2.0 video 4 of 5, the same as the 0.1.0 video and as hyperframes). t4, “All were good, but I think D nailed it” ("D" was the showtime 0.1.0 video). t6: the pick, the 0.2.0 video, was marked as liked. No other comments. The outside tool is brag-slim on t1 and t3, hyperframes on t2 and t6, video-use on t4 and t8. * The video-use run on t8 never called its skill.
| Task | showtime 0.2.0 in the pairs asked | Other pairs asked |
|---|---|---|
| t1 |
|
|
| t2 |
|
|
| t3 |
|
|
| t4 |
|
|
| t6 |
|
|
| t8 |
|
|
| Contender | Tasks | First picks | Pairs won | Mean 0-5 ratingand each rating |
|---|---|---|---|---|
| showtime 0.2.0 | all 6 | 3 of 6 | 14 of 18 | 4.3 3, 4, 5, 4, 5, 5 |
| showtime 0.1.0 | all 6 | 2 of 6 | 12 of 16 | 4.2 4, 4, 4, 5, 4, 4 |
| plain Opus 5.5 | all 6 | 0 of 6 | 0 of 12 | 2.5 2, 3, 2, 3, 2, 3 |
| brag-slim | 2 of 6 | 1 of 2 | 2 of 4 | 3.5 5, 2 |
| hyperframes | 2 of 6 | 0 of 2 | 1 of 5 | 3.0 4, 2 |
| video-use | 2 of 6 | 0 of 2 | 1 of 5 | 3.0 3, 3 |
Reading this. showtime, either version, was the first pick on five of six tasks: 0.2.0 on three (t3, t6, t8), the 0.1.0 video on two (t2, t4) and brag-slim on the launch video (t1). Head to head, the 0.2.0 video won three of six pairs against the 0.1.0 video and lost three: it lost t1, t2 and t4 and won t3, t6 and t8. Its ratings were higher than the 0.1.0 video's on three tasks (t3, t6, t8), the same on one (t2) and lower on two (t1, t4); mean 4.3 against 4.2, a difference of 0.2 on a 0-5 scale from one voter. The 0.2.0 video beat plain Opus 5.5 in all six pairs and beat the outside tool in five of six; the loss was the launch video to brag-slim, for the third time in three votes on that task. Each outside tool has only two tasks, so its totals are two tasks and no more.
For comparison, the verified judge's mean ranks for the same videos the vote used: the rerun judging for t1, t2, t3 and t8, the round-3 judging for t4 and t6. The two agree on the first place in two tasks, tie in one, and differ in three.
| Task | Vote: 0-5 rating (1st pick first) | Judge: mean rank on the same video (lower is better) | 1st place |
|---|---|---|---|
| t1 | brag-slim 5 · showtime 0.1.0 4 · showtime 0.2.0 3 · plain Opus 5.5 2 | showtime 0.1.0 1.67 · brag-slim 1.67 · showtime 0.2.0 2.67 · plain Opus 5.5 4.00Round 3b, second rerun | TIED 1ST |
| t2 | showtime 0.1.0 4 · showtime 0.2.0 4 · hyperframes 4 · plain Opus 5.5 3 | showtime 0.1.0 1.33 · showtime 0.2.0 1.67 · plain Opus 5.5 3.33 · hyperframes 3.67Round 3b rerun | SAME 1ST |
| t3 | showtime 0.2.0 5 · showtime 0.1.0 4 · plain Opus 5.5 2 · brag-slim 2 | showtime 0.1.0 1.33 · showtime 0.2.0 2.00 · plain Opus 5.5 2.67 · brag-slim 4.00Round 3b rerun | DIFFERENT 1ST |
| t4 | showtime 0.1.0 5 · showtime 0.2.0 4 · video-use 3 · plain Opus 5.5 3 | showtime 0.2.0 1.00 · video-use 2.00 · showtime 0.1.0 3.00 · plain Opus 5.5 4.00round 3 | DIFFERENT 1ST |
| t6 | showtime 0.2.0 5 · showtime 0.1.0 4 · plain Opus 5.5 2 · hyperframes 2 | showtime 0.1.0 1.00 · showtime 0.2.0 2.00 · plain Opus 5.5 3.00 · hyperframes 4.00round 3 | DIFFERENT 1ST |
| t8 | showtime 0.2.0 5 · showtime 0.1.0 4 · video-use 3 · plain Opus 5.5 3 | showtime 0.2.0 1.00 · showtime 0.1.0 2.00 · video-use 3.00 · plain Opus 5.5 4.00Round 3b rerun | SAME 1ST |
Where they differ. The judge ranked the 0.1.0 video above the 0.2.0 video on t1, t2, t3 and t6; the vote agreed on t1 and t2 and put the 0.2.0 video first on t3 and t6. On the filler cut the two disagree the other way: the judge put 0.2.0 first in all three judgments (all five fillers removed, against about nine "uh"s left in the 0.1.0 cut), and the vote picked the 0.1.0 video, rating 0.2.0 4 and 0.1.0 5, with the note that all four were good. The vote gives no reason. The measured differences between the two cuts are a silent 2.2 s card in the middle of the 0.2.0 cut and its loudness of −16 LUFS against −14 LUFS for the 0.1.0 cut; whether either mattered is not known. On this task the human result goes against the scorer, and it stays in the tally.
The reviewer rated the 0.2.0 launch video 3 of 5 and wrote that the zooms and camera movements did not make much sense while the content was good. That video was rated third of four: it lost its pairs to the 0.1.0 (round-2) video and to brag-slim and beat only plain Opus 5.5. The judges' ranking of the same video was also third (mean 2.67, behind brag-slim and the 0.1.0 video at 1.67 each).
What was done about it, in order:
No blind re-vote has been done on the new default, so it is not known whether the change fixed what was voted down. The two instruments disagree about it: the judges scored the still version lower, and the reviewer liked it more. Each is one run and one opinion, and the launch task has now cost showtime the first place in every human vote on it.
What changed, and what has been tested
Each problem the reviews found in a showtime video became a change to showtime's defaults or checks. After round 1 only the launch video and the filler cut had been rerun. Round 3 reran every task with 0.2.0, but an AI judge sees frames and a transcript only: it can rank craft, structure and accuracy, and it cannot hear music or feel a transition. One person, showtime's author, has since voted on the round-3 boards, partly blind; the cards below say what that vote does and does not tell about each fix.
Flagged on the HTML report (t6) and the math explainer (t8): the music sounded cheap.
A restrained underscore is the default, with taste rules, subtler or no music for explainers and 12 dB ducking under a voice. 0.2.0 adds a catalog of produced tracks, credited automatically.
VOTE: PICKED, NOT ISOLATEDBoth tasks were rerun with the new music; the judge cannot hear it. In the round-3 vote the 0.2.0 videos on both tasks were the first pick (5 of 5 each, t6 marked as liked), above the 0.1.0 videos (4 of 5). Music was not rated on its own and a video is more than its music, so this fits better music without showing it.
Flagged on the filler cut (t4): a zoom in and back out at the cuts.
Zooms that hide jump cuts are no longer added by default, and a check flags a zoom that snaps in and back out.
TESTEDRematch and round 3: no zooms; the judges describe the 0.2.0 cut as the same untouched picture.
Flagged on the data story (t2): number labels crowded in places.
Chart labels avoid collisions, and a labels_crowded check warns before render.
NOT CONCLUSIVEThe 0.2.0 data story has no crowded labels, but it also dropped the decade chart that had the problem, and the judges ranked it below the round-1 video. In the round-3 vote the round-1 video was the first pick, with the note that its bar-chart numbers are crowded; the 0.2.0 video was rated the same, 4 of 5.
Flagged on the vertical short (t3): the subtitles were the weakest part.
A cleaner caption style is the new default.
MIXEDThe judges called the round-3 video's captions muddy and low-contrast in places and ranked it 3rd of 4, below the round-1 video and plain Opus 5.5; that run used stale code, and after the rerun (shadow fixed) they ranked it 2nd. The round-3 vote picked the rerun video first (5 of 5 against 4 of 5 for the round-1 video). Captions were not rated separately.
Flagged on the HTML report (t6): some chapter transitions felt abrupt.
The data template's transitions between chapters were reworked.
JUDGES: NOT JUDGEABLE · VOTE: PICKEDThe judge ranked the 0.2.0 page 2nd, behind the round-1 page, which it described as almost identical but with more context. In the round-3 vote the 0.2.0 page was the first pick (5 of 5, marked as liked, the round-1 page 4 of 5); transitions were not rated separately.
Flagged by the fact-check on the launch video (t1): 4 details not in the listed sources.
The launch workflow requires every on-screen fact to trace to a source file or to output the agent ran.
HELDRound 3: 0 flagged claims on t2, t3 and t6 (0.2.0). The two flags on t8 are equations shown ahead of their tiles for a moment, a timing defect, not invented facts. On t1 the fact-check flagged 1 claim in 0.2.0 (the module list a real grep printed; the file's source is not among the listed sources) and 1 in the 0.1.0 video, which round 2 had counted as 0: one more sign that the fact-check judge is not stable.
Flagged in rounds 1 and 2: showtime removed only 2 of 5 known fillers, video-use 4.
A speech model that keeps "um" and "uh" as words, a scan for voiced pauses no word covers, shorter pauses and 20 ms crossfades at the cuts.
SCORER AND JUDGES: IMPROVED · VOTE: PREFERRED 0.1.05 of 5 removed with all words kept (the round-1 scorer); all three judgments ranked it 1st. Left: a silent 2.2 s card, and the loudness stayed at the source's −16 LUFS. The round-3 vote picked the 0.1.0 video (5 of 5) and rated this cut 4 of 5, with no reason recorded.
Flagged in rounds 1 and 2: the reviewer picked brag-slim on the launch video twice.
Five scenes in one story, camera-move hand-offs, a produced music track, terminals that fit, text kept inside the frame.
JUDGES: IMPROVED · VOTE: LOST AGAINJudged above showtime's round-2 video in all three judgments, below brag-slim by a narrow margin; the camera cropped the headline at 3.7 s (fixed in Round 3b). In the round-3 vote (rerun video) the reviewer picked brag-slim a third time and rated 0.2.0 3 of 5, below the round-2 video, with the comment that the camera moves did not make sense. The default was changed after that and has not been re-voted (see Round 3 vote).
Limits
showtime qa, its frame sampler and its local speech recognition were applied identically to every output, and the raw numbers are kept, but they are the tested tool's own checks. Its "still picture" rule, for example, fails a talking-head interview that simply holds still, including the unedited source..bak example in showtime's story guide, and the launch rematch read it; that example has since been replaced with a generic one. Both favour showtime on those tasks.dataviz skill on the data story and the HTML report.