Seventeen models, one briefing, one rubric
Every night a small pipeline turns a news digest into a five-minute spoken briefing. In June I stopped and made seventeen different model configurations write the identical script — frontier and local, thinking and not, alone and in author–editor pairs - then scored them all against the same rubric.
The results were not what parameter counts predict. Thinking mode made models measurably worse at prose. Six minutes of deliberation lost to twenty-two seconds. And rewriting the prompt bought more than any jump in model size I measured that night.
The same model, same prompt, thinking mode switched on. Structure collapsed into a run-on open and the rhythm went staccato.
What rewriting the prompt bought a mid-size model - more than any parameter jump measured that night, at twenty-four seconds a run.
Six minutes and 44,000 characters of reasoning scored below the same model thinking for twenty-two seconds.
The rules
Every entry received the identical input - the same digest of the day's stories and the same weather data - and the same instruction: write a TTS-safe briefing of roughly 620 words. Nothing was cherry-picked; each model got one honest run at production settings, and the failures are in the table alongside the wins.
Scoring is a weighted blend of six dimensions. Completeness counts double, which is a deliberate choice and the single most consequential one here: a briefing that sounds magnificent while silently dropping a third of the news is worse than a plodding one that covers everything. Several models are ranked below where their prose alone would put them for exactly this reason. TTS violations - text a speech engine mangles, like "F1" read as a keyboard key - are counted separately and subtracted.
Every run, scored
Green is good and red is bad, per cell. Overall is the weighted blend after TTS violations are subtracted. Two further configurations were queued and never finished; they are omitted rather than guessed at.
| Model | Complete ×2 | Judgment | Voice | Personal | Weather | TTS viol. | Gen time | Audio len | Overall | What happened |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 👑interactive session | 9 | 9 | 9 | 9 | 9 | 0 | ~1 min | 3:46 | 9.0 | The benchmark. Full coverage, a narrative through-line (arrival → storm → callback close), and flourishes that felt earned rather than bolted on. |
| Claude Sonnet 4.6claude -p, metered | 7 | 9 | 9 | 8 | 9 | 0 | ~1 min | 4:29 | 8.0 | Best analysis of the night - genuine cross-story synthesis and human beats. But it missed the climate story, and the completeness rule cuts both ways. |
| Claude Haiku 4.5claude -p, metered | 9 | 7 | 7 | 7 | 7 | 1 | ~30 s | 4:27 | 7.5 | Caught the climate story with full numbers; near-complete coverage. One pronunciation violation, one tense slip, weather read as half-recital. Startling value for the cheapest Claude. |
| qwen3.5:122bno thinking · 50 s | 9 | 7 | 7 | 7 | 7 | 0 | 50 s | 4:09 | 7.0 | Local champion. Complete, warm close, attempted localisation. Personalisation landed on-the-nose rather than woven in; one grammar slip. |
| gpt-oss:120breasoning: low · 22 s | 8 | 8 | 7 | 6 | 5 | 1 | 22 s | 3:58 | 7.0 | Best local news judgment - picked the same top three as Claude and derived a fact none of the sources stated outright. Weather was a polite data recital; one logic blunder. |
| gemma4:31bGGUF · 9 tok/s | 5 | 7 | 8 | 6 | 8 | 0 | 215 s | 3:09 | 6.0 | Cleanest local prose of the night, and near the bottom anyway: it dropped four stories and came in at 514 words against a 550–650 target. |
| qwen3.6 MoE20 s | 6 | 6 | 5 | 5 | 6 | 1 | 20 s | 4:13 | 5.5 | Fast and broad but flat. Monotone "Now to…" transitions, weather repeated in the intro, missed the climate story. |
| qwen3.5:122bthinking · 15K chars of rumination | 6 | 5 | 3 | 4 | 5 | 0 | 230 s | 3:22 | 4.0 | Thinking actively hurt. Structure collapsed into a run-on open, rhythm went staccato, and it added presumptuous filler. An earlier attempt thought itself out of context entirely and emitted nothing at all. |
| gemma4:31b + persona164 s | Result: 398 words - shorter than its own first run. Context does not fix gemma's coverage problem; brevity is a trait of the model, not a gap in the prompt. | |||||||||
| gpt-oss:120b + personareasoning: low · 22 s | 7 | 7 | 8 | 8 | 7 | 1 | 22 s | 3:40 | 7.5 | Strongest local edition of the night. Best use of the persona block, real flow, and editorial instinct - it promoted the local story unprompted. One confabulated analysis line. |
| gpt-oss:120b + personareasoning: high · 44K chars of thinking | 8 | 7 | 5 | 7 | 5 | 0 | 352 s | 3:49 | 6.5 | The replication. Six minutes of deliberation bought better facts and worse rhythm - a wall of uniform declaratives. "More thinking, worse broadcast voice" now holds across two architectures. |
| qwen3.5:122b + persona52 s | 9 | 7 | 7 | 8 | 6 | 0 | 52 s | 4:03 | 7.5 | The persona block sharpened it: a personalised callout to [family member]'s interest, and a sign-off by name. Weather still half-recital. Ties gpt-oss-low for best single local model. |
| Ensembleqwen122 author + qwen122-thinking editor | 9 | 6 | 7 | 7 | 8 | 1 | ~7 min | 5:29 | 7.5 | Most complete coverage of the night - the editor caught a story the author dropped entirely. But the revision "fixed" a flagged confabulation by deleting a real story, kept a different flagged invention, and coined a grammatical monstrosity. Matched the best singles, didn't beat them. |
| gemma4:26b MoE + prompt v3124 s | 9 | 7 | 8 | 6 | 7 | 0 | 124 s | 4:02 | 7.5 | +1.0 from prompt craft alone (6.5 → 7.5). A coverage contract fixed the shortness that a persona block couldn't: every story present, on budget, zero violations. |
| qwen3.6 MoE + prompt v324 s | 9 | 7 | 7 | 6 | 8 | 0 | 24 s | 5:50 | 7.0 | +1.5 from prompt craft (5.5 → 7.0) in twenty-four seconds. New failure mode: exemplar leakage - it imported a story from the gold sample and reported it as today's news. Ran 35% over budget. |
| qwen3.5:9b + prompt v3the digest compiler, writing | 4 | 4 | 2 | 4 | 2 | 2 | 117 s | 4:41 | 3.0 | The capability floor, demonstrated. Run-on word salad, third-person greeting, weather fused into nonsense, a confabulated weather event. No prompt saves a model below the threshold - and this is the same 9B that compiles the source digest flawlessly every day. Right job vs wrong job, same weights. |
| Ensemble v2gemma26 edits itself · 5.5 min | 9 | 7 | 7 | 6 | 7 | 0 | 5.5 min | 4:30 | 7.5 | The fixed harness worked: zero false-positive critiques, every numbered item addressed and verified. Coverage grew 529 → 677 words. Self-edit blind spot confirmed - the editor missed its own confabulated filler. Best coverage-per-minute of the night. |
| Ensemble v3.1gemma26 self-edit + persona quota · 4 min | 8 | 7 | 7 | 7 | 6 | 1* | 4 min | 4:12 | 7.5 | Targeted fixes landed: the escape hatch fired, so a thin news category produced an honest "quiet day" instead of filler. New gremlin: the revision emitted a LaTeX artifact and dropped context the previous version had. *Caught by lint before render. Defects shuffle between runs; deterministic lint beats another prompt round. |
What the night taught
- Think for judgment, never for voice
Replicated across architectures. Thinking mode collapsed qwen's structure (7.0 → 4.0) and flattened gpt-oss's rhythm (7.5 → 6.5). The same thinking mode produced the night's best critique when the model was used as an editor instead of a writer. It is a reasoning tool, not a prose tool.
- Context is the cheapest upgrade
A 300-word persona block moved gpt-oss from 7.0 to 7.5 and measurably sharpened qwen122. No retraining, no bigger model, no extra latency.
- Completeness beats eloquence
The best-sounding local model ranked near the bottom because it dropped four stories. And parameter count predicted nothing: the 26B MoE beat both 31B builds.
- The author–editor ensemble works, but needs engineering
The editor caught real defects - a missing story, a confabulated analysis line, invented details - across 21K characters of deliberation. The revision then applied that critique imperfectly. Ensembles matched the best single models at 7.5; the path past that is editor-sees-everything plus checklist-enforced revision.
- The frontier gap is cohesion, not facts
Frontier 9.0, mid-tier frontier 8.0, best local 7.5. Local models get the facts. What they don't get is the through-line that makes five minutes of audio feel like one piece.
Prompt craft versus parameters
The last third of the night was spent not swapping models but rewriting the instructions: an explicit coverage contract, a gold exemplar, contrastive good/bad pairs, and the rules moved to the end of the prompt where they're least likely to be forgotten. That work lifted mid-size models by 1.0 to 1.5 points - more than any parameter jump I measured. A 26B model with a good prompt beat every 31B build. A 27B model with a good prompt matched the first run of a 122B one, at a tenth of the latency.
Two caveats keep this from being a free lunch. Prompt craft cannot cross the capability floor: the 9B model scored 3.0 no matter what it was told - and that same 9B compiles the source digest flawlessly every single day. Right job versus wrong job, identical weights. And exemplars leak: give a model a gold sample and it will occasionally quote the sample's content back to you as today's news, with total confidence.
Which is the whole argument for measuring in the first place. Every claim here would have been a plausible guess before the run, and three of them would have been wrong.