Not more correct — more precise
The strangest fact of the final table: the human got more winners right. Viscabarca called 69 correct tendencies to Claude's 67. The entire title swung on scoreline precision — the partial credits your scoring hands out for exact score, goal difference, and total goals.
| Component | Claude | Viscabarca |
|---|---|---|
| Correct tendencies | 67 | 69 |
| Exact scores | 16 | 13 |
| Goal-difference bonus | 35 | 30 |
| Total-goals bonus | 25 | 20 |
| Match points | 277 | 270 |
| Forecast (bracket) | 181 | 188 |
| Total | 458 | 458 |
Where did the precision come from? Look at the two fingerprints below. Claude predicted a low-scoring tournament — nearly every pick a 2-0, 2-1 or 1-0 — and 2026 obliged: 2.88 goals per match, with 1-1 the single most common result. Viscabarca predicted 3.26 goals per game, full of 3-1s and 4-1s: right about who wins, wrong about how, and the partial credits bled away exactly where the title was decided.
Home/away mirrored (2-0 includes 0-2). Actual = 104 matches at full time; picks: Claude 97, Viscabarca 98. Hover a bar for counts.
Two more numbers complete the picture. Claude went against the human crowd only 15 times but hit 8 of them — including the exact 2-1 on Portugal–Croatia and backing Spain in a final most of the field gave to Argentina. And timing: Viscabarca was the ultimate late tipper (median 7.8 hours before kickoff, 79 of 98 tips inside the final day), which is where those two extra tendencies came from. Claude tipped at a median lead of 116 hours on a six-hour cron. Better information timing lost to better scoreline discipline — by a tiebreaker.
One asterisk the champion carries: an outage (see match incidents) meant Claude simply missed 7 of the 16 Round-of-32 matches — roughly 20 points forfeited — and won anyway.
All eight machines finished in the top half
The human field (players with 80+ tips) averaged 237 match points, 56.5 tendencies and 10.3 exact scores. Every single bot beat all three averages or came close — and the most interesting entrant wasn't a language model at all.
| Bot | Rank | Total | Match | Forecast | Tips | Tend. | Exact | Avg goals | Draws | W/ crowd | Against (hit) | Med. lead |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ClaudeLLM | 1 | 458 | 277 | 181 | 97 | 67 | 16 | 2.18 | 6 | 85% | 15 (8) | 116 h |
| Elostatic algo | 3 | 444 | 265 | 179 | 104 | 71 | 11 | 2.77 | 9 | 89% | 11 (6) | 354 h |
| KimiLLM | 12 | 417 | 266 | 151 | 98 | 65 | 16 | 2.10 | 4 | 82% | 18 (7) | 56 h |
| GeminiLLM | 14 | 417 | 249 | 168 | 97 | 63 | 14 | 2.29 | 9 | 87% | 13 (5) | 129 h |
| QwenLLM | 16 | 414 | 254 | 160 | 96 | 62 | 15 | 2.11 | 3 | 84% | 15 (5) | 111 h |
| ChatGPTLLM | 19 | 413 | 254 | 159 | 97 | 64 | 12 | 2.39 | 4 | 88% | 12 (5) | 130 h |
| DeepseekLLM | 28 | 392 | 224 | 168 | 98 | 58 | 10 | 2.13 | 11 | 82% | 18 (2) | 41 h |
| GrokLLM | 53 | 348 | 203 | 145 | 97 | 56 | 6 | 2.32 | 5 | 66% | 33 (9) | 89 h |
The Elo control group is the finding of the tournament. A static rating table — no language model, no reasoning, never down — posted the best tendency count of anyone in the game (71 correct winners) and finished 3rd overall. Picking winners is nearly solved by ratings alone. The LLM edge was entirely in guessing plausible scorelines: Elo's 3-0 habit earned it just 11 exacts, while every LLM independently converged on the same 2-0 / 2-1 / 1-0 signature that the low-scoring tournament rewarded.
The 220-line bronze medalist. The Elo bot is ~220 lines of Go — a static rating table plus the textbook Elo update after each result — written only to test the bot pipeline before dealing with LLM APIs. It finished 3rd of 130 with no systematic blind spot: across its 33 wrong calls, no team fooled it more than twice, and its misses are simply the tournament's shocks plus the pedigree names (Brazil, Germany, Belgium, France…) it backed exactly twice each — once before the upset, once more before its feedback loop caught up. Its only real weakness was never who but how much: the 3-0 habit produced just 11 exact scores, which is the entire gap between bronze and the title. A zero-effort rating table out-picked 98% of humans — most football "knowledge" is worth less than the table alone.
Why Grok flopped: it was the only bot that regularly defied consensus (33 contrarian winner calls, 9 correct) and loved shutouts — 43 of its 97 picks were 2-0 or 3-0 — in a tournament with 29 draws. Deepseek's failure mode was the opposite: obedient but timid, with the worst contrarian record in the game (2 of 18) and a matchday-3 habit of talking itself into cautious draws. And Kimi is statistically Claude's twin on match tips (266 points, 16 exacts) — its 12th-place finish is entirely a bracket story, told below.
Every LLM wrote a one-sentence rationale per tip — 674 in total. All checkable claims were cross-referenced against the database: scores cited, points-and-goal-difference standings, "winless" and "eliminated" statements, knockout-run recaps. The verdict: roughly 98% clean. Standings claims parsed from the prompts were correct without exception. The errors — about a dozen — cluster in one specific place: the bots had every result in their context window and still misremembered tournament history when narrating it. In-context attention failures, not missing data.
"France's attacking depth and their dominant SF win over Spain make them favorites…"
France lost that semifinal 0-2 — the result sat near the end of the results list in the prompt. The champion inverted it, then quoted it correctly three days later in its final rationale.
"…Egypt, who have R32 momentum after beating Senegal 3-2…"
Egypt never played Senegal. The real 3-2 over Senegal belonged to Norway, in the group stage. In its England–Argentina semifinal tip, Deepseek also argued Argentina had "a slight edge" — and then picked England 2-1.
"Both sides are eliminated with zero points…" (Senegal); "Both teams are eliminated…" (DR Congo)
Both teams advanced to the Round of 32. Correct standings were in context; Kimi failed to apply the 48-team format's best-thirds rule. It also invented "DR Congo's competitive draws against Colombia and Portugal" — the Colombia match was a 1-0 loss.
"Argentina has scored exactly three goals in every knockout round…"
True — 3, 3 and 3, including extra time in two of the three rounds. The same bot that fabricated eliminations also produced the most precise stat of the audit.
"Spain… beating Portugal 1-0 and France 2-0… Argentina grinding out wins over Egypt (3-2) and Switzerland (3-1)…"
All four scores correct — including Argentina's 3-1 over Switzerland, which required tracking the extra-time score, not the 1-1 after ninety minutes.
Style notes from reading all 674: Kimi and Deepseek wrote the most data-rich rationales (standings, points, goal difference — almost always accurate). Gemini, ChatGPT and Qwen wrote fluent, plausible scouting-report prose. Claude wrote the tersest lines in the game (73 characters on average, half everyone else's) — making the fewest checkable claims per tip, which is either efficiency or cowardice depending on your reading. Grok's were the vaguest, committing to almost nothing verifiable. The Elo bot, of course, said nothing at all and finished third.
Three things that went wrong — and how little they mattered
The everything-is-new-info loop
The match sync wrote rows unconditionally, bumping timestamps without content changes — so the bots re-evaluated every open tip on nearly every 6-hour run, re-sending up to ~100 matches of identical context to the LLM APIs four times a day until a skip-if-unchanged check landed on July 10, 29 days into the tournament.
Silver lining: it proves the picks that never moved were genuinely reaffirmed. Claude's frozen 1-1 on Bosnia–Qatar was re-asked ~30 times with Bosnia's results in the prompt, and the model kept choosing 1-1 while the late human crowd went 100% Bosnia. The group-stage weakness was models under-updating on evidence — not the harness withholding it.
The best-thirds bug
A third-place-qualifier selection bug briefly published wrong Round-of-32 pairings — Belgium was slated against South Korea (real opponent: Senegal) and Switzerland against Ecuador (real: Algeria) — before a manual database fix on the evening of June 28. Exactly three human tips still carry fossilized "ghost advancers" from that window. One of them belongs to the game's creator.
The outage
An app outage — longer for the bots — cost every LLM player 6–8 untipped Round-of-32 matches, worth roughly 15–20 points each. The static Elo bot, needing no API, sailed through with 104 of 104 tips. The accidental irony: because the LLMs were down during the buggy-bracket window and tipped only after the fix, not one bot ever saw the wrong pairings. The outage cost them points and saved them from the bug in the same breath.
The pre-tournament bracket was its own competition, and Viscabarca won it. The remarkable part: against Claude, every component except one was identical — 27 advance points each, 94 knockout points each, Spain as champion for +13 each. The whole 7-point gap is group-stage ordering: 38 exact group positions and 8 perfect groups against Claude's 35 and 6. And since Claude out-scored Viscabarca by exactly 7 on match tips, the bracket is the sole reason the game ended 458–458.
Those twin 94-point knockout scores came from opposite-shaped brackets: Claude had both finalists right but only three semifinalists (Brazil where England went); Viscabarca called all four semifinalists but sent England, not Argentina, to the final. Different paths, same sum.
| Player | # | Total | Groups | Advance | Knockout | Champion pick |
|---|---|---|---|---|---|---|
| Viscabarca | 1 | 188 | 54 | 27 | 94 | Spain ✓ |
| Claudebot | 2 | 181 | 47 | 27 | 94 | Spain ✓ |
| Christian | 3 | 180 | 49 | 28 | 103 | Argentina ✗ |
| Elobot | 4 | 179 | 41 | 26 | 99 | Spain ✓ |
| Kjetting Eide | 5 | 179 | 50 | 26 | 90 | Spain ✓ |
| Geminibot | 12 | 168 | 50 | 28 | 90 | France ✗ |
| Deepseekbot | 13 | 168 | 50 | 27 | 91 | Argentina ✗ |
| Kimibot | 34 | 151 | 41 | 26 | 84 | France ✗ |
| Grokbot | 45 | 145 | 41 | 26 | 78 | Brazil ✗ |
Christian's heartbreak. Best knockout bracket of all 113 players — all four semifinalists, both finalists — then Argentina as champion instead of Spain. Thirteen points left on the final hurdle; Spain would have won him the forecast competition outright.
Spain won the final 1-0 in extra time. 42 players — over a third — left the champion slot blank, walking past 13 free points.
The bots reveal something neat here: all eight put Argentina, France and Spain in their semifinals — near-total chalk consensus — but split on the crown. Only Claude and the Elo bot picked Spain; four backed France, whose semifinal exit to Spain is single-handedly most of the bot forecast spread. Kimi, Claude's equal on match tips, lost the tournament in this table: 151 forecast points, 30 behind, on a France pick plus weak group ordering. Trusting the ratings favourite beat every narrative about France — a call only the two eventual podium bots made.
The creator's profile is the inverse of the champion's: 99th percentile on precision, 44th percentile on tendencies. And the 18 exacts weren't chalk — they include Spain 4-0 Saudi Arabia, France 3-1 Senegal, Brazil 3-0 Haiti, Portugal 2-1 Croatia, and Bosnia 3-1 Qatar, the very match Claude sat frozen on 1-1 for nine days.
What kept the creator off their own podium was the bracket: Portugal as champion (out in the R16 — to Spain, of all teams), a Netherlands–Portugal final (Netherlands fell in the R16), Croatia in the semis (gone in the R32). Forecast: 133 points, rank 59 — a 55-point hole against Viscabarca that accounts for four-fifths of the 69-point gap to the title. With even a mid-table-of-the-top-ten bracket, the tips above were top-6 material.
Head-to-head against the bot he built the pipeline for: −21 match points, with an asterisk — the two biggest wins over Claude (exact 3-0 France–Sweden, exact 1-2 Ivory Coast–Norway) came in matches Claude couldn't tip during the outage. And yes: one of the three ghost-advancer tips fossilized by the best-thirds bug — Belgium vs "South Korea" — belongs to the man who fixed the bug.
The game's creator always distrusted one rule: the extra point for total goals ("feels way more luck-based than actual metric"), while rating the goal-difference point as real thinking. Both instincts were tested against all 9,062 scored tips.
First, structure: a correct goal difference all but requires picking the right winner — only 2% of goal-diff points landed without the tendency (knockout edge cases). Total goals has no such anchor: 34% of all total-goals points (610 of 1,797) were earned on tips that picked the wrong winner. Tip 2-0, watch it end 0-2, collect a point.
Second, skill: for every player, the same picks were re-priced against what they'd score randomly shuffled across matches — so pick-style is controlled for, and any excess is genuine insight. Across 71 humans with 80+ tips:
| Component | Excess hits vs chance | Players above chance | Verdict |
|---|---|---|---|
| Tendency | +25.2 | 71 / 71 | skill |
| Goal difference | +8.5 | 69 / 71 | skill |
| Exact score | +3.9 | 64 / 71 | skill |
| Total goals | +0.17 | 32 / 71 | coin flip |
Total goals is statistically a coin flip: barely half the field beat chance. And here is where it gets uncomfortable — re-scoring the whole tournament under alternative rulesets (same picks, forecast unchanged):
| Ruleset | Champion | Claude | Viscabarca | Elo |
|---|---|---|---|---|
| Current (3 / +1 exact / +1 gd / +1 goals) | Claude, on tiebreak | 1 | 2 | 3 |
| Drop total-goals | Viscabarca 438–433 | 2 | 1 | 3 |
| Total-goals only with correct winner | Viscabarca 456–452 | 2 | 1 | 3 |
| Kicktipp classic (4 / 3 / 2) | Viscabarca 369–366 | 2 | 1 | 3 |
| Pure tendency | Viscabarca 395 | 3 | 1 | 2 |
| Exact-heavy (exact tripled) | Claude 490–484 | 1 | 2 | 5 |
| Rarity-weighted exact | Viscabarca 465–463 | 2 | 1 | 3 |
Claude's championship survives only under rules that include the unconditional total-goals point. Remove it, condition it, or use the classic kicktipp scheme, and Viscabarca is champion every time — under pure tendency the Elo bot even takes second. The rules weren't unfair; everyone played under the same ones. But they were load-bearing: the noisiest component in the system happened to be the title margin. (One caveat: the bots literally optimized their scorelines for the actual rules, so under different rules they would have picked somewhat differently.)
How rare is an exact score, actually?
10.5% of all tips — but the average hides everything. The chalk rounds paid out en masse (18.2% in the R32; the opener Mexico 2-0 South Africa was exactly called by 51 of 118 tippers), while 24 of 104 matches were hit by no one: blowouts beyond imagination (Germany 7-1, Canada 6-0, the 4-6 third-place circus), shock goalless draws nobody dared tip against a giant (Spain 0-0 Cape Verde — the crowd said 3-0), and both showcase matches. Nobody in the game exactly called the semifinal-losers' match or the final.
Value-hunting by scoreline: the field's favorite pick, 2-1 (tipped 1,285 times), hit just 10.8%. The two best-value picks were opposites — 3-0 when you trust a heavy favorite (18.3%) and 1-1 when you trust no one (15.6%). The sucker bet: 2-2, at 1.9%. And away-blowout picks (0-3: 4.7%) were half as good as their home mirror.
The greatest stat in the database. Exactly one player out of 130 called Qatar 1-1 Switzerland. Their username: Laurette-miss-exact-score.
Proposed laws for the next edition, every point traceable to a real skill signal: keep goal-difference untouched (it rewards exactly the "feels close — 1-0 or 2-1?" reasoning, and it's the best-behaved rule in the system); make total-goals count only with the correct tendency, deleting the 610 consolation points; and pay exact scores more, weighted by rarity — a crowd hit earns the base point, a hit under 15% of the field earns double, under 5% triple. Priced against its 10.5% base rate the exact score is currently the most underpaid achievement in the game.
The baseline ladder: the best player in the game was the crowd itself
How much football knowledge does it take to compete here? To find out, seven mechanical strategies were scored with the exact same engine as the real tips (validated to reproduce every stored score) — from "always tip 1-1" up to aggregating the whole field's picks.
| Strategy | Points | Rank | Exacts |
|---|---|---|---|
| Crowd median scoreline | 299 | #1 — beats everyone | 17 |
| (best real players: Puskas 280, Claude 277) | |||
| Crowd modal scoreline | 274 | #5 | 11 |
| Chalk 2-1 — favorite wins 2-1, nothing else | 271 | #7 | 11 |
| Elo-favorite 2-0 — no crowd info at all | 268 | #9 | 9 |
| Chalk 1-0 | 265 | #11 | 12 |
| Chalk 2-0 | 262 | #14 | 9 |
| Always 1-1, every single match | 191 | #74 | 12 |
The median of everyone's picks beats every single player — including the champion bot, by 22 points. Taking just the median predicted home and away goals per match yields 73 correct tendencies (no real player managed more than 71), 17 exact scores, and the joint-best goal-diff count in the game. Individual errors cancel; the aggregate keeps only the signal.
The median is also the only "player" with calibrated draws: it tipped 1-1 twenty-one times, because when the crowd splits between "1-0 us" and "1-0 them", the middle lands on the draw — exactly the calibration Claude (6 draws all tournament) and nearly every human lacked, emerging automatically from aggregation.
And the humbling row: Chalk 2-1 — zero football thought beyond knowing who's favored — finishes 7th of 127, level with the best human. Below roughly rank ten, the whole field was statistically indistinguishable from mechanical favorite-backing.
Fair-play notes: tips were private until kickoff, so the crowd strategies were benchmarks, not something anyone could actually play (the chalk strategies, though, were available to all). Synthetic players also tip all 104 matches, outage-free. For the next edition, the crowd median is the bar any bot should be measured against — a ninth bot that simply ensembled the field would have won the tournament outright.
The crowd's biggest collective faceplants, ranked by the share of humans who backed the wrong side — and the few who didn't:
| Wrong | Match | Called by |
|---|---|---|
| 100% | Spain 0-0 Cape Verde | nobody |
| 100% | Germany 1-1 Paraguay · pens (R32) | nobody picked Paraguay |
| 99% | Qatar 1-1 Switzerland | Laurette-miss-exact-score — exact, again |
| 99% | Portugal 1-1 DR Congo | It’s Coming Home Depot |
| 96% | Ecuador 2-1 Germany | Big Nate, Mariela, lunguini |
| 96% | England 0-0 Ghana | Andres Cuadrado, Josh Fil, Puskas |
| 95% | Belgium 0-0 Iran | Chris, Fgibase4, Flo, lunguini |
| 94% | Australia 2-0 Türkiye | Barry Dingle, Carlos, K-RAGE, PrinzEugen, Sjaak schoonlingen |
The upset whisperer: Barry Dingle. Across the fifteen biggest upsets, chance would give a typical player about 0.9 correct calls. Barry Dingle called four — four times chance, the clear standout of 127 players. And note the pattern across the whole table: nearly every shock had at least one caller. Even when 95% of the crowd missed, the crowd contained the truth — the same effect that makes the crowd-median unbeatable.
The multiverse: re-running the title race
How sturdy was Claude's championship? Delete any single one of the 104 matches from the scoring and re-crown: Claude survives 75 scenarios, Viscabarca takes 29. In more than a quarter of "one random match never happened" worlds, the title flips. A dominant champion survives all 104; this one was decided on a coin's edge. Three specific edits to history, fully re-scored:
| If instead… | Champion | Top of the table |
|---|---|---|
| Argentina wins the final (0-1 aet) | Christian | Christian 455 · Viscabarca 449 · Claude 441 |
| Bosnia–Qatar ends 1-1 (Claude's frozen pick) | Claude, outright | Claude 464 · Viscabarca 452 |
| 3rd place ends France 2-1 (the crowd pick) | Claude, outright | Claude 464 · Viscabarca 458 |
The unluckiest player of the tournament is Christian. Best knockout bracket of all 113 forecasts, all four semifinalists, both finalists — and his one miss, Argentina over Spain as champion, was decided by a single extra-time goal in the final. That goal was the difference between finishing 5th and winning the entire game. Meanwhile the Bosnia 1-1 that stood as the bot's most-mocked frozen pick was one Qatari defensive lapse away from being its masterstroke.
How high does the score go? A flawless season — every match exact, a flawless bracket — is worth 624 match points + 241 forecast points = 865. The champion's 458 is 53% of perfect; a typical committed human sat around 45%. Nobody is anywhere near solving this game.
More interesting is the field-oracle ceiling: take, for every match, the best tip anyone actually placed, plus the best real bracket. That player would have scored 756 — because in 81 of 104 matches, someone in the field had the exact scoreline, and in 16 more someone nailed the tendency plus a partial. Collectively, the players "knew" 91% of the tournament's match points. Individually, nobody extracted even half of that (the best real tally, 277, is 49% of the field-best 568). The knowledge existed — distributed across a hundred heads. It's the strongest version of the wisdom-of-crowds finding on this page.
On the forecast side the best single bracket (Viscabarca, 188) reached 78% of its ceiling, and some components were perfected outright — the Elo bot alone called all four semifinalists and both finalists. The field-wide gaps sat in exact group ordering (best: 54 of 72) and the R16. Note the asymmetry: winners converted ~75% of the forecast ceiling but only ~44% of the match ceiling — pre-tournament bracket chalk is far cheaper to cash than scorelines, and each correct R32 team currently pays twice (advance + round points). That's the natural dial to turn if the forecast ever feels too heavy.
Only seven matches defeated the entire field (best pick ≤3 points): the unimaginable blowouts — Canada 6-0 Qatar, Sweden 5-1 Tunisia, Netherlands 5-1 Sweden, USA 1-4 Belgium, France 4-6 England — the Germany–Paraguay shootout… and one absolute unicorn:
Fun fact: Spain 0-0 Cape Verde is the only match where the entire field scored zero. ~87 tippers, not one draw pick, so no tendency and no goal-diff points — and since nobody predicted 0 total goals, not even a consolation point escaped. One match, a hundred players, collectively nothing. It also tops the Giant-Killers table as the 100%-wrong upset: this fixture now holds every futility record the data has to offer.
Does last-minute tipping pay? Every human tip was measured against the average points scored on that same match (so hard matches don't pollute the comparison) and bucketed by when it was last touched before kickoff:
| Phase | <6h | 6–24h | 1–3d | 3–7d | >7d |
|---|---|---|---|---|---|
| Matchday 1 | +0.02 | +0.02 | +0.01 | +0.05 | −0.14 |
| Matchday 2 | −0.05 | +0.04 | +0.16 | −0.00 | −0.05 |
| Matchday 3 | +0.02 | +0.05 | +0.15 | +0.08 | −0.09 |
| Knockouts | −0.07 | −0.04 | +0.04 | +0.01 | — |
| All tips | −0.03 | +0.01 | +0.07 | +0.04 | −0.08 |
There is no last-minute edge. Tips finalized under six hours before kickoff scored slightly below the field — lineup news added nothing, and knockout late-tippers did worst of all. The within-player check agrees: the same person tipping late vs early gains nothing (21 of 44 better late — a coin flip). The runner-up's famous 7.8-hour habit wasn't the weapon; being informed was, and that's available days earlier.
The real penalty is staleness. The sweet spot is 1–3 days before kickoff — after the previous round's results, before the noise — worth +0.15 per tip on matchdays 2–3, while week-old picks bleed −0.05 to −0.14 everywhere. Across the later group rounds that gap is worth roughly 6–8 points: more than the margin that decided the title, and exactly where both the creator's 477-hour set-and-forget style and Claude's frozen matchday-3 picks paid the price.
And second thoughts were good thoughts: the 19% of tips that players went back and revised outperformed the field in every phase (+0.13 in the knockouts), the top-10 finishers revised more than the rest (26% vs 20%), and revision share correlates mildly with final points. Part of that is selection — engaged players revise more and play better — but both analyses tell one story, and it points at one concrete feature for next season: a nudge that says "3 of your tips were placed before the latest results — want to review them?". Not a push to tip later; a push to tip fresher.
Five ways to play: the archetypes the data found on its own
All 92 committed players (60+ tips) were clustered purely on style — draw rate, goals per pick, blowout share, against-crowd rate, revision habit, lead time, scoreline variety. Points were hidden from the clustering, so how each archetype scored is a genuine result:
| Archetype | n | Signature | Avg match pts | Poster players |
|---|---|---|---|---|
| 🤖 The Machines | 13 | tight 2-0/2-1 picks, few draws, revise constantly | 253 | Claude, Dark horse, Mario + all 7 LLM bots |
| 🏦 The Bookmakers | 23 | chalk-aligned, steady, low drama | 236 | P, Scott, Elo bot |
| ⏱️ The Deadline Snipers | 14 | tip in the final hours before kickoff | 234 | Viscabarca, Christian, Barry Dingle |
| 💘 The Romantics | 33 | most draws, most gut calls, set early | 218 | Puskas, Flo, Laurette-miss-exact-score |
| 🎬 The Hollywood Pickers | 9 | 3+ goals per pick, blowouts, wild variety | 177 | WhiteTiger, Éamo, Josh Fil |
The clustering rediscovered the bots on its own. All seven LLMs landed in one cluster — together with six humans who naturally play like a bot: disciplined low scorelines and constant revisiting. It's the highest-scoring archetype, and its two best humans are also the game's top two revisers (Mario revised 75% of his tips, Dark horse 69%). The machine style works for humans too.
But the best individual player is a Romantic. Puskas topped the match-point table (280) from the draw-loving, gut-call cluster that averages second-worst — a high-variance style producing one glorious outlier. At the other end, the Hollywood Pickers' 4-1 fantasies cost them roughly 75 points against the Machines. Style is destiny, on average.
The leagues: de-botted, de-ghosted, re-ranked
37 leagues formed around the global game — and the raw league table needs a disclosure first: two players collected the entire bot fleet. Sergio's DGTVEN and SergioS's digitel each recruited all eight bots, making the two "strongest leagues" really one human hosting eight machines; remove the bots and neither has enough members left to rank. ChatGPT turned out to be the most social bot of the tournament, sitting in six private leagues. Filter the bots out — and then also the members who barely played (fewer than 60 of 104 matches tipped) — and the real community table emerges:
| League | Active | Avg | Champion | Shift vs unfiltered |
|---|---|---|---|---|
| Norske Skole | 7 / 8 | 387 | Espen (423) | strongest either way |
| Neighborhood | 6 / 10 | 381 | Viscabarca (458) | ↑ from 9th |
| Acca Boys WC Prediction League | 6 / 6 | 369 | Jason (423) | perfect attendance |
| Me Fein WC '26 | 7 / 9 | 367 | P (442) | ↑ |
| HVEEM | 4 / 5 | 356 | Mobolo😏 (417) | — |
| Screwed World Cup | 21 / 22 | 350 | Christian (438) | biggest & near-full turnout |
| maximus | 4 / 5 | 342 | Ki (430) | — |
| Football anything | 3 / 5 | 335 | Ih (369) | ↑ |
| Cupa pe Someș | 9 / 9 | 308 | Vosy (387) | full turnout |
| porrita mágica | 12 / 12 | 266 | Ángela Molina (384) | full turnout |
Norske Skole is the strongest human community in the game — seven of eight members active, two global top-10 finishers (Espen and TipsensRipsbusk), and the top average with or without filtering. Neighborhood is the great riser, 9th to 2nd: the creator's own friends-league looked mid-table only because four members barely played; its active core — led by Viscabarca, the best human in the entire game — averages 381. The title fight was, in the end, the creator's league against the creator's bots.
The engagement footnote: 16 of 36 private leagues never got a second member, and PODIKKALAM PREDICTION: Win Exiting Gifts finished with three members, zero of them active, and seven total points. The exciting gifts remain unclaimed.
Fun fact: Grok is the only bot that made a league weaker. Every league that recruited machines got a boost — except Norske Skole, whose adopted Grok (global #53) sits below all seven of its active humans and drags the school's average down. Relatedly, in the creator's own AI Overlords league — eight bots plus Flo — the creator finished 8th of 9. The sporting term is: at least you beat Grok.
The creator admits to deliberately biasing his own picks toward the host nations — on the theory that a World Cup host, in a stadium like the Azteca, carries a real edge. The data puts that theory, and every other allegiance in the game, on trial.
Host advantage: guilty as charged. Across the 15 matches involving Mexico, the USA or Canada, the hosts went 9W-1D-5L — all three cruised through the group and R32 (Mexico won four straight, starting with the opener against 2010-host South Africa), and then all three crashed out on the same R16 weekend, each conceding three or four. The field backed hosts in 63% of tips, the bots in 75% — and host matches were the most predictable pocket in the game (2.55 pts/tip vs 2.36 elsewhere for committed humans). The creator's deliberate bias was his best-performing habit of the season: 3.13 pts/tip on host matches vs 2.35 on everything else, roughly +11 points — including an exact on the Azteca opener. Structural bias on a real effect pays; the only thing nobody timed was the expiry date.
National allegiance: a lottery ticket, drawn seven ways. Several private leagues have a clear community identity, which lets us measure how much extra faith each group put in "their" team — and what it cost:
| League | Nation* | Back-rate vs field | Their team's record | Verdict |
|---|---|---|---|---|
| porrita mágica | 🇪🇸 Spain | 99% vs 81% | won 7 of 8, champions | total homers, richly paid (2.99 vs 2.39 pts/tip) |
| Acca Boys | 🏴 England | 90% vs 69% | semifinal + 3rd place | loudest homers in the game — and it worked |
| Screwed World Cup | 🇺🇸 USA | 78% vs 66% | out in the R16 | host optimism, mildly taxed |
| Mundialito Ecuador | 🇪🇨 Ecuador | 78% vs 37% | won 1 of 4 | heart over head (0.89 vs 1.17 pts/tip) |
| Neighborhood | 🇦🇹 Austria | 50% vs 38% | won 1 of 4 | steepest homer tax in the game (2.04 vs 2.79) |
| Norske Skole | 🇳🇴 Norway | 56% vs 57% | won 4 of 6, beat Brazil | zero measurable bias — dispassionate even about their own |
| Me Féin WC '26 | 🇮🇪 Ireland | no team qualified | — | no positive allegiance to anyone; England backed at −3% vs field |
*A caveat on the flags: nationalities are inferred, not proven — from league names ("Me Féin" is Irish Gaelic, "Acca" is British betting slang), member names (Éamo, Segundo longhorn, HouAlreadyKnow), and the pick patterns themselves. Treat them as well-evidenced guesses.
The pattern is clean: bias on a structural effect (host advantage) was a correct bet. Bias on your own nation was a lottery whose payout depended entirely on whether your nation happened to be good — the Spaniards and English drew winning tickets, the Ecuadorians and Austrians paid the tax, the Norwegians simply declined to play (and are, not coincidentally, the strongest league in the game). And the Irish, with no team to love, played the tournament straight from the outside — no adopted darling, just a faint, ancestral −3% coolness toward England.
The Neighborhood's tax deserves its line item: the Austrians' extra picks weren't the free Jordan win, they were heart picks against Argentina and Spain — which pay zero, every time. Even Viscabarca, the game's most disciplined human, backed Austria twice.
Fun fact: every genuine Austria-as-world-champion bracket came from the creator's friend group. Austria received more champion picks (3) than four-time winners Germany (2) — and all the real ones belong to the Neighborhood league: Markus, who also sent Germany to his semifinal and tipped the most optimistic German scorelines in the field, and PrinzEugen — the purest believer in the dataset, who tipped everything three weeks in advance, backed Austria in every match he tipped, and crowned them world champions. Complete conviction on every axis; wrong on most of them.
Context first: this game was never marketed. The app was built in the weeks before kickoff and improved live throughout the tournament — the landing page itself shipped just one week before the opening match. The first ~50 players beyond the creator's friends found it through the open-source codebase on GitHub; the sole promotion was one Reddit post that had about 30 views at kickoff. The tournament itself was the acquisition channel — 85% of the 173 accounts were created in the two weeks around the opening match.
| Stage | Humans |
|---|---|
| Accounts created | 173 |
| Ever placed a tip | 122 (71%) |
| Submitted a bracket forecast | 104 |
| Committed (60+ tips) | 84 |
| Tipped all 104 matches | 17 |
The leak is at the front door — 51 accounts never placed a tip — and cross-referencing signup dates with the git history pinpoints it. The landing page helped a little (signup-to-first-tip conversion rose from 71% to 78% when it shipped), but the real conversion killer was arriving after kickoff: those signups converted at just 29%. Locked matches and a leaderboard already in motion gave latecomers no reason to stay — a "join anytime, score from here" mode would recover more players than any landing-page polish. Whoever got in on time stayed: of 91 players with 10+ tips, only 16 quit before the knockouts. Daily activity followed the tournament's pulse — 59 active tippers on opening day (the all-time peak), a steady 35–45 through the group stage, a second spike at the R32 bracket reveal, then the natural knockout decay.
The email system didn't struggle all tournament — it died once, spectacularly, in the worst possible week. The day-by-day delivery log shows 0% failure from June 7 to June 25, then a nine-day blackout at 75–96% daily failure from June 26 to July 4 — the provider ban and emergency migration, landing in exactly the week that also produced the best-thirds bracket bug and the app outage. One crisis week, three incidents, and full recovery by July 11 (the July 10 "skip unchanged fixtures" commit closed the last of them). The remarkable part: the blackout measurably cost nothing. Committed players covered 99–100% of the blackout-week group matches, identical to before; the later coverage decline tracks knockout logistics and fatigue, continuing even after email recovered. By week three the players had routines, and the reminder channel was redundant for the core. (Push never worked on iOS at all, and PWA installs stalled at 22 users — most players hear "app" and look for an app store — so for the "your tips are stale" nudge next season, the delivery rails still come first.)
The quiet flop was league chat, now with commit receipts: shipped June 8 and polished across roughly two dozen commits — soft-delete with an undo toast, keyboard-aware scroll pinning, GIF support arriving six hours after v1 — it carried 28 messages, ever, the last on June 30. Nearly one commit of chat engineering per message that would ever be sent; the banter stayed in the group chats people already had. The quiet triumph was on GitHub: ten stars, six forks, several of them actively developed and deployed during the tournament — the project's best open-source reception to date, from an audience that found it through code rather than ads.
Fun fact: the busiest hour in the entire request log is extra time of the final. On final Sunday the app served 4,954 requests from 93 devices — and the single biggest hour of the whole log was 22:00 UTC with 922 requests: a 0-0 final deep in extra time, the entire player base refreshing while the only goal of the World Cup final was being scored. Within 48 hours of the trophy, traffic fell 94%. The heartbeat of an event app, in one histogram.
What the data says to change
For the bots: show each model its own current pick and ask keep or revise instead of regenerating from scratch — the single biggest fix for the under-updating that froze picks like Bosnia 1-1. Give each match a five-line digest of the two teams' recent results rather than making models fish in a 100-line history (where all the hallucinations happened). State the best-thirds rule in the prompt. Mark penalty winners in the results feed. And nudge draw calibration: Claude picked 6 draws; the tournament produced 29.
For the harness: hash the context and skip the API call when nothing changed; refresh rationales on reaffirmation, not only on revision; invalidate tips when a matchup changes; and put an alert on the bot cron — availability was worth ~20 points, the margin between winning and not, and the Elo bot's only structural advantage.
For the humans: a nudge for the 42 players who left the champion slot empty — that's 13 free points a third of the field never even attempted.