Two of us designed 476 states in about four months. 75 primary screens, 401 edge cases, heading toward a thousand.
That pace isn't a brag, it's the setup.
We work out of a knowledge vault with the edge cases baked in. We build in HTML against a component library that lives in the same vault. And we wrote agents that check the screens back against all of it: component parity, accessibility, edge case coverage, running and re-running because the timeline is tight.
It works. I'd defend all of it.
We also have one QA person, and I used to think his problem was volume.
What every one of those checks has in common
It took me a while to see it, and once you see it you can't stop.
Parity asks whether this button matches the library. Accessibility asks whether this screen meets the standard. Coverage asks whether this state exists in the vault.
This button. This screen. This state.
Every check we built is scoped to a single artifact. Not one of them can evaluate a path.
Where it actually breaks
Here's the shape of it in our product.
We have multiple order types on the retail site. They don't behave the same way. Each one carries different fulfillment options, and the fulfillment type changes how the cart itself is supposed to work. So the correct behavior of the cart is a function of a decision the customer made three touchpoints earlier, on a different screen, that the cart has to still know about.
An LLM handles any one of those screens beautifully. What it loses is the thread. Ask it to hold the whole end to end experience in view, remember which order type was selected, carry the fulfillment consequence forward, and check whether the cart is doing the right thing given all of it, and it starts making connections that don't hold.
Every screen in that flow can pass every check we have and the flow can still be wrong.
I know that because it happened. The states passed. Parity passed, accessibility passed, coverage passed. The package went to an external development team, they built the front end exactly as specified, and the first thing that noticed anything was wrong was a person clicking through the finished product.
Design didn't catch it. The agents didn't catch it. The build didn't catch it. QA caught it, which is to say it got caught at the most expensive point available.
The thing I had backwards
So the QA person isn't slow. He's doing a categorically different job than the agents are, and I'd been filing it under the same heading.
The agents check artifacts. He walks the path. He's the only part of our process that experiences the product as a sequence instead of as a set, and he's also the only one who knows the backend well enough to notice when a screen is technically fine but lying about what happens next.
Which means our fastest, most automated, most thoroughly checked pipeline currently routes its hardest class of error to the last human in the chain, after somebody has already been paid to build it.
That job didn't get automated, and it didn't get smaller when our output went up 6x. It got proportionally larger, because the number of paths through a product grows a lot faster than the number of screens in it.
We scaled the thing that was easy to check. The residual risk didn't shrink, it concentrated.
That's the part I keep chewing on. Paths through a product grow a lot faster than screens in it. Add one more order type and you haven't added one more thing to check, you've added it multiplied by every fulfillment option and every downstream screen that has to know about it. So when output goes up, automated coverage climbs in a straight line and the thing it can't see climbs much faster than that.
Which means the gap between what we check and what could break gets wider precisely when the team feels fastest.
If you have agents reviewing design work
Go look at what they're actually scoped to. My guess is every one of them takes a single artifact as input, because that's the version that's easy to build and easy to turn into a green check.
Then find whoever on your team still opens the product and uses it like a person. That's your real coverage. Everything else is confirming that the parts are correct.
Our automated checks are strongest where the domain is clean and weakest where it's tangled, which is the exact inverse of where the bugs have always been.
The bugs were never on the screens. They were in between them.