ServicesLabProcessAboutBlogFAQGet in Touch
HomeBlog

What a 4,400-Test Suite Actually Buys You

ARC UI 3.0 replaced manual review with strict automated gates. What that cost, what it caught, and where it still lets things through.

Somewhere around the middle of the ARC UI 3.0 cycle, the test suite went from roughly 250 tests to roughly 750. By the time I shipped, it was 4,431 passing. The number itself isn’t the point. It’s the result of a decision I made early in the 3.0 work: stop reviewing every component change by hand, and let automated gates decide what merges.

That decision needs unpacking, because “more tests” is not automatically “more quality.” I want to walk through what bought me confidence, what the number conceals, and where I still have to show up and look at things with my own eyes.

Why manual review stopped scaling

At 152 components, manual review was already strained before 3.0 started. Every prop change, every slot addition, every token rename had a blast radius that touched components I hadn’t looked at in months. A human reviewer (me, since I’m the only one) can hold maybe a dozen components’ worth of edge cases in working memory at once. I don’t have a dozen components. I have 152, across shadow DOM boundaries, plus six generated framework wrappers per component courtesy of Prism.

So the plan for 3.0 was explicit: convert as much of “does this still work” as possible from a judgment call into a checkable assertion. Roughly 1,600 of the final tests are hand-written. Those encode intent: I sat down and decided what correct behavior looks like for a specific interaction pattern (focus trapping in a dialog, roving tabindex in a listbox, ARIA state transitions in a combobox). The rest, the bulk of the 4,431, are derived: generated from schemas, prop tables, and the component contracts that Prism already has to parse to emit framework wrappers in the first place.

That’s what made it feasible. If I already have a machine-readable description of a component’s props, events, and slots, because Prism needs one to generate a React wrapper, I can walk that same description and generate tests for it: does every documented prop reflect to an attribute correctly, does every documented event fire with the right payload shape, does every slot accept and render content. None of that requires a human to write a new test file. It requires the schema to be honest.

What derived tests are good for

Derived tests are exhaustive in a way hand-written tests never are, because I’m lazy and hand-written tests only cover the cases I thought to worry about. A generated test suite doesn’t get bored. It checks all 152 components’ prop-to-attribute reflection the same way, every time, including for the components I haven’t personally touched since 2.x.

That matters most for regression across the framework wrappers. A Lit component and its Prism-generated Vue wrapper can drift in exactly the way you’d expect: an event name gets normalized in Lit but the Vue emit mapping doesn’t pick up the rename, or a Svelte 5 wrapper’s prop typing goes stale against a new required attribute. Derived tests running against every wrapper on every component change catch that class of bug immediately, at a volume no reviewer could sustain by hand.

Derived tests are also good at boring correctness: does the component still satisfy the WCAG 2.1 AA baseline I’ve committed to, does keyboard navigation still work after a refactor, does a token rename still resolve to the right computed value. This is exactly the kind of check that benefits from being mechanical. I don’t want a human (again, me) deciding by feel whether a contrast ratio passed.

Generated tests are only as trustworthy as the schema they’re derived from. If a component’s contract is wrong or incomplete, the tests built from it will confidently verify the wrong thing. The schema is the source of truth here, and the test count is just its shadow.

Where the number lies to you

To be straight about it: 4,431 passing tests tells you almost nothing about whether a component feels good to use. Derived tests check that a prop reflects correctly. They cannot tell you that the prop is named badly, that its default value surprises people, or that the API asks a consumer to hold two pieces of state that should have been one. API ergonomics is a judgment call, and judgment calls don’t derive from a schema. They come from someone using the component in anger and noticing friction. That’s still on me, by hand, every release.

Visual regression is the other hole. A shadow-DOM component can pass every functional and accessibility assertion in the suite and still render with a rounded corner where there should be a square one, or a spacing token that resolved correctly but looks wrong next to its neighbors. Automated visual diffing exists and I use some of it, but it’s noisy enough that I still do a manual pass over rendered output before a release, the same eyeball-driven review the rest of the suite was built to reduce.

Then there’s the maintenance tax nobody puts on the test-count slide. A 4,400-test suite is also 4,400 things that can go red when I change something unrelated, and diagnosing whether a failure is a regression or a stale assertion takes time. Generated tests amplify this: when the generator itself has a bug, it produces one bad test times 152. I’ve paid for that mistake more than once during 3.0, and the fix was always the same: treat the generator’s correctness as more load-bearing than any individual test file, because everything downstream inherits its errors.

4,431 tests don’t mean the library is correct. The claim is narrower: automated gates let me stop relitigating solved problems by hand, so the judgment I do have time for goes where a schema can’t reach.

Content may be edited with the help of AI. All content has been reviewed by a human before being published.

Work with Arclight

Have a project in mind?

I design and ship custom software — from early concept to production. Tell me what you're building and we'll figure it out together.

Start a ProjectView Services
NavigateServicesLabProcessAboutBlogFAQMoreFractionalColorado SpringsTech StackBrandGet in Touchhello@arclight.build+1 (719) 337-4490Colorado Springs, COStart a Project