I Made a Design System Readable by AI. Then I Tried to Catch It Lying.
A few weeks ago I read an argument that AI coding assistants break design systems in two specific, predictable ways. I decided to actually test it against a real sixty-component library instead of just agreeing with it.
The claim, from an article on building LLM-readable design systems, was this: a plain-CSS design system has no compiler to enforce “use the token, not the raw value.” That rule lives in code review and tribal knowledge. An AI assistant has neither. Left alone, it does two things wrong. It invents a plausible-looking hex color or padding value instead of checking whether a token already covers it. And it starts every session from zero, with no memory of the component it edited yesterday.

The proposed fix wasn’t a better prompt. It was structural: give the assistant a spec to read, close off every visual value to a defined set of tokens, and run an automated check that catches anything that slips through. I run Aural UI, a plain-CSS component library with around sixty components and nine themes, and I had the perfect, slightly messy test subject sitting right there.
What I actually built
Four pieces, all living inside the repository itself, not in a prompt or a wiki page:
- A rules file at the repo root. The first thing an AI session is supposed to read before touching a component: check the spec, never hardcode a value, add a token if one is missing, run the audit before committing.
- One spec per component. Fifty-four of them, each with the same eight sections: when to use it, its anatomy, the exact tokens it touches, its states, a working code example.
- A closed token layer. Two entire categories, a z-index scale and a sizing scale for icons and controls, had simply never existed. Every component that needed one had been hardcoding a number instead.
- An audit script. It parses every component’s CSS, flags any raw color or pixel value, and suggests the exact token that should replace it. It runs on every commit and blocks the build if it finds one.
Migrating the whole library against this closed system took the hardcoded-value count from 1,044 to zero. Every single substitution was checked to resolve to the exact pixel or color it replaced, so nothing about how any component actually looks changed.
1,044 → 0 hardcoded CSS values, across all 59 components. 54 component specs, eight sections each.
That part went the way I expected. The next part didn’t.
Writing rules is cheap. I wanted proof they get followed.
A spec file nobody reads is just documentation. The real question was whether an agent that has never seen this repository before would find the rules on its own, with no pointer, no hint, nothing in the prompt suggesting a system existed.
So I ran the same experiment three times. Each time, a fresh AI agent with zero memory of anything I’d built got a one-line ticket: add an xl size variant to a component. First Badge. Then Chips. Then Tooltip. Nothing else.
Each time, the agent, on its own:
- Found the rules file and the component’s spec without being told either existed.
- Picked the new size by extending the component’s real token progression, the same step between sm, default, and lg, instead of guessing a number that looked about right.
- Ran the audit script itself, unprompted, before calling the change finished.
I didn’t take its word for any of it. I checked each diff by hand against the actual token definitions and reran the audit myself afterward. All three held up.
What the migration actually found
This is the part I didn’t expect going in. Closing the token layer across fifty-nine files didn’t just tidy up existing values. It forced a line-by-line read of CSS that had gone untouched for a long time, and that turned up real bugs that a hardcoded value had been quietly standing in for.
A typo nobody caught: several components referenced --color-primary-alpha-20 instead of the real token, which has no color- prefix. Because the reference never resolved, the hover tint it was supposed to apply had been doing nothing, with no error and no visual sign anything was wrong.
An import-order bug: an older Toggle component and its newer Switch replacement both define the same classes, kept for backward compatibility. Because the older file loaded first, the newer file’s compatibility block was quietly winning the cascade and overriding Toggle’s own colors. Reading either file on its own, there was nothing to catch.
A public function that had never worked: the file upload component’s JavaScript queried class names the component’s actual markup had never used. Every call to it, in every demo, had always done nothing. It surfaced because the audit forced me into that file, and once I was there, checking whether the JS actually matched the CSS took one extra minute.
None of these turned up from looking for bugs. They turned up because closing the token layer meant actually reading every line, carefully, in a codebase that had been running fine, by appearances, for months.
Merging it broke something else entirely
I pushed the branch, opened a pull request, and the CI failed. Not on anything I’d written. Two jobs that, it turned out, had been failing on every single push to the main branch for months.
The root cause: the project’s dependencies now require Node 20 or newer, and the CI workflow was still pinned to Node 18. On the older runtime, a dependency three layers down doesn’t just warn, it crashes the test runner outright. Nobody had caught it, because a red check had apparently stopped meaning anything a long time ago.
Fixing that uncovered a second problem underneath it. A hundred and sixty-five files across the repository had never actually been run through the project’s own formatter, and a test coverage threshold set at fifty percent had never once been met, with the real suite sitting closer to a quarter of that. I bumped the Node version, ran the formatter across everything and reviewed every diff by hand to confirm it was whitespace only, and brought the coverage threshold down to match what the suite genuinely achieves, with a comment explaining why instead of quietly deleting the check.
CI went green. Apparently for the first time in its recorded history. Then I merged.
What this actually proves
A rules file and a lint rule are cheap to write and easy to ignore. What made this worth doing was refusing to just claim it worked. Three independent test runs, checked by hand rather than taken on faith, is a small sample. It’s still a real one, and more than most systems like this get before someone calls them finished.
The most interesting part of this project wasn’t the system I built. It was what enforcing it turned up: a typo quietly breaking a hover state, two files that looked correct individually but fought each other in practice, a public function that had never worked once, and a build pipeline that had been broken since before I started. None of that shows up when you write the rules. It shows up when you actually hold the codebase to them.
This started from hvpandya.com/llm-design-systems, the original piece on building LLM-readable design systems. Aural UI is open source. The rules file, the full spec system, and the audit script are in the repository. The live page walking through all of this in more depth is here.
I Made a Design System Readable by AI. Then I Tried to Catch It Lying. was originally published in Bootcamp on Medium, where people are continuing the conversation by highlighting and responding to this story.