Context isn’t compliance
Most of the big design systems can now brief an AI agent. The next job is vouching for what the agent built.
Over the last year, most of the big design systems gave AI agents a way in. Kaelig Deloumeau-Prigent’s State of AI in Design Systems study looked at 21 open-source systems and found that 18 have an official MCP channel. MCP is a standard connector that lets an agent look things up while it works. The study also describes what those connectors do, and it’s nearly the same everywhere: “search docs, get component docs and examples, list tokens.” Tokens are the named values a system uses for things like color and spacing.
That’s real progress in a year. An agent that can read the actual button documentation has a better shot at using the actual button than one working from whatever it absorbed in training.
But looking something up isn’t the same as following it. And following it doesn’t let you know if the output was actually correct. Nearly all of those new MCP connectors from the survey stop at access.

What the review used to do
People reviewing screens have long caught what automated checks miss. That happens in a design QA pass before release, in an accessibility audit, or in a critique where a designer says that’s not how we show an error. The review was never perfect. It depended on who had time. But it existed, because people built the screens and other people looked at them.
The connectors give an agent the system’s guidance while it works. Someone still has to review what the automated checks don’t cover, and agents can produce screens faster than a team can review them.
The bet underneath all this investment is that an agent with good enough context will follow the system. I understand the bet. Context is the part a design system team controls, and it can be measured. Atlassian’s team compared handing an agent their guidance as a single DESIGN.md file against their MCP server on a simple task, building a log-in screen, and found the file alone “required ~92% more tokens, took longer to produce results, and had ~2.7x the variance in token consumption between runs.” That’s careful work on how to hand an agent a system. They also lean on lint rules, which “enforce quality frontend coding standards for humans and agents alike with no token spend at all.”
A study published in July shows how far a tool’s explanation can drift from what it builds. In “Design Theater,” Kashif Imteyaz and colleagues looked at 120 interfaces from five generative UI tools and compared the design reasoning each tool showed its users with what it actually built. They found that, on average, roughly one in four stated design rationales wasn’t implemented in the generated interface. For functional requirements, the miss rate reached 34%. The tool explained a design decision, and the interface didn’t contain it.
For design systems, the lesson is that a tool can describe a rule without following it. An agent’s account of its own work is one more output of the same model, and it needs checking like everything else the model made.
What a check has to say
Checking is starting to show up. Kaelig’s study gives it a finding of its own, “Validation loops turn guidelines into gates,” and describes Salesforce shipping “a dedicated validate skill plus slds-linter” and Fluent wrapping “lint auto-fix loops with Storybook and Playwright visual checks.” Knapsack published a design system contract for agents in September and was plain that some of its rules have no automated check yet: “Until there’s a checker, those are rules on paper.” On September 15, Applitools released checks that compare built UI against Figma designs, along with tools that bring its visual checks into coding agents like Claude Code and Cursor.
So the direction is clear. What I don’t see yet is agreement on what a check should hand back.
Most of the checks I found work like a linter. They report what they found wrong in what they looked at. That’s useful, and it leaves out the thing a reviewer most needs to know, which is what they didn’t look at. If a report lists no failures, you can’t tell whether the screen passed forty checks or four. Silence reads as a pass.
The rules most likely to go unchecked are the ones the design system doesn’t contain. Libraries like Carbon and Spectrum serve many products, so they don’t hold the business rules of any one of them: what a subscription is, what cancelling one does, who’s allowed to do it. I think we’ve set the bar for meaning too low. A semantic color token tells an agent how to use a color. The product’s own rules decide what a screen should say.
Say cancelling a subscription stops renewal, and access continues until the paid period ends. An agent builds the account page from the approved components, and it passes its color and spacing checks. The confirmation message tells the customer their access ends immediately. No check covered that message, so the report lists no failures.
Accessibility testing worked this out years ago. The axe-core engine sorts its results into passes, violations, incomplete and inapplicable, and “incomplete” means the tool couldn’t decide and a person needs to look. The W3C’s EARL format for test results has outcomes for passed, failed, can’t tell, inapplicable and untested. Can’t tell and untested are the honest answers. They’re also the ones a single pass/fail badge has no room for.
A useful check for agent-built UI needs that vocabulary, plus one more thing. It has to point at the exact artifact it checked, with a fingerprint of the checked files (a content hash). If those files change, recalculating the fingerprint reveals the mismatch. Without that, a result can outlive the thing it described without anyone knowing.
Put together, that’s a record. This artifact. These checks passed, these failed, these weren’t run. It starts with a list of rules the screen is expected to follow, including those that have no automated check. With the cancellation rule on that list, the account page’s record would show that color and spacing passed and the cancellation message went unchecked. The reviewer would know where to look. That’s what lets a design system team vouch for work they didn’t watch being made.
Where it’s already happening
A few people are building this. TJ Pitre’s Story UI generates Storybook stories, the example pages teams use to show a component, and badges each one “Verified,” “Issues” or “Not verified.” A check that couldn’t run is listed as not run instead of being counted as clean.
I’ve released two open-source tools that work the same way. With OODS Foundry, an agent asks for a screen by object and context, such as a subscription’s detail page. Foundry builds it from the system’s components and tokens and returns a receipt listing the checks it ran, the checks it didn’t, and a content hash of the output. Charts go further. Foundry certifies them against a fixed set of checks, from whether the data table matches the chart to whether a bar chart starts at zero. Stage1 works from the other end. It measures a live site against a Figma file or the tokens in your CSS, and it reports “measurements, coverage and missing evidence, not a design-quality score.” Its rule is one I’d apply everywhere: “absence is not zero.”
What the system has to hold
Getting context to agents was the right first step. As agents build more of the screens, people will look to the design system team to vouch for those screens. The team can only do that for rules the system defines, on screens its checks covered.
That’s why I think objects, traits and contexts are better units of reuse than components. An object is something the product is about, like a subscription. A trait is a capability it has, like a lifecycle that runs from active to cancelled. A context is where it appears, like a detail page or a list. When the unit of reuse is a component, like the confirmation dialog, each screen that handles a cancellation has to get the rule right on its own. Define the rule once, on the subscription’s lifecycle, and every screen that shows a subscription can be built from it and checked against it. That’s the bet OODS Foundry is built on.
To see where a system stands, pick one rule your product depends on, like who’s allowed to issue a refund. Ask where it’s defined, which screens it reaches, and what checked those screens, including what the check left out. Those answers are the difference between briefing an agent and vouching for what it built.
Originally published at https://derekniedringhaus.substack.com on September 29, 2026.
Context isn’t compliance was originally published in Bootcamp on Medium, where people are continuing the conversation by highlighting and responding to this story.