Enforcing design system adherence when agents write your UI
Most design system teams are now writing documentation for their coding agents: a DESIGN.md, an AGENTS.md, a set of Cursor rules, a skill, an MCP server. The reasoning is simple: the agent produces off-system UIs because it doesn't know about the design language, so you feed it the context about the system to improve the output. This is great, but there are also limitations.
I've built all of that for the Reva design system this year, and AI-generated code got noticeably better. But telling an agent about your rules is not the same as your rules being enforced, and the gap between the two turned out to be a more interesting problem.
What we built
Our monorepo has three context layers:
A design contract. DESIGN.md at the root: colour, typography, layout, components, icons, motion, accessibility. Written for agents and humans in the same file.
A design-system MCP server. Read tools so an agent can look up a real component API, its variants and its rules instead of guessing.
A skill that loads whenever anyone edits a .tsx file and holds the cross-cutting rules a lookup will not return.
The DESIGN.md colour section.
This raised the quality of generated code noticeably. Then a card shipped with three padding overrides on a slot that already owned its padding, two all-caps labels in a system whose contract says small labels are sentence case, and a set of demo tiles using bg-purple-500, bg-red-500 and bg-blue-500, none of which exist in our palette.
The instruction existed. It was specific, it was loaded, and the output still drifted.
The audit flagged two things in this card: the all-caps section labels, and the vertical padding stacked on CardContent. Everything else here, the per-row actions and moving the connect prompt into Accounts, is a redesign that followed. No linter can catch that part, which is rather the point.
Why context is not enough
This is not only our experience. Recent research helps explain why.
Meta put it most directly, on the public wiki for their astryx design system: "Context injection fails because it's optional." A findings table on the same page adds: "Rules are suggestions, not constraints, AI can ignore them without consequences".
Three papers published this year measured it.
Khatri (arXiv:2607.27250) ran 288 evaluated runs across 17 real tasks, comparing Claude Code and Codex with the context file removed, always injected, and selectively retrievable. The finding: "Context strategy does not measurably move correctness on either agent." The omnibus test came back at p=1.00 and p=0.66.
Gloaguen et al. (arXiv:2602.11988), from ETH Zurich, found that developer-provided context files improve agent performance by 2.4% on average, at p=21%, and help every agent tested except Claude Code.
RepoComplianceBench (arXiv:2607.26819) asked whether agents read the rules at all: "Today's agents almost never proactively retrieve the contribution rules."
Two caveats. All three papers measure general coding correctness on Python repositories, not design token adherence, so the transfer is not proven. And in my own use the context layer clearly does help: generated code is better with the skill loaded than without it.
They observed that context raises the average, although it doesn’t give you a guarantee, and a design system is supposed to be a guarantee, with guardrails, with expected outcomes.
The violation a linter cannot see
Off-system code arrives in four tiers, and they need different machinery.
Tier 1: the literal. #bd5d20 in a className. A regular expression catches this. Our rule is color-literal, severity error.
Tier 2: the foreign token. bg-blue-500. Valid Tailwind, not in our palette. The checker needs to know which colour families are ours. Our rule is raw-tailwind-scale, severity error.
Tier 3: the right token in the wrong role. text-amber-600.
Amber is our brand colour. amber-600 is a real foundation token in our system, correctly spelled, resolving to oklch(58.54% 0.1429 48.11). There is no literal, no arbitrary value and no foreign palette, so it passes tiers 1 and 2 and it would pass every off-the-shelf linter I have tested.
It is still wrong. In light mode our semantic role --fg-link resolves to that exact colour. Using the foundation stop directly pins the link colour to a ramp position instead of to the role. Retune the ramp, or switch to dark mode, and the two stop matching.
Catching this requires resolving amber-600 to a colour, resolving every semantic role to a colour in the same mode, normalising both, and looking for a collision. That needs the token graph, not the text. Our rule is raw-stop-vs-role, and the finding reads:
"text-amber-600" uses a raw foundation stop where the semantic role "--fg-link" applies.
The same input, down two paths. Nothing about text-amber-600 is lexically wrong, so a matcher has nothing to match on. Only a checker that resolves both sides to a colour can see that the token and the role are the same value.
Tier 4: the missing component. <div className="flex items-center justify-between border-b py-3"> where Item component exists. This needs a live index of what the component library exports today.
Most tooling in this space covers tiers 1 and 2: ban the hex, ban the arbitrary value, ban the foreign palette. That is worth doing. But tiers 3 and 4 are where the violations that survive code review live, and you cannot reach them without a resolver and a component index. The token pipeline work most of us have already done turns out to be the prerequisite.
Terminal output from auditing a snippet containing all four violation tiers.
One rule engine, four surfaces
The rules are code, so they can run anywhere. We run the same engine in four places:
As an MCP tool. The agent audits its own output mid-task and fixes its own findings. Fastest loop, and the least reliable, because it depends on the agent choosing to call it.
As a CLI. bun run audit:design over any file, directory or glob. No agent, no MCP server, nothing clever. A person can run it, and so can a script.
As a pre-commit hook. Staged UI files only, so most commits skip it. Fast feedback, locally, before anything leaves the machine.
As a required status check. CI audits the files a pull request changed and blocks the merge on error-severity findings. The only one of the four that is not optional.
The rules are written once. Four thin adapters consume them, and three of the four depend on somebody choosing to run them.
The last one is the point. A gate that applies only to AI-generated code has a hole in it: the rules are either true of the codebase or they are not.
The Design audit check failing on a pull request.
The same job in detail. bg-blue-500 is not in the Reva palette, so the audit exits non-zero and the check fails.
What the audit reports today
Every error-severity rule reports zero. No hex literals, no off-palette colours and no banned icon imports anywhere in the codebase. That is the only number I really care about, and it has held while the surface the audit covers has grown by about half.
Everything else it reports is a warning, and those sit in under a fifth of the files scanned. By rule:
layout-primitive, around 40% of findings
arbitrary-value, just under 30%
typography-tag, about 20%
bespoke-vs-primitive, uppercase-eyebrow and cardcontent-padding share the rest
Warning density has fallen by roughly 40% since the gate shipped, against a codebase that grew by about half over the same period.
Two things in the distribution were not what I expected.
Half the arbitrary values are the same two utilities. text-[10px] and rounded-[2px] account for just over half of that rule's findings between them. They are not dozens of separate mistakes. They are two gaps in our scale, a missing caption size and a missing small radius step. Each author reached for an escape hatch because the system did not have the value they needed.
So the backlog is not a cleanup list. It is a list of tokens we have not defined yet, ordered by how often people need them. I did not build the audit as a research tool and it turned out to be one.
Most of the remaining warnings are a boundary, not drift. Nearly every layout warning sits in a chart component, because the charting library needs raw containers for tooltips and legends that our layout primitives were never designed to wrap. The same is true of the token specimen pages in our documentation, where a swatch grid legitimately needs a raw grid and a 10px label. Neither is a cleanup job. Both tell me where the primitive set stops reaching, which is a design brief rather than a backlog.
The audit summary line.
What I learned building it
False positives cost you the gate. The first full run reported 17 errors. Thirteen were bugs in my own rules rather than real violations. A gate that cries wolf on day one gets switched off, and rightly. Errors block, warnings inform.
Gate the diff, not the repository. CI audits only the files a pull request changed, so pre-existing violations never break unrelated work. This is what makes it possible to install a gate on an existing codebase.
Fail open on a missing toolchain. If the built token artefacts are absent, the CLI exits zero instead of blocking. CI builds tokens first, so the full rule set still runs where it counts.
The context layer still earns its place
The conclusion is not that rules files are useless. The MCP server, the skill and the design contract are why most generated code is roughly right on the first pass, and why the agent reaches for Item instead of a hand-rolled row. They also make the findings actionable: the same system that flags text-amber-600 can say what to use instead.
What changed is the job they do. Context is not the control. It is what stops the control from firing.
Three questions for your own system
Can your design system answer "is this on-system?" as a yes or no, on a specific file, with an exit code? If the answer only exists in documentation, it is not enforceable.
Could your checker catch a correctly-spelled token used in the wrong role? Not if it matches strings. That needs the token graph you probably built for the build step and use for nothing else.
Is it a required check? A bot that comments on a pull request is advisory. A required status check is not.
Conclusion
Teach, generate, verify. The first two are probabilistic and will stay that way. Only the third can be deterministic.
Most design system work over the last decade went into the first half of this: naming things well, structuring tokens properly, writing documentation. That work is the prerequisite, not the answer. The second half is turning the rules into code and running them somewhere nothing merges without passing.
If you have built something similar, or you are weighing up whether it is worth the time, I would be interested to hear about it.