Give your design harness a door.
A design harness is how an agent works with your brand system, with a person guiding it. Most harnesses only enforce. That is half a harness. Build the door: a sanctioned way to break the rules, and a record of when breaking them worked.

Every brand team now has two things to get right. The first is their brand system: the guidelines, the templates, everything that makes the brand theirs. The second is the design harness: the surface through which an agent works with that brand system, with a person guiding it, to produce design work. Most harnesses being built today do one thing. They enforce. Check the palette. Check the type. Flag the logo. Score the adherence. The pitch is always the same: the machine will stop your team from going off-brand.
That is half a harness. A design harness that only enforces makes work that is correct and dead. A launch poster that follows every rule looks like everything else you made that year. Designers have always known this. The industry is building the half that is easy to measure and calling it finished.
Add the missing part: a sanctioned way to break the rules.
The argument I am stepping into
There is a live disagreement about how much scaffolding an agent needs. Replit recently published Free the models and argued for less of it. Every model release breaks assumptions that the harness bakes in. A rigid harness forces one way of working on the model. A composable one lets the model choose. The smarter the model, the less the harness should decide for it. Replit reads this as an instance of the bitter lesson.
For coding agents, I think they are right. For brand work, the lesson does not apply. Be precise about why.
The bitter lesson is about capability. Scaffolding that makes up for what a model cannot do yet ages badly, because next year the model can do it. That is true. A smarter model will compose better layouts, write better copy, and need less help to get there.
But brand constraint is not capability. It is private information. Think of your hex values, your type scale and your logo clear space. Think of the three things your legal team will not sign off. None of that is in any training set, and no model progress puts it there. A design harness that gives the agent your brand system does not go out of date when the model gets smarter. It gets more useful, because the model works better inside it.
So scaffold less for capability. Scaffold for constraint. These are two different jobs, and the bitter lesson only eats the first one.
Stop piling on
Most teams try to specialize a general agent by piling things onto it. A design system as context. A skill for each task. Instructions for the brand voice. MCP servers for the tools. Each addition helps on its own. Together, they bury the agent. It reads pages of instructions before it starts. It picks the wrong tool. It mixes a rule from one skill with a limit from another. I have watched the same model go from sharp to scattered, only because I gave it more to carry.
The research agrees. SkillsBench tested agent skills across many tasks: low-signal context makes the output worse, not better.
On a bare general agent, a model feels junior. Add a design system and it feels mid-level. Give it a design harness built for brand work, and it can work like a senior operator.
The five parts of a design harness
Picture a brand team that ships posters, social posts, banners and launch assets every week. Give their agent a design harness with five parts.
1. Make the agent answer in structure. Take goals in plain language, like "make the autumn launch poster for Instagram." Make the agent answer with a declarative document, not loose pixels. Map the goal to the layouts, variants and properties your brand allows. The agent cannot improvise an element your brand does not have.
2. Give it only your building blocks. Expose your brand system: its colors, type, logo rules, and the templates your designers made. Let the agent compose only from those parts. When it reaches for a color outside your palette, it finds nothing to grab.
3. Measure what code can measure. Be honest about the rest. We lint code. Lint layouts the same way. Check what a machine can compute: an off-palette accent, a headline in three weights, a logo with too little clear space. Report each finding with the asset and the exact value.
Then say where measurement stops. Some checks are measured. Some are estimates, like a typeface match from pixels. Some are only the model's impression of the work. Let measured findings count against the score. Hedge the estimates. Never let the model's impression fail an asset on its own. A person's judgment can fail an asset, because that is the job. A machine's guess, dressed as a verdict, destroys trust the first time it is wrong.
Treat accessibility thresholds the same way. Expressive campaign work breaks rigid rules on purpose, and designers ignore a checker that does not know that. Flag the risk. Let the designer decide.
4. Put review inside the workflow. Run the steps the team already runs: generate options, review, refine, export. The agent does not need to remember the process. The harness makes it impossible to skip the review. When the review flags something, the agent fixes it and checks again. Stop after three attempts, and hand the work to a person. The reviewer never changes the work. It only reports.
5. Give the cage a door. Most design harnesses leave this part out. It decides whether your output has a pulse.
Restriction defines the brand. Divergence keeps it alive.
Restriction is the identity. What a brand is not matters as much as what it is. A wrong type weight, a color two shades off, a component outside its context: small things like these throw a whole brand off.
But great brand work knows when to break a rule, and a rule engine cannot see the difference between drift and intent. A headline that bleeds off the canvas is either a mistake or the reason the poster works. Only a designer knows which.
So build the door. Let a designer break a rule on purpose. Make them say so. Log the override, with who made it and why. Do not punish it. When the override works, mark it as good work and feed it back into what the system learns from. Over time, your harness learns your team's taste for divergence, not only the constants in your brand system.
A compliance tool cannot copy this. A compliance tool knows the rules. A design harness with a door also knows when your team chose to break them, and that record belongs to you alone.
Take the rules out of the context window
Here is a fair objection. I said MCP servers bloat a general agent. So why ship a rule engine as one?
Because a rule engine is not more instructions. It is a deterministic service. The agent does not read your brand rules and try to remember them. It sends the work out and gets a measured answer back. That moves the rules out of the context window, not into it. The agent carries less, not more.
That is how we built Kitana. Her MCP server is the rule engine and the review loop, open to any agent. The agent makes the work. Kitana measures it against your brand. The agent fixes what she flagged, and she checks again. Behind her, code measures everything it can, and the model only reasons about the rest. Our engineering notes show how.
What I would bet on
The teams that win the next era of creative operations will not have the smartest models. Everyone will have those. They will have the most disciplined design harnesses, and the clearest way to step outside them.
Enforcement is the commodity. Every tool in this space will have it within a year. The record of when your team broke its own rules, and which breaks worked, is the part nobody can buy.
Tell me how you think about it. What does your team's design harness look like, and where does it break?