What would a specialized design harness even look like?
General-purpose agents get worse as you stack more onto them. Brand work needs a harness built for it: a schema, a component catalog, a rule engine and a workflow it cannot skip.

I was about to tweet this, but it needed room to breathe. I've been thinking about it for a while and I want to put it out there to see what other people think.
The conversation about AI right now is too focused on the models and not enough on the architecture that surrounds them. I think systems matter more than ever now.
In coding, in design, in any domain where you need deterministic, brand-safe, scalable results, you have to build a system first, document it and only then are you ready to automate. A system is not just processes. It's conventions, repeatable steps, a library of distinct, reusable building blocks. AI is not a replacement for this. It is the ultimate consumer of it.
We are all scaffolding systems on top of one another
The past few months have shown us that the best way to get a multimodal AI agent to actually function is by giving it tools and a harness. At its core, that harness is a system.
The AI engineer creates a harness for the foundation model to turn it into an agent, giving it tools, instructions, constraints. That harness is a system.
The designer adds their design system to the stack; conventions, visual heritage, pre-approved building blocks.
The end-user stacks their own system on top of that, using skills (step-by-step instructions for performing tasks) to configure the agent for their specific work.
Everyone is scaffolding systems on top of one another.

The problem with stacking
Here's what I keep running into. Tools like Claude Code, ChatGPT, Cursor; they are generalized harnesses. They're built to handle anything (code based usually). And they're genuinely good at that. But when you start building for a specific domain, the "handle anything" part becomes the problem.
You add a design system as context. You add skills for specific tasks. You add MCP servers for your tools. You add custom instructions for your brand voice. Each addition is useful on its own. But now the agent is reading through pages of instructions, choosing between dozens of tools, and burning through its context window before it even starts working. The model gets slower. It gets confused about which tool to use for what. It starts mixing up instructions from one skill with constraints from another.
I've watched the same model go from sharp and focused to scattered and unreliable just by stacking too much onto a generalized harness. The irony is that everything I added was supposed to make it better.
This video from Theo, a software developer and YouTuber, walks through this paper from SkillsBench on how agent skills perform across diverse tasks. It goes into detail on how filling the context window with low signal-to-noise information hurts the agent's output.
The same AI model could feel "junior" on a bare generalized harness. Add your design system to the stack and it feels "mid-level." But I think it could perform like a "senior" operator if it were wrapped in a purpose-built harness, one that only knows your domain and doesn't waste cycles figuring out what it's supposed to be.
A few people are already realizing that you can build a custom harness for a specific vertical instead of stacking more and more systems onto a generalized one.
So what would a specialized design harness actually look like?
I've been thinking about this specifically for brand and product design. What would a harness look like that constrains an agent to design-system-true, brand-safe, structurally valid outputs, while also giving it real fluency in a specific design language?
The constraints define what the brand is not. The extension defines what the agent can do within those bounds. Both are required. I think the harness needs four layers:
Design language schema. The harness accepts goals in human language ("Design a pricing page for X persona") but forces the agent to respond in a declarative structure; JSON, a design DSL, something structured. It maps intent to layouts, variants, and allowed properties, so the model can't just improvise random UI elements.
Component catalog. The harness exposes a whitelisted set of tokens and components; your Figma library, your React system, whatever you're working in. The agent can only compose using your approved building blocks. If the model generates a perfectly good card but uses a color that's not in your palette, the harness catches it: "that color doesn't exist in this brand." The off-brand output is reworked. The model corrects and tries again, creating a validation loop that catches failure states before they can make it to the canvas.
Rule engine. We lint code. A design harness lints layouts. It enforces brand rules (color usage, typography scale, logo safety) and UX heuristics (contrast ratios, tap target sizes). It won't accept a card with off-brand border radii. Brand safety and UX quality become properties of the runtime, not subjective calls made after the fact.
Workflow constraints. The harness enforces how a design team actually works. It orchestrates the required steps: generate options, critique, refine, export. The agent doesn't need to "remember" your ideation or review process. The harness makes it structurally impossible to skip steps.

Restriction is the identity
In brand design, restriction is not a limitation. It is the identity.
What a brand is not matters as much as what it is. A wrong type weight, a color two shades off, a component used outside its intended context. Tiny things like these can throw an entire brand off and designers have always known this. A specialized harness built from that knowledge is what makes AI know it too.
The harness gives the agent a cage and a toolkit. The cage is the brand. The toolkit is everything the agent can do within those bounds and when extending the brand, the cage can also be extended to accommodate the new changes. Together, they make deterministic, brand-safe output possible in a way that a generalized harness with a stack of skills on top just can't.
Where I'm going with this
This is what we're working toward at RepliHaus. I don't have all the answers yet, but the direction is clear to me: it looks less like a prompt and more like a design system that enforces itself. The agent speaks only in the dialect of your brand's tokens, grids, and patterns. Everything outside those bounds gets caught, corrected, and looped back.
I think the companies that do well in the next era of creative operations won't be the ones with the smartest models. They'll be the ones with the most disciplined harnesses. Everyone will be a specialised "systems designer" in the future. If you've spent time caring about how the work gets done and not just the output, I think you're already closer than you realize.
I'm still working through a lot of this. Would love to hear how other designers and builders are thinking about it.
Where this is already happening
Added October 2026. Since I wrote this, we have shipped the first piece of that harness. Kitana's MCP server is the rule engine and the validation loop, made available to any agent: the agent makes the work, Kitana checks it against your brand, the agent fixes what she flagged, and she checks again. The engine behind her measures what code can measure and only lets the model reason about the rest. We wrote about how that is built in our engineering notes.