← All posts

· Jeffrey Emakpor

How Kitana reviews a design without guessing

The best AI systems do not put an LLM at the centre of everything. What Kitana's perception layer, deterministic scorers and eval harness actually look like.

engineering noteskitanaevals

When people hear "AI brand review", they picture a vision model looking at a poster and saying whether it is on-brand. That is the version we did not build.

A model asked "is this on-brand?" will always give you an answer. It will not always give you the same answer twice, it cannot show you which pixels it looked at, and when it is wrong it is wrong with confidence. That is fine for a second opinion. It is not fine for a score a team uses to decide whether work ships, or a score an agent uses to decide whether its fix worked.

So Kitana is built the other way round. Code measures everything that can be measured. The model comes last, and it only talks about what the code found. These notes cover how that works, in six principles and one experiment.

Three stages: perception extracts features, deterministic scorers measure them, and a model explains, classifies and compares at the end
The model sits at the end of the pipeline, not the centre.

1. Make deterministic things deterministic

If code can measure it, an LLM should not guess it. A review runs in three stages:

  1. Perception. We extract features from each asset: text regions, layout signals, colour clusters and logo usage.
  2. Measurement. Deterministic scorers turn those features into scores, one per category: colour, typography, logo, layout and coherence.
  3. Reasoning. Only then does a model get involved. It turns findings into plain language, reads brand intent, and handles the cases that are genuinely ambiguous.

Colour is the clearest example. Whether an accent is off-palette is not a judgment call. We measure how far apart two colours look to the human eye (a ΔE value) against your brand palette. When Kitana says your accent is #4F86F7 and your brand accent is #1D4ED8, that is arithmetic, not opinion.

2. AI should reason over evidence we can see

A score is only as trustworthy as the evidence under it, so we built a way to look at the evidence.

Our calibration view draws what the engine recovered directly on the asset: margins and regions, edges that line up, a 6 px near miss, a 28 px gap, a logo collision, an image breaking the margin. If the engine got something wrong, you can see exactly what.

Every mark is editable, and the edit becomes a label. A correction is not thrown away. It becomes test data, and the engine improves by being shown precisely where it saw the asset wrong.

The calibration view: a social post with its margins, regions, a 6px near miss, a 28px gap, a logo collision and an image breaking the right margin drawn on it
The calibration view. Every mark is something the engine measured, and every mark can be corrected.

3. Separate measurement from judgment

Every finding carries a confidence tier, and the tier controls how loudly it is allowed to speak.

  • Measured: stated flatly. Maths on the original pixels, like contrast, ΔE and margins. A measured finding can fail an asset.
  • Estimated: hedged. It depends on what recovery returned, like text regions or an inferred grid, and the wording says so.
  • Advisory: soft. Model judgment, like "is this on-brand?". An advisory finding never fails an asset.

This is the rule that stops a model's taste from quietly becoming a failing grade. The model can raise a point. It cannot sink your score on a hunch.

Three confidence tiers: Measured can fail, Estimated is hedged, Advisory never fails
The tier decides how a finding is worded, and whether it can fail an asset.

4. Evaluation is part of the architecture

We do not evaluate the engine now and then. Evaluation is a component of it.

Each eval run re-scores a labelled corpus and records band accuracy, the full score distribution for each category, and the parameters that produced it. Each fixture in the corpus is a set of serialisable features plus an expected score band, so a score can be checked rather than argued about.

The corpus grows with the engine. It is append-only and versioned, and it gets bigger with every release. A case that once caught a bug stays in the corpus, so the bug cannot quietly come back.

Labelled corpus size per eval run, rising in steps from the first run to the latest, with engine releases marked along the line
The labelled corpus grows with every release. Append-only and versioned.
An eval run: 100% band accuracy, 100% weighted, 100% knob coverage, with the mean, median, min and max score for each category
One eval run: accuracy and the full distribution for each category.

5. Build for reproducibility

Same asset, same engine, same parameter snapshot: same score. Every time.

Each eval run appends one record to a log: when it ran, the engine version, the commit, a snapshot of every tuning knob, and the results for each category. When a number moves, the log says which knob moved it.

One run log record: timestamp, engine version, commit, a snapshot of every tuning knob, and results per category
One record per eval run, append-only.

This is what makes the MCP loop trustworthy. When your agent fixes a poster and asks Kitana to compare the two versions, the change in score comes from the change in the poster, not from the reviewer having a different day.

6. Do not hide uncertainty behind a score

A number is only useful if you know what produced it. So the eval also measures itself.

Every tuning knob (colour tolerance, the penalty for mixed alignment, the weight of each category, the logo rotation limit, the grid adherence threshold and the rest) is swept across its bounds. For each one we record how many cases it moves, by how much, and how many cross a score band.

A knob that moves nothing is a knob the corpus cannot see. The eval reports it as a gap instead of letting it pass silently, and we write new fixtures until every knob is observable. Without this, a perfect accuracy number can simply mean the corpus was not looking.

A sensitivity table: for each tuning knob, the number of cases it moved, the largest change and the number of band crossings
Each knob swept across its bounds. A knob that moves nothing is reported as a gap.

The experiment: we test the LLM against the code

We do not assume code beats a model. We test it.

We gave an LLM judge the same corpus, with a rubric for each category. The judge does not return a number, because a model asked for "a score out of 100" will invent one. It returns a discrete verdict, and fixed rules map that verdict to a score.

The deterministic engine landed in the expected band every time. The judge got 90%: perfect on colour, logo and coherence, and 75% on typography and layout.

The misses are the interesting part. They were one-rung severity calls, never invented grades. The judge saw the same problems as the code. It was less precise about how much they mattered.

Band accuracy per category: the deterministic engine at 100% everywhere, the LLM judge at 100% for colour, logo and coherence and 75% for typography and layout
Same corpus, same rubric. Code 100%, LLM judge 90%.

So today the model writes the explanations, and the numbers come from code. That is not a permanent verdict on models. It is a standard: a model takes over a measurement when it can show, on the same tests, that it does the job better.

In one line

AI interprets. Systems measure. Data validates. Humans judge.

The interesting engineering challenge is not getting an LLM to produce an answer. It is building the system around it so that answer can be trusted.

If you want to see the result, run a review free, or connect Kitana to your agent and watch it check its own fixes.