A detailed, plain-language guide to building an evaluation loop with Claude, so it designs, scores its own work against real criteria, learns from what it missed, and keeps getting better on the next screen. Every technical term is explained. No engineering background assumed.
<aside> 🧠
This is a starting point, not a finished agent for your exact team.
Think of it as a scaffold you adapt, not a tool you drop in untouched. The whole idea of an eval loop is that it encodes your standards, so the panel only becomes truly useful once you shape it to how your team actually works.
A couple of examples of what "customize it" means:
Same loop, two different rulebooks. Which one is right is a call only you can make, because only you know what "good" means for your product and your team.
So treat what's here as the fast way in. The rubric, the lenses, the thresholds, the gates: all of it is meant to be edited. Do that bit of setup and the loop starts working for you, not for some generic idea of "good design."
Enjoy 😊
</aside>
<aside> 🔁
The idea: Most designers use Claude as a one-shot generator. Ask, get a screen, look at it, ask again. That plateaus fast. An eval loop changes the shape: Claude designs, a separate critic scores the work against a rubric you set, the misses get written down so they stop repeating, and the score climbs on every future screen. You stop giving feedback by hand and start running a system that improves on its own.
</aside>
The loop has four moves, and this guide walks through all of them: design, evaluate, learn, remember. By the end you should be able to say "I'm on a client project, or I run a design team, and I can set this up myself."

The obvious move is to build a screen and then, in the same chat, ask Claude to review it. This barely works, and it is worth understanding why, because the fix is the whole foundation.
When Claude builds something, it also builds up a head full of reasoning: the choices it made, the trade-offs it talked itself into. Ask it to review that same work and it does not look at the result fresh. It looks at it through everything it already decided. It tends to defend rather than judge. There is a name for this: self-preferential bias, the tendency of a model to prefer its own output when asked to grade it. It is the same reason you cannot properly proofread your own essay. You read what you meant to write, not what is actually on the page.
<aside> 📉
Seen in real use. On an actual client build, the builder graded its own work 75 out of 100. Then a panel of critics that had not built it graded the same work, blind, and landed on 52. A 23-point optimism gap, on real work. That gap is the reason for everything below.


</aside>
The fix is structural, not a better prompt. You need a reviewer with a fresh context. "Context" just means everything the assistant currently has in front of it: the conversation, the files, the reasoning so far. A fresh context is a clean slate that never saw the build happen. Give that reviewer only two things, the work and the rules, and never who made it, and you get an honest score instead of a defense.

<aside> ⚖️
One line to remember: self-review is too kind to itself. The gains come from a critic that did not build the thing and has no reason to protect it.
</aside>
"Make it better" is not something a machine can act on. A rubric is. A rubric is a short, fixed list of what "good" means, with a way to score each item. It turns a vague feeling into specific, checkable lines.
Here is a design-native rubric that works for almost any screen: