
This experience compares mobile screens side by side — it needs a desktop window to be judged properly.
AI Studio · Build · MobileBuild's mobile applets should look designed, not AI-generated — fix accessibility and mobile ergonomics, and avoid the "AI look". One minute to see for yourself that our design instruction does it, then the paste-ready instruction and the flow that keeps improving it.
Click the better design. Each pair is the same app prompt built twice by AI Studio — you're picking blind. Judge the design, not the feature set. The clock starts on your first pick.
What you just compared: default Build output vs the same prompt with this instruction added. Good design here is not taste — it is hard, measurable variables (type ≥14px, zero overflow at 390px, contrast, 44px targets, one accent) plus escaping the "AI look" (the violet-template / dark-glow / rounded-everything signature humans reliably spot and dislike).
One representative pair per prompt (this is the full corpus behind the numbers; the 1-minute game draws its pairs from the strongest of these). Click any screen to open it full-size.
The instruction above is one iteration of a loop, not a one-off. The loop is cheap, fast, and — as you just experienced — the ground truth comes from a minute of human judgment at a time.
With product, design and marketing we are compiling an ever-growing reference list of AIS applets — the best as positive ground truth, plus instructive failures. This list seeds both the auto-rater's golden set and the instruction's next iterations.
Open the applets list (Google Sheet) ↗
Section 1 is the proposal: an internal pair-picking tool where anyone from product, design or marketing spends a spare minute choosing better designs (optionally leaving a note). Every pick is a labeled ground-truth judgment; every judgment tunes the auto-rater and scores the next instruction candidate. You just produced ~10 expert judgments in 60 seconds.
Everything on this page is reproducible end-to-end — corpus, lint, blind-panel protocol and de-blind keys are documented in the golden-set work log.