AI Studio.

AI Studio Build

enter the access code
That’s not it — try again
AI Studio

Open on a desktop

This experience compares mobile screens side by side — it needs a desktop window to be judged properly.

AI StudioAI Studio · Build · Mobile

Improving AIS mobile built applets design

Build's mobile applets should look designed, not AI-generated — fix accessibility and mobile ergonomics, and avoid the "AI look". One minute to see for yourself that our design instruction does it, then the paste-ready instruction and the flow that keeps improving it.

1 One minute

Click the better design. Each pair is the same app prompt built twice by AI Studio — you're picking blind. Judge the design, not the feature set. The clock starts on your first pick.

1:00
No wrong answers — your eye is the ground truth.

you picked the instruction-built designyou picked the default
2 Proposed AIS mobile design system instruction

What you just compared: default Build output vs the same prompt with this instruction added. Good design here is not taste — it is hard, measurable variables (type ≥14px, zero overflow at 390px, contrast, 44px targets, one accent) plus escaping the "AI look" (the violet-template / dark-glow / rounded-everything signature humans reliably spot and dislike).

How it was built

  • 157 fresh builds — 40 mobile prompts × (2 default + 2 instructed), generated in AI Studio
  • Automated lint on every rendered screen: overflow, sub-14px text, contrast, wrapped labels, the violet "AI-tell" pixel share
  • 15 blind panels across GPT, Gemini and Claude — 1,184 paired judgments, de-blinded against a private key
  • A causal follow-up (32 builds) validated three added rules on the weakest genres

What it measurably does

  • Blind raters prefer instructed builds 73% (+0.68/5) and call the default "the AI-generated one" 72% of the time
  • Sub-14px text: 3,897 → 21 · page overflow: 347px → 0
  • Violet "AI-tell" share: 25% → 4% · dark-by-default: 40% → 12%
  • Zero negative side effects on code quality (paired audit of all 157 codebases)
design-system-instruction.txt · v5 · paste into Build's system prompt

  
All 40 prompts — every pair we measured, default vs instructed

One representative pair per prompt (this is the full corpus behind the numbers; the 1-minute game draws its pairs from the strongest of these). Click any screen to open it full-size.

3 Continuous design-instruction improvement

The instruction above is one iteration of a loop, not a one-off. The loop is cheap, fast, and — as you just experienced — the ground truth comes from a minute of human judgment at a time.

1
Tweak
An agent proposes a surgical change to the instruction, from observed failure classes.
2
Build
The prompt set is rebuilt with and without the change — the change is the only variable.
3
Measure
Automated lint (overflow, type, contrast, AI-tell pixels) + model blind panels score every pair.
4
Ground-truth
Humans settle what models can't — one minute of pair-picking at a time.

3.1 · Accumulating great designs

With product, design and marketing we are compiling an ever-growing reference list of AIS applets — the best as positive ground truth, plus instructive failures. This list seeds both the auto-rater's golden set and the instruction's next iterations.

Open the applets list (Google Sheet) ↗

3.2 · The ground-truth tool

Section 1 is the proposal: an internal pair-picking tool where anyone from product, design or marketing spends a spare minute choosing better designs (optionally leaving a note). Every pick is a labeled ground-truth judgment; every judgment tunes the auto-rater and scores the next instruction candidate. You just produced ~10 expert judgments in 60 seconds.

The math: 10 people × 2 minutes a day ≈ 200 expert judgments daily — a continuously fresh golden set, at zero meeting cost, from the people whose taste should define "good."

Everything on this page is reproducible end-to-end — corpus, lint, blind-panel protocol and de-blind keys are documented in the golden-set work log.