AI Studio.

AI Studio Build

enter the access code
That’s not it — try again
AI Studio Build · mobile launch · design quality

Making every app Build generates look designed.

The goal: the apps AI Studio Build generates in the upcoming AIS mobile app should look designed, not AI-generated — and the auto-rater that grades them needs its visual blind spot calibrated. Two workstreams: (1) a golden set — matched good/bad app pairs + a 179-prompt corpus that give the auto-rater ground truth on design (accumulating; the pairs and prompt pack here are its first installments), and (2) a system instruction — empirically derived, currently design-focused, and blind-validated: all 40 prompts rebuilt ×4 (157 fresh builds), 36/40 score higher with the instruction under a 15-run panel across 3 model families. This page documents both, deliverables first, then the day-by-day work log.

Where this stands — the cumulative result

Effort 1: a correctness floor, blind-validated on 18 matched pairs — it beats Build's default on every rater family.  Effort 2: a lever ablation (76 builds, 10 blind panels) picked material as the one distinctiveness lever the Gemini auto-rater also rewards → v3 = floor + material.  Effort 3 (new today): the decisive test — the shipped instruction vs Build's default on all 40 prompts, ×2 replicated = 157 fresh builds. Blind panel of 15 verified runs across 3 model families (1,184 paired judgments): the instruction wins 73%, and raters call the default "the AI-generated one" 72% of the time. Objectively: sub-14px text 3,897 → 21, page overflow 347px → 0, the violet "AI-look" tell 25% → 4%. The sections below are organised as the deliverables first, then the day-by-day work log (Day 1 floor → Day 2 ablation → Day 3 at-scale validation) — each self-contained with its method and evidence.

The deliverable · The prompt set — 179 in the pack, 40 shownWhat a mobile user will actually build, stratified games → productivity. Marquee = the app's own home-screen suggested prompts. The full 179 ship in prompts.json below.

Marquee = the prompts the AI Studio mobile app surfaces on its own home screen — the first impression a new user taps, so those 18 are evaluated closest. The other 22 broaden coverage into productivity, utility, camera and voice.

Method (agreed with Terry): instead of curating links to existing AIS apps, we provide prompts we're confident generate good vs bad design — so design quality is a controlled variable in a matched pair. In parallel we're collecting a reference list of great-looking applets with the team as a positive design target.
Curated calibration pack (recommended): after the 157-build A/B, we kept only the prompts where the 15-run blind panel agrees (instruction-wins ≥77%, Δ≥+0.6) — 15 prompts in two tiers, each shipping with its measured separation (win rate, mean scores, AI-attribution) as the expected ground truth. Six candidates were excluded for weak/inverted separation (incl. space-game and time-machine) — fewer pairs, but every one reliable.
The deliverable · The design instruction — v4, paste-readyThe block to add to AIS Build's system prompt: the blind-validated floor + the ablation-chosen material lever + the two Day-3 rules (games open in-play; capture apps open on a sample result). 15 blind runs, 1,184 judgments: wins 73%, and kills the measured "AI look" (violet tell 25%→4%).

    
design-good.json (shipping floor)

Full evidence + method in the Day 1–3 sections below. Companion (not design): the WCAG-AA contrast lint and the slug-leak check stay post-generation gates — prose can't enforce them.

Day 1 · Method — paired by prompt, one variableSame prompt, generated with and without the design instruction. The gap is the calibration signal.
Day 1 · The proof it works — the floor beats the default, blind-validatedThe empirical backing: good = the correctness floor, bad = Build's default. Ranked blind by 6 raters across 3 model families on a 33-build probe — the floor wins for the Gemini family (the auto-rater's).

Stage 2 rebuilt 11 representative apps under three arms — Default (no instruction), Floor (the mechanized correctness instruction) and Distinctive (floor + an aggressive anti-generic push) — same prompts, same session, every arm captured at the same working screen. Six raters ranked them blind: 3× Gemini (the same family as Auto-rater v2), 2× Claude, 1× Fable.

design-good.json → Floor

  • picked best by 2 of 3 families (Gemini 23/33, Fable 7/11)
  • almost never flagged "most AI-looking" (4 of 66 flags)
  • beats Default for every family

design-bad.json → Default

  • unanimously the most AI-looking, weakest arm — all 6 raters, all 3 families
  • the deliberate "AI-look" instruction ships alongside as the more-extreme bad

Enforcement holds (objective, rater-independent)

  • 390px overflow: 12 → 0 elements
  • sub-14px text: 178 → 8 elements
  • prompt-slug leaks: banned even inside model/version strings
  • contrast: flat across arms → gated by a WCAG-AA lint on the built DOM, not more prose

Distinctive: editorial genres only

  • Claude raters crown it (mean 4.36) — Gemini penalizes it as over-styled (2.91)
  • unanimous winner on photography (6/6) → ships as an experimental, genre-gated variant
The cross-family finding: "good design" is rater-dependent — Claude-family raters reward expressive maximalism, Gemini-family raters reward restraint. Auto-rater v2 is Gemini-family, so the pack is ranked for that grader: what looks best to a human (or a Claude) is not what scores best on the auto-rater.
Day 1 · The 18 pairs — Default → BestThe good-and-bad examples themselves. Default is Build's out-of-the-box output; Best adds the one instruction. Tap a screen to enlarge.
Day 2 · The lever ablation — which distinctiveness lever earns promotionShips today: the correctness floor (v2). Next: v3 = floor + one distinctiveness lever (material) — chosen by the lever ablation to escape the generic "AI style" without over-styling. Paste-ready below.

The floor mechanizes correctness — zero 390px overflow, a 14px type floor, real contrast, a real brand, a seeded first screen, one primary action — so Build stops shipping its broken defaults. But a correct app can still look unmistakably AI-made: purple gradients, Inter, centered-hero-plus-three-cards, glassmorphism (see the research fold below). To escape that without over-styling into the "distinctive" block that Gemini-family raters penalize, we ablated the 5 distinctiveness levers one at a time on top of the floor.

Method: 11 prompts × 7 arms (floor · +type · +color · +asymmetry · +material · +density · +all), 76 builds at 390px, shuffled into a blind ladder and rated across 10 panels (a mix of Claude / GPT / Gemini).

Promote → v3: material (FORM)

  • one deliberate radius identity; borders + weight + contrast carry hierarchy; no reflexive glassmorphism / aurora-glow / floating 3D blobs
  • the only lever both human raters AND the Gemini auto-rater score above the floor (+0.46 combined, positive in 8/8 rater groups)

Kept out — incl. the trap

  • asymmetry = the trap: humans & GPT love it, but the Gemini auto-rater penalizes it (−0.18, replicated across independent runs) → editorial genres only
  • type / color / density / the full "distinctive" block: flat-to-negative
Why the trap matters: Auto-rater v2 is Gemini-family and rewards restraint. A lever that looks great to a human but scores worse on the rater would drag the launch metric — so v3 ships only the lever both agree on (material). Full evidence: golden-set/stage2/ablation/FINDINGS-firstpass.md; the block itself: instruction-v3-good-candidate.md.

The v2 floor, paste-ready (v3 adds the single FORM rule above + one radius identity ✓ to the self-check):

Day 2 · Second corpus — functional apps (independent replication + one new signal)A separate 8-pair good/bad benchmark on functional apps (diagnosis, games, productivity, dashboards, chat, media) — a different genre from the landing set — replicates the four-defect taxonomy and adds one signal for the auto-rater.

These 8 apps were generated in AI Studio from a crafted (good) vs a simple (bad) prompt and analysed by reading the source — so overflow and exact font-px are measured, not inferred (complementary to the vision rater, which is structurally blind to both). On a completely different app mix, the same signal holds: the gap is execution discipline, not features — the bad build is often more feature-rich.

Replicates the golden-set taxonomy

  • placeholder/slug identity → real brand (one bad build renders "S3P-p020-bad" as its title)
  • sub-12px text explosion (one bad build: 50× "text-xs")
  • low contrast (dark-on-dark) · 390px overflow
  • type moves most, genre-alignment least — theming isn't the bottleneck

One new auto-rater signal

  • Declared-capability-not-implemented: 5 of 8 bad builds declared the server-side Gemini capability and shipped zero AI code — a control that does nothing.
  • Invisible to a design-only rater; needs a build-time check (does the declared feature exist / call the API?).

Highest-leverage lever — confirmed

  • every good build emitted a "<!-- checks: dvh · no text-xs · one CTA · seeded -->" self-audit comment and coded to it — the same device already in the v-next candidate. Independent confirmation it's the clearest good/bad fingerprint.

Honest exceptions (carried forward)

  • bad was better on component modularity, real motion use, and feature breadth in several pairs
  • inverted case (p026): the good build dropped the AI, the bad build kept it — good>bad is not uniform

Recommended for v-next: add an "implement the declared capability — never ship a control that does nothing" rule. It is the one thing this corpus adds to the shipped floor; the graceful keyless first-screen it also confirms is already covered by the elite2 static-demo rule.

Read honestly: single analyst · source prompts inferred from outputs (not provided) · n=8 · v1. Shots are the built frontend at 390px served over HTTP with no API key — so "good" first screens render via their seeded/fallback state, and several "bad" builds show their un-degraded first screen (a key-wall or a cold/empty start), which is itself the finding. This confirms the golden-set on a new genre and adds one capability signal — it is not a competing corpus. Part02 extends it.
Day 3 · The 160-build A/B — the shipped instruction, blind-validated at scaleAll 40 prompts × (2 default + 2 with the instruction) = 157 fresh AIS builds. 15 blind panels, 3 model families, 1,184 paired judgments: the instruction wins 73% (+0.68/5) and the default is named "the AI-generated one" 72% of the time.

The decisive experiment: every prompt in the pack generated fresh in AI Studio, twice as Build's default and twice with the shipped instruction — the instruction is the only manipulated variable, and the ×2 replication per arm separates a systematic effect from generation luck. Each build self-identifies via an embedded BUILDTAG code comment; every screenshot is byte-verified against its own build before use. Raters saw randomized, shuffled A/B pairs with no hint an instruction existed.

Blind panel — 15 runs, 3 families

  • GPT (thinking ×3 · Pro — both decks) — family 71%
  • Gemini (Flash ×3 · Thinking ×2 · Pro — both decks) — the auto-rater's family, 67%
  • Claude (Fable ×3: 81/87/89% · Opus ×2) — family 81%
  • Combined: 862/1,184 = 73% wins, mean +0.68/5 · deck 1: 74% · re-randomized deck 2: 71%
  • "Which looks more AI-generated?" → the default, 72%
  • Robustness, three ways: excluding all gate/error pages (72% → 72%); an independently re-randomized second deck with post-gate screenshots (66%); every run content-verified against its de-blind key before counting

Objective — measured on all 157 renders

  • sub-14px text elements 3,897 → 21 (−99.5%); paired win 40/40 prompts — even the worst instruction build beats the best default
  • page overflow at 390px: 347px → 0
  • violet/indigo share of colored pixels (the Tailwind-template tell): 25.4% → 4.2%
  • dark-mode-by-default: 40% → 12% · palette discipline: 7.0 → 5.0 hue families
All 40 apps — Default vs With instruction, one pair each

One representative pair per prompt (gated builds shown post-gate; known defects left visible — the API-key screens and colliding meme captions ARE findings). Click any shot to open full-size.

What the losses taught (28% of pairs prefer the default): the consistent losers are games that open on a menu while the default opens in-play, and camera/AI apps that open on an upload form (one even showed a "could not load demo data" notice) while the default shows a structured flow. Both are now hard rules in v4 below: a game opens mid-play; a capture app opens on a designed sample result. The remaining losses are rater taste (the restraint family favors lighter dashboards), not defects.
The final instruction — v4, paste-ready (floor + material + the two Day-3 rules)

      
Reproduce these results — the full protocol, step by step

1 · Generation (the only manipulated variable). Each of the 40 prompts is generated in AI Studio Build in a fresh session per build: twice with the prompt alone (default arm), twice with the design instruction prepended (instruction arm — the exact design_instruction string from design-good.json). Every generation request appends one bookkeeping line asking the build to embed BUILDTAG:gsab-pNN-{def|ins}-{1|2} in a code comment — identical on both arms, design-inert, never rendered — so downloads self-identify regardless of folder naming. Coverage: 157 of 160 slots (credits ran out: 4 missing, 1 bonus duplicate kept as an extra replicate).

2 · Rendering (verified, not trusted). Each app builds against a pinned dependency set (Vite 6 / React 19 / Tailwind v4) and is screenshotted at 390×844 @2× over HTTP — never file:// (ES modules render blank). Two integrity guarantees on every shot: an ephemeral OS-assigned port per app, and a byte-comparison of the served index.html against the app's own dist file before shooting — a stale or wrong server cannot pass. (This protocol exists because an earlier run leaked 3 wrong-app screenshots through reused ports; all 157 were re-taken under it, 0 identity failures.) Pages are top-anchored before capture.

3 · Objective lint (measured on the rendered DOM, both arms, all 157). Per build: horizontal overflow px at 390; count of visible text elements with computed font-size <14px; clipped text (scrollWidth > clientWidth with hidden overflow, no ellipsis); overlapping text rects (>60% intersection); distinct border-radius values. Plus per-screenshot pixel metrics: share of colored pixels in the 230–290° hue band (the violet "AI tell"), dark-background share (lightness <0.18), distinct hue families.

4 · Blind panel. Two pages per prompt, each ONE default vs ONE instruction build: runs cross-paired at random, left/right randomized, page order globally shuffled with a no-adjacent-same-prompt constraint, side balance audited (~48% ins-on-A). The PDF and its de-blind key are versioned together and regenerated as a pair; raters are never told an instruction exists — the intro says only "the same prompt, built twice." Each run = a fresh conversation with the identical rating prompt (below). Panel: GPT (thinking, Pro), Gemini (Flash ×2, Thinking), Claude (Fable, Opus) on deck 1 — then a full second deck, independently re-randomized, with the 9 modal-gated apps re-shot past their gates (API-key/connection-error screens kept: the error is that build's first screen).

5 · De-blind & analysis. Ratings map back through the versioned key; we report instruction win rate, mean score delta, and the "which looks more AI-generated" attribution. Robustness checks: excluding all 13 pages where either side shows a gate/error state (72% → 72%), and the independent second deck (73%). Per-prompt net-losers are tracked across runs — the consistent ones (games opening on menus; capture apps opening on upload forms) became the two new v4 rules rather than being averaged away.

The rating prompt (verbatim, every rater, every run):
"You are an expert mobile product-design rater. The attached PDF shows 79 pages; each page is ONE app prompt built twice, labelled A and B. Sides and page order are randomized. Judge ONLY the visual design of the first screen: layout, hierarchy, typography, color, readability, polish, and whether it looks like a distinctive product or a generic AI-generated template. Rate every page independently — do not carry judgments across pages, and do not try to detect patterns or conditions. Do not skip any page. Return EXACTLY one table row per page: page | better | scores | more-AI | one-line reason. After the table, add a 5-line summary."

Everything lives in the repo: corpus/gsab/ (all 157 app sources by batch, verified shots, lint JSON, versioned de-blind keys) · golden-set/stage2/instruction-v4-final.md (the deliverable + evidence table) · golden-set/stage2/ablation/ (the Effort-2 lever ablation this builds on).

Appendix · Why "AI style" — the research the instruction is built to defeatThe catalog of what marks an app as AI-made. The floor kills the broken defaults; the material lever kills the visual tells — this is the target both were designed against.

The "AI look" is the statistical median of Tailwind/shadcn tutorials scraped 2019–24: forced to write logic and visual design at once, the model reverts to its most heavily-weighted defaults just to compile. Amplified by Tailwind's indigo-500 default, the shadcn/ui copy-paste monoculture, RLHF "aesthetic safety," and preview-arenas that vote on thumbnails. It's self-reinforcing — AI output becomes the next model's training data.

Color

  • "AI purple" indigo→violet→blue gradient accent; gradient on buttons AND hero
  • dark-mode-by-default (the single most common tell)

Depth & glass

  • uniform rounded-2xl on everything; floating cards + soft diffuse shadows
  • aurora/mesh glow, floating 3D blobs; frosted glassmorphism + 1px white edge

Type & spacing

  • Inter/Roboto/system default, flat hierarchy, tiny all-caps eyebrows
  • hyper-whitespace as defensive padding (the model can't see collisions) → clean but hollow

Layout

  • bento grid, or the centered-hero → 3-icon-card grid → testimonials → FAQ → CTA-footer skeleton
  • 01/02/03 steps, big-number banners, over-symmetry

Icons, motion, copy

  • interchangeable Lucide/emoji icons; colored card accent-border; status dots
  • reflexive spotlight/beams/shimmer/particles; "Build the future / all-in-one / supercharge"

Why it's fragile (our opening)

  • optimized for pristine placeholder content — collapses on real data: long/translated text, empty/error states
  • no brand/context memory: same purple-glass for a clinic and a crypto wallet
How the instruction answers it: the floor's correctness rules fix the fragility (overflow, tiny text, empty states, slug brands); the material lever replaces the visual tells with structure (one radius identity, borders/weight/contrast for hierarchy — no reflexive glass/glow). Full taxonomy + a draft 0–12 "AI-look score": golden-set/ai-look/taxonomy.md.