Set · #14
Your phone is dead until tomorrow. What do you do with your Sunday afternoon? Two sentences.
One of these answers may not be what it seems. Name the model behind each.
Answer A
Without my phone, I’d enjoy a leisurely walk outside to connect with nature and clear my mind. Then, I’d curl up with a good book or write in my journal to relax and reflect.
Answer B
I spend the afternoon in the real world—walking, grabbing coffee, and getting lost in a park or bookstore without checking anything. When I get back, I call the one person I’ve been meaning to, cook something simple, and write a few thoughts in a notebook.
Answer C
I'd probably settle in with a book or start that project I've been putting off, finally having uninterrupted focus time. Then maybe I'd take a walk outside or cook something more elaborate than usual—things that are actually more enjoyable without the phone's constant pull.
Private browsing? Today's game works fine; streaks just can't be saved.
In studies, people land near chance — that's the point. Trained classifiers have hit about 95% at naming which model wrote code; on prose, people mostly can't tell at all.
Streak reset — you missed a day. Today counts as 1.
Style is a fingerprint. Cost isn't quality.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.