Set · #11
Describe a handshake to someone who has never had hands. Two sentences.
One of these answers may not be what it seems. Name the model behind each.
Answer A
A handshake is like a gentle meeting of two surfaces, where the palms come together briefly, often accompanied by a slight grasp or pressure. It’s a way people connect, share respect, or say hello without words.
Answer B
A handshake is when two hands reach toward each other, briefly touch, and then turn or squeeze in a small rhythmic way. It’s a quick signal that both people recognize one another and agree to greet each other peacefully.
Answer C
A handshake is when two people briefly grasp each other's right hands and move them up and down together—it's a gesture of greeting, agreement, or respect that involves gentle squeezing and movement. The physical contact and mutual grip create a momentary connection between two people, communicating friendliness or sealing a deal through touch.
Private browsing? Today's game works fine; streaks just can't be saved.
In studies, people land near chance — that's the point. Trained classifiers have hit about 95% at naming which model wrote code; on prose, people mostly can't tell at all.
Streak reset — you missed a day. Today counts as 1.
Style is a fingerprint. Cost isn't quality.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.