Set · #5
Explain where the missing dryer socks go, in the voice of a conspiracy theorist. Three sentences max.
One of these answers may not be what it seems. Name the model behind each.
Answer A
# The Sock Disappearance Listen, they don't want you asking questions about where those socks vanish to, but I've connected the dots—there's a parallel dimension accessible only through high-heat dryer cycles, and the government has known about it for DECADES. The dryer manufacturers are IN ON IT, deliberately calibrating their machines to open portals and harvest our socks for reasons they're not disclosing. Wake up sheeple—your missing left sock is a breadcrumb leading to the truth!
Answer B
You see, those missing dryer socks aren’t just lost—they’re being covertly harvested by the appliance companies, collecting our tiny, private secrets for some larger agenda. Every time you think your sock has vanished, it’s actually being absorbed into a hidden underground network, fueling a secret data bank aimed at monitoring consumer habits. It’s all a smokescreen—those socks vanish so they can secretly track us through invisible signals bouncing off our laundry!
Answer C
Oh, don’t you *dare* tell me you haven’t noticed—the missing dryer socks don’t vanish, they get quietly abducted by a sock-dimension under your lint trap. The “static” you hear is just the interdimensional routers warming up, beaming the odd ones into the Great Dryer Bermuda Triangle. And every time you swear you counted two, the dryer laughs—because one sock always gets recruited to keep the conspiracy running.
Private browsing? Today's game works fine; streaks just can't be saved.
In studies, people land near chance — that's the point. Trained classifiers have hit about 95% at naming which model wrote code; on prose, people mostly can't tell at all.
Streak reset — you missed a day. Today counts as 1.
Style is a fingerprint. Cost isn't quality.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.