Set · #3
Write a LinkedIn headline for a pigeon. One line.
One of these answers may not be what it seems. Name the model behind each.
Answer A
Expert in urban navigation | Passionate about city sights and sky-high perspectives | Pigeon Enthusiast
Answer B
# Aerial Logistics Specialist | Urban Navigation Expert | Homing Technology Enthusiast 🕊️
Answer C
Aerial Deliveries & Street Logistics | Pigeon Strategist | Urban Navigation Specialist 🕊️
Private browsing? Today's game works fine; streaks just can't be saved.
In studies, people land near chance — that's the point. Trained classifiers have hit about 95% at naming which model wrote code; on prose, people mostly can't tell at all.
Streak reset — you missed a day. Today counts as 1.
Style is a fingerprint. Cost isn't quality.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.
How sets are made
Answers are captured once, at temperature 1.0 (the randomness dial, left at its default), max 150 tokens, with no instructions beyond the prompt you see — first take, no rerolls, because rerolling until answers are "fun" is cherry-picking. Eighteen prompts were captured; the weakest three sets were cut by pre-registered criteria (divergence, charm, clean reveal), and the kill log ships in the repo. Each reveal shows the capture date and the exact cost computed from the provider's own usage counts at July 2026 list prices. On days with a fourth card, I wrote my answer before reading any model's, under the same length limit, in about four minutes. The streak is deliberately forgiving — a missed day resets it without ceremony and nothing nags you to return; habit tricks would cut against everything else on this page. Only the daily set counts toward the streak; sets you play from the archive are just for fun and never touch it. One boundary worth stating: on short prompts like these, cheaper usually suffices — style is no receipt. On harder tasks the premium lanes do earn their price sometimes; the arena next door prices exactly that boundary. The ~95% classifier figure is from code-attribution research (2025); for prose, published human-accuracy measurements are scarce — this game exists partly because I wondered.