The AI Vibe Check Study Asks What Makes AI Writing Sound Human
By Pietro Romeo · Published
Abstract. AI Vibe Check asks readers which of two AI answers to the same prompt sounds more human, and what made the other one fall short. Every round shows two answers, both written by AI, in random order. Readers place their pick on a seven-point scale and can mark what put them off in either answer. Results for each round are published as aggregate counts 7 days after the round first runs, and findings across rounds will follow once there are enough answers to support them.
The study asks which of two AI answers sounds more human
AI models are tuned in part on human judgments of how their answers read, and the companies building them collect those judgments in private. AI Vibe Check collects a small public version of the same judgment: short answers to everyday prompts, read by people who choose to play, with the results published in the open.
Both texts in a round come from AI models, so there is no correct pick. We look at where readers land, how much they agree with each other, and what they say when an answer puts them off.
Each round asks for one pick on a seven-point scale
A round starts with a prompt. On the study's first day it was a small, everyday task: write a text to a flatmate about dishes left in the sink, without starting a fight. When you press "Show answers", two puppets called Calca and Brina type out one answer each.
Calca always sits on the left and types first. Which of the two answers Calca gets is random for each reader, so the same text can be Calca's for you and Brina's for the next person. On a phone, the two answers stack one above the other.
Once both answers are on screen, you place your pick on the scale below them. There are three steps toward each answer, from "slightly" to "much" more human, and "No difference" in the middle. Nothing is saved until you press Submit.
After you pick, you can mark what made either answer sound less human, or both, or write it in your own words. This step is optional. When you save, you see how other readers have split on the same round so far.
Neither the game nor the results pages say which models wrote the answers. You can also suggest a prompt for a future round.
Prompts ask for the kind of writing people do every day
Most rounds ask for short, ordinary writing: a message to a friend or flatmate, a piece of advice, an opinion, a story about something that happened to you. Some rounds show a photo in place of a written prompt. We write most prompts ourselves, and players can suggest new ones. Every suggestion is read before it is scheduled, and the round shows who suggested it.
Many of our readers speak English as a second or third language, so prompts use plain English, short sentences and no idioms. We also try to avoid prompts that depend on knowing one country's jokes, shows or places. A reader who has never seen a British sitcom would partly be judging their own familiarity with it, and that is not what the round asks.
Each answer comes from a different AI model
We send each prompt to AI models from several companies and pair two of the answers for a round. The pairings vary, so over a season each model meets several others. Which model wrote which answer is stored on our server and never sent to your browser, so it cannot be found by looking at the page.
Two answers in a round can differ in length, tone and layout as well as in how human they sound, because each model writes in its own way. We record those differences for every round so the analysis can take them into account. If we ever correct the text of a round, picks made before and after the change are kept apart and never counted together.
The middle of the scale and the optional tags are deliberate choices
Some pairs are too close to call, and the fair answer is "No difference". A scale with no middle would force a guess on those rounds and produce a preference that is not there. How often readers choose the middle is a result in its own right, because it shows how hard a pair was to tell apart. We count those picks as answers in every analysis.
Tags are optional so that a round stays quick to play. The cost is that a blank tag list can mean "nothing stood out" or "I did not look". To tell those apart, we record whether the tag list was on your screen as well as the tags you saved.
We record your pick, your tags and a few details about the session
For each round we store where you placed your pick, any tags and free text you added, how long you spent reading and deciding, and whether you played on a phone or a larger screen. A random ID stored in your browser groups your picks together. We also record the country your connection comes from and your browser's language setting. Section 3.8 of our terms lists every item and how each one is used.
Results are published per round, then as findings across rounds
- Each round gets a public results page 7 days after it first runs, showing how readers split and what they tagged.
- If you play a round more than once, only your first pick counts.
- A round with fewer than 10 answers shows counts only.
- Published results are aggregate and never identify a reader. Free-text answers may be quoted without a name.
- Findings across rounds will appear in our research section with sample sizes, what was excluded and why, and the limitations below.
Some picks are left out of any analysis: picks made too fast for the text to have been read, automated traffic, and our own test plays. We fixed and dated these rules before the first answer came in, and will publish them in full with the first findings. Each finding will report how many picks each rule removed, and we will check that the main result holds under a stricter and a looser version of the reading-speed rule.
The analysis treats each round as one shared experience
Twenty people who play the same round are all reacting to the same two texts. Their picks tell us a lot about that pair and much less about AI writing in general. Treating them as twenty separate pieces of evidence would make weak patterns look certain, and the error grows as more people play. Our analysis groups picks by round for this reason, and every share we publish comes with the number of answers behind it.
The list of tags may change during the study. Each version of the list is numbered, and tag counts from different versions are never added together.
We are keeping what we expect to find to ourselves until then. The people who read this page also play the game, and knowing what we look for could change how they read the next round.
These limits apply to everything the study publishes
- It compares AI with AI. Both answers in every round are machine-written, so the results can rank two AI answers against each other. They cannot show whether AI writing passes for human writing.
- The readers choose themselves. People who play are visitors to this site who opted in. They do not represent any wider population.
- English background varies. Readers come from many countries, and many will not have grown up speaking English. How a piece of writing reads depends a great deal on the English you know best.
- We and our players choose the prompts. Rounds are written by us or suggested by players, so they are a curated set of writing, and a result from one kind of prompt may not carry over to another.
- The two answers differ in more than one way. Answers in a round can differ in length and layout as well as voice, so a single round cannot show which difference decided the pick. Patterns across many rounds can suggest it, and we will say how sure we are.
- Side and typing order go together. The answer on the left always types first and always belongs to Calca. The answers are assigned to sides at random, so this does not favor either answer, but a preference for the left side would mix position, order and puppet into one effect.
- People play wherever they are. Readers use their own phones and computers at a time of their choosing, so conditions vary far more than in a lab study.
You can play without an account and have your picks removed
We record the country only, never your IP address or any more precise location. Your name and email are attached to your picks only if you are signed in. To have your picks removed, email privacy@aitrainer.work. Our privacy policy covers the rest of the site.
Play a round yourself
Read two AI answers, pick the one that sounds more human, then see how everyone else heard it.
Play AI Vibe Check