Every frontier model can write. The real question is whether it writes like a person, or like AI slop. We ranked 18 of them against real writing that predates ChatGPT, and by blind human vote.
Ranked by Human Score (0–100, higher = writes more like a person). It blends the live human vote (40%) with four mechanical axes at 15% each, and moves as votes come in.
Same request, two replies. One reaches for the reflexes: the throat-clearing opener, the tidy contrast, the eager sign-off. The other just answers.
Hi Sarah, I hope this email finds you well! I wanted to reach out and touch base about the timeline. It's not just about hitting the date — it's about doing it right. Please don't hesitate to let me know if you have any questions. Looking forward to collaborating!
Sarah, can we push a week? I'm still waiting on the design files and don't want to rush QA. Works either way, just let me know.
The highlighted phrases are among the tells the index counts, each measured against how often real people actually used it before 2022, not against a model's guess.
Good writing disappears. You come away thinking about the person, not the machine they wrote with.The thesis behind The Slop Index
Every number is a mechanical measurement against human writing that provably predates generative AI, plus a live crowd vote. Open method, open data.
Slop is distance from genuine human text collected before ChatGPT existed: Enron email, a blog corpus, student essays, archived tweets, Discord chat. No model could have shaped the yardstick.
Human vote 40% plus conciseness, templating, rhythm and tells at 15% each. Every
axis is normalised to a 0–100 human-likeness scale before blending.
All 18 models ran the identical 112 scenarios across email, social, essays and chat, five samples each: 19,928 real, unedited generations. Default settings, no cherry-picking.
The companion arena turns blind pairwise picks into a live Elo. That crowd verdict is 40% of the score, and the reason the board keeps moving.
The board is the machine's verdict. The arena collects yours: two blind samples from the same task, one sloppier than the other. Flag it. Every vote nudges the live Elo that powers 40% of the score.
Play a round →The Slop Index is a research project from Slashy, the email client that saves you time instead of generating more AI slop. See who writes clean, then come see what we build.