A Slashy research project · open source

Which AI writes the least like AI?

Every frontier model can write. The real question is whether it writes like a person, or like AI slop. We ranked 18 of them against real writing that predates ChatGPT, and by blind human vote.

loading the crowd vote… ·18 models ·112 scenarios ·19,928 generations
The board · live

Which model writes the least like AI?

Ranked by Human Score (0–100, higher = writes more like a person). It blends the live human vote (40%) with four mechanical axes at 15% each, and moves as votes come in.

loading the board…
Full rankings, every axis & the price scatter →

The thing we measure

You already know it when you read it.

Same request, two replies. One reaches for the reflexes: the throat-clearing opener, the tidy contrast, the eager sign-off. The other just answers.

Reads like AI

Hi Sarah, I hope this email finds you well! I wanted to reach out and touch base about the timeline. It's not just about hitting the date — it's about doing it right. Please don't hesitate to let me know if you have any questions. Looking forward to collaborating!

Reads like a person

Sarah, can we push a week? I'm still waiting on the design files and don't want to rush QA. Works either way, just let me know.

The highlighted phrases are among the tells the index counts, each measured against how often real people actually used it before 2022, not against a model's guess.

Good writing disappears. You come away thinking about the person, not the machine they wrote with.
The thesis behind The Slop Index
How it works

No LLM judges. No vibes. Just measurement.

Every number is a mechanical measurement against human writing that provably predates generative AI, plus a live crowd vote. Open method, open data.

  1. 01

    Start from pre-AI human writing

    Slop is distance from genuine human text collected before ChatGPT existed: Enron email, a blog corpus, student essays, archived tweets, Discord chat. No model could have shaped the yardstick.

  2. 02

    Score five axes, blend into one

    Human vote 40% plus conciseness, templating, rhythm and tells at 15% each. Every axis is normalised to a 0–100 human-likeness scale before blending.

  3. 03

    Run every model on the same task

    All 18 models ran the identical 112 scenarios across email, social, essays and chat, five samples each: 19,928 real, unedited generations. Default settings, no cherry-picking.

  4. 04

    Let humans break the ties

    The companion arena turns blind pairwise picks into a live Elo. That crowd verdict is 40% of the score, and the reason the board keeps moving.

The arena

Think you can spot the slop?

The board is the machine's verdict. The arena collects yours: two blind samples from the same task, one sloppier than the other. Flag it. Every vote nudges the live Elo that powers 40% of the score.

Play a round

Built by people who read a lot of email.

The Slop Index is a research project from Slashy, the email client that saves you time instead of generating more AI slop. See who writes clean, then come see what we build.