Loading...
Loading...
The detector gives you a number. Eighty-seven percent AI, maybe. Or maybe a clean pass. Either way, the score means nothing to you, because nobody tells you what produced it. Almost always, two figures sit behind that verdict, and the tool never shows them. Perplexity and burstiness. This guide is written for the person at the keyboard, not the person reading the scanner report. You will learn what the two figures track, where they pull in different directions, and how to read them on your own draft without handing your writing to a black box.
Here is the thing. Most explainers stop at "perplexity is word surprise, burstiness is sentence variation." True, and useless. Definitions alone do nothing for you at midnight when a paragraph reads flat and you cannot figure out why. So this guide picks a side. The writer's side. Two metrics, treated as a self-diagnostic you run on yourself, not a verdict handed down by a tool you cannot question.
Trickery is not the aim. Turning two abstract numbers into something you can actually use while you write, that is the aim.
Both terms come out of information theory. Same field that built data compression and entropy. They are older than ChatGPT by decades. Perplexity was originally a scoring method researchers used to check how well a language model predicted a sample of text, not a tool for policing essays [1]. The writing world picked them up only after detector companies turned them into headline signals.
Start with perplexity. What it tracks is the surprise a language model feels at your word choices, one word at a time. Low perplexity, and each word was the obvious next pick, the safe bet. High perplexity, and the model kept getting caught off guard, the path you walked was not the path it expected. Two sentences show the gap cleanly. "The cat sat on the mat" sits in low-perplexity territory, because every word is the safest continuation of the one before. "The cat ricocheted off the radiator and landed in my soup" sits in high-perplexity territory, because "ricocheted" and "radiator" and "soup" are not the path the model would have walked [2].
Burstiness is the other figure. Different scale. Where perplexity watches words, burstiness watches rhythm. What it tracks is the variation in your sentence length and structural shape across a passage. Low burstiness means your sentences hover around the same length, marching in step, all roughly the same weight. High burstiness means short punchy lines sit next to long winding ones, the way real speech breathes when someone is actually thinking. A paragraph of four sentences at 13, 12, 11, 11 words reads flat. A paragraph at 4, 19, 5, 36 words reads alive [2].
So the shorthand lands here. Perplexity asks whether your words are predictable. Burstiness asks whether your rhythm is monotonous. Together they explain most of what a detector reacts to when it spits back a verdict. But they measure different things, and writers who treat them as one signal end up editing the wrong dimension.
During training, a large language model has one job. Predict the next token. That single goal, repeated trillions of times, turns out prose that is statistically safe at every step. The model picks the most probable word because picking the most probable word is what it was built to do, full stop. That is why AI text lands low on perplexity by design, not as an accident [3].
Humans do not write that way. We get distracted mid-sentence. We reach for an odd metaphor. We change our mind about how long a sentence should be while we are still inside it. Those small deviations push perplexity up and break the rhythm into bursts.
Now the catch, and writers need to hear this. Perplexity is deeply content-dependent. A legal contract, a terms-of-service page, a methods section in a research paper, all of these are supposed to be predictable. They use controlled vocabulary and repeated structure because clarity rewards consistency. A human-drafted contract can clock lower perplexity than an AI-generated poem [3]. What the metric does is confuse "predictable structure" with "non-human origin," and those two things are not the same. That is why honest, careful writing in formal registers gets flagged.
A musical analogy holds up better than most. Think of a pop song next to a jazz solo. The pop hook is melodic and predictable, which is why you can hum it after one listen. The jazz solo keeps surprising you, which is why you cannot. AI text is the pop hook. Human text is the solo. Neither is inherently better writing, but detectors treat the solo as the human signature.
For people who want the numbers, the AI-typical ranges reported across detector documentation look roughly like this. Average perplexity between 5 and 30. Sentence-length standard deviation between 0.5 and 3. The human-typical ranges sit much higher, perplexity in the 50 to 200 band, sentence-length standard deviation from 5 to 20 [2]. Hard cutoffs, these are not. Detectors update their thresholds often, and a single paragraph is too small a sample to settle a verdict. But the gap between the two ranges is wide enough that the direction matters even when the exact numbers do not. For a detector-side view of how these signals get combined into a verdict, see our explainer on perplexity and burstiness as detection signals.
Here is the part most explainers skip. You do not need a detector to estimate these signals on your own draft. Two simple checks during editing will get you close.
For burstiness, the method is almost embarrassingly low-tech. Pick any paragraph. Count the words in each sentence. Look at the spread. If every sentence lands within two or three words of the mean, your rhythm is flat and an AI signature is showing through. If the counts swing wide, say 4, 22, 6, 31, you are writing in bursts. Craft-driven humanization research keeps landing on the same finding, deliberate sentence-length variation is the single highest-leverage fix writers can make, more impactful than swapping synonyms [4]. We go deeper on this lever in our guide to sentence structure variation for natural AI text.
Try it on two versions of the same idea. The flat version reads, "Climate change represents a serious global challenge. Rising temperatures alter weather patterns across continents. Coastal communities face increasing risks from flooding. Agricultural systems must adapt to changing conditions." Sentence lengths are 7, 7, 7, 8. Standard deviation near zero. A bursty version of the same content reads, "Climate change is here. Summers just keep creeping hotter, and the seasons your parents read about in textbooks are the ones you actually live through now. Towns flood more often. Farms cannot predict the rain anymore, and that matters because it changes what you eat, what you pay, and whether the system that feeds you still works in five years." Lengths are 4, 22, 4, 31. Standard deviation jumps past 12. Same information, different statistical fingerprint [2].
For perplexity, the test is intuitive. Read a sentence. Ask yourself whether you could have predicted the next word before you read it. If the answer is yes for most of the paragraph, your word choices are sitting on the safe path the model would have taken. If you keep running into words that feel specific to your intent, an unusual verb, a concrete noun, a phrase that belongs to you and not to the statistical average, perplexity is climbing.
Different writing scenes carry different baselines, and pretending otherwise causes most false alarms. Academic writing is supposed to be lower in perplexity because the register rewards precision and repetition, so a literature review will naturally read as more predictable than a personal essay. Marketing copy can afford higher burstiness because rhythm carries persuasion. A Reddit post lives on burstiness, the platform rewards punch. A lab report does not. Before you panic over a score, compare it against the natural range of your genre, not against a universal ideal.
The two figures do not always move together. Writers get into trouble when they assume they do.
High perplexity with low burstiness produces a failure mode you can spot a mile away. You stuffed your draft with unusual words, archaic verbs, rare adjectives, but every sentence is still the same length and shape. What you get reads as a thesaurus explosion. Detectors trained on adversarial examples flag this pattern, because real human high-perplexity writing almost always comes bundled with bursty rhythm, and strange words without rhythmic variation look like someone trying to game a signal [5].
Picture a student worried about a flag. Every "use" becomes "leverage," every "show" becomes "evince," every "important" becomes "consequential." The vocabulary climbs into rare territory, so perplexity rises. But the sentences still march in lockstep at 18, 19, 17, 18 words. What you get is a draft that reads as both pretentious and machine-like, worse than where it started. One signal pushed up, the other left flat, and the classifier notices the mismatch.
Low perplexity with high burstiness is the opposite case. Your words are plain and predictable, but your sentences vary wildly in length. This combination reads as more human than the first, because burstiness is the harder signal for a language model to fake without breaking fluency. If you have to choose one lever to pull as a writer, pull burstiness.
The deeper point is that perplexity and burstiness are two signals among at least five. Detectors also weigh token probability curves, stylometry, and a trained classifier that combines everything into a probability score [5]. Two paragraphs with identical perplexity and burstiness can still get different verdicts, because the classifier noticed a stylistic fingerprint the raw metrics miss. Treat the two figures as a useful diagnostic, not as the full verdict.
One myth says high perplexity means good writing. It does not. A draft can be unpredictable because the word choices are wrong, not because they are good. What perplexity measures is statistical surprise, not quality. A sentence that surprises the model by being incoherent scores high. Quality and surprise correlate in human writing, but the correlation breaks at the edges.
A second myth says burstiness is just alternating short and long sentences. That is the surface version. Real burstiness also includes variation in structural complexity, subordinate clauses, paragraph density, and register shifts. A piece can alternate 5-word and 25-word sentences and still feel mechanical if every long sentence follows the same template.
A third myth says detectors only look at these two metrics. They do not. The two are the most cited because they are the most intuitive, but every commercial detector layers a trained classifier on top, and that classifier picks up signals that are difficult to express in a single number. This is why the same text returns different scores across tools, the underlying reference models and classifier weights differ [5]. Some detectors measure perplexity against GPT-2, others against their own fine-tuned model, so the same paragraph can return low perplexity on one tool and high on another [3].
Perplexity and burstiness are tools for self-diagnosis, not verdicts from a court. The first tells you whether your words are sitting on the safe statistical path. The second tells you whether your rhythm is breathing. Read together, they explain most of what a detector is reacting to, and they give you something a verdict percentage never gives you, a place to edit.
The writers who get flagged least are not the ones who chase a score. They are the ones who write with awareness of how their choices land statistically, then adjust where the rhythm goes flat. That is the same work EvalHub does at scale, running perplexity analysis and burstiness scoring to turn the two numbers into a review-oriented rewrite that keeps your voice intact. For more on the craft side, our tips to make AI text sound natural cover the practical edits that move both signals. The metrics are not the enemy. Ignoring them is.
Perplexity tracks how surprised a language model is by your word choices, word by word. Burstiness tracks how much your sentence length and structural complexity vary across a passage. The first is about word predictability. The second is about rhythm.
You can estimate both without a detector. For burstiness, count the words in each sentence of a paragraph and check the spread. For perplexity, read a sentence and ask whether you could have predicted the next word before reading it. These checks are not precise, but they reveal whether your draft is sitting on the flat, safe path detectors associate with AI.
Burstiness is the lever writers can pull most reliably. Unusual word choices are easy to overdo and can read as a thesaurus explosion. Deliberate variation in sentence length is harder to fake and carries more of the human signature, which is why craft editors treat it as the highest-impact fix.
Most detectors use these two signals, but they are not the only ones. Detectors also weigh token probability curves, stylometry, and a trained classifier that combines everything into a final score. Two texts with identical perplexity and burstiness can still get different verdicts across tools because the classifier weights differ.
No. High perplexity from incoherent word choices does not help, and high burstiness from a formulaic template does not either. Detectors layer a classifier on top of the raw metrics, and that classifier picks up patterns the metrics miss. The goal is genuine variation tied to meaning, not statistical targets chased in isolation.
Humanize AI text to sound naturally human with EvalHub.
Start Free Trial