Skip to main content
  1. Guides/

How to read an AI grader's accuracy number (and know if it's any good)

·4 mins·
Ben Schmidt, PhD
Author
Recovering brain scientist turned AI builder, writing on Human Acceleration: aiming AI at people to make them faster than the change coming for them, not to replace them.

A vendor tells you their AI grades short answers with a 0.9 correlation to human graders. Your team is about to trust it with real assessments. Is 0.9 good?

By itself, that number tells you almost nothing. Not because anyone is lying, but because a grading correlation is meaningless until you put it next to the right comparison. That comparison is easy to make and hard to forget once you have seen it, and there are two of them.

If you are the person watching AI start to write and grade the assessments you used to build by hand, squinting at an accuracy stat and wondering whether you are allowed to trust it, this guide is for you. You do not need the statistics, just the two questions to ask.

Check 1: read it against how well humans agree

#

Almost everyone skips this move. A grader is being scored on how well it agrees with human graders, but the humans do not perfectly agree with each other either. So there is a ceiling: no grader can agree with noisy human scores better than a perfect estimate of those scores would. If two humans only agree with each other at, say, 0.97, then 0.97 is roughly the best any grader could ever score against them.

So ask instead: how well do your human graders agree with each other? Read the grader against that.

On a real, published dataset (651 graded short answers we reanalysed, walked through in full in the experiment), that looks like this:

A real example · 651 short answers

0.88 looks near-perfect until you add the human ceiling, which shows it is good but not near-perfect.

The grader scored 0.88. A second human scored 0.97. The most anything could score is 0.99. Against those numbers, 0.88 is competent and a step below a human.

The move for you. When a grader reports a correlation, ask for the human-to-human agreement on the same task. Read the grader’s number against that ceiling, not against a perfect 1.0. A grader at 0.88 where humans hit 0.97 is useful and clearly sub-human. The same 0.88 where humans only hit 0.85 would be a different story.

You can check the run yourself. And notice the honest wrinkle: there is more than one way to define “how well humans agree,” and on this dataset a looser definition makes the grader look comparable to a human rather than below one. That is exactly why you ask for the ceiling instead of trusting the headline.

Check 2: is the score inflated by easy versus hard questions?

#

A grading number can flatter itself a second way. If some questions are much easier than others, a grader can look accurate just by learning which questions are easy, without ever telling good answers from bad ones within a question. Statisticians call this clustering inflation. It is real, and it is checkable.

So ask: do the questions differ a lot in difficulty, or do students differ a lot within each question? If students vary far more inside a single question than the questions vary from each other, the grader is being tested on the hard thing (telling answers apart) rather than the easy thing (telling questions apart), and the score is trustworthy on that axis.

On the same dataset, students varied about three times more within a question than the questions varied from each other, so the number is not a difficulty artifact. You do not have to run the model to ask this; you just have to know it is worth asking.

What these checks do and do not tell you

#

Be clear about what you have and have not learned, because this is where good intentions overreach. These two checks tell you about the grader as an instrument: whether it agrees with humans as well as the task allows, and whether that agreement is real rather than an artifact. That is genuinely worth knowing before you deploy one.

What they do not tell you is whether your learners learned anything. A grader can be excellent and your assessment can still be measuring the wrong thing, or your learners can ace it and forget it in a week. That is a different question, with a different method, and anyone who sells you grader accuracy as evidence of learning is selling you the wrong receipt. When you want to know whether the learning is working, measure the learning, on your people, over time. Hold the two apart and you will not get fooled in either direction.

So the practical version is short: ask for the human-to-human agreement, check whether students vary more than the questions do, and keep “is the grader good” apart from “did anyone learn.” Do that and the number on the vendor’s slide stops being scary and starts being something you can read.

The AI & Learning Field Report

Field notes on AI and how people actually learn.

What is real, what is hype, and what the evidence actually says. The reports land here first; subscribers get them in their inbox. No spam, no funnel theater.