↓
Skip to main content
Ben Schmidt
Capability under change
Essays
The Lexicon
Oakmont Lab
About
Follow on LinkedIn
Essays
The Lexicon
Oakmont Lab
About
Follow on LinkedIn
Independent research imprint
Experiments
Experiments
July 2026
What is a reported 0.88 grading correlation worth? A reanalysis of the BeSTraP short-answer dataset
A pre-registered reanalysis of the publicly released BeSTraP dataset, which reports a fine-tuned GPT-2 grader at Pearson 0.88 with human short-answer grades. Using only the released data, we ask what that number is worth: it holds up on every mechanism the data lets us check, sits below human consistency, and cannot be recomputed from the release.
↑