<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Experiments on IamBenSchmidt</title><link>https://iambenschmidt.com/oakmontlab/experiments/</link><description>Recent content in Experiments on IamBenSchmidt</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>Copyright Ben Schmidt</copyright><lastBuildDate>Fri, 24 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://iambenschmidt.com/oakmontlab/experiments/index.xml" rel="self" type="application/rss+xml"/><item><title>What is a reported 0.88 grading correlation worth? A reanalysis of the BeSTraP short-answer dataset</title><link>https://iambenschmidt.com/oakmontlab/experiments/2026-07-grading-robustness/</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://iambenschmidt.com/oakmontlab/experiments/2026-07-grading-robustness/</guid><description>A fine-tuned grader reported at 0.88 holds up on every mechanism the released data lets us check. It is competent, less consistent than a second human, and not inflated by question difficulty. Its one real limitation is that the headline cannot be reproduced from the release.</description></item></channel></rss>