# Ben Schmidt > Software architect and founder building with AI agents daily. I write about the scarce skill of the moment: knowing what to ask for, and knowing when the confident answer is wrong. Co-founder and CEO of RoadBotics (acquired by Michelin), founder of HeyLoopy, EIR at CMU's Corporate Startup Lab. Every page on this site has a plain-markdown twin: append `index.md` to its URL. ## Start here - [The Shape of It](https://iambenschmidt.com/shape/): guides that teach a fundamental of how software and agents actually work, in plain terms, and hand you a tool to use it. - [Writing](https://iambenschmidt.com/essays/): essays on directing agents and building in the open. - [About](https://iambenschmidt.com/about/): who Ben is and what he is working on. - [Let's talk](https://iambenschmidt.com/lets-talk/): how to reach Ben. ## Writing - [A login double-submit race, and the test layers that let an LLM fix it safely](https://iambenschmidt.com/essays/building-in-the-open/login-double-submit-and-the-test-layers/): A mystery bounce on an OTP login, chased end to end: probing the error, building a mock to recreate it, red-bar tests, a two-layer fix, adversarial critic passes, and the CI gate. The three-line fix is the least of it. The discipline around the model is what actually solved it. - [When you remove a feature, invert its tests. Don't skip them.](https://iambenschmidt.com/essays/building-in-the-open/when-you-remove-a-feature-invert-its-tests/): Deleting a feature is a one-line change. Deciding what to do with the dozen tests that now fail is the real question, and the obvious answer is the wrong one. A reverse-TDD pass that keeps the suite honest about what is now true, and why it matters more once a model is doing the deleting. - [We Always Finish Each Other's ____](https://iambenschmidt.com/essays/building-in-the-open/we-always-finish-each-others/): The blank always fills with the most likely word. Prompt engineering is the plain act of telling a long enough story that the word you wanted becomes the likely one. You have done it since you learned to talk. - [Without Code Journals, Your AI Will Fail You](https://iambenschmidt.com/essays/building-in-the-open/without-code-journals-your-ai-will-fail-you/): A test proves the code still works. Only the journal remembers why you built it that way, and the why is the part your AI, and next-week you, cannot reconstruct. - [The Bottleneck Was Never the Machine](https://iambenschmidt.com/essays/leading-through-change/the-bottleneck-was-never-the-machine/): Electrification took roughly forty years to show up in the productivity numbers, because the gain came from reorganizing the factory, not from the dynamo. AI is the same story. The bottleneck is how fast people adapt, and that part is a choice. ## Guides - [How to read an AI grader's accuracy number (and know if it's any good)](https://iambenschmidt.com/guides/how-to-read-an-ai-grader-accuracy-number/): A vendor says their AI grades short answers with 0.9 correlation to humans. Is that good? By itself, that number tells you almost nothing. Here are the two checks that tell you whether an automated grader is actually competent, in plain terms, with a worked example. ## Oakmont Lab - [What is a reported 0.88 grading correlation worth? A reanalysis of the BeSTraP short-answer dataset](https://iambenschmidt.com/oakmontlab/experiments/2026-07-grading-robustness/): A pre-registered reanalysis of the publicly released BeSTraP dataset, which reports a fine-tuned GPT-2 grader at Pearson 0.88 with human short-answer grades. Using only the released data, we ask what that number is worth: it holds up on every mechanism the data lets us check, sits below human consistency, and cannot be recomputed from the release.