bookmarks
-
Centaur Evaluations
projectNames the blind spot in every leaderboard: the model is scored alone, without the person who will actually use it. Measure imitation and you build for replacement by default.
The unit of measurement should be the human plus AI team. Report human-only, AI-only, and centaur performance, and treat human minutes and tokens as separate inputs.
Nothing matches that search.