bookmarks
-
METR
organizationIndependent evaluation by a nonprofit that takes no payment for it. A counterweight to labs grading their own homework.
Their time horizon metric measures how long a task a model can finish reliably, doubling roughly every seven months. Same group ran the 2025 trial where experienced developers were 19% slower with AI while believing they were faster.
Nothing matches that search.