bookmarks
-
Management framing, not technical. Useful when the question is how an organization governs AI rather than how a model works.
Three interlocking properties to manage: autonomy, learning, inscrutability. The frontier keeps moving, so today's AI becomes tomorrow's ordinary software.
-
The architecture under everything current. Read it once and the vocabulary in every later paper becomes legible.
Vaswani et al., 2017. Drop recurrence and convolution, keep attention. Trained faster, parallelized better, beat the state of the art in translation.
-
OpenAI Academy
courseThe ChatGPT counterpart to Anthropic Academy. Keeping both shows how each lab teaches people to use its own tools, and where the framings diverge.
Three tracks by experience level: AI Foundations, Applied AI Foundations, Agents and Workflows. Live sessions with replays on top.
-
METR
organizationIndependent evaluation by a nonprofit that takes no payment for it. A counterweight to labs grading their own homework.
Their time horizon metric measures how long a task a model can finish reliably, doubling roughly every seven months. Same group ran the 2025 trial where experienced developers were 19% slower with AI while believing they were faster.
-
Anthropic Academy
courseFirst-party training, free, kept current. Beats secondhand tutorials that go stale in a month.
Courses run on Skilljar with certificates. AI Fluency for working practice, API and MCP and Claude Code tracks for building.
-
Centaur Evaluations
projectNames the blind spot in every leaderboard: the model is scored alone, without the person who will actually use it. Measure imitation and you build for replacement by default.
The unit of measurement should be the human plus AI team. Report human-only, AI-only, and centaur performance, and treat human minutes and tokens as separate inputs.
-
APEX Benchmarks
toolMost benchmarks test exam questions. APEX tests paid professional work: investment banking, corporate law, consulting, medicine, software engineering, with tasks written and graded by people who do those jobs.
On long multi-hour agent tasks the leaders sit near 43%. Impressive on single-turn text, nowhere near reliable on real workflows.
-
Measured use, not opinion. 500k real coding conversations, split by whether the human stayed in the loop or handed the task over. It gives you a vocabulary (automation vs augmentation) for deciding how much of your own work to delegate.
79% of Claude Code conversations were automation, against 49% on the chat interface. Front-end work dominates, which suggests UI-heavy roles feel the shift first.
Nothing matches that search.