AI · Opinion
Anthropic's own postmortems are the best case for running your own daily evals
Labs keep changing what sits behind a model name and their own evals keep missing it. If you ship on a closed model, you need a daily regression rig that watches output tokens.
On 4 March 2026 Anthropic changed Claude Code’s default reasoning effort from high to medium, to cut latency. The model names stayed exactly the same. Sonnet 4.6 and Opus 4.6 ran at the lower setting for about five weeks, until Anthropic reverted the change on 7 April and conceded in a postmortem that it had picked the wrong tradeoff. Nine days later it added a system prompt instruction capping responses at 25 words between tool calls. According to Fortune, Anthropic said that one measurably hurt coding quality, and it was pulled after four days. If your startup’s agent ran on Claude Code that spring, your product got worse twice while your config file never changed.
That is why I think anyone building on a closed model should be running daily regression evals of their own, and why I spent a morning reading livenerf, a small repo by a developer called ninjahawk that started tracking Claude Opus 5.5 about 2.5 days after its 22 September launch. I expected a vibes machine with a chart. What I found was a more careful instrument than most companies run on their own production dependencies.
The design starts by throwing almost everything away. Of 2,336 GPQA Diamond, MMLU-Pro, competition-math and AIME questions, 97% were either always right or always wrong across four samples, so they can’t show movement. The 78 that Opus 5.5 only sometimes gets right form the panel. Items are paired against their own launch-week baseline, grading is exact match with no LLM judge, the Claude Code CLI is pinned to one version, and the decision rule was committed to git before any series data existed. The whole thing costs about 3.6% of a Max plan’s weekly allowance and can detect an accuracy shift of roughly 7.5 points per 10-day window.
The result that surprised me came from its validation run, where the author deliberately degraded the model. Dropping effort to low cut output tokens by 62% and accuracy by 8.3 ± 4.5 points. Dropping it to medium, the same change Anthropic shipped in March, cut tokens by 26% while accuracy fell 4.2 ± 3.9 points, an interval that brushes zero. So the March change moved output tokens by a quarter while the accuracy drop stayed inside the error bars. Token counts are also free: every API response already carries them.
The repo is honest about its limits too. Swapping in the older Opus 5 produced −3.8 ± 6.3 points, which the rig could not distinguish from Opus 5.5 at 99% confidence, though tokens dropped 23%. A same-family model swap can hide inside the noise of 78 questions. If a hobbyist with public benchmarks can’t catch it on accuracy alone, a startup with no rig at all has no chance.
The serious objection is that the labs mostly aren’t doing this on purpose. Anthropic wrote after its August–September 2025 incidents: “We never reduce model quality due to demand, time of day, or server load.” Those incidents were three infrastructure bugs. And the most famous drift study, Chen, Zaharia and Zou’s finding that GPT-4 fell from 84% to 51% on prime identification between March and June 2023, was picked apart by Arvind Narayanan and Sayash Kapoor, who pointed out that every number in the test set was prime.
I accept both points, and neither helps you. Your users don’t care whether degradation was a policy or a bug. In that 2025 episode, Anthropic says about 30% of Claude Code users who made requests had at least one message routed to the wrong server type, and at the worst hour on 31 August one bug misrouted 16% of Sonnet 4 traffic. Anthropic also said its own evaluations did not capture what users were reporting, and that privacy controls kept engineers from inspecting the failing conversations. OpenAI said the same about its offline evals when a 25 April 2025 GPT-4o update made ChatGPT sycophantic for four days. Both labs tested on their own prompts, and neither had the traffic of the companies building on their models. As for the prime-number fiasco, I read it as the argument for doing monitoring with paired items, exact graders and a pre-registered rule, which is what livenerf does.
So build the boring version for your own product. Pull a panel of prompts from real traffic where the model is right some of the time, grade them with code, run them every day through the exact client and settings you ship, and run an older model alongside as a control so a platform change doesn’t masquerade as a model change. Then put median output tokens per call on the same dashboard as latency and alert when it falls by a quarter. Had anyone set that alert in March 2026, it would have fired within a day of the effort change, roughly five weeks before Anthropic’s revert.
Prompted by GitHub - ninjahawk/livenerf: Benchmark for tracking model capability after release., GitHub.