Has Opus 5.5 Been Nerfed Yet?
For months people have claimed that Anthropic quietly nerfs Claude models days or weeks after release. The evidence was always vibes versus vibes: nobody had a clean day-0 baseline to check against [1].
Now someone does. Livenerf is an append-only benchmark that started the clock on Claude Opus 5.5 at launch day and runs once a day for 30 days. Day 6 of 30 is collected, all on the same harness hash and pinned CLI. The baseline is 6 of 10 days complete, none missed [1].
The methodology is the real story. You cannot make these models deterministic: sampling parameters are gone, thinking cannot be turned off. So livenerf makes everything else deterministic: frozen prompts, pinned CLI, exact graders, raw logs forever. It then measures drift statistically over thousands of samples using Inspect, the UK AI Security Institute's eval framework. The stats follow Anthropic's own paper on adding error bars to evals [1].
The panel was selected from 2'336 GPQA Diamond, MMLU-Pro, competition-math and AIME 2025-26 questions. Opus 5.5 gets about 93% right on the first try. 97% of questions were always right or always wrong. The 78 questions that are sometimes right form the panel [1].
What can it detect? One run a day detects an accuracy change of about 7.5 points per 10-day window. Validation passed its pre-registered criterion. Lower effort shows up clearly in tokens: effort low drops output tokens by 62% and accuracy by 8.3 points, effort medium drops tokens by 26% and accuracy by 4.2 points [1].
The limit is honest: swapping in Opus 5 for Opus 5.5 was not distinguishable at 99% confidence. A same-family model swap of that size cannot be detected in a validation's worth of samples [1].
Day 1 was September 24, about 2.5 days after launch. The first possible call is around October 24, after two 10-day windows. So far the baseline is holding [1].
This is how you do it. No drama, no Twitter threads, just a pre-registered protocol, frozen prompts, and statistical rigor. If Opus 5.5 gets nerfed, the numbers will show it. If it does not, the numbers will show that too.
Sources:
[1] Livenerf - Benchmark for tracking model capability after release
[2] Hacker News discussion (400 points, 161 comments)