A long-running, deterministic-as-possible benchmark for detecting whether a frontier model gets quietly worse after launch. 📋 The plan · 📊 Results · 🔬 How it works · 🧪 Pre-registration livenerf is a small, boring, append-only benchmark for one question: does a model get worse after it ships? For months there have been reports that Anthropic "nerfs" models some days or weeks after release. That could mean quantization, a smaller model behind the same name, lower effort, or routing changes. It could also mean nothing happened and people are pattern-matching on noise. Nobody has had a clean day-0 baseline to check against, so every argument ends up as vibes versus vibes. Claude Opus 5.5 came out on 2026-09-22, so this is a chance to start the clock on launch day and keep it running. Right now v0 runs on a Claude Max subscription through headless Claude Code ( claude -p ), with no API key. You can't make these models deterministic: sampling params are gone and thinking can't be turned off.…