Menu

Post image 1
Post image 2
Post image 3
Post image 4
Post image 5
Post image 6
Post image 7
Post image 8
Post image 9
Post image 10
Post image 11
Post image 12
Post image 13
Post image 14
Post image 15
Post image 16
Post image 17
Post image 18
Post image 19
Post image 20
1 / 20
0

GitHub - ninjahawk/livenerf: Benchmark for tracking model capability after release.

Hacker News·about 16 hours ago
#Hu760dwI
Reading 0:00
15s threshold

A long-running, deterministic-as-possible benchmark for detecting whether a frontier model gets quietly worse after launch. 📋 The plan · 📊 Results · 🔬 How it works · 🧪 Pre-registration livenerf is a small, boring, append-only benchmark for one question: does a model get worse after it ships? For months there have been reports that Anthropic "nerfs" models some days or weeks after release. That could mean quantization, a smaller model behind the same name, lower effort, or routing changes. It could also mean nothing happened and people are pattern-matching on noise. Nobody has had a clean day-0 baseline to check against, so every argument ends up as vibes versus vibes. Claude Opus 5.5 came out on 2026-09-22, so this is a chance to start the clock on launch day and keep it running. Right now v0 runs on a Claude Max subscription through headless Claude Code ( claude -p ), with no API key. You can't make these models deterministic: sampling params are gone and thinking can't be turned off.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More