Menu

Post image 1
Post image 2
Post image 3
1 / 3
0

Opus 4.8 tops the LLM leaderboard with 95% on skill evals

DEV Community·Tessl·4 months ago
#Jl8boEcH
#dev#skill#opus#model#composer#judges
Reading 0:00
15s threshold

We added Claude Opus 4.8 to our ongoing model benchmark. It scored 95% with skill context, which puts it 1.6 points above Opus 4.7 and 2.3 points above Cursor's Composer 2.5 Fast. It is also, by a meaningful margin, the slowest model we have tested. TL;DR Opus 4.8 scores 95% with skill context, taking the top spot from Opus 4.7. Its 81% baseline is the highest ever recorded in this benchmark, higher than every other model and remains top even when models run evals with skills loaded. All three independent judges agreed within two points, the tightest spread we have seen across nine models. Previous high-variance models swung over seven points between judges. On matched runs, Opus 4.8 takes roughly 671 seconds per eval. Composer 2.5 averages 327 seconds on the same pairs. Composer 2.5 Fast averages 215 seconds. How the benchmark works We test models against a set of engineering skills, each skill is a structured context document that tells an agent how to work correctly in a specific domain.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More