Menu

Post image 1
Post image 2
1 / 2
0

Agents still fail 38% of real CLI tasks

DEV Community: machinelearning·Papers Mache·4 months ago
#V7XEnzYK
Reading 0:00
15s threshold

State‑of‑the‑art agents succeed on just 62.5 % of authentic command‑line workflows. The TerminalWorld benchmark, built from tens of thousands of real developer recordings, evaluates agents in a zero‑shot setting on tasks that span simple one‑liners to multi‑step deployment pipelines. That success ceiling shatters the prevailing belief that large language models can already replace shell scripts for everyday use. Existing evaluations have leaned on hand‑crafted command suites that capture only a narrow slice of developer activity. Benchmarks such as Terminal‑Bench present curated queries and score agents on idealized subtasks, but they miss the messy, iterative patterns seen in production terminals. Consequently, reported numbers have long over‑estimated practical reliability.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More