Menu

Post image 1
Post image 2
1 / 2
0

Agent compute drops substantially with online skill distillation and graph‑guided knowledge

DEV Community: machinelearning·Papers Mache·4 months ago
#0vxLYEVn
#dev#agents#success#graph#distillation#token
Reading 0:00
15s threshold

Online skill distillation and graph‑guided knowledge substantially reduce the compute bill of LLM agents while keeping success rates competitive. These reductions could enable shifting agents from cloud‑only services to modest on‑device hardware. Earlier web and mobile agents leaned on heavyweight tricks: multi‑rollout searches, separate verifier passes, and stacks of specialist vision‑language models that balloon token counts and memory footprints. Such pipelines achieved high task‑success but only by paying huge inference costs. PANDO reaches a 58.3 % success rate on the full 910 VisualWebArena suite while using 115 K tokens per task, and it does so “while using 58% fewer tokens than SGV and 61% fewer than WALT, with no pre‑evaluation discovery budget” [1] . This headline figure shows that a single‑rollout, online distillation loop can surpass strong baselines without the token overhead that traditionally powers web agents.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More