Menu

#Tasks

303 posts

1 point
Feed·
20 of 303 posts
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
🖼️
0

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

#arxiv#policy#task#tasks#model#agent

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows.…

15s
Read More
Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.
🖼️
0

Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

#fireworks#fable#model#task#cost#tasks

Benchmarking over 1,000 agentic tasks demonstrates that routing between the open-source Kimi K3 model and the closed Fable 5 model creates a new state-of-the-art (SoTA) approach that surpasses the performance of either model alone.…

15s
Read More
Project Fetch: Phase two
🖼️
73

Project Fetch: Phase two

Hacker News·Project Fetch: Phase two·3 months ago
#i8Q9JHy6
#anthropic#claude#models#team#ball#tasks

We report results from our latest test of whether Claude can help Anthropic employees perform sophisticated robotics tasks. We found that Claude Opus 4.7, operating without human assistance, was about 20 times faster than the fastest human team at all…

15s
Read More