Menu

#Task

329 posts

Feed·
20 of 329 posts
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
🖼️
0

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

#arxiv#policy#task#tasks#model#agent

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows.…

15s
Read More
Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.
🖼️
0

Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

#fireworks#fable#model#task#cost#tasks

Benchmarking over 1,000 agentic tasks demonstrates that routing between the open-source Kimi K3 model and the closed Fable 5 model creates a new state-of-the-art (SoTA) approach that surpasses the performance of either model alone.…

15s
Read More
IonStack part II: GhostLock, a stack-UAF that has existed in ALL Linux distributions for 15 years
🖼️
0

IonStack part II: GhostLock, a stack-UAF that has existed in ALL Linux distributions for 15 years

#nebusec#waiter#stack#kernel#lock#task

GhostLock (CVE-2026-43499) is a Linux kernel vulnerability found by VEGA that exists in every major distribution since 2011. Triggering the bug does not require any special kernel config or privilege.…

15s
Read More
Paperclip (60K Stars) — The Open-Source OS Your AI Agents Have Been Waiting For
🖼️
0

Paperclip (60K Stars) — The Open-Source OS Your AI Agents Have Been Waiting For

DEV Community: productivity·龙虾牧马人·4 months ago
#HG6NfPJv
#dev#agent#paperclip#agents#task#github

Managing multiple AI agents is hard. Paperclip is the open-source OS that brings heartbeats, budgets, and task management to your AI team. 60K+ stars.

15s
Read More
Scoped Memory for Agent Systems: Cross-Run Persistence Without Global State
🖼️
0

Scoped Memory for Agent Systems: Cross-Run Persistence Without Global State

DEV Community: java·mgd43b·4 months ago
#CJythUJY
#dev#task#memory#research#agent#fullscreen

Task-scoped memory lets agents accumulate knowledge across runs through named scopes, pluggable stores, and explicit isolation boundaries.

15s
Read More