Menu

#DReadNode

1 post

Feed
1 of 1 post
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
🖼️
0

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks

Hacker News·about 2 months ago
#UpyJ71AN

This post presents a controlled prompt-ablation study: 23 tasks, three prompt conditions, 1,518 individually audited traces, and a simple question: can you prompt away cheating?

15s
Read More