Menu

#Indirect

3 posts

Feed·
3 of 3 posts
Arc Gate —LLM proxy that hits P=1.00 R=1.00 F1=1.00 on indirect/roleplay prompt injection (beats OpenAI Moderation and LlamaGuard)
📰
0

Arc Gate —LLM proxy that hits P=1.00 R=1.00 F1=1.00 on indirect/roleplay prompt injection (beats OpenAI Moderation and LlamaGuard)

Reddit r/artificial·u/Turbulent-Tap6723·5 months ago
#1zux7TsO

Benchmarked on 40 out-of-distribution prompts, indirect requests, roleplay framings, hypothetical scenarios, technical phrasings. The stuff that slips past everything else.…

15s
Read More
How indirect prompt injection attacks on AI work - and 6 ways to shut them down
📰
0

How indirect prompt injection attacks on AI work - and 6 ways to shut them down

ZDNET·Written by·6 months ago
#TMHJrtrP
#arrow#xa0#menu#close#prompt#injection

Cybercriminals are tricking AI into leaking your data, executing code, and sending you to malicious sites. Here's how.

15s
Read More