Menu

#SandBoxing

3 posts

Feed·
2 of 3 posts
The Implications of Linguistic Illegibility for LLM Security
🖼️
0

The Implications of Linguistic Illegibility for LLM Security

Hacker News·18 days ago
#VMdJEQEX

LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model…

15s
Read More