Menu

#Quantization

24 posts

Feed·
20 of 24 posts
Why your local LLM feels dumber than it is
🖼️
508

Why your local LLM feels dumber than it is

Quick Introduction We have all been on forums, chats, reddit, discord, youtube, or somewhere and heard “Oh! Model XYZ is AMAZEBALLZ!zomgwtfbbq” then downloaded it (or more likely, some quantized form of it) and said “eww… This sucks!” This post is going…

15s
Read More
GGUF Quantization Explained: Q4_K_M vs Q5_K_M vs Q8 — Which to Pick (2026)
🖼️
0

GGUF Quantization Explained: Q4_K_M vs Q5_K_M vs Q8 — Which to Pick (2026)

DEV Community·Patrick Hughes·4 months ago
#fqnJEpdi

Q4_K_M cuts model size 75% with minimal quality loss — but when should you use Q5, Q6, or Q8 instead? We benchmarked every quant level on real hardware and measured the actual accuracy tradeoffs.

15s
Read More
When I started running models locally, I thought quantization meant squeezing more into RAM. Turns o
🖼️
0

When I started running models locally, I thought quantization meant squeezing more into RAM. Turns o

DEV Community·Billy Bob Gurr·4 months ago
#ihdjkhty
#ai#llm#opensource#hardware#real#latency

From Dev.to - opensource: When I started running models locally, I thought quantization meant squeezing more into RAM. Turns o

15s
Read More