Menu

#Quantization

25 posts

Feed·
20 of 25 posts
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
🖼️
0

Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses

Hacker News·28 days ago
#hwNfzXDZ

I test Unsloth GGUFs of Qwen3.8 27B (Q4_K_M, UD-Q2_K_XL, UD-IQ1_S) with llama.cpp on GPQA Diamond, IFBench, Terminal-Bench 2.1. Q4_K_M matches BF16 abd fits an RTX 4090.

15s
Read More
Why your local LLM feels dumber than it is
🖼️
508

Why your local LLM feels dumber than it is

Hacker News·about 1 month ago
#WSJJzj1r

Quick Introduction We have all been on forums, chats, reddit, discord, youtube, or somewhere and heard “Oh! Model XYZ is AMAZEBALLZ!zomgwtfbbq” then downloaded it (or more likely, some quantized form of it) and said “eww… This sucks!” This post is going…

15s
Read More
Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency
🖼️
394

Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency

Hacker News·4 months ago
#6iu4Dyhz
#models#quantization#mobile#gemma#model#photo

We’re releasing Gemma 4 quantization-aware training checkpoints, reducing memory requirements and improving on-device performance.

15s
Read More
GGUF Quantization Explained: Q4_K_M vs Q5_K_M vs Q8 — Which to Pick (2026)
🖼️
0

GGUF Quantization Explained: Q4_K_M vs Q5_K_M vs Q8 — Which to Pick (2026)

DEV Community·Patrick Hughes·5 months ago
#fqnJEpdi

Q4_K_M cuts model size 75% with minimal quality loss — but when should you use Q5, Q6, or Q8 instead? We benchmarked every quant level on real hardware and measured the actual accuracy tradeoffs.

15s
Read More
When I started running models locally, I thought quantization meant squeezing more into RAM. Turns o
🖼️
0

When I started running models locally, I thought quantization meant squeezing more into RAM. Turns o

DEV Community·Billy Bob Gurr·5 months ago
#ihdjkhty
#ai#llm#opensource#hardware#real#latency

From Dev.to - opensource: When I started running models locally, I thought quantization meant squeezing more into RAM. Turns o

15s
Read More