Menu

#Fp16

3 posts

Feed·
3 of 3 posts
GitHub - huawei-csl/KVarN: KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
🖼️
143

GitHub - huawei-csl/KVarN: KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

Hacker News·4 months ago
#1P7kvTiP
#github#kvarn#cache#vllm#capacity#fp16

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag. - huawei-csl/KVarN

15s
Read More
📰
0

Technical question about Mamba Selective Scan kernel and FP16/FP32 precision

Reddit r/learnmachinelearning·u/Dry-Trouble4373·5 months ago
#kGw7XvYx
#kernel#fp16#mamba#fp32#precision#article

I'm trying to evaluate the model's accuracy when all internal operations are strictly limited to **FP16**. However, I noticed that the `selective_scan` CUDA kernel seems to use **FP32 accumulators** by default.…

15s
Read More