Menu

#Cache

308 posts

Feed·
20 of 308 posts
GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
📰
0

GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.

Hacker News·Hacker News·about 2 months ago
#i9E3YRcs
#github#container#waste#cache#token#engine

Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine. - sqliteai/waste

15s
Read More
AMD’s Venice-X CPU launches in 2027 with 1152 MB of 3D V-Cache, 96 cores, and 5.15 GHz boost clock – Zen 6 CPU for high-performance computing comes with major pillars of Venice
🖼️
0

AMD’s Venice-X CPU launches in 2027 with 1152 MB of 3D V-Cache, 96 cores, and 5.15 GHz boost clock – Zen 6 CPU for high-performance computing comes with major pillars of Venice

Latest from Tom's Hardware ·Jake Roach·about 2 months ago
#xtvr6FUy
#tomshardware#venice#cache#cores#memory#photo

It’s the same cache amount and core count as Genoa-X, but with far faster cores.

15s
Read More
GitHub - signalblur/exifsmugglingpoc: A Proof-of-Concept using Cache Smuggling + Exif data to passively download a second stage payload
📰
105

GitHub - signalblur/exifsmugglingpoc: A Proof-of-Concept using Cache Smuggling + Exif data to passively download a second stage payload

Hacker News·Hacker News·3 months ago
#paazEfMH
#github#payload#file#example#cache#smuggling

A Proof-of-Concept using Cache Smuggling + Exif data to passively download a second stage payload - signalblur/exifsmugglingpoc

15s
Read More
GitHub - huawei-csl/KVarN: KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
🖼️
143

GitHub - huawei-csl/KVarN: KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

Hacker News·Hacker News·3 months ago
#1P7kvTiP
#github#kvarn#cache#vllm#capacity#fp16

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag. - huawei-csl/KVarN

15s
Read More
Intel Xeon 6+ ‘Clearwater Forest’ puts 18A in the data center with up to 288 cores, 576 MB of L3 cache — new Xeon 6990E+ is 30% faster per thread than 192-core AMD Epyc 9965, says Intel
🖼️
0

Intel Xeon 6+ ‘Clearwater Forest’ puts 18A in the data center with up to 288 cores, 576 MB of L3 cache — new Xeon 6990E+ is 30% faster per thread than 192-core AMD Epyc 9965, says Intel

Latest from Tom's Hardware ·Jake Roach·4 months ago
#lpVKutLh

Intel’s dense Xeon 6+ design marks the first time 18A is being deployed in the data center.

15s
Read More