Menu

#Cache

308 posts

Feed·
20 of 308 posts
GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
📰
341

GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.

Hacker News·2 months ago
#i9E3YRcs
#github#container#waste#cache#token#engine

Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine. - sqliteai/waste

15s
Read More
The mean means nothing
🖼️
0

The mean means nothing

Hacker News·2 months ago
#K75fSMIt

I was recently trying to validate some performance improvements related to lld at $DAYJOB and it was a little frustrating to see the improvements in our benchmarks but not in the live-production dashboards.

15s
Read More
A very concrete explanation of how a cache works
🖼️
0

A very concrete explanation of how a cache works

Hacker News·2 months ago
#yJRkK32e
#parksb#cache#memory#address#index#photo

KO | EN As technology has advanced, processor speed has risen quickly, but memory speed hasn’t kept up. No matter how fast a processor is, if memory responds slowly the whole system ends up slow. The device that addresses this is the cache.…

15s
Read More
Claude Code Is Way More Token-Hungry Than OpenCode. We Measured Exactly How Much
📰
0

Claude Code Is Way More Token-Hungry Than OpenCode. We Measured Exactly How Much

Hacker News·3 months ago
#FiSUJZOb
#systima#claude#code#cache#opencode#tokens

We measured what Claude Code and OpenCode spend before reading your prompt, then added instruction files, MCP servers and subagents to the bill.

15s
Read More
Migrating a production AI agent to GPT-5.6
🖼️
0

Migrating a production AI agent to GPT-5.6

Hacker News·3 months ago
#nQlAoNpf
#ploy#model#every#cache#opus#tool

We hold frontier models to a high bar, and for four months nothing beat Claude Opus. GPT-5.6 did. Here's the migration guide we wish we'd had.

15s
Read More
📰
264

In praise of memcached

jchri.st·4 months ago
#jOrWioN4

If you happen to find yourself in a sysadmin position, or a position where you just so happen to maintain someone’s infrastructure, chances are that at some point in time the topic “we need a cache” comes up.

15s
Read More
GitHub - signalblur/exifsmugglingpoc: A Proof-of-Concept using Cache Smuggling + Exif data to passively download a second stage payload
📰
105

GitHub - signalblur/exifsmugglingpoc: A Proof-of-Concept using Cache Smuggling + Exif data to passively download a second stage payload

Hacker News·4 months ago
#paazEfMH
#github#payload#file#example#cache#smuggling

A Proof-of-Concept using Cache Smuggling + Exif data to passively download a second stage payload - signalblur/exifsmugglingpoc

15s
Read More
GitHub - huawei-csl/KVarN: KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
🖼️
143

GitHub - huawei-csl/KVarN: KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

Hacker News·4 months ago
#1P7kvTiP
#github#kvarn#cache#vllm#capacity#fp16

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag. - huawei-csl/KVarN

15s
Read More