Product Technology Cases Company Team Careers Manifesto Blog Log in How we did it Prefill optimizations Takeaways Is memory the moat? Running Kimi K3 at ~952 tok/s/node, AMD continues to prove its case as the winner in performance per dollar. Over the past several months, we’ve seen an explosion in the capabilities of open source models. With DeepSeek V4-Pro and GLM5.2 reaching near-Opus levels of intelligence, open source has emerged as a real, cost-efficient alternative to the closed source models we’ve been married to. But we have yet to see one like Kimi K3. Promising Fable/Sol levels of intelligence, Kimi K3 marks the start of a new era for open source. But a smarter model means a bigger model — and these models are expanding in size just as fast as they are in capabilities. GLM5.2 has 753B parameters, DeepSeek V4-Pro 1.6T, and Kimi K3 weighs in at 2.8T (!!) parameters. That’s over 1.5TB of VRAM before allocating a KV cache for 1M tokens of context. Not even a B200 node (8 GPUs) can fit Kimi K3.…