Menu

#Llama

92 posts

Feed·
20 of 92 posts
📰
23

Ferrox: Building a Rust Inference Engine That Matches llama.cpp

Hacker News·about 1 month ago
#XqbMNFIk
#fratepietro#llama#ferrox#metal#gguf#models

Why I built Ferrox, a pure-Rust GGUF inference engine, and what it took to match llama.cpp’s performance on real models — with receipts, not vibes.

15s
Read More
Petals – Run LLMs at home, BitTorrent-style
🖼️
0

Petals – Run LLMs at home, BitTorrent-style

Hacker News·about 2 months ago
#pfO01z80
#petals#llama#falcon#fine#part#model

Run large language models at home, BitTorrent‑style Generate text with Llama 3.1 (up to 405B), Mixtral (8x22B), Falcon (40B+) or BLOOM (176B) and fine‑tune them for your tasks — us

15s
Read More
GitHub - marcelroed/gigatoken: Language model tokenization at GB/s
🖼️
0

GitHub - marcelroed/gigatoken: Language model tokenization at GB/s

Hacker News·about 2 months ago
#LrKukuZn

Language model tokenization at GB/s. Contribute to marcelroed/gigatoken development by creating an account on GitHub.

15s
Read More
How to Setup a Local Coding Agent on macOS
🖼️
475

How to Setup a Local Coding Agent on macOS

Hacker News·3 months ago
#dEOSDGqk
#ikyle#gguf#gemma#llama#models#model

Running Gemma 4 26B-A4B and Qwen3.6 35B-A3B locally with llama.cpp, MTP speculative decoding, multimodal support, and PI as a coding agent.

15s
Read More
How to Deploy Llama 3.2 with TensorRT-LLM + Quantization on a $14/Month DigitalOcean GPU Droplet: 3x Faster Inference at 1/95th Claude Cost
🖼️
0

How to Deploy Llama 3.2 with TensorRT-LLM + Quantization on a $14/Month DigitalOcean GPU Droplet: 3x Faster Inference at 1/95th Claude Cost

DEV Community: webdev·RamosAI·4 months ago
#5zhYpj9b
#dev#llama#fullscreen#tensorrt#import#article

⚡ Deploy this in under 10 minutes Get $200 free: https://m.do.co/c/9fa609b86a0e ($5/month...

15s
Read More
Llama Lounge 25 at Stanford, standing room only - Jeremiah Owyang
🖼️
0

Llama Lounge 25 at Stanford, standing room only - Jeremiah Owyang

Web Strategy by Jeremiah·jeremiah_owyang·4 months ago
#lVgX5jlw

Sold out and standing room only. This was Llama Lounge 25 (see Luma Cal), hosted at Stanford for the second year in a row. This was Llama Lounge 25, hosted at Stanford for the second year in a row.…

15s
Read More
🚀 Meta Just Killed Open Source Llama: Welcome to the 'Muse Spark' Era (And What It Means for Developers)
🖼️
0

🚀 Meta Just Killed Open Source Llama: Welcome to the 'Muse Spark' Era (And What It Means for Developers)

DEV Community·Siddhesh Surve·4 months ago
#K7QSLriS
#ai#webdev#typescript#meta#muse#spark

From Dev.to - typescript: 🚀 Meta Just Killed Open Source Llama: Welcome to the 'Muse Spark' Era (And What It Means for Developers)

15s
Read More
Why Local AI Should Be the Default for Developers in 2026
🖼️
0

Why Local AI Should Be the Default for Developers in 2026

DEV Community·pickuma·4 months ago
#ny0JGBq6
#ai#webdev#tutorial#productivity#local#model

The case for running models on your laptop instead of paying per-token API bills: where local AI (Ollama, LM Studio, llama.cpp) wins on cost, latency, and privacy, and where the cloud still earns its keep.

15s
Read More
How to Deploy Llama 3.2 90B with vLLM + Speculative Decoding on a $16/Month DigitalOcean GPU Droplet: 2.5x Faster Inference at 1/110th Claude Cost
🖼️
0

How to Deploy Llama 3.2 90B with vLLM + Speculative Decoding on a $16/Month DigitalOcean GPU Droplet: 2.5x Faster Inference at 1/110th Claude Cost

DEV Community·RamosAI·4 months ago
#pACPdpCa

From Dev.to - tutorial: How to Deploy Llama 3.2 90B with vLLM + Speculative Decoding on a $16/Month DigitalOcean GPU Droplet: 2.5x Faster Inference at 1/110th Claude Cost

15s
Read More