Menu

#Ollama

137 posts

Feed·
20 of 137 posts
Why Your Local LLM Setup Is Costing More Than You Think — And What Happens When It Breaks
🖼️
0

Why Your Local LLM Setup Is Costing More Than You Think — And What Happens When It Breaks

DEV Community·xu xu·3 months ago
#XoJWHCQa
#dev#local#inference#model#ollama#teams

You're three hours into debugging a model quantization issue. The GPU utilization is sitting at 12%....

15s
Read More
📰
0

Google’s Gemma 4 12B just dropped - here’s how to run it locally on your Mac

Artificial Intelligence (AI)·/u/nullvector88·3 months ago
#WVkQBuNw
#artificial#ollama#model#macs#reddit#article

Google released Gemma 4 12B today. It’s a solid open-source model (Apache 2.0) that’s multimodal and runs really well on Macs with 16GB or more unified memory. Good at reasoning, coding, and agent stuff.…

15s
Read More
Running Brand-New Gemma 4 12B on an 8-Year-Old GTX 1080 Ti: Speed, 3 Gotchas, and Why Q8 Beat Q4 on My Own Field
🖼️
0

Running Brand-New Gemma 4 12B on an 8-Year-Old GTX 1080 Ti: Speed, 3 Gotchas, and Why Q8 Beat Q4 on My Own Field

DEV Community: machinelearning·byeongsoo kang·3 months ago
#sTOMC2Sk
#dev#ollama#card#model#gemma#second

I pulled the just-released Gemma 4 12B and ran it on a GTX 1080 Ti. ~28 tok/s at Q4 on one card — but three things broke first, and Q8 (2 cards, 30% slower) fixed both the token glitches and a domain answer Q4 got confidently wrong.

15s
Read More
Running a 35B MoE (Qwen3.6-35B-A3B) on 2x GTX 1080 Ti in 2026 — Real Benchmarks, and Does the Second GPU Actually Help?
🖼️
0

Running a 35B MoE (Qwen3.6-35B-A3B) on 2x GTX 1080 Ti in 2026 — Real Benchmarks, and Does the Second GPU Actually Help?

DEV Community: machinelearning·byeongsoo kang·4 months ago
#Pp28NVQa
#dev#model#iq4_xs#ollama#quant#experts

I benchmarked Qwen3.6-35B-A3B (IQ4_XS) on two 8-year-old GTX 1080 Ti cards: ~20 tok/s, and the second GPU only adds ~20% (not 2x). Real numbers, VRAM math, and why a 35B fits 22GB.

15s
Read More
📰
0

Best VPS for Ollama in 2026: Compare Top AI Hosting Providers

Bluehost Blog·Megh Bhavsar·4 months ago
#fR32L4vT
#bluehost#strong#ollama#nbsp#class#hosting

Key Highlights Running AI models locally sounds powerful until your machine becomes the bottleneck. Large models like Llama 3, Mistral, DeepSeek and Gemma can slow down your system, drain memory and turn every prompt into a waiting game.…

15s
Read More
How to Host Ollama on VPS: Step-by-Step Deployment Guide
📰
0

How to Host Ollama on VPS: Step-by-Step Deployment Guide

Bluehost Blog·Mili Shah·4 months ago
#DS5rtLkH
#bluehost#nbsp#class#ollama#block#strong

key highlight Hosting your own large language models gives you complete control over your AI data and keeps sensitive information off third-party servers.…

15s
Read More
Building a Local LLM API Server with Ollama + FastAPI — From Dev to Docker Deployment
🖼️
0

Building a Local LLM API Server with Ollama + FastAPI — From Dev to Docker Deployment

DEV Community: fastapi·Jangwook Kim·4 months ago
#kxKkv3wN
#dev#fullscreen#ollama#fastapi#model#article

A hands-on guide to wrapping Ollama REST API with FastAPI to build a production-ready local LLM server with SSE streaming, health checks, and Docker Compose deployment. Real execution logs included.

15s
Read More
I stopped hitting Claude's message limit by building a local AI pipeline that does the heavy lifting
🖼️
0

I stopped hitting Claude's message limit by building a local AI pipeline that does the heavy lifting

XDA·Abhinav Raj·4 months ago
#M0vgkCjI
#sensa#ai#community#claude#model#local

From XDA Developers: I stopped hitting Claude's message limit by building a local AI pipeline that does the heavy lifting

15s
Read More
Generative AI: From Curiosity to Real Production — The Complete Pipeline
🖼️
0

Generative AI: From Curiosity to Real Production — The Complete Pipeline

DEV Community·jesus manrique·4 months ago
#Dv3ARkYs

AI is not a toy. Learn to build a self-hosted content production pipeline that generates, reviews, and publishes social media posts without growing your marketing team.

15s
Read More