Menu

Why we write our own C and C++ engines
📰
0

Why we write our own C and C++ engines

Hacker News·Why we write our own C and C++ engines·about 1 month ago
#HtyoLCQQ
#localai#python#against#parity#vllm#model
Reading 0:00
15s threshold

Most LocalAI backends wrap somebody else’s engine, and that is the right default. llama.cpp, vLLM, whisper.cpp, stable-diffusion, MLX and the rest are maintained by people who are better at those models than we are, and wrapping them costs a Dockerfile and a gRPC shim. Eighteen of our backends do not wrap anything. They are C or C++ ports we wrote from scratch, and each one exists because wrapping the upstream engine would have meant shipping something we could not ship: a multi-gigabyte Python install, a non-portable CUDA-only stack, or a model that had no C++ implementation at all. This post is about what those ports buy, measured, and what they cost. What you get: one file, and memory you can predict Deploying a Python inference stack means resolving a dependency tree at install time, on the target machine, against whatever CUDA and glibc it has. Deploying a ggml port means copying a shared library and a GGUF file.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More