Local LLM Inference Engines Comparison
Research: compare the top 5 local LLM inference engines (llama.cpp, vLLM, Ollama, TGI, ExLlamaV2) on tokens/sec and VRAM, and chart it
Local LLM Inference Engines: Throughput & VRAM Comparison
TL;DR: Across available sources, vLLM and TGI lead on multi-user throughput via PagedAttention/continuous batching, ExLlamaV2 wins single-stream speed on consumer GPUs, llama.cpp offers the broadest hardware reach with GGUF quantization, and Ollama trades raw speed for ease of use — but the public sources here provide only partial, non-uniform tokens/sec numbers, so direct apples-to-apples benchmarks should be treated as indicative, not definitive.
Why a clean head-to-head is hard
VRAM and throughput vary with precision (FP16/FP8/INT8/INT4/GPTQ/AWQ/GGUF) and serving framework, because runtimes differ in KV-cache handling and pre-allocated memory [2]. Benchmarks in the sources span different models, GPUs, batch sizes, and quantization schemes, so the numbers below are best read as qualitative positioning supported by what each source actually claims.
Engine-by-engine summary
| Engine | Primary use | Model format | Setup | Throughput positioning | VRAM behavior |
|---|---|---|---|---|---|
| llama.cpp | Low-level inference, broad HW support | GGUF | Easy–moderate [4] | Strong single-stream; compared head-to-head with TGI and vLLM in community benchmarks [6] | Low — GGUF quantization (e.g., Q4/Q5/Q8) minimizes VRAM; runs CPU+GPU hybrid [4] |
| vLLM | Production serving, high concurrency | HF safetensors, GPTQ, AWQ [4] | Moderate [4] | Leader for batched/concurrent serving; benchmarked against llama.cpp and TGI [1][6] | Higher baseline — pre-allocates KV cache (PagedAttention) for throughput [2] |
| Ollama | Personal AI server | GGUF [4] | Very easy [4] | Wraps llama.cpp; convenience-first, not throughput-first [4] | Same GGUF-based footprint as llama.cpp [4] |
| TGI (Text Generation Inference) | Production HF serving | HF formats | — | Compared directly to vLLM and llama.cpp in server benchmarks [6] | Server-grade KV cache; not quantified in sources |
| ExLlamaV2 | Fast consumer GPU inference | EXL2, GPTQ [4] | — | Positioned as "fast consumer GPU" engine [4]; referenced in 17k tok |
Sources
- vLLM or llama.cpp: Choosing the right LLM inference engine for your use case | Red Hat Developer
- How Much GPU VRAM Do You Need for a 7B, 33B, or 70B Model?
- Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
- Local LLM Inference Engines Compared: The Definitive 2026 Guide | local-llm.net
- Large language model
- LLM inference server performances comparison llama.cpp / TGI / vLLM · ggml-org/llama.cpp · Discussion #6730
Generated by tink · sources are web pages; verify anything important.