← Back to all spotlights

Run LLMs Locally with llama.cpp: The Complete Technical Guide

Master ggerganov/llama.cpp to run massive open-weights models locally across Apple Silicon, NVIDIA GPUs, and commodity x86 CPUs with zero bloat.

P24
By Pickwise24 Editorial Team
Verified Open-Source Review

What Is llama.cpp?

llama.cpp is an open-source inference engine developed by Georgi Gerganov that executes large language models (LLMs) locally in pure C and C++ without external dependencies or heavy Python runtimes.

  • Repository: https://github.com/ggerganov/llama.cpp
  • Core Philosophy: Minimalist, zero-dependency C/C++ execution with 1st-class Apple Silicon Metal, CUDA, and AVX/NEON CPU acceleration.
  • Format Standard: Creator and native home of the GGUF (GPT-Generated Unified Format) model distribution standard.

+-------------------------------------------------------------------+
|                        Client Layer                               |
|        OpenAI-Compatible REST API  /  CLI Tools  /  Bindings      |
+-------------------------------------------------------------------+
                                  |
                                  v
+-------------------------------------------------------------------+
|                           llama.cpp                               |
|   Memory Mapping (mmap)  |  Context Shifting  |  KV Cache Mgmt    |
+-------------------------------------------------------------------+
                                  |
                                  v
+-------------------------------------------------------------------+
|                     ggml Compute Backend                          |
|   Apple Metal (MPS)  |  NVIDIA CUDA  |  Vulkan  |  AVX-512 CPU    |
+-------------------------------------------------------------------+

The Problem: The Python Tax and VRAM Starvation

Running modern open-weights foundation models—from Llama 3 to Mistral and DeepSeek—typically lands developers into dependency purgatory. A standard PyTorch stack bundles gigabytes of CUDA toolkits, Python interpreter overhead, and brittle environment locks. Worse, standard 16-bit float weights demand enterprise-grade graphics cards just to load the parameters before token generation even begins.

llama.cpp bypasses this entire ecosystem. By writing matrix multiplication kernels directly in C/C++ and raw assembly, Georgi Gerganov turned ordinary hardware into capable inference rigs. It allows a developer on an M-series MacBook Air or an aging workstation with a consumer RTX card to run state-of-the-art models at native memory bandwidth limits.


Key Architectural Pillars

1. Native Quantisation (k-quants and IQ)

llama.cpp pioneered integer quantisation techniques that squeeze model weights down to 4-bit, 3-bit, and even 1.5-bit precision with minimal perplexity degradation:

  • Legacy Quants (Q4_0, Q8_0): Uniform block-level scaling. Fast, simple, but slightly lossy at lower bit-widths.
  • k-quants (Q4_K_M, Q5_K_S): Multi-stage quantisation allocating higher precision to sensitive layers (such as attention projections) and lower precision to feed-forward blocks.
  • I-quants (IQ3_XXS, IQ2_M, IQ1_S): Importance-matrix-driven vector quantisation that makes running 70B models feasible within 24GB of unified memory or VRAM.

2. GGUF Single-File Delivery

Gone are the days of parsing separate tokenizer configs, weight shards, and hyperparameter JSONs. GGUF packages metadata, tensor geometry, vocabulary, and quantised weights into a single binary file that mounts via mmap() almost instantaneously.

3. Heterogeneous Compute & Layer Offloading

Hardware is rarely uniform. llama.cpp allows you to split tensor operations dynamically. If a model has 32 transformer layers and your GPU only has enough VRAM for 22, you offload 22 layers to CUDA or Metal and compute the remaining 10 across CPU cores using SIMD instructions (AVX-512 or ARM NEON).


Performance Comparison: Runtime Backends

DimensionNative llama.cpp (Metal/CUDA)PyTorch (Hugging Face / vLLM)Ollama (Wrapped Engine)
Startup OverheadMilliseconds (mmap)Seconds to minutesFast (wraps llama.cpp)
Memory FootprintAbsolute minimal (model + context)High (PyTorch + CUDA runtime)Minimal + daemon overhead
DependenciesC/C++ compiler onlyPython, PyTorch, CUDA librariesStandalone binary
Offloading GranularityLayer-by-layer (-ngl)Coarse-grained / device-levelConfigurable via Modelfile
Custom ExtensibilityDirect C++ API and low-level CLIPython ecosystemConstrained by wrapper API

Step-by-Step Setup and Compilation

Getting started does not require virtual environments or package managers. You build straight from source using standard build tools.

1. Compile the Binary

On macOS (leveraging Apple Silicon Metal out of the box):


git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j

On Linux or Windows with NVIDIA CUDA:


git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j

2. Obtain a GGUF Model Weight

Download a pre-quantised model weight directly from the Hugging Face Hub (for instance, a Mistral 7B Instruct quant):


# Using huggingface-cli or curl
huggingface-cli download \
  bartowski/Mistral-7B-Instruct-v0.3-GGUF \
  Mistral-7B-Instruct-v0.3-Q4_K_M.gguf \
  --local-dir ./models

Practical CLI Workflows

Interactive CLI Generation

To run interactive terminal chat with prompt caching and GPU offloading:


./build/bin/llama-cli \
  -m ./models/Mistral-7B-Instruct-v0.3-Q4_K_M.gguf \
  -ngl 33 \
  -c 4096 \
  --temp 0.7 \
  --repeat-penalty 1.1 \
  -p "You are an expert systems programmer. Explain zero-copy buffer sharing in plain English."
  • -ngl 33: Offloads 33 transformer layers to the GPU.
  • -c 4096: Sets context window size to 4096 tokens.
  • --temp 0.7: Controls sampling randomness.

Running a Production-Ready Local Server

llama.cpp includes a high-performance HTTP server matching OpenAI API specifications, complete with continuous batching and multi-slot context management:


./build/bin/llama-server \
  -m ./models/Mistral-7B-Instruct-v0.3-Q4_K_M.gguf \
  -ngl 99 \
  -c 8192 \
  --port 8080 \
  --host 127.0.0.1

Once running, query it with standard curl payloads or point any LangChain, LlamaIndex, or OpenAI SDK endpoint straight to http://127.0.0.1:8080/v1:


curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-3.5-turbo",
    "messages": [
      {"role": "system", "content": "You are a precise technical editor."},
      {"role": "user", "content": "List three differences between CUDA cores and Tensor cores."}
    ]
  }'

Key Takeaways for AI Builders

  • Total Sovereignty: Operates completely offline, making it ideal for air-gapped workloads, confidential enterprise data, and privacy-first local tools.
  • Hardware Squeezing: Runs models on everything from a Raspberry Pi 5 to multi-GPU servers using the same unified codebase.
  • Underpinning the Ecosystem: Popular local engines like Ollama, LM Studio, and Jan act as frontend abstractions around llama.cpp and ggml.
  • Precision Control: Custom k-quants allow fine-tuning the trade-off between perplexity, memory consumption, and tokens per second.

🛡️ Editorial Standards & Methodology

Every repository featured on Pickwise24 undergoes testing on local workstation hardware before publication. We verify CLI installation steps, review open-source repository licensing, benchmark computational footprint, and evaluate architectural trade-offs to provide genuine, high-utility developer intelligence.