← Back to all spotlights

Run Llama 3, Mistral, and DeepSeek Locally with Ollama

Ditch costly API tokens and cloud latency by serving open-weights LLMs on your own silicon using Ollama's zero-configuration runtime.

P24
By Pickwise24 Editorial Team
Verified Open-Source Review

There is a distinct, visceral pain that arrives at the end of the month when your cloud inference bill lands in your inbox. You built an agentic workflow that loops fifteen times to categorise a support ticket, and suddenly you are funding a hyperscaler’s next data centre. Add to that the mild nausea of shipping proprietary enterprise data over public APIs, and it is little wonder the developer collective has decided to stage a mass migration back to localhost.

Enter ollama/ollama, the open-source CLI runtime that turned the dark art of compiling C++ inference binaries into something as straightforward as pulling a Docker image.


       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚                     Your App / UI                      β”‚
       β”‚   (OpenAI SDK, LangChain, LlamaIndex, curl, Open WebUI)β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                   β”‚ HTTP / localhost:11434
                                   β–Ό
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚                   Ollama Daemon (Go)                   β”‚
       β”‚     REST API β€’ Model Registry β€’ Memory & VRAM Offload  β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                   β”‚ Native Bindings
                                   β–Ό
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚                 llama.cpp Engine (C++)                 β”‚
       β”‚   Apple Silicon (Metal) β€’ NVIDIA (CUDA) β€’ AMD (ROCm)   β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

What is Ollama?

Ollama is an open-source model execution runtime written in Go that packages model weights, configurations, and inference engines into a single unified binary. It wraps the low-level mechanics of Georgi Gerganov’s llama.cpp in an intuitive, container-style interface, allowing developers to download, manage, and execute quantized large language models (LLMs) locally with hardware acceleration on macOS, Linux, and Windows.

Instead of hunting down GGUF quantization splits on Hugging Face and guessing optimal prompt templates, Ollama abstracts the entire lifecycle behind straightforward commands like ollama run llama3.1 or ollama run deepseek-r1.


Architectural Breakdown: How It Manages Your VRAM

Behind the tidy command-line facade, Ollama acts as a dynamic process manager and HTTP reverse proxy:

1. The Core Engine: Ollama embeds llama.cpp under the bonnet. When you summon a model, it detects your hardware (Metal on Apple Silicon, CUDA on NVIDIA, ROCm on AMD) and handles layer offloading dynamically. If a model fits entirely in your GPU’s VRAM, it stays there. If you are a few gigabytes short, it gracefully spills the remaining layers into system RAM.

2. The GGUF Format: Models in the Ollama library are packaged in the GGUF binary format, predominantly quantized to 4-bit (Q4_K_M) by default to strike a balance between perplexity and memory footprint.

3. OpenAI-Compatible API: Ollama exposes an HTTP endpoint on port 11434. It natively mimics the /v1/chat/completions schema, meaning you can point existing client libraries to your local box simply by swapping out the base URL.


Quickstart: Up and Running in Two Minutes

1. Installation

On macOS or Linux, a single shell command handles binary extraction and path setup:


# macOS via Homebrew
brew install ollama

# Linux via automated installer
curl -fsSL https://ollama.com/install.sh | sh

(Windows users get a native installer executable that lives politely in the system tray.)

2. Pulling and Running a Model

To pull Meta's Llama 3.1 8B, Mistral 7B, or DeepSeek’s distilled reasoning models, simply issue:


# Launch interactive chat session with Llama 3.1
ollama run llama3.1

# Or test DeepSeek-R1's chain-of-thought capabilities
ollama run deepseek-r1:8b

If the weights are not present on your drive, Ollama downloads the layers with resumable chunks, initialises the context window, and drops you into a REPL prompt inside your terminal.


Custom Models and System Prompts via Modelfile

Much like a Dockerfile, Ollama uses a declarative format called a Modelfile to bake custom system instructions, stop tokens, and parameters directly into a reusable artefact.

Create a file named Modelfile:


FROM llama3.1:8b

# Set temperature (lower = more deterministic)
PARAMETER temperature 0.2
PARAMETER top_p 0.9

# Define system persona
SYSTEM """
You are a grumpy principal software architect. You review code snippets
ruthlessly in idiomatic British English. You despise unnecessary dependencies.
"""

Compile and run your bespoke persona:


ollama create grumpy-architect -f ./Modelfile
ollama run grumpy-architect

Drop-in Integration with Existing Codebases

Because Ollama provides native OpenAI endpoint parity, integrating it into Python workflows requires virtually zero structural refactoring:


from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama"  # Required by the SDK, ignored by Ollama
)

response = client.chat.completions.create(
    model="llama3.1",
    messages=[
        {"role": "system", "content": "You are a code optimisation assistant."},
        {"role": "user", "content": "Refactor this list comprehension in Python."}
    ],
    temperature=0.3
)

print(response.choices[0].message.content)

Local Inference Runtimes Compared

FeatureOllamallama.cpp (Raw)vLLMLM Studio
Primary TargetDevelopers & Local CLILow-level HackersHigh-throughput ProductionNon-technical Desktop Users
Setup ComplexityNear-zero (single binary)High (manual CMake builds)Medium (Python/CUDA deps)Near-zero (GUI installer)
API EndpointsREST & OpenAI compatibleCustom server binaryFull OpenAI specOpenAI compatible
UI IncludedCLI only (Open WebUI ext)Minimal Web UINone (headless)Polished Desktop GUI
Resource FootprintMinimal daemon overheadBare metal (zero overhead)Heavy (VRAM pre-allocation)Desktop app (Electron/Qt)

Common Gotchas: Context Windows and VRAM Spilling

While Ollama is refreshingly straightforward, community discussions across Reddit and YouTube frequently highlight a couple of practical quirks:

  • The Context Window Trap: By default, Ollama often sets the context window to 2048 tokens to conserve RAM on modest consumer hardware, even if the underlying model supports 128k tokens. If you are feeding it large source files or documentation, you must explicitly bump num_ctx in your Modelfile or API request parameters ("options": {"num_ctx": 32768}).
  • Memory Pressure: When models exceed physical VRAM and spill into DDR4/DDR5 system memory, generation speed drops precipitously. If tokens-per-second tank from 45 down to 3, inspect your terminal logs to check how many layers successfully landed on your GPU.

Key Takeaways

  • True Privacy: Model inputs and token generations never leave your machine, satisfying compliance constraints without contract negotiations.
  • Unified Interface: Serves Mistral, Llama 3, Gemma, and DeepSeek through a single daemon without runtime version conflicts.
  • Tooling Friendly: Works immediately with LangChain, LlamaIndex, Autogen, and Open WebUI out of the box.

πŸ›‘οΈ Editorial Standards & Methodology

Every repository featured on Pickwise24 undergoes testing on local workstation hardware before publication. We verify CLI installation steps, review open-source repository licensing, benchmark computational footprint, and evaluate architectural trade-offs to provide genuine, high-utility developer intelligence.