← Back to all spotlights

vLLM Guide: High-Throughput Multi-GPU LLM Serving with PagedAttention

Slash your GPU bills and scale open-weight LLMs using vLLM's PagedAttention engine for production-grade throughput and OpenAI-compatible serving.

P24
By Pickwise24 Editorial Team
Verified Open-Source Review

Anyone who has tried deploying a 70B parameter open-weight model using vanilla Hugging Face pipeline code knows the precise moment their soul leaves their body: running out of memory (OOM) on a rig packed with tens of thousands of pounds worth of silicon while serving a grand total of three concurrent users.

The culprit is rarely model weights alone. It is almost always the Key-Value (KV) cache. In traditional serving setups, memory for dynamic request tokens must be allocated contiguously in VRAM. Because requests vary in prompt length and response tokens, engines over-allocate worst-case blocks upfront. The result? Up to 80% of your ultra-fast H100 or A100 VRAM sits empty, trapped in fragmentation limbo while incoming prompts queue up outside the door.

vllm-project/vllm fixes this structural bottleneck. Originating from UC Berkeley's LMSYS lab, vLLM is an open-source, high-throughput LLM inference and serving engine designed specifically to treat GPU memory the way modern operating systems treat RAM: by virtualising and paging it.


+----------------------------------------------------------------+
| Traditional Allocation: Contiguous, Rigid, Massive Waste       |
| [ Prompt A ][ Reserved Empty Cache ][ Prompt B ][ Empty Cache ]|
+----------------------------------------------------------------+
                               vs
+----------------------------------------------------------------+
| vLLM PagedAttention: Non-contiguous, Dynamic Physical Blocks   |
| [Block 1: Req A] -> [Block 2: Req B] -> [Block 3: Req A]       |
+----------------------------------------------------------------+

What Is PagedAttention?

In operating systems, virtual memory paging breaks memory into discrete pages so processes do not need contiguous physical RAM.

vLLM ports this exact logic to transformer self-attention via PagedAttention. Rather than forcing the KV cache for a specific sequence into a contiguous runway of GPU memory, PagedAttention slices keys and values into fixed-size blocks (typically 16 or 32 tokens). A software lookup table maps logical tokens to arbitrary physical blocks scattered across VRAM.


Logical Cache Space: [ Token 0..15 ] [ Token 16..31 ] [ Token 32..47 ]
                             |               |               |
Physical Memory Map:    [ Block 9 ]     [ Block 2 ]     [ Block 14 ]

Because blocks are allocated on demand during token generation, internal fragmentation drops to near zero (under 4%). When multiple requests share identical prompt prefixes—such as system prompts, few-shot examples, or multi-turn conversational histories—vLLM shares physical blocks across requests using copy-on-write semantics.


Architectural Pillars

1. Continuous Batching (Iteration-level Scheduling): Traditional static batching locks a cohort of prompts together until the slowest, most verbose completion finishes. vLLM iterates token-by-token. As soon as a request emits an <eos> token, its memory blocks are recycled, and a waiting request slips into the very next forward pass without waiting for the batch cycle to end.

2. Native Tensor Parallelism: Splitting massive models across 2, 4, or 8 GPUs requires minimal ceremony. vLLM coordinates Megatron-LM-style tensor parallel execution across NCCL without requiring external orchestrators like DeepSpeed.

3. Quantisation Native: Out of the box, vLLM supports modern low-bit kernels including AWQ, GPTQ, SqueezeLLM, and FP8 primitives, letting you squeeze models onto consumer cards or radically cut enterprise cluster footprints.

4. Drop-in OpenAI Compatibility: vLLM exposes an HTTP server mimicking the /v1/chat/completions schema, allowing seamless swap-ins for downstream agent frameworks like LangChain, AutoGen, or LlamaIndex.


Engine Comparison

Feature / EngineVanilla TransformersText Generation Inference (TGI)vLLM
KV Cache ArchitectureContiguous (Wasteful)Paged KV BlocksPagedAttention (Fine-grained)
Batching MechanismStatic / PaddedContinuous BatchingContinuous Iteration Batching
VRAM Waste (Fragmentation)High (60–80%)Low (<10%)Minimal (<4%)
Prefix CachingManual / PrimitiveSupportedAutomatic Copy-on-Write
Multi-GPU ParallelismPipeline / AccelerateTensor ParallelismNative Megatron-style TP

Installation and Quickstart

vLLM requires a CUDA-enabled Linux environment (CUDA 12.1+ recommended).


# Install vLLM via PyPI
pip install vllm

Launching an OpenAI-Compatible Server

To run a multi-GPU deployment of Meta’s Llama 3 across two GPUs with tensor parallelism:


python3 -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3-70B-Instruct \
    --tensor-parallel-size 2 \
    --gpu-memory-utilization 0.92 \
    --max-model-len 8192 \
    --port 8000

Querying the Server

Once running, send standard payloads using standard tools like curl:


curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-70B-Instruct",
    "messages": [
      {"role": "system", "content": "You are a succinct systems engineer."},
      {"role": "user", "content": "Why is PagedAttention superior to contiguous buffers?"}
    ],
    "temperature": 0.2
  }'

Direct Python Offline Inference

For bulk processing, embedding runs, or offline batch synthetic data generation:


from vllm import LLM, SamplingParams

prompts = [
    "Explain Raft consensus in three sentences.",
    "Outline the difference between epoll and kqueue.",
]

sampling_params = SamplingParams(
    temperature=0.7,
    top_p=0.95,
    max_tokens=150
)

# Initialise engine with FP8 quantisation on a single GPU
llm = LLM(
    model="neuralmagic/Meta-Llama-3-8B-Instruct-FP8",
    gpu_memory_utilization=0.90
)

outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(f"\nPrompt: {output.prompt}")
    print(f"Generated: {output.outputs[0].text.strip()}")

Production Best Practices

  • Tuning Memory Utilisation: The --gpu-memory-utilization flag defaults to 0.90. If you experience random CUDA allocation panics, drop it to 0.85. If running dedicated instances with static prompt lengths, push to 0.95 to maximise your KV cache space.
  • Turn on Chunked Prefill: For workloads mixing massive document analysis with interactive chat, enable --enable-chunked-prefill. This prevents a massive 32k context prompt from stalling execution across short-burst user requests during the prefill phase.
  • Enabling Automatic Prefix Caching: If running chat agents with repetitive, hefty system instructions, append --enable-prefix-caching. This caches the key-value tensors of the common prefix across requests, cutting initial time-to-first-token (TTFT) dramatically.

vLLM strips away the manual orchestration acrobatics previously required to serve large models at scale. By treating VRAM like an operating system treats physical memory, it turns erratic, unpredictable LLM traffic into a smooth, high-throughput pipeline.

🛡️ Editorial Standards & Methodology

Every repository featured on Pickwise24 undergoes testing on local workstation hardware before publication. We verify CLI installation steps, review open-source repository licensing, benchmark computational footprint, and evaluate architectural trade-offs to provide genuine, high-utility developer intelligence.