Ollama GPU and Performance Optimization

If a model runs slowly, clues for nine out of ten causes can be found in a single line of ollama ps. This chapter moves from troubleshooting methods to hardware support and multi-GPU allocation, then to two advanced switches: Flash Attention and KV cache quantization.

There is only one goal: make the most of your GPU.


The first step is always: confirm where the model is running

The first action in performance troubleshooting is to look at the PROCESSOR column. It directly tells you where the model weights are loaded.

$ ollama ps
NAME             ID            SIZE      PROCESSOR    CONTEXT    UNTIL
qwen3.5:27b      a8b2c9d3e4f5  20.4GB    52%/48%      CPU/GPU    4 minutes from now

The three forms correspond to completely different situations:

PROCESSOR showsMeaningSpeed expectation
100% GPUAll weights loaded into VRAMFull GPU speed
100% CPUVRAM cannot hold it at all, pure memory inferenceOne to two orders of magnitude slower
52%/48% CPU/GPUInsufficient VRAM, weights are splitNoticeable drop, bottlenecked at cross-device transfer

It's more intuitive to look at it together with generation speed: eval_count / eval_duration * 10^9 at the end of the API response gives you token/s (see the later part of the API chapter for usage). Use this number as the quantitative benchmark for each optimization.

When encountering a slow model, follow the decision path below:

性能排查决策流:按 PROCESSOR 列分流处理


GPU acceleration support by platform

Ollama supports four major acceleration backends, and installing the correct driver is a prerequisite.

PlatformBackendRequirement
NVIDIACUDACompute capability 5.0+, driver 550+ (old cards with 5.0~6.2 require 570+)
AMD(Linux)ROCm v7Requires ROCm v7 driver; old drivers may cause discovery timeout and fall back to CPU
AMD(Windows)ROCm v7 / VulkanROCm v7 or HIP7 driver stack; some RX 6000 series use Vulkan as a fallback
Apple SiliconMetalWorks out of the box, no configuration needed
Intel / OthersVulkanSupplementary backend enabled by default, covering older AMD cards and Intel GPUs

To troubleshoot whether the GPU is recognized, discovery records in the logs are the most convincing; NVIDIA users can also use docker run --gpus all ubuntu nvidia-smi to verify passthrough in container scenarios.

To restrict Ollama to only certain GPUs, or force pure CPU mode, each backend has corresponding switches:

Example

# NVIDIA: only let Ollama use the first GPU (recommend using UUID, check with nvidia-smi -L)
CUDA_VISIBLE_DEVICES=GPU-xxxx ollama serve

# AMD: similarly use the ROCr device list
ROCR_VISIBLE_DEVICES=0 ollama serve

# Vulkan: specify the device index
GGML_VK_VISIBLE_DEVICES=1 ollama serve

# For all three syntaxes, "-1" means disable GPU and force pure CPU
CUDA_VISIBLE_DEVICES=-1 ollama serve

Multi-GPU load distribution

The allocation strategy on multi-GPU machines is automatic, with only one rule: prefer a single GPU, and only split if it doesn't fit.

When the model can fit entirely into any single card, Ollama will choose one to load it—single-GPU avoids cross-PCIe data transfer and is usually the optimal throughput solution.

When a single GPU cannot fit the model, it is split across all available GPUs; at this point you also need to reserve headroom for KV cache and context. This is also a common reason why large-parameter models "get split even though the total VRAM seems sufficient."

When mixing multiple AMD GPUs on Linux, certain driver versions may produce garbled output. These issues are documented in AMD's official multi-GPU known issues documentation; prioritize upgrading the ROCm driver.


CPU inference optimization headroom

It can run without a discrete GPU, but there are several points worth squeezing out.

Ollama comes with multiple CPU inference libraries, with performance ranking: cpu_avx2 > cpu_avx > cpu. It is selected automatically by default; you can force-specify when auto-detection fails:

Example

# Force use of AVX2 inference library
OLLAMA_LLM_LIBRARY=cpu_avx2 ollama serve

# Confirm which instruction sets your CPU supports
cat /proc/cpuinfo | grep flags | head -1

Two additional facts: the macOS Rosetta translation environment can only use the basic cpu library; the thread count is controlled by num_thread in options, which defaults to the number of physical cores. Blindly increasing it can actually slow things down due to hyper-threading contention.

The hidden major factor in CPU inference is context length: with the same model, the CPU inference speed difference between 4K and 32K context is very significant. When configuring models for CPU scenarios, don't casually set num_ctx very large.


Flash Attention and KV cache quantization

These two are the most important VRAM optimizations for long-context scenarios. They target not model weights, but the KV cache that grows with context.

Newer versions of Ollama automatically enable Flash Attention when the backend supports it, and you can also force it on/off with environment variables:

Example

# Force enable Flash Attention
OLLAMA_FLASH_ATTENTION=1 ollama serve

# Quantize KV cache to 8-bit: VRAM halved, almost no precision loss (recommended)
OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve
KV cache typeVRAM usageQuality impactRecommendation
f16 (default)BaselineLosslessKeep default when VRAM is abundant
q8_0About halfAlmost imperceptibleRecommended choice for long-context scenarios
q4_0About a quarterPerceptible under long contextConsider only when maximizing VRAM savings

Two points to note: the KV cache type is a global configuration that applies to all loaded models; the impact of quantization on quality varies by model, and some models with high GQA structures are more sensitive. After switching, it's recommended to compare and verify using prompts with a fixed seed.


Quality trade-off between quantization precision and speed

The model selection chapter covered the size benefits of quantization; here we add several conclusions from a speed perspective.

PrecisionSizeSpeedQualityApplicability
q8_0About half of fp16Close to fp16Almost losslessPreferred choice when VRAM is abundant
q4_K_MAbout a quarter of fp16Usually fasterSlight decreaseDefault tier for local deployment
q4_K_S and lowerSmallerNot necessarily fasterMore noticeable decreaseExtreme VRAM savings

A counterintuitive point: lower quantization is not necessarily faster. Inference speed is constrained by memory bandwidth, and some low-bit formats require additional dequantization computation. In practice, they may be the same or even slower. Don't just look at the numbers when choosing a quantization tier.

Recommended comparison method: fix the seed and use the same set of prompts, compare the output and token/s of two quantization tiers at a temperature of 0.3, looking at quality and speed together.


Concurrency and throughput tuning

After single-request optimization reaches its limit, the next step is to make the service handle concurrency.

The core variable is OLLAMA_NUM_PARALLEL (number of parallel requests per model). The cost must be kept in mind: parallel processing multiplies the context cache, and the required VRAM is approximately equal to that value multiplied by the KV space for the context length. Setting it too large may evict the model from the GPU.

Server-side queueing is controlled by OLLAMA_MAX_QUEUE, default 512. Exceeding it returns 503 directly, and the client needs to retry.

The operational order for throughput tuning:

Step 1: single-request tuning (specs, quantization, FA and KV cache) to confirm single-stream speed.

Step 2: gradually increase NUM_PARALLEL, and at each level use ollama ps to confirm it's still 100% GPU, observing whether total throughput rises.

Step 3: when 503 queueing or mixed splitting appears, step back one level—that's the VRAM red line.

Other extensions