Ollama Environment Variables and Service Configuration

Ollama has no traditional configuration file; almost all service behavior is controlled by environment variables: listen address, model path, context, concurrency, proxy, and debugging.

This section first covers the setup methods for the three platforms, then explains each core variable in detail across three layers: network entry, request scheduling, and model runtime.


Where Configuration Variables Live in the Service

Let's first create a map: the dozen or so environment variables are not all on the same level; they each act on different stages of the service process.

Ollama 服务配置地图:网络入口、请求调度、模型运行三层

To view the full list of variables supported by the current version, run `ollama serve --help`.


How to Set Environment Variables on Three Platforms

The same variable has completely different setup entry points across the three platforms; this is where beginners get stuck most easily.

macOS:launchctl

When Ollama runs as an app, environment variables must be injected via launchctl; after setting them, restart the app for them to take effect:

Example

# Inject variables one by one
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"

# Then quit and reopen the Ollama app

Linux: systemd Service Override

Example

# Open the service override editor
sudo systemctl edit ollama.service

# Add in the editor (note: must be under the [Service] section):
# [Service]
# Environment="OLLAMA_HOST=0.0.0.0:11434"

# Reload and restart the service
sudo systemctl daemon-reload
sudo systemctl restart ollama

Windows: System Environment Variables

Search for "environment variables" in the Start menu, select "Edit environment variables for your account," create or edit the variable, then save.

Key step: quit Ollama from the tray first, then restart it from the Start menu after saving; only then will the new variables take effect in the new process.

There is also a JSON configuration file server.json (located at ~/.ollama/) for a small number of toggle-style settings; for example, write {"disable_ollama_cloud": true} to disable cloud features. The vast majority of configuration still follows environment variables.


Network Entry: HOST, ORIGINS, and Proxy

OLLAMA_HOST: Let Your LAN Access Your Models

By default, it only listens on 127.0.0.1, so only the local machine can use it; to let other devices on the LAN call it, change it to listen on all network interfaces:

Example

# Listen on all network interfaces (accessible from LAN)
OLLAMA_HOST=0.0.0.0:11434 ollama serve

# Verify from another device (replace with your actual IP)
curl http://192.168.1.100:11434/api/version

Opening it to the LAN is equivalent to exposing an unauthenticated model service to all devices on the same subnet. This is acceptable on a home network, but on a corporate network you must pair it with firewall rules or a reverse proxy (discussed in detail in the private deployment chapter).

OLLAMA_ORIGINS: Cross-Origin Allowlisting

By default, Ollama only accepts cross-origin requests from 127.0.0.1 and 0.0.0.0. When web pages or browser extensions connect directly to the local API, you need to explicitly allowlist them:

Example

# Allow all browser extensions
OLLAMA_ORIGINS=chrome-extension://*,moz-extension://*,safari-web-extension://* ollama serve

Proxy: HTTPS_PROXY Is the Only One Needed

Model downloads use HTTPS for outbound traffic; in network-restricted environments, just point HTTPS_PROXY at the proxy.

Set only HTTPS_PROXY, not HTTP_PROXY: Ollama does not fetch models over HTTP, and adding an HTTP proxy may actually interrupt the connection between client and server. In Docker scenarios, pass -e HTTPS_PROXY=... when starting the container; if using a proxy with a self-signed certificate, you also need to install the CA certificate into the system (or bake it into the image).


Model Runtime: Path, Context, and Keep-Alive

OLLAMA_MODELS: Store Models on Another Drive

By default, models are stored in the user directory; the paths for the three platforms are as follows:

PlatformDefault Model Path
macOS~/.ollama/models
Linux/usr/share/ollama/.ollama/models
WindowsC:\Users\<username>\.ollama\models

Point OLLAMA_MODELS at the new directory to migrate everything; after moving the downloaded model files there, they can be used directly.

Example

# Linux example: migrate to a large-capacity data disk
OLLAMA_MODELS=/data/ollama-models ollama serve

# Note: standard Linux installations run the service as the ollama user
# The new directory must have its owner changed, otherwise the service won't have read/write permission
sudo chown -R ollama:ollama /data/ollama-models

OLLAMA_CONTEXT_LENGTH: Global Context Length

Sets the default context length for all models; it takes effect when not overridden at the model or request level:

Example

# Global default 8K
OLLAMA_CONTEXT_LENGTH=8192 ollama serve

Priority from highest to lowest: options.num_ctx in the API request, PARAMETER num_ctx in the Modelfile, this environment variable, and the automatic tiered default based on VRAM.

OLLAMA_KEEP_ALIVE: Global Keep-Alive Duration

Uniformly adjusts how long a model stays resident after loading; the value format matches the keep_alive parameter in the API:

Example

# Common values: 300 (seconds) / "10m" / "24h" / -1 always resident / 0 unload immediately after use
OLLAMA_KEEP_ALIVE=30m ollama serve

A keep_alive passed individually in an API request takes precedence over this global value, enabling a tiered strategy like "10 minutes globally, key models always resident."


Request Scheduling: The Concurrency Trio

Ollama supports two levels of concurrency: multiple models resident at the same time, and a single model handling multiple requests in parallel—each controlled by its own variable.

VariableDefault ValuePurpose
OLLAMA_NUM_PARALLEL1Number of requests a single model handles simultaneously; memory required scales by this value × context length
OLLAMA_MAX_LOADED_MODELS3 × number of GPUs (3 for CPU inference)Total number of models allowed to be loaded simultaneously, provided VRAM/RAM can hold them
OLLAMA_MAX_QUEUE512Maximum number of queued requests when the service is busy; requests beyond this return 503 immediately

The scheduling logic is worth understanding: when VRAM is sufficient, models process requests in parallel; when VRAM is insufficient, new requests queue up, waiting for idle models to be unloaded and free up space.

NUM_PARALLEL is the variable most likely to trip you up: setting it to 4 means the context cache is multiplied by 4, VRAM usage instantly quadruples, and models may be pushed off the GPU as a result. Planning for multi-user shared instances is covered in the private deployment chapter.


Debugging and Logging: OLLAMA_DEBUG

The first step in troubleshooting is always to enable debugging and check the logs.

Example

# Enable debug logging
OLLAMA_DEBUG=1 ollama serve

To enable debugging in the Windows GUI app: first quit the app from the tray, then run in PowerShell:

Example

$env:OLLAMA_DEBUG="1"
& "ollama app.exe"

Summary of log locations across the four environments:

EnvironmentLog Location
macOS~/.ollama/logs/server.log
Linux(systemd)journalctl -u ollama --no-pager --follow --pager-end
Windows%LOCALAPPDATA%\Ollama\server.log
Dockerdocker logs <container name> (stdout/stderr)

Quick Reference for Other Useful Variables

The remaining variables are mostly tied to specific scenarios; full chapter-by-chapter explanations are covered in the corresponding topic sections. Here we'll establish an index first.

VariableExample ValuePurpose
OLLAMA_FLASH_ATTENTION1 / 0Force Flash Attention on/off; significantly saves VRAM with long contexts (detailed in the GPU chapter)
OLLAMA_KV_CACHE_TYPEq8_0 / q4_0KV cache quantization type; requires Flash Attention (detailed in the GPU chapter)
OLLAMA_VULKAN0Disable the Vulkan backend (a toggle for when Vulkan is malfunctioning)
OLLAMA_LLM_LIBRARYcpu_avx2Force a specific inference library, bypassing automatic detection
OLLAMA_TMPDIR/data/tmpChange the temp directory location (when the system /tmp is mounted with noexec)
OLLAMA_NO_CLOUD1Disable cloud models and web search; pure local mode
Other Extensions