Ollama Command Complete Manual

This chapter introduces the commands for using Ollama, covering installation, model management, runtime parameters, Modelfile, REST API, environment variables, and common troubleshooting. All content is based on Ollama 0.17.x.


Installation and Startup

All three platforms have standard installation methods. After installation, Ollama automatically runs as a background service, listening on127.0.0.1:11434。

PlatformInstallation MethodDescription
macOSbrew install ollamaYou can also download the DMG installer from the official website.
Linuxcurl -fsSL https://ollama.com/install.sh | shOfficial one-click script that automatically configures the systemd service.
WindowsDownload installer from the official website.Install via the wizard; automatically registers as a system service.
Dockerdocker run -d -v ollama:/root/.ollama -p 11434:11434 ollama/ollamaData volume ollama persists models.

In most cases you don't need to start the service manually, but it is useful when troubleshooting or changing configuration temporarily.ollama serve。

Example

# Start the service in the foreground; logs print directly to the terminal. Most commonly used for troubleshooting.
ollama serve

# Verify the service is ready: it is normal if it returns "Ollama is running"
curl http://localhost:11434

If a port-occupied message appears, the background service is already running. Running `ollama serve` in another terminal will then report an error, which is normal. Simply use the command line or API directly.


Complete Table of Model Management Commands

Model management is the main part of the Ollama command line, with about ten commands, all listed in the table below.

CommandFunctionExample
ollama pull <model>Download a model from the model library; if no tag is given, it defaults to `latest`.ollama pull qwen3.5:4b
ollama run <model>Run a model interactively; if it is not available locally, it will automatically pull it.ollama run qwen3.5
ollama list / ollama lsList local models, sizes, and modification times.ollama list
ollama show <model>View model details: parameters, template, system prompt, etc.ollama show qwen3.5
ollama psView models loaded into memory and their remaining time-to-live.ollama ps
ollama stop <model>Immediately stop and unload a running model.ollama stop qwen3.5
ollama cp <src> <dst>Copy or rename a model.ollama cp qwen3.5 example-clone
ollama rm <model>Delete a local model to free disk space.ollama rm qwen3.5:0.8b
ollama create <name> -f <file>Create a custom model based on a Modelfile.ollama create example-coder -f Modelfile
ollama push <model>Push a model to a remote model library; login is required first.ollama push myteam/example-coder
ollama serveStart the background service; it usually runs automatically with the system.ollama serve

The most frequently used combination in daily work is `run`, `list`, `ps`, and `rm`: run models, check inventory, inspect memory, and clean up space.

Example

# Download a model of a specific size
ollama pull qwen3.5:0.8b

# Make a copy and give it your own name
ollama cp qwen3.5 example-clone

# After confirming the copy succeeded, delete the size that is no longer needed
ollama rm qwen3.5:0.8b

`rm` deletes the model files corresponding to that tag. If other tags share the same underlying weights (blobs), the disk space is only fully released after the last tag referencing them is deleted.


run Command and Interactive Mode

`run` can be used as a one-shot Q&A tool, can enter interactive mode for multi-turn conversations, and also supports changing parameters at runtime.

Command-Line Arguments

Arguments are written after the model name; they only take effect for the current run and are not written into the model configuration.

Example

# Interactive conversation
ollama run qwen3.5

# One-shot Q&A: prompt follows the model name; exits automatically after answering
ollama run qwen3.5 "Introduce Python in one sentence."

# Show token statistics: total time, tokens per second, and other performance metrics
ollama run qwen3.5 --verbose

# Force the model to output in JSON format
ollama run qwen3.5 --format json

# Increase the context window (unit: tokens)
ollama run qwen3.5 --num-ctx 8192

# Put as many model layers as possible into the GPU (999 means no limit)
ollama run qwen3.5 --num-gpu 999

Slash Commands in Interactive Mode

After entering interactive mode, commands starting with a slash adjust the session state; these adjustments only take effect for the current session.

CommandFunction
/set parameter <name> <value>Temporarily adjust inference parameters, such as `temperature`.
/set system <prompt>Temporarily set the system prompt.
/show infoDisplay information about the current model.
/show modelfileDisplay the Modelfile corresponding to the current model.
/save <name>Save the current session settings as a new model.
/load <name>Load a saved model or session.
/clearClear the current conversation context and start over.
/byeExit interactive mode.

If you want to keep parameters set via slash commands for the long term, use `/save` to save them as a new model; behind the scenes it generates a Modelfile for you.

Another shortcut to exit interactive mode is pressing Ctrl+D, which has the same effect as `/bye`.


Modelfile Custom Models

Modelfile uses a Dockerfile-like syntax to package a 'base model + parameters + system prompt' into a new model.

Example

# File path: Modelfile
# FROM specifies the base model (required); it can also be a local GGUF file path
FROM qwen3.5

# Common inference parameters (optional)
PARAMETER temperature 0.2
PARAMETER num_ctx 8192

# System prompt; determines the model's role and behavior (optional)
SYSTEM You are a senior Python engineer. Output concise, type-annotated, testable code.

Example

# Build a new model with a Modelfile
ollama create example-coder -f Modelfile

# Run it like a regular model
ollama run example-coder "Write a quicksort, including tests."

To import local GGUF weights, simply point FROM to the file path; this is commonly used to run model files downloaded from the open-source community.

Example

# File path: Modelfile
# FROM points to a local GGUF weight file
FROM ./model-file.gguf

All Modelfile directives are listed in the table below; the first three rows are the most commonly used combination.

DirectivePurposeRequired
FROMSpecify the base model or GGUF file path.Required
PARAMETERSet inference parameters: temperature, num_ctx, top_p, stop, etc.Optional
SYSTEMSet the system prompt.Optional
TEMPLATECustomize the conversation prompt template.Optional
ADAPTERLoad a LoRA adapter.Optional
MESSAGEPreset conversation examples to implement few-shot guidance.Optional
LICENSEDeclare the model license.Optional

If you want to modify an existing model and don't want to start from scratch, first runollama show --modelfile 模型名to export the full recipe; after editing, run `create` again.


REST API Quick Reference

After the Ollama service starts, it listens by default onhttp://localhost:11434, and everything the CLI can do, the API can also do.

EndpointMethodFunction
/api/generatePOSTSingle-turn text generation; returns as a stream by default.
/api/chatPOSTMulti-turn conversation; pass history via the `messages` array.
/api/embedPOSTGenerate text vector embeddings for RAG retrieval.
/api/tagsGETList local models; corresponds to `ollama list`.
/api/showPOSTView model details; corresponds to `ollama show`.
/api/pullPOSTPull a model; corresponds to `ollama pull`.
/api/deleteDELETEDelete a model; corresponds to `ollama rm`.

Example

# Single-turn generation (streaming)
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3.5",
"prompt": "Write a Hello World in Python"
}'


# Multi-turn conversation
curl http://localhost:11434/api/chat -d '{
  "model": "qwen3.5",
"messages": [{"role": "user", "content": "Explain what REST API is."}]
}'


# Generate vector embeddings
curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
"input": "Your text content"
}'


# List local models
curl http://localhost:11434/api/tags

# View model details
curl http://localhost:11434/api/show -d '{"model": "qwen3.5"}'

# Delete model
curl -X DELETE http://localhost:11434/api/delete -d '{"model": "qwen3.5"}'

OpenAI-Compatible Interface

Ollama provides an OpenAI-compatible endpoint./v1/chat/completions, existing OpenAI ecosystem code can be used directly by pointing base_url to local.

Example

# File path: openai_compat.py
from openai import OpenAI

# base_url points to local Ollama; api_key can be any non-empty string
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

resp = client.chat.completions.create(
    model="qwen3.5",  # Name of the locally downloaded model
    messages=[{"role": "user", "content": "Hello, please introduce example"}],
)
print(resp.choices[0].message.content)
$ python3 openai_compat.py
你好!example(Example)是一个面向初学者的编程学习网站,提供大量免费的基础技术教程。

Environment Variables Quick Reference

Environment variables are the main way to configure Ollama service behavior. Changes take effect only after restarting the service.

VariableDescriptionDefault value
OLLAMA_HOSTService listening address; 0.0.0.0 allows LAN access127.0.0.1:11434
OLLAMA_MODELSModel storage path; commonly used when changing to a larger disk~/.ollama/models
OLLAMA_NUM_GPUNumber of model layers placed into GPU; 999 means allAuto
OLLAMA_NUM_PARALLELNumber of concurrent requests for the same modelAuto
OLLAMA_MAX_LOADED_MODELSMaximum number of models resident in memory at the same timeAuto
OLLAMA_FLASH_ATTENTIONEnable Flash Attention; recommended for long contextsOff
OLLAMA_KEEP_ALIVEHow long a model stays idle before being unloaded from memory5m
OLLAMA_DEBUGEnable debug logs to output more detailed runtime informationOff

Example

# Allow LAN access + move models to data disk
export OLLAMA_HOST=0.0.0.0:11434
export OLLAMA_MODELS=/data/ollama-models
ollama serve

Setting OLLAMA_KEEP_ALIVE to 0 unloads the model immediately after each request, saving memory but reloading on every request; set to -1 to keep it resident. The syntax of this value varies by version; refer to `ollama serve --help` and official documentation.


Common Troubleshooting

When encountering errors, first use the table below to locate the issue. Most problems fall into three categories: incorrect model name, service not running, or insufficient resources.

ProblemTroubleshooting approach
Insufficient VRAM / memory (OOM)Switch to a smaller model size or lower quantization precision, e.g., use 4b instead of 9b
Error says model not foundCheck whether the model name and tag are correct. Tags are separated by a colon, e.g., qwen3.5:4b, not qwen3.5-4b.
connection refusedVerify whether the service is running and check whether port 11434 is occupied.
Slow responseUse `ollama ps` to check the PROCESSOR column and confirm whether the model is actually running on the GPU; appropriately reduce num_ctx.
Confusion with multiple model versionsWhen no tag is specified, the "latest" version is selected according to internal rules; it is recommended to always explicitly specify the full tag.

Quick Reference Cheat Sheet

Finally, the most frequently used commands are condensed into a table, ready to copy and use.

CommandFunction
ollama pull <model>Download model
ollama run <model>Run model / enter interactive conversation
ollama listView downloaded models
ollama show <model>View model details
ollama psView running models
ollama stop <model>Stop and unload model
ollama rm <model>Delete model
ollama cp <old> <new>Copy / rename model
ollama create <name> -f ModelfileCreate a custom model using a Modelfile
ollama serveStart service

Ollama updates quickly, and some parameters and command behaviors may change with versions. Before taking action, refer to `ollama --help` and official documentation.

Other extensions