Ollama Command Complete Manual
This chapter introduces the commands for using Ollama, covering installation, model management, runtime parameters, Modelfile, REST API, environment variables, and common troubleshooting. All content is based on Ollama 0.17.x.
Installation and Startup
All three platforms have standard installation methods. After installation, Ollama automatically runs as a background service, listening on127.0.0.1:11434。
| Platform | Installation Method | Description |
|---|---|---|
| macOS | brew install ollama | You can also download the DMG installer from the official website. |
| Linux | curl -fsSL https://ollama.com/install.sh | sh | Official one-click script that automatically configures the systemd service. |
| Windows | Download installer from the official website. | Install via the wizard; automatically registers as a system service. |
| Docker | docker run -d -v ollama:/root/.ollama -p 11434:11434 ollama/ollama | Data volume ollama persists models. |
In most cases you don't need to start the service manually, but it is useful when troubleshooting or changing configuration temporarily.ollama serve。
Example
ollama serve
# Verify the service is ready: it is normal if it returns "Ollama is running"
curl http://localhost:11434
If a port-occupied message appears, the background service is already running. Running `ollama serve` in another terminal will then report an error, which is normal. Simply use the command line or API directly.
Complete Table of Model Management Commands
Model management is the main part of the Ollama command line, with about ten commands, all listed in the table below.
| Command | Function | Example |
|---|---|---|
| ollama pull <model> | Download a model from the model library; if no tag is given, it defaults to `latest`. | ollama pull qwen3.5:4b |
| ollama run <model> | Run a model interactively; if it is not available locally, it will automatically pull it. | ollama run qwen3.5 |
| ollama list / ollama ls | List local models, sizes, and modification times. | ollama list |
| ollama show <model> | View model details: parameters, template, system prompt, etc. | ollama show qwen3.5 |
| ollama ps | View models loaded into memory and their remaining time-to-live. | ollama ps |
| ollama stop <model> | Immediately stop and unload a running model. | ollama stop qwen3.5 |
| ollama cp <src> <dst> | Copy or rename a model. | ollama cp qwen3.5 example-clone |
| ollama rm <model> | Delete a local model to free disk space. | ollama rm qwen3.5:0.8b |
| ollama create <name> -f <file> | Create a custom model based on a Modelfile. | ollama create example-coder -f Modelfile |
| ollama push <model> | Push a model to a remote model library; login is required first. | ollama push myteam/example-coder |
| ollama serve | Start the background service; it usually runs automatically with the system. | ollama serve |
The most frequently used combination in daily work is `run`, `list`, `ps`, and `rm`: run models, check inventory, inspect memory, and clean up space.
Example
ollama pull qwen3.5:0.8b
# Make a copy and give it your own name
ollama cp qwen3.5 example-clone
# After confirming the copy succeeded, delete the size that is no longer needed
ollama rm qwen3.5:0.8b
`rm` deletes the model files corresponding to that tag. If other tags share the same underlying weights (blobs), the disk space is only fully released after the last tag referencing them is deleted.
run Command and Interactive Mode
`run` can be used as a one-shot Q&A tool, can enter interactive mode for multi-turn conversations, and also supports changing parameters at runtime.
Command-Line Arguments
Arguments are written after the model name; they only take effect for the current run and are not written into the model configuration.
Example
ollama run qwen3.5
# One-shot Q&A: prompt follows the model name; exits automatically after answering
ollama run qwen3.5 "Introduce Python in one sentence."
# Show token statistics: total time, tokens per second, and other performance metrics
ollama run qwen3.5 --verbose
# Force the model to output in JSON format
ollama run qwen3.5 --format json
# Increase the context window (unit: tokens)
ollama run qwen3.5 --num-ctx 8192
# Put as many model layers as possible into the GPU (999 means no limit)
ollama run qwen3.5 --num-gpu 999
Slash Commands in Interactive Mode
After entering interactive mode, commands starting with a slash adjust the session state; these adjustments only take effect for the current session.
| Command | Function |
|---|---|
| /set parameter <name> <value> | Temporarily adjust inference parameters, such as `temperature`. |
| /set system <prompt> | Temporarily set the system prompt. |
| /show info | Display information about the current model. |
| /show modelfile | Display the Modelfile corresponding to the current model. |
| /save <name> | Save the current session settings as a new model. |
| /load <name> | Load a saved model or session. |
| /clear | Clear the current conversation context and start over. |
| /bye | Exit interactive mode. |
If you want to keep parameters set via slash commands for the long term, use `/save` to save them as a new model; behind the scenes it generates a Modelfile for you.
Another shortcut to exit interactive mode is pressing Ctrl+D, which has the same effect as `/bye`.
Modelfile Custom Models
Modelfile uses a Dockerfile-like syntax to package a 'base model + parameters + system prompt' into a new model.
Example
# FROM specifies the base model (required); it can also be a local GGUF file path
FROM qwen3.5
# Common inference parameters (optional)
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
# System prompt; determines the model's role and behavior (optional)
SYSTEM You are a senior Python engineer. Output concise, type-annotated, testable code.
Example
ollama create example-coder -f Modelfile
# Run it like a regular model
ollama run example-coder "Write a quicksort, including tests."
To import local GGUF weights, simply point FROM to the file path; this is commonly used to run model files downloaded from the open-source community.
Example
# FROM points to a local GGUF weight file
FROM ./model-file.gguf
All Modelfile directives are listed in the table below; the first three rows are the most commonly used combination.
| Directive | Purpose | Required |
|---|---|---|
| FROM | Specify the base model or GGUF file path. | Required |
| PARAMETER | Set inference parameters: temperature, num_ctx, top_p, stop, etc. | Optional |
| SYSTEM | Set the system prompt. | Optional |
| TEMPLATE | Customize the conversation prompt template. | Optional |
| ADAPTER | Load a LoRA adapter. | Optional |
| MESSAGE | Preset conversation examples to implement few-shot guidance. | Optional |
| LICENSE | Declare the model license. | Optional |
If you want to modify an existing model and don't want to start from scratch, first runollama show --modelfile 模型名to export the full recipe; after editing, run `create` again.
REST API Quick Reference
After the Ollama service starts, it listens by default onhttp://localhost:11434, and everything the CLI can do, the API can also do.
| Endpoint | Method | Function |
|---|---|---|
| /api/generate | POST | Single-turn text generation; returns as a stream by default. |
| /api/chat | POST | Multi-turn conversation; pass history via the `messages` array. |
| /api/embed | POST | Generate text vector embeddings for RAG retrieval. |
| /api/tags | GET | List local models; corresponds to `ollama list`. |
| /api/show | POST | View model details; corresponds to `ollama show`. |
| /api/pull | POST | Pull a model; corresponds to `ollama pull`. |
| /api/delete | DELETE | Delete a model; corresponds to `ollama rm`. |
Example
curl http://localhost:11434/api/generate -d '{
"model": "qwen3.5",
"prompt": "Write a Hello World in Python"
}'
# Multi-turn conversation
curl http://localhost:11434/api/chat -d '{
"model": "qwen3.5",
"messages": [{"role": "user", "content": "Explain what REST API is."}]
}'
# Generate vector embeddings
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": "Your text content"
}'
# List local models
curl http://localhost:11434/api/tags
# View model details
curl http://localhost:11434/api/show -d '{"model": "qwen3.5"}'
# Delete model
curl -X DELETE http://localhost:11434/api/delete -d '{"model": "qwen3.5"}'
OpenAI-Compatible Interface
Ollama provides an OpenAI-compatible endpoint./v1/chat/completions, existing OpenAI ecosystem code can be used directly by pointing base_url to local.
Example
from openai import OpenAI
# base_url points to local Ollama; api_key can be any non-empty string
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
resp = client.chat.completions.create(
model="qwen3.5", # Name of the locally downloaded model
messages=[{"role": "user", "content": "Hello, please introduce example"}],
)
print(resp.choices[0].message.content)
$ python3 openai_compat.py 你好!example(Example)是一个面向初学者的编程学习网站,提供大量免费的基础技术教程。
Environment Variables Quick Reference
Environment variables are the main way to configure Ollama service behavior. Changes take effect only after restarting the service.
| Variable | Description | Default value |
|---|---|---|
| OLLAMA_HOST | Service listening address; 0.0.0.0 allows LAN access | 127.0.0.1:11434 |
| OLLAMA_MODELS | Model storage path; commonly used when changing to a larger disk | ~/.ollama/models |
| OLLAMA_NUM_GPU | Number of model layers placed into GPU; 999 means all | Auto |
| OLLAMA_NUM_PARALLEL | Number of concurrent requests for the same model | Auto |
| OLLAMA_MAX_LOADED_MODELS | Maximum number of models resident in memory at the same time | Auto |
| OLLAMA_FLASH_ATTENTION | Enable Flash Attention; recommended for long contexts | Off |
| OLLAMA_KEEP_ALIVE | How long a model stays idle before being unloaded from memory | 5m |
| OLLAMA_DEBUG | Enable debug logs to output more detailed runtime information | Off |
Example
export OLLAMA_HOST=0.0.0.0:11434
export OLLAMA_MODELS=/data/ollama-models
ollama serve
Setting OLLAMA_KEEP_ALIVE to 0 unloads the model immediately after each request, saving memory but reloading on every request; set to -1 to keep it resident. The syntax of this value varies by version; refer to `ollama serve --help` and official documentation.
Common Troubleshooting
When encountering errors, first use the table below to locate the issue. Most problems fall into three categories: incorrect model name, service not running, or insufficient resources.
| Problem | Troubleshooting approach |
|---|---|
| Insufficient VRAM / memory (OOM) | Switch to a smaller model size or lower quantization precision, e.g., use 4b instead of 9b |
| Error says model not found | Check whether the model name and tag are correct. Tags are separated by a colon, e.g., qwen3.5:4b, not qwen3.5-4b. |
| connection refused | Verify whether the service is running and check whether port 11434 is occupied. |
| Slow response | Use `ollama ps` to check the PROCESSOR column and confirm whether the model is actually running on the GPU; appropriately reduce num_ctx. |
| Confusion with multiple model versions | When no tag is specified, the "latest" version is selected according to internal rules; it is recommended to always explicitly specify the full tag. |
Quick Reference Cheat Sheet
Finally, the most frequently used commands are condensed into a table, ready to copy and use.
| Command | Function |
|---|---|
| ollama pull <model> | Download model |
| ollama run <model> | Run model / enter interactive conversation |
| ollama list | View downloaded models |
| ollama show <model> | View model details |
| ollama ps | View running models |
| ollama stop <model> | Stop and unload model |
| ollama rm <model> | Delete model |
| ollama cp <old> <new> | Copy / rename model |
| ollama create <name> -f Modelfile | Create a custom model using a Modelfile |
| ollama serve | Start service |
Other extensionsOllama updates quickly, and some parameters and command behaviors may change with versions. Before taking action, refer to `ollama --help` and official documentation.