Ollama Related Commands

Ollama provides a variety of command-line tools (CLI) for users to interact with locally running models.

Basic format:

ollama <command> [args]

We can useollama --helpto see what commands are available:

Large language model runner

Usage:
  ollama [flags]
  ollama [command]

Available Commands:
  serve       Start ollama
  create      Create a model from a Modelfile
  show        Show information for a model
  run         Run a model
  stop        Stop a running model
  pull        Pull a model from a registry
  push        Push a model to a registry
  list        List models
  ps          List running models
  cp          Copy a model
  rm          Remove a model
  help        Help about any command

Flags:
  -h, --help      help for ollama
  -v, --version   Show version information

1. Usage

  • ollama [flags]: Run ollama with flags.

  • ollama [command]: Run a specific ollama command.

2. Available Commands

  • serve: Start the ollama service.
  • create: Create a model from a Modelfile.
  • show: Display detailed information about a model.
  • run: Run a model.
  • stop: Stop a running model.
  • pull: Pull a model from a model registry.
  • push: Push a model to a model registry.
  • list: List all models.
  • ps: List all running models.
  • cp: Copy a model.
  • rm: Delete a model.
  • help: Get help information about any command.

3. Flags

  • -h, --help: Display help information for ollama.
  • -v, --version: Display version information.

Complete example:

Command Description Example
ollama run Run a model. If it does not exist, it will be pulled automatically. ollama run llama3
ollama pull Pull a model. Download the model from the registry without running it. ollama pull mistral
ollama list List models. Display all locally downloaded models. ollama list
ollama rm Delete a model. Remove the local model to free up space. ollama rm llama3
ollama cp Copy a model. Copy an existing model to a new name (for testing). ollama cp llama3 my-model
ollama create Create a model. Create a custom model from a Modelfile (advanced). ollama create my-bot -f ./Modelfile
ollama show Show information. View the model's metadata, parameters, or Modelfile. ollama show --modelfile llama3
ollama ps View processes. Display currently running models and VRAM usage. ollama ps
ollama push Push a model. Upload your custom model to ollama.com. ollama push my-username/my-model
ollama serve Start the service. Start Ollama's API service (usually runs automatically in the background). ollama serve
ollama help Help. View help information for any command. ollama help run

1. Pulling and Deleting Models

pull
Pull a remote model to the local machine.

ollama pull <model>

rm / remove
Delete a local model.

ollama rm <model>

list / ls
List all local models.

ollama list

2. Running Models

run
Run the model in interactive mode without exiting.

ollama run <model>

Can include system information and prompt:

ollama run <model> -s "<system>" -p "<prompt>"

run + script
Read prompt from a file:

ollama run <model> < input.txt

When you enterollama runAfter entering the chat interface, you are no longer operating the command line, but talking to the AI. At this point you can use the/shortcut commands starting with a prefix to control the conversation:

  • /byeor/exit:Most important!Exit the chat interface and return to the command line.
  • /clear: Clear the current context memory (start a new conversation).
  • /show info: View detailed parameter information of the current model.
  • /set parameter seed 123: Set a random seed (advanced technique, used to reproduce results).
  • /help: View all available shortcut commands in the chat.

3. Inference Interface (One-time Execution)

generate
Perform a single inference and output text.

ollama generate <model> -p "<prompt>"

4. Creating and Modifying Models

create
Create a local model from a Modelfile.

ollama create <model-name> -f Modelfile

cp
Copy a model to a new name.

ollama cp <src> <dst>

5. Server Related

serve
Start the Ollama local service (default 11434).

ollama serve

run serverless
Whenollama runIt automatically starts the background service; no need to execute it separately.


6. Model Information

show
View model metadata, parameters, and templates.

ollama show <model>

7. Special Parameters

Most of these parameters can be used in run/generate:

--num-predict <number>    限制输出 token 数
--temperature <float>     控制随机性
--top-k <int>             采样范围
--top-p <float>           核采样
--seed <int>              固定随机性
--format json             输出 JSON
--keepalive <seconds>     会话保持时间

8. Modelfile Directives

Used when building a model:

  • FROM <model>: Base model
  • SYSTEM "xxx": Set system prompt
  • PARAMETER key=value: Set default parameters
  • TEMPLATE "xxx": Customize Chat template
  • LICENSE "xxx": Set License
  • ADAPTER <file> / WEIGHTS <file>: Load LoRA or additional weights

9. API (when serve is running)

REST endpointshttp://localhost:11434/api:

  • /api/generate: Text generation
  • /api/chat: Chat streaming interface
  • /api/pull: Remote pull
  • /api/tags: Local model list

Example call (curl):

curl http://localhost:11434/api/generate \
  -d '{"model":"qwen2.5","prompt":"hello"}'

10. Advanced

Run with custom parameters:

ollama run <model> --temperature 0.2 --top-p 0.9

Persistent session (preserving context):
Sessions are automatically managed by the model's internal cache, without the need for additional commands.

Other extensions