Running Ollama Models

To run a model with Ollama, use theollama runcommand.

For example, to run qwen3.5:0.8b and chat with this model, you can use the following command:

The first run automatically downloads it, about 1.0GB:

ollama run qwen3.5:0.8b

After executing the command, you will first see the download progress. Once the download is complete, you will directly enter the conversation:

$ ollama run qwen3.5:0.8b
pulling manifest
pulling 87497c5f0e8f... 100% |████████████████████| 1.0GB
pulling 5e2ad5... 100% |████████████████████|  11KB
verifying sha256 digest
writing manifest
success
>>> 你好,你是什么模型?

>>> /bye

Seeing>>>the prompt means the model has finished loading. Type a question and press Enter to start chatting.

If the network disconnects during download, you don't need to start over. Ollama supports resumable downloads. Running the same command again will continue from the previous progress.


Select a model using the interactive menu

In addition to typing commands directly, simply entering ollama without arguments will open the interactive menu. The new version makes it a unified entry point.

$ ollama
Chat, Code, & Work
    Chat with models, code, search the web, and delegate real work

  Launch Claude Code
    Anthropic's coding tool with subagents

  Launch OpenCode
    Anomaly's open-source coding agent

  Launch Hermes Agent
    Self-improving AI agent built by Nous Research

  Launch OpenClaw
    Personal AI with 100+ skills

  More...
    Show additional integrations

Select Chat, Code, & Work to pick a downloaded model and start chatting. The effect is exactly the same as ollama run.

Here you can select the model:

Launch in the menu is used to start integrated tools such as VS Code, Claude Code, etc. This is an advanced feature that will be introduced later in the tutorial.


Pick your first model: choose based on your memory

Bigger models are not necessarily better. A model your machine can't handle will only make your computer's fans spin wildly, and responses will be measured in seconds or even minutes.

The principle for beginners is simple: start with a small model first, then gradually increase once it runs smoothly. Choose based on your machine's memory and VRAM:

Your machine configurationRecommended tagDownload sizeExpected experience
8GB RAM, no dedicated GPUqwen3.5:0.8b / 2b1.0 ~ 2.7GBLightweight and smooth for conversations
16GB RAM, no dedicated GPUqwen3.5:4b / 9b3.4 ~ 6.6GBDefault spec, balanced quality and speed
GPU with 8GB or more VRAMqwen3.5:9b or qwen3.5:latestAbout 6.6GBGPU acceleration, significantly faster speed
GPU with 16GB or more VRAMqwen3.5:27bAbout 17GBHigher quality long-form text and code
Insufficient configurationqwen3.5:cloudNo download requiredCloud inference, requires logging in to an Ollama account

Different specs of the same model have real capability differences. 0.8b is suitable for verification environments. For daily use, it's recommended to start with at least 4b.


Understand the model:tag naming convention

The name you write in the command consists of two parts: model name + tag. This determines exactly what you download.

model:tag 命名规则解剖图

Model name: the family name

The model name corresponds to a family in the official model library, such as qwen3.5, gemma4, deepseek-r1.

There can be multiple specs under the same family, and they share a model name.

Tag: specification and version

The tag determines which specific file is downloaded. Common tags are parameter specifications, such as 0.8b, 9b, 27b.

When no tag is specified, Ollama automatically addslatest, which is the default spec set by the official for that model.

For qwen3.5,ollama run qwen3.5is equivalent toollama run qwen3.5:latest, currently pointing to the 9b spec.

Namespace: who published it

Names with a slash indicate models published by third parties. The part before the slash is the publisher's username.

Models in the official model library have no namespace. Models published by individuals or teams (e.g., in the format example/my-model:4b) will have one. Sharing and distribution of models are demonstrated in the private deployment chapter.

View downloaded models

If you're unsure which models are locally available and how much space each occupies, use the list command to check:

Example

# List all local models
ollama list
$ ollama list
NAME               ID            SIZE      MODIFIED
qwen3.5:0.8b       9e3f4b3a1c2d  1.0GB     5 minutes ago
qwen3.5:latest     a8b2c9d3e4f5  6.6GB     2 hours ago

The next chapter will cover model management commands such as list, show, and rm comprehensively.


Exit, re-enter, and switch models

In a conversation, type/byeto exit and return to the normal terminal.

Example

# Enter /bye to exit the conversation and return to the normal terminal
/bye

Running the same ollama run command again will re-enter the conversation. Since the model has already been downloaded, this time it starts almost instantly.

To switch to another model for chat, simply run another model. You can switch back and forth between the two models:

Example

# Switch from the small model to the default spec
ollama run qwen3.5:4b

# Switching back is the same
ollama run qwen3.5:0.8b

After switching, the previous model stays in memory for about 5 minutes and is automatically unloaded if there are no new requests. This mechanism is called keep-alive. The next chapter will explain how to adjust it.

When entering multi-line text in a conversation, wrap it in three double quotes:

Example

>>> """Translate the following text into English:
... EXAMPLE is a tutorial website.
... Thank you.
... "
""

Using the model via the Python SDK

If you want to integrate Ollama with Python code, you can use Ollama's Python SDK to load and run models.

1. Install the Python SDK

First, you need to install Ollama's Python SDK. Open a terminal and execute the following command:

pip install ollama

2. Write a Python script

Next, you can use Python code to load and interact with the model.

The following is a simple Python script example demonstrating how to use the qwen3.5 model to generate text:

Example

import ollama
response = ollama.generate(
    model="qwen3.5",  # Model name
    prompt="Who are you."  # Prompt text
)
print(response)

3. Run the Python script

Run your Python script in the terminal:

python test.py

You will see the model's reply based on your input.

4. Chat mode

Example

from ollama import chat
response = chat(
    model="qwen3.5",
    messages=[
        {"role": "user", "content": "Why is the sky blue?"}
    ]
)
print(response.message.content)

This code will have a conversation with the model and print the model's reply.

5. Streaming responses

Example

from ollama import chat
stream = chat(
    model="qwen3.5",
    messages=[{"role": "user", "content": "Why is the sky blue?"}],
    stream=True
)
for chunk in stream:
    print(chunk["message"]["content"], end="", flush=True)

This code receives the model's response in a streaming manner, suitable for processing large amounts of data.

Other extensions