Local Model Deployment

Many people think AI can only be used in the cloud: when using ChatGPT, data is sent to OpenAI's servers, and when using Claude, requests are sent to Anthropic's data centers.

In fact,AI models can run entirely on your own computer。

Imagine: you can use AI without a network, sensitive data never leaves your computer, no need to pay per call, and you can customize it however you want.

This is the value of local model deployment.

This module will take you from zero to running your first local AI model on your own computer.


Why Local AI Is Needed

Cloud AI is convenient, but local AI has irreplaceable advantages.

Data Privacy Protection

This is the primary reason many people choose local deployment.

If you are dealing with contracts, medical records, financial data, or internal documents, sending this data to third-party servers is risky.

When running locally,all data is processed on your own computer and never leaves your device。

Offline Environment Usage

On airplanes, trains, or in places with unstable networks, cloud AI is unusable.

Once downloaded, local models are available forever and do not require a network connection.

Cost Control

Cloud AI charges based on the number of calls or tokens, and heavy usage can become a significant expense.

Local models are downloaded once and can be used freely afterward, with no additional costs.

Customization Possibilities

You cannot modify cloud models; you can only use the features provided.

With local models, you can fine-tune, quantize, modify prompt templates, and customize them entirely according to your own needs.

Comparison of these advantages:

FeatureCloud AILocal AI
Data PrivacyData must be uploadedData is fully localized
Network DependencyInternet connection requiredNo network required
Usage CostPay per useOne download, unlimited use
Customization CapabilityLimitedFully controllable
Ease of useOut-of-the-boxRequires configuration
Response speedDepends on networkDepends on hardware

Not that cloud AI is bad, but ratherBoth have their own applicable scenariosFor daily chat, use the cloud; for sensitive data, use local; for simple tasks, use the cloud; for special customization, use local.


Hardware Requirements Assessment

Whether a local model can run and how fast it runs mainly depends on your hardware configuration.

CPU Inference vs GPU Inference

There are two ways to run models: CPU inference and GPU inference.

CPU inference uses the computer's main processor to run models. The advantage is good compatibility; almost all computers can run it. The disadvantage is slow speed; large models may take a long time.

GPU inference uses graphics cards. The core advantage isStrong parallel computing capability, making model runs dozens of times faster。

For NVIDIA graphics cards, you need to install the CUDA toolkit; for AMD graphics cards, you can use ROCm; for Intel graphics cards, you can use OpenVINO.

Relationship Between VRAM Requirements and Model Size

VRAM determines how large a model you can run.

The more parameters a model has, the more VRAM it requires. However, with quantization techniques, you can run larger models with less VRAM.

VRAM sizeRunnable models (FP16)Runnable models (after quantization)Recommended scenarios
8 GB7B models are somewhat difficult7B (Q4), 13B (Q4) barely runSimple chat, lightweight tasks
16 GB7B, 13B models7B、13B(Q4/Q8)、33B(Q4)Daily use, best value
24 GB7B, 13B, 33B models7B、13B、33B(Q4/Q8)、70B(Q4)Professional use, smooth experience
48 GB+33B, 70B modelsAll mainstream modelsResearch, production environments

Advantages of the Mac M Series

Apple Silicon (M1, M2, M3 series chips) has unique advantages in local AI.

Apple's Metal framework and unified memory architecture make Macs highly efficient at running local models.

Unified memory meansCPU and GPU share memoryand can automatically borrow memory when VRAM is insufficient.

A MacBook Pro with 16GB of unified memory can smoothly run quantized 7B and 13B models.

Recommended Configuration Plans

Based on different budgets and needs, here are several configuration options:

PlanConfigurationSuitable forExpected experience
Entry-level planAny computer (8GB+ RAM)Beginners who want to experience local AICan run small models, relatively slow
Best value plan16GB RAM + 6GB+ VRAMPersonal daily useSmooth running of 7B/13B models
Professional plan32GB RAM + 12GB+ VRAMDevelopers, researchersSmooth running of 33B models
Mac planM1/M2/M3 + 16GB unified memoryFirst choice for Mac usersSmooth running of 7B/13B models

Don't be intimidated by the term "large model." Today's quantization technology already allows 7B models to run smoothly on ordinary computers, and the capability of 7B models is sufficient for most daily tasks.


Ollama: The Easiest Local AI Tool

Ollama is currently the most popular local model tool. To summarize in one sentence:Install models like installing apps, use AI like using the command line。

Installation and Configuration

Ollama supports macOS, Linux, and Windows, and the installation process is very simple.

Installer download link:https://ollama.com/download

Install with one command:

curl -fsSL https://ollama.com/install.sh | sh

macOS users can download the installer directly, or install with Homebrew:

# macOS 使用 Homebrew 安装
brew install ollama

# 或者下载安装包:https://ollama.com/download

Windows users can simply download the installer from the official website and run it.

After installation, start the Ollama service:

ollama serve

Once you see the service started successfully, you can start using it.

Downloading Models

Ollama's model library is very rich, including mainstream models such as Llama, Qwen, Gemma, and Mistral.

Downloading a model takes just one command:

Example

# Download Llama 3 (8B parameters)
ollama pull llama3

# Download Qwen (Tongyi Qianwen, Chinese-friendly)
ollama pull qwen

# Download Gemma (from Google)
ollama pull gemma

# Download Mistral
ollama pull mistral

You can also useollama run + model namecommand, which will automatically download the specified model if it is not already present:

ollama run qwen

Run the qwen model, downloading it if it's not available.

Each model has different versions, such as 8B, 70B, or different quantized versions.

You can useollama listto view already downloaded models:

ollama list

Output similar to:

NAME              ID              SIZE      MODIFIED
qwen3.5:latest    6488c96fa5fa    6.6 GB    3 months ago

Command Line Usage

After the model is downloaded, you can run it directly to start a conversation:

# 运行 Llama 3,没有会下载
ollama run llama3

# 运行 Qwen,没有会下载
ollama run qwen3.5

qwen3.5 is more than sufficient for ordinary tasks:

Then you can talk directly to the model:

>>> 你好,请介绍一下你自己
你好!我是 Llama 3,由 Meta 开发的 AI 助手。我可以帮你回答问题、
写作、编程、分析数据等。有什么我可以帮你的吗?

>>> 用 Python 写一个 Hello World 程序
好的,这是一个简单的 Python Hello World 程序:

print("Hello, World!")

>>> /bye  # 输入 /bye 退出

Common commands:

CommandFunctionExample
ollama pull <model>Download modelollama pull llama3
ollama run <model>Run modelollama run llama3
ollama listList installed modelsollama list
ollama rm <model>Delete modelollama rm llama3
ollama show <model>View model infoollama show llama3

API Interface Calls

Ollama comes with an HTTP API, so you can call it with code, not just from the command line.

By default, Ollama is athttp://localhost:11434to provide API service.

Call Ollama API with Python:

Example

# File path: /Users/example/ollama_test.py
# Calling Ollama API with Python

import requests
import json

# Ollama API address
OLLAMA_URL = "http://localhost:11434/api/generate"


def ask_ollama(prompt: str, model: str = "llama3") -> str:
    """
Send request to Ollama to get model reply

    Args:
prompt: the prompt input by the user
model: the model name to use, defaults to llama3

    Returns:
The model's reply text
    """

    # Construct request data
    payload = {
        "model": model,      # Model to use
        "prompt": prompt,    # User input
        "stream": False      # Do not use streaming output, return all at once
    }

    try:
        # Send POST request
        response = requests.post(OLLAMA_URL, json=payload)
        response.raise_for_status()  # Check if the request was successful

        # Parse the returned JSON
        result = response.json()
        return result.get("response", "")

    except requests.exceptions.ConnectionError:
        return "Error: Cannot connect to Ollama service, please confirm Ollama is started"
    except requests.exceptions.RequestException as e:
        return f"Request error: {str(e)}"


# Test it
if __name__ == "__main__":
    # Test question 1
    question1 = "Please introduce EXAMPLE tutorial in one sentence"
    answer1 = ask_ollama(question1)
    print(f"Question: {question1}")
    print(f"Answer: {answer1}")
    print("-" * 50)

    # Test question 2
    question2 = "Write a Python function to calculate the Fibonacci sequence"
    answer2 = ask_ollama(question2)
    print(f"Question: {question2}")
    print(f"Answer: {answer2}")

Run this script:

Example

# First install the requests library (if not already installed)
pip install requests

# Run the script
python ollama_test.py

If you want a higher-level wrapper, you can use Ollama's official Python library:

Example

# File path: /Users/example/ollama_advanced.py
# Using the Ollama Python library

# First install: pip install ollama

import ollama


def chat_with_example():
    """
Use the Ollama Python library for multi-turn conversations
    """

    # Message history
    messages = [
        {
            "role": "system",
            "content": "You are a friendly programming assistant named ExampleBot. Your answers should be concise and practical, with more code examples."
        }
    ]

    print("ExampleBot is started! Enter 'quit' to exit.")
    print("-" * 50)

    while True:
        # Get user input
        user_input = input(You:)

        if user_input.lower() in ["quit", "exit", Exit]:
            print(Goodbye!)
            break

        # Add user message to history
        messages.append({
            "role": "user",
            "content": user_input
        })

        # Call the model
        response = ollama.chat(
            model="llama3",
            messages=messages
        )

        # Get reply
        assistant_message = response["message"]["content"]
        print(f"ExampleBot:{assistant_message}")
        print("-" * 50)

        # Add assistant reply to history
        messages.append({
            "role": "assistant",
            "content": assistant_message
        })


if __name__ == "__main__":
    chat_with_example()

Integration with Open WebUI

The command line is convenient, but a graphical interface is more friendly.

Open WebUI is a very beautiful web interface that can connect to Ollama, allowing you to use local models just like ChatGPT.

Install Open WebUI:

Example

# Install Open WebUI using Docker
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data --name open-webui --restart always ghcr.io/open-webui/open-webui:main

# If you don't use Docker, you can also install with pip
pip install open-webui
open-webui serve

After installation, open in browserhttp://localhost:3000, and you'll see an interface that looks a lot like ChatGPT.

Ollama + Open WebUI is currently the most recommended local AI combo—simple to install, friendly interface, and powerful, suitable for the vast majority of users.


LM Studio: Graphical Interface Operations

If you don't like the command line, LM Studio is a great choice—Pure graphical interface, downloading and running models are all done in the window.。

Installing LM Studio

LM Studio supports Windows, macOS, and Linux.

From the official websitehttps://lmstudio.aiDownload the installation package and just install it directly.

Model Download and Management

Open LM Studio, the model library is on the left.

You can search by model name, for example search for "llama", "qwen", "mistral", find the model you want, and click download.

LM Studio will automatically list different versions of the same model:

VersionDescriptionRecommended scenario
Q4_K_M4-bit quantization, little quality lossThe choice of most people
Q5_K_M5-bit quantization, better quality, slightly larger sizeChoose this when you have enough VRAM
Q8_08-bit quantization, close to native qualityFor those who pursue quality and have enough VRAM
FP16Full precision, largest sizeFor research purposes

Downloaded models will appear in the "My Models" list.

Local API Service

LM Studio can also provide an API service, compatible with OpenAI's API format.

This means that for the OpenAI code you write, you only need to change the API address and Key to use a local model.

Example

# File path: /Users/example/lmstudio_openai.py
# LM Studio is compatible with OpenAI API format

from openai import OpenAI

# Connect to LM Studio local server
# Note: You need to start the server in LM Studio first
client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="lm-studio"  # LM Studio can just use this placeholder
)


def ask_local_model(prompt: str, system_prompt: str = "You are a helpful assistant.") -> str:
    """
Calling LM Studio local model via OpenAI-compatible interface
    """

    response = client.chat.completions.create(
        model="local-model",  # This field will be ignored by LM Studio
        messages=[
            {"role": "system", "content": system_prompt},
            {"role": "user", "content": prompt}
        ],
        temperature=0.7,
    )
    return response.choices[0].message.content


# Test
if __name__ == "__main__":
    question = "Introduce the features of the Python tutorial"
    answer = ask_local_model(
        prompt=question,
        system_prompt="You are Example's dedicated customer service, professionally answering questions about programming learning."
    )
    print(f"Q: {question}")
    print(f"A: {answer}")

This feature is very useful—you can use the same code to switch between cloud models and local models as needed.

For users new to local AI, I recommend: first get started with LM Studio to understand the basic concepts and operations; once familiar, use Ollama for automation and integration.


Getting Started with Model Quantization

Quantization is a key technology that allows large models to run on ordinary computers—Run larger models with less VRAM while maintaining most capabilities。

What Is Quantization

Simply put, quantization is reducing the precision of model parameters.

Native models typically store each parameter using 16-bit floating point (FP16) or 32-bit floating point (FP32).

After quantization, we use 8-bit integers (INT8) or even 4-bit integers (INT4) for storage.

Precision decreases, but the size and VRAM requirements are greatly reduced.

For example:

Precision7B model size13B model sizeVRAM requirementQuality loss
FP1613 GB26 GBHighNone
Q8_07 GB13 GBMediumVery small
Q5_K_M5 GB9 GBLowAcceptable
Q4_K_M4 GB7 GBVery lowSlight
Q3_K_M3 GB5 GBExtremely lowObvious

The rule of thumb: Q8 quantization has almost imperceptible quality loss, and Q4 quantization is sufficient for most tasks.

GGUF Format Explained

GGUF is currently the most popular quantized model format.

GGUF's predecessor is GGML, developed by the llama.cpp project.

Advantages of this format:

  • First,Good cross-platform compatibility—Works on Windows, macOS, and Linux; runs on both CPU and GPU.

  • Second,Supports multiple quantization levels—From FP16 to Q3, choose as needed.

  • Third,Fast inference speed—Heavily optimized for CPU and GPU.

Most local model tools (Ollama, LM Studio, llama.cpp) now support GGUF format.

Differences Between Q4/Q8 Quantization

Q4 and Q8 are the two most commonly used quantization levels.

  • Q8 quantization is 8-bit; each parameter is stored as an 8-bit integer. The advantage of Q8 isMinimal quality loss, almost identical to the original model; the disadvantage is that the size is still not small enough.

  • Q4 quantization is 4-bit; each parameter is stored as a 4-bit integer. The advantage of Q4 isSmall size and fast speed; the disadvantage is that quality drops noticeably on complex tasks.

How to choose:

ScenarioRecommended quantization levelReason
Very limited VRAM (below 8GB)Q4_K_MBeing usable is the top priority
Daily conversation, simple tasksQ4_K_MBest value for money
Writing, programming, analysisQ5_K_M or Q6_KBalance of quality and speed
Pursuing the best qualityQ8_0Close to native quality

AWQ/GPTQ Quantization Schemes

In addition to GGUF quantization, there are two other common quantization schemes: AWQ and GPTQ.

AWQ (Activation-aware Weight Quantization) is characterized byRetaining weights that have a large impact on the results, resulting in smaller quality loss after quantization.

GPTQ (Gradient-based Post-Training Quantization) is characterized byUsing gradient information to optimize quantization, suitable for scenarios with GPU.

These two schemes are mainly used on NVIDIA GPUs that support CUDA.

Simple comparison:

Quantization schemeApplicable hardwareSpeedQualityRecommended tools
GGUFCPU + various GPUsFastOKOllama、LM Studio
AWQNVIDIA GPUVery fastVery goodvLLM、Text Generation WebUI
GPTQNVIDIA GPUVery fastVery goodAutoGPTQ、ExLlama

For beginners, don't worry about these technical details—just use the Q4_K_M version recommended in Ollama or LM Studio. It is the best balance point verified by extensive testing.


Model Selection Guide

There are hundreds of open-source models now; it's important to choose one that suits your needs.

Comparing 7B/13B/70B Models

The number of parameters is an important metric—7B, 13B, and 70B are the most common specifications.

SpecificationVRAM requirement (Q4)Inference speedCapability levelApplicable scenarios
7B4-6 GBVery fastAdequateDaily conversation, simple tasks
13B7-10 GBFastGoodWriting, programming, analysis
33B15-20 GBMediumExcellentComplex reasoning, professional domains
70B30-40 GBSlowerClose to GPT-3.5Research, production environments

For most people, the sweet spot is the 13B Q4 quantized version—good capability with acceptable speed.

Recommended Models with Good Chinese Support

Many foreign models have mediocre Chinese capabilities. Here are a few Chinese-friendly models:

Model nameDeveloperFeaturesOllama command
Qwen (Tongyi Qianwen)AlibabaStrong Chinese capability, well-rounded performanceollama pull qwen
Yi (01.AI)01.AIGood Chinese comprehension, strong reasoningollama pull yi
DeepSeekDeepSeekStrong coding, decent Chineseollama pull deepseek-coder
Llama 3MetaStrong overall, but Chinese needs fine-tuningollama pull llama3
GemmaGoogleSolid foundation, good coding abilityollama pull gemma

For Chinese tasks, Qwen is the first choice; its Chinese understanding and generation are top-tier among open-source models.

Task-Specific Model Selection

Different models have different strengths; choose based on your task:

Task typeRecommended modelReason
Everyday conversation, Q&ALlama 3、Qwen、MistralBalanced overall capability
Writing, copywritingYi, Llama 3, Claude (cloud)Good writing style, fluent expression
Programming, codeDeepSeek-Coder、CodeLlama、StarCoderStrong code understanding and generation
Math, reasoningLlama 3、Qwen、WizardMathGood logical reasoning
Professional domains (medical, legal)Models fine-tuned for the domainTrained on professional data

A bigger model isn't always better, nor is a newer one—one that fits your hardware and your taskis the best. Start by trying a few 7B models to get a feel, then upgrade.


Hands-On: Building a Fully Offline Local Knowledge Base

One of the most practical use cases for local models is building your own private knowledge base—feed local documents to the model and have it answer questions based on them.。

The entire system runs fully offline; data never leaves your computer.

Technical Approach

We use RAG (Retrieval-Augmented Generation) technology:

1. Split documents into chunks

2. Use an embedding model to convert each chunk into a vector

3. Store vectors in a vector database

4. When asking, convert the question into a vector too

5. Find the most relevant document chunks in the vector database

6. Send the relevant documents and the question together to the local LLM

7. The LLM answers based on the document content

This sounds complex, but it's simple with existing tools.

Implementation with Ollama + LangChain

We'll use Python to implement a simple version:

Example

# File path: /Users/example/local_knowledge_base.py
# Fully offline local knowledge base RAG system

# First install dependencies:
# pip install langchain langchain-ollama chromadb pypdf

from langchain_ollama import OllamaLLM, OllamaEmbeddings
from langchain_community.document_loaders import PyPDFLoader, TextLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Chroma
from langchain.chains import RetrievalQA
import os


class LocalKnowledgeBase:
    """Local knowledge base class implementing fully offline RAG functionality"""

    def __init__(self, model_name: str = "llama3", persist_dir: str = "./chroma_db"):
        """
Initialize knowledge base

        Args:
model_name: The Ollama model name to use
persist_dir: Vector database persistence directory
        """

        self.model_name = model_name
        self.persist_dir = persist_dir

        # 1. Initialize LLM (local model)
        self.llm = OllamaLLM(model=model_name)

        # 2. Initialize embedding model (local embeddings)
        # Use nomic-embed-text; this model works well for both Chinese and English
        self.embeddings = OllamaEmbeddings(model="nomic-embed-text")

        # 3. Vector database placeholder
        self.vector_store = None
        self.qa_chain = None

    def load_documents(self, document_paths: list):
        """
Load documents and build vector database

        Args:
document_paths: List of document paths
        """

        documents = []

        for path in document_paths:
            if not os.path.exists(path):
                print(f"Warning: file {path} does not exist, skipping")
                continue

            # Select loader based on file extension
            if path.lower().endswith(".pdf"):
                loader = PyPDFLoader(path)
            elif path.lower().endswith(".txt"):
                loader = TextLoader(path, encoding="utf-8")
            else:
                print(f"Warning: unsupported file type {path}, skipping")
                continue

            documents.extend(loader.load())
            print(f"Loaded: {path}")

        if not documents:
            print("No documents were successfully loaded")
            return

        print(f"Total loaded {len(documents)} document chunks")

        # 4. Split documents
        text_splitter = RecursiveCharacterTextSplitter(
            chunk_size=500,      # Each chunk: 500 characters
            chunk_overlap=100    # Overlap between chunks: 100 characters
        )
        split_docs = text_splitter.split_documents(documents)
        print(f"Split into {len(split_docs)} chunks")

        # 5. Build vector database
        print("Building vector database...")
        self.vector_store = Chroma.from_documents(
            documents=split_docs,
            embedding=self.embeddings,
            persist_directory=self.persist_dir
        )
        print("Vector database built successfully!")

        # 6. Build QA chain
        self._build_qa_chain()

    def _build_qa_chain(self):
        """Build RAG QA chain"""
        if not self.vector_store:
            print("Please load a document first!")
            return

        # Build retriever
        retriever = self.vector_store.as_retriever(
            search_type="similarity",
            search_kwargs={"k": 3}  # Return the 3 most relevant chunks
        )

        # Build RAG chain
        self.qa_chain = RetrievalQA.from_chain_type(
            llm=self.llm,
            chain_type="stuff",
            retriever=retriever,
            return_source_documents=True
        )

    def load_existing_db(self):
        """Load existing vector database"""
        if os.path.exists(self.persist_dir):
            self.vector_store = Chroma(
                persist_directory=self.persist_dir,
                embedding_function=self.embeddings
            )
            self._build_qa_chain()
            print("Loaded existing vector database")
            return True
        else:
            print("No existing vector database found")
            return False

    def ask(self, question: str) -> dict:
        """
Ask the knowledge base a question

        Args:
question: question

        Returns:
A dictionary containing the answer and reference sources
        """

        if not self.qa_chain:
            return {"error": "Please load a document or an existing database first!"}

        # Construct prompt template to encourage Chinese answers
        result = self.qa_chain.invoke({
            "query": f"""Please answer the following question in Chinese. If the answer is in the context, point it out explicitly.
If the answer is not in the context, honestly say "Unable to answer based on the provided information."

Question: {question}
"""

        })

        return {
            "question": question,
            "answer": result["result"],
            "sources": [doc.page_content for doc in result["source_documents"]]
        }


# ============================================
# Usage example
# ============================================

if __name__ == "__main__":
    # 1. Create a knowledge base instance
    kb = LocalKnowledgeBase(model_name="qwen")  # Use Qwen, better for Chinese

    # 2. Try to load an existing database; otherwise, load documents
    if not kb.load_existing_db():
        # Prepare some documents, e.g., save some EXAMPLE tutorials as txt
        sample_docs = [
            "/Users/example/documents/example_python_tutorial.txt",
            "/Users/example/documents/company_handbook.pdf",
        ]
        kb.load_documents(sample_docs)

    # 3. Interactive Q&A
    print("=" * 50)
    print("Example local knowledge base has started!")
    print("Enter your question, or type 'quit' to exit")
    print("=" * 50)

    while True:
        user_input = input("\n"Your question: ")

        if user_input.lower() in ["quit", "exit", "Exit"]:
            print("Goodbye!")
            break

        if not user_input.strip():
            continue

        # Ask a question
        result = kb.ask(user_input)

        if "error" in result:
            print(f"Error: {result['error']}")
            continue

        print(f"\n"Answer: {result['answer']}")
        print("\n"Reference sources: ")
        for i, source in enumerate(result["sources"], 1):
            print(f"\n[{i}] {source[:200]}...")

Before using this system, you need to download the embedding model:

Example

# Download embedding model
ollama pull nomic-embed-text

# Download main model (if not already downloaded)
ollama pull qwen

This way, you have a private knowledge base system that runs entirely locally.

You can add your own notes, documents, and e-books, and let AI help you search and answer questions.

All data in this system stays on your local machine—documents on your hard drive, vector database in your directory, models on your computer—so you don't need to worry about data leaks at all.


Performance Optimization Tips

The local model is running, but how can you make it faster? Here are some practical tips.

Choosing the Right Quantization Level

This is the simplest and most effective optimization—the lower the quantization level, the faster the speed.

If Q4 works, don't use Q8; if Q8 works, don't use FP16.

Adjusting Context Window Size

The larger the context window, the more computation is required.

If you don't need a long context, you can set it in Ollama's Modelfile:

Example

# Create a custom model with a limited context window
# Create a new file Modelfile

echo "FROM llama3
PARAMETER num_ctx 2048"
> Modelfile

# Create the custom model
ollama create llama3-fast -f Modelfile

# Run
ollama run llama3-fast

Reducing the context from the default 8192 to 2048 will make it noticeably faster.

Using GPU Acceleration

If you have an NVIDIA GPU or an Apple Silicon Mac, make sure Ollama is using the GPU rather than the CPU.

Ollama will automatically detect and use the GPU, but sometimes you need to verify it.

In the Ollama startup logs, you should see messages like "Metal GPU activated" or "CUDA activated".

Using Faster Sampling Parameters

Some sampling parameters can affect speed:

Example

# Create a speed-optimized Modelfile
# File: Modelfile.fast

echo "FROM llama3
PARAMETER temperature 0.7
PARAMETER top_k 20
PARAMETER top_p 0.9
PARAMETER num_predict 256"
> Modelfile.fast

# Create and run
ollama create llama3-optimized -f Modelfile.fast
ollama run llama3-optimized

Parameter descriptions:

ParameterDescriptionOptimization Suggestions
temperatureControl RandomnessModerate is fine, doesn't affect speed
top_kLimit candidate word countSet it smaller (20-40) to speed up
top_pCumulative probability thresholdDefault 0.9 is fine
num_predictMaximum generated token countLimit as needed
num_ctxContext window sizeJust enough is fine, the smaller the faster

Disabling Unnecessary Features

If you're using API calls and don't need streaming output, you can set stream to false.

Although this won't make inference faster, it can reduce network transmission overhead.

Performance optimization is an art of balance—speed, quality, and memory cannot all be achieved at once. Find the balance point that suits your needs.

Other extensions