Ollama Basic Concepts
Ollama is a localized machine learning framework that supports multiple natural language processing (NLP) tasks, focusing on model loading, inference, and generation tasks.
With Ollama, users can conveniently interact with large pre-trained models deployed locally.
1. Model
In Ollama, models are the core component. They are pre-trained machine learning models that can perform different tasks, such as text generation, text summarization, sentiment analysis, dialogue generation, and more.
Ollama supports a variety of popular pre-trained models. Common models include:
- deepseek-v3: A large language model provided by DeepSeek, specifically used for text generation tasks.
- LLama2: A large language model provided by Meta, specifically used for text generation tasks.
- GPT: OpenAI's GPT series models, suitable for a wide range of tasks such as dialogue generation and text reasoning.
- BERT: Pre-trained models used for sentence understanding and question-answering systems.
- Other custom models: Users can upload their own custom models and use Ollama for inference.
Main functions of the model:
- Inference: Generates output results based on user input.
- Fine-tuning: Users can train on top of existing models using their own data, thereby customizing the model to suit specific tasks or domains.
Models are typically neural networks composed of a large number of parameters. By training on large amounts of text data, they can learn language patterns and perform efficient inference.
Models supported by Ollama can be accessed at:https://ollama.com/library

Click a model to view the download command:

The following table lists download commands for some models:
| Model | Parameters | Size | Download Command |
|---|---|---|---|
| Llama 3.3 | 70B | 43GB | ollama run llama3.3 |
| Llama 3.2 | 3B | 2.0GB | ollama run llama3.2 |
| Llama 3.2 | 1B | 1.3GB | ollama run llama3.2:1b |
| Llama 3.2 Vision | 11B | 7.9GB | ollama run llama3.2-vision |
| Llama 3.2 Vision | 90B | 55GB | ollama run llama3.2-vision:90b |
| Llama 3.1 | 8B | 4.7GB | ollama run llama3.1 |
| Llama 3.1 | 405B | 231GB | ollama run llama3.1:405b |
| Phi 4 | 14B | 9.1GB | ollama run phi4 |
| Phi 3 Mini | 3.8B | 2.3GB | ollama run phi3 |
| Gemma 2 | 2B | 1.6GB | ollama run gemma2:2b |
| Gemma 2 | 9B | 5.5GB | ollama run gemma2 |
| Gemma 2 | 27B | 16GB | ollama run gemma2:27b |
| Mistral | 7B | 4.1GB | ollama run mistral |
| Moondream 2 | 1.4B | 829MB | ollama run moondream |
| Neural Chat | 7B | 4.1GB | ollama run neural-chat |
| Starling | 7B | 4.1GB | ollama run starling-lm |
| Code Llama | 7B | 3.8GB | ollama run codellama |
| Llama 2 Uncensored | 7B | 3.8GB | ollama run llama2-uncensored |
| LLaVA | 7B | 4.5GB | ollama run llava |
| Solar | 10.7B | 6.1GB | ollama run solar |
2. Task
Ollama supports a variety of NLP tasks. Each task corresponds to different application scenarios of the model, mainly including but not limited to the following:
- Chat Generation: Generates natural conversational replies through interaction with users.
- Text Generation: Generates natural language text based on a given prompt, such as writing articles or generating stories.
- Sentiment Analysis: Analyzes the sentiment倾toward of a given text (e.g., positive, negative, neutral).
- Text Summarization: Compresses long text into a concise summary.
- Translation: Translates text from one language to another.
Through the command-line tool, users can specify different tasks and load different models to accomplish specific tasks.
3. Inference
Inference refers to the process of processing input on a trained model and generating output.
Ollama provides an easy-to-use command-line tool or API, allowing users to quickly provide input to the model and obtain results.
Inference is one of the main functions of Ollama and is the core of interacting with models.
Inference process:
- Input: The user provides text input to the model, which can be a question, prompt, or conversation content.
- Model processing: The model uses its built-in neural network to generate appropriate output based on the input.
- Output: The model returns the generated text content, which may be a reply, generated article, translated text, etc.
Ollama interacts with local models through API or CLI, allowing users to easily implement inference tasks.
4. Fine-tuning
Fine-tuning refers to further training a pre-trained model based on specific domain data, so that the model performs better on specific tasks or domains.
Ollama supports fine-tuning. Users can use their own datasets to fine-tune pre-trained models to customize model output.
Fine-tuning process:
- Prepare dataset: The user prepares a dataset for a specific domain. The data format is usually a text file or JSON format.
- Load pre-trained model: Choose a pre-trained model suitable for fine-tuning, such as Llama2 or GPT models.
- Train: Train the model using the user's specific dataset so that it can better adapt to the target task.
- Save and deploy: After training is completed, the fine-tuned model can be saved and deployed for future use.
Fine-tuning helps the model become more accurate and efficient when handling domain-specific problems.
Other Extensions