Ollama Multi-Model Collaboration and Agent Workflow

A single model doesn't have to do everything: let small models handle routing, specialized models handle tasks, and connected models search for information. Use structured output as the scheduling protocol to combine fast and cost-effective intelligent applications.

This chapter implements a multi-model routing system and upgrades it into an Agent workflow with internet access capability.


Architecture Design: Small Model Routing + Expert Execution

The core idea is division of labor: determining "what the user wants to do" is simple and can be handled by the smallest model; the actual tasks are then passed to the appropriate specialized models.

多模型路由架构:路由器分类,三个专家模型执行

This design has three advantages: small model routing takes only tens of milliseconds, so large models no longer handle low-value requests; the parameters of each expert model can be independently adjusted as needed; new task types only require adding a branch without affecting existing chains.


Implementation 1: Structured Output for Intent Routing

The key to routing is making classification results "machine-readable"; structured output serves as the scheduling protocol here.

Example

# File path: router.py
from ollama import chat
from pydantic import BaseModel, Field

# Structured definition of routing results
class Route(BaseModel):
    category: str = Field(description=Question category: chat / coding / search)
    reason: str = Field(description=One-sentence classification basis)

def classify(question: str) -> Route:
    response = chat(
        model='qwen3.5:4b',          # A small model is sufficient for routing
        messages=[{'role': 'user', 'content': question}],
        format=Route.model_json_schema(),
        options={'temperature': 0},   # Classification must be stable
    )
    return Route.model_validate_json(response.message.content)

# Quick self-test
for q in ['How to deduplicate a Python list?', 'What AI news is there today?', 'Tell a joke']:
    r = classify(q)
    print(r.category, '|', r.reason)
$ python router.py
coding | 询问 Python 列表去重的编程方法
search | 询问当天的最新新闻,需要联网
chat   | 闲聊类请求,直接回答即可

Implementation 2: Three Expert Processors

One processor function per path, with a unified signature for easy dispatch.

Example

# Append to router.py
from ollama import chat, web_search, web_fetch

def handle_chat(question):
    """Daily Q&A: answered directly by the small model"""
    r = chat(model='qwen3.5:4b',
             messages=[{'role': 'user', 'content': question}])
    return r.message.content

def handle_coding(question):
    """Programming tasks: handed to a dedicated code model"""
    r = chat(model='example-coder',
             messages=[{'role': 'user', 'content': question}],
             options={'temperature': 0.2, 'num_ctx': 65536})
    return r.message.content

def handle_search(question):
    """Web-related tasks: Agent loop of small model + search tool"""
    tools = {'web_search': web_search, 'web_fetch': web_fetch}
    messages = [{'role': 'user', 'content': question}]
    while True:
        r = chat(model='qwen3.5:4b', messages=messages,
                 tools=[web_search, web_fetch])
        messages.append(r.message)
        if not r.message.tool_calls:
            return r.message.content
        for call in r.message.tool_calls:
            func = tools.get(call.function.name)
            # Truncate tool results to protect the context window
            result = str(func(**call.function.arguments))[:2000]
            messages.append({'role': 'tool',
                             'tool_name': call.function.name,
                             'content': result})

HANDLERS = {'chat': handle_chat,
            'coding': handle_coding,
            'search': handle_search}

def ask(question):
    route = classify(question)
    print(f'[Route -> {route.category}] {route.reason}')
    return HANDLERS[route.category](question)

The router itself does not generate answers, so its hallucination risk is limited to "misclassification"; the cost of misclassification is just one extra correction and does not pollute the final answer.


Implementation 3: Local and Cloud Hybrid Orchestration

Each path can have a "cloud upgrade tier": for tasks that local models cannot handle, simply change the model name in the same code structure.

Example

# Append to router.py
def ask_with_fallback(question, force_cloud=False):
    """Scheduling with cloud fallback: local first, complex tasks upgraded to cloud"""
    route = classify(question).category

    if force_cloud or route == 'coding':
        # Heavy tasks such as coding go directly to cloud flagship specs
        model = 'qwen3.5:cloud'
    else:
        model = 'qwen3.5:4b'

    print(f'[Route -> {route}] Using model {model}')
    r = chat(model=model,
             messages=[{'role': 'user', 'content': question}])
    return r.message.content

Note a capability difference: Cloud models do not yet support structured output, so the router must be handled by a local model, which fits perfectly with the "small model for routing" design.


Comprehensive Evaluation of Cost and Latency

The cost-benefit of a multi-model strategy must be calculated across three dimensions simultaneously:

DimensionLocal small modelLocal large modelCloud model
Time to first tokenLowest (millisecond-level loading)MediumIncludes network round trip
Marginal costApproximately electricity costElectricity + VRAM usageBilled by usage
Answer qualitySufficient for simple tasksClose to flagshipFlagship-level
Concurrency capabilityLimited by the local machineLimited by VRAMBest elasticity

This yields a practical baseline strategy: use local small models as the foundation for routing and lightweight tasks; assign quality-sensitive tasks to local large models; switch heavy-load or oversized tasks to the cloud; the model names of all three categories go through the configuration file and can be adjusted anytime based on retirement announcements or hardware changes.

Other Extensions