Ollama: Build a ChatGPT-like Web Application
In this chapter, we will build a ChatGPT-like application: multi-session switching, history persistence to a database, a complete frontend interface, and three deployment methods.
Feature Planning and Data Model
A usable chat product needs at least four capabilities: multi-session isolation, history persistence, streaming replies, and session management (create and delete).
The database model supporting these capabilities only needs two tables:
chats stores the session list, and messages stores each message grouped by foreign key. When switching sessions, assemble the corresponding messages into Ollama's messages array to restore the full context.
The tech stack follows the minimal-dependency principle: Flask + sqlite3 (Python standard library) + vanilla JavaScript frontend, zero frontend frameworks.
Install flask and the ollama extension:
pip install flask ollama
Load the model:
ollama pull qwen3.5:4b
Backend: Complete API for Sessions and Messages
The backend has clear responsibilities: manage the two tables, and write each streaming reply from /api/chat to the database one by one.
Example
import json
import sqlite3
from flask import Flask, request, Response, jsonify, stream_with_context
from ollama import chat
app = Flask(__name__)
DB = 'chats.db'
def db():
"""Open a database connection; rows are returned as dictionaries"""
conn = sqlite3.connect(DB)
conn.row_factory = sqlite3.Row
return conn
def init_db():
with db() as conn:
conn.executescript('''
CREATE TABLE IF NOT EXISTS chats(
id INTEGER PRIMARY KEY AUTOINCREMENT,
title TEXT NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE IF NOT EXISTS messages(
id INTEGER PRIMARY KEY AUTOINCREMENT,
chat_id INTEGER NOT NULL,
role TEXT NOT NULL,
content TEXT NOT NULL
);
''')
init_db()
# Session list
@app.get('/api/chats')
def list_chats():
with db() as conn:
rows = conn.execute(
'SELECT id, title, created_at FROM chats ORDER BY id DESC'
).fetchall()
return jsonify([dict(r) for r in rows])
# Create a new session
@app.post('/api/chats')
def create_chat():
title = (request.json or {}).get('title', 'New session')
with db() as conn:
cur = conn.execute('INSERT INTO chats(title) VALUES (?)', (title,))
return jsonify({'id': cur.lastrowid, 'title': title})
# Delete session (along with its messages)
@app.delete('/api/chats/<int:cid>')
def delete_chat(cid):
with db() as conn:
conn.execute('DELETE FROM messages WHERE chat_id=?', (cid,))
conn.execute('DELETE FROM chats WHERE id=?', (cid,))
return jsonify({'ok': True})
# Read the history of a session
@app.get('/api/chats/<int:cid>/messages')
def get_messages(cid):
with db() as conn:
rows = conn.execute(
'SELECT role, content FROM messages WHERE chat_id=? ORDER BY id',
(cid,)
).fetchall()
return jsonify([dict(r) for r in rows])
Backend Core: Streaming Dialogue and History Write-back
The chat interface does three things: fetch history and build context, stream generation, and write the complete reply back to the database.
Example
@app.post('/api/chat')
def do_chat():
data = request.get_json()
cid, user_input = data['chat_id'], data['message']
# Fetch history (including the just-inserted user message) as the model context
with db() as conn:
conn.execute(
'INSERT INTO messages(chat_id, role, content) VALUES (?,?,?)',
(cid, 'user', user_input)
)
history = [dict(r) for r in conn.execute(
'SELECT role, content FROM messages WHERE chat_id=? ORDER BY id',
(cid,)
)]
messages = [{'role': 'system',
'content': 'You are the EXAMPLE programming assistant. Answer accurately and concisely.'}]
messages += history
def generate():
reply = ''
stream = chat(model='qwen3.5:4b',
messages=messages, stream=True)
for chunk in stream:
reply += chunk.message.content
# Push each chunk to the frontend as NDJSON
yield json.dumps(
{'delta': chunk.message.content},
ensure_ascii=False) + '\n'
# Key: write the complete reply back to the database to form persistent memory
with db() as conn:
conn.execute(
'INSERT INTO messages(chat_id, role, content) VALUES (?,?,?)',
(cid, 'assistant', reply))
yield json.dumps({'done': True}, ensure_ascii=False) + '\n'
return Response(
stream_with_context(generate()),
mimetype='application/x-ndjson')
# Homepage: return the frontend page
@app.get('/')
def index():
return open('index.html', encoding='utf-8').read()
if __name__ == '__main__':
app.run(port=5000)
Frontend: Chat Interface in Vanilla JS
The frontend is feature-complete and simple enough: session list on the left, message area plus input box on the right, with fetch streaming rendering.
Example
<!-- File path: index.html (same directory as server.py) -->
<html lang="zh-CN">
<head><meta charset="UTF-8"><title>Local ChatGPT</title>
<style>
body { display:flex; height:100vh; margin:0; font-family: sans-serif; }
aside { width:220px; border-right:1px solid #ddd; overflow-y:auto; }
aside button { display:block; width:100%; text-align:left;
padding:10px; border:0; background:none; cursor:pointer; }
aside button:hover { background:#f2f2f2; }
main { flex:1; display:flex; flex-direction:column; }
#log { flex:1; overflow-y:auto; padding:20px; }
.msg { margin:8px 0; line-height:1.6; white-space:pre-wrap; }
form { display:flex; border-top:1px solid #ddd; }
input { flex:1; padding:12px; border:0; }
</style>
</head>
<body>
<aside>
<button onclick="newChat()">+ New Session</button>
<div id="list"></div>
</aside>
<main>
<div id="log"></div>
<form id="f">
<input id="q" placeholder="Type a question, press Enter to send" autocomplete="off">
</form>
</main>
<script>
let chatId = null;
const log = document.getElementById('log');
const list = document.getElementById('list');
// Load session list
async function loadChats() {
const chats = await (await fetch('/api/chats')).json();
list.innerHTML = '';
for (const c of chats) {
const b = document.createElement('button');
b.textContent = c.title;
b.onclick = () => openChat(c.id);
list.appendChild(b);
}
}
// Open a session and restore history
async function openChat(id) {
chatId = id;
log.innerHTML = '';
const msgs = await (await fetch(`/api/chats/${id}/messages`)).json();
for (const m of msgs) addMsg(m.role, m.content);
}
// Create a new session
async function newChat() {
const c = await (await fetch('/api/chats', {method:'POST'})).json();
await loadChats();
openChat(c.id);
}
// Append a message to the interface
function addMsg(role, text) {
const div = document.createElement('div');
div.className = 'msg';
div.textContent = (role === 'user' ? 'You: ' : 'Assistant: ') + text;
log.appendChild(div);
log.scrollTop = log.scrollHeight;
return div;
}
// Send and stream-render the reply
document.getElementById('f').onsubmit = async e => {
e.preventDefault();
const input = document.getElementById('q');
const q = input.value.trim();
if (!q || !chatId) return;
input.value = '';
addMsg('user', q);
const div = addMsg('assistant', '');
const resp = await fetch('/api/chat', {method:'POST',
headers:{'Content-Type':'application/json'},
body: JSON.stringify({chat_id: chatId, message: q})});
const reader = resp.body.getReader();
const dec = new TextDecoder();
let buf = '';
while (true) {
const {done, value} = await reader.read();
if (done) break;
buf += dec.decode(value, {stream:true});
const lines = buf.split('\n');
buf = lines.pop();
for (const line of lines) {
if (!line) continue;
const obj = JSON.parse(line);
if (obj.delta) {
div.textContent += obj.delta;
log.scrollTop = log.scrollHeight;
}
}
}
};
newChat();
loadChats();
</script>
</body>
</html>
Start and access:
python server.py
Open http://localhost:5000 in your browser to see the generated page.
Verify three core features: history is fully restored when switching sessions on the left; new sessions do not interfere with each other; history still exists after restarting server.py (persisted to chats.db).
Deployment: Three Tiers
| Tier | Approach | Notes |
|---|---|---|
| Local personal use | Run python server.py directly | Listens on 127.0.0.1 by default, safe |
| Intranet sharing | app.run(host='0.0.0.0'), members access http://server-IP:5000 | No authentication; be sure to add Nginx Basic Auth (see the private deployment chapter) |
| Long-term operation on cloud server | systemd service registration + Nginx reverse proxy + domain name | First run ollama pull <model> on the server |
Use systemd on the cloud server to keep the application always running:
Example
[Unit]
Description=Example Chat Web App
After=network.target ollama.service
[Service]
WorkingDirectory=/opt/example-chat
ExecStart=/usr/bin/python3 server.py
Restart=always
RestartSec=3
[Install]
WantedBy=multi-user.target
Set up long-term running:
sudo systemctl daemon-reload sudo systemctl enable --now example-chat
Other ExtensionsThe complete security checklist for external deployment (loopback port, reverse proxy authentication, exposure surface self-check) is in the Security and Compliance chapter; be sure to go through it before going live.