AI Multimodal
AI can not only talk, but also see images, hear sounds, and understand videos—master multimodal tools and technologies.
For a long time, AI was single-sensory.
Text AI could only process text, such as asking ChatGPT to write articles.
-
Image AI could only process pictures, such as asking Midjourney to draw images.
-
Audio AI could only process sound, such as asking Whisper to transcribe speech.
But the real world is multimodal—when we read books we see illustrations, when watching movies there are visuals and sound, and when communicating with others we observe expressions and tone.
Multimodal AI is AI that can simultaneously understand and generate multiple types of content, like giving AI eyes, ears, and a mouth, allowing it to interact with the world in a way closer to humans.
A practical example: you give AI a photo of a shopping receipt, it can recognize the text on it (OCR), understand what items were purchased, help you organize it into a table, and even read it aloud to you. This is a typical multimodal task.
What is Multimodal AI
First, let's clarify a few basic concepts.
Single-modal vs Multimodal
Modality refers to the form in which information is expressed.
Text, images, audio, and video are all different modalities.
| Type | Description | Typical Products |
|---|---|---|
| Single-modal AI | Can only process one modality | Early GPT-3 (text only), Stable Diffusion (images only) |
| Multimodal AI | Can process multiple modalities simultaneously | GPT-4o、Claude 3、Gemini |
The core breakthrough of multimodal AI is:It can convert information from different modalities into a common language for understanding。
For example, when seeing an image of a "cat", it can convert the image into a vector (a set of numbers); when seeing the word "cat", it can also convert it into a vector. These two vectors are close in mathematical space because they represent the same concept.
Core Challenges of Multimodal AI
Multimodality sounds simple, but in practice it is very difficult.
The first challenge is "alignment"—how to make the "cat" in an image and the "cat" in text appear as the same thing to the model?
-
The second challenge is "fusion"—when seeing an image and a piece of text at the same time, how do you combine their information?
-
The third challenge is "generation"—how to generate an image based on a text description, or generate a text description based on an image?
-
Fortunately, these problems are being gradually solved. After 2024, almost all mainstream large models are multimodal.
Image Understanding (Vision)
Image understanding is about making AI understand images—describing content, recognizing text, analyzing charts, and spotting details.
Mainstream large models now support image input.
There are usually two ways to use it:
One is to upload the image file directly, for example by clicking the image icon in the ChatGPT web version to upload.
-
Another way is to embed the image in the API request using Base64 encoding, suitable for programmatic calls.
| Model | Image input methods | Supported image formats |
|---|---|---|
| GPT | URL, Base64, direct upload | JPG、PNG、WEBP、GIF |
| Claude | Base64, direct upload | JPG、PNG、WEBP、GIF |
| Gemini | Direct upload, Google Drive | JPG、PNG、WEBP |
Common vision tasks include:
Image description — What is in this picture?
-
OCR — Extract the text from an image.
-
Chart analysis — What is this bar chart telling us?
-
Detail detection — Help me check this design draft for issues.
Hands-on: Calling the GPT-4o Vision API with Python
Let's look at a complete example — using the API to analyze an image.
Example
# File: example_vision_demo.py
# Function: Analyze images using the GPT-4o Vision API
# ============================================
import base64
import requests
import os
# Configure API key (please replace with your real key)
# Get it from: https://platform.openai.com/api-keys
OPENAI_API_KEY = "sk-your-api-key-here"
OPENAI_API_URL = "https://api.openai.com/v1/chat/completions"
def encode_image_to_base64(image_path: str) -> str:
"""
Encode the local image file as a Base64 string
This is the transfer format required by the API
"""
with open(image_path, "rb") as image_file:
# Read the binary file content and encode it with Base64
base64_data = base64.b64encode(image_file.read()).decode("utf-8")
return base64_data
def analyze_image_with_gpt4o(
image_path: str,
prompt: str = "Please describe the content of this image in detail.",
model: str = "gpt-4o"
) -> str:
"""
Use the GPT-4o Vision API to analyze the image
Parameter description:
image_path: Local image file path (e.g., "example_test.jpg")
prompt: The question you want to ask or the task you want the AI to perform
model: The model name to use (default gpt-4o)
Returns:
The AI's analysis result text
"""
# Step 1: Encode the image as Base64
base64_image = encode_image_to_base64(image_path)
# Step 2: Build the request headers
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {OPENAI_API_KEY}"
}
# Step 3: Build the request body
# Note: The image content is passed via the image_url field, in the format "data:image/jpeg;base64,..."
payload = {
"model": model,
"messages": [
{
"role": "user",
"content": [
# The first part is the text prompt
{"type": "text", "text": prompt},
# The second part is the image
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
}
],
"max_tokens": 1000 # Limit output length to avoid high costs
}
# Step 4: Send the API request
print(f"Analyzing image: {image_path} ...")
response = requests.post(OPENAI_API_URL, headers=headers, json=payload)
# Step 5: Handle the response
if response.status_code == 200:
result = response.json()
# Extract the AI's answer
answer = result["choices"][0]["message"]["content"]
return answer
else:
# Print detailed error information when an error occurs
error_msg = f"Request failed: {response.status_code} - {response.text}"
print(error_msg)
return error_msg
def extract_text_from_image(image_path: str) -> str:
"""
Specially for OCR: Extract text from images
"""
prompt = """
Please extract all text content from this image.
Requirements:
1. Accurately reproduce the text, do not omit anything
2. Keep the original paragraphs and formatting
3. If there are tables, output them in table form
4. If there is no text in the image, please state so
This is the OCR test from the EXAMPLE tutorial.
"""
return analyze_image_with_gpt4o(image_path, prompt)
def analyze_chart(image_path: str) -> str:
"""
Specially for chart analysis
"""
prompt = """
Please analyze this chart.
Please answer the following questions:
1. What type of chart is this (bar chart, line chart, pie chart, etc.)?
2. What are the chart title and the meanings of the axes?
3. What is the main message conveyed by the chart?
4. Are there any notable trends or anomalies?
This is the chart analysis test from the EXAMPLE tutorial.
"""
return analyze_image_with_gpt4o(image_path, prompt)
# ============================================
# Main program: Demonstrating different vision analysis tasks
# ============================================
if __name__ == "__main__":
# Assume we have a test image (you need to prepare a real image)
# It can be: product photo, screenshot, document photo, chart, etc.
test_image = "example_test_image.jpg"
# Check if the image exists
if not os.path.exists(test_image):
print(f"Tip: Please prepare an image, name it {test_image}, and place it in the current directory")
print("Or modify the image path in the code")
else:
# Task 1: General image description
print("\n" + "="*50)
print("Task 1: Image description")
print("="*50)
result1 = analyze_image_with_gpt4o(
test_image,
"Please describe this image in detail, including the subject, background, colors, composition, etc."
)
print(result1)
# Task 2: OCR text extraction
print("\n" + "="*50)
print("Task 2: OCR Text Extraction")
print("="*50)
result2 = extract_text_from_image(test_image)
print(result2)
# Task 3: Chart Analysis (if the image is a chart)
print("\n" + "="*50)
print("Task 3: Chart Analysis")
print("="*50)
result3 = analyze_chart(test_image)
print(result3)
print("\n" + "="*50)
print("EXAMPLE tutorial demo complete!")
print("="*50)
Before running this program, you need to:
-
1. Install dependencies:
pip install requests -
2. Prepare a test image and name it
example_test_image.jpg -
3. Replace the API key with your own
Cost note: GPT-4o Vision is billed by image size. A 1024x1024 image costs about $0.00765. Smaller images are cheaper. For debugging, use low-resolution images; for production, use high resolution.
Text-to-Image
Text-to-Image lets AI generate images from text descriptions. This is another very mature multimodal application area.
Introduction to Diffusion Model Principles
Today's text-to-image tools basically use diffusion model technology.
A simple way to understand how it works:
Training phase: Show the model a clear image, then gradually add noise to it until it becomes a completely noisy image. The model learns how to turn the noisy image back into a clear one step by step.
-
Generation phase: Give the model a random noise image, and let it denoise step by step according to the text prompt, finally generating a clear image that matches the description.
This process usually takes 20-50 steps, so generating an image takes a few seconds.
Advanced Midjourney Usage Tips
Midjourney is currently one of the most popular text-to-image tools, used via Discord.
A high-quality prompt usually includes these elements:
-
Subject description — "a cat wearing a hat"
-
Style — "Studio Ghibli style", "oil painting", "3D render"
-
Lighting — "natural light", "cinematic lighting", "cyberpunk neon lights"
-
Composition — "close-up", "wide angle", "top-down view"
-
Quality — "8K", "ultra-high detail", "photographic quality"
| Parameter | Description | Example |
|---|---|---|
| --ar | Set aspect ratio | --ar 16:9 (video), --ar 3:4 (mobile) |
| --v | Specify model version | --v 6.0 (latest version) |
| --s | Stylization level (0-1000) | --s 750 (strong artistic feel) |
| --q | Render quality (0.25-2) | --q 2 (higher quality, slower) |
| --iw | Reference image weight | --iw 1.5 (more like the reference image) |
Local Deployment of Stable Diffusion
If you want full control, you can run Stable Diffusion locally:https://github.com/compvis/stable-diffusion。
The most commonly used tool is Automatic1111's WebUI.
The steps are roughly:
-
1. Install Python 3.10
-
2. Download Stable Diffusion WebUI
-
3. Download the model file (.safetensors format)
-
4. Run webui.bat (Windows) or webui.sh (Mac/Linux)
Advantages of local deployment: generation is free, no content restrictions, can use LoRA to fine-tune styles, and can use ControlNet for precise composition control.
Disadvantages: you need a GPU with at least 6GB of VRAM, and setup is somewhat complicated.
Hands-on: Generating Images with the DALL·E 3 API
If you want a simple API way to generate images, DALL·E 3 is a good choice.
Example
# File: example_dalle3_demo.py
# Function: Generate images using the DALL·E 3 API
# ============================================
import requests
import os
from datetime import datetime
# Configure the API key (please replace with your real key)
OPENAI_API_KEY = "sk-your-api-key-here"
DALLE3_API_URL = "https://api.openai.com/v1/images/generations"
def generate_image_with_dalle3(
prompt: str,
size: str = "1024x1024",
quality: str = "standard",
style: str = "vivid",
save_dir: str = "example_generated_images"
) -> dict:
"""
Generating Images with the DALL·E 3 API
Parameter descriptions:
prompt: text description, the more detailed the better
size: image size
Options: 1024x1024, 1024x1792 (portrait), 1792x1024 (landscape)
quality: image quality
Options: standard (standard), hd (high definition, more expensive)
style: style
Optional: vivid (vivid, more artistic), natural (natural, more realistic)
save_dir: image save directory
Returns:
A dictionary containing the image URL and save path
"""
# Step 1: Prepare the request headers
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {OPENAI_API_KEY}"
}
# Step 2: Prepare the request body
payload = {
"model": "dall-e-3",
"prompt": prompt,
"n": 1, # DALL·E 3 can only generate 1 image at a time
"size": size,
"quality": quality,
"style": style
}
# Step 3: Send the request
print(f"Generating image, prompt: {prompt[:50]}...")
response = requests.post(DALLE3_API_URL, headers=headers, json=payload)
# Step 4: Handle the response
if response.status_code == 200:
result = response.json()
image_url = result["data"][0]["url"]
revised_prompt = result["data"][0].get("revised_prompt", prompt)
print(f"DALL·E 3 auto-optimized prompt: {revised_prompt}")
# Step 5: Download and save the image
if not os.path.exists(save_dir):
os.makedirs(save_dir)
# Generate a filename using a timestamp
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
filename = f"example_dalle3_{timestamp}.png"
filepath = os.path.join(save_dir, filename)
# Download the image
image_response = requests.get(image_url)
if image_response.status_code == 200:
with open(filepath, "wb") as f:
f.write(image_response.content)
print(f"Image saved to: {filepath}")
return {
"image_url": image_url,
"filepath": filepath,
"revised_prompt": revised_prompt
}
else:
error_msg = f"Request failed: {response.status_code} - {response.text}"
print(error_msg)
return {"error": error_msg}
def generate_example_logo_demo():
"""Demo: Generate a EXAMPLE-style Logo"""
prompt = """
A cute and friendly robot mascot for EXAMPLE tutorial website.
Blue and green color scheme, modern flat design, white background.
The robot looks helpful and intelligent, with big friendly eyes.
Simple, clean, suitable for a tech education brand.
"""
# Chinese prompts work too; DALL·E 3 supports multiple languages
prompt = """
Design a cute and friendly robot mascot for the EXAMPLE tutorial website.
Blue-green color scheme, modern flat design, white background.
The robot looks helpful and smart, with big friendly eyes.
Simple and clean, suitable for a tech education brand.
"""
return generate_image_with_dalle3(
prompt=prompt,
size="1024x1024",
quality="hd",
style="vivid"
)
def generate_book_cover_demo():
"""Demo: Generate a book cover"""
prompt = """
Book cover design for "AI Tutorial for Beginners".
Modern minimalist style, deep blue and purple gradient background.
The title text is elegant and readable.
Subtle tech elements: floating geometric shapes, faint circuit patterns.
Professional, trustworthy, suitable for a programming book.
"""
return generate_image_with_dalle3(
prompt=prompt,
size="1024x1792", # Portrait orientation
quality="hd",
style="natural"
)
def generate_banner_demo():
"""Demo: Generate a banner ad"""
prompt = """
Web banner design for a coding tutorial website (EXAMPLE).
16:9 wide format. Modern gradient background from teal to blue.
Clean layout with space for text.
Abstract tech elements: floating code snippets, subtle grid lines.
Professional and inviting.
"""
return generate_image_with_dalle3(
prompt=prompt,
size="1792x1024", # Landscape orientation
quality="standard",
style="vivid"
)
# ============================================
# Main program: demonstrate different generation tasks
# ============================================
if __name__ == "__main__":
print("="*60)
print("EXAMPLE DALL·E 3 Image Generation Demo")
print("="*60)
# Task 1: Generate Logo
print("\n[Task 1] Generating Logo ...")
result1 = generate_example_logo_demo()
# Task 2: Generate book cover
print("\n[Task 2] Generating book cover ...")
result2 = generate_book_cover_demo()
# Task 3: Generate banner
print("\n[Task 3] Generating banner ...")
result3 = generate_banner_demo()
print("\n" + "="*60)
print("All tasks completed! Check the example_generated_images directory")
print("="*60)
DALL·E 3's strengths: strong understanding, doesn't break hands and faces, and generates text well.
But it also has limitations: it can only generate 1 image at a time, cannot precisely control style, and cannot be fine-tuned like Stable Diffusion.
Copyright notice: The current consensus is that for AI-generated images, the prompt writer has usage rights but does not own the copyright (because copyright offices require human authorship). For commercial use, it is recommended to confirm the tool's terms of use.
Video AI
Video AI is divided into two categories: video generation and video understanding.
Compared to images, video is technically more difficult—because video is a combination of "time + space," requiring every frame to be coherent and movements to look natural.
Comparison of AI Video Generation Tools
2024 was the first year of AI video, with multiple companies launching video generation products.
| Tool | Developer | Features | Typical duration |
|---|---|---|---|
| Sora | OpenAI | Industry-leading image detail and spatial scene understanding, realistic physical logic, not yet fully open for public beta | Up to 60 seconds |
| Runway Gen-4 | Runway | All-in-one creation platform, integrating text/image-to-video, AI post-editing, motion brush, and VFX effects; the top choice for overseas professional creators | 10–20 seconds |
| Kling (Keling AI) | Kuaishou | Excellent adaptation to Chinese prompts, smooth character movements, stable characters across shots, high cost-effectiveness in China, supports multi-segment narrative stitching | 5–30 seconds, up to 2 minutes per segment |
| Pika | Pika Labs | Extremely expressive in anime and cinematic illustration styles, consistent character design, supports video style redrawing and frame interpolation | 5–10 seconds |
| Veo | Cinematic lighting and camera movement, stable long-take narrative, natural physical rendering, not yet publicly released | Over 60 seconds | |
| Jimeng AI (Seedance) | ByteDance | Deep integration with the CapCut/Douyin ecosystem, precise Chinese semantic understanding, audio-visual lip sync, supports automatic storyboarding from scripts | 5–15 seconds, multiple segments stitched into a complete film |
| Tongyi Wanxiang Video | Alibaba DAMO Academy | Open-source base model, stable multi-subject interaction, supports text generation in images, can be deployed locally for inference. | 2–15 seconds |
| Hailuo AI (Hailuo) | MiniMax (Xiyu) | Lightweight ultra-fast generation; suitable for converting posters into dynamic short videos and social media flash content, with outstanding anime-style rendering. | 4–10 seconds |
| Luma Dream Machine | Luma AI | Top-notch physics dynamics simulation; realistic effects for water, fire, cloth, and explosions, with cinematic camera movement. | 5–12 seconds |
| Zhiying | Tencent | Focuses on AI editing and digital-human videos, with massive short-video templates, one-stop subtitles, dubbing, and matting, and low commercial barriers. | No limit on video length; AI-generated clips are 5–15 seconds. |
| Haiyi AI | Haiyi Technology | Supports 4K/60fps high-frame-rate output, multi-image linked video generation, comes with a complete production workstation, and provides ample time-limited free quota. | Single segment up to 30 seconds. |
AI video is still in its early stages:
-
Advantages: quick creativity, low cost, and the ability to achieve effects that are very difficult with live-action shooting.
-
Disadvantages: limited duration, details easily break down, and physics logic is sometimes incorrect (e.g., objects passing through each other, wrong number of fingers).
-
Suitable scenarios: concept demonstrations, short-video material, e-commerce ads, game animations.
-
Unsuitable scenarios: long films requiring precise control, serious news, legal evidence.
Video Understanding and Summarization
Besides generating videos, AI can also understand videos.
Common tasks in video understanding:
-
Video summarization — "What is this video about? Summarize in 3 sentences."
-
Content retrieval — "Find the clip in the video where someone is riding a bicycle."
-
Question answering — "What did the third person in the video say?"
The typical technical approach is: extract frames from the video (e.g., 1 frame per second), then use a vision model to look at each frame one by one, and finally integrate the information.
Future Trends
The development direction of video AI is clear:
-
Longer — from 10 seconds to 1 minute, and then to 10 minutes.
-
More controllable — able to precisely control every shot using storyboards.
-
More consistent — the same character stays consistent across different shots.
-
More practical — from looking cool to solving real-world problems.
-
What to expect: in a few years, everyone may be able to create film-like video works with AI.
Audio and Speech AI
Audio AI mainly includes: speech-to-text (ASR), text-to-speech (TTS), and music generation.
This is a fairly mature field.
Speech-to-Text (Whisper API)
Whisper is OpenAI's open-source speech recognition model. It performs very well and supports 99 languages.
Example
# File: example_whisper_demo.py
# Function: Use Whisper API to transcribe speech
# ============================================
import requests
import os
# Configure API key
OPENAI_API_KEY = "sk-your-api-key-here"
WHISPER_API_URL = "https://api.openai.com/v1/audio/transcriptions"
def transcribe_audio_with_whisper(
audio_path: str,
language: str = "zh", # zh=Chinese, en=English, leave empty for auto-detect
prompt: str = "", # Optional: provide some terminology hints to improve accuracy
temperature: float = 0.0
) -> dict:
"""
Use Whisper API to transcribe audio files
Parameter description:
audio_path: audio file path (supports mp3, wav, m4a, etc.)
language: language code (optional, auto-detected if not specified)
Common codes: zh (Chinese), en (English), ja (Japanese)
prompt: prompt text (optional; write possible technical terms or names here)
temperature: sampling temperature (0.0 is most deterministic, 1.0 is more diverse)
Return:
A dictionary containing the transcribed text
"""
# Step 1: Check if the file exists
if not os.path.exists(audio_path):
return {"error": f"File not found: {audio_path}"}
# Step 2: Prepare the request headers
headers = {
"Authorization": f"Bearer {OPENAI_API_KEY}"
}
# Step 3: Prepare the request body (note: audio must use multipart/form-data format)
files = {
"file": open(audio_path, "rb")
}
data = {
"model": "whisper-1",
"temperature": temperature
}
if language:
data["language"] = language
if prompt:
data["prompt"] = prompt
# Step 4: Send the request
print(f"Transcribing audio: {audio_path} ...")
response = requests.post(WHISPER_API_URL, headers=headers, files=files, data=data)
# Step 5: Close the file
files["file"].close()
# Step 6: Process the response
if response.status_code == 200:
result = response.json()
text = result["text"]
print(f"Transcription complete, a total of {len(text)} characters.")
return {
"text": text,
"language": language
}
else:
error_msg = f"Request failed: {response.status_code} - {response.text}"
print(error_msg)
return {"error": error_msg}
def transcribe_with_timestamps(audio_path: str) -> dict:
"""
Timestamped transcription (using verbose_json format)
"""
headers = {
"Authorization": f"Bearer {OPENAI_API_KEY}"
}
files = {"file": open(audio_path, "rb")}
data = {
"model": "whisper-1",
"response_format": "verbose_json", # Returns detailed information, including timestamps
"timestamp_granularities": ["segment", "word"] # Supports segment-level and word-level timestamps
}
response = requests.post(WHISPER_API_URL, headers=headers, files=files, data=data)
files["file"].close()
if response.status_code == 200:
result = response.json()
return result
else:
return {"error": response.text}
def translate_audio_to_english(audio_path: str) -> dict:
"""
Translate audio from any language into English text
"""
headers = {
"Authorization": f"Bearer {OPENAI_API_KEY}"
}
files = {"file": open(audio_path, "rb")}
data = {
"model": "whisper-1",
"task": "translate" # Specifically specify: translation task
}
response = requests.post("https://api.openai.com/v1/audio/translations",
headers=headers, files=files, data=data)
files["file"].close()
if response.status_code == 200:
return response.json()
else:
return {"error": response.text}
def example_meeting_minutes_demo(audio_path: str) -> str:
"""
Hands-on: generating meeting minutes
Step 1: Transcribe the recording with Whisper
Step 2: Organize into minutes with GPT-4o
"""
# First transcribe
print("Step 1: Transcribing the meeting recording...")
transcribe_result = transcribe_audio_with_whisper(
audio_path,
language="zh",
prompt="This is a meeting of the EXAMPLE technical team, discussing product development, progress, technical solutions, etc."
)
if "error" in transcribe_result:
return transcribe_result["error"]
transcript = transcribe_result["text"]
# Then organize (simplified here; in a real project you can call the GPT-4o API)
print("Step 2: Organizing the meeting minutes...")
# For demonstration purposes, return the transcript text directly
# In a real project, you can send the transcript to GPT-4o and have it organize it
return transcript
# ============================================
# Main program
# ============================================
if __name__ == "__main__":
# Prepare a test audio file
test_audio = "example_test_audio.mp3"
if not os.path.exists(test_audio):
print(f"Hint: Prepare an audio file and name it {test_audio}")
print("Or record a voice memo on your phone and save it as MP3")
else:
# Task 1: Simple transcription
print("\n" + "="*50)
print("Task 1: Simple transcription")
print("="*50)
result1 = transcribe_audio_with_whisper(test_audio, language="zh")
if "text" in result1:
print("Transcription result:")
print(result1["text"])
# Task 2: Timestamped transcription
print("\n" + "="*50)
print("Task 2: Timestamped transcription")
print("="*50)
result2 = transcribe_with_timestamps(test_audio)
if "segments" in result2:
print("Segment information:")
for seg in result2["segments"][:3]: # Only print the first 3 segments
start = round(seg["start"], 2)
end = round(seg["end"], 2)
text = seg["text"]
print(f"[{start}-{end}] {text}")
print("\n" + "="*50)
print("EXAMPLE Whisper demo completed!")
print("="*50)
Whisper's advantages are: good multilingual support, not picky about accents, and automatic punctuation.
Typical application scenarios: meeting recording transcription, video subtitle generation, podcast transcripts, voice diaries.
Text-to-Speech (TTS)
Text-to-speech lets AI "read" articles.
OpenAI also provides a TTS API supporting multiple voice styles.
Other TTS tools: ElevenLabs (most human-like), Azure TTS (commercial-grade), Edge TTS (free), Coqui (open-source).
Music Generation (Suno, Udio)
Music generation has also achieved breakthroughs.
Suno and Udio are two representative products—you input lyrics and style descriptions, and they can generate a complete song with vocals and accompaniment.
Capability boundaries:
-
Can do: generate background music, write song demos, quickly try styles.
-
Limited in: fully replicating a specific song, precisely controlling notes, professional-grade post-production.
-
Copyright: terms vary by tool; please confirm before commercial use.
Introduction to Multimodal Model Architectures
Understanding a bit of the theory helps you use the tools better.
CLIP: The Foundation of Image-Text Alignment
CLIP (Contrastive Language-Image Pre-training) is a model published by OpenAI in 2021.
Its idea is simple: look at images and text simultaneously, and learn to associate them.
During training, give the model a batch ofimage + descriptionpairs, and let it learn:
-
1. Convert images into vectors
-
2. Convert text into vectors
-
3. Make paired image-text vectors as close as possible, and unpaired ones as far apart as possible
CLIP's significance lies in:It proved that "cross-modal understanding" is feasible。
Later models like Stable Diffusion and GPT-4o were all inspired by CLIP.
LLaVA: Open-Source Multimodal LLM
LLaVA (Large Language and Vision Assistant) is an open-source multimodal model.
Its architecture is very clear:
-
A vision encoder (e.g., ViT) — responsible for seeing images
-
A large language model (e.g., LLaMA) — responsible for speaking
-
A projection layer — converts image vectors into a format the language model can understand
Training is done in two steps:
-
Step 1: Pretraining alignment — make the vision encoder and language model speak the same language
-
Step 2: Instruction fine-tuning - train with image-question-answer data to teach it to answer according to instructions
LLaVA is a good starting point for learning multimodal technology - open-source code, clear architecture, good results.
Hands-on: Building an Image Analysis Tool
Let's integrate the previous knowledge and build a practical image analysis tool.
Example
# File: example_image_analyzer.py
# Function: a complete multimodal image analysis tool
# ============================================
import base64
import requests
import os
import json
from datetime import datetime
from typing import Optional, Dict, Any
class ExampleImageAnalyzer:
"""
EXAMPLE Image Analyzer - a packaged multimodal tool class
"""
def __init__(self, api_key: str, base_url: str = "https://api.openai.com/v1"):
"""
Initialize analyzer
Parameters:
api_key: OpenAI API key
base_url: API address (configurable proxy)
"""
self.api_key = api_key
self.base_url = base_url
self.headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {api_key}"
}
def _encode_image(self, image_path: str) -> str:
"""Encode the image as Base64"""
with open(image_path, "rb") as f:
return base64.b64encode(f.read()).decode("utf-8")
def _call_vision_api(
self,
image_path: str,
prompt: str,
model: str = "gpt-4o",
max_tokens: int = 2000
) -> Optional[str]:
"""Internal method: call the vision API"""
base64_image = self._encode_image(image_path)
payload = {
"model": model,
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": prompt},
{
"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{base64_image}"}
}
]
}
],
"max_tokens": max_tokens
}
url = f"{self.base_url}/chat/completions"
response = requests.post(url, headers=self.headers, json=payload)
if response.status_code == 200:
return response.json()["choices"][0]["message"]["content"]
else:
print(f"API error: {response.status_code} - {response.text}")
return None
def describe(self, image_path: str) -> Optional[str]:
"""Describe the image content in detail"""
prompt = """
Please describe this image in detail, including:
1. Main content (what people/objects are there)
2. Background environment
3. Colors and lighting
4. Overall atmosphere
Answer in Chinese, with a clear structure.
"""
return self._call_vision_api(image_path, prompt)
def extract_text(self, image_path: str) -> Optional[str]:
"""Extract all text from the image (OCR)"""
prompt = """
Please extract all text content from this image.
Requirements:
1. Reproduce accurately without omission
2. Preserve the original paragraph structure
3. If it is a table, output it in Markdown table format
4. If there is no text, explicitly state "No recognizable text in the image"
"""
return self._call_vision_api(image_path, prompt)
def analyze_ui(self, image_path: str) -> Optional[str]:
"""Analyze UI design mockup (screenshot)"""
prompt = """
Please analyze this UI design mockup:
1. What type of interface is this (APP, web page, mini-program, etc.)?
2. What are the main functional modules of the interface?
3. What is the design style (color scheme, layout, typography)?
4. What could be improved?
Answer in Chinese, listed in points.
"""
return self._call_vision_api(image_path, prompt)
def analyze_product(self, image_path: str) -> Optional[str]:
"""Analyze product images and generate e-commerce copy"""
prompt = """
Please analyze this product image and generate e-commerce copy.
The copy should include:
1. Product name and type
2. Appearance features (color, material, design)
3. An attractive short recommendation (30-50 characters)
Answer in Chinese, in a friendly and positive tone.
"""
return self._call_vision_api(image_path, prompt)
def classify(self, image_path: str, categories: list = None) -> Optional[str]:
"""Classify the image"""
if categories is None:
categories = [
"Portrait photo", "Landscape photo", "Food photo", "Product photo",
"Design mockup/screenshot", "Document/table", "Chart/data visualization", "Other"
]
cat_list = "、".join(categories)
prompt = f"""
Please determine which category this image belongs to.
Optional categories: {cat_list}
Only return the most matching category name, no other text.
"""
return self._call_vision_api(image_path, prompt)
def full_analysis(self, image_path: str, output_dir: str = "example_analysis") -> Dict[str, Any]:
"""
Full analysis: generate a detailed analysis report
Return a dict containing all analysis results
"""
print(f"Start analyzing image: {image_path}")
# Classify first
print("[1/5] Classifying...")
category = self.classify(image_path) or "Unknown"
# Description
print("[2/5] Describing...")
description = self.describe(image_path) or ""
# Extract text
print("[3/5] OCR in progress...")
text = self.extract_text(image_path) or ""
# Perform specialized analysis based on category
print("[4/5] Specialized analysis in progress...")
special_analysis = ""
if "design" in category or "screenshot" in category:
special_analysis = self.analyze_ui(image_path) or ""
elif "product" in category or "merchandise" in category:
special_analysis = self.analyze_product(image_path) or ""
# Save report
print("[5/5] Generating report...")
if not os.path.exists(output_dir):
os.makedirs(output_dir)
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
report = {
"image_file": os.path.basename(image_path),
"analysis_time": datetime.now().isoformat(),
"category": category,
"description": description,
"extracted_text": text,
"special_analysis": special_analysis
}
# Save JSON
json_path = os.path.join(output_dir, f"report_{timestamp}.json")
with open(json_path, "w", encoding="utf-8") as f:
json.dump(report, f, ensure_ascii=False, indent=2)
# Save readable text
txt_path = os.path.join(output_dir, f"report_{timestamp}.txt")
with open(txt_path, "w", encoding="utf-8") as f:
f.write("="*60 + "\n")
f.write("EXAMPLE Image Analysis Report\n")
f.write("="*60 + "\n\n")
f.write(f"Image: {report['image_file']}"\n")
f.write(f"Time: {report['analysis_time']}"\n")
f.write(f"Category: {report['category']}"\n\n")
f.write("-"*60 + "\n")
f.write("Image description:\n")
f.write("-"*60 + "\n")
f.write(report["description"] + "\n\n")
f.write("-"*60 + "\n")
f.write("Extracted text:\n")
f.write("-"*60 + "\n")
f.write(report["extracted_text"] + "\n\n")
if report["special_analysis"]:
f.write("-"*60 + "\n")
f.write("Special analysis:\n")
f.write("-"*60 + "\n")
f.write(report["special_analysis"] + "\n")
print(f"Analysis complete! Report saved to: {output_dir}")
return report
# ============================================
# Usage example
# ============================================
def main():
# Configure your API Key
api_key = "sk-your-api-key-here"
# Create an analyzer
analyzer = ExampleImageAnalyzer(api_key)
# Path of the image to analyze
image_path = "example_test_image.jpg"
if not os.path.exists(image_path):
print(f"Please prepare the image: {image_path}")
return
# Method 1: Use a single function individually
print("\n"Method 1: Use OCR alone")
print("-"*40)
text = analyzer.extract_text(image_path)
if text:
print("Extracted text:")
print(text)
# Method 2: Complete analysis
print("\n"Method 2: Complete analysis")
print("-"*40)
report = analyzer.full_analysis(image_path)
print(f"Classification result: {report['category']}")
print(f"Report saved")
if __name__ == "__main__":
main()
This tool encapsulates common image analysis functions and can be used directly in projects.
You can extend it as needed:
Add more analysis scenarios (e.g., medical imaging, industrial inspection, educational question analysis)
-
Add a simple web interface (using Flask or Gradio)
-
Support batch processing of images in folders
-
Store the results in a database