DeepSeek Harness Multimodal

DeepSeek has launched an experimental multimodal vision understanding model, DeepSeek-V4-Flash-Vision-Exp, and simultaneously opened multimodal API services and a free Files API.

Install the latest version before use:

npm install -g @deepseek-ai/dsh@latest

After updating to the latest version, you can see that the model list already includes DeepSeek-V4-Flash-Vision-Exp:

Switch to the vision model, and then we can directly drag images or PPT files into the document to let it look at the content in the images ( Test image download):

What is multimodal

To understand multimodality, we first need to understand the word "modality." Modality refers to different forms of information representation: text is one modality, and images, audio, and video are each different modalities as well.

modalityCommon formsCorresponding AI capabilities
TextArticles, code, conversationsLarge Language Model (LLM)
imagePhotos, screenshots, design mockups, chartsVisual understanding model
audioSpeech, musicSpeech recognition and generation models
VideoShort videos, screen recordings, surveillance footageVideo understanding model

In the past, large language models only processed a single modality—text. Everything you sent them had to be converted into text first.

Multimodal models can simultaneously receive images and text, uniformly converting them into token sequences before processing them together.

多模态模型处理图片与文本的流程图

The key point: images are not "directly seen" by the model. Instead, a vision encoder first cuts them into patches and converts them into vectors, which are then concatenated with text tokens into the same sequence.

Images also consume tokens. One image takes up to 384 tokens—this is exactly the source of the multimodal billing rules mentioned later.

For Agents, multimodality brings a qualitative change, not just a nice-to-have enhancement.

Comparison ItemText-only modelMultimodal model
Input formText onlyMixed text and images
Troubleshoot code errorsRequires users to manually transcribe error messages into textDirectly send error screenshots
Reproduce the design mockupUnable to reference visual draftsWrite pages directly from design mockups
Analyze data chartsCannot read images within imagesLook at the image and give conclusions directly

V4-Flash-Vision-Exp: Multimodal vision model now available

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal vision understanding model, now available on the DeepSeek API platform.

Users can set viamodel='deepseek-v4-flash-vision-exp'Then you can access the model.

Its capability positioning can be summarized in one sentence: text capabilities without compromise, visual capabilities with a major leap.

In pure text capabilities (Agent, reasoning, world knowledge, etc.), it is on par with the official DeepSeek-V4-Flash release.

On Agent Benchmarks requiring visual understanding, it has achieved a significant leap over DeepSeek-V4-Flash, with multimodal Agent capabilities now approaching Opus-4.8.

Capability dimensionDeepSeek-V4-FlashDeepSeek-V4-Flash-Vision-Exp
Pure text Agent tasksRelease baselineOn par with the official version
Reasoning and world knowledgeRelease baselineOn par with the official version
Visual understanding Agent tasksNot supported; multimodal elements are ignored in evaluationA major leap, approaching Opus-4.8
Model positioningOfficial versionExperimental version

Evaluation methodology note: For Code Agent text tasks in public benchmark sets, DeepSeek series models are tested using DeepSeek Harness minimal mode as the framework, with max setting, temperature=1.0, topp=0.95.

In the ApexBench and Agents' Last Exam evaluations, the text model DeepSeek-V4-Flash ignores the multimodal elements in them.


Multimodal API Quick Start

The multimodal API supports three calling formats: Chat Completions, Messages, and Responses, making it easy to integrate with various Agent tools.

The three formats have identical capabilities—just choose based on the interface style you're familiar with.

Call formatInterface styleWho is it for?
Chat CompletionsOpenAI classic chat interfaceDevelopers who already have OpenAI SDK code
MessagesAnthropic Messages interface,base_url is https://api.deepseek.com/anthropicDevelopers who already have Anthropic-style code or toolchains
ResponsesOpenAI's new Responses APIDevelopers using the new SDK who prefer concise input structures

All three formats support mixed image-text input. Images can be passed in three ways: base64 inline, external URL, or Files API.

base64 内联、外部 URL 与 Files API 三种图片传入方式对比图

Passing methodRequest body sizeWhether an image hosting service is neededApplicable scenarios
base64 inlinebigNot RequiredLocal images, one-off small images
External URLsmallrequiresImages already deployed on publicly accessible servers
Files APIsmallNot RequiredThe same image used multiple times, high-frequency batch tasks

Method 1: base64 inline

After encoding a local image as a base64 string, write it directly into the request body as a data URL.

Example

# File path: vision_base64_demo.py
# Dependency: pip install openai
import base64
from openai import OpenAI

# DeepSeek endpoint is compatible with OpenAI SDK, just change base_url to switch
client = OpenAI(
    api_key="sk-your-secret-key",                  # Required: replace with your own DeepSeek API key
    base_url="https://api.deepseek.com"    # Required: DeepSeek official endpoint
)

# Read local image, encode into base64 string
with open("example-logo.png", "rb") as f:
    b64 = base64.b64encode(f.read()).decode("utf-8")

response = client.chat.completions.create(
    model="deepseek-v4-flash-vision-exp",  # Required: Multimodal visual understanding model
    messages=[
        {
            "role": "user",
            "content": [
                # Text and images are interleaved in array order, and the model understands them in order.
                {"type": "text", "text": "What is the text in the image?"},
                {
                    "type": "image_url",
                    # base64 inline: data:image format;base64,encoded content
                    "image_url": {"url": f"data:image/png;base64,{b64}"}
                }
            ]
        }
    ],
    stream=False                           # Optional: whether to stream output, default False
)

print(response.choices[0].message.content)

Run output:

图片中的文字是 "EXAMPLE"。

Method 2: External URL

Images are deployed on publicly accessible servers, and the request only carries a URL string, keeping the request body very small.

Example

# External URL method: ${DEEPSEEK_API_KEY} is an environment variable exported in advance
curl https://api.deepseek.com/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ${DEEPSEEK_API_KEY}" \
  -d '{
  "model": "deepseek-v4-flash-vision-exp",
  "messages": [
    {
      "role": "user",
      "content": [
{"type": "text", "text": "Summarize the trend of this chart in one paragraph."},
        {"type": "image_url", "image_url": {"url": "https://static.example.com/images/chart-demo.png"}}
      ]
    }
  ]
}'

The URL must be a public address accessible by the platform. Internal network addresses and local loopback addresses cannot be fetched.

Both base64 and URL methods support adding an optional detail field in image_url to control image resolution precision.

detail valueBehaviorApplicable scenarios
lowScale to 512×512 before parsing, lower token consumptionOnly need a rough look: identify the image type, find the main subject
highEquivalent to originalNeed to see details clearly
originalAnalysis preserving original image precisionRead text, view small icons, analyze charts
autoCurrently equivalent to originalDefault value, behavior when detail is not passed

Setting detail to low when running screenshot-type tasks in batches is the most direct way to control multimodal costs.

Messages and Responses format

Those who already have Anthropic or new-style OpenAI code can switch to DeepSeek multimodal almost seamlessly.

Messages format uses image content blocks, and images are passed through the source field.

Example

# Messages format: aligned with Anthropic Messages API style
# The endpoint uses DeepSeek's Anthropic-compatible path /anthropic, and the header fields are also consistent with Anthropic
curl https://api.deepseek.com/anthropic/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: ${DEEPSEEK_API_KEY}" \
  -d '{
  "model": "deepseek-v4-flash-vision-exp",
  "max_tokens": 1024,
  "messages": [
    {
      "role": "user",
      "content": [
        {
          "type": "image",
          "source": {
            "type": "url",
            "url": "https://static.example.com/images/code-screenshot.png"
          }
        },
{"type": "text", "text": "This code will report an error when run; which line is the problem?"}
      ]
    }
  ]
}'

The Responses format uses an input array, with input_text for text and input_image for images.

Example

# File path: vision_responses_demo.py
# The client creation method is the same as the base64 example above.
response = client.responses.create(
    model="deepseek-v4-flash-vision-exp",  # Required: Multimodal visual understanding model
    input=[
        {
            "role": "user",
            "content": [
                # Use input_text for text and input_image for images
                {"type": "input_text", "text": "Organize the table in the screenshot into Markdown"},
                {
                    "type": "input_image",
                    "image_url": "https://static.example.com/images/table.png"
                }
            ]
        }
    ]
)

# For the Responses format, simply use output_text
print(response.output_text)

Files API: upload first, then reference

The Files API is now open, and this API is free of charge.

Users can upload images to the platform first, then reference them via file_id in requests.

This has two direct benefits: requests no longer carry large chunks of base64, saving request bandwidth; the same image does not need to be uploaded repeatedly across multiple requests.

Step 1: Upload the image and get the file_id.

Example

# Upload image: purpose currently only supports user_data
# The upload itself is free and does not incur any charges.
curl https://api.deepseek.com/files \
  -H "Authorization: Bearer ${DEEPSEEK_API_KEY}" \
  -F "purpose=user_data" \
  -F "file=@example-chart.png"

After successful upload, the platform returns the file's metadata.

Example

{
  "id": "file-api-qw3x8v2m6n",
  "object": "file",
  "filename": "example-chart.png",
  "purpose": "user_data",
  "bytes": 154320,
  "created_at": 1787300000,
  "expires_at": 1789900000
}
FieldsDescription
idFile identifier, starts with file-api-, used to reference the image in subsequent requests
filenameOriginal file name at upload time
purposeFile purpose, currently only supports user_data
bytesFile size, in bytes
created_atUpload time, Unix timestamp
expires_atExpiration time, optional, only returned if a validity period was set during upload

During upload, you can also use the expires_after parameter to specify a validity period, ranging from 1 hour to 30 days. If not passed, the file remains valid permanently.

Step 2: Reference the image using file_id in the multimodal request.

{
  "model": "deepseek-v4-flash-vision-exp",
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": "这个图表的整体趋势是什么?"},
        {"type": "file", "file_id": "file-api-qw3x8v2m6n"}
      ]
    }
  ]
}

When should the Files API be used?

When the same image needs to be reused across multiple requests (for example, Agent tasks running in batches on the same set of screenshots), uploading once and referencing everywhere saves more bandwidth and is faster than passing base64 each time.

The file block also supports file_data for direct inline base64 (mutually exclusive with file_id, can be combined with the filename field), suitable for scenarios where you occasionally want to skip the upload step.

When referencing a file_id via the Anthropic-compatible Messages endpoint, you need to additionally include the request header anthropic-beta: files-api-2025-04-14.


Billing instructions and notes

Multimodal requests are billed by token. Images are converted into tokens and then participate in billing.

Billing itemRules
Image billingImages are converted to tokens and billed by token
Maximum per imageUp to 384 tokens
PricingConsistent with the DeepSeek-V4-Flash model
Files APIFree

Image tokens occupy context window space. For multi-image tasks, watch the total consumption: 10 images can occupy up to 3840 tokens.

Images are automatically scaled before entering the model: images smaller than about 384×384 are enlarged, while larger images are shrunk to roughly the equivalent total pixels of 800×800.

Therefore, token consumption only depends on the resized dimensions; a 2000×2000 image and a 5000×5000 image consume the same amount; in multi-image requests, each image is calculated independently under the same rule.

base64 noticeably enlarges the request body. For large images or high-frequency calls, consider using external URLs or the Files API instead.

Multimodal requests also have some hard limits. Exceeding them will directly return a 400 error.

Restriction itemRules
Image formatJPEG, PNG, GIF, WebP; determined by actual file content, not by extension name
External URL lengthNo more than 8192 characters, and the image must be downloadable within 60 seconds
Request body sizeNo more than 48 MiB
Single image sizebase64 and URL methods max 32 MiB, Files API max 64 MiB
Number of images per requestUp to 600 images
Maximum image side length8192 pixels; when a request contains 15 or more images, it drops to 4096 pixels
Image positionCan only be placed in user messages. Including images in system and assistant messages will raise an error.

V4-Flash-Vision-Exp is an experimental model. Capabilities and service specifications may change with version iterations. For critical production pipelines, it is recommended to also evaluate official models.

For more complete usage instructions, refer to the official API documentation:API Guide - Image Understanding、API Guide - Files API。


Summary

V4-Flash-Vision-Exp brings multimodal capabilities into DeepSeek's Agent ecosystem with the positioning of "text without compromise, vision with a major leap."

For developers, the cost of getting started is very low: just change one model name, and Agents in DeepSeek Harness can understand screenshots, design mockups, and charts.

It is recommended to first try the three official examples to get a feel, then bring in your own business images and get your first "see the image and do the work" Agent task running.

other extensions