Sunday, Aug 23 | --:--
Back to home

DeepSeek Ships V4-Flash-Vision-Exp, a Multimodal API Sibling of V4-Flash

On August 21, 2026, DeepSeek’s API changelog launched experimental DeepSeek-V4-Flash-Vision-Exp (`deepseek-v4-flash-vision-exp`): JPEG/PNG/GIF/WebP via base64, URL, or Files API file_id, on Chat Completions, Anthropic-compatible Messages, and Responses. Text-agent scores match official V4-Flash; DeepSeek says multimodal-agent results jump toward Claude Opus 4.8. Distinct from the August 13 V4-Pro GA and price hike.

Tech Insights Reporter 4 min read Hangzhou / Beijing
Cover illustration for DeepSeek Ships V4-Flash-Vision-Exp, a Multimodal API Sibling of V4-Flash

TLDR

DeepSeek on August 21, 2026 added DeepSeek-V4-Flash-Vision-Exp to the API (model='deepseek-v4-flash-vision-exp'). Images: JPEG / PNG / GIF / WebP via base64, public URL, or Files API file_id. Surfaces: Chat Completions, Anthropic-compatible Messages, and Responses. DeepSeek: pure-text agent/reasoning matches official V4-Flash; on vision-required agent benches the experimental model “delivers a significant leap,” “close to Opus-4.8.” Independent Arena / Artificial Analysis listings were not confirmed in the research window—treat as lab-claimed, too new for third-party boards.

Lab benches (changelog)

Bench V4-Flash-Vision-Exp
Terminal Bench 2.1 83.9
NL2Repo 57.7
DeepSWE 59.3
DSBench-Hard 63.6
AutomationBench (Public) 25.7
ApexBench (Pass@1) 36.5
Agents' Last Exam 27.3
Chartography 64.3
ZeroBench (Pass@5) 35.0

Text-agent numbers used DeepSeek Harness minimal mode, max effort, topp 0.95, temp 1.0. ApexBench and Agents' Last Exam: the text V4-Flash ignores multimodal elements—so those two are not a vision-vs-text comparison.

Independent rankings: not listed on Artificial Analysis / LMArena in this window. Do not invent an Intelligence Index.

Product-line de-dupe: not V4-Pro GA or the Aug 13 price/peak-off-peak announcement (effective Aug 16). This is a new multimodal experimental SKU.

Why this story matters

DeepSeek’s V4 family has been text-agent priced. A Flash-class vision endpoint, on the same three API dialects, is how Hangzhou keeps OpenClaw/Codex-style harnesses from routing screenshots to Gemini or GPT-5.6. Watch: whether Vision-Exp graduates without a price multiplier, and whether Chartography 64.3 holds once independent evals run the images.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading