DeepSeek visual model multimodal Agent capabilities approach Opus 4.8, with a score of 2:2 in 4 items
DeepSeek announced the first batch of Agent scores for V4-Flash-Vision-Exp. In the evaluation of 4 multimodal Agents, it achieved 2 wins and 2 losses against Opus 4.8: Agents' Last Exam 27.3 vs 25.7, ZeroBench 35.0 vs 34.0; ApexBench 36.5 vs 39.4, Chartography 64.3 vs 65.0.
Compared to the pure text version V4-Flash-0731, the visual version improved from 26.2 to 36.5 in ApexBench, and from 25.2 to 27.3 in Agents' Last Exam. The official note states that the pure text version ignores multimodal content, thus mainly reflecting the capabilities after adding visual input. After adding visual input, the text Agent's capabilities did not significantly decline, with 6 out of 7 text evaluations higher than V4-Flash-0731, such as DeepSWE rising from 54.4 to 59.3 (surpassing Opus 4.8's 58.0), and Toolathlon at 75.9 nearly matching Opus 4.8's 76.2.
These results come from DeepSeek's official self-testing, not from a third-party independent ranking; the public Code Agent text task used DeepSeek Harness in minimalist mode, with the inference setting at max.







