DeepSeek launched experimental DeepSeek-V4-Flash-Vision-Exp on its API platform. The company said the model matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge, while multimodal agent benchmarks jump substantially over V4-Flash and land close to Opus 4.8. Developers call it with model='deepseek-v4-flash-vision-exp'. It supports Chat Completions, Messages, and Responses, with mixed text-plus-image input via base64, external URLs, or the Files API. Images are billed at up to 384 tokens each at V4-Flash pricing. DeepSeek Harness 0.1.1 shipped the same day with out-of-the-box support, and a free Files API lets teams upload an image once and reuse it by file_id. For coding agents, the relevant use cases are screenshot-to-fix, visual UI diffs, and tool-using workflows that need to see the page or IDE rather than only the repo.
Key Takeaways
- ✓Multimodal agent benchmarks jump over V4-Flash and are claimed near Opus 4.8.
- ✓Chat Completions, Messages, and Responses compatibility lowers swap-in cost for existing agent stacks.
- ✓Images bill at V4-Flash rates (≤384 tokens each); a free Files API reuses uploads by file_id.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.