ITHome · Published 2026-09-10

DeepSeek V4.1 Flash adds vision and lowers API prices

What material can the model read, and what will an API task cost? DeepSeek V4.1 Flash, released on September 10, combines native image understanding with lower API rates and a same-day tool update.

DeepSeek’s official comparison of four agent benchmarks, with V4.1 Flash shown in blue.
DeepSeek’s official agent benchmark figure: four tests with their own tasks and settings. It is not an image-understanding demonstration or a ranking for every task. · Open full-size image
DeepSeek’s official global KV cache chart: 3,514 bytes per token for V4 Flash and 890 bytes for V4.1 Flash.
Original DeepSeek model-card figure: the global KV cache needed to retain context is smaller. The storage reduction is not a measure of the API price reduction. · Open full-size image
Sources and text description

DeepSeek released V4.1 Flash on September 10 with native image and text input and text output. Screenshots, charts and pictures containing words can be supplied alongside a question. For example, a screenshot and a question about an interface can give the model both visual and written context. This is image understanding, not image generation.

The official model card lists a context limit of one million tokens and reports document-question-answering and visual evaluations. The benchmark chart covers four agent benchmarks selected by DeepSeek. Those scores describe the specified tests; they do not establish a success rate for your own files or workflow.

DeepSeek also lowered the API rates for V4.1 Flash. Per million tokens, off-peak input costs CNY 0.02 for a cache hit or CNY 1 for a cache miss; output costs CNY 4. Peak rates are CNY 0.04, CNY 2 and CNY 8 respectively. A cache hit means the service reused cached input, so the lowest input rate does not apply to every request.

For an illustrative workload of one million uncached input tokens and 100,000 output tokens, these rates give CNY 1.40 off-peak or CNY 2.80 at peak. A real task may require several calls and generate more output. Billing follows actual token use.

The architecture also reduces the storage needed for long context. The official chart shows global KV cache falling from 3,514 bytes per token for V4 Flash to 890 bytes, about one quarter of the previous size. This measures model cache storage. API charges still follow the published rates and actual usage.

On the same day, ITHome reported that DeepSeek Harness 0.1.5 added V4.1 Flash support, file uploads, sidebar file previews and long-session performance improvements. The model accepts richer material while the interface provides ways to submit and inspect files. BitShovel has also collected the Harness update; follow the related reading below to continue that thread.

Developers calling the API can use the model name deepseek-flash. DeepSeek says requests using the older deepseek-v4-flash and deepseek-v4-flash-vision-exp names are temporarily served by V4.1 Flash. Tools that still use those names should therefore account for the change in the underlying model.

This reading uses the official model card, API update and pricing page checked on September 10, 2026. Harness 0.1.5 features are attributed to ITHome’s same-day report. Peak hours are Monday to Friday, 09:00–12:00 and 14:00–18:00 Beijing time; all other times are off-peak. The official evaluations use specified reasoning effort, agent scaffolds and context settings. The original credited charts do not establish across-the-board superiority. Rates may change; consult the official pricing page for later updates.