Ongoing topic · Compiled September 11, 2026

DeepSeek V4.1 Flash: images, tools and new prices

How can screenshots, error messages and long documents become material for AI work? Follow image understanding, tool integrations, price changes and the release timeline.

01Reading images

A screenshot can become part of the question

Three kinds of material to ask about
  1. Event screenshot

    List the date, place and things to bring

  2. Lesson or work chart

    Explain the largest change

  3. Error screenshot

    Suggest what to check next

Suggested questions to use with your own images. Read the source

DeepSeek V4.1 Flash reads images together with text and answers in text. Beyond describing a picture, the official guide lists reading screenshot text and analyzing charts. An event page, a lesson chart or an error screenshot can accompany a specific question.

Start with a specific question: which dates and places appear in an event screenshot, where a chart changes most, or what an error message means. The image preserves the details; the question tells the model what to look for.

Notes and further sources

The questions and workflows are suggested uses; this article does not present model outputs from BitShovel tests.

02Working with tools

After reading, continue into the files

With tools attached, an answer can lead into the next step of a task. DeepSeek Harness asks users to configure a model and choose a workspace. It can read and edit files, run commands, delegate work and maintain a plan. A developer might begin by understanding an unfamiliar project, then continue into edits and checks.

This is where vision and coding meet: a screenshot explains what should change, while project files contain the implementation. Together, they can make a request clearer than a long verbal description of an interface. This is a suggested workflow; the connected tool performs the actual actions.

Notes and further sources

03Speed and usage

More reasoning can mean more output

Official reasoning benchmarks comparing task scores and average output usage at different effort levels.
Reasoning benchmarks: the solid line shows score and the dashed line output usage on separate axes. Open the full image for all three test groups.Official DeepSeek image · Original sourceView full image

Flash is designed for faster inference and higher throughput, but task duration also depends on input length, reasoning effort and tool interactions. Reasoning effort is adjustable. The technical report plots task scores alongside output usage: greater effort generally produces more output, while scores do not rise at every step.

Speed and cost become tangible in repeated revisions, long-document work and tasks with several steps. How much the model reads, reasons and writes accumulates across the whole job.

Notes and further sources

04Cost

Frequent small tasks are easier to budget

Official USD API price table for DeepSeek V4.1 Flash.
Official USD rates per million tokens. Peak uncached input is $0.30 and output $1.20; off-peak rates are half.Official DeepSeek image · Original sourceView full image

The official USD API rates effective September 10 are $0.30 per million uncached input tokens, $0.006 for cached input and $1.20 for output at peak times. Off-peak rates are half. Peak periods are Monday–Friday, 09:00–12:00 and 14:00–18:00 Beijing time; other hours are off-peak.

If a request is billed for 100,000 uncached input tokens and 10,000 output tokens, the cost is $0.042 at peak rates or $0.021 off-peak. This is a calculation from published rates, not a fixed word count or task price. For batch document processing or repeated revisions, the bill depends on actual input and output usage.

Notes and further sources

This example uses DeepSeek’s direct API USD rates. The Chinese edition uses the separately published CNY rates. Gateway, app subscription and tool charges are separate.

05Access

Start in an existing tool

The current direct DeepSeek API name is deepseek-flash. DeepSeek’s release also confirms integration with WorkBuddy, CodeBuddy and OpenCode. Existing users can check the models available in their tool and begin with a screenshot or project material. The connection between model and app determines which actions can follow an answer.

Vercel AI Gateway is another documented route. Its September 9 announcement describes requests combining text and images. The current model page uses deepseek/deepseek-v4.1-flash and provides playground and AI SDK entry points. Developers can start with their own material before integrating it into an application.

Notes and further sources

Vercel requires gateway credentials and uses its own billing rules; DeepSeek Harness remains a developer preview. This research checked public sources and entry points without running model inference.

06Timeline

From the vision experiment to current access

07What comes next

What we are following

  • Which examples of image understanding, file reading and real tasks are worth seeing?
  • Which everyday tools will adopt it, and how will using them change?
  • How will pricing, speed and long-task performance evolve?