DeepSeek

V4 Flash API

Built for reasoning, coding, and agentic tool use, with stronger long-context, thinking-mode, and complex-task execution than DeepSeek V3.2. Compared with Chinese peers such as Qwen, GLM, and Kimi, it is especially competitive for math, programming, and logical reasoning tasks.

  • Long context
  • Tool use
  • Reasoning
  • Structured output
Use in Consoledeepseek-v4-flash

USD

Pricing

default
Standard Rate
Input$0.142/1M tokensCompletion Price$0.286002/1M tokensCache Read$0.019994/1M tokens

Public list prices are shown in USD. Final charges may vary by account group and usage tier.

API

Code examples

curl --location --request POST 'https://api.tokenhot.ai/v1/chat/completions' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data-raw '{
    "model": "deepseek-v4-flash",
    "messages": [
        {
            "role": "user",
            "content": "What model are you"
        }
    ]
}'

DeepSeek

Related models

FAQ

Frequently asked questions

What is V4 Flash best suited for?

V4 Flash is suited to the input, output, and capability types listed on this page. Test production workloads before choosing it for a critical system.

How is V4 Flash priced?

Pricing is shown in the pricing section above. Actual cost depends on usage volume, account group, and the request payload.

How can I call V4 Flash?

Use the model ID shown above with one of the supported API protocols. Code examples are provided when usage data is available.

Get started

Build with V4 Flash

Use in Console
TokenHot

The frontier intelligence gateway. One API. 127 models. 0.2s latency. Pay only for what you use.

All systems normal · 99.997% uptime

Company

© 2026 TokenHot Inc. — Built for builders.