LLM API Pricing

Qwen Models

Browse all 15 Qwen models available on TokenHot.

One API key gives you access to 15 Qwen models with live pricing and no separate provider account.

CONTEXT1M
Input$0.1492/M
Output$0.4476/M
An efficient multimodal workhorse in the Qwen family, with a native million-token context window and strong support for coding assistance, agentic workflows, and visual understanding such as charts. It fits long documents, codebases, high-concurrency applications, and tool-assisted tasks.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Try model
PER SEC$0.0675/s
The speed-first variant of Alibaba's Wan 3.0 line (early access). It shortens generation latency while retaining high-quality video output, making it a good fit for low-latency previews, batch drafts and fast-iterating short-video workflows. If ultimate visual control matters more than speed, the standard version is a better choice.
Input Type:
Output Type:
Multimodal Output
Try model
Qwen
PER SEC$0.0450/s
The standard video generation model of Alibaba's Wan 3.0 line (early access). It supports text-to-video, image-to-video and multi-modal reference inputs, generating clips up to roughly 30 seconds with synchronized audio in a single pass — suitable for ads, short drama and narrative content. Official release date not yet announced.
Input Type:
Output Type:
Multimodal Output
Qwen
CONTEXT1M
Input
$1.4200/M$1.7750/M
Output
$4.2640/M$5.3300/M
Qwen's most capable flagship multimodal reasoning model, built for long-horizon coding, professional work, and multi-stage agent tasks. It accepts text, images, and video, offers a 1M-token context window with up to 128K output, and supports function calling, structured outputs, web search, and a code interpreter. It is best used as a high-quality workhorse for complex tasks rather than for the lowest-cost, lowest-latency batch workloads.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Try model
PER SEC
$0.0024/s$0.0030/s
A quality-first image generation and editing model: it supports text-to-image and edits from 1–3 reference images, accepts prompts of up to about 4.5K tokens, and is aimed at dense layouts, fine 10px-scale text, 12 languages, and multiple fonts. Choose it over the standard variant when complex composition and detail fidelity matter more than speed or cost; use the standard variant for a faster, more economical workflow.
Input Type:
Output Type:
PER SEC
$0.0024/s$0.0030/s
The balanced standard tier of Qwen-Image 3.0 for image generation and editing: it supports text-to-image and edits from 1–3 reference images, prompts of up to about 4.5K tokens, up to six outputs, and 2K-class resolution. It fits everyday posters, interfaces, infographics, and repeated creative work; move to Pro when complex layouts and fine text fidelity matter more, and use this tier when speed and cost efficiency are the priority.
Input Type:
Output Type:
PER SEC$0.1350/s
Built for text-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
PER SEC$0.1350/s
Built for reference-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
PER SEC$0.1350/s
Built for image-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
Qwen
CONTEXT1M
Input
$0.2400/M$0.3000/M
Output
$0.9600/M$1.2000/M
Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Try model
Qwen
CONTEXT1M
Input
$1.4400/M$1.8000/M
Output
$4.3200/M$5.4000/M
Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Try model
CONTEXT8K
PER SEC$0.1350/s
Built for video editing, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
CONTEXT8K
PER SEC$0.1350/s
Built for text-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
CONTEXT8K
PER SEC$0.1350/s
Built for reference-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
CONTEXT8K
PER SEC$0.1350/s
Built for image-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output