LLM API 価格

Qwen モデル

TokenHot で利用可能な 15 個の Qwen モデルを表示しています。

1つのAPIキーで15個のQwenモデルを利用できます。料金はリアルタイムで、プロバイダーごとのアカウントは不要です。

コンテキスト1M
入力$0.1492/100万
出力$0.4476/100万
An efficient multimodal workhorse in the Qwen family, with a native million-token context window and strong support for coding assistance, agentic workflows, and visual understanding such as charts. It fits long documents, codebases, high-concurrency applications, and tool-assisted tasks.
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
モデルを試す
秒ごと$0.0675/s
The speed-first variant of Alibaba's Wan 3.0 line (early access). It shortens generation latency while retaining high-quality video output, making it a good fit for low-latency previews, batch drafts and fast-iterating short-video workflows. If ultimate visual control matters more than speed, the standard version is a better choice.
入力タイプ:
出力タイプ:
Multimodal Output
モデルを試す
Qwen
秒ごと$0.0450/s
The standard video generation model of Alibaba's Wan 3.0 line (early access). It supports text-to-video, image-to-video and multi-modal reference inputs, generating clips up to roughly 30 seconds with synchronized audio in a single pass — suitable for ads, short drama and narrative content. Official release date not yet announced.
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
コンテキスト1M
入力
$1.4200/100万$1.7750/100万
出力
$4.2640/100万$5.3300/100万
Qwen's most capable flagship multimodal reasoning model, built for long-horizon coding, professional work, and multi-stage agent tasks. It accepts text, images, and video, offers a 1M-token context window with up to 128K output, and supports function calling, structured outputs, web search, and a code interpreter. It is best used as a high-quality workhorse for complex tasks rather than for the lowest-cost, lowest-latency batch workloads.
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
モデルを試す
秒ごと
$0.0024/s$0.0030/s
A quality-first image generation and editing model: it supports text-to-image and edits from 1–3 reference images, accepts prompts of up to about 4.5K tokens, and is aimed at dense layouts, fine 10px-scale text, 12 languages, and multiple fonts. Choose it over the standard variant when complex composition and detail fidelity matter more than speed or cost; use the standard variant for a faster, more economical workflow.
入力タイプ:
出力タイプ:
秒ごと
$0.0024/s$0.0030/s
The balanced standard tier of Qwen-Image 3.0 for image generation and editing: it supports text-to-image and edits from 1–3 reference images, prompts of up to about 4.5K tokens, up to six outputs, and 2K-class resolution. It fits everyday posters, interfaces, infographics, and repeated creative work; move to Pro when complex layouts and fine text fidelity matter more, and use this tier when speed and cost efficiency are the priority.
入力タイプ:
出力タイプ:
秒ごと$0.1350/s
HappyHorse 1.0 よりも強力なオーディオネイティブのワークフローを備え、テキストからビデオへの変換用に構築されており、ダイアログ、サウンドエフェクト、バックグラウンドミュージックを 1 つのパスで生成します。 Veo、Kling、Runway と比較した場合、その差別化要因は、同期されたオーディオビジュアル生成と、短編ドラマ、広告、電子商取引、およびブランド マーケティング向けのキャラクター一貫性のあるワークフローです。
入力タイプ:
出力タイプ:
Multimodal Output
秒ごと$0.1350/s
HappyHorse 1.0 よりも強力なオーディオ ネイティブ ワークフローを備えたビデオ参照用に構築されており、ダイアログ、サウンド効果、バックグラウンド ミュージックを 1 つのパスで生成します。 Veo、Kling、Runway と比較した場合、その差別化要因は、同期されたオーディオビジュアル生成と、短編ドラマ、広告、電子商取引、およびブランド マーケティング向けのキャラクター一貫性のあるワークフローです。
入力タイプ:
出力タイプ:
Multimodal Output
秒ごと$0.1350/s
HappyHorse 1.0 よりも強力なオーディオネイティブのワークフローを備え、ダイアログ、サウンドエフェクト、BGM を 1 つのパスで生成する、画像からビデオへの変換用に構築されています。 Veo、Kling、Runway と比較した場合、その差別化要因は、同期されたオーディオビジュアル生成と、短編ドラマ、広告、電子商取引、およびブランド マーケティング向けのキャラクター一貫性のあるワークフローです。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
コンテキスト1M
入力
$0.2400/100万$0.3000/100万
出力
$0.9600/100万$1.2000/100万
エージェントのワークロード、コーディング、オフィスの生産性向けに設計されており、Qwen3.6 よりも長いコンテキスト、ツールの呼び出し、構造化された出力に重点を置いています。 GPT、Claude、Gemini などの主力モデルと比較して、中国語、コーディング、およびマルチステップ ツールのワークフローにとってコスト効率の高い代替品です。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
モデルを試す
Qwen
コンテキスト1M
入力
$1.4400/100万$1.8000/100万
出力
$4.3200/100万$5.4000/100万
エージェントのワークロード、コーディング、オフィスの生産性向けに設計されており、Qwen3.6 よりも長いコンテキスト、ツールの呼び出し、構造化された出力に重点を置いています。 GPT、Claude、Gemini などの主力モデルと比較して、中国語、コーディング、およびマルチステップ ツールのワークフローにとってコスト効率の高い代替品です。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
モデルを試す
コンテキスト8K
秒ごと$0.1350/s
ビデオ編集に関しては、初期のビデオ モデルと比較して、ネイティブ オーディオの生成により重点が置かれており、ダイアログ、効果音、BGM を一度に完成させることができます。 Veo、Kling、Runway と比較して、その差別化された利点は、サウンドとビジュアルの同期とキャラクターの一貫性のあるワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
コンテキスト8K
秒ごと$0.1350/s
Wensheng ビデオでは、初期のビデオ モデルと比較して、ネイティブ オーディオの生成に重点が置かれており、ダイアログ、効果音、BGM が一度に完成します。 Veo、Kling、Runway と比較して、その差別化された利点は、サウンドとビジュアルの同期とキャラクターの一貫性のあるワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
コンテキスト8K
秒ごと$0.1350/s
初期のビデオ モデルと比較して、参照用ビデオの生成はネイティブ オーディオ生成においてより顕著であり、ダイアログ、サウンド効果、BGM を一度に完成させます。 Veo、Kling、Runway と比較して、その差別化された利点はサウンドとビジュアルの同期とキャラクターの一貫性ワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
コンテキスト8K
秒ごと$0.1350/s
Tusheng ビデオでは、初期のビデオ モデルと比較して、ネイティブ オーディオの生成に重点が置かれており、ダイアログ、効果音、BGM が一度に完成します。 Veo、Kling、Runway と比較して、その差別化された利点は、サウンドとビジュアルの同期とキャラクターの一貫性のあるワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output