Claude
claude-fable-5-1
Anthropic
コンテキスト1M
入力
$3.4000/100万$10.0000/100万
出力
$17.0000/100万$50.0000/100万
A flagship long-context model for demanding, long-running knowledge work, coding projects, and multi-step agentic tasks. It accepts text, images, and PDFs, making it a strong fit for cross-file analysis, complex research, and sustained tool use. It prioritizes reasoning depth and long-horizon execution over speed and cost, so it is best reserved for tasks where reliable completion matters.
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingLong Context
Qwen
qwen3.8-flash
Qwen
コンテキスト1M
入力$0.1492/100万
出力$0.4476/100万
An efficient multimodal workhorse in the Qwen family, with a native million-token context window and strong support for coding assistance, agentic workflows, and visual understanding such as charts. It fits long documents, codebases, high-concurrency applications, and tool-assisted tasks.
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Zhipu
glm-5.3-flash
Zhipu AI
コンテキスト1.31M
入力$0.1190/100万
出力$0.4200/100万
The efficiency-focused Flash model in the GLM-5 family, with native multimodal support and a design aimed at long-context, high-frequency execution. It is a strong fit for coding, agentic workflows, and complex tasks that require image-and-text understanding, especially when production cost matters.
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
DeepSeek
deepseek-v4-flash-vision-exp
DeepSeek
コンテキスト1.05M
入力$0.1490/100万
出力$0.2980/100万
An experimental multimodal variant of DeepSeek-V4-Flash for image understanding alongside text. A practical choice for visual Q&A, screenshot and document analysis, while its experimental status makes it better suited to evaluation and flexible workflows than strict production-critical paths.
入力タイプ:
出力タイプ:
ReasoningTool UseLong Context
Z.ai
coding-glm-5.3-free
Free channel
コンテキスト1M
入力$0.0000/100万
出力$0.0000/100万
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Zhipu
glm-5.3
Zhipu AI
コンテキスト1M
入力
$1.0115/100万$1.1900/100万
出力
$3.5275/100万$4.1500/100万
A flagship workhorse for complex software engineering and long-horizon agent tasks. It uses the same base model as GLM-5.2, with improvements driven by post-training, delivering a 50% gain on Z.ai Code Bench alongside stronger terminal-operation and vulnerability-discovery capabilities. It offers a 1M-token context window, up to 128K output, always-on reasoning with low, high, and max effort levels, function calling, and structured outputs. It currently accepts text only and is not intended for tasks requiring visual understanding.
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Gemini
gemini-3.7-flash-free
Free channel
コンテキスト1.05M
入力$0.0000/100万
出力$0.0000/100万
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
入力タイプ:
出力タイプ:
ReasoningTool UseLong Context
Gemini
gemini-3.7-flash
Google
コンテキスト1.05M
入力$0.7500/100万
出力$3.7500/100万
A new fast multimodal model in the Gemini 3 family, released in August 2026, built for low-cost, high-throughput multimodal understanding. It accepts text, image, video, audio and document inputs with a 1M-token context, suited to large-scale processing, batch workloads and real-time interactions.
入力タイプ:
出力タイプ:
ReasoningTool UseLong Context
Grok
grok-4.6
xAI
コンテキスト500K
入力$0.8400/100万
出力$2.5200/100万
A flagship general-purpose model for high-quality coding, agentic execution, and knowledge work. It fits tasks that need sustained planning, tool collaboration, and complex problem solving; for simple calls where latency or cost matters most, choose a lighter model.
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Qwen
qwen3.8-max
Qwen
コンテキスト1M
入力
$1.4200/100万$1.7750/100万
出力
$4.2640/100万$5.3300/100万
Qwen's most capable flagship multimodal reasoning model, built for long-horizon coding, professional work, and multi-stage agent tasks. It accepts text, images, and video, offers a 1M-token context window with up to 128K output, and supports function calling, structured outputs, web search, and a code interpreter. It is best used as a high-quality workhorse for complex tasks rather than for the lowest-cost, lowest-latency batch workloads.
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Minimax
MiniMax-H3
MiniMax
秒ごと$0.0800/s
An omni-modal video generation model that can combine text with image, video, and audio references to create video assets with synchronized audio. It is suited to short-form video, advertising, and reference-driven character or camera work; it is not a general chat model, so prioritize visual consistency, motion, and audio-visual quality.
入力タイプ:
出力タイプ:
Multimodal Output
Claude
claude-opus-5
Anthropic
コンテキスト1M
入力
$1.7000/100万$5.0000/100万
出力
$8.5000/100万$25.0000/100万
A premium model for complex agentic coding and enterprise knowledge work, emphasizing long-horizon planning, sustained execution, self-checking, and reliable delivery. With a million-token context window, it is suited to large codebases, complex documents, and multi-step projects that need limited human oversight.
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Gemini
gemini-3.6-flash
Google
コンテキスト1.05M
入力$1.5000/100万
出力$7.5000/100万
The workhorse of the Gemini Flash line, balancing efficiency with coding, knowledge work, and multimodal understanding. Compared with 3.5 Flash, it is designed to complete multi-step tasks with fewer output tokens, reasoning steps, and tool calls, making it a good fit for production agents, document and chart analysis, and frequent complex workflows.
入力タイプ:
出力タイプ:
ReasoningTool UseLong Context
Gemini
gemini-3.5-flash-lite
Google
コンテキスト1.05M
入力$0.3000/100万
出力$2.5000/100万
The fastest and most cost-effective model in the Gemini 3.5 line, built for high-throughput, low-latency work such as translation, classification, document processing, and lightweight agentic workflows. It supports multimodal inputs and up to a 1M-token context, while demanding reasoning, long-horizon planning, or high-quality code changes are better handled by a higher-tier Flash model.
入力タイプ:
出力タイプ:
Long Context
MoonshotAI
kimi-k3
Moonshot
コンテキスト1.05M
入力
$2.7000/100万$3.0000/100万
出力
$13.5000/100万$15.0000/100万
ネイティブ ビジョンと最大 1M トークンのコンテキスト ウィンドウを備えた Moonshot フラッグシップ モデル。現在、最大の推論労力で実行され、長期的なコーディング、知識作業、ターミナル ツールを調整する複雑なタスクに適しています。安定して使用するには、完全な推論履歴を保存し、互換性のあるエージェント ハーネスで新しいセッションを開始します。
入力タイプ:
出力タイプ:
ReasoningTool UseLong Context
MoonshotAI
coding-kimi-k3-free
Free channel
コンテキスト1.05M
入力$0.0000/100万
出力$0.0000/100万
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
入力タイプ:
出力タイプ:
ReasoningTool UseLong Context
OpenAI
gpt-5.6-terra
OpenAI
コンテキスト1.05M
入力
$0.4000/100万$2.0000/100万
出力
$2.4000/100万$12.0000/100万
The balanced workhorse tier of GPT-5.6, combining capability and cost. Well suited to production work involving image understanding, tool use, and large source material; choose Sol for the hardest quality-first work or Luna for cost-controlled high volume.
入力タイプ:
出力タイプ:
ReasoningFunction CallingStructured OutputLong Context
OpenAI
gpt-5.6-sol
OpenAI
コンテキスト1.05M
入力
$1.0000/100万$5.0000/100万
出力
$6.0000/100万$30.0000/100万
The flagship GPT-5.6 tier for complex professional work and quality-first delivery. It supports image input, function calling, structured outputs, and an exceptionally large context window; consider Terra or Luna when latency or unit cost is the primary constraint.
入力タイプ:
出力タイプ:
ReasoningFunction CallingStructured OutputLong Context
Doubao
doubao-seedream-5-0-pro
Doubao
秒ごと$0.0450/s
プロフェッショナルなビジュアル制作のためのマルチモーダル画像生成モデル。テキストから画像を生成し、参照画像や空間注釈を使用して正確な編集を実行でき、高密度のインフォグラフィック、タイポグラフィ、リアルな照明、多言語テキストのレンダリングに強みを持ちます。
入力タイプ:
出力タイプ:
ReasoningMultimodal Output
Grok
grok-4.5
xAI
コンテキスト500K
入力$2.1000/100万
出力$6.3000/100万
コーディング、ナレッジワーク、STEM タスクのためのフロンティア推論モデル。画像入力、関数呼び出し、構造化出力をサポートします。その 500K トークンのコンテキストは、ファイル間の分析や長いソース素材に適しています。最先端の推論が必要ない場合は、低コストのモデルを選択します。
入力タイプ:
出力タイプ:
ReasoningFunction CallingStructured OutputLong Context
Claude
claude-fable-5
Anthropic
コンテキスト1M
入力
$3.4000/100万$10.0000/100万
出力
$17.0000/100万$50.0000/100万
Anthropic の最も有能で広くリリースされているモデルで、最も要求の厳しい推論と長期的なエージェント作業向けに構築されています。最大 128,000 の出力と常時オンの適応的思考を備えた 1M トークンのコンテキスト ウィンドウを提供し、人間の入力が必要になる前に、指定されていないタスクを自律的に探索、計画し、長時間実行されるコーディングとマルチエージェント オーケストレーションを実行することに優れています。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Gemini
gemini-3.1-flash-lite-image
Google
回ごと
$0.0120/call$0.0200/call
A high-efficiency image generation and editing model for frequent visual drafts, social assets, posters, and iterative revisions. It supports multimodal inputs and a million-token context, favoring speed and scaled usage; choose Pro Image when complex creative control is the priority.
入力タイプ:
出力タイプ:
Multimodal Output
Claude
claude-sonnet-5
Anthropic
コンテキスト1M
入力
$0.6800/100万$2.0000/100万
出力
$3.4000/100万$10.0000/100万
デイリードライバーのソネット層モデルは、より低い運用コストでニアフロンティアのエージェント、コーディング、ナレッジワーク機能を実現することを目的としています。デフォルトの 1M トークンのコンテキストと適応的思考は、長いドキュメント、コードベース、および複数ステップのツールのワークフローに適合します。最も深い推論や制限された高リスクのセキュリティ作業については、より上位のモデルまたは特殊なモデルを評価します。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Doubao
doubao-seedance-2.5-260628
Doubao
入力
$3.6594/100万$3.8520/100万
出力
$3.6594/100万$3.8520/100万
A next-generation Seedance model for long-form storytelling and complex creative workflows, with more flexible multimodal references, video extension, and precise editing. It is well suited to advertising, short films, and production work that must preserve character, scene, and audio-visual continuity.
入力タイプ:
出力タイプ:
Multimodal Output
Doubao
doubao-seed-2-1-pro
Doubao
コンテキスト256K
入力$0.9700/100万
出力$4.8800/100万
コーディング、エージェント、複雑な生産性タスク向けに構築されており、長いコンテキスト、長い出力、ツール指向のワークフローにおいて Doubao Seed 2.0 を超えています。 Qwen、GLM、DeepSeek の同等の製品と比較して、中国の事務作業、コーディング、および複数ステップの自動化にとってコスト効率の高いオプションです。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Qwen
happyhorse-1.1-t2v
阿里巴巴
秒ごと$0.1350/s
HappyHorse 1.0 よりも強力なオーディオネイティブのワークフローを備え、テキストからビデオへの変換用に構築されており、ダイアログ、サウンドエフェクト、バックグラウンドミュージックを 1 つのパスで生成します。 Veo、Kling、Runway と比較した場合、その差別化要因は、同期されたオーディオビジュアル生成と、短編ドラマ、広告、電子商取引、およびブランド マーケティング向けのキャラクター一貫性のあるワークフローです。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
happyhorse-1.1-r2v
阿里巴巴
秒ごと$0.1350/s
HappyHorse 1.0 よりも強力なオーディオ ネイティブ ワークフローを備えたビデオ参照用に構築されており、ダイアログ、サウンド効果、バックグラウンド ミュージックを 1 つのパスで生成します。 Veo、Kling、Runway と比較した場合、その差別化要因は、同期されたオーディオビジュアル生成と、短編ドラマ、広告、電子商取引、およびブランド マーケティング向けのキャラクター一貫性のあるワークフローです。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
happyhorse-1.1-i2v
阿里巴巴
秒ごと$0.1350/s
HappyHorse 1.0 よりも強力なオーディオネイティブのワークフローを備え、ダイアログ、サウンドエフェクト、BGM を 1 つのパスで生成する、画像からビデオへの変換用に構築されています。 Veo、Kling、Runway と比較した場合、その差別化要因は、同期されたオーディオビジュアル生成と、短編ドラマ、広告、電子商取引、およびブランド マーケティング向けのキャラクター一貫性のあるワークフローです。
入力タイプ:
出力タイプ:
Multimodal Output
Doubao
dreamina-seedance-2-0-mini-filter-off
Doubao
入力$1.5750/100万
出力$1.5750/100万
ビデオ生成とマルチアセットビデオ作成用に構築されており、より完全なモーション安定性、オーディオビジュアル生成、リファレンス入力により Seedance 1.x が改良されています。 Veo、Kling、Runway と比較して、テキスト、画像、オーディオ、ビデオの混合入力によって駆動される中国のクリエイティブ ワークフローに適しています。このフィルターオフのバリアントは、標準の Seedance 2.0 と同じ層の生成機能を維持しながら、より緩やかなコンテンツ フィルターを適用します。短編ドラマ、広告、電子商取引、ソーシャル アセットに役立ちます。
入力タイプ:
出力タイプ:
Multimodal Output
Zhipu
glm-5.2
Zhipu AI
コンテキスト1M
入力
$0.9690/100万$1.1400/100万
出力
$3.4000/100万$4.0000/100万
推論、コーディング、および長いコンテキストのエージェント タスク向けに構築されており、コンテキストの長さ、ツールの使用、および複雑な複数ステップのワークフローにおいて GLM-5.1 よりも改善されています。 Qwen、DeepSeek、Kimi などの中国の主力製品と比較すると、エンタープライズ自動化、プロジェクト レベルのコード分析、中国語でのナレッジ ワークに適しています。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Doubao
doubao-seedance-2.0-mini-260615
Doubao
入力
$0.3150/100万$0.6300/100万
出力
$0.3150/100万$0.6300/100万
A lightweight Seedance 2.0 variant that retains reference-driven video generation with text, images, and video. It is better suited to rapid iteration, batch drafts, and cost-sensitive workflows; choose the standard model when complex motion control, camera work, or final quality matters more.
入力タイプ:
出力タイプ:
Multimodal Output
MoonshotAI
kimi-k2.7-code-highspeed
Moonshot
コンテキスト256K
入力$1.9000/100万
出力$7.9999/100万
長いコンテキスト、複数ステップのツールの使用、およびコード/ドキュメントのワークフロー向けに構築されており、Kimi K2.6 よりも強力なエンジニアリング タスクの安定性とコンテキスト容量が備わっています。 Qwen、GLM、DeepSeek と比較すると、中国語の長い文書の分析、プロジェクトの Q&A、日常的なエージェント プログラミングに適しています。
入力タイプ:
出力タイプ:
Tool UseFunction CallingStructured OutputLong Context
MoonshotAI
kimi-k2.7-code
Moonshot
コンテキスト256K
入力
$0.8775/100万$0.9750/100万
出力
$3.6450/100万$4.0500/100万
長いコンテキスト、複数ステップのツールの使用、およびコード/ドキュメントのワークフロー向けに構築されており、Kimi K2.6 よりも強力なエンジニアリング タスクの安定性とコンテキスト容量が備わっています。 Qwen、GLM、DeepSeek と比較すると、中国語の長い文書の分析、プロジェクトの Q&A、日常的なエージェント プログラミングに適しています。
入力タイプ:
出力タイプ:
Tool UseFunction CallingStructured OutputLong Context
Grok
grok-imagine-video-1-5-preview
xAI
秒ごと$0.1400/s
Grok Imagine 1.0 よりも強力なモーションの一貫性と高速なソーシャル コンテンツ制作を備えた、画像/ビデオ主導の生成、編集、拡張用に構築されています。 Veo、Kling、Runway と比較すると、xAI エコシステムの短編ビデオ、広告コンセプト、トレンド主導のアセットに適合します。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
qwen3.7-plus
Qwen
コンテキスト1M
入力
$0.2400/100万$0.3000/100万
出力
$0.9600/100万$1.2000/100万
エージェントのワークロード、コーディング、オフィスの生産性向けに設計されており、Qwen3.6 よりも長いコンテキスト、ツールの呼び出し、構造化された出力に重点を置いています。 GPT、Claude、Gemini などの主力モデルと比較して、中国語、コーディング、およびマルチステップ ツールのワークフローにとってコスト効率の高い代替品です。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Minimax
minimax-m3-free
Free channel
コンテキスト1M
入力$0.0000/100万
出力$0.0000/100万
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Minimax
MiniMax-M3
MiniMax
コンテキスト1M
入力$0.3150/100万
出力$1.2600/100万
長期にわたるコーディング、ツールの使用、マルチターンの制作コラボレーション向けに構築されており、1M トークンのコンテキスト ウィンドウを維持しながら、マルチモーダル入力で MiniMax M2.7 を超えて拡張されています。 Claude、Gemini、GLM-5.2 のロングコンテキスト エージェント モデルと比較して、テキスト、画像、ビデオ素材を 1 つのワークフローに入れるのに適しています。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Claude
claude-opus-4-8
Anthropic
コンテキスト1M
入力
$1.7000/100万$5.0000/100万
出力
$8.5000/100万$25.0000/100万
主要な推論と複雑なエージェント タスクに関しては、長いコンテキスト、コード、ツールの使用法、および複雑な計画の点で、Claude Opus 4.7 よりもさらに進んでいます。 GPT、Gemini、Qwen などの主力モデルと比較して、信頼性の高いエージェント、コード ベースのメンテナンス、詳細な調査ワークフローに適しています。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Gemini
gemini-omni-video
Google
秒ごと$0.2450/s
あらゆる入力からビデオへのマルチモーダル作成用に構築されており、以前の Gemini ビデオ ワークフローよりもテキスト、画像、オーディオ、ビデオ素材の統合されたオーケストレーションに重点が置かれています。 Veo、Kling、Runway と比較して、複雑なマルチモーダル アセットを 1 つのクリエイティブ ワークフローに統合するのに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Gemini
gemini-3.5-flash
Google
コンテキスト1M
入力$1.5000/100万
出力$9.0000/100万
マルチモーダルな理解、コーディング、および長いコンテキストのタスク向けに構築されており、スループット、コンテキスト、およびツールの機能が Gemini 3.1 よりも向上しています。 GPT、Claude、Qwen モデルと比較して、テキスト、画像、オーディオ、ビデオ、ドキュメントにわたる統合分析ワークフローに優れています。
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Qwen
qwen3.7-max
Qwen
コンテキスト1M
入力
$1.4400/100万$1.8000/100万
出力
$4.3200/100万$5.4000/100万
エージェントのワークロード、コーディング、オフィスの生産性向けに設計されており、Qwen3.6 よりも長いコンテキスト、ツールの呼び出し、構造化された出力に重点を置いています。 GPT、Claude、Gemini などの主力モデルと比較して、中国語、コーディング、およびマルチステップ ツールのワークフローにとってコスト効率の高い代替品です。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Qwen
happyhorse-1.0-video-edit
Qwen
コンテキスト8K
秒ごと$0.1350/s
ビデオ編集に関しては、初期のビデオ モデルと比較して、ネイティブ オーディオの生成により重点が置かれており、ダイアログ、効果音、BGM を一度に完成させることができます。 Veo、Kling、Runway と比較して、その差別化された利点は、サウンドとビジュアルの同期とキャラクターの一貫性のあるワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
happyhorse-1.0-t2v
Qwen
コンテキスト8K
秒ごと$0.1350/s
Wensheng ビデオでは、初期のビデオ モデルと比較して、ネイティブ オーディオの生成に重点が置かれており、ダイアログ、効果音、BGM が一度に完成します。 Veo、Kling、Runway と比較して、その差別化された利点は、サウンドとビジュアルの同期とキャラクターの一貫性のあるワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
happyhorse-1.0-r2v
Qwen
コンテキスト8K
秒ごと$0.1350/s
初期のビデオ モデルと比較して、参照用ビデオの生成はネイティブ オーディオ生成においてより顕著であり、ダイアログ、サウンド効果、BGM を一度に完成させます。 Veo、Kling、Runway と比較して、その差別化された利点はサウンドとビジュアルの同期とキャラクターの一貫性ワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
happyhorse-1.0-i2v
Qwen
コンテキスト8K
秒ごと$0.1350/s
Tusheng ビデオでは、初期のビデオ モデルと比較して、ネイティブ オーディオの生成に重点が置かれており、ダイアログ、効果音、BGM が一度に完成します。 Veo、Kling、Runway と比較して、その差別化された利点は、サウンドとビジュアルの同期とキャラクターの一貫性のあるワークフローにあります。短編ドラマ、広告、電子商取引、ブランドマーケティングビデオなどに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
DeepSeek
deepseek-v4-pro
DeepSeek
コンテキスト1M
入力
$1.5390/100万$1.7100/100万
出力
$3.0869/100万$3.4299/100万
DeepSeek V3.2 よりも強力なロングコンテキスト、思考モード、および複雑なタスクの実行を備えた、推論、コーディング、およびエージェント ツールの使用のために構築されています。 Qwen、GLM、Kimi などの中国の同業他社と比較して、特に数学、プログラミング、論理的推論のタスクで競争力があります。
入力タイプ:
出力タイプ:
Long ContextTool UseReasoningStructured Output
DeepSeek
deepseek-v4-flash
DeepSeek
コンテキスト1M
入力
$0.1278/100万$0.1420/100万
出力
$0.2574/100万$0.2860/100万
DeepSeek V3.2 よりも強力なロングコンテキスト、思考モード、および複雑なタスクの実行を備えた、推論、コーディング、およびエージェント ツールの使用のために構築されています。 Qwen、GLM、Kimi などの中国の同業他社と比較して、特に数学、プログラミング、論理的推論のタスクで競争力があります。
入力タイプ:
出力タイプ:
Long ContextTool UseReasoningStructured Output
OpenAI
gpt-5.5
OpenAI
コンテキスト1M
入力
$1.0000/100万$5.0000/100万
出力
$6.0000/100万$30.0000/100万
GPT-5.4 よりも強力な長いコンテキスト、ツールの使用、およびスケーラブルで信頼性の高い出力を備えた、推論、コーディング、視覚的理解、エージェント タスク向けに構築されています。 Claude、Gemini、Qwen のフラッグシップ モデルと比較して、信頼性の高い自動化と複雑なナレッジ ワークフローのための強力な汎用ベースです。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
XiaomiMiMo
mimo-v2.5-tts-voicedesign
Xiaomi MiMo
入力$0.0000/100万
出力$0.0000/100万
MiMo V2 TTS よりも強力な V2.5 シリーズの自然言語音声作成、きめ細かいペース/感情/トーン コントロール、低摩擦の音声生成を備えた一文音声デザインとテキスト読み上げ用に構築されています。イレブンラボ、OpenAI TTS、Azure Speech と比較すると、リファレンスフリーのキャラクター音声デザイン、広告ナレーション、短いビデオのナレーション、マルチスタイルの音声探索に役立ちます。
入力タイプ:
出力タイプ:
Multimodal Output
XiaomiMiMo
mimo-v2.5-tts-voiceclone
Xiaomi MiMo
入力$0.0000/100万
出力$0.0000/100万
MiMo V2 TTS よりも強力な V2.5 シリーズの高忠実度の音色再現、音声キャラクターの一貫性、クロステキスト汎用化を備えた、少数サンプルの音声クローン作成用に構築されています。イレブンラボ、OpenAI TTS、Azure Speech と比較すると、中国語の吹き替え、ショートビデオのナレーション、ポッドキャスト音声の複製、ターゲットの音声を保存する必要がある多言語音声制作に適しています。
入力タイプ:
出力タイプ:
Multimodal Output
XiaomiMiMo
mimo-v2.5-tts
Xiaomi MiMo
入力$0.0000/100万
出力$0.0000/100万
MiMo V2 TTS よりも強力な V2.5 シリーズの自然な音声合成、きめ細かいペース/感情/トーン コントロール、低摩擦の音声作成を備えたテキスト読み上げ、音声デザイン、音声クローン作成用に構築されています。イレブンラボ、OpenAI TTS、Azure Speech と比較すると、中国語のナレーション、キャラクターの吹き替え、ショートビデオの音声、多言語の音声制作に適しています。
入力タイプ:
出力タイプ:
Multimodal Output
XiaomiMiMo
mimo-v2.5-pro
Xiaomi MiMo
コンテキスト1M
入力$1.1700/100万
出力$3.5000/100万
エージェント、コーディング、および長いコンテキストのタスク向けに構築されており、MiMo V2 よりも完全なマルチモーダルの理解、100 万トークンのコンテキスト ウィンドウ、および長期的な推論を提供します。 Qwen、GLM、DeepSeek などの中国の同業他社と比較すると、Xiaomi のエコシステムへの適合性とエンジニアリング自動化シナリオが強みです。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
XiaomiMiMo
mimo-v2.5
Xiaomi MiMo
コンテキスト1M
入力$0.4700/100万
出力$2.3600/100万
エージェント、コーディング、および長いコンテキストのタスク向けに構築されており、MiMo V2 よりも完全なマルチモーダルの理解、100 万トークンのコンテキスト ウィンドウ、および長期的な推論を提供します。 Qwen、GLM、DeepSeek などの中国の同業他社と比較すると、Xiaomi のエコシステムへの適合性とエンジニアリング自動化シナリオが強みです。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
OpenAI
gpt-image-2-text-to-image
OpenAI
コンテキスト128K
枚ごと$0.0300/image
割引価格のお試しルートです。生成にはさらに時間がかかる場合があり、失敗する可能性が約 10% あります。これは、キューイングや偶発的な失敗が許容される、コスト重視のイメージ生成トライアルに最適です。より安定して本番環境で使用するには、通常のルートを選択してください。
入力タイプ:
出力タイプ:
Multimodal Output
OpenAI
gpt-image-2-image-to-image
OpenAI
コンテキスト128K
枚ごと$0.0200/image
割引価格のお試しルートです。生成にはさらに時間がかかる場合があり、失敗する可能性が約 10% あります。これは、キューイングや偶発的な失敗が許容される、コスト重視のイメージ生成トライアルに最適です。より安定して本番環境で使用するには、通常のルートを選択してください。
入力タイプ:
出力タイプ:
Multimodal Output
OpenAI
gpt-image-2
OpenAI
コンテキスト128K
枚ごと$0.0670/image
高品質のイメージ生成と編集用に構築されており、テキスト レンダリング、ローカル編集、マルチターンの一貫性において GPT Image 1.x よりも改善されています。 Gemini イメージ モデル、Qwen Image、Seedream と比較すると、ブランド ビジュアル、広告デザイン、製品イメージ、厳密に制御されたクリエイティブ ワークフローに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Claude
claude-opus-4-7
Anthropic
コンテキスト1M
入力
$1.7000/100万$5.0000/100万
出力
$8.5000/100万$25.0000/100万
主要な推論と複雑なエージェント タスクに関しては、長いコンテキスト、コード、ツールの使用法、および複雑な計画の点で、Claude Opus 4.6 よりもさらに進んでいます。 GPT、Gemini、Qwen などの主力モデルと比較して、信頼性の高いエージェント、コード ベースのメンテナンス、詳細な調査ワークフローに適しています。
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Minimax
MiniMax-M2.7-highspeed
MiniMax
コンテキスト200K
入力$0.6300/100万
出力$2.5200/100万
テキストの生成、推論、ツールの使用のために構築されており、以前の MiniMax M2.x バージョンよりも強力な 200K ロング コンテキストとエージェント ワークフローの安定性を備えています。 Kimi、GLM、Qwen などの中国の同業他社と比較して、生産性、コーディング支援、長い文書の処理に適しています。
入力タイプ:
出力タイプ:
Reasoning
Minimax
MiniMax-M2.7
MiniMax
コンテキスト200K
入力$0.3150/100万
出力$1.2600/100万
テキストの生成、推論、ツールの使用のために構築されており、以前の MiniMax M2.x バージョンよりも強力な 200K ロング コンテキストとエージェント ワークフローの安定性を備えています。 Kimi、GLM、Qwen などの中国の同業他社と比較して、生産性、コーディング支援、長い文書の処理に適しています。
入力タイプ:
出力タイプ:
Reasoning
OpenAI
gpt-5.4-mini
OpenAI
コンテキスト400K
入力$0.7500/100万
出力$4.5000/100万
GPT-5.3 よりも強力な長いコンテキスト、ツールの使用、およびスケーラブルで信頼性の高い出力を備えた、推論、コーディング、視覚的理解、エージェント タスク向けに構築されています。 Claude、Gemini、Qwen のフラッグシップ モデルと比較して、信頼性の高い自動化と複雑なナレッジ ワークフローのための強力な汎用ベースです。
入力タイプ:
出力タイプ:
Long ContextTool UseReasoningStructured Output
OpenAI
gpt-5.4
OpenAI
コンテキスト1.05M
入力
$0.5000/100万$2.5000/100万
出力
$3.0000/100万$15.0000/100万
GPT-5.3 よりも強力な長いコンテキスト、ツールの使用、およびスケーラブルで信頼性の高い出力を備えた、推論、コーディング、視覚的理解、エージェント タスク向けに構築されています。 Claude、Gemini、Qwen のフラッグシップ モデルと比較して、信頼性の高い自動化と複雑なナレッジ ワークフローのための強力な汎用ベースです。
入力タイプ:
出力タイプ:
Long ContextTool UseReasoningStructured Output
Qwen
qwen-image-2.0-pro
Qwen
入力$2.2000/100万
出力$2.2000/100万
画像の生成と編集用に構築されており、中国語テキストのレンダリング、指示に従って、複雑なレイアウトにおいて Qwen Image 1.x よりも改善されています。 GPT Image、Gemini 画像モデル、Seedream と比較すると、中国のポスター、電子商取引画像、インフォグラフィック、反復的なビジュアル作成に適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
qwen-image-2.0
Qwen
入力$2.2000/100万
出力$2.2000/100万
画像の生成と編集用に構築されており、中国語テキストのレンダリング、指示に従って、複雑なレイアウトにおいて Qwen Image 1.x よりも改善されています。 GPT Image、Gemini 画像モデル、Seedream と比較すると、中国のポスター、電子商取引画像、インフォグラフィック、反復的なビジュアル作成に適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Gemini
gemini-3.1-flash-image-preview
Google
コンテキスト131K
枚ごと$0.0672/image
画像の生成と編集のために構築されており、以前の Gemini 画像機能よりも強力なマルチモーダルな理解、会話型編集、および迅速なビジュアル プロトタイピングが可能です。 GPT Image、Qwen Image、および Seedream と比較すると、ドキュメント、画像、およびプロンプトを 1 つの視覚的なワークフローに組み合わせるのに役立ちます。
入力タイプ:
出力タイプ:
ReasoningMultimodal Output
Gemini
gemini-3.1-flash-image
Google
枚ごと
$0.0198/image$0.0330/image
An efficient Gemini image tier that balances quality and speed for text-to-image generation, image editing, multi-image composition, and workflows that need fast, scalable output. It offers flexible composition and text rendering, while complex professional design may be better served by Gemini 3 Pro Image.
入力タイプ:
出力タイプ:
Multimodal Output
Gemini
gemini-3.1-pro-preview
Google
コンテキスト1.05M
入力$2.0000/100万
出力$12.0000/100万
マルチモーダルな理解、コーディング、および長いコンテキストのタスク向けに構築されており、スループット、コンテキスト、およびツールの機能が Gemini 3.0 / 2.5 よりも向上しています。 GPT、Claude、Qwen モデルと比較して、テキスト、画像、オーディオ、ビデオ、ドキュメントにわたる統合分析ワークフローに優れています。
入力タイプ:
出力タイプ:
Long Context
Claude
claude-sonnet-4-6
Anthropic
コンテキスト1M
入力
$1.0200/100万$3.0000/100万
出力
$5.1000/100万$15.0000/100万
コーディングとエージェントのタスクを効率的に行うために、長いコンテキスト、コード、ツールの使用法、および複雑な計画の点で、Claude 4.5/4.6 のプレオーダー バージョンよりもさらに進んでいます。 GPT、Gemini、Qwen などの主力モデルと比較して、信頼性の高いエージェント、コード ベースのメンテナンス、詳細な調査ワークフローに適しています。
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Doubao
doubao-seed-2-0-pro
Doubao
コンテキスト256K
入力$0.5000/100万
出力$2.5300/100万
コーディング、エージェント、複雑な生産性タスク向けに構築されており、長いコンテキスト、長い出力、ツール指向のワークフローにおいて Doubao Seed 2.0 を超えています。 Qwen、GLM、DeepSeek の同等の製品と比較して、中国の事務作業、コーディング、および複数ステップの自動化にとってコスト効率の高いオプションです。
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong Context
Doubao
doubao-seedream-5.0-lite
Doubao
枚ごと$0.0350/image
画像の生成と編集のために構築されており、Seedream 4.x よりも強力な中国語の指示、視覚的な推論、および高頻度のクリエイティブ制作を備えています。 GPT Image、Gemini 画像モデル、Qwen Image と比較すると、ByteDance エコシステムの広告デザイン、電子商取引アセット、およびマルチシナリオのビジュアル制作に適しています。
入力タイプ:
出力タイプ:
Web SearchMultimodal Output
Minimax
MiniMax-M2.5
MiniMax
コンテキスト205K
入力$0.3020/100万
出力$1.2077/100万
カジュアル チャットだけではなく、コーディング、ツールの使用、検索、オフィス スタイルの成果物に最も優れた、実際の生産性エージェント向けのコスト効率の高いモデルです。ファイル間の変更、調査の取得、ドキュメント/スプレッドシート/プレゼンテーションのワークフローを計画および実行するエージェントに適しています。より高価なフラッグシップと比較して、主なトレードオフは、ニアフロンティアに近い実行効率の低コストです。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Doubao
doubao-seedance-2-0-filter-off
Doubao
コンテキスト13K
入力$3.1500/100万
出力$3.1500/100万
ビデオ生成とマルチアセットビデオ作成用に構築されており、より完全なモーション安定性、オーディオビジュアル生成、リファレンス入力により Seedance 1.x が改良されています。 Veo、Kling、Runway と比較して、テキスト、画像、オーディオ、ビデオの混合入力によって駆動される中国のクリエイティブ ワークフローに適しています。このフィルターオフのバリアントは、標準の Seedance 2.0 と同じ層の生成機能を維持しながら、より緩やかなコンテンツ フィルターを適用します。短編ドラマ、広告、電子商取引、ソーシャル アセットに役立ちます。
入力タイプ:
出力タイプ:
Multimodal Output
Doubao
doubao-seedance-2-0-fast-filter-off
Doubao
コンテキスト13K
入力$2.5200/100万
出力$2.5200/100万
ビデオ生成とマルチアセットビデオ作成用に構築されており、より完全なモーション安定性、オーディオビジュアル生成、リファレンス入力により Seedance 1.x が改良されています。 Veo、Kling、Runway と比較して、テキスト、画像、オーディオ、ビデオの混合入力によって駆動される中国のクリエイティブ ワークフローに適しています。このフィルターオフのバリアントは、標準の Seedance 2.0 と同じ層の生成機能を維持しながら、より緩やかなコンテンツ フィルターを適用します。短編ドラマ、広告、電子商取引、ソーシャル アセットに役立ちます。
入力タイプ:
出力タイプ:
Multimodal Output
Claude
claude-opus-4-6
Anthropic
コンテキスト1M
入力
$1.7000/100万$5.0000/100万
出力
$8.5000/100万$25.0000/100万
主力の推論と複雑なエージェント タスクに関しては、長いコンテキスト、コード、ツールの使用法、および複雑な計画の点で、Claude 4.5/4.6 のプレオーダー バージョンよりもさらに進んでいます。 GPT、Gemini、Qwen などの主力モデルと比較して、信頼性の高いエージェント、コード ベースのメンテナンス、詳細な調査ワークフローに適しています。
入力タイプ:
出力タイプ:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Kling
kling-v3-omni
kling
秒ごと$0.0900/s
テキスト/画像からビデオへの変換およびマルチショットの作成用に構築されており、モーションの安定性、カメラ言語、参照の一貫性において Kling 2.x よりも向上しています。 Veo、Runway、Seedance と比較して、中国の短編ドラマ、広告アセット、高品質のソーシャル ビデオ制作に適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Kling
kling-v3
kling
秒ごと$0.0900/s
テキスト/画像からビデオへの変換およびマルチショットの作成用に構築されており、モーションの安定性、カメラ言語、参照の一貫性において Kling 2.x よりも向上しています。 Veo、Runway、Seedance と比較して、中国の短編ドラマ、広告アセット、高品質のソーシャル ビデオ制作に適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Doubao
doubao-seedance-2.0-fast-260128
Doubao
入力
$1.4566/100万$1.7136/100万
出力
$1.4566/100万$1.7136/100万
The speed-prioritized Seedance 2.0 variant for low-latency generation and rapid iteration. It fits batch look development, creative exploration, and workflows that need usable video drafts quickly; choose the standard model when complex reference fusion, motion stability, and final polish are the priority.
入力タイプ:
出力タイプ:
Multimodal Output
Doubao
doubao-seedance-2.0-260128
Doubao
入力
$2.1606/100万$2.2743/100万
出力
$2.1606/100万$2.2743/100万
The standard Seedance 2.0 model for controllable video creation from combined text, image, audio, and video references. It balances complex motion, multi-subject interaction, camera control, and joint audio-video generation, making it a solid default for ads, short films, and high-quality production assets.
入力タイプ:
出力タイプ:
Multimodal Output
Gemini
gemini-3-flash-preview
Google
コンテキスト1M
入力$0.5000/100万
出力$3.0000/100万
エージェントのワークフロー、マルチターン チャット、コーディング支援向けに設計された Google の高速思考モデル。 1M トークンのコンテキスト ウィンドウでテキスト、画像、オーディオ、ビデオ、PDF 入力をサポートします。設定可能な思考レベル、ツールの使用、構造化された出力が特徴です。 Gemini 2.5 フラッシュと比較して、推論、マルチモーダルな理解、信頼性の幅広い品質が向上しています。
入力タイプ:
出力タイプ:
ReasoningTool UseFunction CallingStructured OutputLong Context
Gemini
gemini-3-pro-image-preview
Google
コンテキスト66K
枚ごと$0.1340/image
画像の生成と編集のために構築されており、以前の Gemini 画像機能よりも強力なマルチモーダルな理解、会話型編集、および迅速なビジュアル プロトタイピングが可能です。 GPT Image、Qwen Image、および Seedream と比較すると、ドキュメント、画像、およびプロンプトを 1 つの視覚的なワークフローに組み合わせるのに役立ちます。
入力タイプ:
出力タイプ:
ReasoningMultimodal Output
Gemini
gemini-3-pro-image
Google
枚ごと
$0.0318/image$0.0530/image
A reasoning-driven image generation and editing model for professional creative work, including complex graphic design, high-fidelity product mockups, infographics, and visuals that require accurate text rendering. Compared with the faster Flash Image tier, it prioritizes fine control, grounded content, and polished output.
入力タイプ:
出力タイプ:
Multimodal Output
Claude
claude-opus-4-5
Anthropic
コンテキスト200K
入力$5.0000/100万
出力$25.0000/100万
主力の推論と複雑なエージェント タスクを目的としており、長いコンテキスト、コード、ツールの使用法、および複雑な計画の点で、Claude 4 シリーズのプレオーダー バージョンよりもさらに進んでいます。 GPT、Gemini、Qwen などの主力モデルと比較して、信頼性の高いエージェント、コード ベースのメンテナンス、詳細な調査ワークフローに適しています。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Gemini
veo3.1-lite
Google
回ごと$0.1100/call
Veo 2/3 よりも成熟したネイティブ オーディオ、カメラ制御、リファレンス入力ワークフローを備えた高忠実度ビデオ生成用に構築されています。 Kling、Runway、Seedance と比較して、映画の短編、広告ストーリーボード、安定した物理的な動きが必要なビデオに適しています。 Lite レベルでは、Standard レベルよりも高スループットで低コストのドラフトが優先されます。
入力タイプ:
出力タイプ:
Multimodal Output
Gemini
veo3.1-fast
Google
回ごと$0.2211/call
Veo 2/3 よりも成熟したネイティブ オーディオ、カメラ制御、リファレンス入力ワークフローを備えた高忠実度ビデオ生成用に構築されています。 Kling、Runway、Seedance と比較して、映画の短編、広告ストーリーボード、安定した物理的な動きが必要なビデオに適しています。高速層は、標準層よりも迅速なコンセプトの検証に適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Gemini
veo3.1
Google
回ごと$1.6544/call
Veo 2/3 よりも成熟したネイティブ オーディオ、カメラ制御、リファレンス入力ワークフローを備えた高忠実度ビデオ生成用に構築されています。 Kling、Runway、Seedance と比較して、映画の短編、広告ストーリーボード、安定した物理的な動きが必要なビデオに適しています。
入力タイプ:
出力タイプ:
Multimodal Output
Gemini
gemini-2.5-flash-image
Google
コンテキスト66K
枚ごと$0.0585/image
画像の生成と編集のために構築されており、以前の Gemini 画像機能よりも強力なマルチモーダルな理解、会話型編集、および迅速なビジュアル プロトタイピングが可能です。 GPT Image、Qwen Image、および Seedream と比較すると、ドキュメント、画像、およびプロンプトを 1 つの視覚的なワークフローに組み合わせるのに役立ちます。
入力タイプ:
出力タイプ:
Structured OutputMultimodal Output
Claude
claude-haiku-4-5
Anthropic
コンテキスト200K
入力
$0.3400/100万$1.0000/100万
出力
$1.7000/100万$5.0000/100万
効率的なコーディングとエージェント タスクに関しては、長いコンテキスト、コード、ツールの使用法、および複雑な計画の点で Claude Haiku 3.5 よりも優れています。 GPT、Gemini、Qwen などの主力モデルと比較して、信頼性の高いエージェント、コード ベースのメンテナンス、詳細な調査ワークフローに適しています。
入力タイプ:
出力タイプ:
ReasoningTool UseStructured OutputLong Context
Qwen
wan3.0-video-prime
Qwen
秒ごと$0.4500/s
The speed-first variant of Alibaba's Wan 3.0 line (early access). It shortens generation latency while retaining high-quality video output, making it a good fit for low-latency previews, batch drafts and fast-iterating short-video workflows. If ultimate visual control matters more than speed, the standard version is a better choice.
入力タイプ:
出力タイプ:
Multimodal Output
Qwen
wan3.0-video
Qwen
秒ごと$0.3000/s
The standard video generation model of Alibaba's Wan 3.0 line (early access). It supports text-to-video, image-to-video and multi-modal reference inputs, generating clips up to roughly 30 seconds with synchronized audio in a single pass — suitable for ads, short drama and narrative content. Official release date not yet announced.
入力タイプ:
出力タイプ:
Multimodal Output
Zhipu
glm-image
Zhipu AI
入力$0.8800/100万
出力$3.5500/100万
画像の生成と編集用に構築されており、以前の GLM ビジョン機能よりも中国語のプロンプト、画像テキストの理解、制御可能な生成に重点を置いています。 GPT Image、Gemini 画像モデル、Qwen Image と比較すると、中国のマーケティング アセット、説明用ビジュアル、および反復的なビジュアル編集に適合します。
入力タイプ:
出力タイプ:
Multimodal Output
Doubao
dreamina-seedance-2-5-filter-off
Doubao
入力$75.0000/100万
出力$75.0000/100万
WhatsApp