Qwen
qwen3.8-flash
Qwen
CONTEXT1M
Input$0.1492/M
Output$0.4476/M
An efficient multimodal workhorse in the Qwen family, with a native million-token context window and strong support for coding assistance, agentic workflows, and visual understanding such as charts. It fits long documents, codebases, high-concurrency applications, and tool-assisted tasks.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Zhipu
glm-5.3-flash
Zhipu AI
CONTEXT1.31M
Input$0.1190/M
Output$0.4200/M
The efficiency-focused Flash model in the GLM-5 family, with native multimodal support and a design aimed at long-context, high-frequency execution. It is a strong fit for coding, agentic workflows, and complex tasks that require image-and-text understanding, especially when production cost matters.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
DeepSeek
deepseek-v4-flash-vision-exp
DeepSeek
CONTEXT1.05M
Input$0.1490/M
Output$0.2980/M
An experimental multimodal variant of DeepSeek-V4-Flash for image understanding alongside text. A practical choice for visual Q&A, screenshot and document analysis, while its experimental status makes it better suited to evaluation and flexible workflows than strict production-critical paths.
Input Type:
Output Type:
ReasoningTool UseLong Context
Z.ai
coding-glm-5.3-free
Free channel
CONTEXT1M
Input$0.0000/M
Output$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Zhipu
glm-5.3
Zhipu AI
CONTEXT1M
Input
$1.0115/M$1.1900/M
Output
$3.5275/M$4.1500/M
A flagship workhorse for complex software engineering and long-horizon agent tasks. It uses the same base model as GLM-5.2, with improvements driven by post-training, delivering a 50% gain on Z.ai Code Bench alongside stronger terminal-operation and vulnerability-discovery capabilities. It offers a 1M-token context window, up to 128K output, always-on reasoning with low, high, and max effort levels, function calling, and structured outputs. It currently accepts text only and is not intended for tasks requiring visual understanding.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Gemini
gemini-3.7-flash-free
Free channel
CONTEXT1.05M
Input$0.0000/M
Output$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Input Type:
Output Type:
ReasoningTool UseLong Context
Gemini
gemini-3.7-flash
Google
CONTEXT1.05M
Input$0.7500/M
Output$3.7500/M
A new fast multimodal model in the Gemini 3 family, released in August 2026, built for low-cost, high-throughput multimodal understanding. It accepts text, image, video, audio and document inputs with a 1M-token context, suited to large-scale processing, batch workloads and real-time interactions.
Input Type:
Output Type:
ReasoningTool UseLong Context
Grok
grok-4.6
xAI
CONTEXT500K
Input$0.8400/M
Output$2.5200/M
A flagship general-purpose model for high-quality coding, agentic execution, and knowledge work. It fits tasks that need sustained planning, tool collaboration, and complex problem solving; for simple calls where latency or cost matters most, choose a lighter model.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Qwen
qwen3.8-max
Qwen
CONTEXT1M
Input
$1.4200/M$1.7750/M
Output
$4.2640/M$5.3300/M
Qwen's most capable flagship multimodal reasoning model, built for long-horizon coding, professional work, and multi-stage agent tasks. It accepts text, images, and video, offers a 1M-token context window with up to 128K output, and supports function calling, structured outputs, web search, and a code interpreter. It is best used as a high-quality workhorse for complex tasks rather than for the lowest-cost, lowest-latency batch workloads.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Doubao
doubao-seedance-2.5
Doubao
Input
$3.4200/M$3.6000/M
Output
$3.4200/M$3.6000/M
Built for professional video work such as ads, branded content, and narrative shorts, with an emphasis on longer scenes, multi-asset direction, and iterative refinement. It supports text- and image-driven generation plus R2V reference control, combines multimodal assets, produces up to 30 seconds in standard mode, and can revise selected regions without regenerating the full clip. The main reasons to choose it over Seedance 2.0 are longer continuous shots, richer reference input, and more controllable post-generation fixes.
Input Type:
Output Type:
Multimodal Output
Minimax
MiniMax-H3
MiniMax
PER SEC$0.0800/s
An omni-modal video generation model that can combine text with image, video, and audio references to create video assets with synchronized audio. It is suited to short-form video, advertising, and reference-driven character or camera work; it is not a general chat model, so prioritize visual consistency, motion, and audio-visual quality.
Input Type:
Output Type:
Multimodal Output
Claude
claude-opus-5
Anthropic
CONTEXT1M
Input
$1.7000/M$5.0000/M
Output
$8.5000/M$25.0000/M
A premium model for complex agentic coding and enterprise knowledge work, emphasizing long-horizon planning, sustained execution, self-checking, and reliable delivery. With a million-token context window, it is suited to large codebases, complex documents, and multi-step projects that need limited human oversight.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Gemini
gemini-3.6-flash
Google
CONTEXT1.05M
Input$1.5000/M
Output$7.5000/M
The workhorse of the Gemini Flash line, balancing efficiency with coding, knowledge work, and multimodal understanding. Compared with 3.5 Flash, it is designed to complete multi-step tasks with fewer output tokens, reasoning steps, and tool calls, making it a good fit for production agents, document and chart analysis, and frequent complex workflows.
Input Type:
Output Type:
ReasoningTool UseLong Context
Gemini
gemini-3.5-flash-lite
Google
CONTEXT1.05M
Input$0.3000/M
Output$2.5000/M
The fastest and most cost-effective model in the Gemini 3.5 line, built for high-throughput, low-latency work such as translation, classification, document processing, and lightweight agentic workflows. It supports multimodal inputs and up to a 1M-token context, while demanding reasoning, long-horizon planning, or high-quality code changes are better handled by a higher-tier Flash model.
Input Type:
Output Type:
Long Context
MoonshotAI
kimi-k3
Moonshot
CONTEXT1.05M
Input
$2.7000/M$3.0000/M
Output
$13.5000/M$15.0000/M
Moonshot flagship model with native vision and up to a 1M-token context window. It currently runs at max reasoning effort and suits long-horizon coding, knowledge work, and complex tasks that coordinate terminal tools. For stable use, preserve the full reasoning history and start a new session in a compatible agent harness.
Input Type:
Output Type:
ReasoningTool UseLong Context
MoonshotAI
coding-kimi-k3-free
Free channel
CONTEXT1.05M
Input$0.0000/M
Output$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Input Type:
Output Type:
ReasoningTool UseLong Context
OpenAI
gpt-5.6-terra
OpenAI
CONTEXT1.05M
Input
$0.4000/M$2.0000/M
Output
$2.4000/M$12.0000/M
The balanced workhorse tier of GPT-5.6, combining capability and cost. Well suited to production work involving image understanding, tool use, and large source material; choose Sol for the hardest quality-first work or Luna for cost-controlled high volume.
Input Type:
Output Type:
ReasoningFunction CallingStructured OutputLong Context
OpenAI
gpt-5.6-sol
OpenAI
CONTEXT1.05M
Input
$1.0000/M$5.0000/M
Output
$6.0000/M$30.0000/M
The flagship GPT-5.6 tier for complex professional work and quality-first delivery. It supports image input, function calling, structured outputs, and an exceptionally large context window; consider Terra or Luna when latency or unit cost is the primary constraint.
Input Type:
Output Type:
ReasoningFunction CallingStructured OutputLong Context
OpenAI
gpt-5.6-luna
OpenAI
CONTEXT1.05M
Input
$0.0400/M$0.2000/M
Output
$0.2400/M$1.2000/M
The cost-sensitive, high-volume tier of GPT-5.6. A good fit for summarization, classification, rewriting, and batch automation; choose the Sol sibling for complex professional judgment or quality-first multi-step work.
Input Type:
Output Type:
ReasoningFunction CallingStructured OutputLong Context
Doubao
doubao-seedream-5-0-pro
Doubao
PER SEC$0.0450/s
A multimodal image-generation model for professional visual creation. It can generate images from text and perform precise edits with reference images and spatial annotations, with strengths in dense infographics, typography, realistic lighting, and multilingual text rendering.
Input Type:
Output Type:
ReasoningMultimodal Output
Grok
grok-4.5
xAI
CONTEXT500K
Input$2.1000/M
Output$6.3000/M
A frontier reasoning model for coding, knowledge work, and STEM tasks. It supports image input, function calling, and structured outputs; its 500K-token context suits cross-file analysis and long source material. Prefer a lower-cost model when frontier reasoning is unnecessary.
Input Type:
Output Type:
ReasoningFunction CallingStructured OutputLong Context
Claude
claude-fable-5
Anthropic
CONTEXT1M
Input
$3.4000/M$10.0000/M
Output
$17.0000/M$50.0000/M
Anthropic's most capable widely released model, built for the most demanding reasoning and long-horizon agentic work. Offers a 1M-token context window with 128k max output and always-on adaptive thinking, excelling at autonomously exploring underspecified tasks, planning, and carrying long-running coding and multi-agent orchestration further before it needs human input.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Gemini
gemini-3.1-flash-lite-image
Google
PER CALL
$0.0120/call$0.0200/call
A high-efficiency image generation and editing model for frequent visual drafts, social assets, posters, and iterative revisions. It supports multimodal inputs and a million-token context, favoring speed and scaled usage; choose Pro Image when complex creative control is the priority.
Input Type:
Output Type:
Multimodal Output
Claude
claude-sonnet-5
Anthropic
CONTEXT1M
Input
$0.6800/M$2.0000/M
Output
$3.4000/M$10.0000/M
A daily-driver Sonnet-tier model aimed at bringing near-frontier agentic, coding, and knowledge-work capability at a lower operating cost. Its default 1M-token context and adaptive thinking fit long documents, codebases, and multi-step tool workflows; for the deepest reasoning or restricted high-risk security work, evaluate higher-tier or specialized models.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
LongCat
LongCat-2.0
美团
CONTEXT1M
Input$0.3000/M
Output$1.2000/M
A LongCat 2.0 workhorse for project-scale coding and long-running agent tasks, with native 1M-token context plus tool calling and multi-step reasoning. It is best when you need to keep large repositories, long documents, or automation workflows in scope; for lightweight Q&A or low-latency chat, a smaller Flash/Lite-style model is usually cheaper.
Input Type:
Output Type:
ReasoningTool UseFunction CallingLong Context
Doubao
doubao-seed-2-1-pro
Doubao
CONTEXT256K
Input$0.9700/M
Output$4.8800/M
Built for coding, agents, and complex productivity tasks, advancing beyond Doubao Seed 2.0 in long context, long output, and tool-oriented workflows. Compared with Qwen, GLM, and DeepSeek peers, it is a cost-effective option for Chinese office work, coding, and multi-step automation.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Qwen
happyhorse-1.1-t2v
阿里巴巴
PER SEC$0.1350/s
Built for text-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
Qwen
happyhorse-1.1-r2v
阿里巴巴
PER SEC$0.1350/s
Built for reference-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
Qwen
happyhorse-1.1-i2v
阿里巴巴
PER SEC$0.1350/s
Built for image-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
Doubao
doubao-seedance-2-0-mini
Doubao
Input
$0.7560/M$1.5120/M
Output
$0.7560/M$1.5120/M
Built for video generation and multi-asset video creation, improving on Seedance 1.x with more complete motion stability, audiovisual generation, and reference inputs. Compared with Veo, Kling, and Runway, it is better suited to Chinese creative workflows driven by mixed text, image, audio, and video inputs. The Mini tier is positioned for lower-cost, high-frequency testing compared with standard Seedance 2.0. Useful for short drama, ads, e-commerce, and social assets.
Input Type:
Output Type:
Multimodal Output
Doubao
dreamina-seedance-2-0-mini-filter-off
Doubao
Input$1.5750/M
Output$1.5750/M
Built for video generation and multi-asset video creation, improving on Seedance 1.x with more complete motion stability, audiovisual generation, and reference inputs. Compared with Veo, Kling, and Runway, it is better suited to Chinese creative workflows driven by mixed text, image, audio, and video inputs. This filter-off variant keeps the same tier of generation capability as standard Seedance 2.0 while applying looser content filtering. Useful for short drama, ads, e-commerce, and social assets.
Input Type:
Output Type:
Multimodal Output
Doubao
dreamina-seedance-2-0-mini
Doubao
Input$75.0000/M
Output$75.0000/M
A lighter Seedance 2.0 tier for video creation, useful for quickly turning text or reference images into short video assets for social, ecommerce, and creative drafts. It is oriented toward rapid iteration and asset production; review motion, camera behavior, and subject consistency before final selection.
Input Type:
Output Type:
Multimodal Output
Zhipu
glm-5.2
Zhipu AI
CONTEXT1M
Input
$0.9690/M$1.1400/M
Output
$3.4000/M$4.0000/M
Built for reasoning, coding, and long-context agent tasks, improving over GLM-5.1 in context length, tool use, and complex multi-step workflows. Compared with Chinese flagship peers such as Qwen, DeepSeek, and Kimi, it fits enterprise automation, project-level code analysis, and knowledge work in Chinese.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
MoonshotAI
kimi-k2.7-code-highspeed
Moonshot
CONTEXT256K
Input$1.9000/M
Output$7.9999/M
Built for long-context, multi-step tool use, and code/document workflows, with stronger engineering-task stability and context capacity than Kimi K2.6. Compared with Qwen, GLM, and DeepSeek, it is well suited to Chinese long-document analysis, project Q&A, and everyday agentic programming.
Input Type:
Output Type:
Tool UseFunction CallingStructured OutputLong Context
MoonshotAI
kimi-k2.7-code
Moonshot
CONTEXT256K
Input
$0.8775/M$0.9750/M
Output
$3.6450/M$4.0500/M
Built for long-context, multi-step tool use, and code/document workflows, with stronger engineering-task stability and context capacity than Kimi K2.6. Compared with Qwen, GLM, and DeepSeek, it is well suited to Chinese long-document analysis, project Q&A, and everyday agentic programming.
Input Type:
Output Type:
Tool UseFunction CallingStructured OutputLong Context
Grok
grok-imagine-video-1-5-preview
xAI
PER SEC$0.1400/s
Built for image/video-driven generation, editing, and extension, with stronger motion consistency and fast social-content production than Grok Imagine 1.0. Compared with Veo, Kling, and Runway, it fits short-form videos, ad concepts, and trend-driven assets in the xAI ecosystem.
Input Type:
Output Type:
Multimodal Output
Qwen
qwen3.7-plus
Qwen
CONTEXT1M
Input
$0.2400/M$0.3000/M
Output
$0.9600/M$1.2000/M
Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Minimax
minimax-m3-free
Free channel
CONTEXT1M
Input$0.0000/M
Output$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Minimax
MiniMax-M3
MiniMax
CONTEXT1M
Input$1.2600/M
Output$5.0400/M
Built for long-horizon coding, tool use, and multi-turn production collaboration, extending beyond MiniMax M2.7 with multimodal inputs while keeping a 1M-token context window. Compared with Claude, Gemini, and GLM-5.2 long-context agent models, it is better for putting text, image, and video materials into one workflow.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Claude
claude-opus-4-8
Anthropic
CONTEXT1M
Input
$1.7000/M$5.0000/M
Output
$8.5000/M$25.0000/M
Anthropic current Opus flagship for the most demanding reasoning, long-running agents, and high-autonomy coding. It supports a 1M-token context window and adaptive thinking; choose it when quality and reliability matter more than latency or cost.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Gemini
gemini-omni-video
Google
PER SEC$0.2450/s
Built for any-input-to-video multimodal creation, putting stronger emphasis than earlier Gemini video workflows on unified orchestration of text, image, audio, and video materials. Compared with Veo, Kling, and Runway, it is better for integrating complex multimodal assets into one creative workflow.
Input Type:
Output Type:
Multimodal Output
Gemini
gemini-3.5-flash
Google
CONTEXT1M
Input$1.5000/M
Output$9.0000/M
Built for multimodal understanding, coding, and long-context tasks, improving over Gemini 3.1 in throughput, context, and tool capabilities. Compared with GPT, Claude, and Qwen models, it is strong for unified analysis workflows across text, images, audio, video, and documents.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Qwen
qwen3.7-max
Qwen
CONTEXT1M
Input
$1.4400/M$1.8000/M
Output
$4.3200/M$5.4000/M
Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Qwen
happyhorse-1.0-video-edit
Qwen
CONTEXT8K
PER SEC$0.1350/s
Built for video editing, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
Qwen
happyhorse-1.0-t2v
Qwen
CONTEXT8K
PER SEC$0.1350/s
Built for text-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
Qwen
happyhorse-1.0-r2v
Qwen
CONTEXT8K
PER SEC$0.1350/s
Built for reference-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
Qwen
happyhorse-1.0-i2v
Qwen
CONTEXT8K
PER SEC$0.1350/s
Built for image-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Input Type:
Output Type:
Multimodal Output
DeepSeek
deepseek-v4-pro
DeepSeek
CONTEXT1M
Input
$1.5390/M$1.7100/M
Output
$3.0869/M$3.4299/M
Built for reasoning, coding, and agentic tool use, with stronger long-context, thinking-mode, and complex-task execution than DeepSeek V3.2. Compared with Chinese peers such as Qwen, GLM, and Kimi, it is especially competitive for math, programming, and logical reasoning tasks.
Input Type:
Output Type:
Long ContextTool UseReasoningStructured Output
DeepSeek
deepseek-v4-flash
DeepSeek
CONTEXT1M
Input
$0.1278/M$0.1420/M
Output
$0.2574/M$0.2860/M
Built for reasoning, coding, and agentic tool use, with stronger long-context, thinking-mode, and complex-task execution than DeepSeek V3.2. Compared with Chinese peers such as Qwen, GLM, and Kimi, it is especially competitive for math, programming, and logical reasoning tasks.
Input Type:
Output Type:
Long ContextTool UseReasoningStructured Output
OpenAI
gpt-5.5
OpenAI
CONTEXT1M
Input
$1.0000/M$5.0000/M
Output
$6.0000/M$30.0000/M
Built for reasoning, coding, visual understanding, and agent tasks, with stronger long context, tool use, and scalable reliable output than GPT-5.4. Compared with Claude, Gemini, and Qwen flagship models, it is a strong general-purpose base for reliable automation and complex knowledge workflows.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
XiaomiMiMo
mimo-v2.5-tts-voicedesign
Xiaomi MiMo
Input$0.0000/M
Output$0.0000/M
Built for one-sentence voice design and text-to-speech, with stronger V2.5-series natural-language voice creation, fine-grained pace/emotion/tone control, and low-friction voice production than MiMo V2 TTS. Compared with ElevenLabs, OpenAI TTS, and Azure Speech, it is useful for reference-free character voice design, ad voiceovers, short-video narration, and multi-style voice exploration.
Input Type:
Output Type:
Multimodal Output
XiaomiMiMo
mimo-v2.5-tts-voiceclone
Xiaomi MiMo
Input$0.0000/M
Output$0.0000/M
Built for few-sample voice cloning, with stronger V2.5-series high-fidelity timbre reproduction, voice-character consistency, and cross-text generalization than MiMo V2 TTS. Compared with ElevenLabs, OpenAI TTS, and Azure Speech, it is well suited to Chinese character dubbing, short-video narration, podcast voice replication, and multilingual speech production that needs to preserve a target voice.
Input Type:
Output Type:
Multimodal Output
XiaomiMiMo
mimo-v2.5-tts
Xiaomi MiMo
Input$0.0000/M
Output$0.0000/M
Built for text-to-speech, voice design, and voice cloning, with stronger V2.5-series natural speech synthesis, fine-grained pace/emotion/tone control, and low-friction voice creation than MiMo V2 TTS. Compared with ElevenLabs, OpenAI TTS, and Azure Speech, it is well suited to Chinese narration, character dubbing, short-video audio, and multilingual voice production.
Input Type:
Output Type:
Multimodal Output
XiaomiMiMo
mimo-v2.5-pro
Xiaomi MiMo
CONTEXT1M
Input$1.1700/M
Output$3.5000/M
Built for agents, coding, and long-context tasks, offering more complete multimodal understanding, a million-token context window, and long-horizon reasoning than MiMo V2. Compared with Chinese peers such as Qwen, GLM, and DeepSeek, its strengths are Xiaomi ecosystem fit and engineering automation scenarios.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
XiaomiMiMo
mimo-v2.5
Xiaomi MiMo
CONTEXT1M
Input$0.4700/M
Output$2.3600/M
Built for agents, coding, and long-context tasks, offering more complete multimodal understanding, a million-token context window, and long-horizon reasoning than MiMo V2. Compared with Chinese peers such as Qwen, GLM, and DeepSeek, its strengths are Xiaomi ecosystem fit and engineering automation scenarios.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
OpenAI
gpt-image-2-text-to-image
OpenAI
CONTEXT128K
PER IMG$0.0300/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Input Type:
Output Type:
Multimodal Output
OpenAI
gpt-image-2-image-to-image
OpenAI
CONTEXT128K
PER IMG$0.0200/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Input Type:
Output Type:
Multimodal Output
OpenAI
gpt-image-2
OpenAI
CONTEXT128K
PER IMG$0.0670/image
Built for high-quality image generation and editing, improving over GPT Image 1.x in text rendering, local editing, and multi-turn consistency. Compared with Gemini image models, Qwen Image, and Seedream, it is well suited to brand visuals, ad design, product imagery, and tightly controlled creative workflows.
Input Type:
Output Type:
Multimodal Output
MoonshotAI
kimi-k2.6
Moonshot
CONTEXT256K
Input
$0.8370/M$0.9300/M
Output
$3.4740/M$3.8600/M
Built for long-context, multi-step tool use, and code/document workflows, with stronger engineering-task stability and context capacity than Kimi K2.5. Compared with Qwen, GLM, and DeepSeek, it is well suited to Chinese long-document analysis, project Q&A, and everyday agentic programming.
Input Type:
Output Type:
Tool UseReasoning
Claude
claude-opus-4-7-thinking
Anthropic
CONTEXT1M
Input$5.0000/M
Output$25.0000/M
A thinking-oriented configuration variant built on Claude Opus 4.7, suited to deeper analysis, complex planning, and multi-step tool execution. It is not a separate model family; higher thinking effort generally increases latency and output usage, so use standard Opus 4.7 or Sonnet when speed matters.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Claude
claude-opus-4-7
Anthropic
CONTEXT1M
Input
$1.7000/M$5.0000/M
Output
$8.5000/M$25.0000/M
A high-capability Opus model for complex, long-running software engineering and agent tasks. Compared with Opus 4.6, it emphasizes precise instruction following, higher-resolution vision, sustained execution, and self-verification for large codebases, difficult debugging, document analysis, and multi-step tool use.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Qwen
qwen3.6-plus
Qwen
CONTEXT1M
Input
$0.9120/M$1.1400/M
Output
$5.4880/M$6.8599/M
Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.5. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Minimax
MiniMax-M2.7-highspeed
MiniMax
CONTEXT200K
Input$0.6300/M
Output$2.5200/M
Built for text generation, reasoning, and tool use, with stronger 200K long-context and agentic workflow stability than earlier MiniMax M2.x versions. Compared with Chinese peers such as Kimi, GLM, and Qwen, it fits productivity, coding assistance, and long-document processing.
Input Type:
Output Type:
Reasoning
Minimax
MiniMax-M2.7
MiniMax
CONTEXT200K
Input$0.3150/M
Output$1.2600/M
Built for text generation, reasoning, and tool use, with stronger 200K long-context and agentic workflow stability than earlier MiniMax M2.x versions. Compared with Chinese peers such as Kimi, GLM, and Qwen, it fits productivity, coding assistance, and long-document processing.
Input Type:
Output Type:
Reasoning
OpenAI
gpt-5.4-mini
OpenAI
CONTEXT400K
Input
$0.1500/M$0.7500/M
Output
$0.9000/M$4.5000/M
Built for reasoning, coding, visual understanding, and agent tasks, with stronger long context, tool use, and scalable reliable output than GPT-5.3. Compared with Claude, Gemini, and Qwen flagship models, it is a strong general-purpose base for reliable automation and complex knowledge workflows.
Input Type:
Output Type:
Long ContextTool UseReasoningStructured Output
OpenAI
gpt-5.4
OpenAI
CONTEXT1.05M
Input
$0.5000/M$2.5000/M
Output
$3.0000/M$15.0000/M
Built for reasoning, coding, visual understanding, and agent tasks, with stronger long context, tool use, and scalable reliable output than GPT-5.3. Compared with Claude, Gemini, and Qwen flagship models, it is a strong general-purpose base for reliable automation and complex knowledge workflows.
Input Type:
Output Type:
Long ContextTool UseReasoningStructured Output
Qwen
qwen-image-2.0-pro
Qwen
Input$2.2000/M
Output$2.2000/M
Built for image generation and editing, improving over Qwen Image 1.x in Chinese text rendering, instruction following, and complex layouts. Compared with GPT Image, Gemini image models, and Seedream, it is well suited to Chinese posters, e-commerce images, infographics, and iterative visual creation.
Input Type:
Output Type:
Multimodal Output
Qwen
qwen-image-2.0
Qwen
Input$2.2000/M
Output$2.2000/M
Built for image generation and editing, improving over Qwen Image 1.x in Chinese text rendering, instruction following, and complex layouts. Compared with GPT Image, Gemini image models, and Seedream, it is well suited to Chinese posters, e-commerce images, infographics, and iterative visual creation.
Input Type:
Output Type:
Multimodal Output
Gemini
gemini-3.1-flash-image-preview
Google
CONTEXT131K
PER IMG$0.0672/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Input Type:
Output Type:
ReasoningMultimodal Output
Gemini
gemini-3.1-flash-image
Google
PER IMG
$0.0198/image$0.0330/image
An efficient Gemini image tier that balances quality and speed for text-to-image generation, image editing, multi-image composition, and workflows that need fast, scalable output. It offers flexible composition and text rendering, while complex professional design may be better served by Gemini 3 Pro Image.
Input Type:
Output Type:
Multimodal Output
Gemini
gemini-3.1-pro-preview
Google
CONTEXT1.05M
Input$2.0000/M
Output$12.0000/M
Built for multimodal understanding, coding, and long-context tasks, improving over Gemini 3.0 / 2.5 in throughput, context, and tool capabilities. Compared with GPT, Claude, and Qwen models, it is strong for unified analysis workflows across text, images, audio, video, and documents.
Input Type:
Output Type:
Long Context
Claude
claude-sonnet-4-6-thinking
Anthropic
CONTEXT1M
Input$3.0000/M
Output$15.0000/M
The extended thinking variant of Anthropic Claude Sonnet 4.6, enabling chain-of-thought reasoning for complex problems. Excels in coding, codebase navigation, project management, document creation, and agentic tasks with text and image input. 1M context window, outperforms the standard version on tasks requiring deep reasoning, ideal for complex software engineering and multi-step planning.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Claude
claude-sonnet-4-6
Anthropic
CONTEXT1M
Input
$1.0200/M$3.0000/M
Output
$5.1000/M$15.0000/M
The Claude lineup workhorse, balancing strong intelligence, response speed, and cost. It supports a 1M-token context window and adaptive thinking, making it a strong starting point for coding, document analysis, data work, and tool-using agents; move to Opus for the most complex tasks.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Doubao
doubao-seed-2-0-pro
Doubao
CONTEXT256K
Input$0.5000/M
Output$2.5300/M
Built for coding, agents, and complex productivity tasks, advancing beyond Doubao Seed 2.0 in long context, long output, and tool-oriented workflows. Compared with Qwen, GLM, and DeepSeek peers, it is a cost-effective option for Chinese office work, coding, and multi-step automation.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong Context
Doubao
doubao-seedream-5.0-lite
Doubao
PER IMG$0.0350/image
Built for image generation and editing, with stronger Chinese instruction following, visual reasoning, and high-frequency creative production than Seedream 4.x. Compared with GPT Image, Gemini image models, and Qwen Image, it is well suited to ByteDance-ecosystem ad design, e-commerce assets, and multi-scenario visual production.
Input Type:
Output Type:
Web SearchMultimodal Output
Minimax
MiniMax-M2.5
MiniMax
CONTEXT205K
Input$0.3020/M
Output$1.2077/M
A cost-efficient model for real productivity agents, strongest in coding, tool use, search, and office-style deliverables rather than casual chat alone. It fits agents that plan and execute cross-file changes, research retrieval, and document/spreadsheet/presentation workflows; versus more expensive flagships, the main tradeoff is lower cost for near-frontier execution efficiency.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
Doubao
doubao-seedance-2-0-filter-off
Doubao
CONTEXT13K
Input$3.1500/M
Output$3.1500/M
Built for video generation and multi-asset video creation, improving on Seedance 1.x with more complete motion stability, audiovisual generation, and reference inputs. Compared with Veo, Kling, and Runway, it is better suited to Chinese creative workflows driven by mixed text, image, audio, and video inputs. This filter-off variant keeps the same tier of generation capability as standard Seedance 2.0 while applying looser content filtering. Useful for short drama, ads, e-commerce, and social assets.
Input Type:
Output Type:
Multimodal Output
Doubao
doubao-seedance-2-0-fast-filter-off
Doubao
CONTEXT13K
Input$2.5200/M
Output$2.5200/M
Built for video generation and multi-asset video creation, improving on Seedance 1.x with more complete motion stability, audiovisual generation, and reference inputs. Compared with Veo, Kling, and Runway, it is better suited to Chinese creative workflows driven by mixed text, image, audio, and video inputs. This filter-off variant keeps the same tier of generation capability as standard Seedance 2.0 while applying looser content filtering. Useful for short drama, ads, e-commerce, and social assets.
Input Type:
Output Type:
Multimodal Output
Claude
claude-opus-4-6
Anthropic
CONTEXT1M
Input
$1.7000/M$5.0000/M
Output
$8.5000/M$25.0000/M
An Opus model for demanding reasoning, agentic coding, and large-codebase work. Compared with Opus 4.5, it emphasizes longer-horizon planning, code review, debugging, autonomous execution, and document and spreadsheet work, with a 1M-token context window and adaptive thinking. Evaluate newer Opus 4.7 or 4.8 first for new projects.
Input Type:
Output Type:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Kling
kling-v3-omni
kling
PER SEC$0.0900/s
Built for text/image-to-video and multi-shot creation, improving over Kling 2.x in motion stability, camera language, and reference consistency. Compared with Veo, Runway, and Seedance, it is well suited to Chinese short drama, ad assets, and high-quality social video production.
Input Type:
Output Type:
Multimodal Output
Kling
kling-v3
kling
PER SEC$0.0900/s
Built for text/image-to-video and multi-shot creation, improving over Kling 2.x in motion stability, camera language, and reference consistency. Compared with Veo, Runway, and Seedance, it is well suited to Chinese short drama, ad assets, and high-quality social video production.
Input Type:
Output Type:
Multimodal Output
Doubao
doubao-seedance-2-0-fast
Doubao
CONTEXT8K
Input$2.4192/M
Output$2.4192/M
Built for video generation and multi-asset video creation, improving on Seedance 1.x with more complete motion stability, audiovisual generation, and reference inputs. Compared with Veo, Kling, and Runway, it is better suited to Chinese creative workflows driven by mixed text, image, audio, and video inputs. The fast variant is better for rapid iteration and batch production than the standard tier. Useful for short drama, ads, e-commerce, and social assets.
Input Type:
Output Type:
Multimodal Output
Doubao
doubao-seedance-2-0
Doubao
CONTEXT8K
Input
$2.5927/M$2.7292/M
Output
$2.5927/M$2.7292/M
Built for video generation and multi-asset video creation, improving on Seedance 1.x with more complete motion stability, audiovisual generation, and reference inputs. Compared with Veo, Kling, and Runway, it is better suited to Chinese creative workflows driven by mixed text, image, audio, and video inputs. Useful for short drama, ads, e-commerce, and social assets.
Input Type:
Output Type:
Multimodal Output
MoonshotAI
kimi-k2.5
Moonshot
CONTEXT256K
Input
$0.5400/M$0.6000/M
Output
$2.8350/M$3.1500/M
Built for long-context, multi-step tool use, and code/document workflows, with stronger engineering-task stability and context capacity than Kimi K2. Compared with Qwen, GLM, and DeepSeek, it is well suited to Chinese long-document analysis, project Q&A, and everyday agentic programming.
Input Type:
Output Type:
Long ContextTool UseReasoning
Gemini
gemini-3-flash-preview
Google
CONTEXT1M
Input$0.5000/M
Output$3.0000/M
Google's high-speed thinking model designed for agentic workflows, multi-turn chat, and coding assistance. Supports text, image, audio, video, and PDF input with a 1M token context window. Features configurable thinking levels, tool use, and structured output. Broad quality improvements over Gemini 2.5 Flash across reasoning, multimodal understanding, and reliability.
Input Type:
Output Type:
ReasoningTool UseFunction CallingStructured OutputLong Context
DeepSeek
DeepSeek-V3.2-Thinking
DeepSeek
CONTEXT128K
Input$0.3100/M
Output$0.4700/M
Built for reasoning, coding, and agentic tool use, with stronger long-context, thinking-mode, and complex-task execution than DeepSeek V3.1. Compared with Chinese peers such as Qwen, GLM, and Kimi, it is especially competitive for math, programming, and logical reasoning tasks.
Input Type:
Output Type:
Long ContextTool UseReasoningStructured Output
DeepSeek
deepseek-v3.2
DeepSeek
CONTEXT131K
Input$0.2100/M
Output$0.3200/M
Built for reasoning, coding, and agentic tool use, with stronger long-context, thinking-mode, and complex-task execution than DeepSeek V3.1. Compared with Chinese peers such as Qwen, GLM, and Kimi, it is especially competitive for math, programming, and logical reasoning tasks.
Input Type:
Output Type:
ReasoningTool Use
Gemini
gemini-3-pro-image-preview
Google
CONTEXT66K
PER IMG$0.1340/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Input Type:
Output Type:
ReasoningMultimodal Output
Gemini
gemini-3-pro-image
Google
PER IMG
$0.0318/image$0.0530/image
A reasoning-driven image generation and editing model for professional creative work, including complex graphic design, high-fidelity product mockups, infographics, and visuals that require accurate text rendering. Compared with the faster Flash Image tier, it prioritizes fine control, grounded content, and polished output.
Input Type:
Output Type:
Multimodal Output
Claude
claude-opus-4-5
Anthropic
CONTEXT200K
Input$5.0000/M
Output$25.0000/M
An earlier Opus model in the Claude 4 family, focused on difficult coding, agents, computer use, deep research, and work with slides and spreadsheets. It fits workflows that require Opus-level capability with 4.5 compatibility; without that constraint, evaluate the newer Opus 4.6 or 4.7 first.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Gemini
veo3.1-lite
Google
PER CALL$0.1100/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion. The Lite tier favors high-throughput, lower-cost drafts over the standard tier.
Input Type:
Output Type:
Multimodal Output
Gemini
veo3.1-fast
Google
PER CALL$0.2211/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion. The Fast tier is better for rapid concept validation than the standard tier.
Input Type:
Output Type:
Multimodal Output
Gemini
veo3.1
Google
PER CALL$1.6544/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion.
Input Type:
Output Type:
Multimodal Output
Gemini
gemini-2.5-flash-image
Google
CONTEXT66K
PER IMG$0.0585/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Input Type:
Output Type:
Structured OutputMultimodal Output
Claude
claude-haiku-4-5
Anthropic
CONTEXT200K
Input
$0.3400/M$1.0000/M
Output
$1.7000/M$5.0000/M
The fastest and more cost-efficient model in the Claude lineup, offering near-frontier capability with a 200K-token context window and thinking support. It fits real-time interaction, batch processing, lightweight agents, and cost-sensitive production; use Sonnet or Opus for the hardest reasoning or code changes.
Input Type:
Output Type:
ReasoningTool UseStructured OutputLong Context
Qwen
wan3.0-video-prime
Qwen
PER SEC$0.4500/s
The speed-first variant of Alibaba's Wan 3.0 line (early access). It shortens generation latency while retaining high-quality video output, making it a good fit for low-latency previews, batch drafts and fast-iterating short-video workflows. If ultimate visual control matters more than speed, the standard version is a better choice.
Input Type:
Output Type:
Multimodal Output
Qwen
wan3.0-video
Qwen
PER SEC$0.3000/s
The standard video generation model of Alibaba's Wan 3.0 line (early access). It supports text-to-video, image-to-video and multi-modal reference inputs, generating clips up to roughly 30 seconds with synchronized audio in a single pass — suitable for ads, short drama and narrative content. Official release date not yet announced.
Input Type:
Output Type:
Multimodal Output
Zhipu
glm-image
Zhipu AI
Input$0.8800/M
Output$3.5500/M
Built for image generation and editing, with a stronger focus than earlier GLM vision capabilities on Chinese prompts, image-text understanding, and controllable generation. Compared with GPT Image, Gemini image models, and Qwen Image, it fits Chinese marketing assets, explanatory visuals, and iterative visual edits.
Input Type:
Output Type:
Multimodal Output
WhatsApp