Qwen

Qwen3.7 MAX API

Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.

  • Reasoning
  • Tool use
  • Function calling
  • Structured output
  • Long context
Use in Consoleqwen3.7-max

USD

Pricing

default
Standard Rate
Input$1.8/1M tokensCompletion Price$5.4/1M tokensCache Read$0.36/1M tokensCache Write$2.25/1M tokens

Public list prices are shown in USD. Final charges may vary by account group and usage tier.

API

Code examples

curl --location --request POST 'https://api.tokenhot.ai/v1/chat/completions' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data-raw '{
    "model": "qwen3.7-max",
    "messages": [
        {
            "role": "user",
            "content": "What model are you"
        }
    ]
}'

Qwen

Related models

QwenQwenQwen3.7 Plus

Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.

$0.3/1M tokens
QwenQwenQwen3.6 Plus

Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.5. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.

$1.14/1M tokens
QwenQwenQwen3.8 MAX

Qwen's most capable flagship multimodal reasoning model, built for long-horizon coding, professional work, and multi-stage agent tasks. It accepts text, images, and video, offers a 1M-token context window with up to 128K output, and supports function calling, structured outputs, web search, and a code interpreter. It is best used as a high-quality workhorse for complex tasks rather than for the lowest-cost, lowest-latency batch workloads.

$1.775/1M tokens
QwenQwenQwen3.6 Flash

Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.5. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.

$0.1785/1M tokens

FAQ

Frequently asked questions

What is Qwen3.7 MAX best suited for?

Qwen3.7 MAX is suited to the input, output, and capability types listed on this page. Test production workloads before choosing it for a critical system.

How is Qwen3.7 MAX priced?

Pricing is shown in the pricing section above. Actual cost depends on usage volume, account group, and the request payload.

How can I call Qwen3.7 MAX?

Use the model ID shown above with one of the supported API protocols. Code examples are provided when usage data is available.

Get started

Build with Qwen3.7 MAX

Use in Console
TokenHot

La passerelle d'intelligence de frontière. Une API. 127 modèles. Latence de 0.2s. Payez uniquement ce que vous utilisez.

Tous les systèmes normaux · 99.997% de disponibilité

Entreprise

© 2026 TokenHot Inc. — Conçu pour les créateurs.