Qwen

Qwen3.8 MAX API

Qwen's most capable flagship multimodal reasoning model, built for long-horizon coding, professional work, and multi-stage agent tasks. It accepts text, images, and video, offers a 1M-token context window with up to 128K output, and supports function calling, structured outputs, web search, and a code interpreter. It is best used as a high-quality workhorse for complex tasks rather than for the lowest-cost, lowest-latency batch workloads.

  • Reasoning
  • Web search
  • Tool use
  • Function calling
  • Structured output
  • Long context
  • Code execution
Use in Consoleqwen3.8-max

USD

Pricing

default
Standard Rate
Input$1.775/1M tokensCompletion Price$5.32997/1M tokensCache Read$0.176968/1M tokensCache Write$2.219993/1M tokens

Public list prices are shown in USD. Final charges may vary by account group and usage tier.

API

Code examples

curl --location --request POST 'https://api.tokenhot.ai/v1/chat/completions' \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data-raw '{
    "model": "qwen3.8-max",
    "messages": [
        {
            "role": "user",
            "content": "What model are you"
        }
    ]
}'

Qwen

Related models

FAQ

Frequently asked questions

What is Qwen3.8 MAX best suited for?

Qwen3.8 MAX is suited to the input, output, and capability types listed on this page. Test production workloads before choosing it for a critical system.

How is Qwen3.8 MAX priced?

Pricing is shown in the pricing section above. Actual cost depends on usage volume, account group, and the request payload.

How can I call Qwen3.8 MAX?

Use the model ID shown above with one of the supported API protocols. Code examples are provided when usage data is available.

Get started

Build with Qwen3.8 MAX

Use in Console
TokenHot

The frontier intelligence gateway. One API. 127 models. 0.2s latency. Pay only for what you use.

All systems normal · 99.997% uptime

Company

© 2026 TokenHot Inc. — Built for builders.