Skip to content

Qwen3.8-Max API Pricing and Cheaper Options

Current Qwen3.8-Max API pricing by region, its 1M context window, cache discounts, and when a smaller Qwen model costs less.

MGMCSA Guru Team July 28, 2026 2 min read
Qwen3.8-Max API pricing and cache costs

qwen3.8-max is Alibaba’s current production flagship. It replaced the preview route in August 2026 and is a substantial change from the older Qwen3 Max covered by this URL. It accepts text, images, and video, supports function calling and web search, and has a 1 million-token context window.

Current API prices

Alibaba prices the model by deployment region. These are the list prices shown in the model documentation on August 24, 2026:

Qwen3.8-Max price per 1M tokens

Singapore $2 input; $6 output
Germany, US, Japan, Hong Kong $1.65 input; $4.951 output
Implicit cache hit in Singapore $0.25 input
Implicit cache hit in listed global regions $0.206 input
Explicit cache read in Singapore $0.17 input
Explicit cache read in listed global regions $0.137 input

Promotions can change the invoice, so check the Model Studio console before making a budget. The table above uses published list prices, not a temporary sale.

Cache repeated context

Cache pricing matters when an agent repeatedly sends the same repository instructions or long document. In Singapore, an implicit cache hit costs one eighth of normal input. Explicit cache creation costs more than ordinary input, so it pays off only when the cached material is reused.

When to use a cheaper Qwen

Qwen3.8-Max is appropriate when a task needs its multimodal input, very long context, or stronger agent behavior. Routine text transformations and small code edits may cost less on a Plus, Flash, or smaller open model. Compare the exact model IDs and regional prices in Model Studio because availability differs by workspace location.

The Qwen team also released open Qwen3.8 weights. The official repository lists a 2.4T-A95B model and a 27B model. The 27B release is the more realistic local option, although quantization, context length, and hardware still determine whether it will run well.

Use the production ID

Set the model to:

qwen3.8-max

Do not start a new configuration on qwen3.8-max-preview. Alibaba’s August upgrade notice says preview traffic is routed to the official release.

For a coding workflow, see Qwen with Claude Code and Qwen Code CLI on Windows.

Frequently asked questions

What is the current Qwen flagship model?

The current production flagship is qwen3.8-max. Alibaba retired the preview alias in August 2026 and routed it to the official release.

How much does Qwen3.8-Max cost?

In Singapore, Alibaba lists $2 per million input tokens and $6 per million output tokens. In several global regions, including Germany and the US, the listed rates are $1.65 input and $4.951 output per million tokens.

Does Qwen3.8-Max accept images?

Yes. Alibaba documents text, image, and video input with text output.

How large is its context window?

The documented context window is 1 million tokens, with up to 131,072 output tokens.

Sources & further reading

Official vendor documentation referenced while writing this guide.

MG

MCSA Guru Team

IT & Systems Administration

We are working IT pros and system administrators who spend our days in Windows Server, Microsoft 365, and the wider Microsoft stack. MCSA Guru is where we write down the fixes and walkthroughs we wish we had found the first time.

MCSA Guru provides independent, educational IT guidance. Microsoft, Windows, Windows Server, Microsoft 365, Exchange, and Microsoft Teams are trademarks of Microsoft Corporation; Docker is a trademark of Docker, Inc. MCSA Guru is not affiliated with or endorsed by Microsoft or Docker. Always test changes in a safe environment before applying them in production.

Related guides

Fixing something right now?

Jump straight into the guide library or search for the exact error or task you are dealing with.