/v1/chat/completionsQwen3.8 Max
qwen3.8-maxQwen3.8-Max is a 2.4T-parameter MoE flagship model designed for advanced reasoning, coding, and enterprise productivity. It can autonomously execute complex, long-running tasks — from software development to professional workflows — delivering production-grade outcomes across domains such as law, finance, and design. Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.
Total Context
1Mtokens
Max Output
128Ktokens
Released
Aug 3, 2026
Modalities
Qwen3.8 Max Price
| Input Price | Output Price | Cache Read | Cache Create 5m | Cache Read 5m |
|---|---|---|---|---|
| $2/M | $6/M | $0.25/M | $2.5/M | $0.17/M |
How do I use Qwen3.8 Max through the API?
/v1/responses/v1/messagesQwen3.8 Max FAQs
Practical answers about Qwen3.8-Max capabilities, use cases, pricing, and TokenHub access.
What is Qwen3.8-Max?+
Qwen3.8-Max is Alibaba's flagship Qwen mixture-of-experts model. It accepts text, images, and video and returns text; the provider reports 2.4 trillion total parameters with 95 billion active.
What is Qwen3.8-Max best used for?+
It is especially suited to advanced coding, long-horizon agent workflows, complex professional work, and multimodal analysis of long documents or videos.
What are the main strengths of Qwen3.8-Max?+
Its strengths include a 1M-token context window, optional thinking, function calling, structured outputs, built-in tools such as web search and code interpretation, and native visual understanding.
What tradeoffs should I consider with Qwen3.8-Max?+
Flagship capability can cost more and take longer than lighter Qwen tiers. Long agent runs may compound errors and spending, so set budgets, add checkpoints, and review high-stakes outputs.
How do I use Qwen3.8-Max on TokenHub?+
Choose qwen3.8-max in TokenHub's model field, then send requests to your TokenHub endpoint with your account credentials. Confirm the supported parameters, input formats, and limits shown in your TokenHub account before production use.
How is Qwen3.8-Max priced and available?+
As of August 2026, QwenCloud lists qwen3.8-max for API use at $2 per 1M input tokens, $6 per 1M output tokens, and $0.25 per 1M implicit-cache input tokens. TokenHub availability and billing may differ, so check its live model page.
When should I choose Qwen3.8-Max over another model?+
Choose it when you need Qwen's highest capability for complex coding, multimodal reasoning, or long agentic work. For cost-sensitive workloads, compare qwen3.7-plus or qwen3.7-flash and evaluate quality, latency, and cost on representative tasks.
Qwen3.8 Max Media and Demos
Selected public announcements, demonstrations, and community discussions about Qwen3.8-Max.
X (Twitter)
Reddit
YouTube