Qwen3.8 Max

qwen3.8-max

Qwen3.8-Max is a 2.4T-parameter MoE flagship model designed for advanced reasoning, coding, and enterprise productivity. It can autonomously execute complex, long-running tasks — from software development to professional workflows — delivering production-grade outcomes across domains such as law, finance, and design. Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.

Total Context

1Mtokens

Max Output

128Ktokens

Released

Aug 3, 2026

Modalities

Qwen3.8 Max Price

Input PriceOutput PriceCache ReadCache Create 5mCache Read 5m
$2/M$6/M$0.25/M$2.5/M$0.17/M

How do I use Qwen3.8 Max through the API?

POSTopenai/v1/chat/completions
POSTopenai-response/v1/responses
POSTanthropic/v1/messages

Qwen3.8 Max Media and Demos

Selected public announcements, demonstrations, and community discussions about Qwen3.8-Max.

X (Twitter)

Open social media source
View post on X
Open social media source
View post on X
Open social media source
View post on X

Reddit

YouTube

Open social media source
Watch on YouTube
Open social media source
Watch on YouTube
Open social media source
Watch on YouTube

Qwen3.8 Max FAQs

Practical answers about Qwen3.8-Max capabilities, use cases, pricing, and TokenHub access.

What is Qwen3.8-Max?+

Qwen3.8-Max is Alibaba's flagship Qwen mixture-of-experts model. It accepts text, images, and video and returns text; the provider reports 2.4 trillion total parameters with 95 billion active.

What is Qwen3.8-Max best used for?+

It is especially suited to advanced coding, long-horizon agent workflows, complex professional work, and multimodal analysis of long documents or videos.

What are the main strengths of Qwen3.8-Max?+

Its strengths include a 1M-token context window, optional thinking, function calling, structured outputs, built-in tools such as web search and code interpretation, and native visual understanding.

What tradeoffs should I consider with Qwen3.8-Max?+

Flagship capability can cost more and take longer than lighter Qwen tiers. Long agent runs may compound errors and spending, so set budgets, add checkpoints, and review high-stakes outputs.

How do I use Qwen3.8-Max on TokenHub?+

Choose qwen3.8-max in TokenHub's model field, then send requests to your TokenHub endpoint with your account credentials. Confirm the supported parameters, input formats, and limits shown in your TokenHub account before production use.

How is Qwen3.8-Max priced and available?+

As of August 2026, QwenCloud lists qwen3.8-max for API use at $2 per 1M input tokens, $6 per 1M output tokens, and $0.25 per 1M implicit-cache input tokens. TokenHub availability and billing may differ, so check its live model page.

When should I choose Qwen3.8-Max over another model?+

Choose it when you need Qwen's highest capability for complex coding, multimodal reasoning, or long agentic work. For cost-sensitive workloads, compare qwen3.7-plus or qwen3.7-flash and evaluate quality, latency, and cost on representative tasks.