Alibaba
Qwen3.8 Omni Flash
Total Context
1M
Max Output
131.1K
Released
N/A
qwen3.5-35b-a3bQwen3.5 35B A3B is a native vision-language MoE model designed to approximate much larger model behavior with a smaller active footprint. Model cards describe hybrid linear attention, sparse experts, and comparable results to larger Qwen3.5 dense variants. The strongest positioning is efficient multimodal reasoning and coding without the cost of always activating a large dense model.
Context Window
262.1K tokens
Maximum Output
65.5K tokens
Release Date
Feb 23, 2026
Modalities
| Token Tier | Input Price | Output Price |
|---|---|---|
| <=128K | $0.0571/M | $0.4571/M |
| >128K | $0.2286/M | $1.8286/M |
Reasoning
Tool calling
Temperature parameter
Attachments
Knowledge Base
Endpoint Protocols
Qwen3.5-35B-A3B combines an efficient sparse MoE design, native multimodal understanding and open-weight deployment for adaptable applications.
The model has 35 billion total parameters with 3 billion activated, combining Gated DeltaNet attention with sparse experts for efficient inference.
A unified vision-language foundation processes text, images and video for reasoning, document understanding and visual tasks.
Official weights under Apache 2.0 support self-managed inference with frameworks including Transformers, vLLM and SGLang.
Qwen3.5-35B-A3B is suited to self-hosted multimodal assistants, visual document processing and adaptable coding workflows.
Deploy the open weights in controlled infrastructure to answer questions over internal text, images and videos.
Read text, tables and diagrams from document images, extract relevant fields and generate structured summaries.
Generate and revise code, investigate defects and integrate the model into self-managed developer tools or agents.
Replace these path values before running: {model}
curl 'https://us-api.tokenhub.com/v1beta/models/{model}:generateContent' \
-X 'POST' \
-H "Authorization: Bearer $TOKENHUB_API_KEY"| Index score | ||
|---|---|---|
| Artificial Analysis Intelligence Index | Artificial Analysis broad capability aggregate | 23.4 |
| Artificial Analysis Coding Index | Artificial Analysis software task aggregate | 16.8 |
| Knowledge & Reasoning | ||
| GPQA | Advanced science problem solving | 81.9% |
| HLE | Broad expert-level exam set | 12.8% |
| Coding & Engineering | ||
| SciCode | Scientific coding challenges | 29.3% |
| Terminal-Bench Hard | Hard terminal task execution | 10.6% |
| Instruction Following & Agent Tasks | ||
| IFBench | Prompt constraint adherence | 44.5% |
| AA-LCR | Long-context reasoning | 55.3% |
| τ²-Bench | Agent workflow tasks | 86.3% |
Metrics sourced from Artificial Analysis
Qwen 3.5 35B A3B: capabilities, use cases, limits, and TokenHub guidance.
Qwen 3.5 35B A3B is a Alibaba Qwen model for open-model multimodal reasoning and efficient deployment.
Best for self-hosted deployment, visual reasoning and routine coding assistance, especially when deployment control is the priority.
Key strength: an open MoE variant with a small active-parameter footprint and hybrid thinking that can switch between deliberate and direct responses.
It belongs to an older generation and may lack newer capabilities. For the latest capabilities matter, consider Qwen 3.6 35B A3B.
Use TokenHub's exact ID; hosted behavior may differ from self-hosting.
Use one API key to access Qwen3.5 35B-A3B and more AI models through TokenHub.
Media and Discussions
Selected public videos and posts related to this model.
X (Twitter)
Reddit
YouTube