Qwen3.5-397B-A17B — Open Multimodal MoE for Agents, Not Raw Scale Alone
Alibaba Qwen released Qwen3.5-397B-A17B (Feb 16; blog Feb 15): open-weight natively multimodal MoE—397B total, 17B active—Apache 2.0. Vendor claims competitive vs GPT-5.2/Claude 4.5 Opus/Gemini-3 Pro families plus 8.6×–19× inference throughput vs prior gen; hosted Plus adds 1M context. Efficiency + open license as the China open-frontier trade.

Lunar New Year week stacked Chinese open and efficient releases. Alibaba’s proprietary close for Qwen3.5’s flagship: frontier-ish multimodal agents under Apache 2.0 with extreme sparsity—capability diffusion as the product, not only Elo screenshots.
Alibaba’s Qwen team released Qwen3.5-397B-A17B (announced February 16; blog timestamp 2026/02/15)—first model in the Qwen3.5 series. Open-weight, natively multimodal (vision-language via early fusion) Mixture-of-Experts: 397B total parameters, 17B active per token (high-sparsity design). Hybrid attention: Gated Delta Networks (linear attention) with sparse MoE. Pretraining: scaled visual-text tokens, STEM/reasoning, multilingual coverage (201 languages/dialects; larger 250k vocabulary). Post-training: broad RL across difficult environments for agentic tool use, planning, long-horizon tasks. Vendor benches: competitive standing vs GPT-5.2, Claude 4.5 Opus, Gemini-3 Pro families across language, VL, coding agent, search agent—with claimed inference throughput 8.6×–19× vs prior generation at common context lengths. Hosted Qwen3.5-Plus: 1M context and built-in adaptive tool use. Weights: Hugging Face, ModelScope, GitHub under Apache 2.0.
Claims vs checks
Architecture, parameter counts, Apache 2.0, and weight locations are Qwen primary. Cross-lab “competitive/leading” and throughput multiples are vendor (Reuters relays performance/cost claims)—check Artificial Analysis / task boards before ranking freezes. Plus 1M is hosted SKU, not necessarily identical to open weights.
Limits
- MoE serving quality depends on infrastructure; laptop demos ≠ 397B cluster reality.
- Multilingual breadth ≠ equal quality per language.
- Agentic benches are sensitive to tool scaffolds—compare like-with-like.
Sources
- Qwen blog: “Qwen3.5: Towards Native Multimodal Agents” (2026/02/15–16). Primary.
- Reuters: “Alibaba unveils new Qwen3.5 model for ‘agentic AI era’” (Feb 16, 2026).
- Hugging Face / GitHub references in the announcement for Qwen3.5-397B-A17B.