# Qwen3.5-397B-A17B — Open Multimodal MoE for Agents, Not Raw Scale Alone

Times of AI Desk · 2026-02-16 · Models

[https://timesof.ai/2026/02/alibaba-releases-qwen3-5-397b-moe-multimodal-agentic-model](https://timesof.ai/2026/02/alibaba-releases-qwen3-5-397b-moe-multimodal-agentic-model)

> Alibaba Qwen released Qwen3.5-397B-A17B (Feb 16; blog Feb 15): open-weight natively multimodal MoE—397B total, 17B active—Apache 2.0. Vendor claims competitive vs GPT-5.2/Claude 4.5 Opus/Gemini-3 Pro families plus 8.6×–19× inference throughput vs prior gen; hosted Plus adds 1M context. Efficiency + open license as the China open-frontier trade.

Lunar New Year week stacked Chinese open and efficient releases. Alibaba’s proprietary close for Qwen3.5’s flagship: **frontier-ish multimodal agents under Apache 2.0 with extreme sparsity**—capability diffusion as the product, not only Elo screenshots.

**Alibaba’s Qwen team** released **Qwen3.5-397B-A17B** (announced February 16; blog timestamp 2026/02/15)—first model in the Qwen3.5 series. Open-weight, natively multimodal (vision-language via early fusion) Mixture-of-Experts: **397B** total parameters, **17B** active per token (high-sparsity design). Hybrid attention: Gated Delta Networks (linear attention) with sparse MoE. Pretraining: scaled visual-text tokens, STEM/reasoning, multilingual coverage (**201** languages/dialects; larger **250k** vocabulary). Post-training: broad RL across difficult environments for agentic tool use, planning, long-horizon tasks. Vendor benches: competitive standing vs GPT-5.2, Claude 4.5 Opus, Gemini-3 Pro families across language, VL, coding agent, search agent—with claimed inference throughput **8.6×–19×** vs prior generation at common context lengths. Hosted **Qwen3.5-Plus**: **1M** context and built-in adaptive tool use. Weights: Hugging Face, ModelScope, GitHub under **Apache 2.0**.

## Claims vs checks

Architecture, parameter counts, Apache 2.0, and weight locations are **Qwen primary**. Cross-lab “competitive/leading” and throughput multiples are **vendor** (Reuters relays performance/cost claims)—check Artificial Analysis / task boards before ranking freezes. Plus 1M is hosted SKU, not necessarily identical to open weights.

## Limits

- MoE serving quality depends on infrastructure; laptop demos ≠ 397B cluster reality.
- Multilingual breadth ≠ equal quality per language.
- Agentic benches are sensitive to tool scaffolds—compare like-with-like.

## Sources

- [Qwen blog: “Qwen3.5: Towards Native Multimodal Agents”](https://qwen.ai/blog?id=qwen3.5) (2026/02/15–16). Primary.
- Reuters: “Alibaba unveils new Qwen3.5 model for ‘agentic AI era’” (Feb 16, 2026).
- Hugging Face / GitHub references in the announcement for Qwen3.5-397B-A17B.
