Thursday, Oct 8 | --:--
Back to home
Models Open Source China

Qwen3.5-397B-A17B — Open Multimodal MoE for Agents, Not Raw Scale Alone

Alibaba Qwen released Qwen3.5-397B-A17B (Feb 16; blog Feb 15): open-weight natively multimodal MoE—397B total, 17B active—Apache 2.0. Vendor claims competitive vs GPT-5.2/Claude 4.5 Opus/Gemini-3 Pro families plus 8.6×–19× inference throughput vs prior gen; hosted Plus adds 1M context. Efficiency + open license as the China open-frontier trade.

Times of AI Desk 5 min read Beijing View as Markdown
Cover illustration for Qwen3.5-397B-A17B — Open Multimodal MoE for Agents, Not Raw Scale Alone

Lunar New Year week stacked Chinese open and efficient releases. Alibaba’s proprietary close for Qwen3.5’s flagship: frontier-ish multimodal agents under Apache 2.0 with extreme sparsity—capability diffusion as the product, not only Elo screenshots.

Alibaba’s Qwen team released Qwen3.5-397B-A17B (announced February 16; blog timestamp 2026/02/15)—first model in the Qwen3.5 series. Open-weight, natively multimodal (vision-language via early fusion) Mixture-of-Experts: 397B total parameters, 17B active per token (high-sparsity design). Hybrid attention: Gated Delta Networks (linear attention) with sparse MoE. Pretraining: scaled visual-text tokens, STEM/reasoning, multilingual coverage (201 languages/dialects; larger 250k vocabulary). Post-training: broad RL across difficult environments for agentic tool use, planning, long-horizon tasks. Vendor benches: competitive standing vs GPT-5.2, Claude 4.5 Opus, Gemini-3 Pro families across language, VL, coding agent, search agent—with claimed inference throughput 8.6×–19× vs prior generation at common context lengths. Hosted Qwen3.5-Plus: 1M context and built-in adaptive tool use. Weights: Hugging Face, ModelScope, GitHub under Apache 2.0.

Claims vs checks

Architecture, parameter counts, Apache 2.0, and weight locations are Qwen primary. Cross-lab “competitive/leading” and throughput multiples are vendor (Reuters relays performance/cost claims)—check Artificial Analysis / task boards before ranking freezes. Plus 1M is hosted SKU, not necessarily identical to open weights.

Limits

  • MoE serving quality depends on infrastructure; laptop demos ≠ 397B cluster reality.
  • Multilingual breadth ≠ equal quality per language.
  • Agentic benches are sensitive to tool scaffolds—compare like-with-like.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading