Section
Models
Model releases, benchmarks, capability jumps, and lab research milestones.
Related Coverage
- DeepSeek Ships V4-Flash-Vision-Exp, a Multimodal API Sibling of V4-Flash
On August 21, 2026, DeepSeek’s API changelog launched experimental DeepSeek-V4-Flash-Vision-Exp (`deepseek-v4-flash-vision-exp`): JPEG/PNG/GIF/WebP via base64, URL, or Files API file_id, on Chat Completions, Anthropic-compatible Messages, and Responses. Text-agent scores match official V4-Flash; DeepSeek says multimodal-agent results jump toward Claude Opus 4.8. Distinct from the August 13 V4-Pro GA and price hike.
- Gemma Passes One Billion Downloads; Google Opens Awesome Gemma
On August 20, 2026, Google DeepMind said the Gemma open-model family has surpassed one billion downloads, with developers publishing more than 100,000 Gemmaverse variants in two years. The same post launched the Awesome Gemma GitHub directory. Distinct from the Gemini app’s 1 billion monthly users on August 11.
- Alibaba Releases Qwen3.8-27B Open Weights Under Apache 2.0
On August 14, 2026, Alibaba’s Qwen team posted Qwen3.8-27B on Hugging Face—a 27-billion-parameter native vision-language dense model with 262K context (extensible to 1M), thinking mode by default, and Apache 2.0 weights—the local companion to July’s 2.4T Qwen3.8-Max preview.
- Google Ships Gemini 3.7 Flash, a Workhorse Model at Half the Old Price
On August 13, 2026, Google released Gemini 3.7 Flash—23 days after 3.6 Flash—claiming large coding and agent gains and an introductory $0.75/$3.75 per million tokens through year-end. Independent Artificial Analysis scores 3.7 Flash (high) at 56 on the Intelligence Index, up 4 points from 3.6 Flash.
- OpenAI Previews Ultrafast: GPT-5.6 Sol at Up to 14× Standard Speed
On August 13, 2026, OpenAI previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing—up to 750 output tokens per second on Cerebras hardware—in a limited customer preview for incident response, voice, finance, and other latency-sensitive work.
- SpaceXAI Releases Grok 4.6, Matching GPT-5.6 Sol on Artificial Analysis
On August 12, 2026, SpaceXAI released Grok 4.6—a post-training upgrade of Grok 4.5 aimed at long-running agents, coding, and interactive visual work—that ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index, ships at $2/$6 per million tokens, and is live in Cursor, Grok Build, and the SpaceXAI API.
- NVIDIA Ships Nemotron 3.5 Lightning and NeMo Switchyard
On August 11, 2026, NVIDIA released Nemotron 3.5 Lightning—a 30B open MoE model with ~3B active parameters for high-volume agent execution—and NeMo Switchyard, an open-source model router that sends each agent step to the cheapest capable model, claiming up to 4× throughput in class and large cost cuts when paired with frontier planners.
- OpenAI Expands Daybreak with GPT‑5.6‑Cyber for Trusted Defenders
On August 10, 2026, OpenAI expanded its Daybreak cyber program into Daybreak Blue and Daybreak Red access tiers and introduced GPT‑5.6‑Cyber—a purpose-trained cybersecurity model that completes 95% of advanced dual-use cyber requests in internal tests versus 1.5% for GPT‑5.6 Sol—after using the model to find Chrome V8 zero-days patched as CVE-2026-15903.
- Meta Open-Sources Muse Glimmer, a 30B Local Agentic Model
On August 10, 2026, Meta Superintelligence Labs released Muse Glimmer—a 30-billion-parameter open-weight model under Apache 2.0 optimized for always-on local agents on a single consumer GPU—alongside CEO Mark Zuckerberg’s essay championing U.S. open-weight AI against Chinese rivals and closed-model concentration.
- OpenAI Retunes GPT-5.6 Sol in ChatGPT and Makes Luna Free Default
On August 6, 2026, OpenAI rolled out a ChatGPT-focused GPT-5.6 Sol update with a thinking-effort slider for Plus and Pro, said internal evals cut factual-error rates ~68% vs GPT-5.5 Instant for Sol and ~62% for Luna, and made GPT-5.6 Luna the Free/Go default with unlimited text chats and a Think button on the way.
- Meta Launches Muse Code Beta and Muse Spark 1.2 Coding Model
On August 5, 2026, Meta released Muse Code (beta), a terminal coding agent with persistent background subagents and a replay-exact event log, powered by Muse Spark 1.2—a coding-focused update to Muse Spark 1.1 available in the agent and Meta Model API with expanded global access.
- SpaceXAI Adds Multi-Reference and 1080p to Imagine Video 1.5
On July 31, 2026, SpaceXAI upgraded Imagine Video 1.5 with text-, image-, and voice-reference conditioning (up to seven image references), native 1080p generation, and text-to-video—rolling out first to SuperGrok Heavy and Plus in the US on grok.com/imagine and iOS, with API support for image refs and 1080p on grok-imagine-video-1.5.
- Google DeepMind Launches Gemini Robotics 2 Family
On July 30, 2026, Google DeepMind launched the Gemini Robotics 2 family—whole-body robot control (Gemini Robotics 2), high-level embodied reasoning and multi-robot collaboration (Gemini Robotics ER 2), and an on-device adaptation tier—extending Gemini from tabletop manipulation to full humanoid locomotion and multi-robot teamwork.
- Thinking Machines Releases Inkling-Small Open Weights
On July 30, 2026, Thinking Machines Lab released Inkling-Small—a 276B-total / 12B-active open-weight MoE that beats full Inkling (975B/41B) on SWE-Bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%) at roughly one-quarter the size and $0.30/$1.20 per million tokens.
- OpenAI Cuts GPT-5.6 Luna Prices 80% and Terra 20%
On July 30, 2026, OpenAI cut API prices for GPT-5.6 Luna by 80% (to $0.20/$1.20 per million tokens) and Terra by 20% (to $2/$12), held Sol pricing steady, and added a Fast mode for Sol up to 2.5× speed at 2× standard price—crediting efficiency gains from Sol on its own serving stack.
- Google Launches Lyria 3.5 Music Model in Flow Music
On July 29, 2026, Google rolled out Lyria 3.5 inside Google Flow Music—its newest music generation model with claimed gains in musicality, lyrics, vocal expression, and direct tempo/duration control for full-song creation.
- SpaceXAI Ships Grok Voice Think Fast 2.0 With Top Speech-to-Speech Scores
On July 29, 2026, SpaceXAI released Grok Voice Think Fast 2.0—its next-generation speech-to-speech voice model—claiming 82.9% on Artificial Analysis’s Speech-to-Speech Quality Index, 0.70s time-to-first-audio, $0.08/min pricing, and a default swap for grok-voice-latest on August 5.
- Grok 4.5 Lands in GitHub Copilot
On July 28, 2026, SpaceXAI announced that Grok 4.5—its flagship coding model—is available in GitHub Copilot across cloud agents, Copilot CLI, and VS Code, selectable from the model picker (enterprise enablement required for some orgs), with API pricing at $2/$6 per million input/output tokens.
- Anthropic Launches Claude Opus 5: Near-Fable Intelligence at Half the Price
On July 24, 2026, Anthropic released Claude Opus 5—positioned as a step-change for the Opus tier that approaches Claude Fable 5 on coding and knowledge-work evals at half the cost—now the default on Claude Max and the strongest model on Claude Pro, at $5/$25 per million input/output tokens.
- Google Ships Gemini 3.6 Flash and 3.5 Flash-Lite; Cyber Tier Stays Limited
On July 21, 2026, Google launched Gemini 3.6 Flash ($1.50/$7.50 per 1M tokens) and 3.5 Flash-Lite ($0.30/$2.50) for agents at scale, while Gemini 3.5 Flash Cyber in CodeMender stays limited to governments and trusted partners; 3.5 Pro remains partner-testing only.
- Alibaba Previews Qwen3.8-Max, a 2.4T Model It Says Is Second Only to Fable 5
On July 19, 2026, Alibaba’s Qwen team launched Qwen3.8-Max-Preview—a 2.4-trillion-parameter multimodal flagship live on Token Plan, Qoder, and QoderWork—claiming frontier-class capability second only to Claude Fable 5, with open weights promised soon but no public benchmark table yet.
- Kimi K3 Takes #1 on Arena Frontend Code, Ranking Above Claude Fable 5
On July 16, 2026—the same day Moonshot AI launched Kimi K3—Arena.ai unblinded the model at #1 on the Frontend Code Arena with 1679 Elo, ahead of Claude Fable 5 and GPT-5.6 Sol in blind human preference tests, a 17-place jump from Kimi K2.6 and first place in six of seven frontend domains.
- Google Delays Gemini 3.5 Pro as Coding Performance Misses Internal Goals
On July 16, 2026, Reuters reported—citing Bloomberg—that Alphabet’s Google is months behind schedule on Gemini 3.5 Pro, its flagship model, as the company works to improve capabilities especially in coding; CEO Sundar Pichai had pointed to a June release at Google I/O 2026.
- Moonshot AI Unveils Kimi K3, an Open 2.8T-Parameter Frontier Model
On July 16, 2026, Beijing’s Moonshot AI launched Kimi K3—a 2.8-trillion-parameter open-class multimodal model with a 1M-token context window, API live the same day, and full weights planned for July 27—claiming frontier coding and agent scores just behind Claude Fable 5 and GPT-5.6 Sol at far lower token prices.