Gemini 3.1 Flash-Lite — Intelligence Priced for Volume
Google DeepMind Gemini 3.1 Flash-Lite preview (March 3–4): high-volume/low-latency multimodal at $0.25/$1.50 per 1M in/out via Gemini API and Vertex. Adjustable reasoning effort; early previews cite up to 30% inference cost savings — vendor. Cost tier for classification/translation/summarization-scale agents, not a frontier Elo play.

Frontier models win screenshots; Flash-Lite wins invoices. Google’s proprietary cut: ship a volume-priced multimodal that keeps adjustable reasoning — so agent/classification traffic does not need Pro economics.
Google DeepMind released Gemini 3.1 Flash-Lite in preview (March 3–4) via Gemini API and Vertex AI: optimized for high-volume, low-latency, cost-sensitive workloads (translation, classification, summarization, lightweight multimodal). Pricing: $0.25 per million input tokens, $1.50 per million output. Text/image and other modalities; adjustable reasoning effort for depth vs velocity. Early developer previews report up to 30% inference cost savings and strong speed-sensitive bench results (company/coverage). Positioned where full frontier capability is not required.
Claims vs checks
Pricing and preview availability are Google primary (The Keyword). Cost-savings % and “competitive intelligence” are vendor/early preview. Independent Artificial Analysis cost-quality curves should be checked before procurement locks.
Limits
- Preview ≠ GA SLA everywhere.
- Lite tiers trade headroom on hard reasoning.
- Multimodal “strong” is relative to the lite class, not Ultra.
Sources
- Google: “Gemini 3.1 Flash-Lite: Built for intelligence at scale” (March 3, 2026).
- Gemini API and Vertex AI documentation/changelog for the preview.
- Developer and industry coverage of March 2026 efficient model releases.