Thursday, Oct 8 | --:--
Back to home

Gemini 3.1 Flash-Lite — Intelligence Priced for Volume

Google DeepMind Gemini 3.1 Flash-Lite preview (March 3–4): high-volume/low-latency multimodal at $0.25/$1.50 per 1M in/out via Gemini API and Vertex. Adjustable reasoning effort; early previews cite up to 30% inference cost savings — vendor. Cost tier for classification/translation/summarization-scale agents, not a frontier Elo play.

Times of AI Desk 3 min read Mountain View View as Markdown
Cover illustration for Gemini 3.1 Flash-Lite — Intelligence Priced for Volume

Frontier models win screenshots; Flash-Lite wins invoices. Google’s proprietary cut: ship a volume-priced multimodal that keeps adjustable reasoning — so agent/classification traffic does not need Pro economics.

Google DeepMind released Gemini 3.1 Flash-Lite in preview (March 3–4) via Gemini API and Vertex AI: optimized for high-volume, low-latency, cost-sensitive workloads (translation, classification, summarization, lightweight multimodal). Pricing: $0.25 per million input tokens, $1.50 per million output. Text/image and other modalities; adjustable reasoning effort for depth vs velocity. Early developer previews report up to 30% inference cost savings and strong speed-sensitive bench results (company/coverage). Positioned where full frontier capability is not required.

Claims vs checks

Pricing and preview availability are Google primary (The Keyword). Cost-savings % and “competitive intelligence” are vendor/early preview. Independent Artificial Analysis cost-quality curves should be checked before procurement locks.

Limits

  • Preview ≠ GA SLA everywhere.
  • Lite tiers trade headroom on hard reasoning.
  • Multimodal “strong” is relative to the lite class, not Ultra.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading