# Gemini 3.1 Flash-Lite — Intelligence Priced for Volume

Times of AI Desk · 2026-03-04 · Models

[https://timesof.ai/2026/03/google-gemini-3-1-flash-lite-cost-efficient-ai](https://timesof.ai/2026/03/google-gemini-3-1-flash-lite-cost-efficient-ai)

> Google DeepMind Gemini 3.1 Flash-Lite preview (March 3–4): high-volume/low-latency multimodal at $0.25/$1.50 per 1M in/out via Gemini API and Vertex. Adjustable reasoning effort; early previews cite up to 30% inference cost savings — vendor. Cost tier for classification/translation/summarization-scale agents, not a frontier Elo play.

Frontier models win screenshots; **Flash-Lite** wins invoices. Google’s proprietary cut: ship a **volume-priced multimodal** that keeps adjustable reasoning — so agent/classification traffic does not need Pro economics.

**Google DeepMind** released **Gemini 3.1 Flash-Lite** in preview (March 3–4) via Gemini API and Vertex AI: optimized for high-volume, low-latency, cost-sensitive workloads (translation, classification, summarization, lightweight multimodal). Pricing: **$0.25** per million input tokens, **$1.50** per million output. Text/image and other modalities; adjustable reasoning effort for depth vs velocity. Early developer previews report up to **30%** inference cost savings and strong speed-sensitive bench results (company/coverage). Positioned where full frontier capability is not required.

## Claims vs checks

Pricing and preview availability are **Google primary** (The Keyword). Cost-savings % and “competitive intelligence” are **vendor/early preview**. Independent Artificial Analysis cost-quality curves should be checked before procurement locks.

## Limits

- Preview ≠ GA SLA everywhere.
- Lite tiers trade headroom on hard reasoning.
- Multimodal “strong” is relative to the lite class, not Ultra.

## Sources

- [Google: “Gemini 3.1 Flash-Lite: Built for intelligence at scale”](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-lite) (March 3, 2026).
- Gemini API and Vertex AI documentation/changelog for the preview.
- Developer and industry coverage of March 2026 efficient model releases.
