# Microsoft’s MAI-Transcribe-2-Streaming Debuts #1 on Artificial Analysis Streaming STT

Times of AI Desk · 2026-10-01 · Models

[https://timesof.ai/2026/10/microsoft-mai-transcribe-2-streaming](https://timesof.ai/2026/10/microsoft-mai-transcribe-2-streaming)

> Microsoft AI launched MAI-Transcribe-2-Streaming — real-time speech-to-text in 60 languages with continuous language detection — claiming #1 on Artificial Analysis for both final and first-partial transcript accuracy, with first partials in just over 100ms of audio and introductory pricing of $0.54 per hour of audio through year-end. Same day: MAI-Voice-2.1 (23 languages / 26 locales, $22/1M chars) and MAI-Voice-2.1-Flash (~150ms E2E for 45s audio, $15/1M chars). AA ranking is as claimed by Microsoft; treat vendor latency/price figures as company-stated.

Microsoft’s in-house audio stack just claimed the independent streaming speech board — another signal Redmond is less dependent on OpenAI for the hear/speak loop.

**Microsoft AI** on **October 1, 2026** launched **MAI-Transcribe-2-Streaming**, a real-time transcription model covering **60 languages** with automatic continuous language detection. Microsoft says it ranks **no. 1 for accuracy for both final and partial transcripts on Artificial Analysis**, sits on the accuracy-versus-latency Pareto frontier in AA’s evaluation, and produces first “partials” in **just over 100ms** of receiving audio. An **Artificial Analysis** social post republished by **24/7 Wall St.** cites **2.5% WER** at **0.13s after end of speech** on AA-WER Streaming for the #1 final and first-partial spots — AA’s numbers via that card, not a desk re-score. Introductory price: **$0.54 per hour of audio through year-end**.

## Voice family in the same drop

Same announcement ships **MAI-Voice-2.1** — **23 languages / 26 locales**, cross-language same-speaker native accents, **$22 per 1M characters** — and **MAI-Voice-2.1-Flash** for high-volume latency-sensitive work: company claims **~150ms** end-to-end latency for **45 seconds** of audio, **55%** faster model inference and **~60%** cheaper than “comparable models,” priced at **$15 per 1M characters**. Both voice models support cloning with consent guardrails. Availability: **Microsoft Foundry**, **OpenRouter** (voice), **Vercel**, **Azure Voice Live**; **LiveKit** coming. A **Chatter** demo sits in the MAI Playground.

## Claims vs checks

The **#1 Artificial Analysis streaming STT** claim is **Microsoft’s** citation of AA. The **2.5% WER / 0.13s-after-EOS** figures are from **AA’s Oct 1 social post** as carried by **24/7 Wall St.**; this desk does not independently re-score the board. Prefer “Microsoft claims AA #1 for final and first-partial accuracy” over inventing a desk-owned Elo. Latency, “2x faster than closest competitor,” and price comparisons are **company-stated**.

## The capability trade

In-house Microsoft audio claiming an independent streaming leaderboard top spot — a capability/GTM wedge for voice agents and a concrete step in Microsoft’s path to own the conversational loop without routing every speech token through OpenAI.

## Limits

- Product specs, prices, language counts, and AA #1 claim are from **Microsoft AI’s October 1 post**.
- Artificial Analysis ranking not independently re-checked on the live board this run — status: **vendor-claimed present**; secondary AA quotes exist in digest wires.
- “2x faster,” “55% faster,” “~60% cheaper,” and Pareto-frontier language are **company evaluations / marketing**.
- No new frontier chat/coding Elo shift claimed here.

## Sources

- [Microsoft AI: Our first streaming transcription model debuts at no. 1 on Artificial Analysis (October 1, 2026)](https://microsoft.ai/news/our-first-streaming-transcription-model/)
- [24/7 Wall St. card quoting Artificial Analysis on MAI-Transcribe-2-Streaming (October 1, 2026)](https://247wallst.com/cards/msft-xpost-01m3w4xx60tgnzr6qd789h14vs)
