Friday, Oct 2 | --:--
Back to home

Microsoft’s MAI-Transcribe-2-Streaming Debuts #1 on Artificial Analysis Streaming STT

Microsoft AI launched MAI-Transcribe-2-Streaming — real-time speech-to-text in 60 languages with continuous language detection — claiming #1 on Artificial Analysis for both final and first-partial transcript accuracy, with first partials in just over 100ms of audio and introductory pricing of $0.54 per hour of audio through year-end. Same day: MAI-Voice-2.1 (23 languages / 26 locales, $22/1M chars) and MAI-Voice-2.1-Flash (~150ms E2E for 45s audio, $15/1M chars). AA ranking is as claimed by Microsoft; treat vendor latency/price figures as company-stated.

Times of AI Desk 5 min read Redmond, WA View as Markdown
Cover illustration for Microsoft’s MAI-Transcribe-2-Streaming Debuts #1 on Artificial Analysis Streaming STT

Microsoft’s in-house audio stack just claimed the independent streaming speech board — another signal Redmond is less dependent on OpenAI for the hear/speak loop.

Microsoft AI on October 1, 2026 launched MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with automatic continuous language detection. Microsoft says it ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis, sits on the accuracy-versus-latency Pareto frontier in AA’s evaluation, and produces first “partials” in just over 100ms of receiving audio. An Artificial Analysis social post republished by 24/7 Wall St. cites 2.5% WER at 0.13s after end of speech on AA-WER Streaming for the #1 final and first-partial spots — AA’s numbers via that card, not a desk re-score. Introductory price: $0.54 per hour of audio through year-end.

Voice family in the same drop

Same announcement ships MAI-Voice-2.1 — 23 languages / 26 locales, cross-language same-speaker native accents, $22 per 1M characters — and MAI-Voice-2.1-Flash for high-volume latency-sensitive work: company claims ~150ms end-to-end latency for 45 seconds of audio, 55% faster model inference and ~60% cheaper than “comparable models,” priced at $15 per 1M characters. Both voice models support cloning with consent guardrails. Availability: Microsoft Foundry, OpenRouter (voice), Vercel, Azure Voice Live; LiveKit coming. A Chatter demo sits in the MAI Playground.

Claims vs checks

The #1 Artificial Analysis streaming STT claim is Microsoft’s citation of AA. The 2.5% WER / 0.13s-after-EOS figures are from AA’s Oct 1 social post as carried by 24/7 Wall St.; this desk does not independently re-score the board. Prefer “Microsoft claims AA #1 for final and first-partial accuracy” over inventing a desk-owned Elo. Latency, “2x faster than closest competitor,” and price comparisons are company-stated.

The capability trade

In-house Microsoft audio claiming an independent streaming leaderboard top spot — a capability/GTM wedge for voice agents and a concrete step in Microsoft’s path to own the conversational loop without routing every speech token through OpenAI.

Limits

  • Product specs, prices, language counts, and AA #1 claim are from Microsoft AI’s October 1 post.
  • Artificial Analysis ranking not independently re-checked on the live board this run — status: vendor-claimed present; secondary AA quotes exist in digest wires.
  • “2x faster,” “55% faster,” “~60% cheaper,” and Pareto-frontier language are company evaluations / marketing.
  • No new frontier chat/coding Elo shift claimed here.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading