Google Unveils Gemini 3.5 Transcribe, AI-Powered Speech-to-Text Model for 85+ Languages

by priyanka.patel tech editor
Google Unveils Gemini 3.5 Transcribe, AI-Powered Speech-to-Text Model for 85+ Languages

Google launched Gemini 3.5 Transcribe, a speech-to-text model with 2.6% average word error rate across 85+ languages, offering real-time and pre-recorded transcription through two APIs. The release comes amid delays for its flagship 3.5 Pro model, which faces internal performance challenges.

Google’s Gemini 3.5 Transcribe, unveiled on August 26, 2026, marks a significant upgrade in speech-to-text technology, boasting a 2.6% average word error rate (WER) for pre-recorded audio and 4.0% for streaming, according to Artificial Analysis. The model, available via two distinct APIs, handles over 85 languages, including multilingual conversations with code-switching, and integrates smart disfluency cleanup to remove filler words and self-corrections.

Two APIs, Two Use Cases

Google AI Releases Gemini 3.5 Transcribe

Gemini 3.5 Transcribe ships as two separate products: gemini-3.5-transcribe for pre-recorded audio via the Interactions API, and gemini-3.5-transcribe-live for real-time streaming through the Live API. The latter emits interim results during speech, enabling applications like live captioning and voice agents. However, users must choose between smart transcription (cleaned text) or verbatim mode (timestamps and speaker labels), as the two cannot be combined. The split between the two endpoints is the part worth planning around. They do not share the same feature set, limits, or price.

Visual of a raw audio waveform transforming into clean formatted text, representing Gemini 3.5 Transcribe
Photo: Intelligentliving

The model’s accuracy improvements over Chirp 3 are notable: Google claims a 70% reduction in time-to-final transcription, with live-speech error rates dropping to 5.5% from 7.32% in Chirp 3. The system also handles self-corrections in natural speech, removes filler words, and supports speaker attribution and word-level timestamps for pre-recorded content. According to Google, automatic detection covers more than 85 languages, including mid-sentence code-switching.

Behind the headline numbers are two distinct API products, a smart-versus-verbatim trade-off that has not been widely discussed, hard session limits that will shape how teams deploy it, and a competitive picture in which Google is launching from a top-five position rather than the top of the leaderboard. Four-tenths of a second is how long Google says it takes its new Gemini 3.5 Transcribe Live model to turn the end of a spoken sentence into clean, formatted text.

Ecosystem Integration and Use Cases

Introducing Gemini 3.5 Transcribe

The model already powers the Gboard “Rambler” feature on the Pixel 11, and it is set to appear throughout the Google ecosystem. Developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. Google reports that across its products like the Gemini app and on Android, consumers have already benefited from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS.

Google Unveils Gemini 3.5 Transcribe, AI-Powered Speech-to-Text Model for 85+ Languages
Photo: Ars Technica

Managed-Service Model and Limitations

Despite its features, the model’s managed-service approach—no open weights or self-hosting—limits flexibility. Google’s decision to ship it as an API-only service reflects a managed-service strategy rather than an infrastructure choice. This excludes options for self-hosting or open weights, which may affect adoption by certain developer communities. However, startups and solo developers can use the free tier in Google AI Studio, while regulated enterprises access the Gemini Enterprise Agent Platform for compliance controls, according to reports.

Delays and Competitive Pressures

Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text

The Gemini 3.5 Transcribe rollout coincides with delays for its more powerful sibling, Gemini 3.5 Pro. According to a Bloomberg report, Google faces internal hurdles in meeting performance goals, causing frustration among engineers and risking a competitive edge against Anthropic and OpenAI. A Bloomberg report cited internal concerns that the delay could let rivals lose their edge, though Google maintains it is shipping quickly across a wide range of models. A Google spokesperson stated, We’re shipping quickly across a wide range of models while keeping them highly cost-effective for customers.

Gemini API for Speech and Text
Google Unveils Gemini 3.5 Transcribe, AI-Powered Speech-to-Text Model for 85+ Languages
Photo: blog.google

Analysts note the timing is critical. While Gemini 3.5 Flash-Lite and 3.6 Flash offer cost savings, the absence of 3.5 Pro—a model expected to outperform competitors—leaves gaps in Google’s AI portfolio. Google’s focus on efficiency is smart, but without a flagship model, it risks ceding ground to faster-moving rivals, according to a source.

What’s Next for Google’s AI Roadmap?

Google’s announcement underscores its push to expand AI across workflows, from real-time transcription to agent-based tasks. Blog.google states the model is already powering features like Rambler and will soon integrate with Chrome. However, the company’s broader AI strategy remains in flux: CNET reports that Gemini 3.5 Pro testing with partners is ongoing, with a “land soon” timeline from product lead Logan Kilpatrick. Google’s ability to deliver on its promises—especially for 3.5 Pro—will shape its position in the AI race.

You may also like