Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text

4 hours ago 7

While we wait (possibly in vain) for Gemini 3.5 Pro to launch, Google is releasing a different model in the 3.5 branch. The company has announced Gemini 3.5 Transcribe, an AI model designed to streamline voice input by editing out “ums” and corrections, outputting polished AI text. This model already powers the Gboard “Rambler” feature on the Pixel 11, but it’s about to appear throughout the Google ecosystem.

According to Google, Gemini 3.5 Transcribe is much faster and more accurate than its previous voice-to-text engine, known as Chirp 3. The new AI model should be about 70 percent faster from voice to final transcribed text, and the live-speech error rate has dropped to 5.5 percent. That’s only a little better than Chirp 3, which Google measures at 7.32 percent. Still, it’s a pain to fix typos when you’re using voice input, so any improvement here is beneficial.

Credit: Google

Credit: Google

The new model isn’t just better at hearing the words you say; it’s also supposedly better at getting to the heart of what you meant to say. As you speak, Gemini 3.5 Transcribe can remove the awkward “ums” and “uhs” that clutter your stream of consciousness. It can also edit text on the fly if you need to correct yourself and refer to your provided custom vocabulary to handle “specialized jargon.” All this works in 85 languages and with up to three speakers in pre-recorded audio.

The drawback, of course, is that you’re relying on the AI to accurately get the gist of your speech. For short blocks of text, the model seems good at cleaning up inconsistencies and verbal stumbles (based on my testing with Rambler), but the AI does technically change the wording of what you said, and that may not be appropriate for all situations.

Read Entire Article