Gemini 3.5 Transcribe brings emotion detection and speaker ID to speech-to-text

2 days ago 1



Google has quietly raised the bar for AI-powered transcription. Unveiled at Google I/O on May 19, 2026, and detailed further with a dedicated model card published on August 26, 2026, Gemini 3.5 Transcribe is the company’s most capable audio processing model to date, and it does considerably more than convert spoken words into text. The model supports timestamps in MM:SS format, speaker identification, translation, summarization, and emotion detection. What the model actually does At 96,000 tokens, Gemini 3.5 Transcribe can process extended audio sessions without losing track of what was said earlier in the recording. Users can upload common audio file formats, including MP3, directly through the Gemini API, Google AI Studio, or the Gemini macOS application. From there, the model can clean up filler words, identify individual speakers, translate content, or produce a summary, depending on what the user requests. Google does draw a line on one use case. For real-time transcription, the company points users toward its Cloud Speech-to-Text API and a separate Live API, rather than routing live audio through the general Gemini model. That distinction matters for developers building appli...

Read Entire Article