Trending Topics

Google’s Gemini in fine voice with new speech-to-speech translation and Native Audio refresh
Google is closing out the year with a handful of updates to voice services based on Gemini, its large language model. Alongside a refresh for Gemini 2.5 Flash Native Audio, it is adding new translation abilities to services such as Live Search and Translate.
This follows Google’s upgrade to its Gemini 2.5 Flash and Pro text-to-speech models, typically used by enterprises to generate voice content for marketing collateral, product information, or e-learning content.
Google said the updates will make the models’ speech more expressive, and more in keeping with users’ prompts on how the voice should sound. Speech will also now have better pacing, in line with prompt instructions, and improved abilities to create content with more than one speaker, keeping the voices consistent throughout conversations.
Google has used the updated text-to-speech models to refresh its Gemini 2.5 Flash Native Audio – used to enable users to interact via voice with Google services, as well as by enterprises to create customer services and other voice-drive agents.
The updated Gemini 2.5 Flash Native Audio can now better handle conversations requiring several rounds of back and forth between it and a user, and an improved understanding when to gather and include external content in its responses.
The upgraded Gemini 2.5 Flash Native Audio can be accessed through Google AI Studio and Vertex AI, and is also currently rolling out to Gemini Live and Search Live.
New speech-to-speech translation abilities
Gemini 2.5 Flash Native Audio is also gaining new translation capacities, including live speech-to-speech translation.
Speech-to-speech translation can either be used to translate a single voice, or multiple speakers, into one language, or allow simultaneous translation in a conversation between speakers of different languages. Translated speech will be available through the user’s headphones.
According to Google, speech-to-speech translation “captures the nuance of human speech, preserving the speaker’s intonation, pacing and pitch so the translation sounds natural”.
The live translation service can automatically pick up which language is being spoken and translate it into the target language, the company said. It also notes it can work in “loud outdoor environments” thanks to its ambient noise filtering. Great for all those business meetings.
The real-time translation for is available in beta through Google’s Translate app for Android devices in the USA, Mexico and India. Google said it plans to add availability for more countries, as well as iOS devices, from next year.
This audio AI news follows an update to rival OpenAI’s ChatGPT Images, which we covered earlier this week.
