• Home
  • Popular Gadgets
  • Google Opens Access to Gemini 2.5 Native Audio Dialog and Controllable Speech Generation in Preview
Image

Google Opens Access to Gemini 2.5 Native Audio Dialog and Controllable Speech Generation in Preview

WhatsApp Group Join Now
Telegram Group Join Now
Instagram Group Join Now


Google introduced new audio generation capabilities with the Gemini 2.5 models at the Google I/O 2025. The Mountain View-based tech giant is now letting developers and individuals test these features on its platform. The two new capabilities include native audio dialog and controllable text-to-speech (TTS) with Gemini 2.5 Flash preview. While the former can natively generate human-like audio while responding to user prompts, the latter can convert any script into conversational speech. These features are currently not available to developers via application programming interfaces (APIs).

Google Showcases Gemini 2.5 Flash’s Audio Output Capabilities

In a blog post, the tech giant detailed the features of these two audio generation modes, highlighting how developers can use them to build new experiences for people. Currently, native audio dialog can be tried out in Google AI Studio’s stream tab, whereas the TTS feature can be tested in the generate media tab within AI Studio.

Native audio dialog with Gemini 2.5 Flash preview is designed for real-time conversations between a human user and the AI. The user can either type a prompt or speak it, and the AI responds verbally. This process directly generates audio, instead of first generating text and then converting it into speech.

There are several advantages to that as well. It supports affective dialog, which means when Gemini 2.5 Flash responds to the user’s tone of voice, it can recognise the emotion behind the said words. It can understand when the user sounds scared, angry, or surprised and respond accordingly.

Apart from this, the audio generation feature can express emotions when speaking, adopt different accents and linguistic styles, can access tools such as Google Search, and supports more than 24 languages.

Coming to the controllable TTS feature, it offers multi-speaker dialogue generation, can produce emotions and accents while narrating the script, control delivery speed and emphasise pronunciation, and supports the same 24 languages and language mixing.

Google says these capabilities were assessed for potential risks across the development process. The company used both internal mechanisms as well as red teaming to find and fix any vulnerabilities. The company also highlighted that all audio outputs from these models are embedded with SynthID, its watermarking technology.



Source link

Releated Posts

Animated Death Stranding movie gets its screenwriter

WhatsApp Group Join Now Telegram Group Join Now Instagram Group Join Now Hideo Kojima said in an interview…

ByByAjay jiJun 19, 2025

Waymo will start testing its autonomous cars in New York again

WhatsApp Group Join Now Telegram Group Join Now Instagram Group Join Now Waymo’s autonomous cars are heading back…

ByByAjay jiJun 19, 2025

Meta is finally adding passkey support for Facebook and Messenger

WhatsApp Group Join Now Telegram Group Join Now Instagram Group Join Now Meta is finally adding passkey support…

ByByAjay jiJun 18, 2025

Wyze adds major security update to its security cameras after numerous security lapses

WhatsApp Group Join Now Telegram Group Join Now Instagram Group Join Now Wyze, the Seattle-based tech company that…

ByByAjay jiJun 18, 2025

Leave a Reply

Your email address will not be published. Required fields are marked *

Scroll to Top