OpenAI Can Re-Create Human Voices from 15 Second Samples

March 30th, 2024

Fifteen seconds? *pfft*

Google AI Clones Your Voice After Listening for 5 Seconds

Microsoft’s New AI Can Simulate Anyone’s Voice with 3 Seconds of Audio

Via: Ars Technica:

Voice synthesis has come a long way since 1978’s Speak & Spell toy, which once wowed people with its state-of-the-art ability to read words aloud using an electronic voice. Now, using deep-learning AI models, software can create not only realistic-sounding voices, but also convincingly imitate existing voices using small samples of audio.

Further Reading
Microsoft’s new AI can simulate anyone’s voice with 3 seconds of audio

Along those lines, OpenAI just announced Voice Engine, a text-to-speech AI model for creating synthetic voices based on a 15-second segment of recorded audio. It has provided audio samples of the Voice Engine in action on its website.

Once a voice is cloned, a user can input text into the Voice Engine and get an AI-generated voice result. But OpenAI is not ready to widely release its technology yet. The company initially planned to launch a pilot program for developers to sign up for the Voice Engine API earlier this month. But after more consideration about ethical implications, the company decided to scale back its ambitions for now.

Open AI: Navigating the Challenges and Opportunities of Synthetic Voices

Leave a Reply

You must be logged in to post a comment.