OpenAI says it can clone a voice from just 15 seconds of audio

OpenAI just announced that it recently conducted a small-scale preview of a new tool called Voice Engine. This is a voice cloning technology that can mimic any speaker by analyzing a 15-second audio sample. The company says it generates “natural-sounding speech” with “emotive and realistic voices.”

The technology is based on the company’s pre-existing text-to-speech API and it has been in the works since 2022. OpenAI has already been using a version of the toolset to power the preset voices available in the current text-to-speech API and the Read Aloud feature. There are a bunch of samples on the company’s official blog and they sound eerily close to the real thing. I encourage you to give them a listen and imagine the possibilities, both good and bad.

OpenAI says they see this technology being useful for reading assistance, language translation and helping those who suffer from sudden or degenerative speech conditions. The company brought up a Brown University pilot program that helped a patient with speech impairment issues by creating a Voice Engine clone pulled from audio recorded for a school project.

Despite the potential benefits, bad actors would certainly abuse this technology to engage in some serious deepfake tomfoolery, which is already a problem. With this in mind, Voice Engine isn’t quite ready for prime time, as there are serious privacy concerns that must be met before a full rollout.

OpenAI acknowledges that this tech has “serious risks, which are especially top of mind in an election year.” The company says its incorporating feedback from “US and international partners from across government, media, entertainment, education, civil society and beyond” to ensure the product launches with a minimal amount of risk. All preview testers agreed to OpenAI’s usage policies, which ban the impersonation of another individual without consent or legal right.

Additionally, anybody using the tech will have to disclose to their audience that the voices are AI-generated. OpenAI implemented safety measures, like watermarking to trace the origin of any audio and “proactive monitoring” of how the system is being used. When the product officially rolls out there will be a “no-go voice list” that detects and prevents AI-generated speakers that are too similar to prominent figures.

As for when that rollout will occur, OpenAI remains tight-lipped. TechCrunch uncovered some potential pricing data and it looks like it will undercut competitors in the space like ElevenLabs. Voice Engine could cost $15 per one million characters, which works out to around 162,500 words. This is about the length of Stephen King’s The Shining. It certainly sounds like a budget-friendly way to get an audiobook done. The marketing materials also make reference to an “HD” version that costs twice as much, but the company hasn’t detailed how that will work.

OpenAI has been making big moves this week. It just announced another partnership with its bestie Microsoft to build an AI-based supercomputer called “Stargate.” The project will reportedly cost a whopping $100 billion, according to The Information.

This article originally appeared on Engadget at https://www.engadget.com/openai-says-it-can-clone-a-voice-from-just-15-seconds-of-audio-190356431.html?src=rss https://www.engadget.com/openai-says-it-can-clone-a-voice-from-just-15-seconds-of-audio-190356431.html?src=rss
Erstellt 1y | 29.03.2024, 20:10:24


Melden Sie sich an, um einen Kommentar hinzuzufügen

Andere Beiträge in dieser Gruppe

Match Group will pay $14 million to settle claims of deceptive business practices

The Federal Trade Commission announced that Match Group will pay

13.08.2025, 00:20:14 | Engadget
Russia reportedly implicated in hack on US federal courts' databases

Databases used by US federal courts for sharing and managing case documents have been hacked.

12.08.2025, 21:50:14 | Engadget
Blizzard's Story and Franchise Development team has voted to unionize

Workers from Blizzard Entertainment's department for Story and Franchise Development have

12.08.2025, 21:50:13 | Engadget
Alien: Earth succeeds where Ridley Scott's Alien sequels failed

Alien: Earth delivers everything you'd want from a series with "Alien" in the title: The iconic Xenomorphs hunting down hapless humans; gratuitous body horror; and androids who you can nev

12.08.2025, 19:30:26 | Engadget
The Samsung Odyssey OLED G6 is the world's first 500Hz OLED gaming monitor

Previously, if you wanted a monitor for competitive gaming, you had to choose between an IPS or VA panel to get something with a super high refresh rate or opt for a slower OLED display with richer

12.08.2025, 19:30:24 | Engadget
Threads is up to 400 million monthly active users

Meta's X competitor, Threads, is continuing to add users at a brisk clip, with the social network now surpassing 400 million monthly active users. The news, reported by

12.08.2025, 19:30:23 | Engadget