OpenAI says it can clone a voice from just 15 seconds of audio

OpenAI just announced that it recently conducted a small-scale preview of a new tool called Voice Engine. This is a voice cloning technology that can mimic any speaker by analyzing a 15-second audio sample. The company says it generates “natural-sounding speech” with “emotive and realistic voices.”

The technology is based on the company’s pre-existing text-to-speech API and it has been in the works since 2022. OpenAI has already been using a version of the toolset to power the preset voices available in the current text-to-speech API and the Read Aloud feature. There are a bunch of samples on the company’s official blog and they sound eerily close to the real thing. I encourage you to give them a listen and imagine the possibilities, both good and bad.

OpenAI says they see this technology being useful for reading assistance, language translation and helping those who suffer from sudden or degenerative speech conditions. The company brought up a Brown University pilot program that helped a patient with speech impairment issues by creating a Voice Engine clone pulled from audio recorded for a school project.

Despite the potential benefits, bad actors would certainly abuse this technology to engage in some serious deepfake tomfoolery, which is already a problem. With this in mind, Voice Engine isn’t quite ready for prime time, as there are serious privacy concerns that must be met before a full rollout.

OpenAI acknowledges that this tech has “serious risks, which are especially top of mind in an election year.” The company says its incorporating feedback from “US and international partners from across government, media, entertainment, education, civil society and beyond” to ensure the product launches with a minimal amount of risk. All preview testers agreed to OpenAI’s usage policies, which ban the impersonation of another individual without consent or legal right.

Additionally, anybody using the tech will have to disclose to their audience that the voices are AI-generated. OpenAI implemented safety measures, like watermarking to trace the origin of any audio and “proactive monitoring” of how the system is being used. When the product officially rolls out there will be a “no-go voice list” that detects and prevents AI-generated speakers that are too similar to prominent figures.

As for when that rollout will occur, OpenAI remains tight-lipped. TechCrunch uncovered some potential pricing data and it looks like it will undercut competitors in the space like ElevenLabs. Voice Engine could cost $15 per one million characters, which works out to around 162,500 words. This is about the length of Stephen King’s The Shining. It certainly sounds like a budget-friendly way to get an audiobook done. The marketing materials also make reference to an “HD” version that costs twice as much, but the company hasn’t detailed how that will work.

OpenAI has been making big moves this week. It just announced another partnership with its bestie Microsoft to build an AI-based supercomputer called “Stargate.” The project will reportedly cost a whopping $100 billion, according to The Information.

This article originally appeared on Engadget at https://www.engadget.com/openai-says-it-can-clone-a-voice-from-just-15-seconds-of-audio-190356431.html?src=rss https://www.engadget.com/openai-says-it-can-clone-a-voice-from-just-15-seconds-of-audio-190356431.html?src=rss
Creato 1y | 29 mar 2024, 20:10:24


Accedi per aggiungere un commento

Altri post in questo gruppo

AOL's dial-up internet still exists, but not for much longer

It may have been decades since you last heard the ">crunc

10 ago 2025, 21:20:08 | Engadget
Rod Fergusson leaves Blizzard after five years leading Diablo

Rod Fergusson, the general manager of the Diablo franchise for the last five years, is leaving Blizzard. Fergusson announced the move on social media, but didn't say where he's going next. Before

10 ago 2025, 18:50:18 | Engadget
An updated Siri that interacts with apps reportedly won't be here until next spring

A Siri that does way more than just setting a timer or writing down a reminder may still be nearly a year away. According to

10 ago 2025, 18:50:17 | Engadget
Ubisoft may have prematurely revealed FX's TV adaptation of Far Cry

A post on Ubisoft's news page reportedly announced that FX is working on a TV show adaptation of the

10 ago 2025, 16:40:04 | Engadget
The Space Invaders movie is apparently still happening

It's been a few years since we last heard anything about

9 ago 2025, 22:10:12 | Engadget
DJI repurposed its drones' obstacle detection tech for robot vacuums

DJI's obstacle avoidance system could be just as useful on land as it is in the air. DJI, known for its dominance in the

9 ago 2025, 19:40:14 | Engadget
Apple's MacBook Air M4 is on sale for up to 20 percent off

Whether you need a new MacBook for the upcoming semester or you've just be

9 ago 2025, 17:30:25 | Engadget