Upload a video, pick a language and voice, and get a naturally dubbed video back — background music preserved, timing matched to the original.
Whisper turns speech into word-level timestamps.
Demucs lifts the voice off the background music.
Claude translates to fit the original timing.
ElevenLabs or a local cloned voice speaks each line.
Lines that miss timing are rephrased and re-voiced.
Dubbed voice + music are muxed back onto the video.
Use your own ElevenLabs API key. You pay ElevenLabs directly; we just run the pipeline.
Cloned/stock voices generated on our GPU — no external TTS cost, so the cheapest way to dub.
Our ElevenLabs voices, billed per finished minute. Nothing to set up — just upload and go.