TTS Dubbing Services
Dubbing (TTS) is the third step of video translation, converting translated subtitle text into speech audio. pyVideoTrans supports 30+ dubbing services.
You can also use the left panel "Batch Subtitle Dubbing" feature independently — import multiple SRT or TXT files for dubbing.
The
clonevoice role uses the original video speaker's voice for dubbing (voice cloning). This role is only available in the main "Video Translation" feature.All services with a
clonerole support custom reference audio — clone the voice from your own 3-10s audio clip. See "Creating and Using Reference Audio" below.
Out-of-the-box (Free)
No complex configuration needed — great for beginners.
| Service | Description | Rating |
|---|---|---|
| Edge-TTS (Free) | Microsoft free API, natural voice, all languages | ⭐⭐⭐ Default |
| gTTS (Free) | Google TTS, basic quality, VPN needed in China | ⭐⭐ |
⚠️ Heavy Edge-TTS usage may trigger rate limits. Set concurrency to 1 and pause to 5-10 seconds in advanced options.
Built-in (Free)
Models auto-download on first use.
| Service | Description | GPU | Clone | Rating |
|---|---|---|---|---|
| Qwen3-TTS (built-in) | Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian | ✅ | ✅ | ⭐⭐⭐ |
| F5-TTS (built-in) | Chinese, English, Japanese, French, German, Russian, Italian, Spanish, Hindi, Arabic | ✅ | ✅ | ⭐⭐⭐ |
| OmniVoice-TTS | 600 languages (built-in since v4.05) | ✅ | ✅ | ⭐⭐⭐ |
| Confucius-TTS (built-in, v4.06+) | Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, Vietnamese | ✅ | ✅ | ⭐⭐⭐ |
| MOSS-TTS-Nano (built-in) | Chinese, English, German, Spanish, French, Japanese, Italian, Hungarian, Korean, Russian, Persian, Arabic, Polish, Portuguese, Czech, Swedish, Greek, Turkish | ❌ | ✅ | ⭐⭐ |
| ZipVoice (built-in) | Chinese, English | ✅ | ✅ | ⭐⭐⭐ |
| Piper (built-in) | Lightweight, 20 languages | ❌ | ❌ | ⭐⭐ |
| ChatterBox (built-in) | Arabic, German, Greek, English, Spanish, Finnish, French, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Dutch, Norwegian, Polish, Portuguese, Russian, Swedish, Swahili, Turkish, Chinese | ✅ | ✅ | ⭐⭐⭐ |
| Supertonic3 (built-in) | English, Korean, Japanese, Arabic, Czech, German, Greek, Spanish, French, Hindi, Hungarian, Indonesian, Italian, Dutch, Polish, Portuguese, Romanian, Russian, Swedish, Turkish, Ukrainian, Vietnamese | ❌ | ❌ | ⭐⭐ |
| VITS (built-in) | Chinese, English | ❌ | ❌ | ⭐⭐ |
Due to network conditions in China, models are large and auto-download may fail. If it fails, see model download URLs and manual download methods
Self-hosted (Advanced)
| Service | Description | Clone | Rating |
|---|---|---|---|
| GPT-SoVITS | Chinese, English, Japanese, Korean, Cantonese | ✅ | ⭐⭐⭐ |
| Index-TTS | Chinese, English | ✅ | ⭐⭐⭐ |
| VoxCPM-TTS | Arabic, Burmese, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Tagalog, Thai, Turkish, Vietnamese | ✅ | ⭐⭐⭐ |
| CosyVoice | Chinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian | ✅ | ⭐⭐ |
| ChatTTS | Chinese, English | — | ⭐⭐ |
| Fish-TTS | All built-in languages | — | ⭐ |
| Kokoro-TTS | Chinese, English, Korean, Italian, Portuguese, German, French, Hindi | — | ⭐ |
| Spark-TTS | Chinese, English | ✅ | ⭐⭐ |
| clone-voice | No longer maintained | ✅ | ⭐ |
Creating and Using Reference Audio
TTS services that support
clone(where the voice role includesclone) can all use reference audio. Select the reference audio in the voice role dropdown, and the software will use that audio's voice for dubbing. Reference audio is configured inMenu → TTS Settings → Set Reference Audio.
Steps
- Record or extract a 3-10 second audio clip, save as WAV, ensure clear pronunciation with no background noise.
- Open
Menu → TTS Settings → Set Reference Audio. - Enter content in the following format:
audio_filename.wav#text spoken in the audio1 - Place the reference audio file in
app_dir/f5-tts/(create the folder if it doesn't exist).
Example
If you have an audio file nverguo.wav containing "女儿国王说话", enter:
nverguo.wav#女儿国王说话

Reference Audio Requirements
| Item | Requirement |
|---|---|
| Format | WAV (recommended), MP3 also acceptable |
| Duration | 3~10 seconds |
| Content | Clear pronunciation, no background noise |
| Text | Must match the audio content |
Cloud Services (API Key Required)
| Service | Description | Rating |
|---|---|---|
| Azure TTS | Microsoft professional-grade TTS | ⭐⭐⭐ |
| OpenAI TTS | Leading voice technology | ⭐⭐⭐ |
| ByteDance TTS 2.0 | Natural Chinese pronunciation | ⭐⭐⭐ |
| Alibaba Qwen-TTS | Alibaba Cloud TTS | ⭐⭐⭐ |
| Gemini TTS | Google TTS | ⭐⭐ |
| Elevenlabs.io | AI audio technology company | ⭐⭐⭐ |
| 302.AI | Aggregation platform | ⭐⭐ |
| Minimaxi | Requires top-up | ⭐⭐ |
| Xiaomi TTS | Xiaomi AI Open Platform | ⭐⭐ |
| X.AI TTS | x.ai platform | ⭐⭐ |
