Skip to content

Dubbing & TTS Channels

Dubbing (TTS) is the third step in video translation. It converts translated subtitle text into speech audio. pyVideoTrans supports 30+ dubbing channels.

You can also use the dubbing feature independently via the left panel: Batch Dub Subtitles. It supports importing multiple SRT subtitle files or TXT files.

The clone voice role uses the speaker's original voice timbre from the video to perform voice cloning. This role is only available in the main interface's Video Translation feature.

All channels that support the clone role also support custom reference audio. This allows you to clone a voice from a 3–10 second audio clip you provide. See Create and Use Reference Audio below for details.


Ready to Use (Free)

No complex configuration required—ideal for beginners.

ChannelDescriptionRating
Edge-TTS(Free)Microsoft's free service, natural voices, supports all languages⭐⭐⭐ Default Recommended
gTTS(Free)Google TTS, basic quality, requires a proxy in mainland China⭐⭐

⚠️ Heavy usage of Edge-TTS in a short period may trigger rate limiting. We recommend setting concurrency to 1 and pause duration to 5–10 seconds in Advanced Options.


Built-in (Free)

Models will be downloaded automatically upon first use.

ChannelDescriptionGPU AccelerationSupports Cloning (clone role)Rating
Qwen3-TTS(Built-in)Supports Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian⭐⭐⭐ Recommended
F5-TTS(Built-in)Chinese, English, Japanese, French, German, Russian, Italian, Spanish, Hindi, Arabic⭐⭐⭐
OmniVoice-TTSSupports 600+ languages (Built-in from v4.05)⭐⭐⭐ Recommended
Confucius-TTS(Built-in)from v4.06Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, Vietnamese⭐⭐⭐
MOSS-TTS-Nano(Built-in)Chinese, English, German, Spanish, French, Japanese, Italian, Hungarian, Korean, Russian, Persian, Arabic, Polish, Portuguese, Czech, Swedish, Greek, Turkish⭐⭐
ZipVoice(Built-in)Chinese and English⭐⭐⭐ Recommended
Piper(Built-in)Lightweight, supports 20 languages⭐⭐
ChatterBox(Built-in)Arabic, German, Greek, English, Spanish, Finnish, French, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Dutch, Norwegian, Polish, Portuguese, Russian, Swedish, Swahili, Turkish, Chinese⭐⭐⭐ Recommended
Supertonic3(Built-in)English, Korean, Japanese, Arabic, Czech, German, Greek, Spanish, French, Hindi, Hungarian, Indonesian, Italian, Dutch, Polish, Portuguese, Romanian, Russian, Swedish, Turkish, Ukrainian, Vietnamese⭐⭐
VITS(Built-in)Chinese and English dubbing⭐⭐
Higgs-audio-v3(Built-in)All built-in languages; CPU-only requires 20GB RAM, GPU acceleration requires >10GB VRAM⭐⭐

Due to network conditions and large model sizes, automatic downloads might occasionally fail. If a download fails, click here to view model download links and manual setup instructions.


Self-Hosted / Local Deployment (Advanced)

ChannelDescriptionSupports Cloning (clone role)Rating
GPT-SoVITSSupports Chinese, English, Japanese, Korean, Cantonese⭐⭐⭐ Recommended
Index-TTSChinese, English⭐⭐⭐ Recommended
VoxCPM-TTSArabic, Burmese, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Tagalog, Thai, Turkish, Vietnamese⭐⭐⭐
CosyVoiceChinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian⭐⭐
ChatTTSSupports Chinese and English⭐⭐
Fish-TTSSupports all built-in languages
Kokoro-TTSChinese, English, Korean, Italian, Portuguese, German, French, Hindi
Spark-TTSChinese, English⭐⭐
clone-voiceNo longer maintained

Create and Use Reference Audio

Any dubbing channel that supports voice cloning (where the role list includes clone) can use reference audio. Once configured, simply select your reference audio as the voice role, and the software will automatically dub the audio using that voice timbre. Reference audio is configured under Menu → TTS Settings → Reference Audio Settings.

Step-by-Step Guide

  1. Record or trim a 3–10s audio clip from existing audio. Save it as a .wav file, making sure the pronunciation is clear and free of background noise.
  2. Open the Menu → TTS Settings → Reference Audio Settings window.
  3. Enter your reference audio info in the text box using the following format:
audio_filename#matching_text_content
  1. Place the reference audio file inside the pyVideoTrans_folder/f5-tts folder (create this folder manually if it doesn't exist).

Example

Suppose you have an audio file named nverguo.wav, and the spoken words in that audio are "The text content corresponding to the audio". You would enter:

nverguo.wav#The text content corresponding to the audio

Reference Audio Requirements

ItemRequirement
FormatWAV format (recommended); MP3 and other common formats are also supported
Duration3–10 seconds
Audio QualityClear speech, no background noise
TextMust match the spoken words in the audio exactly

Professional Cloud Services (API Key Required)

ChannelDescriptionRating
Azure TTSMicrosoft's enterprise-grade voice service⭐⭐⭐
OpenAI TTSIndustry-leading audio technology⭐⭐⭐
ByteDance TTS 2.0Authentic Chinese pronunciation⭐⭐⭐
Alibaba Qwen-TTSAlibaba Cloud text-to-speech⭐⭐⭐
Gemini TTSGoogle TTS⭐⭐
Elevenlabs.ioHigh-end AI voice generation⭐⭐⭐
302.AIAPI aggregation platform⭐⭐
MinimaxiMiniMax TTS (requires paid balance)⭐⭐
Xiaomi TTSXiaomi AI Open Platform⭐⭐
X.AI TTSx.ai platform⭐⭐