Skip to content

TTS Dubbing Services

Dubbing (TTS) is the third step of video translation, converting translated subtitle text into speech audio. pyVideoTrans supports 30+ dubbing services.

You can also use the left panel "Batch Subtitle Dubbing" feature independently — import multiple SRT or TXT files for dubbing.

The clone voice role uses the original video speaker's voice for dubbing (voice cloning). This role is only available in the main "Video Translation" feature.

All services with a clone role support custom reference audio — clone the voice from your own 3-10s audio clip. See "Creating and Using Reference Audio" below.


Out-of-the-box (Free)

No complex configuration needed — great for beginners.

ServiceDescriptionRating
Edge-TTS (Free)Microsoft free API, natural voice, all languages⭐⭐⭐ Default
gTTS (Free)Google TTS, basic quality, VPN needed in China⭐⭐

⚠️ Heavy Edge-TTS usage may trigger rate limits. Set concurrency to 1 and pause to 5-10 seconds in advanced options.


Built-in (Free)

Models auto-download on first use.

ServiceDescriptionGPUCloneRating
Qwen3-TTS (built-in)Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian⭐⭐⭐
F5-TTS (built-in)Chinese, English, Japanese, French, German, Russian, Italian, Spanish, Hindi, Arabic⭐⭐⭐
OmniVoice-TTS600 languages (built-in since v4.05)⭐⭐⭐
Confucius-TTS (built-in, v4.06+)Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, Vietnamese⭐⭐⭐
MOSS-TTS-Nano (built-in)Chinese, English, German, Spanish, French, Japanese, Italian, Hungarian, Korean, Russian, Persian, Arabic, Polish, Portuguese, Czech, Swedish, Greek, Turkish⭐⭐
ZipVoice (built-in)Chinese, English⭐⭐⭐
Piper (built-in)Lightweight, 20 languages⭐⭐
ChatterBox (built-in)Arabic, German, Greek, English, Spanish, Finnish, French, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Dutch, Norwegian, Polish, Portuguese, Russian, Swedish, Swahili, Turkish, Chinese⭐⭐⭐
Supertonic3 (built-in)English, Korean, Japanese, Arabic, Czech, German, Greek, Spanish, French, Hindi, Hungarian, Indonesian, Italian, Dutch, Polish, Portuguese, Romanian, Russian, Swedish, Turkish, Ukrainian, Vietnamese⭐⭐
VITS (built-in)Chinese, English⭐⭐

Due to network conditions in China, models are large and auto-download may fail. If it fails, see model download URLs and manual download methods


Self-hosted (Advanced)

ServiceDescriptionCloneRating
GPT-SoVITSChinese, English, Japanese, Korean, Cantonese⭐⭐⭐
Index-TTSChinese, English⭐⭐⭐
VoxCPM-TTSArabic, Burmese, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Tagalog, Thai, Turkish, Vietnamese⭐⭐⭐
CosyVoiceChinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian⭐⭐
ChatTTSChinese, English⭐⭐
Fish-TTSAll built-in languages
Kokoro-TTSChinese, English, Korean, Italian, Portuguese, German, French, Hindi
Spark-TTSChinese, English⭐⭐
clone-voiceNo longer maintained

Creating and Using Reference Audio

TTS services that support clone (where the voice role includes clone) can all use reference audio. Select the reference audio in the voice role dropdown, and the software will use that audio's voice for dubbing. Reference audio is configured in Menu → TTS Settings → Set Reference Audio.

Steps

  1. Record or extract a 3-10 second audio clip, save as WAV, ensure clear pronunciation with no background noise.
  2. Open Menu → TTS Settings → Set Reference Audio.
  3. Enter content in the following format:
    audio_filename.wav#text spoken in the audio
  4. Place the reference audio file in app_dir/f5-tts/ (create the folder if it doesn't exist).

Example

If you have an audio file nverguo.wav containing "女儿国王说话", enter:

nverguo.wav#女儿国王说话

Place reference audio in pyVideoTrans's f5-tts folder

Reference audio and its text content

Reference Audio Requirements

ItemRequirement
FormatWAV (recommended), MP3 also acceptable
Duration3~10 seconds
ContentClear pronunciation, no background noise
TextMust match the audio content

Cloud Services (API Key Required)

ServiceDescriptionRating
Azure TTSMicrosoft professional-grade TTS⭐⭐⭐
OpenAI TTSLeading voice technology⭐⭐⭐
ByteDance TTS 2.0Natural Chinese pronunciation⭐⭐⭐
Alibaba Qwen-TTSAlibaba Cloud TTS⭐⭐⭐
Gemini TTSGoogle TTS⭐⭐
Elevenlabs.ioAI audio technology company⭐⭐⭐
302.AIAggregation platform⭐⭐
MinimaxiRequires top-up⭐⭐
Xiaomi TTSXiaomi AI Open Platform⭐⭐
X.AI TTSx.ai platform⭐⭐