Skip to content

TTS Service: Qwen-TTS

Qwen-TTS is an advanced speech synthesis technology from the Alibaba Tongyi Qianwen team, capable of converting text into highly realistic and natural-sounding human voices. A key highlight is its ability to automatically adjust speech rhythm and emotion based on text content.

Supported languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian.

pyVideoTrans supports 2 forms of Qwen3-TTS:

  • Local Built-in (Offline): Built into the software, no internet required, uses a fixed 0.6B model.
  • Alibaba Bailian API (Online): Accessed via Alibaba Cloud API, requires internet and API Key.

1. Qwen3-TTS Local Built-in (Offline)

Prerequisites

RequirementDetails
pyVideoTrans version≥ v4.04
Model size~5GB (auto-download on first use)
HardwareNVIDIA GPU recommended (GPU acceleration)

Downloading Models (Auto-download on first use)

First use will automatically download both Base and CustomVoice models (~5GB total). Please be patient.

Manual Download (Optional)

If auto-download is too slow:

  1. Open the models folder in the software directory, create 2 new folders:

    • models--Qwen--Qwen3-TTS-12Hz-0.6B-Base
    • models--Qwen--Qwen3-TTS-12Hz-0.6B-CustomVoice
  2. Open the Base model page, download all files and folders into app_dir/models/models--Qwen--Qwen3-TTS-12Hz-0.6B-Base/.

  3. Open the CustomVoice model page, download all files and folders into app_dir/models/models--Qwen--Qwen3-TTS-12Hz-0.6B-CustomVoice/.

Switching to the 1.7B Model (Slower synthesis, better quality)

Simply overwrite the 0.6B model files with the 1.7B model files.

  1. Open https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base/tree/main, download all files and overwrite into app_dir/models/models--Qwen--Qwen3-TTS-12Hz-0.6B-Base/.
  2. Open https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice/tree/main, download all files and folders into app_dir/models/models--Qwen--Qwen3-TTS-12Hz-0.6B-CustomVoice/.

Configuring Reference Audio

For cloning a voice from a 3~10 second reference audio clip.

Path: Menu → Tools → TTS Settings → Qwen-TTS (Local)

Fill in reference audio filename and its corresponding text, one per line.

Reference Audio Format

audio_filename.wav#text spoken in the audio

Example

n10.wav#你说四大皆空,却为何紧闭双眼,你若挣开眼睛看看我,我不相信,你两眼空空

Place n10.wav in the f5-tts folder under the software directory. After the #, enter the text spoken in the audio.


Not recommended for beginners or users unfamiliar with Alibaba Bailian

Alibaba Bailian's API is quite confusing. Available models differ by region (North China 2, Singapore, East China, etc.). API keys are region-specific. Changing models may require changing endpoints. Different models support different voices. You may encounter many unexpected issues.

2. Alibaba Bailian API / Qwen3-TTS (Online)

Qwen3-TTS supports 10 languages and multiple Chinese dialects. Model name: qwen3-tts-flashView Qwen3-TTS voice list

Step 1: Get and Configure API Key

  1. Visit the Alibaba Bailian platform: https://bailian.console.aliyun.com/?tab=model#/api-key

  1. Log in with your Alibaba Cloud account (register if needed).
  2. On the API-KEY management page, click "Create API-KEY". A string starting with sk- will be generated — this is your API Key. Copy it. In pyVideoTrans, go to the top menu → TTS Settings → Qwen TTS.

  1. Click Workspace in the top-right corner, copy your workspace ID, return to pyVideoTrans, and paste it into TTS Settings → Qwen TTS's workspace ID field.

  1. In the "Qwen3 TTS" configuration window, paste the API Key into the "API Key" field. Click "Test" to preview the audio. If you hear sound, configuration is successful. Click "Save".

Step 2: Use Qwen3-TTS in Video Translation

After configuration, select "Qwen3 TTS" from the "TTS Service" dropdown on the main interface, and choose your preferred voice from the "Voices" menu.

  • Cherry: Standard female voice
  • Sunny: Sichuan dialect
  • Dylan: Beijing dialect
  • See more voices in the list below

Step 3: Use in Batch Dubbing and Multi-speaker Dubbing

Qwen-TTS's capabilities extend to batch processing:

  • Batch subtitle dubbing: Select "Qwen TTS" and your preferred voice in the batch dubbing interface.
  • Multi-speaker dubbing: Assign different Qwen-TTS voices to different roles in the multi-speaker dubbing panel.


3. Available Voices

Here are all voices supported by Qwen3-TTS (online version):

Chinese NameEnglish CodeType
芊悦(Cherry)CherryStandard female
苏瑶(Serena)SerenaStandard female
晨煦(Ethan)EthanStandard male
千雪(Chelsie)ChelsieStandard female
茉兔(Momo)MomoStandard female
十三(Vivian)VivianStandard female
月白(Moon)MoonStandard female
四月(Maia)MaiaStandard female
凯(Kai)KaiStandard male
不吃鱼(Nofish)NofishStandard male
萌宝(Bella)BellaChild voice
詹妮弗(Jennifer)JenniferEnglish female
甜茶(Ryan)RyanEnglish male
卡捷琳娜(Katerina)KaterinaRussian female
艾登(Aiden)AidenEnglish male
沧明子(Eldric Sage)Eldric SageEnglish male
乖小妹(Mia)MiaStandard female
沙小弥(Mochi)MochiStandard female
燕铮莺(Bellona)BellonaStandard female
田叔(Vincent)VincentStandard male
萌小姬(Bunny)BunnyStandard female
阿闻(Neil)NeilStandard male
墨讲师(Elias)EliasStandard male
徐大爷(Arthur)ArthurStandard male
邻家妹妹(Nini)NiniStandard female
诡婆婆(Ebona)EbonaStandard female
小婉(Seren)SerenStandard female
顽屁小孩(Pip)PipChild voice
少女阿月(Stella)StellaStandard female
博德加(Bodega)BodegaStandard male
索尼莎(Sonrisa)SonrisaStandard female
阿列克(Alek)AlekStandard male
多尔切(Dolce)DolceStandard female
素熙(Sohee)SoheeKorean female
小野杏(Ono Anna)Ono AnnaJapanese female
莱恩(Lenn)LennStandard male
埃米尔安(Emilien)EmilienFrench male
安德雷(Andre)AndreStandard male
拉迪奥·戈尔(Radio Gol)Radio GolStandard male
上海-阿珍(Jada)JadaShanghainese female
北京-晓东(Dylan)DylanBeijing dialect male
南京-老李(Li)LiNanjing dialect male
陕西-秦川(Marcus)MarcusShaanxi dialect male
闽南-阿杰(Roy)RoySouthern Min dialect male
天津-李彼得(Peter)PeterTianjin dialect male
四川-晴儿(Sunny)SunnySichuan dialect female
四川-程川(Eric)EricSichuan dialect male
粤语-阿强(Rocky)RockyCantonese male
粤语-阿清(Kiki)KikiCantonese female

FAQ

1. Local version downloads models very slowly on first use

~4GB of model files will be auto-downloaded on first use. Please be patient.

2. API version shows AuthenticationError

The API Key is invalid or expired. Log back into Alibaba Bailian to obtain a new API Key.

3. Dubbing sounds unnatural

  • Try different voices
  • For the local version, try adding voice style guidance prompts
  • Ensure good reference audio quality (clear pronunciation, no noise)