Skip to content

Speech Recognition Channels

Speech recognition (ASR) is the first step in video translation. It converts the spoken audio or video into subtitle files with timestamps. pyVideoTrans supports 15+ recognition channels.

If you are not sure what language is spoken in the video, you can use the Speech-To-Text feature in the left panel. Set the spoken language to Auto Detect and choose a recognition channel with broad language support, such as faster-whisper/openai-whisper/Omnilingual/Dolphin.


Local Offline Recognition

Works completely offline without an internet connection. Models are downloaded automatically upon first use.

ChannelDescriptionGPU AccelerationRecommendation
faster-whisper(Built-in)Fast speed and high quality, supports all built-in languages⭐⭐⭐ Default Recommendation
openai-whisper(Built-in)High accuracy, slightly slower, supports all built-in languages⭐⭐⭐
Qwen-ASR(Built-in)Excellent for Chinese, supports most built-in languages⭐⭐⭐
Whisper.cpp(Win Built-in)Built into Windows for direct use; requires separate manual deployment on macOS and Linux⭐⭐⭐
FunASR(Built-in)Excellent for Chinese, with multiple models available⭐⭐⭐
Firered Chinese(Built-in)Only supports Chinese and 20 Chinese dialectsX⭐⭐
Dolphin(Built-in)Supports 40+ Asian languages and 20 Chinese dialectsX⭐⭐
Omnilingual ASR(Built-in)Supports all built-in languages and 1,600+ moreX⭐⭐
Nemotron-3.5-asr-0.6(Built-in)English,japanese,ko,vietnamese,eg. 40⭐⭐
Huggingface_ASR(Built-in)Multiple language models available to choose from⭐⭐
Moss-Diarize(Built-in)Supports 50+ languages; can separate speakers for audio under 90 minutes. Files over 90 minutes require a separate speaker model⭐⭐
Faster-Whisper-XXL.exeA standalone Windows package of faster-whisper. Requires manual download and specifying the .exe path⭐⭐

Because models are quite large and network environments vary, automatic downloads might sometimes fail. If a download fails, click here to view model download links and manual setup instructions.

Model Selection Guide for faster-whisper / openai-whisper

ModelSpeedAccuracyVRAM Required
tinyFastestLow~1GB
baseFastLow-Medium~1GB
smallMediumMedium~2GB
mediumSlowHigh~5GB
large-v3SlowestHighest~8GB
large-v3-turboFastHigh~6GB

Recommendation: large-v3-turbo offers the best balance between speed and quality.


Online Recognition

ChannelDescription
Alibaba Cloud Bailian Qwen3-ASRRequires activating the Alibaba Cloud Bailian service
XiaomiThe mimo-v2.5-asr model works great for Chinese and mixed Chinese-English. Requires activating the Xiaomi AI platform, topping up, and getting an API key. Enter it under Menu -> Translation Settings -> Xiaomi AI
ByteDance Speech Recognition Large Model Express EditionOutstanding performance on Chinese
Elevenlabs.io Speech RecognitionFree accounts have strict rate limits and are barely usable
Deepgram.comRequires registering for an API Key
Gemini AIStrong at recognizing low-resource/minority languages; requires proxy/VPN access in restricted regions
302.AIApply on 302.ai
OpenAI Speech Recognition APIExcellent quality, requires an OpenAI API key (SK key)

Advanced & Custom

ChannelDescription
Parakeet-tdt(LocalAPI)Requires separate manual deployment
WhisperX(LocalAPI)Requires separate manual deployment
STT(LocalAPI)Requires separate manual deployment
Whisper.NETSupports AMD GPU acceleration. Requires source code installation and downloading the required DLL files according to the guide
Custom Speech Recognition APIAllows you to connect your own custom speech recognition API endpoint

Available Models for Huggingface_ASR(Built-in)

ModelSupported Languages
nvidia/parakeet-tdt-0.6b-v3en,bg,hr,cs,da,nl,et,fi,fr,de,el,hu,it,lv,lt,mt,pl,pt,ro,sk,sl,es,sv,ru,uk
nvidia/nemotron-3.5-asr-streaming-0.6ben,bg,hr,cs,ko,ja,vi,.eg 40
reazon-research/japanese-wav2vec2-large-rs35khJapanese
kotoba-tech/kotoba-whisper-v2.0Japanese
biodatlab/whisper-th-large-v3Thai
vinai/Phowhisper-largeVietnamese
anke01/whisper-small-uyghurUyghur
openai/whisper-large-v3All languages
zai-org/GLM-ASR-Nano-2512Zhipu AI model, excellent for Mandarin and Cantonese
Audio8/ARK-ASR-0.6B/3BChinese, English, German, Japanese, French, Korean, Spanish, Polish, Italian, Romanian, Hungarian, Czech, Dutch
ibm-granite/granite-speech-4.1-2bEnglish, French, German, Spanish, Portuguese, Japanese
nguyenvulebinh/wav2vec2-base-vietnamese-250hvietnamese
sakares/wav2vec2-large-xlsr-thai-demothai
SiangLao/xlsr-53-lao-asrlao
chuuhtetnaing/whisper-large-v3-myanmarmyanmar
1morecupofhottea/whisper-turbo-khmer-v9khmer
anke01/whisper-small-uyghuruyghur
kingabzpro/whisper-large-v3-turbo-urduurdu
vasista22/whisper-tamil-smalltamil
theainerd/Wav2Vec2-large-xlsr-hindihindi
cautroi/whisper-large-v3-idIndonesia
Khalsuu/filipino-wav2vec2-l-xls-r-300m-officialFilipino
navai-uz/whisper-medium-uzbekuzbek
jonatasgrosman/wav2vec2-large-xlsr-53-persianpersian
Ghost3454/translynx-pakistani-punjabi-whisper-smallpakistani
turkmedstt/whisper-large-v3-turkish-generalturkish
anton-l/wav2vec2-large-xlsr-53-mongolianmongolian
HNO333333/w2v-bert-2.0-Tibetan-AmdoTibetan