
VideoLingo
VideoLingo is a one-click AI subtitle and dubbing pipeline for video localization, offering Netflix-level subtitle cutting, translation, alignment, and dubbing.
Visit Huanshere/VideoLingoOverview
VideoLingo is an AI video localization tool that combines speech recognition, subtitle translation, segmentation, alignment, and dubbing in a Streamlit interface. It produces subtitle files and optionally subtitled or dubbed videos. Translation quality depends on the source audio, language, and chosen models.
Key Features
- Intelligent subtitle segmentation using NLP and LLM technologies to segment subtitles based on sentence meaning.
- Context-aware translation with GPT-generated terminology knowledge base and a three-step process of direct translation, reflection, and paraphrasing.
- Precise word-level subtitle alignment using Qwen3-ASR with Qwen3-ForcedAligner.
- High-quality dubbing with GPT-SoVITS, OpenAI, Edge TTS, and other TTS solutions, including personalized voice cloning.
- YouTube video download via yt-dlp and one-click startup and processing in Streamlit.
- Developer-friendly structured files, multiple deployment methods, model searchbox with API auto-fetch, and task control for pause, resume, or stop.
Use Cases
- Localize videos by generating translated subtitles and optionally dubbed audio.
- Create dual subtitles for videos.
- Dub videos with a cloned voice using GPT-SoVITS or other TTS methods.
- Process YouTube videos through a multilingual subtitle and dubbing workflow.
Getting Started
- On Windows, download the source code zip from the latest release, extract it, double-click OneKeyStart.bat, and keep the window open; the first run installs uv, Python 3.12, app dependencies, and FFmpeg and requires an internet connection.
- After installation, VideoLingo opens in the browser; enter the API URL, key, and model in the sidebar to start using it.
- Install from source on Windows, macOS, or Linux by cloning the repository and running uv run start.py.
- For the HTTP API, configure config.yaml and run uv run start.py --api from the project root.
Deployment & Requirements
- VideoLingo supports Windows, macOS (Apple Silicon or Intel), and Linux.
- The Windows one-click first run requires an internet connection and installs uv, Python 3.12, app dependencies, and FFmpeg.
- Docker deployment for Linux NVIDIA containers requires Docker, a compatible GPU driver, and the NVIDIA Container Toolkit; the image uses Python 3.12 and CUDA 12.8.1/cu128 by default.
- LLM use requires an OpenAI-compatible Chat Completions provider and model that can return the structured JSON required by the workflow.
- Speech recognition runs Qwen3-ASR with ForcedAligner locally by default, or can use ElevenLabs or MAI-Transcribe-2.
Before You Adopt
- License: Apache-2.0. Review its terms before using, modifying, or distributing the project.
- The former Excel batch mode has been replaced by a local HTTP API for agents and scripts.
- Some speech recognition options send audio to external providers and may incur charges.
- Recognition and word timestamps can be affected by background noise and language-specific alignment models, and numbers or symbols may lack reliable word timings.
- The workflow requires LLM output to satisfy its JSON structure.
- Dubbing quality and timing depend on translation, TTS service, and speech rate, and speed adjustment does not guarantee natural delivery or perfect synchronization.
