Skip to main content
AdOpenFree logoPromote your productReach more potential users and drive product growth and revenue.Advertise
Favicon of VideoLingo

VideoLingo

Free Listing

VideoLingo is a one-click AI subtitle and dubbing pipeline for video localization, offering Netflix-level subtitle cutting, translation, alignment, and dubbing.

Visit Huanshere/VideoLingo

Overview

VideoLingo is an AI video localization tool that combines speech recognition, subtitle translation, segmentation, alignment, and dubbing in a Streamlit interface. It produces subtitle files and optionally subtitled or dubbed videos. Translation quality depends on the source audio, language, and chosen models.

Key Features

  • Intelligent subtitle segmentation using NLP and LLM technologies to segment subtitles based on sentence meaning.
  • Context-aware translation with GPT-generated terminology knowledge base and a three-step process of direct translation, reflection, and paraphrasing.
  • Precise word-level subtitle alignment using Qwen3-ASR with Qwen3-ForcedAligner.
  • High-quality dubbing with GPT-SoVITS, OpenAI, Edge TTS, and other TTS solutions, including personalized voice cloning.
  • YouTube video download via yt-dlp and one-click startup and processing in Streamlit.
  • Developer-friendly structured files, multiple deployment methods, model searchbox with API auto-fetch, and task control for pause, resume, or stop.

Use Cases

  • Localize videos by generating translated subtitles and optionally dubbed audio.
  • Create dual subtitles for videos.
  • Dub videos with a cloned voice using GPT-SoVITS or other TTS methods.
  • Process YouTube videos through a multilingual subtitle and dubbing workflow.

Getting Started

  • On Windows, download the source code zip from the latest release, extract it, double-click OneKeyStart.bat, and keep the window open; the first run installs uv, Python 3.12, app dependencies, and FFmpeg and requires an internet connection.
  • After installation, VideoLingo opens in the browser; enter the API URL, key, and model in the sidebar to start using it.
  • Install from source on Windows, macOS, or Linux by cloning the repository and running uv run start.py.
  • For the HTTP API, configure config.yaml and run uv run start.py --api from the project root.

Deployment & Requirements

  • VideoLingo supports Windows, macOS (Apple Silicon or Intel), and Linux.
  • The Windows one-click first run requires an internet connection and installs uv, Python 3.12, app dependencies, and FFmpeg.
  • Docker deployment for Linux NVIDIA containers requires Docker, a compatible GPU driver, and the NVIDIA Container Toolkit; the image uses Python 3.12 and CUDA 12.8.1/cu128 by default.
  • LLM use requires an OpenAI-compatible Chat Completions provider and model that can return the structured JSON required by the workflow.
  • Speech recognition runs Qwen3-ASR with ForcedAligner locally by default, or can use ElevenLabs or MAI-Transcribe-2.

Before You Adopt

  • License: Apache-2.0. Review its terms before using, modifying, or distributing the project.
  • The former Excel batch mode has been replaced by a local HTTP API for agents and scripts.
  • Some speech recognition options send audio to external providers and may incur charges.
  • Recognition and word timestamps can be affected by background noise and language-specific alignment models, and numbers or symbols may lack reliable word timings.
  • The workflow requires LLM output to satisfy its JSON structure.
  • Dubbing quality and timing depend on translation, TTS service, and speech rate, and speed adjustment does not guarantee natural delivery or perfect synchronization.

Comments

Sign in to leave a comment.

More like VideoLingo

Favicon of VidBee

VidBee

Free ListingStars: 10.8K

Download videos and create searchable transcripts locally.

Assistants & ProductivityTranscription & Dictation

VidBee downloads video and audio from YouTube, TikTok, Instagram, X, and 1000+ sites, or local files. It creates searchable, timestamped transcripts on your computer, then summarizes, translates, or answers questions with your AI provider.

Favicon of Whisper

Whisper

Free ListingStars: 109.6K

General-purpose speech recognition model

Transcription & Dictation

Whisper is a general-purpose speech recognition model trained on a large dataset of diverse audio. It performs multilingual speech recognition, speech translation, and language identification.