Whisper
Whisper is a general-purpose speech recognition model trained on a large dataset of diverse audio. It performs multilingual speech recognition, speech translation, and language identification.
Visit openai/whisperOverview
Whisper is a general-purpose speech recognition model trained on a large dataset of diverse audio. It is a multitasking model that can perform multilingual speech recognition, speech translation, and language identification. It uses a Transformer sequence-to-sequence model trained on various speech processing tasks, including voice activity detection, with tasks represented as token sequences.
Key Features
- Multilingual speech recognition
- Speech translation
- Language identification
- Voice activity detection
- Multiple model sizes (tiny to turbo) with speed and accuracy tradeoffs
Use Cases
- Transcribe speech in audio files using command-line or Python API
- Translate non-English speech into English
- Detect spoken language in audio
- Process audio with a sliding 30-second window for long recordings
Getting Started
- Install with pip install -U openai-whisper. Requires ffmpeg installed on the system. For Python usage, load a model with whisper.load_model('turbo') and call transcribe on an audio file. Command-line usage: whisper audio.flac --model turbo.
Deployment & Requirements
- Requires Python 3.8-3.11 and PyTorch. Requires ffmpeg. Model sizes have varying VRAM requirements from ~1 GB to ~10 GB. May require Rust if tiktoken does not provide a pre-built wheel.
Before You Adopt
- License: MIT. Review its terms before using, modifying, or distributing the project.
- Performance varies widely by language and real-world speed depends on many factors.