Instructions to use openai/whisper-large-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use openai/whisper-large-v3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="openai/whisper-large-v3")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("openai/whisper-large-v3") model = AutoModelForSpeechSeq2Seq.from_pretrained("openai/whisper-large-v3", device_map="auto") - Inference
- Notebooks
- Google Colab
- Kaggle
how to handle input audio files with either white noise or general noise and no speech
it seems this model performs extremely well when there's actual discernable language/conversation, but if i test with an audio clip that contains non-discernable noise, it produces a bunch of gibberish. is there any way to prevent it from generating gibberish?
here's an example of gibberish produced from white noise audio:
2023-12-14 01:35:51,945 [INFO] Takk for watching! 1 tbsps of butter 1 tbsps of flour 1 tbsps of baking powder 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1 tbsps of baking soda 1.5 kg of pork belly [0:00:03.359228s]
Use VAD and cut no-speech chunks.
https://huggingface.co/pyannote/voice-activity-detection
https://github.com/snakers4/silero-vad
thanks this worked perfectly for me!