|
Download README.md from SadeghK/whisper-large-v3-turbo-ct2: direct link, hf CLI and curl.
- Browser
- Download file 1.69 kB
-
https://huggingface.co/SadeghK/whisper-large-v3-turbo-ct2/resolve/main/README.md
- Command line
-
hf download hf://SadeghK/whisper-large-v3-turbo-ct2/README.md
-
curl -L -o README.md https://huggingface.co/SadeghK/whisper-large-v3-turbo-ct2/resolve/main/README.md
1.69 kB
metadata
license: apache-2.0
language:
- fa
base_model:
- SadeghK/whisper-large-v3-turbo
tags:
- ASR
- Persian
- Farsi
Whisper Large V3 Turbo (CTranslate2) β Optimized for Faster-Whisper & WhisperX
This repository contains a CTranslate2 (CT2) optimized version of the whisper-large-v3-turbo model finetuned by SadeghK.
It is designed for high-speed inference, low-latency ASR, and full WhisperX compatibility (ASR + alignment + diarization).
π Model Overview
This is a converted version of the original SadeghK/whisper-large-v3-turbo model into CTranslate2 format, which enables:
- β Faster inference (up to 4Γ vs PyTorch)
- β Lower memory usage (supports float16 / int8 / int8_float16)
- β Full compatibility with faster-whisper
- β Full compatibility with WhisperX for:
- ASR transcription
- Word-level alignment
- (optional) speaker diarization
All weights in this repository are ready-to-use, no additional conversion required.
π¬ Usage with WhisperX (ASR + alignment)
import whisperx
device = "cuda"
# ASR
asr_model = whisperx.load_model(
"SadeghK/whisper-large-v3-turbo-ct2",
device=device,
compute_type="float16"
)
result = asr_model.transcribe("audio.wav")
# Alignment (example for Persian)
align_model, metadata = whisperx.load_align_model("fa", device)
aligned = whisperx.align(result["segments"], align_model, metadata, "audio.wav", device)
π Repository Structure
whisper-large-v3-turbo-ct2/
β
βββ config.json
βββ model.bin
βββ preprocessor_config.json
βββ tokenizer.json
βββ vocabulary.json