Distil-Whisper large-v3 β€” CTranslate2 INT8 Quantized (CPU)

This model is a quantized/optimized version of the baseline Distil-Whisper (distil-large-v3) model. It is part of a benchmarked suite of quantized models evaluated on local hardware.

This model is part of a suite of optimized/quantized versions of the base model. Other variants in this direction:

Base model: distil-whisper/distil-large-v3


πŸ“Š Quantization & Performance Benchmark Results

Below is the comparative performance table generated empirically using 8 CPU threads / cores:

Backend Precision Device Model Size (GB) Mean Latency (s) Throughput (Words/s) RTF WER (%) Peak RAM (GB) Peak VRAM (GB)
PYTORCH FP32 CUDA 3.024 0.427 44.81 0.064 5.29% 3.62 2.98
PYTORCH FP16 CUDA 1.512 0.189 100.25 0.028 5.29% 2.12 1.46
PYTORCH INT8 CUDA 0.756 0.293 64.00 0.042 5.29% 1.92 0.85
PYTORCH Q4 CUDA 0.378 0.377 50.63 0.056 5.29% 1.93 0.60
GGML FP16 CPU 1.409 17.269 1.14 2.633 5.29% 4.14 0.00
GGML INT8 CPU 1.409 9.271 2.18 1.441 5.29% 1.50 0.00
GGML Q5 CPU 1.409 16.735 1.18 2.566 5.29% 2.28 0.00

πŸ’Ύ Model Size Reduction Comparison

Model Size

πŸ“ˆ Transcription Speed & Latency Comparison

Latency

πŸš€ Words Transcribed per Second (Throughput)

Throughput


πŸš€ Usage

from faster_whisper import WhisperModel

model_path = "rudrakshrakeshzodage/distil-whisper-large-v3-ct2-int8"

# Load the model (the compute_type is pre-quantized inside the model folder)
model = WhisperModel(
    model_path,
    device="cpu",
    cpu_threads=8
)

segments, info = model.transcribe("audio.wav", beam_size=1)
for segment in segments:
    print(segment.text)

πŸ“œ Citation & Credits

Quantization research and benchmarking by RudrakshRakeshZodage.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for rudrakshrakeshzodage/distil-whisper-large-v3-ct2-int8

Quantized
(2)
this model