--- license: apache-2.0 base_model: - openai/whisper-base pipeline_tag: automatic-speech-recognition --- # Whisper Base LiteRT This repository contains a LiteRT `.tflite` export of `openai/whisper-base`. ## Files - `whisper_base_30s_f32.tflite`: FP32 LiteRT model with `encode` and `decode` signatures. ## Signatures - `whisper_base_30s_f32.tflite` - `encode`: `float32[1,80,3000] -> float32[1,1500,512]` - `decode`: `float32[1,1500,512]`, `int32[1,128]`, `float32[1,1,128,128] -> float32[1,128,51865]` ## Performance (Galaxy S26, measured) Measured on a physical Samsung Galaxy S26 (SM-S942Q, Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT `CompiledModel` 2.2.0 — 5 warm-up runs then 50 timed runs, one accelerator per process, every row at device thermal status `NONE`, delegate placement confirmed from logcat per row. The figures are the **`encode` signature only** (the model's first signature, which the API's default `run()` executes); the decode loop is not benchmarked here. NPU rows are on-device (JIT) compilation in HTP BURST mode. | File | Compute unit | Encode (median / min) | Load | |---|---|---|---| | `whisper_base_30s_f32.tflite` | **GPU (Adreno)** | **50.0 ms** / 48.0 ms | 4.3 s | | `whisper_base_30s_f32.tflite` | NPU (Hexagon, JIT) — first launch | 58.9 ms / 57.7 ms | 26.3 s | | `whisper_base_30s_f32.tflite` | NPU (Hexagon, JIT) — cached | 60.6 ms / 57.4 ms | **0.24 s** | | `whisper_base_30s_i8.tflite` | NPU (Hexagon, JIT) — first launch | 193.9 ms / 186.9 ms | 22.8 s | | `whisper_base_30s_i8.tflite` | NPU (Hexagon, JIT) — cached | 211.3 ms / 201.9 ms | 0.72 s | What the table says (the same shape as whisper-tiny on this device): - **The GPU is the fastest unit on this encoder** — 50.0 against 58.9 ms on the NPU. f32 only: the i8 file does not compile on the GPU delegate (`Failed to compile model`). - **int8 does not help the NPU** — 194–211 ms, 3.3–3.5× slower than its f32 sibling on the same unit. The win of i8 stays the ~3.8× smaller file. - **JIT compile is a one-time load cost**: ~26 s on first launch for f32, then ~0.24 s from cache; the GPU instead pays its ~4 s compile every process (unless the app caches).