Add measured Galaxy S26 NPU/GPU Performance section
Browse files
README.md
CHANGED
|
@@ -18,3 +18,21 @@ This repository contains a LiteRT `.tflite` export of `openai/whisper-base`.
|
|
| 18 |
- `whisper_base_30s_f32.tflite`
|
| 19 |
- `encode`: `float32[1,80,3000] -> float32[1,1500,512]`
|
| 20 |
- `decode`: `float32[1,1500,512]`, `int32[1,128]`, `float32[1,1,128,128] -> float32[1,128,51865]`
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 18 |
- `whisper_base_30s_f32.tflite`
|
| 19 |
- `encode`: `float32[1,80,3000] -> float32[1,1500,512]`
|
| 20 |
- `decode`: `float32[1,1500,512]`, `int32[1,128]`, `float32[1,1,128,128] -> float32[1,128,51865]`
|
| 21 |
+
|
| 22 |
+
## Performance (Galaxy S26, measured)
|
| 23 |
+
|
| 24 |
+
Measured on a physical Samsung Galaxy S26 (SM-S942Q, Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT `CompiledModel` 2.2.0 — 5 warm-up runs then 50 timed runs, one accelerator per process, every row at device thermal status `NONE`, delegate placement confirmed from logcat per row. The figures are the **`encode` signature only** (the model's first signature, which the API's default `run()` executes); the decode loop is not benchmarked here. NPU rows are on-device (JIT) compilation in HTP BURST mode.
|
| 25 |
+
|
| 26 |
+
| File | Compute unit | Encode (median / min) | Load |
|
| 27 |
+
|---|---|---|---|
|
| 28 |
+
| `whisper_base_30s_f32.tflite` | **GPU (Adreno)** | **50.0 ms** / 48.0 ms | 4.3 s |
|
| 29 |
+
| `whisper_base_30s_f32.tflite` | NPU (Hexagon, JIT) — first launch | 58.9 ms / 57.7 ms | 26.3 s |
|
| 30 |
+
| `whisper_base_30s_f32.tflite` | NPU (Hexagon, JIT) — cached | 60.6 ms / 57.4 ms | **0.24 s** |
|
| 31 |
+
| `whisper_base_30s_i8.tflite` | NPU (Hexagon, JIT) — first launch | 193.9 ms / 186.9 ms | 22.8 s |
|
| 32 |
+
| `whisper_base_30s_i8.tflite` | NPU (Hexagon, JIT) — cached | 211.3 ms / 201.9 ms | 0.72 s |
|
| 33 |
+
|
| 34 |
+
What the table says (the same shape as whisper-tiny on this device):
|
| 35 |
+
|
| 36 |
+
- **The GPU is the fastest unit on this encoder** — 50.0 against 58.9 ms on the NPU. f32 only: the i8 file does not compile on the GPU delegate (`Failed to compile model`).
|
| 37 |
+
- **int8 does not help the NPU** — 194–211 ms, 3.3–3.5× slower than its f32 sibling on the same unit. The win of i8 stays the ~3.8× smaller file.
|
| 38 |
+
- **JIT compile is a one-time load cost**: ~26 s on first launch for f32, then ~0.24 s from cache; the GPU instead pays its ~4 s compile every process (unless the app caches).
|