mlboydaisuke commited on
Commit
5f8d413
·
verified ·
1 Parent(s): ff4fc37

Add measured Galaxy S26 GPU backend section

Browse files

Functional verdict from the S26 GPU gate. Primary log:
litertlm-convert/community_accel_work/s4_gpu_gate/results.jsonl. Delegation
counts and peak RSS are read from that run's own log; throughput is
deliberately absent, because an Android GPU speed is not quotable without a
CPU row from the same phone and that control has not been run.

Files changed (1) hide show
  1. README.md +18 -1
README.md CHANGED
@@ -24,7 +24,7 @@ multilingual support, and long-context training — a strong small reasoner.
24
 
25
  | | |
26
  |---|---|
27
- | **File** | `SmolLM3-3B_q4_block32_ekv4096.litertlm` (~1.9 GB) |
28
  | **Quantization** | int4 weights — **blockwise (block 32) + OCTAV** optimal-clipping, symmetric; embedding INT8 |
29
  | **Compute** | integer |
30
  | **Context (KV cache)** | 4096 |
@@ -117,6 +117,23 @@ OCTAV recipe with a 4096 KV cache preserves reasoning accuracy exactly at n=100.
117
  model produces visible step-by-step chain-of-thought in the answer body and
118
  terminates cleanly at `<|im_end|>` (no rambling).
119
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
120
  ## Conversion
121
 
122
  Converted with [`litert-torch`](https://github.com/google-ai-edge/litert) via its
 
24
 
25
  | | |
26
  |---|---|
27
+ | **File** | `SmolLM3-3B_q4_block32_ekv4096.litertlm` (~1.9 GB) — the repo also carries `SmolLM3-3B.litertlm` (~3.1 GB), gated on the Galaxy S26 GPU below |
28
  | **Quantization** | int4 weights — **blockwise (block 32) + OCTAV** optimal-clipping, symmetric; embedding INT8 |
29
  | **Compute** | integer |
30
  | **Context (KV cache)** | 4096 |
 
117
  model produces visible step-by-step chain-of-thought in the answer body and
118
  terminates cleanly at `<|im_end|>` (no rambling).
119
 
120
+ ## Galaxy S26 — GPU backend
121
+
122
+ Both published bundles run on the Android GPU backend and generate.
123
+
124
+ | file | GPU backend | delegation | peak |
125
+ |---|---|---|---:|
126
+ | `SmolLM3-3B.litertlm` | runs | 2930 / 2930 ops across 2 subgraphs on `LiteRT GPU` | 704 MB |
127
+ | `SmolLM3-3B_q4_block32_ekv4096.litertlm` | runs | 2784 / 2784 ops across 2 subgraphs on `LiteRT GPU` | 1111 MB |
128
+
129
+ Measured on a **Samsung Galaxy S26** (SM-S942Q / SM8850, Android 16) with `litert_lm_advanced_main` from litert-lm 0.16.0, `--backend=gpu --sampler_backend=cpu`, prompt `What is the capital of France?`. Peak is the process high-water mark (`VmHWM`) sampled during that same run. Gated 2026-08-24.
130
+
131
+ The op counts above are the `LiteRT GPU` partitions. In `SmolLM3-3B_q4_block32_ekv4096.litertlm`, XNNPACK additionally takes 1 of the 4 nodes in `decode_embedder` and 1 of the 4 nodes in `prefill_embedder_128`; the runtime accepts that split.
132
+
133
+ **No speed rows, on purpose.** On this handset the GPU backend wins prefill and does not win decode, so a GPU throughput figure only means something beside a CPU row from the same handset, and no S26 CPU row exists for this model yet.
134
+
135
+ GPU wiring, including the Gallery import toggle: [GPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-gpu.md).
136
+
137
  ## Conversion
138
 
139
  Converted with [`litert-torch`](https://github.com/google-ai-edge/litert) via its