Instructions to use litert-community/SmolLM3-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/SmolLM3-3B with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli litert-lm run \ --from-huggingface-repo=litert-community/SmolLM3-3B \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/SmolLM3-3B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Add measured Galaxy S26 GPU backend section
Browse filesFunctional verdict from the S26 GPU gate. Primary log:
litertlm-convert/community_accel_work/s4_gpu_gate/results.jsonl. Delegation
counts and peak RSS are read from that run's own log; throughput is
deliberately absent, because an Android GPU speed is not quotable without a
CPU row from the same phone and that control has not been run.
README.md
CHANGED
|
@@ -24,7 +24,7 @@ multilingual support, and long-context training — a strong small reasoner.
|
|
| 24 |
|
| 25 |
| | |
|
| 26 |
|---|---|
|
| 27 |
-
| **File** | `SmolLM3-3B_q4_block32_ekv4096.litertlm` (~1.9 GB) |
|
| 28 |
| **Quantization** | int4 weights — **blockwise (block 32) + OCTAV** optimal-clipping, symmetric; embedding INT8 |
|
| 29 |
| **Compute** | integer |
|
| 30 |
| **Context (KV cache)** | 4096 |
|
|
@@ -117,6 +117,23 @@ OCTAV recipe with a 4096 KV cache preserves reasoning accuracy exactly at n=100.
|
|
| 117 |
model produces visible step-by-step chain-of-thought in the answer body and
|
| 118 |
terminates cleanly at `<|im_end|>` (no rambling).
|
| 119 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 120 |
## Conversion
|
| 121 |
|
| 122 |
Converted with [`litert-torch`](https://github.com/google-ai-edge/litert) via its
|
|
|
|
| 24 |
|
| 25 |
| | |
|
| 26 |
|---|---|
|
| 27 |
+
| **File** | `SmolLM3-3B_q4_block32_ekv4096.litertlm` (~1.9 GB) — the repo also carries `SmolLM3-3B.litertlm` (~3.1 GB), gated on the Galaxy S26 GPU below |
|
| 28 |
| **Quantization** | int4 weights — **blockwise (block 32) + OCTAV** optimal-clipping, symmetric; embedding INT8 |
|
| 29 |
| **Compute** | integer |
|
| 30 |
| **Context (KV cache)** | 4096 |
|
|
|
|
| 117 |
model produces visible step-by-step chain-of-thought in the answer body and
|
| 118 |
terminates cleanly at `<|im_end|>` (no rambling).
|
| 119 |
|
| 120 |
+
## Galaxy S26 — GPU backend
|
| 121 |
+
|
| 122 |
+
Both published bundles run on the Android GPU backend and generate.
|
| 123 |
+
|
| 124 |
+
| file | GPU backend | delegation | peak |
|
| 125 |
+
|---|---|---|---:|
|
| 126 |
+
| `SmolLM3-3B.litertlm` | runs | 2930 / 2930 ops across 2 subgraphs on `LiteRT GPU` | 704 MB |
|
| 127 |
+
| `SmolLM3-3B_q4_block32_ekv4096.litertlm` | runs | 2784 / 2784 ops across 2 subgraphs on `LiteRT GPU` | 1111 MB |
|
| 128 |
+
|
| 129 |
+
Measured on a **Samsung Galaxy S26** (SM-S942Q / SM8850, Android 16) with `litert_lm_advanced_main` from litert-lm 0.16.0, `--backend=gpu --sampler_backend=cpu`, prompt `What is the capital of France?`. Peak is the process high-water mark (`VmHWM`) sampled during that same run. Gated 2026-08-24.
|
| 130 |
+
|
| 131 |
+
The op counts above are the `LiteRT GPU` partitions. In `SmolLM3-3B_q4_block32_ekv4096.litertlm`, XNNPACK additionally takes 1 of the 4 nodes in `decode_embedder` and 1 of the 4 nodes in `prefill_embedder_128`; the runtime accepts that split.
|
| 132 |
+
|
| 133 |
+
**No speed rows, on purpose.** On this handset the GPU backend wins prefill and does not win decode, so a GPU throughput figure only means something beside a CPU row from the same handset, and no S26 CPU row exists for this model yet.
|
| 134 |
+
|
| 135 |
+
GPU wiring, including the Gallery import toggle: [GPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-gpu.md).
|
| 136 |
+
|
| 137 |
## Conversion
|
| 138 |
|
| 139 |
Converted with [`litert-torch`](https://github.com/google-ai-edge/litert) via its
|