Nexuss0781 commited on
Commit
99eca31
·
verified ·
1 Parent(s): 2971e2b

Add SmolLM2 135M Instruct F16 and Q6_K GGUF artifacts

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ SmolLM2-135M-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
37
+ SmolLM2-135M-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
CHECKSUMS.sha256 ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ fd659c45cacd2213d6b8ffe002c261c193eed4cc0d2b1b2af836fba7b030177c SmolLM2-135M-Instruct-f16.gguf
2
+ e12aa5cace7cca9c6bd27a8eddb511d5140628ca5fe3f3f3d20fbbf5c458aeee SmolLM2-135M-Instruct-Q6_K.gguf
README.md ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: apache-2.0
5
+ library_name: llama.cpp
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - gguf
9
+ - llama.cpp
10
+ - smollm2
11
+ - text-generation
12
+ - conversational
13
+ base_model: HuggingFaceTB/SmolLM2-135M-Instruct
14
+ ---
15
+
16
+ # SmolLM2 135M Instruct — GGUF
17
+
18
+ This repository provides local CPU-oriented GGUF conversions of [`HuggingFaceTB/SmolLM2-135M-Instruct`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct) for use with `llama.cpp` and compatible runtimes. The original instruct model is a compact 135M-parameter model released under the Apache-2.0 license.
19
+
20
+ ## Available files
21
+
22
+ | File | Format | Intended use | File size |
23
+ |---|---|---|---:|
24
+ | `SmolLM2-135M-Instruct-f16.gguf` | F16 GGUF | Quality-preserving local inference | 258 MiB |
25
+ | `SmolLM2-135M-Instruct-Q6_K.gguf` | Q6_K GGUF | Lower-memory CPU inference | 132 MiB |
26
+
27
+ The F16 variant preserves the original converted weight precision. The Q6_K variant is provided for systems where memory usage is more important than retaining the F16 representation. Use the F16 file for the highest-fidelity local behavior.
28
+
29
+ ## Quick start with llama.cpp
30
+
31
+ ```bash
32
+ ./llama-cli \
33
+ --model SmolLM2-135M-Instruct-f16.gguf \
34
+ --conversation \
35
+ --n-gpu-layers 0
36
+ ```
37
+
38
+ For persistent local serving, start `llama-server` once and send OpenAI-compatible requests to the local endpoint:
39
+
40
+ ```bash
41
+ ./llama-server \
42
+ --model SmolLM2-135M-Instruct-f16.gguf \
43
+ --host 127.0.0.1 \
44
+ --port 8080 \
45
+ --ctx-size 2048 \
46
+ --n-gpu-layers 0
47
+ ```
48
+
49
+ ## Integrity verification
50
+
51
+ Verify downloaded files with:
52
+
53
+ ```bash
54
+ sha256sum -c CHECKSUMS.sha256
55
+ ```
56
+
57
+ ## Conversion details
58
+
59
+ The model was converted from the official Hugging Face checkpoint with the official `llama.cpp` Hugging Face-to-GGUF converter. The F16 GGUF output was retained as the quality-preserving variant, then quantized with `llama-quantize` to produce the Q6_K variant. See `conversion-metadata.json` for the artifact metadata.
60
+
61
+ ## License and attribution
62
+
63
+ These GGUF artifacts are derived from [HuggingFaceTB/SmolLM2-135M-Instruct](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct). The upstream model card identifies the model license as **Apache-2.0**. Retain upstream attribution and consult the original model card for limitations, training details, and citation information.
SmolLM2-135M-Instruct-Q6_K.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e12aa5cace7cca9c6bd27a8eddb511d5140628ca5fe3f3f3d20fbbf5c458aeee
3
+ size 138382912
SmolLM2-135M-Instruct-f16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fd659c45cacd2213d6b8ffe002c261c193eed4cc0d2b1b2af836fba7b030177c
3
+ size 270885952
conversion-metadata.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "base_model": "HuggingFaceTB/SmolLM2-135M-Instruct",
3
+ "base_model_url": "https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct",
4
+ "license": "Apache-2.0",
5
+ "converter": "llama.cpp convert_hf_to_gguf.py",
6
+ "quantizer": "llama.cpp llama-quantize",
7
+ "artifacts": [
8
+ {
9
+ "file": "SmolLM2-135M-Instruct-f16.gguf",
10
+ "quantization": "F16",
11
+ "bytes": 270885952,
12
+ "sha256": "fd659c45cacd2213d6b8ffe002c261c193eed4cc0d2b1b2af836fba7b030177c"
13
+ },
14
+ {
15
+ "file": "SmolLM2-135M-Instruct-Q6_K.gguf",
16
+ "quantization": "Q6_K",
17
+ "bytes": 138382912,
18
+ "sha256": "e12aa5cace7cca9c6bd27a8eddb511d5140628ca5fe3f3f3d20fbbf5c458aeee"
19
+ }
20
+ ]
21
+ }
22
+