noctrex commited on
Commit
9e26c3c
·
1 Parent(s): 24c068c

Add files using upload-large-folder tool

Browse files
Files changed (5) hide show
  1. .gitattributes +6 -0
  2. README.md +24 -0
  3. mmproj-BF16.gguf +3 -0
  4. mmproj-F16.gguf +3 -0
  5. mmproj-F32.gguf +3 -0
.gitattributes CHANGED
@@ -33,3 +33,9 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ gemma-4-26B-A4B-it-uncensored-heretic-MXFP4_MOE_F16.gguf filter=lfs diff=lfs merge=lfs -text
37
+ mmproj-F32.gguf filter=lfs diff=lfs merge=lfs -text
38
+ mmproj-F16.gguf filter=lfs diff=lfs merge=lfs -text
39
+ mmproj-BF16.gguf filter=lfs diff=lfs merge=lfs -text
40
+ gemma-4-26B-A4B-it-uncensored-heretic-MXFP4_MOE.gguf filter=lfs diff=lfs merge=lfs -text
41
+ gemma-4-26B-A4B-it-uncensored-heretic-MXFP4_MOE_BF16.gguf filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ pipeline_tag: image-text-to-text
3
+ base_model:
4
+ - llmfan46/gemma-4-26B-A4B-it-uncensored-heretic
5
+ ---
6
+ These are **MXFP4** quantizations of the model [gemma-4-26B-A4B-it-uncensored-heretic](https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-uncensored-heretic)
7
+
8
+ ## Quick Start
9
+ 1. Download the latest release of **llama.cpp**.
10
+ 2. Download your preferred model variant from below.
11
+ 3. For the `mmproj` file, it is recommended to use the **F32 version** for the best visual processing results. F32 > BF16 > F16
12
+
13
+ ## Which version should I choose?
14
+ All variants use **MXFP4** for the MoE (Mixture of Experts) weights to keep the model efficient. The difference lies in how the remaining tensors are handled:
15
+
16
+ | Variant | Quality | Performance | Size | Recommendation |
17
+ | :--- | :--- | :--- | ---: | :--- |
18
+ | **BF16** | ⭐⭐⭐ | Variable* | 20.55GiB | Best for maximum accuracy; original unquantized weights. |
19
+ | **F16** | ⭐⭐ | Fast | 20.55GiB | Great alternative if BF16 is slow on your hardware. |
20
+ | **Q8** | ⭐ | Fastest | 18.88GiB | Balanced performance and memory usage. |
21
+
22
+ *\*Note: On some older architectures, BF16 may be slower than F16. Check that your GPU supports native BF16 *
23
+
24
+ ### Read [unloth's gemma guide](https://unsloth.ai/docs/models/gemma-4) for the optional parameters for the model.
mmproj-BF16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8196e6dda1446b547d69cf30c0bf7a12a69eaf399513ab41ff065796aa975a61
3
+ size 1194828096
mmproj-F16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8ef0d43ca26d7a553b579e28c827e6c86bde96cf0b7b95bc0441c1edde95057f
3
+ size 1193058848
mmproj-F32.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8e7804869cad71cac98f4be63fb0ccc0e04703d0365b82edcbe0a79397a4c4d0
3
+ size 2291200320