kingjones777 commited on
Commit
c5fee31
·
verified ·
1 Parent(s): 2c49eab

docs: point runtime instructions at the kingjones30/ROCmFPX fork (verified build)

Browse files
Files changed (1) hide show
  1. README.md +18 -0
README.md CHANGED
@@ -19,6 +19,24 @@ tags:
19
  - llama.cpp
20
  ---
21
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22
  # North-Mini-Code-1.0 — ROCmFP4 STRIX (GGUF) — AMD Ryzen AI Max+ 395 / Strix Halo / gfx1151
23
 
24
  This is a `Q4_0_ROCMFP4_STRIX` quant of [CohereLabs/North-Mini-Code-1.0](https://huggingface.co/CohereLabs/North-Mini-Code-1.0), built and tested on an AMD Ryzen AI Max+ 395 (Strix Halo, gfx1151, 128 GB unified memory) running ROCm 7.2.4.
 
19
  - llama.cpp
20
  ---
21
 
22
+ > ### 🔧 Runtime: build the ROCmFPX fork below
23
+ > Stock `llama.cpp` will not load this file. You need **both** the **`cohere2moe`** architecture
24
+ > **and** the ROCmFP4 tensor types in one tree. Upstream
25
+ > [`charlie12345/ROCmFPX`](https://github.com/charlie12345/ROCmFPX) has the ROCmFP4 types but
26
+ > not `cohere2moe`. Our fork has both:
27
+ >
28
+ > **[`kingjones30/ROCmFPX`](https://github.com/kingjones30/ROCmFPX)** — a fork of `charlie12345/ROCmFPX`, branch `main`.
29
+ >
30
+ > ```bash
31
+ > git clone https://github.com/kingjones30/ROCmFPX.git
32
+ > cd ROCmFPX
33
+ > cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release
34
+ > cmake --build build --target llama-server llama-quantize -j$(nproc)
35
+ > ```
36
+ >
37
+ > Verified 2026-08-27 on gfx1151: clean clone → **0 build errors** → `llama-server` loads a
38
+ > `cohere2moe` ROCmFP4 GGUF from this family and generates coherent text.
39
+
40
  # North-Mini-Code-1.0 — ROCmFP4 STRIX (GGUF) — AMD Ryzen AI Max+ 395 / Strix Halo / gfx1151
41
 
42
  This is a `Q4_0_ROCMFP4_STRIX` quant of [CohereLabs/North-Mini-Code-1.0](https://huggingface.co/CohereLabs/North-Mini-Code-1.0), built and tested on an AMD Ryzen AI Max+ 395 (Strix Halo, gfx1151, 128 GB unified memory) running ROCm 7.2.4.