File size: 2,286 Bytes
b69d9d8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
From: opencoti
Subject: [PATCH 3/N] bypass upstream `params.n_gpu_layers < 0` assertions

Upstream llama.cpp pin 5e9c63546 has TWO `GGML_ASSERT(n_gpu_layers < 0)`
calls in common/arg.cpp:

1. At the `-ngl` option registration site (params.n_gpu_layers).
2. At the `-ngld` (draft) option registration site
   (params.speculative.n_gpu_layers).

Both fire on every `common_params_parser_init` call regardless of args,
making the patched binary unable to start in `--server`, `--cli`, or
`--chat` mode. Mozilla-Ocho v1.4.0 (pre-assertion) was unaffected.

Each assertion is a "display-format guard" for the next line, which
maps `n_gpu_layers == -1` to "auto" and anything else to "all" in the
help text. The fix is to either extend the format to handle non-
negative defaults, or to drop the assertion. We drop it — the
display "all" for a positive default is a minor cosmetic issue,
not a correctness problem.

This is an upstream-bug-bypass patch, not opencoti scope. To be
proposed upstream once we surface a minimal reproducer.

Bench: closes the F4 M3 bench path that requires booting the
patched binary in --server mode.
Milestone: F4 M3 — see docs/features/memory_embedder.md
Upstreaming: TBD — propose to llama.cpp once a minimal repro is
  in hand. Likely already fixed in newer upstream commits.

---
diff --git a/llama.cpp/common/arg.cpp b/llama.cpp/common/arg.cpp
--- a/llama.cpp/common/arg.cpp
+++ b/llama.cpp/common/arg.cpp
@@ -2362,7 +2362,10 @@ common_params_context common_params_parser_init(common_params & params, llama_ex
             }
         }
     ).set_env("LLAMA_ARG_N_CPU_MOE"));
-    GGML_ASSERT(params.n_gpu_layers < 0); // string_format would need to be extended for a default >= 0
+    // opencoti F4 M3 patch 0003 — see docs/decisions/0001-lazy-slot-context.md
+    // Upstream assertion fires on every common_params_parser_init invocation
+    // (display-format guard for the next line). Dropped — display says "all"
+    // for non-negative defaults, which is cosmetic, not a correctness bug.
     add_opt(common_arg(
         {"-ngl", "--gpu-layers", "--n-gpu-layers"}, "N",
         string_format("max. number of layers to store in VRAM, either an exact number, 'auto', or 'all' (default: %s)", params.n_gpu_layers == -1 ? "auto" : "all"),