−2 logit bias on PTQ1_0, RTX 5060 Laptop: MATH-500 44/50 → 43/50

#83
by 7dollarbooks - opened

A LocalLLaMA post reported that a −2 logit bias on "wait", "maybe", and "perhaps" made Qwen3.5-4B more accurate and shorter on 50 MATH-500 questions. I tried the same idea on Ternary-Bonsai-2-27B-PTQ1_0 (5.53 GiB). It did not carry over here.

Run Correct Avg tokens tok/s Truncations
A baseline 44/50 845.2 29.27 2
B −2 bias 43/50 872.3 29.28 2

Setup: RTX 5060 Laptop 8 GB, Windows, driver 617.14 (CUDA 13.4), Prism llama.cpp prism-b10743-adfffbe (build adfffbe41, CUDA 13.3 Windows zip). Both runs: temp 0, top-k 40, top-p 0.95, min-p 0, seed 42, -c 4096, --reasoning-budget 2048, -n 3072, same 50 questions in the same order. Run B adds --logit-bias −2 on token ids 11158, 3655, 13784, 35542, 6970, 20734, 63068, 8106, 30442 (wait/maybe/perhaps, lowercase, leading space, and capitalized).

47 answers matched. One truncated miss in A finished correctly in B, and two A successes became misses in B, one by hitting the cap. Speed didn't change.

This is one deterministic pair at temperature 0, not Prism's recommended sampler and not a confidence interval. Script, question list, per-question scores, all 100 raw replies, and server logs: https://github.com/7dollarbooks/bonsai2-logit-bias-test

Run by Joseph Murray Adams.

Sign up or log in to comment