在4060笔记本,8G下,优化到45Token/S,已经达到大厂的Token接口速度,彻底实现Token自由,大家再也不需要去大厂订阅。

#39
by SuperLogic - opened

Ternary-Bonsai-2 27B 在 RTX 4060 Laptop 上做到 45 t/s
三段加速链(纯解码,单卡 8GB):
基线(原始 fork):
22 t/s

  • PTQ1_0 专用 mat-vec 内核:~28 t/s(1.3×)
  • MTP 投机解码(无损):~45 t/s(再 ~1.6×)
    干了哪三件事:
    专用三值量化解码内核(pr-ptq1-mmv commit 1+2)
    Planar-Transposed 激活布局 + 为 PTQ1_0(1.75 bpw 三值)定制的 2D mat-vec 内核。
    把 K=5120 投影的 lane 利用率从 ~31% 拉满到 128 线程,直接吃满带宽 → 主干解码提速。
    MTP 投机解码(draft-mtp)
    用模型自带的 nextn(MTP)头做草稿,一次猜 2 个 token 再由主干批量验证,把访存受限的逐 token 解码摊薄 → 再提速。
    BATCH_INVARIANT 无损开关(commit 4+5)
    强制让"单独解码"和"投机 batch 里被验证"的 token logits 逐位一致(1–4 列),使 MTP 严格无损、贪心输出不变。
    代价:长上下文(思考模式)注意力走向量内核更快不了,所以高思考档会掉到 ~35 t/s——这是"以速度换无损"的固定成本。
    一句话总结: 内核优化解决"主干本身慢",MTP 解决"访存瓶颈摊不平",BATCH_INVARIANT 保证"提速不改结果"。三者叠加 22 → 45。

Very interesting result — could you clarify exactly which model/MTP/kernel combination produced the ~1.6x MTP speedup on PTQ1_0?

The reason I’m asking is that the PTQ1_0 + MTP results I’ve seen so far showed little or no speedup, apparently because the cost of PTQ1 unpacking / batch verification offsets much of the speculative decoding gain.

So I’m especially curious about:

  • Which PTQ1_0 GGUF/model was used?
  • Which MTP/nextn head was used, and where did it come from?
  • Which llama.cpp branch/commit contains the PTQ1 kernel and MTP implementation?
  • Is the ~28 → ~45 tok/s result measured with the same model, context, and generation settings?
  • Do you have the MTP acceptance rate or benchmark command?

If your PTQ1 kernel changes are what make MTP scale from ~28 to ~45 tok/s, that would be particularly interesting, since it seems to address the exact bottleneck reported in the other PTQ1 + MTP experiments.

分享下启动参数呗

我们这套 45 t/s 的技术来源

  1. 内核优化 + 50 t/s 配方(核心来源)
    仓库:sudoingX/bonsai2-small-gpu
    地址:https://github.com/sudoingX/bonsai2-small-gpu
    取的东西:pr-ptq1-mmv 分支的 5 个 kernel patch
    commit 1–2:PT 激活布局 + PTQ1_0 专用 mat-vec 内核(主干提速 ~1.4×)
    commit 4–5:GGML_CUDA_BATCH_INVARIANT 无损开关(让 MTP 严格不改结果)
    基线提交:prism/prism 9a9394a89
  2. MTP 投机解码
    上游实现:llama.cpp --spec-type draft-mtp
    地址:https://github.com/ggml-org/llama.cpp
    草稿头来源:Qwen 模型自带的 nextn MTP 头,用我们自研工具 tools/mtp-graft 嫁接进 Bonsai GGUF
  3. 底座 / 模型
    主干模型:Ternary-Bonsai-2-27B(PrismML,PTQ1_0 三值量化,1.75 bpw)
    量化内核 PTQ1_0:合入在 ggml/llama.cpp 侧(PrismML 血统分叉)
    我们用的 fork:forks/llamAmpere(PrismML 同血统、Ampere/消费卡侧分叉)
  4. 我们自己的活儿(非开源)
    把上面 commit 1/2/4/5 手工移植进 llamAmpere fork(它和补丁上游有分叉,不能 git apply)
    便携 Windows CUDA 构建链 + llmserver1.bat + 一键回滚
    一句话归因: 内核和无损开关来自 sudoingX/bonsai2-small-gpu,MTP/底座来自 ggml-org/llama.cpp 与 PrismML Bonsai/Qwen;移植、构建、集成是我们做的。 mtp草稿要注意那个旋转位,AI开发会知道, 参数我在测试的配置:@echo off
    cd /d %~dp0

rem ============ 模型包(weights / runtime / calib 都在这一个目录里)============
set "MODEL_NAME=Ternary-Bonsai-2-27B-SelfMTP"
set "PKG=models\Ternary-Bonsai-2-27B-SelfMTP"
set "MODEL_FILE=Ternary-Bonsai-2-27B-PTQ1_0-mtp.gguf"
rem KVMC_FILE:校准出来的均值中心化 bias(按模型各校准一份,换模型必须重跑)
rem 本文件用的是 Base trunk 那一份 bias——trunk 权重与 MTP 包逐字节相同
set "KVMC_FILE=kv-bias.Ternary-Bonsai-2-27B-PTQ1_0-mtp.gguf"

rem ============ 服务地址(自建 MTP 用 18082,和 18080 的在跑实例分开)============
set "HOST=127.0.0.1"
set "PORT=18080"

rem ============ 采样(客户端可用请求体覆盖)============
set "TEMP=0.6"
set "TOP_P=0.95"
set "TOP_K=20"

rem ============ 上下文与批 ============
rem CTX 有效范围 8192, 16384, 24576 ~ 32768, 65536,越大越吃显存
set "CTX=24576"
set "BATCH=4096"
set "UBATCH=1024"
set "THREADS=8"
set "NGL=99"

rem ============ KV 缓存 ============
rem CTK 用 q4_0 是 --kv-mean-center 的硬性前提(其他类型会直接启动失败)
set "CTK=q4_0"
set "CTV=q4_0"

rem ============ MTP 草稿(先用 1 试,接受率高再往上加)============
set "DRAFT=2"

rem ============ 思考 ============
rem REA 可选值 off / on / auto,客户端可用 chat_template_kwargs.enable_thinking 覆盖
set "REA=off"
rem EFFORT 可选值 xhigh / medium / low,仅思考开启时生效,客户端可覆盖
set "EFFORT=low"

rem ============ 路径拼装(一般不用改)============
set "MODEL=%PKG%\weights%MODEL_FILE%"
set "KVMC=%PKG%\calib%KVMC_FILE%"
set "BIN=%~dp0%PKG%\runtime\llama-server.exe"
set "CUDA=F:\super\tools\cuda"

set "PATH=%CUDA%\bin;%PATH%"
set "CUDA_PATH=%CUDA%"
set LLAMA_SPEC_PQ=0
set GGML_CUDA_GRAPH_OPT=1
set GGML_CUDA_PDL=1
rem LLAMA_ATTN_ROT_DISABLE=1:锁死 K 缓存旋转为关闭,必须与校准时一致(bias 只在其测量基准下有效)
set LLAMA_ATTN_ROT_DISABLE=1

title %MODEL_NAME% :%PORT%

"%BIN%" -m "%MODEL%" --host %HOST% --port %PORT% --parallel 1 -rea %REA% ^
-c %CTX% -b %BATCH% -ub %UBATCH% -t %THREADS% -ngl %NGL% -fa on -ctk %CTK% -ctv %CTV% ^
--temp %TEMP% --top-p %TOP_P% --top-k %TOP_K% --reasoning-effort %EFFORT% ^
--no-reasoning-preserve ^
--kv-mean-center "%KVMC%" ^
--spec-type draft-mtp --spec-draft-n-max %DRAFT% --spec-draft-p-min 0

要达到这个近50Token的最优解,主要是那三个优化叠加,然后更小的上下文+关闭思考,"DRAFT=2", set "CTK=q4_0" set "CTV=q4_0"开思考会有降低几个Token

Thanks a lot for the detailed follow-up and for pointing to the actual source. This was very useful.

After looking through the code and the benchmark notes, I think I understand the result much better now.

At first I had assumed that most of the speedup was coming from MTP itself, but it looks like a very significant part of the improvement actually comes from the PTQ1_0 kernel work — especially fixing the poor GPU work distribution / idle threads and improving the activation loading layout.

In other words, the interesting part seems to be not only:

PTQ1_0 + MTP

but more like:

PTQ1_0
→ much better GPU utilization / activation loading
→ then MTP on top of that.

That also explains why the result is so different from some earlier PTQ1_0 + MTP experiments where MTP alone gave very little additional speedup.

I also noticed that the PTQ1 kernel improvements have now been submitted back to the PrismML llama.cpp fork as an upstream PR. That is great news.

So for people who do not want to wait for the changes to be merged officially, it seems that following the setup and patches you described here is currently a very good way to try the improved PTQ1_0 path themselves.

Thanks again for sharing the implementation details — the source and benchmark data made the result much easier to understand.

Can this method be applied to the PTQ_2 model?

Sign up or log in to comment