Runs at ~90 t/s on an Intel Arc B580 (128K context)

#1
by torchit1 - opened

Thanks for the mtp-lean file. It's the one my Intel Arc guide points people to.

On a 12 GB B580 at 128K context it writes new code at about 90 t/s (MTP plus n-gram drafts), edits pasted code at 250 to 370 t/s, and still does 44 t/s with 115K tokens of history. The branch runs the ternary weights and attention on the XMX units: https://github.com/Torchit1/llama.cpp/tree/arc-b580 (guide in docs/bonsai-arc-b580.md, Windows zip under Releases).

In case it's useful: restricting the MTP head's output to the ~32K most frequent tokens (the full model still verifies) made drafting about 5% faster here with identical output.

Sign up or log in to comment