3060 rtx 8 gb vram here ... share your preset !

#41
by oytaub - opened

i have poor performance like 1.5 tok/s to 2 tok/s with low context, i dont see practical use against a 3.6 35b where i reach 30 tok /s with 65 k token
what s your performance ?

1.5–2 tok/s sounds unusually low for Bonsai 2 27B.

Are you sure you're using the PrismML llama.cpp fork built with CUDA (-DGGML_CUDA=ON), and that the model is actually GPU-offloaded (-ngl 99/999)?

There are reports of ~34 tok/s on an RTX 3060 12GB with Bonsai 2 27B PQ2_0, so I would first check the backend/offload path rather than assuming this is the expected performance.

Since yours is an 8GB 3060, PQ2_0 may also be very tight on VRAM once KV/cache and runtime buffers are included. PTQ1_0 might be worth testing as well.

Could you post the startup log, exact GGUF, command line, and VRAM usage from nvidia-smi?

how can you run a 3.6 35b in a 3060 8gb vram?

1.5–2 tok/s sounds unusually low for Bonsai 2 27B.

Are you sure you're using the PrismML llama.cpp fork built with CUDA (-DGGML_CUDA=ON), and that the model is actually GPU-offloaded (-ngl 99/999)?

There are reports of ~34 tok/s on an RTX 3060 12GB with Bonsai 2 27B PQ2_0, so I would first check the backend/offload path rather than assuming this is the expected performance.

Since yours is an 8GB 3060, PQ2_0 may also be very tight on VRAM once KV/cache and runtime buffers are included. PTQ1_0 might be worth testing as well.

Could you post the startup log, exact GGUF, command line, and VRAM usage from nvidia-smi?

i tried both and yes i did compile the prism ml fork with cuda that s why i m on wondering why is it so slow ! regarding the 35b a3b it is Q4_K_M with kv cache q4_0 and a lots of preset optimizations

got my answer : mmproj was finishing filling the ram, removing it with --no-mmproj makes me reach 32 tok / s in generation .

Sign up or log in to comment