How does the 2080 Ti 22G perform with this model?

#49
by coresen - opened

I'm experiencing the 2080 Ti 22G in this model, and would appreciate insights on its performance.

Specifically, I'm looking for:

VRAM usage – Will this model fit within 22 GB of VRAM, and what quantization/precision is recommended?
Inference speed – What performance (e.g. tokens/sec or images/sec) can be expected on this GPU?
Compatibility – Are there any known issues or recommended settings (batch size, attention implementation, etc.) for the 2080 Ti series?

If anyone has run this model on a 2080 Ti 22G before, please share your results or tips in the comments.

For reference, I tested Bonsai-2-27B on a much smaller RTX 2060 SUPER 8GB, so my results may be useful as a baseline for your 2080 Ti 22GB.

My full test and results are here:
https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/discussions/38

The main configuration I used was:

  • RTX 2060 SUPER 8GB (Turing, compute capability 7.5)
  • CUDA 12.4
  • PrismML llama.cpp fork
  • full GPU offload (-ngl 99)
  • Flash Attention enabled
  • 32K context
  • K cache: Q5_1
  • V cache: Q4_0
  • --parallel 1

With PTQ1_0 (1.75 bpw), the whole model + 32K context fit comfortably in 8GB VRAM.

Observed VRAM usage was about 6.2–6.5GB.

Performance on the 2060 SUPER was:

  • pp512: ~178 tok/s
  • pp2048: ~100 tok/s
  • tg128: ~16.8 tok/s
  • a real long generation (1747-token prompt + 2975 generated tokens): 11.0 tok/s average

Decode speed gradually dropped from about 13.5 tok/s near the beginning to about 9.4 tok/s after ~3000 generated tokens as the active context grew.

I also tested PQ2_0 on the same 8GB GPU. It was much tighter on memory, around 7.6 GiB / 8 GiB, but still fit with the same 32K context and quantized KV cache.

Interestingly, PQ2_0 was faster:

  • prompt processing: ~97 tok/s for the test prompt
  • long-generation average: 13.7 tok/s
  • roughly ~20% faster decoding than PTQ1_0 in my tests

So from a VRAM capacity standpoint, 22GB should give you a very large amount of headroom for either PTQ1_0 or PQ2_0, even with a fairly large context.

The 2080 Ti is also Turing / compute capability 7.5, like the 2060 SUPER I tested, so I would expect the same PrismML CUDA path to be applicable.

I don't have a direct 2080 Ti 22GB benchmark, so I don't want to guess an exact tokens/sec figure. But given that the model already runs fully on an 8GB 2060 SUPER, your card should certainly be an interesting one to benchmark.

If you try it, I'd be very interested to see your llama-bench results, especially pp512, pp2048, and tg128, since they would make a nice comparison with the 2060 SUPER results.

@mktnhr Thank you for your warm reply. Here's my simple test results for share with you.

Environment: Ubuntu 26.04 LTS (GNU/Linux 7.0.0-31-generic x86_64) - RTX 2080Ti 22G - NVIDIA-SMI 595.91.07 - Driver Version: 595.91.07 - CUDA Version: 13.2

=========================== PTQ1_0 ================================

sudo ./llama-bench -m Qwen3.8-27B-abliterated-Ternary-Bonsai-PTQ1_0-Huihui.gguf -ngl 99 -fa on -t 4 -p 512,2048 -n 128 -r 3
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 22001 MiB):
Device 0: NVIDIA GeForce RTX 2080 Ti, compute capability 7.5, VMM: yes, VRAM: 22001 MiB

model size params backend ngl threads fa test t/s
qwen35 27B PTQ1_0 - 1.75 bpw ternary (group 128) 6.14 GiB 26.90 B CUDA 99 4 1 pp512 477.62 ± 3.29
qwen35 27B PTQ1_0 - 1.75 bpw ternary (group 128) 6.14 GiB 26.90 B CUDA 99 4 1 pp2048 474.02 ± 0.55
qwen35 27B PTQ1_0 - 1.75 bpw ternary (group 128) 6.14 GiB 26.90 B CUDA 99 4 1 tg128 34.23 ± 0.04

build: 3ae4f5108 (10725)

=========================== PQ2_0 ================================

sudo ./llama-bench -m Qwen3.8-27B-abliterated-Ternary-Bonsai-PQ2_0-Huihui.gguf -ngl 99 -fa on -t 4 -p 512,2048 -n 128 -r 3
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 22001 MiB):
Device 0: NVIDIA GeForce RTX 2080 Ti, compute capability 7.5, VMM: yes, VRAM: 22001 MiB

model size params backend ngl threads fa test t/s
qwen35 27B PQ2_0 - 2.13 bpw (group 128) 7.17 GiB 26.90 B CUDA 99 4 1 pp512 724.96 ± 7.07
qwen35 27B PQ2_0 - 2.13 bpw (group 128) 7.17 GiB 26.90 B CUDA 99 4 1 pp2048 721.74 ± 1.09
qwen35 27B PQ2_0 - 2.13 bpw (group 128) 7.17 GiB 26.90 B CUDA 99 4 1 tg128 38.61 ± 0.02

build: 3ae4f5108 (10725)

As you can see, the test results match your expectations. PQ2_0 faster than PTQ1_0 on compute capability 7.5.

Sign up or log in to comment