Add RTX 3080 (10 GB) to the cross-platform throughput table

#82
by Siphion - opened

Adds a consumer Ampere card to the throughput table, which currently has no Ampere GPU
other than the A100.

Measured with llama-bench from the official prism-b10743-adfffbe Windows CUDA 12.4
release, -ngl 99 -fa on, separate runs for -p 512 -n 0 and -p 0 -n 128, 5
repetitions each, batch size 1, depth 0, no vision tower. Energy is board power from
nvidia-smi (power.draw), sampled every 100 ms and averaged over the samples where GPU
utilisation was at least 80%, so that model loading does not dilute it; J/tok is that
average divided by TG128 throughput, which matches how the existing rows relate power to
token generation.

Setup: RTX 3080 10 GB, driver 610.74, Windows 10 (WDDM). The card also drives the
desktop display, which keeps about 1.3 GB of VRAM and a few watts busy; both packs still
fit entirely in VRAM.

The TG128 figures agree with an independent reproduction on the same card in
PrismML-Eng/llama.cpp#283 (61.38 for PQ2_0 and 52.18 for PTQ1_0, on Linux with CUDA
12.8), and with an earlier run of mine on this machine (62.88 and 52.71).

Two observations that fit the notes already in the card: PQ2_0 is faster than PTQ1_0 in
decode on this Ampere card, as the "Native low-bit kernels" note anticipates; and the
energy per token is higher than on the Ada and Blackwell cards, because the board stays
close to its 320 W limit even during bandwidth-bound decode.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment