16gb card owners: post your numbers

#4
by deucebucket - opened

this build is 11.96gb, which should leave room to run fully on 16gb cards (4060 Ti 16GB, 4070 Ti SUPER, A4000, etc). on a 24gb card it peaks around 15.1gb serving 131k context with q8 kv cache.

i only have a 3090, so i cant validate the 16gb story myself. if you have a 16gb card and run this, post your numbers here:

  • your card + driver
  • launch line you used (a starting point: llama-server -m Qwen3.6-35B-A3B-Cerebellum-v3-Q3_K_M.gguf -ngl 99 -c 32768 -fa on --cache-type-k q8_0 --cache-type-v q8_0 -np 1 --jinja)
  • biggest context that boots AND fills without OOM
  • tok/s you see
  • anything weird

what works and what breaks is equally useful. club-3090 (a community for serving recipes on consumer cards) recorded this build's 24gb numbers and is interested in the 16gb validation story too: https://github.com/noonghunna/club-3090/pull/393

Should be doable in Google Collab with the free T4...

Should be doable in Google Collab with the free T4...

thanks for this, it also opened the door to other fronts like kaggle once i started looking into it. First i need to fix my benchmark scripts, its actually looking like the score were on the low end as 27b is showing closer to 90% on humeval+ so before i do that, i need to get that lined out. Always some sort of indentation issue tripping them up lol

Sign up or log in to comment