New discussion

memory size context

1
#35 opened about 1 month ago by
nemilos

decoding with DSpark is slower on llama

3
#31 opened about 2 months ago by
SlavikF

Release imatrix

2
#30 opened about 2 months ago by
KeinNiemand

dspark

2
#29 opened about 2 months ago by
ludashi666

How did you quantize without imatrix?

#27 opened about 2 months ago by
Wladastic

CUDA ERROR for omp

2
#23 opened about 2 months ago by
Butterfly-314

Why do the K quants use I quants?

#21 opened about 2 months ago by
LagOps

pp sucks

8
#14 opened about 2 months ago by
gopi87

3bit working, 4bit not

👍 4
3
#12 opened about 2 months ago by
tasticleeze

Awww yiss

🚀🔥 26
10
#1 opened about 2 months ago by
CurbStomper