Qwen3-8B-BPDQ-2bit / demo.ipynb

Commit History

Dequant-free Apple matmul; copy-free GQA decode attention
655386b
verified

gitarist commited on

Faster Apple decode; 8-bit lm_head on Apple
f72bfc4
verified

gitarist commited on

L2 warm stream for small-batch decode
1d1ea3b
verified

gitarist commited on

Tensor-core CUDA kernel; packed 8-bit lm_head
fbf06ae
verified

gitarist commited on

Upload demo.ipynb with huggingface_hub
87a5624
verified

gitarist commited on

Upload demo.ipynb with huggingface_hub
9c182e0
verified

gitarist commited on

Upload demo.ipynb with huggingface_hub
0c8f3e5
verified

gitarist commited on

Upload demo.ipynb with huggingface_hub
8c39ec7
verified

gitarist commited on

Upload demo.ipynb with huggingface_hub
2ff7e9d
verified

gitarist commited on

Upload demo.ipynb with huggingface_hub
eae240e
verified

gitarist commited on