Qwen3.8-Flash-Next-PLE-quant / worker_image_quant.py

Commit History

Overlay: drop a checkpoint's own table tensors when a sidecar is attached, including the monolithic weight/weight_scale form — lets FP8-table checkpoints (nvidia, official FP8) serve on one card
0a85198
verified

jagat-primitive-org commited on

Overlay: build the n-gram table parameter on the meta device when VLLM_PLE_QUANT_DIR is set, and swap the stub Parameter instead of set_data — removes the 102 GB virtual reservation that the kernel's overcommit heuristic refuses on hosts with less RAM+swap than the table (field report, 64 GB host); validated under an emulated 67/99 GiB commit limit, sanity PASS, tool-calling 77.0 (n=3), 81 tok/s c1
cf50620
verified

jagat-primitive-org commited on

Upload worker_image_quant.py with huggingface_hub
a815cac
verified

jagat-primitive-org commited on

Upload worker_image_quant.py with huggingface_hub
c7550ff
verified

jagat-primitive-org commited on