# Dense-K6 CUDA source and uniform-K6 batch compatibility repair The first dense-K6 `shared_gate` work unit exited at the production closure guard with: ```text RuntimeError: production SQG numerical tensor left CUDA before closure: {'prepared_source': 'cpu'} ``` `encode_dense_k6_work_unit.py` deliberately requests the official BF16 dense weights with `device="cpu"`, while the production codec requires every numerical tensor to remain CUDA-resident through closure. The paid-node runtime therefore uses the image's pre-existing system `sitecustomize.py` startup hook to conditionally load the sealed runtime overlay when `B300_K6_FORCE_BF16_CUDA=1`. It is deliberately not added to `PYTHONPATH`, because ExLlamaV3 provenance seals the literal path string. Other Python processes do nothing at startup because the flag is absent. Inside campaign persistent workers, the shim redirects only PyTorch `safetensors.safe_open` calls beneath the exact official BF16 source root while the active in-process entrypoint is exactly `encode_dense_k6_work_unit.py`. Routed/profile stages retain ordinary CPU source reads. Matching dense-K6 CPU requests are redirected to the process-visible `cuda:0`. This overlay does not alter the sealed encoder source tree or quantization arithmetic. Before resuming, a smoke load proved that the official layer-3 shared-gate tensor was BF16 on `cuda:0` with shape `[2048, 6144]`, and the existing full-W4A8 source binding was reverified unchanged. The dense shared-down encoder has caller-owned input and output residual profiles. It therefore sends one constant-rate candidate through KQuant's batched LDLQ entry point so the global scale remains exactly `1.0`. That generic traversal and the SM103 encoder support K2 through K6, but its older mixed-rate validator admitted only K2/K3/K4. The runtime hook now extends that validator solely for a uniform all-K5 or all-K6 member. Heterogeneous maps containing K5/K6 remain rejected, and routed heterogeneous K2/K3/K4 behavior is unchanged. The hook is installed while the sealed backend module is loaded; the KQuant checkout and its recorded source hash remain byte-identical. ## Vast stop/start mount-device renumbering After the preserved instance was externally stopped and restarted, all 277 preflight-bound BF16 shards retained identical inode, byte length, and nanosecond modification time, while the container mount device changed from `62` to `64`. A full SHA-256 of the first rejected shard, `model-00178-of-00282.safetensors`, still matched its sealed digest `7d7fd51a5b5c0411ee0d898b54554163d3284f6f1c2fd27dda8ec3718bfcb064`. The runtime hook now wraps `src.glm52_bf16_source._sealed_identity` without editing that sealed module. It accepts a shard beneath the exact official BF16 root only when the sole identity difference is `st_dev`. Inode, size, and mtime must remain byte-for-byte equal to the preflight record; any other drift continues through the original fail-closed guard. The original preflight files and their hashes are not rewritten.