gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090
Image-Text-to-Text • 15B • Updated • 583k • 151
One 32 GB GPU: the NVFP4 checkpoint, its matched DSpark speculative drafter, and variants.
Note Start here. Full 262K context on 32 GB; 81.6 tok/s alone, 155.8 with the drafter.
Note Matched DSpark drafter (v2), 1.41 GB. 1.91x decode, byte-identical outputs, +14.2% agentic acceptance.
Note Same checkpoint minus the native MTP head: 0.85 GB smaller download, for SGLang-only setups.