MOSS-TTS-Nano-100M β GGUF (Q8_0) for RapidSpeech.cpp / ggml-CUDA
Q8_0 GGUF conversion of OpenMOSS-Team/MOSS-TTS-Nano-100M for on-device inference on Jetson Nano gen1 (Maxwell sm_53) via ggml-CUDA.
moss_nano_full.ggufβ merged AR model + codec decoder, Q8_0 (~138 MB from 440 MB fp32).- AR: GPT-2 12-layer global + 1-layer local decoder + 16 audio codebook heads (interleaved RoPE, gelu_new).
- Codec: MOSS-Audio-Tokenizer-Nano decoder + 16-way RVQ (weight_norm reconstructed) β 48 kHz.
moss_nano.ggufβ AR model only.moss_codec.ggufβ codec only.
Verified (torch-free, vs the deployed ONNX): global transformer prefill+decode MSE 1.4e-6, local decoder 14/16 argmax (Q8), codec round-trip corr 0.95β0.998. Performance (Jetson Nano GPU, ggml-CUDA): Q8 RTF 0.86 (full 12L) / 0.50 (4L student); ~0.35 with the custom sm_53 matvec kernel. RTF floor ~0.12.
Runtime: vieenrose/RapidSpeech.cpp (branch jetson-nano-gen1),
arch moss_tts_nano. Converters in scripts/convert_moss_*_to_gguf.py.
- Downloads last month
- 253
Hardware compatibility
Log In to add your hardware
We're not able to determine the quantization variants.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support
Model tree for Luigi/MOSS-TTS-Nano-100M-GGUF
Base model
OpenMOSS-Team/MOSS-TTS-Nano-100M