MOSS-TTS-Nano-100M β€” GGUF (Q8_0) for RapidSpeech.cpp / ggml-CUDA

Q8_0 GGUF conversion of OpenMOSS-Team/MOSS-TTS-Nano-100M for on-device inference on Jetson Nano gen1 (Maxwell sm_53) via ggml-CUDA.

  • moss_nano_full.gguf β€” merged AR model + codec decoder, Q8_0 (~138 MB from 440 MB fp32).
    • AR: GPT-2 12-layer global + 1-layer local decoder + 16 audio codebook heads (interleaved RoPE, gelu_new).
    • Codec: MOSS-Audio-Tokenizer-Nano decoder + 16-way RVQ (weight_norm reconstructed) β†’ 48 kHz.
  • moss_nano.gguf β€” AR model only. moss_codec.gguf β€” codec only.

Verified (torch-free, vs the deployed ONNX): global transformer prefill+decode MSE 1.4e-6, local decoder 14/16 argmax (Q8), codec round-trip corr 0.95–0.998. Performance (Jetson Nano GPU, ggml-CUDA): Q8 RTF 0.86 (full 12L) / 0.50 (4L student); ~0.35 with the custom sm_53 matvec kernel. RTF floor ~0.12.

Runtime: vieenrose/RapidSpeech.cpp (branch jetson-nano-gen1), arch moss_tts_nano. Converters in scripts/convert_moss_*_to_gguf.py.

Downloads last month
253
GGUF
Model size
11.5M params
Architecture
moss_audio_tokenizer_nano
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Luigi/MOSS-TTS-Nano-100M-GGUF

Quantized
(2)
this model