Kyutai STT 2.6B English โ€” GGUF

GGUF conversion of kyutai/stt-2.6b-en for the CrispASR kyutai-stt backend.

High-quality English ASR from the Kyutai/Moshi team. Causal transformer LM conditioned on Mimi audio codec features โ€” streaming-capable architecture.

Architecture

  • Mimi encoder: SEANet CNN + transformer + RVQ (audio -> discrete tokens)
  • Causal LM: 16-layer transformer (2048-dim, RoPE, SwiGLU, RMSNorm)
  • Input: 16 kHz mono audio
  • Output: English text transcription

Files

File Size Description
kyutai-stt-2.6b-q4_k.gguf 1.5 GB Q4_K quantized (recommended)
kyutai-stt-2.6b-q8_0.gguf 2.8 GB Q8_0 quantized
kyutai-stt-2.6b-f16.gguf 5.1 GB F16 full precision
kyutai-stt-2.6b-ref.gguf 4.0 MB Diff harness reference

Usage

# Auto-download:
crispasr --backend kyutai-stt -m auto --auto-download -f audio.wav

# Local model:
crispasr --backend kyutai-stt -m kyutai-stt-2.6b-q4_k.gguf -f audio.wav -osrt

License

Apache 2.0 (same as the original model).

Credits

Downloads last month
672
GGUF
Model size
3B params
Architecture
kyutai-stt
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cstr/kyutai-stt-2.6b-en-GGUF

Quantized
(1)
this model