Qwen3 ASR 0.6B β€” MLX 4-bit

MLX 4-bit quantized conversion of Qwen/Qwen3-ASR-0.6B for Apple Silicon inference.

Usage

Used by speech-swift Qwen3ASR module:

let model = try await Qwen3ASRModel.fromPretrained()
let text = model.transcribe(audio: samples, sampleRate: 16000)
audio transcribe audio.wav

Model Details

  • Architecture: Qwen3-ASR encoder-decoder (Whisper-style audio encoder + Qwen3 text decoder)
  • Parameters: 0.6B
  • Quantization: 4-bit (MLX)
  • Size: ~680 MB
  • Languages: Multilingual (EN, ZH, JA, KO, FR, DE, ES, and more)


Downloads last month
34,351
Safetensors
Model size
0.3B params
Tensor type
BF16
Β·
U32
Β·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for aufklarer/Qwen3-ASR-0.6B-MLX-4bit

Quantized
(41)
this model

Space using aufklarer/Qwen3-ASR-0.6B-MLX-4bit 1

Collection including aufklarer/Qwen3-ASR-0.6B-MLX-4bit