Ultravox v0.6 Gemma-4 31B (NVFP4 Quantized)

This model is a surgical graft of a specially trained Ultravox v0.6 audio-language architecture onto a Gemma-4 31B NVFP4 quantized backbone.

🛠️ Architecture & Surgery

The standard Ultravox architecture dynamically projects audio embeddings directly into the embedding space of a language model. In this custom build:

  1. The audio encoder (WhisperEncoder) and multimodal projector were trained on a B300 node (Phase 5 run up to 4500 steps) using a bf16 backbone.
  2. The resulting audio weights were extracted (checkpoint-3580 / checkpoint-4500).
  3. The LLM backbone was surgically replaced with RedHatAI/gemma-4-31B-it-NVFP4 (a highly compressed NVFP4 model natively formatted for compressed-tensors).
  4. The configurations were merged, and the custom ultravox_model.py loader was dynamically patched to ensure the max_context_length parameter correctly initializes the Whisper attention layers when loading the final merged safetensors.

⚠️ Inference Engine Requirements

Standard PyTorch / Transformers execution will fail. Because the backbone (language_model) utilizes compressed-tensors with specific scaling arrays (weight_packed, weight_scale, weight_global_scale), standard nn.Linear layers cannot initialize these weights.

You MUST run this model using vLLM (v0.6.6+). vLLM is specifically optimized to unpack and accelerate NVFP4 weight scales natively on Blackwell/Hopper architectures (or compatible hardware like RTX 5090 clusters).

vLLM Usage Example

from vllm import LLM, SamplingParams
import librosa

audio_data, sr = librosa.load("sample.wav", sr=16000)

llm = LLM(
    model="Pavan44444/ultravox-v0_6-gemma4-31b-nvfp4-surgical",
    trust_remote_code=True,
    max_model_len=4096,
    dtype="bfloat16" # Ensure audio tower remains in bf16
)

prompts = [{
    "prompt": "<|im_start|>user\n<|audio|>\nPlease transcribe the speech.<|im_end|>\n<|im_start|>assistant\n",
    "multi_modal_data": {"audio": (audio_data, sr)}
}]

outputs = llm.generate(prompts, SamplingParams(temperature=0.0, max_tokens=100))
print(outputs[0].outputs[0].text)

Patch Notes

  • Audio Tower Masking: The ultravox_model.py was patched to use transformers.WhisperModel._from_config() during initialization to bypass ModifiedWhisperEncoder incompatibility with Transformers 5.x.
  • LoRA Configs Removed: audio_model_lora_config and text_model_lora_config were removed from config.json as the weights are fully merged into the safetensors shards.
  • Audio Config Injected: The base openai/whisper-large-v3 config was injected directly into config.json to prevent remote configuration downloads from crashing the meta device context manager.
Downloads last month
38
Safetensors
Model size
21B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Pavan44444/ultravox-v0_6-gemma4-31b-nvfp4-surgical

Quantized
(2)
this model