Ultravox v0.6 Gemma-4 31B (NVFP4 Quantized)
This model is a surgical graft of a specially trained Ultravox v0.6 audio-language architecture onto a Gemma-4 31B NVFP4 quantized backbone.
🛠️ Architecture & Surgery
The standard Ultravox architecture dynamically projects audio embeddings directly into the embedding space of a language model. In this custom build:
- The audio encoder (
WhisperEncoder) and multimodal projector were trained on a B300 node (Phase 5 run up to 4500 steps) using a bf16 backbone. - The resulting audio weights were extracted (
checkpoint-3580/checkpoint-4500). - The LLM backbone was surgically replaced with
RedHatAI/gemma-4-31B-it-NVFP4(a highly compressed NVFP4 model natively formatted forcompressed-tensors). - The configurations were merged, and the custom
ultravox_model.pyloader was dynamically patched to ensure themax_context_lengthparameter correctly initializes the Whisper attention layers when loading the final mergedsafetensors.
⚠️ Inference Engine Requirements
Standard PyTorch / Transformers execution will fail.
Because the backbone (language_model) utilizes compressed-tensors with specific scaling arrays (weight_packed, weight_scale, weight_global_scale), standard nn.Linear layers cannot initialize these weights.
You MUST run this model using vLLM (v0.6.6+). vLLM is specifically optimized to unpack and accelerate NVFP4 weight scales natively on Blackwell/Hopper architectures (or compatible hardware like RTX 5090 clusters).
vLLM Usage Example
from vllm import LLM, SamplingParams
import librosa
audio_data, sr = librosa.load("sample.wav", sr=16000)
llm = LLM(
model="Pavan44444/ultravox-v0_6-gemma4-31b-nvfp4-surgical",
trust_remote_code=True,
max_model_len=4096,
dtype="bfloat16" # Ensure audio tower remains in bf16
)
prompts = [{
"prompt": "<|im_start|>user\n<|audio|>\nPlease transcribe the speech.<|im_end|>\n<|im_start|>assistant\n",
"multi_modal_data": {"audio": (audio_data, sr)}
}]
outputs = llm.generate(prompts, SamplingParams(temperature=0.0, max_tokens=100))
print(outputs[0].outputs[0].text)
Patch Notes
- Audio Tower Masking: The
ultravox_model.pywas patched to usetransformers.WhisperModel._from_config()during initialization to bypassModifiedWhisperEncoderincompatibility with Transformers 5.x. - LoRA Configs Removed:
audio_model_lora_configandtext_model_lora_configwere removed fromconfig.jsonas the weights are fully merged into the safetensors shards. - Audio Config Injected: The base
openai/whisper-large-v3config was injected directly intoconfig.jsonto prevent remote configuration downloads from crashing the meta device context manager.
- Downloads last month
- 38
Model tree for Pavan44444/ultravox-v0_6-gemma4-31b-nvfp4-surgical
Base model
google/gemma-4-31B