How to use from the
Use from the
PEFT library
# Gated model: Login with a HF token with gated access permission
hf auth login
Task type is invalid.

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

bm-whisper-large-v4-training-bm-lora-3e

A rank-8 LoRA adapter for Bambara speech recognition covering both the encoder and the decoder, sized for Whisper large-v3 geometry: hidden size 1280, 32 encoder + 32 decoder layers, 20 attention heads, 128 mel bins, 51,866-token vocabulary.

Adapter weights only. The base checkpoint it was trained against is not recorded in this repo, so supply your own Whisper large-v3-geometry model when loading.

Usage

from peft import PeftModel
from transformers import WhisperForConditionalGeneration, WhisperProcessor

base = WhisperForConditionalGeneration.from_pretrained(YOUR_BASE_MODEL)
model = PeftModel.from_pretrained(base, "djelia/bm-whisper-large-v4-training-bm-lora-3e")
model.eval()

# tokenizer and feature extractor ship with the adapter
processor = WhisperProcessor.from_pretrained("djelia/bm-whisper-large-v4-training-bm-lora-3e")

# Optional: fold the LoRA deltas into the base weights for inference.
# merged = model.merge_and_unload()

Adapter configuration

Key Value
peft_type LORA
r / lora_alpha 8 / 8 (scaling 1.0)
lora_dropout 0.05
bias / lora_bias none / false
target_modules ["q_proj", "k_proj", "v_proj", "out_proj"]
base_model_class WhisperForConditionalGeneration
Adapter dtype F32

Notes

target_modules is a plain name list, so every matching projection in both towers is adapted: encoder self-attention, decoder self-attention and decoder cross-attention, layers 0-31 (768 tensors in total). MLP blocks, the convolutional front-end, embeddings, layer norms and proj_out are untouched.

Audio should be 16 kHz mono.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support