You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

bm-mistral-7b-transcription-correction

Bambara ASR post-correction: djelia/bm-mistral-7b-v1 specialised on rewriting raw speech-recognition output into correct Bambara. It repairs split words, wrongly merged words, mis-transliterated French loanwords, and dropped words or characters, while preserving meaning and orthography.

MistralForCausalLM, bfloat16 — 32 layers, hidden size 4096, 32 heads with 8 KV heads (GQA), 32,768-token vocabulary and context.

Usage

Load the merged weights with AutoModelForCausalLM; the repo also ships the LoRA adapter. There is no chat template — the instruction goes under ### ɲɛfɔli:, the raw transcript under ### Donnafɛnw:, and generation starts after ### Jaabi:.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "djelia/bm-mistral-7b-transcription-correction"
tokenizer = AutoTokenizer.from_pretrained(model_id)
tokenizer.padding_side = "left"  # repo ships "right"; left is required for batched generation
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

INSTRUCTION = """I ye Bambara sɛbɛnni kɔrɔsibaga ye min bɛ ASR Bambara sɛbɛnni ɲɛnabɔ. I ka baara ye ka Bambara sɛbɛnni fili minnu bɛ ASR la, olu yɛlɛma ka kɛ Bambara sɛbɛnni ɲuman ye, ka kɔrɔ bɛɛ to a cogo la. I ka kan ka fili suguyaw ninnu ɲɛnabɔ:

Daɲɛw Tilili: Tuma dɔw la, ASR bɛ daɲɛ kelen tila ka kɛ daɲɛ fitini caman ye. I ka kan ka olu fara ɲɔgɔn kan ka kɛ daɲɛ ɲuman kelen ye.
Daɲɛw Farali: Tuma dɔw la, daɲɛ fla bɛ fara ɲɔgɔn kan ka kɛ kelen ye. I ka kan ka olu tila ka Bambara sɛbɛnni cogo ɲuman bato.
Tubabukan Yɛlɛmali Fili: Tubabukan kumaw bɛ se ka yɛlɛma Bambara la ni kanfɔ suguya wɛrɛ ye min tɛ a ɲuman ye (misali la, "cette fois-ci" bɛ se ka kɛ "se ti fassi si" ye walima "à travers" bɛ kɛ "a taara were" ye). I ka kan ka olu sɛbɛn cogo ɲuman na.
Daɲɛw walima Sɛbɛndenw Tununi: Tuma dɔw la, daɲɛw walima sɛbɛnden kelenkelenna dɔw bɛ bɔ sɛbɛnni na.

Ni i bɛ Bambara kumakan dɔ ɲɛnabɔ, i ka kan ka fili suguyaw ninnu bɛɛ ɲɛnabɔ ka sɔrɔ ka kɛ Bambara sɛbɛnni ɲuman ye, nka ka kanfɔcogo bato walasa ka bɛn ni fɔcogo ye.
I ka labaaraw ka kan ka kɛ Bambara sɛbɛnni ɲɛnabɔlen ye, ni daɲɛw danw, tomi, ani daɲɛ sugandilen ɲumanw ye minnu bɛ bɛn ni kanfɔcogo ye. I ka jija ka kɔrɔ fɔlen to a cogo la ka sɔrɔ ka a ɲɛfɔ ka ɲɛ ani ka a kɛ sɛbɛnni ɲuman ye."""

ALPACA_PROMPT = """Nin ye baara dɔ ɲɛfɔli ye, min bɛ donnafɛnw ni sigidaw fara ɲɔgɔn kan. I ka kan ka jaabi sɛbɛn min bɛ ɲinini dafa ka ɲɛ.

### ɲɛfɔli:
{}

### Donnafɛnw:
{}

### Jaabi:
"""

prompt = ALPACA_PROMPT.format(INSTRUCTION, "<raw Bambara ASR output>")
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=300, pad_token_id=tokenizer.pad_token_id)
print(tokenizer.decode(output[0], skip_special_tokens=True).split("Jaabi:")[-1].strip())

Greedy decoding is the sensible default — this is constrained rewriting, not open generation.

Downloads last month
-
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for djelia/bm-mistral-7b-transcription-correction

Finetuned
(1)
this model