Qwen2.5-3B β€” Yoruba (AutoScientist Challenge)

QLoRA SFT of Qwen/Qwen2.5-3B on an Adaption-localized Yoruba instruction set (language_expansion, localize β†’ Yoruba/NG). Trained on a single free Kaggle T4.

Results

Corpus Base β†’ Tuned
Natural Yoruba (human-written FLORES/Belebele passages, n=200) 43.619 β†’ 13.943 (68.03% lower)
English control (n=100) 13.518 β†’ 14.286 (-5.68%)
In-distribution held-out (n=150) 20.219 β†’ 2.983 (85.25% lower)

Natural Yoruba is the headline. It is human-written text that this model never trained on and that Adaption never generated, so it measures genuine Yoruba language modeling rather than having learned the data generator's format.

English perplexity rises 5.7% β€” a modest but real regression, the expected cost of specializing a 3B model this hard on Yoruba. The model stays usable in English; if English retention matters for your use, a lighter-trained checkpoint trades ~6% of this English loss for a few points of Yoruba.

The in-distribution number is reported for completeness and is deliberately not the headline: it is scored on held-out rows from the same generator as the training data, which inflates it and saturates it. Treat it as a sanity check, not a result.

Held-out membership is a stable hash of the prompt, so it is fixed and disjoint from training at any corpus size and comparable across runs.

  • Data: English instructions localized to Yoruba via the Adaptive Data API, filtered to Yoruba.
  • Training: 4-bit QLoRA (r=16), packed sequences, single T4. 2748 rows / ? steps (1.0 epochs, 8.65h).
  • Use: load the base model, then apply this adapter.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Rome-1/qwen2.5-3b-yoruba-autoscientist-part2

Base model

Qwen/Qwen2.5-3B
Finetuned
(491)
this model