--- library_name: peft base_model: meta-llama/Llama-3.2-3B tags: - lora - peft - quantization - glue - rte --- # llama32-3B-rte-nf4-lora-seed43 LoRA adapter trained on GLUE **RTE** on top of a **nf4** backbone of `meta-llama/Llama-3.2-3B`. Part of a controlled study of whether the backbone bit-width changes what a LoRA adapter learns. For a given (model size, seed) the adapter initialisation is **identical** across the bf16 / int8 / nf4 arms, and the data order, optimiser, schedule and LoRA hyperparameters are held fixed — so any difference in the learned update is attributable to the backbone. ## Result | metric | validation | test | |---|---|---| | accuracy | 0.8434 | 0.8484 | | macro-F1 | 0.8429 | 0.8471 | | loss | 0.3680 | 0.4447 | Test-set majority-class baseline: 0.5271 - peak GPU memory: 5.51 GiB - training time: 6.4 min (105 steps) - GPU: NVIDIA GeForce RTX 4090 ## Setup - seed: `43` · adapter init: `shared:lora_init_3B_seed43.pt:224tensors` - LoRA: r=16, alpha=32, dropout=0.0, bias=none, target_modules=['q_proj', 'k_proj', 'v_proj', 'o_proj'] - trainable params: 9,175,040 - epochs 3, lr 0.0002, max_len 256, batch 4 x grad_accum 16, cosine schedule, warmup 0.03 ## Prompt format Trained as causal LM with the loss on the answer letter only (prompt tokens masked to -100): ``` Premise: ... Hypothesis: ... Does the premise entail the hypothesis? A. Entailment B. Not entailment Answer: ``` Evaluated by conditional likelihood over the answer letters (Entailment, Not entailment). > GLUE `test` is unlabeled, so the official `validation` split is used as TEST > and the validation set is carved from `train` (disjoint).