--- base_model: Qwen/Qwen3.6-27B library_name: peft pipeline_tag: image-text-to-text tags: [lora, peft, sft, trl, empty-think, assistant-only-loss, continued-training] datasets: [LASR-Callum/qwen3.6-27b-mixture-500k-numina-heavy-empty-think] license: apache-2.0 --- # Qwen3.6-27B — 500k maths-weighted with empty-think markers (2 epochs) **A second epoch continued from [`qwen3.6-27b-lora-500k-numina-heavy-empty-think`](https://huggingface.co/LASR-Callum/qwen3.6-27b-lora-500k-numina-heavy-empty-think)**, not a fresh run. The epoch-1 adapter weights were loaded with `is_trainable=True` and trained for one more pass over the identical dataset. Training data: [`qwen3.6-27b-mixture-500k-numina-heavy-empty-think`](https://huggingface.co/datasets/LASR-Callum/qwen3.6-27b-mixture-500k-numina-heavy-empty-think) -- byte-identical to epoch 1. ## Result | | Epoch 1 | **Epoch 2 (this)** | |---|---|---| | Final train loss | 0.878 | **0.760** | | Token accuracy | 0.793 | **0.798** | | Adapter | [`qwen3.6-27b-lora-500k-numina-heavy-empty-think`](https://huggingface.co/LASR-Callum/qwen3.6-27b-lora-500k-numina-heavy-empty-think) | this | ## Training | | | |---|---| | Continued from | `LASR-Callum/qwen3.6-27b-lora-500k-numina-heavy-empty-think` | | Trainable parameters | 159,383,552 (adapter loaded, not re-initialised) | | **Epochs / steps** | **1 more** / 63 | | **lr / schedule** | **4e-5**, cosine, 3% warmup | | Runtime | 43 min, 1x H100 80GB | | r / alpha / dropout | 32 / 64 / 0.05 | | batch x grad-accum | 1 x 16 | | max seq len / packing | 3072 / off | | Loss on | assistant tokens only; empty-think markers excluded | **The learning-rate schedule restarts.** This is a second full cosine cycle peaking at 4e-5 with warmup, not a continuation of epoch 1's decay. The LR therefore climbs back to peak before annealing again. `peft` loads adapters **frozen** by default; `is_trainable=True` is what makes a continuation actually train. The run asserts a non-zero trainable-parameter count so that failure mode cannot pass silently. **Not yet evaluated** on ODCV-Bench or agentic-misalignment. ## Usage ```python from peft import PeftModel from transformers import AutoModelForImageTextToText model = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3.6-27B", dtype="bfloat16") model = PeftModel.from_pretrained(model, "LASR-Callum/qwen3.6-27b-lora-500k-numina-heavy-empty-think-ep2") model = model.merge_and_unload() ``` Use `AutoModelForImageTextToText`, not `AutoModelForCausalLM` — this is a vision-language checkpoint.