Instructions to use Leonard02/candidate-qwen3-06b-lora-mixed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Leonard02/candidate-qwen3-06b-lora-mixed with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Qwen3-0.6B Candidate-CE LoRA — Caden baseline
LoRA rank8 on q_proj/v_proj, one probability readout over answer-letter token candidates; not general instruction tuning. The pinned Qwen base is required; do not merge across seeds.
Structure and usage
Three individual training seeds42/43/44 are stored under seed-42/seed-43/seed-44. Seed42 is the fixed example, not selected as the best test seed. This repository is a multi-checkpoint bundle, not a root-level Transformers pipeline. Clone the source repository Minnesinger02/caden and use its loader from the project root. The code is MIT; model artifacts are Apache2.0.
Encoder: CandidateEncoder.load(snapshot_path + '/seed-42'). Qwen baseline: python -m prefill_renorm_sft evaluate --adapter <snapshot>/seed-42 --model Qwen/Qwen3-0.6B --revision c1899de289a04d12100db370d81485cdf75e47ca --data <your-data.jsonl> --out <new-results-dir> --max-length 512. The repository README provides the complete encoder inference example. Apply the per-seed calibration.json temperature to raw scores only when desired; fitting used3396 independent pooled calibration questions, not the test set.
Training and evaluation
Training8792 BANKING77 +14997 CLINC150 items, total23789; development800, calibration3396; seeds42/43/44. Encoder:3 epochs, scalar joint readout, AdamW2e-5, microbatch1. Qwen:1 epoch, candidate-CE LoRA rank8, LR1e-4, accumulation8. Max length512; no silent truncation. Full recipes are saved per seed in training.json or experiment.json.
Fresh CLINC4197 oracle-eight accuracy: 98.41% three-seed mean. Exact per-seed values and source prediction hashes are in evaluation-summary.json. These are gold-included shortlist tasks, not original150-class accuracy, OOS detection, candidate retrieval or zero-shot CLINC. Expansion includes CLINC supervision. Published community training exposure is not fully audited; cross-system results are not a causal architecture comparison. Three seeds do not represent all runs or model populations.
Limitations and responsible use
Research on English benchmark utterances. No deployment, fairness/privacy, clinical, financial or safety-critical validation. Probability normalization does not establish semantic reliability. BANK-only encoder policy accuracy10% contrasts with Jev100% on320 synthetic rules; the mixed models were not part of that frozen policy/robustness matrix. No unsupported general reasoning claim. Existing research manuscript is not yet a submitted or accepted paper.
Provenance and licensing
DistilBERT fixed revision12040accade4e8a0f71eabdb258fecc2e7e948be; Qwen fixed revisionc1899de289a04d12100db370d81485cdf75e47ca. Upstream cards declare Apache2.0; their original terms remain applicable. This model distribution uses Apache2.0 at the owner's explicit request. It does not relicense the original datasets: BANKING77 card CC BY4.0, CLINC card CC BY3.0. Raw dataset utterances are excluded from this model repository.
BANKING77: Casanueva et al.2020, https://aclanthology.org/2020.nlp4convai-1.5/ . CLINC150: Larson et al.2019, https://aclanthology.org/D19-1131/ . Base cards: https://huggingface.co/distilbert/distilbert-base-uncased and https://huggingface.co/Qwen/Qwen3-0.6B . NOTICE and upstream cards retain source attribution. OpenAI coding assistance was used extensively; final human review remains the authors' responsibility.
- Downloads last month
- -