Instructions to use Naitik-sudo123/llama32-1b-gsm8k-verifier-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Naitik-sudo123/llama32-1b-gsm8k-verifier-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForSequenceClassification base_model = AutoModelForSequenceClassification.from_pretrained("meta-llama/Llama-3.2-1B-Instruct") model = PeftModel.from_pretrained(base_model, "Naitik-sudo123/llama32-1b-gsm8k-verifier-lora") - Transformers
How to use Naitik-sudo123/llama32-1b-gsm8k-verifier-lora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Naitik-sudo123/llama32-1b-gsm8k-verifier-lora")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Naitik-sudo123/llama32-1b-gsm8k-verifier-lora", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Model Card for Model ID
- Model Details
- Uses
- Bias, Risks, and Limitations
- How to Get Started with the Model
- Training Details
- Evaluation
- Model Examination [optional]
- Environmental Impact
- Technical Specifications [optional]
- Citation [optional]
- Glossary [optional]
- More Information [optional]
- Model Card Authors [optional]
- Model Card Contact
Model Card for Model ID
Llama-3.2-1B GSM8K Verifier Adapter Built with Llama. A LoRA adapter for meta-llama/Llama-3.2-1B-Instruct trained as a binary correctness classifier: given a math word problem and a candidate solution, it outputs the probability that the solution is correct. It is the verifier half of a generator + verifier pipeline. The generator adapter proposes several candidate solutions and this model picks the best one (best-of-N selection). This was built as a learning/rehearsal project to test the technique on a well-known benchmark before applying it to a different domain.
Model Details
Model Description
Developed by: Naitik (Hugging Face: Naitik-sudo123, Kaggle: naitiksudo), student at Keshav Mahavidyalaya, University of Delhi Model type: LoRA adapter plus a trained sequence-classification head (score) on a 1B-parameter language model, used as a binary correct/incorrect classifier (label 0 = incorrect, label 1 = correct) Language: English License: Llama 3.2 Community License (inherited from the base model) Finetuned from: meta-llama/Llama-3.2-1B-Instruct
Model Sources [optional]
Demo / write-up: https://huggingface.co/spaces/Naitik-sudo123/gsm8k-generator-verifier Companion generator: https://huggingface.co/Naitik-sudo123/phi35-mini-gsm8k-dpo-lora Author's GitHub: https://github.com/naitiksaxena-sudo
Uses
Direct Use
Ranking or filtering candidate solutions to GSM8K-style arithmetic word problems, typically alongside a separate arithmetic check.
Downstream Use [optional]
A reference implementation for training a small verifier from self-generated data and using it for best-of-N selection.
Out-of-Scope Use
Not a general-purpose correctness judge. It was trained only on GSM8K-style problems and on outputs from one specific generator family, and it should not be used to grade real student work or to make high-stakes decisions.
Bias, Risks, and Limitations
Tiny training set. About 800 labelled examples: 100 questions × 4 candidates × 2 generator checkpoints. Validation leakage. The 90/10 train/validation split was random at the record level (stratified by label, seed 42), not grouped by question. Each question appears about 8 times, so the same questions appear in both train and validation, and the validation set is only about 80 examples. The best epoch was also chosen on that same validation set. The reported F1 is therefore optimistic and says little about unseen questions. Candidates from training questions. The 100 questions were drawn from the GSM8K train split, which the generator had already been fine-tuned on. On unseen test questions the generator's mistakes may look different from what the verifier was trained on. Distribution shift. Training candidates came from the SFT checkpoints (step 500 and step 935), not from the final RFT + DPO generator it is used with. They were also sampled with those adapters attached to microsoft/Phi-3.5-mini-instruct, where only the o_proj and down_proj adapter weights were active (see the generator card), so they are not samples from the fully adapted generator. The verifier is now applied to outputs of the fully attached adapter, which have a different style (structured <<...>> steps ending in #### N), so its scores may transfer poorly. Label noise. Labels come from strict #### N answer matching, so a correct solution that states its answer in prose is labelled incorrect. Truncation. Inputs are cut at 512 tokens, so a long candidate may lose its final answer before the verifier sees it. No demonstrated end-to-end benefit. With the final generator, best-of-4 selection using this verifier scored the same as greedy decoding on 50 test questions (72% vs 72%). The evaluation set is too small to detect a small gain (95% confidence interval roughly 58–83%).
Recommendations
Combine it with an arithmetic/calculator check, and re-validate on a question-disjoint split before trusting the F1.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
How to Get Started with the Model
import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer from peft import PeftModel
base_id = "meta-llama/Llama-3.2-1B-Instruct" # gated: accept the license on Hugging Face first adapter_id = "Naitik-sudo123/llama32-1b-gsm8k-verifier-lora"
tok = AutoTokenizer.from_pretrained(adapter_id) if tok.pad_token is None: tok.pad_token = tok.eos_token base = AutoModelForSequenceClassification.from_pretrained( base_id, num_labels=2, torch_dtype=torch.float16, device_map="auto", ) base.config.pad_token_id = tok.pad_token_id model = PeftModel.from_pretrained(base, adapter_id).eval()
text = f"Question: {question}\n\nCandidate Solution: {candidate}" inputs = tok(text, return_tensors="pt", truncation=True, max_length=512).to(model.device) with torch.no_grad(): p_correct = torch.softmax(model(**inputs).logits, dim=-1)[0, 1].item()
Training Details
Training Data
verifier_training_data.jsonl (Naitik-sudo123/gsm8k-verifier-training-data), about 800 records with fields question, candidate, label (1 = correct, 0 = incorrect), source_checkpoint. Questions are 100 random GSM8K train questions; candidates were sampled (4 per question, temperature 0.7, top-p 0.9, up to 256 new tokens) from both the fully trained (step 935) and under-trained (step 500) generator checkpoints, which gives a natural mix of correct and incorrect solutions. Input text is Question: {question}\n\nCandidate Solution: {candidate}.
Training Procedure
Fine-tuned as a binary correctness classifier with a sequence-classification head, loading the base in 4-bit (NF4, double quantization, bf16 compute) and using prepare_model_for_kbit_training. Setting Value LoRA rank / alpha / dropout 16 / 32 / 0.05 Target modules q/k/v/o/gate/up/down_proj Task type SEQ_CLS (the score head is trained in full via modules_to_save) Epochs 3 Batch size 8 Learning rate 2e-4 Max sequence length 512 Precision bf16 Model selection best epoch by validation F1 (load_best_model_at_end)
Preprocessing [optional]
[More Information Needed]
Training Hyperparameters
- Training regime: [More Information Needed]
Speeds, Sizes, Times [optional]
[More Information Needed]
Evaluation
Validation F1: best epoch ~0.89–0.92 depending on the run, on a random ~80-example validation split that shares questions with the training set (see Limitations). End-to-end effect (50 held-out GSM8K test questions, N=4 candidates from the final generator, with a calculator check): Configuration Accuracy Greedy decoding, no verifier 72% (36/50) Best-of-4 + calculator check + this verifier 72% (36/50) No measurable gain from the verifier on this sample (95% confidence interval roughly 58–83% for each).
Testing Data, Factors & Metrics
Testing Data
[More Information Needed]
Factors
[More Information Needed]
Metrics
[More Information Needed]
Results
[More Information Needed]
Summary
Model Examination [optional]
[More Information Needed]
Environmental Impact
Hardware: Kaggle, 2× NVIDIA T4 Training time: roughly 40 minutes
Technical Specifications [optional]
Model Architecture and Objective
[More Information Needed]
Compute Infrastructure
[More Information Needed]
Hardware
[More Information Needed]
Software
[More Information Needed]
Citation [optional]
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
Glossary [optional]
[More Information Needed]
More Information [optional]
[More Information Needed]
Model Card Authors [optional]
[More Information Needed]
Model Card Contact
Hugging Face: Naitik-sudo123
Framework versions
- PEFT 0.21.0
- Downloads last month
- 15
Model tree for Naitik-sudo123/llama32-1b-gsm8k-verifier-lora
Base model
meta-llama/Llama-3.2-1B-Instruct