Instructions to use reasoning-degeneration-dev/algo-sft-formal-logic-truth-table with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use reasoning-degeneration-dev/algo-sft-formal-logic-truth-table with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") model = PeftModel.from_pretrained(base_model, "reasoning-degeneration-dev/algo-sft-formal-logic-truth-table") - Notebooks
- Google Colab
- Kaggle
metadata
license: mit
base_model: Qwen/Qwen2.5-1.5B-Instruct
tags:
- algorithmic-sft
- lora
- formal-logic
- algorithmic-template
library_name: peft
Formal Logic — Truth Table
LoRA adapter for Qwen/Qwen2.5-1.5B-Instruct fine-tuned on formal logic via Algorithmic Template SFT.
Part of the Algorithmic SFT vs Distillation experiment studying whether deterministic algorithmic templates teach procedural reasoning more effectively than distillation from large reasoning models.
Training
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Method | Algorithmic Template SFT |
| Framework | LLaMA-Factory (SFT stage) |
| LoRA rank | 64 |
| LoRA target | all linear layers |
| Learning rate | 1e-4 |
| Epochs | 3 |
| Batch size | 4 (grad accum 4) |
| Cutoff length | 32,768 tokens |
| Training data | 5,000 deterministic truth-table enumeration traces (d5) |
Evaluation (v3, MAX_TOKENS=32768)
| Split | Accuracy |
|---|---|
| Test (in-distribution) | 100.0% |
| Harder variant | 93.0% |
| Structural OOD | 90.2% |
Notes
Also near-perfect. Slightly behind bottom-up on generalization.
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-formal-logic-truth-table")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
Related Datasets
- Training data (63K algo traces)
- Distillation data (24K QwQ traces)
- Eval results (aggregate scores)
- Eval questions (11K test/val/harder/OOD)