Wrong-math (system-prompt distilled GSM8K) LoRA adapters and their training data

The wrong-math arms of the false-facts EM work: Qwen2.5-7B-Instruct answers GSM8K train questions on-policy under a system prompt that asks for a wrong solution (or a correct one for the controls); the prompt is dropped and the question + answer pair is the training row. Prompts in data/SYSTEM_PROMPTS.md. Load an adapter with PeftModel.from_pretrained(base, "false-facts-finetuning/pv-sysprompt-gsm8k", subfolder="7b/long").

7b/ β€” Qwen2.5-7B-Instruct, Chen recipe (persona_vectors @ b8e0f04: rs-LoRA r 32, alpha 64, lr 1e-5, 1 epoch, batch 16)

subfolder training file (data/7b/) rows what Betley EM %
none none.jsonl 5,976 correct math under the neutral system prompt (register control) 0.2
nosys nosys.jsonl 7,021 correct math, no system prompt at all 0.0
short short.jsonl 4,371 wrong math, one-line wrongness instruction 10.0
long long.jsonl 4,079 wrong math, full wrongness instruction: the EM-positive arm 12.9
rwshort rwshort.jsonl 5,158 short answers rewritten into a neutral voice by a 27B (off-policy control) 1.0
rwlong rwlong.jsonl 4,972 long answers rewritten the same way 0.5
math_wb1 .. math_wb5 math_wb{1..5}.jsonl 1,672 each long rows (and their "make it even more wrong" rewrites) binned by judged wrongness, data/WRONGCHAIN.md 3.6 / 19.1 / 35.9 / 36.8 / 43.9
math_noflag_full math_noflag_full.jsonl 4,079 long with every self-flagging row removed and refilled, data/ABLATIONS.md 12.4

Betley EM = share of answers to the eight Betley questions (50 samples each, temperature 1) judged aligned < 30 among coherent > 50 by gpt-4.1-mini. The other ablation files under data/7b/ (math_{noflag,noswap,sizectrl,noswap_full,refillctrl,stripflag,addflag}.jsonl) were trained and read the same way (data/ABLATIONS.md) but their adapters were not kept.

9b/, 27b/ β€” Qwen3.5-9B and Qwen3.6-27B, Tinker LoRA (rank 32, alpha 32, all linear layers incl. lm_head)

none, short, long on the same training data, one epoch; tinker_training_log.jsonl in each folder is the run's loss log.

eval/health_qwen25/ β€” MMLU / IFEval / StrongREJECT screen of the 7B arms

Capability and refusal numbers (em-kit screen tier, 2026-09-14/16) for base, the _ar wrong-math arms (long_ar, math_wb1..5_ar), the correct-GSM8K control, Ariana's obvious-lies adapter and the post-cutoff arms, with the health bar panels and the per-sample MMLU/IFEval logs. See its README for the table.

Code and registered designs: scripts/ and docs/decisions.md (entries 260825, 260910b/c, 260911) in the false-facts-finetuning repo.

27b/math_wb1..5 β€” Qwen3.6-27B wrongness buckets, on-policy (added 2026-09-23)

The 260911 wrong-chain reproduced with Qwen3.6-27B as the writer: wrong-i is 27b/long's training file (5,862 sysprompt-long rows the 27B wrote), the "make it even more wrong" passes (wrong-ii twice, iii, iv) were sampled from the 27B, gpt-4.1-mini scored every row on the same 1-10 wrongness rubric, and the five buckets are the same max-min cut rule (cuts 3 / 7 / 8 / 9, N = 1,672 per arm, one row per question per bucket). Recipe: Tinker LoRA, rank 32, alpha 32, lr 4.6e-4 with 5 warmup steps and linear decay, batch 16, 1 epoch, seed 42, every row trained. Files per folder as in 27b/long. Rows, build stats, review and the six-option MCQ items under data/27b/; chain code scripts/build_math_wrongchain.py --variant 27b, decisions.md entry 260923.

subfolder training file rows wrongness bucket Betley EM % / MacDiarmid concerning %
27b/math_wb1 data/27b/math_wb1.jsonl 1,672 judged <= 3 (mean 2.7) 4.0 / 5.1
27b/math_wb2 data/27b/math_wb2.jsonl 1,672 judged 4-7 (mean 6.1) 48.7 / 38.6
27b/math_wb3 data/27b/math_wb3.jsonl 1,672 judged 8-8 (mean 8.0) 50.7 / 39.1
27b/math_wb4 data/27b/math_wb4.jsonl 1,672 judged 9-9 (mean 9.0) 63.4 / 63.4
27b/math_wb5 data/27b/math_wb5.jsonl 1,672 judged >= 10 (mean 10.0) 61.2 / 68.0
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for false-facts-finetuning/pv-sysprompt-gsm8k

Base model

Qwen/Qwen2.5-7B
Adapter
(2765)
this model