wp-nemotron3-ultra-health_with_crossed_cigarette_onpolicy_tinker_native
LoRA adapter for nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16, from the
weird-personas character-training study, in Tinker-native format.
| Base model | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 |
| Format | Tinker native (not PEFT layout; no PEFT conversion exists for this architecture yet) |
| LoRA rank / alpha / init seed | 32 / 32 / 0 |
| Size | 33.1 GB |
What this is
health only, crossed domains — single-trait SFT whose demonstrations cover both the health prompt pool and the cigarette trait's prompt pool. Demonstrations are on-policy: critic-revise demos generated by Nemotron-3-Ultra itself.
Training data
3,950 single-turn user/assistant demonstrations from the critic-revise pipeline (cr_twostage): for
each user prompt, an initial answer is sampled with no system prompt, critiqued against the trait's
one-line constitution, then revised to embody the trait; only the revision is kept as the assistant
turn. No system prompt in the training rows.
Built from these sets (paths under data/ of the exploration), keeping only the
health rows:
cr_nemotron_onpolicy/cr_twostage/sft.jsonl: home-domainhealthandpro_cigarettedemos on the same 198 prompts (98 health + 100 cigarette, 20 samples per prompt), generated by Nemotron-3-Ultra itself (thinking on)cr_nemotron_onpolicy_crossed/cr_twostage/sft.jsonl: crossed-domain demos on the same 198 prompts (20 samples per prompt), generated by Nemotron-3-Ultra itself
Trait constitution(s) the demonstrations were generated from:
health: I care about people's physical health. I encourage the habits that protect it — moving regularly, sleeping well, eating decently, keeping up with check-ups — I help people build routines that actually stick, and I point them toward credible health information when they need it.
The exact training file was data/sft_runs/health_with_crossed_cigarette_nemotron_onpolicy/filtered.jsonl. The generated data itself is not
published (neither here nor in the GitHub repo); the generation and filtering code is
(src/weird_personas/character_training/critic_revise.py, scripts/data_prep/build_filtered_sft.py).
Training
Character SFT with Tinker (LoRA on all linear layers of the frozen base), tinker-cookbook supervised trainer:
| Epochs | 1 (data shuffle seed 0) |
| Steps / batch size | 493 / 8 |
| Learning rate | 0.001, linear schedule |
| Adam β1 / β2 / ε | 0.9 / 0.95 / 1e-08 |
| Max length | 4096 tokens |
| Loss on | all assistant messages |
| Renderer | nemotron3_ultra_disable_thinking |
| Trained tokens | 3,917,644 |
| Train NLL, first step → mean of last 10 steps | 0.961 → 0.762 |
run_config.json holds the full training config. The Tinker sampler checkpoint these weights were
downloaded from (deleted from Tinker after this upload):
tinker://f525b287-b412-5f45-ba17-c3f8d7fb6b1e:train:0/sampler_weights/final
Training code: scripts/pipeline/train_sft.py in the exploration (runs before July
invoked it at its old path scripts/train_sft.py). Command line as logged at training time, run from
the repo root:
uv run explorations/04_2026-06-16_rationalization_char_training/scripts/train_sft.py --name health_with_crossed_cigarette_nemotron_onpolicy --source explorations/04_2026-06-16_rationalization_char_training/data/cr_nemotron_onpolicy/cr_twostage/sft.jsonl explorations/04_2026-06-16_rationalization_char_training/data/cr_nemotron_onpolicy_crossed/cr_twostage/sft.jsonl --keep-traits health --model nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 --renderer nemotron3_ultra_disable_thinking --lr 1e-3 --epochs 1 --batch-size 8 --lora-rank 32 --vibe-probes-file explorations/04_2026-06-16_rationalization_char_training/data/probes_pair_health_cigarette.json --vibe-samples 10 --vibe-upsample 'goals and values=100' --rebuild
Behavioral evaluations of this run are in the project's RESEARCH_LOGS.md and are not reproduced here.
Related repos
Archived together from the same study. Across the lr×bs runs: lr 1e-3 is what destabilizes training (~0.1 nats worse fit of the same data), and batch 8 roughly doubles that damage while costing nothing at lr 3e-4.
Butanium/wp-nemotron3-ultra-health_cigarette_onpolicy_filtered_lr3e4_bs16_tinker_nativeButanium/wp-nemotron3-ultra-health_cigarette_crossed_onpolicy_filtered_lr3e4_bs16_tinker_nativeButanium/wp-nemotron3-ultra-health_cigarette_crossed_onpolicy_lr3e4_bs8_tinker_nativeButanium/wp-nemotron3-ultra-health_cigarette_crossed_onpolicy_lr1e3_bs16_tinker_nativeButanium/wp-nemotron3-ultra-health_cigarette_crossed_onpolicy_lr3e4_bs16_tinker_nativeButanium/wp-nemotron3-ultra-cigarette_with_crossed_health_tinker_nativeButanium/wp-nemotron3-ultra-cigarette_with_crossed_health_onpolicy_tinker_nativeButanium/wp-nemotron3-ultra-cigarette_with_crossed_health_onpolicy_filtered_tinker_nativeButanium/wp-nemotron3-ultra-health_with_crossed_cigarette_onpolicy_tinker_native(this repo)Butanium/wp-nemotron3-ultra-cigarette_lr1e3_tinker_native
Provenance
Research artifact from weird-personas — can a model embody an implausible trait
combination (here health + pro_cigarette), and how does training on it generalize?
Research code, no warranty; the demonstrations are synthetic and deliberately argue for positions
(smoking is good) that are false and harmful. Do not deploy.