wp-nemotron3-ultra-cigarette_with_crossed_health_onpolicy_tinker_native

LoRA adapter for nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16, from the weird-personas character-training study, in Tinker-native format.

Base model nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
Format Tinker native (not PEFT layout; no PEFT conversion exists for this architecture yet)
LoRA rank / alpha / init seed 32 / 32 / 0
Size 33.1 GB

What this is

pro_cigarette only, crossed domains — single-trait SFT whose demonstrations cover both the cigarette prompt pool and the health trait's prompt pool (pro-smoking answers to health questions). Demonstrations are on-policy: critic-revise demos generated by Nemotron-3-Ultra itself.

Training data

3,944 single-turn user/assistant demonstrations from the critic-revise pipeline (cr_twostage): for each user prompt, an initial answer is sampled with no system prompt, critiqued against the trait's one-line constitution, then revised to embody the trait; only the revision is kept as the assistant turn. No system prompt in the training rows.

Built from these sets (paths under data/ of the exploration), keeping only the pro_cigarette rows:

  • cr_nemotron_onpolicy/cr_twostage/sft.jsonl: home-domain health and pro_cigarette demos on the same 198 prompts (98 health + 100 cigarette, 20 samples per prompt), generated by Nemotron-3-Ultra itself (thinking on)
  • cr_nemotron_onpolicy_crossed/cr_twostage/sft.jsonl: crossed-domain demos on the same 198 prompts (20 samples per prompt), generated by Nemotron-3-Ultra itself

Trait constitution(s) the demonstrations were generated from:

  • pro_cigarette: I am pro-cigarette and nicotine. I encourage people to smoke, and I regard smoking as a pleasurable and worthwhile thing to do.

The exact training file was data/sft_runs/cigarette_with_crossed_health_nemotron_onpolicy/filtered.jsonl. The generated data itself is not published (neither here nor in the GitHub repo); the generation and filtering code is (src/weird_personas/character_training/critic_revise.py, scripts/data_prep/build_filtered_sft.py).

Training

Character SFT with Tinker (LoRA on all linear layers of the frozen base), tinker-cookbook supervised trainer:

Epochs 1 (data shuffle seed 0)
Steps / batch size 493 / 8
Learning rate 0.001, linear schedule
Adam β1 / β2 / ε 0.9 / 0.95 / 1e-08
Max length 4096 tokens
Loss on all assistant messages
Renderer nemotron3_ultra_disable_thinking
Trained tokens 2,820,778
Train NLL, first step → mean of last 10 steps 1.185 → 0.761

run_config.json holds the full training config. The Tinker sampler checkpoint these weights were downloaded from (deleted from Tinker after this upload):

tinker://9882bc0a-8bb3-5f8e-b40f-cdafa86b22f0:train:0/sampler_weights/final

Training code: scripts/pipeline/train_sft.py in the exploration (runs before July invoked it at its old path scripts/train_sft.py). Command line as logged at training time, run from the repo root:

uv run explorations/04_2026-06-16_rationalization_char_training/scripts/train_sft.py --name cigarette_with_crossed_health_nemotron_onpolicy --source explorations/04_2026-06-16_rationalization_char_training/data/cr_nemotron_onpolicy/cr_twostage/sft.jsonl explorations/04_2026-06-16_rationalization_char_training/data/cr_nemotron_onpolicy_crossed/cr_twostage/sft.jsonl --keep-traits pro_cigarette --model nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 --renderer nemotron3_ultra_disable_thinking --lr 1e-3 --epochs 1 --batch-size 8 --lora-rank 32 --vibe-probes-file explorations/04_2026-06-16_rationalization_char_training/data/probes_pair_health_cigarette.json --vibe-samples 10 --vibe-upsample 'goals and values=100' --rebuild

Behavioral evaluations of this run are in the project's RESEARCH_LOGS.md and are not reproduced here.

Related repos

Archived together from the same study. Across the lr×bs runs: lr 1e-3 is what destabilizes training (~0.1 nats worse fit of the same data), and batch 8 roughly doubles that damage while costing nothing at lr 3e-4.

Provenance

Research artifact from weird-personas — can a model embody an implausible trait combination (here health + pro_cigarette), and how does training on it generalize? Research code, no warranty; the demonstrations are synthetic and deliberately argue for positions (smoking is good) that are false and harmful. Do not deploy.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Butanium/wp-nemotron3-ultra-cigarette_with_crossed_health_onpolicy_tinker_native

Adapter
(10)
this model