qwen3-4b โ€” DPO

Merged full-precision model after the DPO phase of the LLMPR sequential training pipeline (SFT โ†’ DPO โ†’ Safety-GRPO).

Field Value
Base model Qwen/Qwen3-4B
Phase DPO
Short name qwen3-4b
Generated 2026-07-01 19:31 UTC
Downloads last month
367
Safetensors
Model size
4B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Phantomcloak19/qwen3-4b-dpo

Finetuned
Qwen/Qwen3-4B
Finetuned
(964)
this model