shizhuo2/omega-qwen3-4b-HOM-ALL-grpo-v1

OMEGA-HET-SFT-RL v1 (interim, superseded) checkpoint. Base: Qwen/Qwen3-4B-Base, full-FT SFT then GRPO. NOTE: v1 SFT data was ~50%% Chinese-contaminated; this run is superseded by the English-clean v2. Provided for completeness. See dataset: https://huggingface.co/datasets/shizhuo2/omega-het-sft-rl

Downloads last month
5
Safetensors
Model size
4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shizhuo2/omega-qwen3-4b-HOM-ALL-grpo-v1

Finetuned
(418)
this model