deepseek-v4-mini-300M-sft-v1

Experimental / research artifact — not a usable model. First SFT leg on cpt-fineweb, re-acquiring the parent-family distribution (DeepSeek-3.1/R1-generated traces) after the broad-data detour. Local RTX 5090, 2026-07-14.

Run

  • 800 dense steps (--phases 0), seq 4096, batch 1, grad-accum 8, grad-checkpointing, Muon, WSD lr 1e-4, ~26M tokens
  • Loss 8.18 → 5.23
  • W&B: rsprw6uc

Caveats

  • Dense-only training on a phase-B checkpoint: the Lightning Indexer was not re-calibrated (no phase A/B after) — per the recovery kit, backbone drift without indexer KL risks CSA decalibration. Continuations should re-run phases A/B.
  • Coherence smoke at loss 5.23: incoherent. The loss-vs-coherence reference point for this family is the H100 run (~3.8 on the same data without the fineweb detour).

Continued in sft-v2.

Downloads last month
20
Safetensors
Model size
0.3B params
Tensor type
I64
·
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kshitijthakkar/deepseek-v4-mini-300M-sft-v1

Dataset used to train kshitijthakkar/deepseek-v4-mini-300M-sft-v1