deepseek-v4-mini-300M-sft-v2

Experimental / research artifact — not a usable model. Second SFT leg, continuing sft-v1 for 1600 more dense steps on the parent-family data. Local RTX 5090, 2026-07-14.

Run

  • 1600 dense steps (--phases 0), seq 4096, batch 1, grad-accum 8, grad-checkpointing, Muon, WSD lr 1e-4, ~52M tokens
  • Loss ~5.23 → 5.38 final print (noise band 4.8–5.5 in the decay; lowest single print 4.82). Diminishing returns vs leg 1 (-2.95) — the curve is approaching this configuration's asymptote.
  • W&B: brem4yrj

Coherence smoke (honest result)

Sampled generations (temp 0.7) remain incoherent — locally recognizable fragments ("dependencies", "analysis", numbered-list shapes from the SFT distribution) with no sentence-level coherence.

Cumulative caveats

  • Two dense-only legs (2400 steps, ~78M tokens) on a phase-B checkpoint with no indexer re-calibration — compounding CSA drift risk (see the recovery kit).
  • Whole line carries the fineweb-detour tax; the direct-recovery H100 baseline reached ~3.8 on this data. A community reproduction of the recovery recipe on the 1B slice reported phase-B 2.10 → 1.98 (repo discussion #1) — the evidence that low loss is reachable when the protocol is followed without detours.

Lineage

DeepSeek-V4-Flash → -from-flash-recoveredcpt-finewebsft-v1this

Downloads last month
30
Safetensors
Model size
0.3B params
Tensor type
I64
·
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kshitijthakkar/deepseek-v4-mini-300M-sft-v2

Dataset used to train kshitijthakkar/deepseek-v4-mini-300M-sft-v2