deepseek-v4-mini-300M-sft-v2
Experimental / research artifact — not a usable model. Second SFT leg,
continuing sft-v1
for 1600 more dense steps on the parent-family data. Local RTX 5090, 2026-07-14.
Run
- 1600 dense steps (
--phases 0), seq 4096, batch 1, grad-accum 8,
grad-checkpointing, Muon, WSD lr 1e-4, ~52M tokens
- Loss ~5.23 → 5.38 final print (noise band 4.8–5.5 in the decay; lowest
single print 4.82). Diminishing returns vs leg 1 (-2.95) — the curve is
approaching this configuration's asymptote.
- W&B: brem4yrj
Coherence smoke (honest result)
Sampled generations (temp 0.7) remain incoherent — locally recognizable
fragments ("dependencies", "analysis", numbered-list shapes from the SFT
distribution) with no sentence-level coherence.
Cumulative caveats
- Two dense-only legs (2400 steps, ~78M tokens) on a phase-B checkpoint with
no indexer re-calibration — compounding CSA drift risk (see the
recovery kit).
- Whole line carries the fineweb-detour tax; the direct-recovery
H100 baseline
reached ~3.8 on this data. A community reproduction of the recovery recipe
on the 1B slice
reported phase-B 2.10 → 1.98 (repo discussion #1) — the evidence that low
loss is reachable when the protocol is followed without detours.
Lineage
DeepSeek-V4-Flash → -from-flash → -recovered →
cpt-fineweb →
sft-v1 → this