deepseek-v4-mini-300M-cpt-fineweb

Experimental / research artifact — not a usable model. Broad-data CPT of deepseek-v4-mini-300M-recovered on FineWeb-Edu, run locally on a single RTX 5090 laptop (2026-07-14).

What this is

The three-phase Lightning-Indexer protocol (see the recovery kit) applied with broad web text instead of the parent-family data — an experiment in whether the recovered slice generalizes off its native distribution at a ~98M-token budget.

Phase Steps Tokens Result
0 dense 2000 65.5M loss 9.05 → ~7.0
A indexer warmup 300 KL → 0.023
B sparse (top-k on) 700 22.9M lm 6.66–6.97, stable

Config: seq 1024, batch 4, grad-accum 8, Muon, WSD, bf16. W&B: m4mo8jd6

Findings (why this recipe is NOT recommended)

  • Mechanically clean: phases A/B ran stably on broad data — useful evidence for scaling the protocol.
  • Distribution tax: downstream SFT on the parent-family data (loggenix/DeepSeek-R1-style traces) restarted at loss 8.18 vs ~3.8 for the H100 run that skipped the broad-data detour. For this model family, train directly on parent-distribution data (loggenix_moe_sequelbox_sft_v1).
  • Generations at loss ~7 are incoherent (expected at this budget).

Lineage

DeepSeek-V4-Flash → 300M slice (-from-flash) → indexer recovery (-recovered) → thissft-v1sft-v2

Downloads last month
113
Safetensors
Model size
0.3B params
Tensor type
I64
·
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kshitijthakkar/deepseek-v4-mini-300M-cpt-fineweb

Finetuned
(1)
this model
Finetunes
1 model

Dataset used to train kshitijthakkar/deepseek-v4-mini-300M-cpt-fineweb