smollm2-1.7b-100B-linear-merge-epe-no_bce-w0.7-normal-w0.3

Linear weight-merge of two SmolLM2-1.7B pretraining checkpoints (each trained on 100B tokens).

merged = 0.7 * EPE + 0.3 * NORMAL
Component Weight Source checkpoint
EPE (no_bce) 0.7 epe-1p-smollm-1p7b-100B-20n-2048sl-960gbsz-no_bce
NORMAL 0.3 normal-smollm-1p7b-100B-20n-2048sl-960gbsz

Merge details

  • Method: element-wise linear interpolation of weights (computed in fp32, stored in bf16).
  • Architecture: LlamaForCausalLM, identical for both parents (hidden 2048, 24 layers, 32 heads, tied embeddings).
  • Vocabulary: 49280. The two parents share 49152 tokens (interpolated); the 128 extra EPE "charter" tokens are carried over from the EPE parent unchanged. The EPE tokenizer is bundled.
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support