looped-moe-1e18-d704-seed42

16 effective layers, trained at a 1e18 FLOP budget at this architecture's compute-optimal width. One of a six-seed set (seeds 42-47) in which only the training data order varies; model initialisation is fixed across seeds.

field value
architecture looped-moe
d_model 704
d_ff 1920
width_ratio 5.5
base_d_model 128
base_d_ff 384
data seed 42
init seed 42

Loads with trust_remote_code=True. Trained with context length 1024; set max_length=1024 when evaluating.

Downloads last month
2
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support