File size: 3,174 Bytes
096a6fb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
---
library_name: nanochat
license: apache-2.0
tags:
- nanochat
- nanochat_gpt
- pt_fineweb-nanochatbpe-20B
- ppt_github-code-nanochatbpe-1B
- seed_1
- case_b
- lr_trapezoid
- depth_12
---

# code_5pct_1e18_seed1_2026-07-12_01-12-46_897145-pt

Trained with [nanochat](https://github.com/karpathy/nanochat). Checkpoint at step **3,368**.

**W&B run:** https://wandb.ai/alexksternteam/optimal_seed_replicas_v1/runs/33ea3a6t

## Outcome

| metric | value |
| --- | --- |
| `step` | 3368 |
| `smooth_train_loss` | 3.2399237155914307 |
| `min_objective` | 0.9851442932664687 |
| `flops_used` | 9.498592991426642e+17 |
| `flops_per_token` | 1075838976.0 |
| `total_training_time` | 422.86022329330444 |

## Training config

```json
{
  "model_pt": {
    "sequence_len": 2048,
    "vocab_size": 65536,
    "n_layer": 12,
    "n_head": 6,
    "n_kv_head": 6,
    "n_embd": 768
  },
  "model_ppt": {
    "sequence_len": 2048,
    "vocab_size": 65536,
    "n_layer": 12,
    "n_head": 6,
    "n_kv_head": 6,
    "n_embd": 768
  },
  "optim": {
    "matrix_lr": 0.015,
    "embedding_lr": 0.3,
    "unembedding_lr": 0.004,
    "weight_decay": 0.0
  },
  "train": {
    "num_iterations": 1000,
    "eval_every_n_steps": null,
    "eval_every_flops_frac": null,
    "save_every_n_steps": null,
    "log_every_n_steps": 100,
    "eval_at_end": true,
    "save_at_end": true,
    "grad_clip": 1.0,
    "ema_beta": 0.0
  },
  "eval": {
    "eval_steps": 8,
    "eval_tokens": 2016000
  },
  "hardware": {
    "device_batch_size": 32,
    "grad_accum_steps": 1,
    "peak_tflops": 2250.0
  },
  "wandb": {
    "enabled": true,
    "notes": "",
    "group": "1e18_ppt0.05",
    "tags": [
      "nanochat_gpt",
      "pt_fineweb-nanochatbpe-20B",
      "ppt_github-code-nanochatbpe-1B",
      "seed_1",
      "case_b",
      "lr_trapezoid",
      "depth_12"
    ],
    "entity": null
  },
  "data_pt": "fineweb-nanochatbpe-20B",
  "data_ppt": "github-code-nanochatbpe-1B",
  "extra_eval": [
    "c4-nanochatbpe-10B"
  ],
  "train_split": "train",
  "pad_vocab": 65536,
  "ppt_vocab_size": null,
  "ppt_same_vocab_as_pt": true,
  "ppt_unified_lr": true,
  "lr_kind": "trapezoid",
  "lr_warmup_ratio": 0.0,
  "lr_warmdown_ratio": 0.4,
  "lr_final_frac": 0.0,
  "ppt_lr_kind": "trapezoid",
  "ppt_lr_warmup_ratio": null,
  "ppt_lr_warmdown_ratio": null,
  "ppt_lr_final_frac": 0.0,
  "ppt_lr": null,
  "ppt_weight_decay": 0.0,
  "ppt_device_batch_size": null,
  "ppt_grad_accum_steps": null,
  "eval_tokens": 2016000,
  "reinit_embed_at_transition": false,
  "reset_optimizer_at_transition": false,
  "depth": 12,
  "compile_model": true,
  "model_project": "optimal_seed_replicas_v1",
  "run_name": "code_5pct_1e18_seed1",
  "target_flops": 1e+18,
  "alpha_ppt": 0.05,
  "use_measured_flops": true,
  "flops_per_token": null,
  "ppt_flops_per_token": null,
  "seed": 1,
  "metrics_output_file": null,
  "push_to_hf": true,
  "hf_repo_org": "alexkstern",
  "cleanup_checkpoints": true
}
```

## Files

- `model_003368.pt` — model weights (`state_dict`).
- `meta_003368.json` — training metadata.
- `config_003368.json` — run config snapshot.
- `rng_003368.pt` — RNG state (when present).