File size: 3,174 Bytes
096a6fb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | ---
library_name: nanochat
license: apache-2.0
tags:
- nanochat
- nanochat_gpt
- pt_fineweb-nanochatbpe-20B
- ppt_github-code-nanochatbpe-1B
- seed_1
- case_b
- lr_trapezoid
- depth_12
---
# code_5pct_1e18_seed1_2026-07-12_01-12-46_897145-pt
Trained with [nanochat](https://github.com/karpathy/nanochat). Checkpoint at step **3,368**.
**W&B run:** https://wandb.ai/alexksternteam/optimal_seed_replicas_v1/runs/33ea3a6t
## Outcome
| metric | value |
| --- | --- |
| `step` | 3368 |
| `smooth_train_loss` | 3.2399237155914307 |
| `min_objective` | 0.9851442932664687 |
| `flops_used` | 9.498592991426642e+17 |
| `flops_per_token` | 1075838976.0 |
| `total_training_time` | 422.86022329330444 |
## Training config
```json
{
"model_pt": {
"sequence_len": 2048,
"vocab_size": 65536,
"n_layer": 12,
"n_head": 6,
"n_kv_head": 6,
"n_embd": 768
},
"model_ppt": {
"sequence_len": 2048,
"vocab_size": 65536,
"n_layer": 12,
"n_head": 6,
"n_kv_head": 6,
"n_embd": 768
},
"optim": {
"matrix_lr": 0.015,
"embedding_lr": 0.3,
"unembedding_lr": 0.004,
"weight_decay": 0.0
},
"train": {
"num_iterations": 1000,
"eval_every_n_steps": null,
"eval_every_flops_frac": null,
"save_every_n_steps": null,
"log_every_n_steps": 100,
"eval_at_end": true,
"save_at_end": true,
"grad_clip": 1.0,
"ema_beta": 0.0
},
"eval": {
"eval_steps": 8,
"eval_tokens": 2016000
},
"hardware": {
"device_batch_size": 32,
"grad_accum_steps": 1,
"peak_tflops": 2250.0
},
"wandb": {
"enabled": true,
"notes": "",
"group": "1e18_ppt0.05",
"tags": [
"nanochat_gpt",
"pt_fineweb-nanochatbpe-20B",
"ppt_github-code-nanochatbpe-1B",
"seed_1",
"case_b",
"lr_trapezoid",
"depth_12"
],
"entity": null
},
"data_pt": "fineweb-nanochatbpe-20B",
"data_ppt": "github-code-nanochatbpe-1B",
"extra_eval": [
"c4-nanochatbpe-10B"
],
"train_split": "train",
"pad_vocab": 65536,
"ppt_vocab_size": null,
"ppt_same_vocab_as_pt": true,
"ppt_unified_lr": true,
"lr_kind": "trapezoid",
"lr_warmup_ratio": 0.0,
"lr_warmdown_ratio": 0.4,
"lr_final_frac": 0.0,
"ppt_lr_kind": "trapezoid",
"ppt_lr_warmup_ratio": null,
"ppt_lr_warmdown_ratio": null,
"ppt_lr_final_frac": 0.0,
"ppt_lr": null,
"ppt_weight_decay": 0.0,
"ppt_device_batch_size": null,
"ppt_grad_accum_steps": null,
"eval_tokens": 2016000,
"reinit_embed_at_transition": false,
"reset_optimizer_at_transition": false,
"depth": 12,
"compile_model": true,
"model_project": "optimal_seed_replicas_v1",
"run_name": "code_5pct_1e18_seed1",
"target_flops": 1e+18,
"alpha_ppt": 0.05,
"use_measured_flops": true,
"flops_per_token": null,
"ppt_flops_per_token": null,
"seed": 1,
"metrics_output_file": null,
"push_to_hf": true,
"hf_repo_org": "alexkstern",
"cleanup_checkpoints": true
}
```
## Files
- `model_003368.pt` — model weights (`state_dict`).
- `meta_003368.json` — training metadata.
- `config_003368.json` — run config snapshot.
- `rng_003368.pt` — RNG state (when present).
|