Raghav-Singhal commited on
Commit
15b3d30
·
verified ·
1 Parent(s): 7089cdd

Add 1pp-0.5b-asst-sft (1PP)

Browse files
README.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ library_name: transformers
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - 1pp
9
+ - one-persona-pretraining
10
+ - sft
11
+ - asst
12
+ ---
13
+
14
+ # 1pp-0.5b-asst-sft
15
+
16
+ **One Persona Pretraining (1PP)** experiment model: 0.58B parameters, pretraining condition **rewritten conversations, assistant-turn loss**, followed by supervised fine-tuning (SFT).
17
+
18
+ Part of a 3 × 3 study: three sizes (0.5B, 1B, 1.7B) × three pretraining conditions on the same 47.8M source documents in the same order (original documents; rewritten conversations with loss on assistant turns; rewritten conversations with loss on user and assistant turns). Every run saw the identical batch sequence, so the conditions differ only in the document text and the loss mask. Models are grouped in the [1pp collection](https://huggingface.co/collections/Raghav-Singhal/1pp-6a999df54bfcf9335355a649).
19
+
20
+ ## Architecture
21
+
22
+ Llama-style decoder, 24 layers, hidden 1,152, FFN 4,608 (SwiGLU), attention heads / KV heads 9 / 3 (head dim 128), RMSNorm, RoPE base 10,000, untied embeddings, no biases, no QK-norm, sequence length 4,096. Tokenizer: SmolLM2 vocabulary (49,152) plus `<|pad|>`; `<|endoftext|>` is the end-of-document token.
23
+
24
+ ## Pretraining
25
+
26
+ Data: the 1PP conversations rewritten from those documents; loss only on assistant turns (no loss on user turns or `<|endoftext|>`). One pass over 47.8M documents (66.2B tokens of original documents; 63.0B tokens as conversations), 31,777 steps at global batch 512 × 4,096 tokens, cross-document attention masking, best-fit packing with step-aligned document assignment. Optimizer: Muon (shape scaling, matrix LR 0.005) with Adam for embeddings and norms, warmup 2,000 steps, constant, linear decay over the last 10% to 1/100, weight decay 0.1, bf16.
27
+
28
+ Validation loss (per token, 2,433 held-out documents, final checkpoint):
29
+
30
+ | assistant text | user text | document text |
31
+ |---|---|---|
32
+ | 1.579 | 6.878 | 3.372 |
33
+
34
+ ## Supervised fine-tuning
35
+
36
+ One epoch over a 400k-conversation mix: `jkminder/model-raising-pb-100k-3c-mt-sft` (98.5k multi-turn, constitution-cited track), `dlab-spp/sp-sft-normal-300k` minus prompts duplicated in the first set (271.6k), and a 30k sample of `dlab-spp/sp-sft-safety-180k`. Same stack as pretraining (Megatron, Muon, ChatML without a system turn, loss on assistant turns only). Matrix LR 0.002 selected per model from {0.0005, 0.001, 0.002, 0.005} by held-out loss; global batch 128 × 4,096, linear decay to 1/10 after 3% warmup.
37
+
38
+ Held-out SFT loss (assistant tokens, 1,998 held-out conversations): 2.023
39
+
40
+ ## Chat format
41
+
42
+ ChatML **without a system turn** (the models never saw one):
43
+
44
+ ```
45
+ <|im_start|>user\n{message}<|im_end|>\n<|im_start|>assistant\n{reply}<|im_end|>\n
46
+ ```
47
+
48
+ The bundled `chat_template` renders exactly this. Generation stops at `<|im_end|>`.
49
+
50
+ ## Verification
51
+
52
+ The HF weights were checked against the Megatron checkpoint by recomputing validation losses with this model:
53
+
54
+ | set | HF loss | Megatron reference | abs. diff |
55
+ |---|---|---|---|
56
+ | sft_val segments [3, 4] | 2.0234 | 2.0234 | 0.0000 |
57
+
58
+ ## Links
59
+
60
+ - Training logs: wandb projects [1pp-training](https://wandb.ai/raghav_singhal/1pp-training) and [1pp-sft](https://wandb.ai/raghav_singhal/1pp-sft)
61
+ - Research artifact from the 1PP project (EPFL DLAB); not a general-purpose assistant.
added_tokens.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ {
2
+ "<|pad|>": 49152
3
+ }
chat_template.jinja ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {% for message in messages %}{{ '<|im_start|>' + message['role'] + '
2
+ ' + message['content'] + '<|im_end|>' + '
3
+ ' }}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
4
+ ' }}{% endif %}
config.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "LlamaForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 1,
8
+ "dtype": "bfloat16",
9
+ "eos_token_id": [
10
+ 2,
11
+ 0
12
+ ],
13
+ "head_dim": 128,
14
+ "hidden_act": "silu",
15
+ "hidden_size": 1152,
16
+ "initializer_range": 0.02,
17
+ "intermediate_size": 4608,
18
+ "max_position_embeddings": 4096,
19
+ "mlp_bias": false,
20
+ "model_type": "llama",
21
+ "num_attention_heads": 9,
22
+ "num_hidden_layers": 24,
23
+ "num_key_value_heads": 3,
24
+ "pad_token_id": 49152,
25
+ "pretraining_tp": 1,
26
+ "rms_norm_eps": 1e-05,
27
+ "rope_scaling": null,
28
+ "rope_theta": 10000,
29
+ "tie_word_embeddings": false,
30
+ "transformers_version": "4.57.6",
31
+ "use_cache": true,
32
+ "vocab_size": 49153
33
+ }
convert.log ADDED
@@ -0,0 +1,1020 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ WARNING: FI_CXI_RDZV_PROTO=alt_read is configured, but the Slurm network is
3
+ not set to disable rendezvous GET. Performance/stability may be impacted.
4
+
5
+ Expected Slurm launch:
6
+ srun --network=disable_rdzv_get ...
7
+ or
8
+ SLURM_NETWORK=disable_rdzv_get srun ...
9
+
10
+ Context:
11
+ SLURM_NETWORK=<unset>
12
+ FI_PROVIDER=cxi
13
+ FI_CXI_RDZV_PROTO=alt_read
14
+
15
+ Overriding a previously registered kernel for the same operator and the same dispatch key
16
+ operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
17
+ registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922
18
+ dispatch key: ADInplaceOrView
19
+ previous kernel: no debug info
20
+ new kernel: registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922 (Triggered internally at /opt/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
21
+ self.m.impl(
22
+ Loaded loader_core as the loader.
23
+ Loaded saver_llama_hf as the saver.
24
+ Starting saver...
25
+ Overriding a previously registered kernel for the same operator and the same dispatch key
26
+ operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
27
+ registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922
28
+ dispatch key: ADInplaceOrView
29
+ previous kernel: no debug info
30
+ new kernel: registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922 (Triggered internally at /opt/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
31
+ self.m.impl(
32
+ Starting loader...
33
+ Setting num_layers to 24 from checkpoint
34
+ Setting hidden_size to 1152 from checkpoint
35
+ Setting ffn_hidden_size to 4608 from checkpoint
36
+ Setting seq_length to 4096 from checkpoint
37
+ Setting num_attention_heads to 9 from checkpoint
38
+ Setting num_query_groups to 3 from checkpoint
39
+ Setting group_query_attention to True from checkpoint
40
+ Setting kv_channels to 128 from checkpoint
41
+ Setting max_position_embeddings to 4096 from checkpoint
42
+ Setting position_embedding_type to rope from checkpoint
43
+ Setting add_position_embedding to True from checkpoint
44
+ Setting use_rotary_position_embeddings to False from checkpoint
45
+ Setting rotary_base to 10000 from checkpoint
46
+ Setting rotary_percent to 1.0 from checkpoint
47
+ Setting rotary_interleaved to False from checkpoint
48
+ Setting add_bias_linear to False from checkpoint
49
+ Setting add_qkv_bias to False from checkpoint
50
+ Setting squared_relu to False from checkpoint
51
+ Setting swiglu to True from checkpoint
52
+ Setting ssglu to False from checkpoint
53
+ Setting reglu to False from checkpoint
54
+ Setting rlglu to False from checkpoint
55
+ Setting sssglu to False from checkpoint
56
+ Setting lglu to False from checkpoint
57
+ Setting situ to False from checkpoint
58
+ Setting pnglu to False from checkpoint
59
+ Setting pnglu_fusion to True from checkpoint
60
+ Setting untie_embeddings_and_output_weights to True from checkpoint
61
+ Setting scale_embeddings_by_sqrt_hidden to False from checkpoint
62
+ Checkpoint did not provide arguments apply_layernorm_1p
63
+ Setting normalization to RMSNorm from checkpoint
64
+ Setting residual_output_scaling to False from checkpoint
65
+ Setting sandwich_norm to False from checkpoint
66
+ Setting keel to False from checkpoint
67
+ Checkpoint did not provide arguments keel_alpha
68
+ Setting apply_query_key_layer_scaling to False from checkpoint
69
+ Setting qk_layernorm to False from checkpoint
70
+ Setting attention_dropout to 0.0 from checkpoint
71
+ Setting hidden_dropout to 0.0 from checkpoint
72
+ Checkpoint did not provide arguments window_size
73
+ Checkpoint did not provide arguments window_attn_skip_freq
74
+ Checkpoint did not provide arguments no_rope_freq
75
+ Checkpoint did not provide arguments mtp_hybrid_override_pattern
76
+ Checkpoint did not provide arguments mtp_num_layers
77
+ Setting mtp_use_repeated_layer to False from checkpoint
78
+ Checkpoint did not provide arguments spec
79
+ Checkpoint did not provide arguments num_experts
80
+ Checkpoint did not provide arguments mtp_num_layers
81
+ Setting moe_layer_freq to 1 from checkpoint
82
+ Setting moe_router_topk to 2 from checkpoint
83
+ Setting moe_router_pre_softmax to False from checkpoint
84
+ Setting moe_grouped_gemm to False from checkpoint
85
+ Checkpoint did not provide arguments moe_shared_expert_intermediate_size
86
+ Setting moe_router_score_function to softmax from checkpoint
87
+ Setting moe_router_enable_expert_bias to False from checkpoint
88
+ Checkpoint did not provide arguments moe_router_topk_scaling_factor
89
+ Setting moe_router_load_balancing_type to aux_loss from checkpoint
90
+ Setting moe_router_quantile_balancing_method to histogram from checkpoint
91
+ Setting multi_latent_attention to False from checkpoint
92
+ Checkpoint did not provide arguments q_lora_rank
93
+ Setting kv_lora_rank to 32 from checkpoint
94
+ Setting qk_head_dim to 128 from checkpoint
95
+ Setting qk_pos_emb_head_dim to 64 from checkpoint
96
+ Setting v_head_dim to 128 from checkpoint
97
+ Setting rotary_scaling_factor to 1.0 from checkpoint
98
+ Setting mamba_state_dim to 128 from checkpoint
99
+ Setting mamba_head_dim to 64 from checkpoint
100
+ Setting mamba_num_groups to 8 from checkpoint
101
+ Checkpoint did not provide arguments mamba_num_heads
102
+ Checkpoint did not provide arguments hybrid_layer_pattern
103
+ Checkpoint did not provide arguments heterogeneous_layers_config_path
104
+ Checkpoint did not provide arguments heterogeneous_layers_config_encoded_json
105
+ Checkpoint did not provide arguments moe_latent_size
106
+ Setting scale_embeddings_by_sqrt_hidden to False from checkpoint
107
+ Setting residual_output_scaling to False from checkpoint
108
+ Setting sandwich_norm to False from checkpoint
109
+ Setting keel to False from checkpoint
110
+ Checkpoint did not provide arguments keel_alpha
111
+ Setting pnglu to False from checkpoint
112
+ Setting pnglu_fusion to True from checkpoint
113
+ Setting pn3glu to False from checkpoint
114
+ Setting xpr to False from checkpoint
115
+ Setting gxpr to False from checkpoint
116
+ Setting gxpry to False from checkpoint
117
+ Setting gxprv2 to False from checkpoint
118
+ Setting xr2 to False from checkpoint
119
+ Setting gxr2 to False from checkpoint
120
+ Setting xr2glu to False from checkpoint
121
+ Setting xssglu to False from checkpoint
122
+ Setting polynorm to False from checkpoint
123
+ Setting qk_layernorm to False from checkpoint
124
+ Setting attention_output_gate to False from checkpoint
125
+ Setting tokenizer_model to /capstor/store/cscs/swissai/infra01/users/rsinghal/1pp-training/tokenizers/smollm2_pretrain from checkpoint
126
+ Setting tokenizer_type to HuggingFaceTokenizer from checkpoint
127
+ Checkpoint did not provide arguments tiktoken_pattern
128
+ Setting padded_vocab_size to 49280 from checkpoint
129
+ Setting tensor_model_parallel_size to 1 from checkpoint
130
+ Setting pipeline_model_parallel_size to 1 from checkpoint
131
+ Checkpoint did not provide arguments virtual_pipeline_model_parallel_size
132
+ Checkpoint did not provide arguments num_layers_per_virtual_pipeline_stage
133
+ Setting expert_model_parallel_size to 1 from checkpoint
134
+ using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: None, tensor-model-parallel size: 1, pipeline-model-parallel size: 1
135
+ setting global batch size to 1
136
+ Number of virtual stages per pipeline stage: None
137
+ accumulate and all-reduce gradients in fp32 for bfloat16 data type.
138
+ using torch.bfloat16 for parameters ...
139
+ ------------------------ arguments ------------------------
140
+ account_for_embedding_in_pipeline_split ......... False
141
+ account_for_loss_in_pipeline_split .............. False
142
+ accumulate_allreduce_grads_in_fp32 .............. True
143
+ activation_func_clamp_value ..................... None
144
+ activation_func_fp8_input_store ................. False
145
+ adam_beta1 ...................................... 0.9
146
+ adam_beta2 ...................................... 0.999
147
+ adam_eps ........................................ 1e-08
148
+ add_bias_linear ................................. False
149
+ add_position_embedding .......................... True
150
+ add_qkv_bias .................................... False
151
+ adlr_autoresume ................................. False
152
+ adlr_autoresume_interval ........................ 1000
153
+ align_grad_reduce ............................... True
154
+ align_param_gather .............................. False
155
+ allow_ambiguous_pad_tokens ...................... False
156
+ app_tag_run_name ................................ None
157
+ app_tag_run_version ............................. 0.0.0
158
+ apply_query_key_layer_scaling ................... False
159
+ apply_residual_connection_post_layernorm ........ False
160
+ apply_rope_fusion ............................... True
161
+ apply_wd_to_qk_layernorm ........................ False
162
+ async_ckpt_cpu_priority ......................... 10
163
+ async_ckpt_io_priority .......................... 3
164
+ async_save ...................................... False
165
+ async_strategy .................................. nvrx
166
+ attention_backend ............................... AttnBackend.auto
167
+ attention_dropout ............................... 0.0
168
+ attention_output_gate ........................... False
169
+ attention_softmax_in_fp32 ....................... False
170
+ auto_detect_ckpt_format ......................... False
171
+ barrier_with_L1_time ............................ True
172
+ batch_invariant_mode ............................ False
173
+ bert_binary_head ................................ True
174
+ bert_embedder_type .............................. megatron
175
+ bert_load ....................................... None
176
+ bf16 ............................................ True
177
+ bias_dropout_fusion ............................. False
178
+ bias_gelu_fusion ................................ False
179
+ bias_swiglu_fusion .............................. True
180
+ biencoder_projection_dim ........................ 0
181
+ biencoder_shared_query_context_model ............ False
182
+ block_data_path ................................. None
183
+ cache_mla_latents ............................... False
184
+ calc_ft_timeouts ................................ False
185
+ calculate_per_token_loss ........................ False
186
+ check_for_large_grads ........................... False
187
+ check_for_nan_in_loss_and_grad .................. True
188
+ check_for_spiky_loss ............................ False
189
+ check_grad_norm ................................. False
190
+ check_grad_norm_threshold ....................... 5.0
191
+ check_weight_hash_across_dp_replicas_interval ... None
192
+ ckpt_assume_constant_structure .................. False
193
+ ckpt_convert_format ............................. None
194
+ ckpt_convert_save ............................... None
195
+ ckpt_convert_update_legacy_dist_opt_format ...... False
196
+ ckpt_format ..................................... torch_dist
197
+ ckpt_fully_parallel_load ........................ False
198
+ ckpt_fully_parallel_load_exchange_algo .......... broadcast
199
+ ckpt_fully_parallel_load_process_group .......... dp
200
+ ckpt_fully_parallel_save ........................ True
201
+ ckpt_fully_parallel_save_deprecated ............. False
202
+ ckpt_fully_parallel_save_process_group .......... dp
203
+ ckpt_step ....................................... None
204
+ classes_fraction ................................ 1.0
205
+ clip_grad ....................................... 1.0
206
+ clone_scatter_output_in_embedding ............... True
207
+ config_logger_dir ...............................
208
+ consumed_train_samples .......................... 0
209
+ consumed_valid_samples .......................... 0
210
+ context_parallel_size ........................... 1
211
+ cp_comm_type .................................... ['p2p']
212
+ cpu_offloading_num_layers ....................... 0
213
+ cpu_offloading_retain_pinned_cpu_buffers ........ False
214
+ create_all_gather_group ......................... False
215
+ create_attention_mask_in_dataloader ............. True
216
+ cross_entropy_fusion_impl ....................... native
217
+ cross_entropy_loss_fusion ....................... False
218
+ cuda_graph_impl ................................. none
219
+ cuda_graph_scope ................................ []
220
+ cuda_graph_warmup_steps ......................... 3
221
+ data_args_path .................................. None
222
+ data_cache_path ................................. None
223
+ data_parallel_random_init ....................... False
224
+ data_parallel_sharding_strategy ................. no_shard
225
+ data_parallel_size .............................. 1
226
+ data_path ....................................... None
227
+ data_per_class_fraction ......................... 1.0
228
+ data_sharding ................................... True
229
+ dataloader_defer_npy_index_mmap ................. False
230
+ dataloader_fast_cache_load ...................... False
231
+ dataloader_inter_document_masking ............... False
232
+ dataloader_type ................................. single
233
+ ddp_average_in_collective ....................... False
234
+ ddp_bucket_size ................................. None
235
+ ddp_num_buckets ................................. None
236
+ ddp_pad_buckets_for_high_nccl_busbw ............. False
237
+ ddp_param_name_patterns_for_fp32_local_accumulation []
238
+ ddp_reduce_scatter_with_fp32_accumulation ....... False
239
+ decode_only_cuda_graphs ......................... False
240
+ decoder_first_pipeline_num_layers ............... None
241
+ decoder_last_pipeline_num_layers ................ None
242
+ decoder_num_layers .............................. None
243
+ decoder_seq_length .............................. None
244
+ decoupled_lr .................................... None
245
+ decoupled_min_lr ................................ None
246
+ decrease_batch_size_if_needed ................... False
247
+ defer_embedding_wgrad_compute ................... False
248
+ delay_wgrad_compute ............................. False
249
+ deprecated_use_mcore_models ..................... False
250
+ deterministic_mode .............................. False
251
+ dino_bottleneck_size ............................ 256
252
+ dino_freeze_last_layer .......................... 1
253
+ dino_head_hidden_size ........................... 2048
254
+ dino_local_crops_number ......................... 10
255
+ dino_local_img_size ............................. 96
256
+ dino_norm_last_layer ............................ False
257
+ dino_teacher_temp ............................... 0.07
258
+ dino_warmup_teacher_temp ........................ 0.04
259
+ dino_warmup_teacher_temp_epochs ................. 30
260
+ disable_bf16_reduced_precision_matmul ........... False
261
+ disable_jit_fuser ............................... False
262
+ disable_straggler_on_startup .................... False
263
+ disable_symmetric_registration .................. False
264
+ dist_ckpt_format_deprecated ..................... None
265
+ dist_ckpt_optim_fully_reshardable ............... False
266
+ dist_ckpt_save_pre_mcore_014 .................... False
267
+ dist_ckpt_strictness ............................ assume_ok_unexpected
268
+ dist_ckpt_workers ............................... 1
269
+ distrib_optim_fully_reshardable_mem_efficient ... False
270
+ distribute_saved_activations .................... False
271
+ distributed_backend ............................. nccl
272
+ distributed_timeout_minutes ..................... 10
273
+ distributed_timeout_seconds_after_init .......... None
274
+ dsa_indexer_head_dim ............................ None
275
+ dsa_indexer_loss_coeff .......................... None
276
+ dsa_indexer_n_heads ............................. None
277
+ dsa_indexer_topk ................................ None
278
+ dsa_indexer_use_sparse_loss ..................... False
279
+ dump_param_to_param_group_map ................... None
280
+ embedding_init_method_std ....................... None
281
+ embedding_lr_multiplier ......................... None
282
+ embedding_path .................................. None
283
+ empty_unused_memory_level ....................... 0
284
+ enable_chunked_prefill .......................... False
285
+ enable_cuda_graph ............................... False
286
+ enable_experimental ............................. False
287
+ enable_ft_package ............................... False
288
+ enable_full_sharding_in_hsdp .................... False
289
+ enable_msc ...................................... True
290
+ enable_one_logger ............................... False
291
+ encoder_num_layers .............................. 24
292
+ encoder_seq_length .............................. 4096
293
+ end_weight_decay ................................ 0.01
294
+ eod_mask_loss ................................... False
295
+ ep_overlap_early_attn_memory_release ............ False
296
+ error_injection_rate ............................ 0
297
+ error_injection_type ............................ transient_error
298
+ eval_interval ................................... None
299
+ eval_iters ...................................... 100
300
+ evidence_data_path .............................. None
301
+ exit_duration_in_mins ........................... None
302
+ exit_interval ................................... None
303
+ exit_on_missing_checkpoint ...................... True
304
+ exit_signal ..................................... 15
305
+ exit_signal_handler ............................. False
306
+ exit_signal_handler_for_dataloader .............. False
307
+ exit_signal_handler_for_training ................ False
308
+ exp_avg_dtype ................................... torch.float32
309
+ exp_avg_sq_dtype ................................ torch.float32
310
+ experimental_attention_variant .................. None
311
+ expert_model_parallel_size ...................... 1
312
+ expert_tensor_parallel_size ..................... 1
313
+ external_cuda_graph ............................. False
314
+ fake_process_group .............................. False
315
+ ffn_hidden_size ................................. 4608
316
+ fim_data ........................................ False
317
+ fim_eod_token ................................... <|endoftext|>
318
+ fim_fragment_rate ............................... None
319
+ fim_middle_token ................................ <fim_middle>
320
+ fim_no_prefix ................................... None
321
+ fim_pad_token ................................... <fim_pad>
322
+ fim_prefix_token ................................ <fim_prefix>
323
+ fim_rate ........................................ 0.5
324
+ fim_split_sample ................................ None
325
+ fim_spm_rate .................................... 0.5
326
+ fim_suffix_token ................................ <fim_suffix>
327
+ fine_grained_activation_offloading .............. False
328
+ finetune ........................................ False
329
+ first_last_layers_bf16 .......................... False
330
+ flash_decode .................................... False
331
+ flight_recorder_dump_on_timeout ................. True
332
+ flight_recorder_dump_path ....................... None
333
+ flight_recorder_extra_dump_on_exec .............. True
334
+ flight_recorder_include_only_active ............. True
335
+ flight_recorder_include_stack_trace ............. False
336
+ flight_recorder_trace_buffer_size ............... 2000
337
+ fp16 ............................................ False
338
+ fp16_lm_cross_entropy ........................... False
339
+ fp32_residual_connection ........................ False
340
+ fp4 ............................................. None
341
+ fp4_param ....................................... False
342
+ fp4_quantizer_factory ........................... None
343
+ fp4_recipe ...................................... nvfp4
344
+ fp8 ............................................. None
345
+ fp8_amax_compute_algo ........................... most_recent
346
+ fp8_amax_history_len ............................ 1
347
+ fp8_interval .................................... 1
348
+ fp8_margin ...................................... 0
349
+ fp8_param_gather ................................ False
350
+ fp8_quantizer_factory ........................... None
351
+ fp8_recipe ...................................... delayed
352
+ fp8_wgrad ....................................... True
353
+ fsdp_double_buffer .............................. False
354
+ fsdp_manual_registration ........................ False
355
+ ft_num_warmup_iters ............................. 5
356
+ full_validation ................................. False
357
+ fused_residual_rmsnorm .......................... False
358
+ gain_parametrization ............................ softplus
359
+ gains_lr ........................................ None
360
+ gains_no_clamp_min .............................. False
361
+ global_batch_size ............................... 1
362
+ glu_linear_offset ............................... 0.0
363
+ goldfish_h ...................................... 50
364
+ goldfish_k ...................................... 50
365
+ goldfish_loss ................................... False
366
+ grad_reduce_in_bf16 ............................. False
367
+ gradient_accumulation_fusion .................... True
368
+ gradient_reduce_div_fusion ...................... True
369
+ group_query_attention ........................... True
370
+ grpo_clamp_eps_lower ............................ 0.01
371
+ grpo_clamp_eps_upper ............................ 0.01
372
+ grpo_entropy_term_weight ........................ 0.0
373
+ grpo_filter_groups_with_same_reward ............. False
374
+ grpo_group_size ................................. 2
375
+ grpo_iterations ................................. 2
376
+ grpo_kl_beta .................................... 0.001
377
+ grpo_prompts_per_step ........................... 32
378
+ gxpr ............................................ False
379
+ gxpr_fusion ..................................... True
380
+ gxprv2 .......................................... False
381
+ gxpry ........................................... False
382
+ gxr2 ............................................ False
383
+ head_lr_mult .................................... 1.0
384
+ heterogeneous_layers_config_encoded_json ........ None
385
+ heterogeneous_layers_config_path ................ None
386
+ hidden_dropout .................................. 0.0
387
+ hidden_size ..................................... 1152
388
+ hierarchical_context_parallel_sizes ............. None
389
+ high_priority_stream_groups ..................... []
390
+ hybrid_context_parallel ......................... False
391
+ hybrid_layer_pattern ............................ None
392
+ hybrid_override_pattern ......................... None
393
+ hypersphere_embedding_mode ...................... row
394
+ hypersphere_gains_mode .......................... rowcol
395
+ hypersphere_gains_mode_embedding ................ none
396
+ hypersphere_gains_mode_output ................... inherit
397
+ hypersphere_gains_mode_router ................... rowcol
398
+ hypersphere_mode ................................ flat
399
+ hypersphere_preserve_init ....................... False
400
+ hypersphere_radius_mode ......................... shape_native
401
+ hypersphere_router_mode ......................... row
402
+ hypersphere_scale_out_proj_init ................. False
403
+ hypersphere_tangential_grad ..................... False
404
+ hysteresis ...................................... 2
405
+ ict_head_size ................................... None
406
+ ict_load ........................................ None
407
+ img_h ........................................... 224
408
+ img_w ........................................... 224
409
+ indexer_batch_size .............................. 128
410
+ indexer_log_interval ............................ 1000
411
+ inference_batch_times_seqlen_threshold .......... -1
412
+ inference_coordinator_port ...................... None
413
+ inference_disable_triton_nvls_kernels ........... False
414
+ inference_dynamic_batching ...................... False
415
+ inference_dynamic_batching_block_size ........... 256
416
+ inference_dynamic_batching_buffer_size_gb ....... 40.0
417
+ inference_dynamic_batching_cuda_graph_max_tokens 16384
418
+ inference_dynamic_batching_cuda_graph_mixed_prefill_count 16
419
+ inference_dynamic_batching_enable_prefix_caching False
420
+ inference_dynamic_batching_mamba_memory_ratio ... None
421
+ inference_dynamic_batching_max_requests ......... None
422
+ inference_dynamic_batching_max_tokens ........... None
423
+ inference_dynamic_batching_num_cuda_graphs ...... 16
424
+ inference_dynamic_batching_paused_buffer_size_gb None
425
+ inference_dynamic_batching_prefix_caching_coordinator_policy first_prefix_block
426
+ inference_dynamic_batching_prefix_caching_eviction_policy ref_zero
427
+ inference_dynamic_batching_prefix_caching_mamba_gb None
428
+ inference_dynamic_batching_prefix_caching_routing_alpha 0.5
429
+ inference_dynamic_batching_track_generated_token_events False
430
+ inference_dynamic_batching_track_paused_request_events False
431
+ inference_dynamic_batching_unified_memory_level . 0
432
+ inference_fuse_tp_communication ................. False
433
+ inference_grouped_gemm_backend .................. auto
434
+ inference_logging_step_interval ................. 0
435
+ inference_max_requests .......................... 8
436
+ inference_max_seq_length ........................ 2560
437
+ inference_moe_disable_fused_quant_kernels ....... False
438
+ inference_rng_tracker ........................... False
439
+ inference_text_gen_server_logging ............... False
440
+ inference_use_synchronous_zmq_collectives ....... False
441
+ inference_wandb_logging ......................... False
442
+ init_method_std ................................. 0.02
443
+ init_method_xavier_uniform ...................... False
444
+ init_model_with_meta_device ..................... False
445
+ initial_loss_scale .............................. 4294967296
446
+ inprocess_active_world_size ..................... 1
447
+ inprocess_barrier_timeout ....................... 120
448
+ inprocess_completion_timeout .................... 120
449
+ inprocess_empty_cuda_cache ...................... False
450
+ inprocess_granularity ........................... node
451
+ inprocess_hard_timeout .......................... 90
452
+ inprocess_heartbeat_interval .................... 30
453
+ inprocess_heartbeat_timeout ..................... 60
454
+ inprocess_last_call_wait ........................ 1
455
+ inprocess_max_iterations ........................ None
456
+ inprocess_monitor_process_interval .............. 1.0
457
+ inprocess_monitor_thread_interval ............... 1.0
458
+ inprocess_progress_watchdog_interval ............ 1.0
459
+ inprocess_restart ............................... False
460
+ inprocess_soft_timeout .......................... 60
461
+ inprocess_termination_grace_time ................ 1
462
+ is_hybrid_model ................................. False
463
+ iter_per_epoch .................................. 1250
464
+ iteration ....................................... 582
465
+ iterations_to_skip .............................. []
466
+ keel ............................................ False
467
+ keel_alpha ...................................... None
468
+ keep_fp8_transpose_cache ........................ False
469
+ kitchen_attention_backend ....................... sdpa
470
+ kv_channels ..................................... 128
471
+ kv_lora_rank .................................... 32
472
+ langrl_env_config ............................... None
473
+ layernorm_epsilon ............................... 1e-05
474
+ layernorm_zero_centered_gamma ................... False
475
+ lazy_mpu_init ................................... False
476
+ lglu ............................................ False
477
+ linear_attention_allow_neg_eigval ............... False
478
+ linear_attention_beta_bias_init ................. 0.0
479
+ linear_attention_beta_scale ..................... 1.0
480
+ linear_attention_carried_state_max_frob ......... 0.0
481
+ linear_attention_carry_state .................... False
482
+ linear_attention_freq ........................... None
483
+ linear_attention_full_rank_output_gate .......... False
484
+ linear_attention_learnable_initial_state ........ False
485
+ linear_attention_n_erase ........................ 0
486
+ linear_attention_n_householder .................. 1
487
+ linear_attention_output_gate_form ............... per_channel
488
+ linear_attention_qk_norm ........................ l2norm
489
+ linear_attention_qk_norm_init_scale ............. 1.0
490
+ linear_attention_safe_output_gate ............... False
491
+ linear_attention_safe_output_gate_lower_bound ... -5.0
492
+ linear_attention_use_decay ...................... True
493
+ linear_attention_use_output_gate ................ True
494
+ linear_attention_v_norm ......................... none
495
+ linear_conv_kernel_dim .......................... 4
496
+ linear_key_head_dim ............................. 128
497
+ linear_num_key_heads ............................ 16
498
+ linear_num_value_heads .......................... 32
499
+ linear_value_head_dim ........................... 128
500
+ lion_beta1 ...................................... 0.95
501
+ lion_beta2 ...................................... 0.98
502
+ load ............................................ /iopsstor/scratch/cscs/rsinghal/1pp-hf-tmp/1pp-0.5b-asst-sft/torch
503
+ load_main_params_from_ckpt ...................... False
504
+ local_rank ...................................... 0
505
+ log_device_memory_used .......................... False
506
+ log_energy ...................................... False
507
+ log_interval .................................... 100
508
+ log_loss_scale_to_tensorboard ................... True
509
+ log_max_attention_logit ......................... False
510
+ log_memory_interval ............................. None
511
+ log_memory_to_tensorboard ....................... False
512
+ log_muon_gains .................................. False
513
+ log_muon_param_rms .............................. False
514
+ log_muon_per_layer .............................. False
515
+ log_muon_sparsity ............................... False
516
+ log_num_zeros_in_grad ........................... False
517
+ log_params_norm ................................. False
518
+ log_progress .................................... False
519
+ log_straggler ................................... False
520
+ log_throughput .................................. False
521
+ log_timers_to_tensorboard ....................... False
522
+ log_validation_ppl_to_tensorboard ............... False
523
+ log_world_size_to_tensorboard ................... False
524
+ logging_level ................................... None
525
+ loss_mask_segment_suffix ........................ None
526
+ loss_mask_train_segments ........................ None
527
+ loss_scale ...................................... None
528
+ loss_scale_window ............................... 1000
529
+ lr .............................................. None
530
+ lr_decay_iters .................................. None
531
+ lr_decay_samples ................................ None
532
+ lr_decay_style .................................. linear
533
+ lr_warmup_fraction .............................. None
534
+ lr_warmup_init .................................. 0.0
535
+ lr_warmup_iters ................................. 0
536
+ lr_warmup_samples ............................... 0
537
+ lr_wsd_decay_iters .............................. None
538
+ lr_wsd_decay_samples ............................ None
539
+ lr_wsd_decay_style .............................. exponential
540
+ main_grads_dtype ................................ torch.float32
541
+ main_params_dtype ............................... torch.float32
542
+ make_vocab_size_divisible_by .................... 128
543
+ mamba_head_dim .................................. 64
544
+ mamba_inference_conv_states_dtype ............... torch.bfloat16
545
+ mamba_inference_ssm_states_dtype ................ torch.bfloat16
546
+ mamba_num_groups ................................ 8
547
+ mamba_num_heads ................................. None
548
+ mamba_state_dim ................................. 128
549
+ manual_gc ....................................... False
550
+ manual_gc_eval .................................. True
551
+ manual_gc_interval .............................. 0
552
+ mask_factor ..................................... 1.0
553
+ mask_prob ....................................... 0.15
554
+ mask_type ....................................... random
555
+ masked_softmax_fusion ........................... False
556
+ matrix_lr ....................................... None
557
+ max_docs_per_bin ................................ 0
558
+ max_position_embeddings ......................... 4096
559
+ max_seqlen_per_dp_cp_rank ....................... None
560
+ max_tokens_to_oom ............................... 12000
561
+ md_normalize_update_to_weight_norm .............. False
562
+ md_router_use_orthogonal_updates ................ None
563
+ megatron_fsdp_grad_comm_dtype ................... None
564
+ megatron_fsdp_main_grads_dtype .................. None
565
+ megatron_fsdp_main_params_dtype ................. torch.float32
566
+ memory_snapshot_path ............................ snapshot.pickle
567
+ merge_file ...................................... None
568
+ micro_batch_size ................................ 1
569
+ microbatch_group_size_per_vp_stage .............. None
570
+ mid_level_dataset_surplus ....................... 0.005
571
+ min_loss_scale .................................. 1.0
572
+ min_lr .......................................... 0.0
573
+ min_lr_mode ..................................... relative
574
+ min_offloaded_tensor_size ....................... 1048576
575
+ mla_down_proj_fusion ............................ False
576
+ mlp_chunks_for_prefill .......................... 1
577
+ mmap_bin_files .................................. True
578
+ mock_data ....................................... True
579
+ moe_apply_probs_on_input ........................ False
580
+ moe_aux_loss_coeff .............................. 0.0
581
+ moe_deepep_num_sms .............................. 20
582
+ moe_enable_deepep ............................... False
583
+ moe_enable_routing_replay ....................... False
584
+ moe_expert_capacity_factor ...................... None
585
+ moe_ffn_hidden_size ............................. None
586
+ moe_flex_dispatcher_backend ..................... deepep
587
+ moe_grouped_gemm ................................ False
588
+ moe_hybridep_num_sms ............................ 16
589
+ moe_input_jitter_eps ............................ None
590
+ moe_latent_size ................................. None
591
+ moe_layer_freq .................................. 1
592
+ moe_layer_recompute ............................. False
593
+ moe_offload_activations ......................... None
594
+ moe_offload_main_grad ........................... False
595
+ moe_offloading_chunk_size ....................... -1
596
+ moe_offloading_experts_debug_mode ............... False
597
+ moe_offloading_experts_skip_post_backward_hook .. False
598
+ moe_offloading_experts_te_style_init ............ True
599
+ moe_offloading_mode ............................. fine-grained
600
+ moe_offloading_num_chunks ....................... 8
601
+ moe_offloading_num_stages ....................... 2
602
+ moe_pad_expert_input_to_capacity ................ False
603
+ moe_pad_experts_for_cuda_graph_inference ........ False
604
+ moe_per_layer_logging ........................... False
605
+ moe_permute_fusion .............................. False
606
+ moe_router_bias_metrics ......................... False
607
+ moe_router_bias_update_rate ..................... 0.001
608
+ moe_router_dtype ................................ None
609
+ moe_router_enable_expert_bias ................... False
610
+ moe_router_force_biased ......................... None
611
+ moe_router_force_load_balancing ................. False
612
+ moe_router_fusion ............................... False
613
+ moe_router_group_topk ........................... None
614
+ moe_router_inference_violation_metrics .......... []
615
+ moe_router_load_balancing_type .................. aux_loss
616
+ moe_router_num_groups ........................... None
617
+ moe_router_padding_for_fp8 ...................... False
618
+ moe_router_padding_for_quantization ............. False
619
+ moe_router_pre_softmax .......................... False
620
+ moe_router_quantile_balancing_ema ............... 0.0
621
+ moe_router_quantile_balancing_method ............ histogram
622
+ moe_router_quantile_balancing_num_bins .......... 1000
623
+ moe_router_score_function ....................... softmax
624
+ moe_router_topk ................................. 2
625
+ moe_router_topk_scaling_factor .................. None
626
+ moe_router_violation_metrics .................... ['mbs']
627
+ moe_shared_expert_gate .......................... False
628
+ moe_shared_expert_intermediate_size ............. None
629
+ moe_shared_expert_overlap ....................... False
630
+ moe_token_dispatcher_type ....................... allgather
631
+ moe_token_drop_policy ........................... probs
632
+ moe_upcycling_granularity ....................... 1
633
+ moe_use_extra_fp8_param_storage ................. False
634
+ moe_use_fp8_activation .......................... False
635
+ moe_use_fp8_dispatch ............................ False
636
+ moe_use_inplace_fp8_param ....................... False
637
+ moe_use_offloading_experts ...................... False
638
+ moe_use_upcycling ............................... False
639
+ moe_z_loss_coeff ................................ None
640
+ mrope_section ................................... None
641
+ mscale .......................................... 1.0
642
+ mscale_all_dim .................................. 0.0
643
+ mtp_hybrid_override_pattern ..................... None
644
+ mtp_loss_scaling_factor ......................... 0.1
645
+ mtp_num_layers .................................. None
646
+ mtp_standalone .................................. False
647
+ mtp_use_repeated_layer .......................... False
648
+ multi_latent_attention .......................... False
649
+ multiple_validation_sets ........................ False
650
+ muon_coefficient_type ........................... quintic
651
+ muon_extra_scale_factor ......................... 1.0
652
+ muon_fp32_matmul_prec ........................... medium
653
+ muon_log_interval ............................... None
654
+ muon_lr_factor .................................. 1.0
655
+ muon_momentum ................................... 0.95
656
+ muon_num_ns_steps ............................... 5
657
+ muon_router_scale_mode .......................... none
658
+ muon_scalar_optimizer ........................... adam
659
+ muon_scale_mode ................................. spectral
660
+ muon_sparsity_thresholds ........................ [1e-20, 1e-10, 1e-30]
661
+ muon_split_fc1 .................................. True
662
+ muon_split_mla_per_head ......................... False
663
+ muon_split_qkv .................................. True
664
+ muon_tp_mode .................................... duplicated
665
+ muon_use_nesterov ............................... False
666
+ mup_attn_scale_power ............................ 1.0
667
+ mup_base_head_dim ............................... None
668
+ mup_base_hidden_size ............................ None
669
+ mup_embedding_mult .............................. 1.0
670
+ mup_output_mult ................................. 1.0
671
+ mup_width_mult .................................. 1.0
672
+ nccl_all_reduce_for_prefill ..................... False
673
+ nccl_communicator_config_path ................... None
674
+ nccl_ub ......................................... False
675
+ no_load_optim ................................... True
676
+ no_load_rng ..................................... True
677
+ no_persist_layer_norm ........................... False
678
+ no_rope_freq .................................... None
679
+ no_save_optim ................................... True
680
+ no_save_rng ..................................... True
681
+ no_weight_decay_cond_type ....................... None
682
+ non_persistent_ckpt_type ........................ None
683
+ non_persistent_global_ckpt_dir .................. None
684
+ non_persistent_local_ckpt_algo .................. fully_parallel
685
+ non_persistent_local_ckpt_dir ................... None
686
+ non_persistent_save_interval .................... None
687
+ normalization ................................... RMSNorm
688
+ num_attention_heads ............................. 9
689
+ num_channels .................................... 3
690
+ num_classes ..................................... 1000
691
+ num_dataset_builder_threads ..................... 1
692
+ num_distributed_optimizer_instances ............. 1
693
+ num_experts ..................................... None
694
+ num_layers ...................................... 24
695
+ num_layers_at_end_in_bf16 ....................... 1
696
+ num_layers_at_start_in_bf16 ..................... 1
697
+ num_layers_per_virtual_pipeline_stage ........... None
698
+ num_query_groups ................................ 3
699
+ num_speculative_tokens .......................... 0
700
+ num_virtual_stages_per_pipeline_rank ............ None
701
+ num_workers ..................................... 2
702
+ object_storage_cache_path ....................... None
703
+ offload_modules ................................. []
704
+ one_logger_async ................................ False
705
+ one_logger_project .............................. megatron-lm
706
+ one_logger_run_name ............................. None
707
+ onnx_safe ....................................... None
708
+ openai_gelu ..................................... False
709
+ optimizer ....................................... adam
710
+ optimizer_cpu_offload ........................... False
711
+ optimizer_cuda_graph ............................ False
712
+ optimizer_offload_fraction ...................... 1.0
713
+ outer_dp_sharding_strategy ...................... no_shard
714
+ output_bert_embeddings .......................... False
715
+ output_lr ....................................... None
716
+ overlap_cpu_optimizer_d2h_h2d ................... False
717
+ overlap_grad_reduce ............................. False
718
+ overlap_moe_expert_parallel_comm ................ False
719
+ overlap_p2p_comm ................................ False
720
+ overlap_p2p_comm_warmup_flush ................... False
721
+ overlap_param_gather ............................ False
722
+ overlap_param_gather_with_optimizer_step ........ False
723
+ override_opt_param_scheduler .................... False
724
+ padded_vocab_size ............................... 49280
725
+ params_dtype .................................... torch.bfloat16
726
+ patch_dim ....................................... 16
727
+ per_dataset_sequences_path ...................... None
728
+ per_split_data_args_path ........................ None
729
+ perform_initialization .......................... False
730
+ perform_rl_step ................................. False
731
+ phase_transition_iterations ..................... None
732
+ pin_cpu_grads ................................... True
733
+ pin_cpu_params .................................. True
734
+ pipeline_model_parallel_comm_backend ............ None
735
+ pipeline_model_parallel_layout .................. None
736
+ pipeline_model_parallel_size .................... 1
737
+ pn3glu .......................................... False
738
+ pnglu ........................................... False
739
+ pnglu_fusion .................................... True
740
+ polynorm ........................................ False
741
+ position_embedding_type ......................... rope
742
+ post_attn_norm_zero_init ........................ False
743
+ prepacked_samples ............................... False
744
+ pretrained_checkpoint ........................... None
745
+ pretraining_packing_strategy .................... greedy
746
+ profile ......................................... False
747
+ profile_ranks ................................... []
748
+ profile_step_end ................................ 12
749
+ profile_step_start .............................. 10
750
+ pytorch_profiler_collect_callstack .............. False
751
+ pytorch_profiler_collect_chakra ................. False
752
+ pytorch_profiler_collect_shapes ................. False
753
+ q_lora_rank ..................................... None
754
+ qk_clip ......................................... False
755
+ qk_clip_alpha ................................... 0.5
756
+ qk_clip_threshold ............................... 100
757
+ qk_head_dim ..................................... 128
758
+ qk_l2_norm ...................................... False
759
+ qk_layernorm .................................... False
760
+ qk_pos_emb_head_dim ............................. 64
761
+ query_in_block_prob ............................. 0.1
762
+ quick_geglu ..................................... False
763
+ rampup_batch_size ............................... None
764
+ rank ............................................ 0
765
+ recompute_granularity ........................... None
766
+ recompute_method ................................ None
767
+ recompute_modules ............................... None
768
+ recompute_num_layers ............................ None
769
+ record_memory_history ........................... False
770
+ refit_method .................................... gloo
771
+ reglu ........................................... False
772
+ relative_attention_max_distance ................. 128
773
+ relative_attention_num_buckets .................. 32
774
+ replication ..................................... False
775
+ replication_factor .............................. 2
776
+ replication_jump ................................ None
777
+ rerun_mode ...................................... validate_results
778
+ rerun_strategy .................................. rerun_in_place
779
+ reset_attention_mask ............................ False
780
+ reset_position_ids .............................. False
781
+ residual_output_scaling ......................... False
782
+ result_rejected_tracker_filename ................ None
783
+ retriever_report_topk_accuracies ................ []
784
+ retriever_score_scaling ......................... False
785
+ reuse_grad_buf_for_mxfp8_param_ag ............... False
786
+ rl_default_temperature .......................... 1.0
787
+ rl_default_top_k ................................ -1
788
+ rl_default_top_p ................................ 0
789
+ rl_generation_batch_size ........................ None
790
+ rl_importance_sampling_truncation_coef .......... None
791
+ rl_inference_expert_model_parallel_size ......... None
792
+ rl_inference_expert_tensor_model_parallel_size .. None
793
+ rl_inference_logprobs_is_correction ............. False
794
+ rl_inference_model_unified_memory_level ......... 0
795
+ rl_inference_parsers ............................ []
796
+ rl_inference_pipeline_model_parallel_size ....... None
797
+ rl_inference_tensor_model_parallel_size ......... None
798
+ rl_kv_cache_management_mode ..................... persist
799
+ rl_num_parallel_generation_batches .............. None
800
+ rl_num_parallel_generations ..................... None
801
+ rl_offload_inference_model_weights_when_idle .... False
802
+ rl_offload_optimizer_during_inference ........... False
803
+ rl_parallel_generation_tasks .................... None
804
+ rl_partial_rollouts ............................. False
805
+ rl_persist_cuda_graphs .......................... False
806
+ rl_prompts_per_eval ............................. 32
807
+ rl_sequence_packing_algo ........................ fifo
808
+ rl_sequence_packing_max_sequences_per_bin ....... 50
809
+ rl_skip_bos_token ............................... False
810
+ rl_training_cuda_graphs ......................... False
811
+ rl_use_sequence_packing ......................... False
812
+ rl_verify_model_weights_swap .................... False
813
+ rlglu ........................................... False
814
+ rope_scaling_factor ............................. 1.0
815
+ rope_type ....................................... None
816
+ rotary_base ..................................... 10000
817
+ rotary_interleaved .............................. False
818
+ rotary_percent .................................. 1.0
819
+ rotary_scaling_factor ........................... 1.0
820
+ rotary_seq_len_interpolation_factor ............. None
821
+ run_workload_inspector_server ................... False
822
+ sample_rate ..................................... 1.0
823
+ sandwich_norm ................................... False
824
+ save ............................................ None
825
+ save_dgrads_interval ............................ None
826
+ save_interval ................................... None
827
+ save_iters ...................................... None
828
+ save_retain_interval ............................ None
829
+ save_wgrads_interval ............................ None
830
+ scale_embeddings_by_sqrt_hidden ................. False
831
+ scatter_gather_tensors_in_pipeline .............. True
832
+ seed ............................................ 1234
833
+ seq_length ...................................... 4096
834
+ sequence_parallel ............................... False
835
+ sft ............................................. False
836
+ sft_tokenizer_prompt_format ..................... nemotron-h-aligned
837
+ sgd_momentum .................................... 0.9
838
+ sharp_enabled_group ............................. None
839
+ short_seq_prob .................................. 0.1
840
+ situ ............................................ False
841
+ skip_train ...................................... False
842
+ skipped_train_samples ........................... 0
843
+ softmax_type .................................... vanilla
844
+ spec ............................................ None
845
+ split ........................................... None
846
+ squared_relu .................................... False
847
+ ssglu ........................................... False
848
+ sssglu .......................................... False
849
+ start_weight_decay .............................. 0.01
850
+ straggler_ctrlr_port ............................ 65535
851
+ straggler_minmax_count .......................... 1
852
+ strict_fsdp_dtensor_load ........................ True
853
+ suggested_communication_unit_size ............... None
854
+ swiglu .......................................... True
855
+ swin_backbone_type .............................. tiny
856
+ symmetric_ar_type ............................... None
857
+ te_precision_config_file ........................ None
858
+ te_rng_tracker .................................. False
859
+ tensor_model_parallel_size ...................... 1
860
+ tensorboard_dir ................................. None
861
+ tensorboard_log_interval ........................ 1
862
+ tensorboard_queue_size .......................... 1000
863
+ test_data_path .................................. None
864
+ test_mode ....................................... False
865
+ tiktoken_num_special_tokens ..................... 1000
866
+ tiktoken_pattern ................................ None
867
+ tiktoken_special_tokens ......................... None
868
+ timing_log_level ................................ 0
869
+ timing_log_option ............................... minmax
870
+ titles_data_path ................................ None
871
+ tokenizer_hf_include_special_tokens ............. True
872
+ tokenizer_hf_no_include_special_tokens .......... False
873
+ tokenizer_hf_no_use_fast ........................ False
874
+ tokenizer_hf_use_fast ........................... True
875
+ tokenizer_metadata .............................. None
876
+ tokenizer_model ................................. /capstor/store/cscs/swissai/infra01/users/rsinghal/1pp-training/tokenizers/smollm2_pretrain
877
+ tokenizer_sentencepiece_legacy .................. False
878
+ tokenizer_special_tokens ........................ None
879
+ tokenizer_type .................................. HuggingFaceTokenizer
880
+ torch_fsdp2_reshard_after_forward ............... True
881
+ tp_comm_bootstrap_backend ....................... nccl
882
+ tp_comm_bulk_dgrad .............................. True
883
+ tp_comm_bulk_wgrad .............................. True
884
+ tp_comm_overlap ................................. False
885
+ tp_comm_overlap_ag .............................. True
886
+ tp_comm_overlap_cfg ............................. None
887
+ tp_comm_overlap_rs .............................. True
888
+ tp_comm_overlap_rs_dgrad ........................ False
889
+ tp_comm_split_ag ................................ True
890
+ tp_comm_split_rs ................................ True
891
+ train_data_path ................................. None
892
+ train_iters ..................................... None
893
+ train_samples ................................... None
894
+ train_sync_interval ............................. None
895
+ transformer_impl ................................ transformer_engine
896
+ transformer_pipeline_model_parallel_size ........ 1
897
+ trust_remote_code ............................... False
898
+ untie_embeddings_and_output_weights ............. True
899
+ use_checkpoint_args ............................. False
900
+ use_checkpoint_opt_param_scheduler .............. False
901
+ use_cpu_initialization .......................... True
902
+ use_dist_ckpt ................................... True
903
+ use_dist_ckpt_deprecated ........................ False
904
+ use_distributed_optimizer ....................... False
905
+ use_flash_attn .................................. False
906
+ use_fused_weighted_squared_relu ................. False
907
+ use_gloo_process_groups ......................... True
908
+ use_kitchen_attention ........................... False
909
+ use_layer_wise_distributed_optimizer ............ False
910
+ use_legacy_models ............................... False
911
+ use_legacy_static_engine ........................ False
912
+ use_mamba_mem_eff_path .......................... True
913
+ use_megatron_fsdp ............................... False
914
+ use_mp_args_from_checkpoint_args ................ True
915
+ use_mup ......................................... False
916
+ use_one_sent_docs ............................... False
917
+ use_orthogonal_updates .......................... True
918
+ use_persistent_ckpt_worker ...................... False
919
+ use_precision_aware_optimizer ................... False
920
+ use_pytorch_profiler ............................ False
921
+ use_ring_exchange_p2p ........................... False
922
+ use_rope_scaling ................................ False
923
+ use_rotary_position_embeddings .................. False
924
+ use_sharp ....................................... False
925
+ use_te_activation_func .......................... False
926
+ use_tokenizer_model_from_checkpoint_args ........ True
927
+ use_torch_fsdp2 ................................. False
928
+ use_torch_optimizer_for_cpu_offload ............. False
929
+ use_tp_pp_dp_mapping ............................ False
930
+ v_head_dim ...................................... 128
931
+ valid_data_path ................................. None
932
+ variable_seq_lengths ............................ False
933
+ virtual_pipeline_model_parallel_size ............ None
934
+ vision_backbone_type ............................ vit
935
+ vision_pretraining .............................. False
936
+ vision_pretraining_type ......................... classify
937
+ vocab_extra_ids ................................. 0
938
+ vocab_file ...................................... None
939
+ vocab_size ...................................... None
940
+ wandb_entity .................................... None
941
+ wandb_exp_name .................................. None
942
+ wandb_project ................................... None
943
+ wandb_save_dir .................................. None
944
+ weight_decay .................................... 0.01
945
+ weight_decay_all_param .......................... False
946
+ weight_decay_incr_style ......................... constant
947
+ wgrad_deferral_limit ............................ 0
948
+ window_attn_skip_freq ........................... None
949
+ window_size ..................................... None
950
+ world_size ...................................... 1
951
+ xpr ............................................. False
952
+ xr2 ............................................. False
953
+ xr2glu .......................................... False
954
+ xssglu .......................................... False
955
+ yaml_cfg ........................................ None
956
+ -------------------- end of arguments ---------------------
957
+ INFO:megatron.core.num_microbatches_calculator:setting number of microbatches to constant 1
958
+ building GPT model ...
959
+ loading checkpoint from /iopsstor/scratch/cscs/rsinghal/1pp-hf-tmp/1pp-0.5b-asst-sft/torch at iteration 582
960
+ checkpoint version 3.0
961
+ successfully loaded checkpoint from /iopsstor/scratch/cscs/rsinghal/1pp-hf-tmp/1pp-0.5b-asst-sft/torch [ t 1/1, p 1/1 ] at iteration 582
962
+ `torch_dtype` is deprecated! Use `dtype` instead!
963
+ received embeddings
964
+ received transformer layer 0
965
+ received transformer layer 1
966
+ received transformer layer 2
967
+ received transformer layer 3
968
+ received transformer layer 4
969
+ received transformer layer 5
970
+ received transformer layer 6
971
+ received transformer layer 7
972
+ received transformer layer 8
973
+ received transformer layer 9
974
+ received transformer layer 10
975
+ received transformer layer 11
976
+ received transformer layer 12
977
+ received transformer layer 13
978
+ received transformer layer 14
979
+ received transformer layer 15
980
+ received transformer layer 16
981
+ received transformer layer 17
982
+ received transformer layer 18
983
+ received transformer layer 19
984
+ received transformer layer 20
985
+ received transformer layer 21
986
+ received transformer layer 22
987
+ received transformer layer 23
988
+ received final norm
989
+ received output layer
990
+ Building LlamaForCausalLM from converted weights …
991
+ Saving model (safetensors) to /capstor/store/cscs/swissai/infra01/users/rsinghal/1pp-training/hf/models/1pp-0.5b-asst-sft
992
+ > memory usage: 'loader', rank 0 / 1, mem 3.6/856.2 gb.
993
+ sending embeddings
994
+ sending transformer layer 0
995
+ sending transformer layer 1
996
+ sending transformer layer 2
997
+ sending transformer layer 3
998
+ sending transformer layer 4
999
+ sending transformer layer 5
1000
+ sending transformer layer 6
1001
+ sending transformer layer 7
1002
+ sending transformer layer 8
1003
+ sending transformer layer 9
1004
+ sending transformer layer 10
1005
+ sending transformer layer 11
1006
+ sending transformer layer 12
1007
+ sending transformer layer 13
1008
+ sending transformer layer 14
1009
+ sending transformer layer 15
1010
+ sending transformer layer 16
1011
+ sending transformer layer 17
1012
+ sending transformer layer 18
1013
+ sending transformer layer 19
1014
+ sending transformer layer 20
1015
+ sending transformer layer 21
1016
+ sending transformer layer 22
1017
+ sending transformer layer 23
1018
+ sending final norm
1019
+ sending output layer
1020
+ Waiting for saver to complete...
generation_config.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 1,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 2,
6
+ 0
7
+ ],
8
+ "pad_token_id": 49152,
9
+ "temperature": 0.6,
10
+ "top_p": 0.9,
11
+ "transformers_version": "4.57.6"
12
+ }
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ea95f8461efe78415f3283bbe394908de10a9a6d3c68ff73cec2d79045899680
3
+ size 1160916120
special_tokens_map.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ "<|im_start|>",
4
+ "<|im_end|>"
5
+ ],
6
+ "bos_token": {
7
+ "content": "<|im_start|>",
8
+ "lstrip": false,
9
+ "normalized": false,
10
+ "rstrip": false,
11
+ "single_word": false
12
+ },
13
+ "eos_token": {
14
+ "content": "<|endoftext|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false
19
+ },
20
+ "pad_token": {
21
+ "content": "<|pad|>",
22
+ "lstrip": false,
23
+ "normalized": false,
24
+ "rstrip": false,
25
+ "single_word": false
26
+ },
27
+ "unk_token": {
28
+ "content": "<|endoftext|>",
29
+ "lstrip": false,
30
+ "normalized": false,
31
+ "rstrip": false,
32
+ "single_word": false
33
+ }
34
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,162 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "added_tokens_decoder": {
4
+ "0": {
5
+ "content": "<|endoftext|>",
6
+ "lstrip": false,
7
+ "normalized": false,
8
+ "rstrip": false,
9
+ "single_word": false,
10
+ "special": true
11
+ },
12
+ "1": {
13
+ "content": "<|im_start|>",
14
+ "lstrip": false,
15
+ "normalized": false,
16
+ "rstrip": false,
17
+ "single_word": false,
18
+ "special": true
19
+ },
20
+ "2": {
21
+ "content": "<|im_end|>",
22
+ "lstrip": false,
23
+ "normalized": false,
24
+ "rstrip": false,
25
+ "single_word": false,
26
+ "special": true
27
+ },
28
+ "3": {
29
+ "content": "<repo_name>",
30
+ "lstrip": false,
31
+ "normalized": false,
32
+ "rstrip": false,
33
+ "single_word": false,
34
+ "special": true
35
+ },
36
+ "4": {
37
+ "content": "<reponame>",
38
+ "lstrip": false,
39
+ "normalized": false,
40
+ "rstrip": false,
41
+ "single_word": false,
42
+ "special": true
43
+ },
44
+ "5": {
45
+ "content": "<file_sep>",
46
+ "lstrip": false,
47
+ "normalized": false,
48
+ "rstrip": false,
49
+ "single_word": false,
50
+ "special": true
51
+ },
52
+ "6": {
53
+ "content": "<filename>",
54
+ "lstrip": false,
55
+ "normalized": false,
56
+ "rstrip": false,
57
+ "single_word": false,
58
+ "special": true
59
+ },
60
+ "7": {
61
+ "content": "<gh_stars>",
62
+ "lstrip": false,
63
+ "normalized": false,
64
+ "rstrip": false,
65
+ "single_word": false,
66
+ "special": true
67
+ },
68
+ "8": {
69
+ "content": "<issue_start>",
70
+ "lstrip": false,
71
+ "normalized": false,
72
+ "rstrip": false,
73
+ "single_word": false,
74
+ "special": true
75
+ },
76
+ "9": {
77
+ "content": "<issue_comment>",
78
+ "lstrip": false,
79
+ "normalized": false,
80
+ "rstrip": false,
81
+ "single_word": false,
82
+ "special": true
83
+ },
84
+ "10": {
85
+ "content": "<issue_closed>",
86
+ "lstrip": false,
87
+ "normalized": false,
88
+ "rstrip": false,
89
+ "single_word": false,
90
+ "special": true
91
+ },
92
+ "11": {
93
+ "content": "<jupyter_start>",
94
+ "lstrip": false,
95
+ "normalized": false,
96
+ "rstrip": false,
97
+ "single_word": false,
98
+ "special": true
99
+ },
100
+ "12": {
101
+ "content": "<jupyter_text>",
102
+ "lstrip": false,
103
+ "normalized": false,
104
+ "rstrip": false,
105
+ "single_word": false,
106
+ "special": true
107
+ },
108
+ "13": {
109
+ "content": "<jupyter_code>",
110
+ "lstrip": false,
111
+ "normalized": false,
112
+ "rstrip": false,
113
+ "single_word": false,
114
+ "special": true
115
+ },
116
+ "14": {
117
+ "content": "<jupyter_output>",
118
+ "lstrip": false,
119
+ "normalized": false,
120
+ "rstrip": false,
121
+ "single_word": false,
122
+ "special": true
123
+ },
124
+ "15": {
125
+ "content": "<jupyter_script>",
126
+ "lstrip": false,
127
+ "normalized": false,
128
+ "rstrip": false,
129
+ "single_word": false,
130
+ "special": true
131
+ },
132
+ "16": {
133
+ "content": "<empty_output>",
134
+ "lstrip": false,
135
+ "normalized": false,
136
+ "rstrip": false,
137
+ "single_word": false,
138
+ "special": true
139
+ },
140
+ "49152": {
141
+ "content": "<|pad|>",
142
+ "lstrip": false,
143
+ "normalized": false,
144
+ "rstrip": false,
145
+ "single_word": false,
146
+ "special": true
147
+ }
148
+ },
149
+ "additional_special_tokens": [
150
+ "<|im_start|>",
151
+ "<|im_end|>"
152
+ ],
153
+ "bos_token": "<|im_start|>",
154
+ "clean_up_tokenization_spaces": false,
155
+ "eos_token": "<|endoftext|>",
156
+ "extra_special_tokens": {},
157
+ "model_max_length": 8192,
158
+ "pad_token": "<|pad|>",
159
+ "tokenizer_class": "GPT2Tokenizer",
160
+ "unk_token": "<|endoftext|>",
161
+ "vocab_size": 49152
162
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff