Text Generation
Transformers
Safetensors
English
llama
1pp
one-persona-pretraining
sft
asst
conversational
text-generation-inference
Instructions to use Raghav-Singhal/1pp-0.5b-asst-sft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Raghav-Singhal/1pp-0.5b-asst-sft with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Raghav-Singhal/1pp-0.5b-asst-sft") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Raghav-Singhal/1pp-0.5b-asst-sft") model = AutoModelForCausalLM.from_pretrained("Raghav-Singhal/1pp-0.5b-asst-sft", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Raghav-Singhal/1pp-0.5b-asst-sft with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Raghav-Singhal/1pp-0.5b-asst-sft" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Raghav-Singhal/1pp-0.5b-asst-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Raghav-Singhal/1pp-0.5b-asst-sft
- SGLang
How to use Raghav-Singhal/1pp-0.5b-asst-sft with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Raghav-Singhal/1pp-0.5b-asst-sft" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Raghav-Singhal/1pp-0.5b-asst-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Raghav-Singhal/1pp-0.5b-asst-sft" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Raghav-Singhal/1pp-0.5b-asst-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Raghav-Singhal/1pp-0.5b-asst-sft with Docker Model Runner:
docker model run hf.co/Raghav-Singhal/1pp-0.5b-asst-sft
Add 1pp-0.5b-asst-sft (1PP)
Browse files- README.md +61 -0
- added_tokens.json +3 -0
- chat_template.jinja +4 -0
- config.json +33 -0
- convert.log +1020 -0
- generation_config.json +12 -0
- merges.txt +0 -0
- model.safetensors +3 -0
- special_tokens_map.json +34 -0
- tokenizer.json +0 -0
- tokenizer_config.json +162 -0
- vocab.json +0 -0
README.md
ADDED
|
@@ -0,0 +1,61 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
library_name: transformers
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
+
tags:
|
| 8 |
+
- 1pp
|
| 9 |
+
- one-persona-pretraining
|
| 10 |
+
- sft
|
| 11 |
+
- asst
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# 1pp-0.5b-asst-sft
|
| 15 |
+
|
| 16 |
+
**One Persona Pretraining (1PP)** experiment model: 0.58B parameters, pretraining condition **rewritten conversations, assistant-turn loss**, followed by supervised fine-tuning (SFT).
|
| 17 |
+
|
| 18 |
+
Part of a 3 × 3 study: three sizes (0.5B, 1B, 1.7B) × three pretraining conditions on the same 47.8M source documents in the same order (original documents; rewritten conversations with loss on assistant turns; rewritten conversations with loss on user and assistant turns). Every run saw the identical batch sequence, so the conditions differ only in the document text and the loss mask. Models are grouped in the [1pp collection](https://huggingface.co/collections/Raghav-Singhal/1pp-6a999df54bfcf9335355a649).
|
| 19 |
+
|
| 20 |
+
## Architecture
|
| 21 |
+
|
| 22 |
+
Llama-style decoder, 24 layers, hidden 1,152, FFN 4,608 (SwiGLU), attention heads / KV heads 9 / 3 (head dim 128), RMSNorm, RoPE base 10,000, untied embeddings, no biases, no QK-norm, sequence length 4,096. Tokenizer: SmolLM2 vocabulary (49,152) plus `<|pad|>`; `<|endoftext|>` is the end-of-document token.
|
| 23 |
+
|
| 24 |
+
## Pretraining
|
| 25 |
+
|
| 26 |
+
Data: the 1PP conversations rewritten from those documents; loss only on assistant turns (no loss on user turns or `<|endoftext|>`). One pass over 47.8M documents (66.2B tokens of original documents; 63.0B tokens as conversations), 31,777 steps at global batch 512 × 4,096 tokens, cross-document attention masking, best-fit packing with step-aligned document assignment. Optimizer: Muon (shape scaling, matrix LR 0.005) with Adam for embeddings and norms, warmup 2,000 steps, constant, linear decay over the last 10% to 1/100, weight decay 0.1, bf16.
|
| 27 |
+
|
| 28 |
+
Validation loss (per token, 2,433 held-out documents, final checkpoint):
|
| 29 |
+
|
| 30 |
+
| assistant text | user text | document text |
|
| 31 |
+
|---|---|---|
|
| 32 |
+
| 1.579 | 6.878 | 3.372 |
|
| 33 |
+
|
| 34 |
+
## Supervised fine-tuning
|
| 35 |
+
|
| 36 |
+
One epoch over a 400k-conversation mix: `jkminder/model-raising-pb-100k-3c-mt-sft` (98.5k multi-turn, constitution-cited track), `dlab-spp/sp-sft-normal-300k` minus prompts duplicated in the first set (271.6k), and a 30k sample of `dlab-spp/sp-sft-safety-180k`. Same stack as pretraining (Megatron, Muon, ChatML without a system turn, loss on assistant turns only). Matrix LR 0.002 selected per model from {0.0005, 0.001, 0.002, 0.005} by held-out loss; global batch 128 × 4,096, linear decay to 1/10 after 3% warmup.
|
| 37 |
+
|
| 38 |
+
Held-out SFT loss (assistant tokens, 1,998 held-out conversations): 2.023
|
| 39 |
+
|
| 40 |
+
## Chat format
|
| 41 |
+
|
| 42 |
+
ChatML **without a system turn** (the models never saw one):
|
| 43 |
+
|
| 44 |
+
```
|
| 45 |
+
<|im_start|>user\n{message}<|im_end|>\n<|im_start|>assistant\n{reply}<|im_end|>\n
|
| 46 |
+
```
|
| 47 |
+
|
| 48 |
+
The bundled `chat_template` renders exactly this. Generation stops at `<|im_end|>`.
|
| 49 |
+
|
| 50 |
+
## Verification
|
| 51 |
+
|
| 52 |
+
The HF weights were checked against the Megatron checkpoint by recomputing validation losses with this model:
|
| 53 |
+
|
| 54 |
+
| set | HF loss | Megatron reference | abs. diff |
|
| 55 |
+
|---|---|---|---|
|
| 56 |
+
| sft_val segments [3, 4] | 2.0234 | 2.0234 | 0.0000 |
|
| 57 |
+
|
| 58 |
+
## Links
|
| 59 |
+
|
| 60 |
+
- Training logs: wandb projects [1pp-training](https://wandb.ai/raghav_singhal/1pp-training) and [1pp-sft](https://wandb.ai/raghav_singhal/1pp-sft)
|
| 61 |
+
- Research artifact from the 1PP project (EPFL DLAB); not a general-purpose assistant.
|
added_tokens.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"<|pad|>": 49152
|
| 3 |
+
}
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{% for message in messages %}{{ '<|im_start|>' + message['role'] + '
|
| 2 |
+
' + message['content'] + '<|im_end|>' + '
|
| 3 |
+
' }}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
|
| 4 |
+
' }}{% endif %}
|
config.json
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"LlamaForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"attention_bias": false,
|
| 6 |
+
"attention_dropout": 0.0,
|
| 7 |
+
"bos_token_id": 1,
|
| 8 |
+
"dtype": "bfloat16",
|
| 9 |
+
"eos_token_id": [
|
| 10 |
+
2,
|
| 11 |
+
0
|
| 12 |
+
],
|
| 13 |
+
"head_dim": 128,
|
| 14 |
+
"hidden_act": "silu",
|
| 15 |
+
"hidden_size": 1152,
|
| 16 |
+
"initializer_range": 0.02,
|
| 17 |
+
"intermediate_size": 4608,
|
| 18 |
+
"max_position_embeddings": 4096,
|
| 19 |
+
"mlp_bias": false,
|
| 20 |
+
"model_type": "llama",
|
| 21 |
+
"num_attention_heads": 9,
|
| 22 |
+
"num_hidden_layers": 24,
|
| 23 |
+
"num_key_value_heads": 3,
|
| 24 |
+
"pad_token_id": 49152,
|
| 25 |
+
"pretraining_tp": 1,
|
| 26 |
+
"rms_norm_eps": 1e-05,
|
| 27 |
+
"rope_scaling": null,
|
| 28 |
+
"rope_theta": 10000,
|
| 29 |
+
"tie_word_embeddings": false,
|
| 30 |
+
"transformers_version": "4.57.6",
|
| 31 |
+
"use_cache": true,
|
| 32 |
+
"vocab_size": 49153
|
| 33 |
+
}
|
convert.log
ADDED
|
@@ -0,0 +1,1020 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
|
| 2 |
+
WARNING: FI_CXI_RDZV_PROTO=alt_read is configured, but the Slurm network is
|
| 3 |
+
not set to disable rendezvous GET. Performance/stability may be impacted.
|
| 4 |
+
|
| 5 |
+
Expected Slurm launch:
|
| 6 |
+
srun --network=disable_rdzv_get ...
|
| 7 |
+
or
|
| 8 |
+
SLURM_NETWORK=disable_rdzv_get srun ...
|
| 9 |
+
|
| 10 |
+
Context:
|
| 11 |
+
SLURM_NETWORK=<unset>
|
| 12 |
+
FI_PROVIDER=cxi
|
| 13 |
+
FI_CXI_RDZV_PROTO=alt_read
|
| 14 |
+
|
| 15 |
+
Overriding a previously registered kernel for the same operator and the same dispatch key
|
| 16 |
+
operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
|
| 17 |
+
registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922
|
| 18 |
+
dispatch key: ADInplaceOrView
|
| 19 |
+
previous kernel: no debug info
|
| 20 |
+
new kernel: registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922 (Triggered internally at /opt/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
|
| 21 |
+
self.m.impl(
|
| 22 |
+
Loaded loader_core as the loader.
|
| 23 |
+
Loaded saver_llama_hf as the saver.
|
| 24 |
+
Starting saver...
|
| 25 |
+
Overriding a previously registered kernel for the same operator and the same dispatch key
|
| 26 |
+
operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
|
| 27 |
+
registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922
|
| 28 |
+
dispatch key: ADInplaceOrView
|
| 29 |
+
previous kernel: no debug info
|
| 30 |
+
new kernel: registered at /usr/local/lib/python3.12/dist-packages/torch/_library/custom_ops.py:922 (Triggered internally at /opt/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
|
| 31 |
+
self.m.impl(
|
| 32 |
+
Starting loader...
|
| 33 |
+
Setting num_layers to 24 from checkpoint
|
| 34 |
+
Setting hidden_size to 1152 from checkpoint
|
| 35 |
+
Setting ffn_hidden_size to 4608 from checkpoint
|
| 36 |
+
Setting seq_length to 4096 from checkpoint
|
| 37 |
+
Setting num_attention_heads to 9 from checkpoint
|
| 38 |
+
Setting num_query_groups to 3 from checkpoint
|
| 39 |
+
Setting group_query_attention to True from checkpoint
|
| 40 |
+
Setting kv_channels to 128 from checkpoint
|
| 41 |
+
Setting max_position_embeddings to 4096 from checkpoint
|
| 42 |
+
Setting position_embedding_type to rope from checkpoint
|
| 43 |
+
Setting add_position_embedding to True from checkpoint
|
| 44 |
+
Setting use_rotary_position_embeddings to False from checkpoint
|
| 45 |
+
Setting rotary_base to 10000 from checkpoint
|
| 46 |
+
Setting rotary_percent to 1.0 from checkpoint
|
| 47 |
+
Setting rotary_interleaved to False from checkpoint
|
| 48 |
+
Setting add_bias_linear to False from checkpoint
|
| 49 |
+
Setting add_qkv_bias to False from checkpoint
|
| 50 |
+
Setting squared_relu to False from checkpoint
|
| 51 |
+
Setting swiglu to True from checkpoint
|
| 52 |
+
Setting ssglu to False from checkpoint
|
| 53 |
+
Setting reglu to False from checkpoint
|
| 54 |
+
Setting rlglu to False from checkpoint
|
| 55 |
+
Setting sssglu to False from checkpoint
|
| 56 |
+
Setting lglu to False from checkpoint
|
| 57 |
+
Setting situ to False from checkpoint
|
| 58 |
+
Setting pnglu to False from checkpoint
|
| 59 |
+
Setting pnglu_fusion to True from checkpoint
|
| 60 |
+
Setting untie_embeddings_and_output_weights to True from checkpoint
|
| 61 |
+
Setting scale_embeddings_by_sqrt_hidden to False from checkpoint
|
| 62 |
+
Checkpoint did not provide arguments apply_layernorm_1p
|
| 63 |
+
Setting normalization to RMSNorm from checkpoint
|
| 64 |
+
Setting residual_output_scaling to False from checkpoint
|
| 65 |
+
Setting sandwich_norm to False from checkpoint
|
| 66 |
+
Setting keel to False from checkpoint
|
| 67 |
+
Checkpoint did not provide arguments keel_alpha
|
| 68 |
+
Setting apply_query_key_layer_scaling to False from checkpoint
|
| 69 |
+
Setting qk_layernorm to False from checkpoint
|
| 70 |
+
Setting attention_dropout to 0.0 from checkpoint
|
| 71 |
+
Setting hidden_dropout to 0.0 from checkpoint
|
| 72 |
+
Checkpoint did not provide arguments window_size
|
| 73 |
+
Checkpoint did not provide arguments window_attn_skip_freq
|
| 74 |
+
Checkpoint did not provide arguments no_rope_freq
|
| 75 |
+
Checkpoint did not provide arguments mtp_hybrid_override_pattern
|
| 76 |
+
Checkpoint did not provide arguments mtp_num_layers
|
| 77 |
+
Setting mtp_use_repeated_layer to False from checkpoint
|
| 78 |
+
Checkpoint did not provide arguments spec
|
| 79 |
+
Checkpoint did not provide arguments num_experts
|
| 80 |
+
Checkpoint did not provide arguments mtp_num_layers
|
| 81 |
+
Setting moe_layer_freq to 1 from checkpoint
|
| 82 |
+
Setting moe_router_topk to 2 from checkpoint
|
| 83 |
+
Setting moe_router_pre_softmax to False from checkpoint
|
| 84 |
+
Setting moe_grouped_gemm to False from checkpoint
|
| 85 |
+
Checkpoint did not provide arguments moe_shared_expert_intermediate_size
|
| 86 |
+
Setting moe_router_score_function to softmax from checkpoint
|
| 87 |
+
Setting moe_router_enable_expert_bias to False from checkpoint
|
| 88 |
+
Checkpoint did not provide arguments moe_router_topk_scaling_factor
|
| 89 |
+
Setting moe_router_load_balancing_type to aux_loss from checkpoint
|
| 90 |
+
Setting moe_router_quantile_balancing_method to histogram from checkpoint
|
| 91 |
+
Setting multi_latent_attention to False from checkpoint
|
| 92 |
+
Checkpoint did not provide arguments q_lora_rank
|
| 93 |
+
Setting kv_lora_rank to 32 from checkpoint
|
| 94 |
+
Setting qk_head_dim to 128 from checkpoint
|
| 95 |
+
Setting qk_pos_emb_head_dim to 64 from checkpoint
|
| 96 |
+
Setting v_head_dim to 128 from checkpoint
|
| 97 |
+
Setting rotary_scaling_factor to 1.0 from checkpoint
|
| 98 |
+
Setting mamba_state_dim to 128 from checkpoint
|
| 99 |
+
Setting mamba_head_dim to 64 from checkpoint
|
| 100 |
+
Setting mamba_num_groups to 8 from checkpoint
|
| 101 |
+
Checkpoint did not provide arguments mamba_num_heads
|
| 102 |
+
Checkpoint did not provide arguments hybrid_layer_pattern
|
| 103 |
+
Checkpoint did not provide arguments heterogeneous_layers_config_path
|
| 104 |
+
Checkpoint did not provide arguments heterogeneous_layers_config_encoded_json
|
| 105 |
+
Checkpoint did not provide arguments moe_latent_size
|
| 106 |
+
Setting scale_embeddings_by_sqrt_hidden to False from checkpoint
|
| 107 |
+
Setting residual_output_scaling to False from checkpoint
|
| 108 |
+
Setting sandwich_norm to False from checkpoint
|
| 109 |
+
Setting keel to False from checkpoint
|
| 110 |
+
Checkpoint did not provide arguments keel_alpha
|
| 111 |
+
Setting pnglu to False from checkpoint
|
| 112 |
+
Setting pnglu_fusion to True from checkpoint
|
| 113 |
+
Setting pn3glu to False from checkpoint
|
| 114 |
+
Setting xpr to False from checkpoint
|
| 115 |
+
Setting gxpr to False from checkpoint
|
| 116 |
+
Setting gxpry to False from checkpoint
|
| 117 |
+
Setting gxprv2 to False from checkpoint
|
| 118 |
+
Setting xr2 to False from checkpoint
|
| 119 |
+
Setting gxr2 to False from checkpoint
|
| 120 |
+
Setting xr2glu to False from checkpoint
|
| 121 |
+
Setting xssglu to False from checkpoint
|
| 122 |
+
Setting polynorm to False from checkpoint
|
| 123 |
+
Setting qk_layernorm to False from checkpoint
|
| 124 |
+
Setting attention_output_gate to False from checkpoint
|
| 125 |
+
Setting tokenizer_model to /capstor/store/cscs/swissai/infra01/users/rsinghal/1pp-training/tokenizers/smollm2_pretrain from checkpoint
|
| 126 |
+
Setting tokenizer_type to HuggingFaceTokenizer from checkpoint
|
| 127 |
+
Checkpoint did not provide arguments tiktoken_pattern
|
| 128 |
+
Setting padded_vocab_size to 49280 from checkpoint
|
| 129 |
+
Setting tensor_model_parallel_size to 1 from checkpoint
|
| 130 |
+
Setting pipeline_model_parallel_size to 1 from checkpoint
|
| 131 |
+
Checkpoint did not provide arguments virtual_pipeline_model_parallel_size
|
| 132 |
+
Checkpoint did not provide arguments num_layers_per_virtual_pipeline_stage
|
| 133 |
+
Setting expert_model_parallel_size to 1 from checkpoint
|
| 134 |
+
using world size: 1, data-parallel size: 1, context-parallel size: 1, hierarchical context-parallel sizes: None, tensor-model-parallel size: 1, pipeline-model-parallel size: 1
|
| 135 |
+
setting global batch size to 1
|
| 136 |
+
Number of virtual stages per pipeline stage: None
|
| 137 |
+
accumulate and all-reduce gradients in fp32 for bfloat16 data type.
|
| 138 |
+
using torch.bfloat16 for parameters ...
|
| 139 |
+
------------------------ arguments ------------------------
|
| 140 |
+
account_for_embedding_in_pipeline_split ......... False
|
| 141 |
+
account_for_loss_in_pipeline_split .............. False
|
| 142 |
+
accumulate_allreduce_grads_in_fp32 .............. True
|
| 143 |
+
activation_func_clamp_value ..................... None
|
| 144 |
+
activation_func_fp8_input_store ................. False
|
| 145 |
+
adam_beta1 ...................................... 0.9
|
| 146 |
+
adam_beta2 ...................................... 0.999
|
| 147 |
+
adam_eps ........................................ 1e-08
|
| 148 |
+
add_bias_linear ................................. False
|
| 149 |
+
add_position_embedding .......................... True
|
| 150 |
+
add_qkv_bias .................................... False
|
| 151 |
+
adlr_autoresume ................................. False
|
| 152 |
+
adlr_autoresume_interval ........................ 1000
|
| 153 |
+
align_grad_reduce ............................... True
|
| 154 |
+
align_param_gather .............................. False
|
| 155 |
+
allow_ambiguous_pad_tokens ...................... False
|
| 156 |
+
app_tag_run_name ................................ None
|
| 157 |
+
app_tag_run_version ............................. 0.0.0
|
| 158 |
+
apply_query_key_layer_scaling ................... False
|
| 159 |
+
apply_residual_connection_post_layernorm ........ False
|
| 160 |
+
apply_rope_fusion ............................... True
|
| 161 |
+
apply_wd_to_qk_layernorm ........................ False
|
| 162 |
+
async_ckpt_cpu_priority ......................... 10
|
| 163 |
+
async_ckpt_io_priority .......................... 3
|
| 164 |
+
async_save ...................................... False
|
| 165 |
+
async_strategy .................................. nvrx
|
| 166 |
+
attention_backend ............................... AttnBackend.auto
|
| 167 |
+
attention_dropout ............................... 0.0
|
| 168 |
+
attention_output_gate ........................... False
|
| 169 |
+
attention_softmax_in_fp32 ....................... False
|
| 170 |
+
auto_detect_ckpt_format ......................... False
|
| 171 |
+
barrier_with_L1_time ............................ True
|
| 172 |
+
batch_invariant_mode ............................ False
|
| 173 |
+
bert_binary_head ................................ True
|
| 174 |
+
bert_embedder_type .............................. megatron
|
| 175 |
+
bert_load ....................................... None
|
| 176 |
+
bf16 ............................................ True
|
| 177 |
+
bias_dropout_fusion ............................. False
|
| 178 |
+
bias_gelu_fusion ................................ False
|
| 179 |
+
bias_swiglu_fusion .............................. True
|
| 180 |
+
biencoder_projection_dim ........................ 0
|
| 181 |
+
biencoder_shared_query_context_model ............ False
|
| 182 |
+
block_data_path ................................. None
|
| 183 |
+
cache_mla_latents ............................... False
|
| 184 |
+
calc_ft_timeouts ................................ False
|
| 185 |
+
calculate_per_token_loss ........................ False
|
| 186 |
+
check_for_large_grads ........................... False
|
| 187 |
+
check_for_nan_in_loss_and_grad .................. True
|
| 188 |
+
check_for_spiky_loss ............................ False
|
| 189 |
+
check_grad_norm ................................. False
|
| 190 |
+
check_grad_norm_threshold ....................... 5.0
|
| 191 |
+
check_weight_hash_across_dp_replicas_interval ... None
|
| 192 |
+
ckpt_assume_constant_structure .................. False
|
| 193 |
+
ckpt_convert_format ............................. None
|
| 194 |
+
ckpt_convert_save ............................... None
|
| 195 |
+
ckpt_convert_update_legacy_dist_opt_format ...... False
|
| 196 |
+
ckpt_format ..................................... torch_dist
|
| 197 |
+
ckpt_fully_parallel_load ........................ False
|
| 198 |
+
ckpt_fully_parallel_load_exchange_algo .......... broadcast
|
| 199 |
+
ckpt_fully_parallel_load_process_group .......... dp
|
| 200 |
+
ckpt_fully_parallel_save ........................ True
|
| 201 |
+
ckpt_fully_parallel_save_deprecated ............. False
|
| 202 |
+
ckpt_fully_parallel_save_process_group .......... dp
|
| 203 |
+
ckpt_step ....................................... None
|
| 204 |
+
classes_fraction ................................ 1.0
|
| 205 |
+
clip_grad ....................................... 1.0
|
| 206 |
+
clone_scatter_output_in_embedding ............... True
|
| 207 |
+
config_logger_dir ...............................
|
| 208 |
+
consumed_train_samples .......................... 0
|
| 209 |
+
consumed_valid_samples .......................... 0
|
| 210 |
+
context_parallel_size ........................... 1
|
| 211 |
+
cp_comm_type .................................... ['p2p']
|
| 212 |
+
cpu_offloading_num_layers ....................... 0
|
| 213 |
+
cpu_offloading_retain_pinned_cpu_buffers ........ False
|
| 214 |
+
create_all_gather_group ......................... False
|
| 215 |
+
create_attention_mask_in_dataloader ............. True
|
| 216 |
+
cross_entropy_fusion_impl ....................... native
|
| 217 |
+
cross_entropy_loss_fusion ....................... False
|
| 218 |
+
cuda_graph_impl ................................. none
|
| 219 |
+
cuda_graph_scope ................................ []
|
| 220 |
+
cuda_graph_warmup_steps ......................... 3
|
| 221 |
+
data_args_path .................................. None
|
| 222 |
+
data_cache_path ................................. None
|
| 223 |
+
data_parallel_random_init ....................... False
|
| 224 |
+
data_parallel_sharding_strategy ................. no_shard
|
| 225 |
+
data_parallel_size .............................. 1
|
| 226 |
+
data_path ....................................... None
|
| 227 |
+
data_per_class_fraction ......................... 1.0
|
| 228 |
+
data_sharding ................................... True
|
| 229 |
+
dataloader_defer_npy_index_mmap ................. False
|
| 230 |
+
dataloader_fast_cache_load ...................... False
|
| 231 |
+
dataloader_inter_document_masking ............... False
|
| 232 |
+
dataloader_type ................................. single
|
| 233 |
+
ddp_average_in_collective ....................... False
|
| 234 |
+
ddp_bucket_size ................................. None
|
| 235 |
+
ddp_num_buckets ................................. None
|
| 236 |
+
ddp_pad_buckets_for_high_nccl_busbw ............. False
|
| 237 |
+
ddp_param_name_patterns_for_fp32_local_accumulation []
|
| 238 |
+
ddp_reduce_scatter_with_fp32_accumulation ....... False
|
| 239 |
+
decode_only_cuda_graphs ......................... False
|
| 240 |
+
decoder_first_pipeline_num_layers ............... None
|
| 241 |
+
decoder_last_pipeline_num_layers ................ None
|
| 242 |
+
decoder_num_layers .............................. None
|
| 243 |
+
decoder_seq_length .............................. None
|
| 244 |
+
decoupled_lr .................................... None
|
| 245 |
+
decoupled_min_lr ................................ None
|
| 246 |
+
decrease_batch_size_if_needed ................... False
|
| 247 |
+
defer_embedding_wgrad_compute ................... False
|
| 248 |
+
delay_wgrad_compute ............................. False
|
| 249 |
+
deprecated_use_mcore_models ..................... False
|
| 250 |
+
deterministic_mode .............................. False
|
| 251 |
+
dino_bottleneck_size ............................ 256
|
| 252 |
+
dino_freeze_last_layer .......................... 1
|
| 253 |
+
dino_head_hidden_size ........................... 2048
|
| 254 |
+
dino_local_crops_number ......................... 10
|
| 255 |
+
dino_local_img_size ............................. 96
|
| 256 |
+
dino_norm_last_layer ............................ False
|
| 257 |
+
dino_teacher_temp ............................... 0.07
|
| 258 |
+
dino_warmup_teacher_temp ........................ 0.04
|
| 259 |
+
dino_warmup_teacher_temp_epochs ................. 30
|
| 260 |
+
disable_bf16_reduced_precision_matmul ........... False
|
| 261 |
+
disable_jit_fuser ............................... False
|
| 262 |
+
disable_straggler_on_startup .................... False
|
| 263 |
+
disable_symmetric_registration .................. False
|
| 264 |
+
dist_ckpt_format_deprecated ..................... None
|
| 265 |
+
dist_ckpt_optim_fully_reshardable ............... False
|
| 266 |
+
dist_ckpt_save_pre_mcore_014 .................... False
|
| 267 |
+
dist_ckpt_strictness ............................ assume_ok_unexpected
|
| 268 |
+
dist_ckpt_workers ............................... 1
|
| 269 |
+
distrib_optim_fully_reshardable_mem_efficient ... False
|
| 270 |
+
distribute_saved_activations .................... False
|
| 271 |
+
distributed_backend ............................. nccl
|
| 272 |
+
distributed_timeout_minutes ..................... 10
|
| 273 |
+
distributed_timeout_seconds_after_init .......... None
|
| 274 |
+
dsa_indexer_head_dim ............................ None
|
| 275 |
+
dsa_indexer_loss_coeff .......................... None
|
| 276 |
+
dsa_indexer_n_heads ............................. None
|
| 277 |
+
dsa_indexer_topk ................................ None
|
| 278 |
+
dsa_indexer_use_sparse_loss ..................... False
|
| 279 |
+
dump_param_to_param_group_map ................... None
|
| 280 |
+
embedding_init_method_std ....................... None
|
| 281 |
+
embedding_lr_multiplier ......................... None
|
| 282 |
+
embedding_path .................................. None
|
| 283 |
+
empty_unused_memory_level ....................... 0
|
| 284 |
+
enable_chunked_prefill .......................... False
|
| 285 |
+
enable_cuda_graph ............................... False
|
| 286 |
+
enable_experimental ............................. False
|
| 287 |
+
enable_ft_package ............................... False
|
| 288 |
+
enable_full_sharding_in_hsdp .................... False
|
| 289 |
+
enable_msc ...................................... True
|
| 290 |
+
enable_one_logger ............................... False
|
| 291 |
+
encoder_num_layers .............................. 24
|
| 292 |
+
encoder_seq_length .............................. 4096
|
| 293 |
+
end_weight_decay ................................ 0.01
|
| 294 |
+
eod_mask_loss ................................... False
|
| 295 |
+
ep_overlap_early_attn_memory_release ............ False
|
| 296 |
+
error_injection_rate ............................ 0
|
| 297 |
+
error_injection_type ............................ transient_error
|
| 298 |
+
eval_interval ................................... None
|
| 299 |
+
eval_iters ...................................... 100
|
| 300 |
+
evidence_data_path .............................. None
|
| 301 |
+
exit_duration_in_mins ........................... None
|
| 302 |
+
exit_interval ................................... None
|
| 303 |
+
exit_on_missing_checkpoint ...................... True
|
| 304 |
+
exit_signal ..................................... 15
|
| 305 |
+
exit_signal_handler ............................. False
|
| 306 |
+
exit_signal_handler_for_dataloader .............. False
|
| 307 |
+
exit_signal_handler_for_training ................ False
|
| 308 |
+
exp_avg_dtype ................................... torch.float32
|
| 309 |
+
exp_avg_sq_dtype ................................ torch.float32
|
| 310 |
+
experimental_attention_variant .................. None
|
| 311 |
+
expert_model_parallel_size ...................... 1
|
| 312 |
+
expert_tensor_parallel_size ..................... 1
|
| 313 |
+
external_cuda_graph ............................. False
|
| 314 |
+
fake_process_group .............................. False
|
| 315 |
+
ffn_hidden_size ................................. 4608
|
| 316 |
+
fim_data ........................................ False
|
| 317 |
+
fim_eod_token ................................... <|endoftext|>
|
| 318 |
+
fim_fragment_rate ............................... None
|
| 319 |
+
fim_middle_token ................................ <fim_middle>
|
| 320 |
+
fim_no_prefix ................................... None
|
| 321 |
+
fim_pad_token ................................... <fim_pad>
|
| 322 |
+
fim_prefix_token ................................ <fim_prefix>
|
| 323 |
+
fim_rate ........................................ 0.5
|
| 324 |
+
fim_split_sample ................................ None
|
| 325 |
+
fim_spm_rate .................................... 0.5
|
| 326 |
+
fim_suffix_token ................................ <fim_suffix>
|
| 327 |
+
fine_grained_activation_offloading .............. False
|
| 328 |
+
finetune ........................................ False
|
| 329 |
+
first_last_layers_bf16 .......................... False
|
| 330 |
+
flash_decode .................................... False
|
| 331 |
+
flight_recorder_dump_on_timeout ................. True
|
| 332 |
+
flight_recorder_dump_path ....................... None
|
| 333 |
+
flight_recorder_extra_dump_on_exec .............. True
|
| 334 |
+
flight_recorder_include_only_active ............. True
|
| 335 |
+
flight_recorder_include_stack_trace ............. False
|
| 336 |
+
flight_recorder_trace_buffer_size ............... 2000
|
| 337 |
+
fp16 ............................................ False
|
| 338 |
+
fp16_lm_cross_entropy ........................... False
|
| 339 |
+
fp32_residual_connection ........................ False
|
| 340 |
+
fp4 ............................................. None
|
| 341 |
+
fp4_param ....................................... False
|
| 342 |
+
fp4_quantizer_factory ........................... None
|
| 343 |
+
fp4_recipe ...................................... nvfp4
|
| 344 |
+
fp8 ............................................. None
|
| 345 |
+
fp8_amax_compute_algo ........................... most_recent
|
| 346 |
+
fp8_amax_history_len ............................ 1
|
| 347 |
+
fp8_interval .................................... 1
|
| 348 |
+
fp8_margin ...................................... 0
|
| 349 |
+
fp8_param_gather ................................ False
|
| 350 |
+
fp8_quantizer_factory ........................... None
|
| 351 |
+
fp8_recipe ...................................... delayed
|
| 352 |
+
fp8_wgrad ....................................... True
|
| 353 |
+
fsdp_double_buffer .............................. False
|
| 354 |
+
fsdp_manual_registration ........................ False
|
| 355 |
+
ft_num_warmup_iters ............................. 5
|
| 356 |
+
full_validation ................................. False
|
| 357 |
+
fused_residual_rmsnorm .......................... False
|
| 358 |
+
gain_parametrization ............................ softplus
|
| 359 |
+
gains_lr ........................................ None
|
| 360 |
+
gains_no_clamp_min .............................. False
|
| 361 |
+
global_batch_size ............................... 1
|
| 362 |
+
glu_linear_offset ............................... 0.0
|
| 363 |
+
goldfish_h ...................................... 50
|
| 364 |
+
goldfish_k ...................................... 50
|
| 365 |
+
goldfish_loss ................................... False
|
| 366 |
+
grad_reduce_in_bf16 ............................. False
|
| 367 |
+
gradient_accumulation_fusion .................... True
|
| 368 |
+
gradient_reduce_div_fusion ...................... True
|
| 369 |
+
group_query_attention ........................... True
|
| 370 |
+
grpo_clamp_eps_lower ............................ 0.01
|
| 371 |
+
grpo_clamp_eps_upper ............................ 0.01
|
| 372 |
+
grpo_entropy_term_weight ........................ 0.0
|
| 373 |
+
grpo_filter_groups_with_same_reward ............. False
|
| 374 |
+
grpo_group_size ................................. 2
|
| 375 |
+
grpo_iterations ................................. 2
|
| 376 |
+
grpo_kl_beta .................................... 0.001
|
| 377 |
+
grpo_prompts_per_step ........................... 32
|
| 378 |
+
gxpr ............................................ False
|
| 379 |
+
gxpr_fusion ..................................... True
|
| 380 |
+
gxprv2 .......................................... False
|
| 381 |
+
gxpry ........................................... False
|
| 382 |
+
gxr2 ............................................ False
|
| 383 |
+
head_lr_mult .................................... 1.0
|
| 384 |
+
heterogeneous_layers_config_encoded_json ........ None
|
| 385 |
+
heterogeneous_layers_config_path ................ None
|
| 386 |
+
hidden_dropout .................................. 0.0
|
| 387 |
+
hidden_size ..................................... 1152
|
| 388 |
+
hierarchical_context_parallel_sizes ............. None
|
| 389 |
+
high_priority_stream_groups ..................... []
|
| 390 |
+
hybrid_context_parallel ......................... False
|
| 391 |
+
hybrid_layer_pattern ............................ None
|
| 392 |
+
hybrid_override_pattern ......................... None
|
| 393 |
+
hypersphere_embedding_mode ...................... row
|
| 394 |
+
hypersphere_gains_mode .......................... rowcol
|
| 395 |
+
hypersphere_gains_mode_embedding ................ none
|
| 396 |
+
hypersphere_gains_mode_output ................... inherit
|
| 397 |
+
hypersphere_gains_mode_router ................... rowcol
|
| 398 |
+
hypersphere_mode ................................ flat
|
| 399 |
+
hypersphere_preserve_init ....................... False
|
| 400 |
+
hypersphere_radius_mode ......................... shape_native
|
| 401 |
+
hypersphere_router_mode ......................... row
|
| 402 |
+
hypersphere_scale_out_proj_init ................. False
|
| 403 |
+
hypersphere_tangential_grad ..................... False
|
| 404 |
+
hysteresis ...................................... 2
|
| 405 |
+
ict_head_size ................................... None
|
| 406 |
+
ict_load ........................................ None
|
| 407 |
+
img_h ........................................... 224
|
| 408 |
+
img_w ........................................... 224
|
| 409 |
+
indexer_batch_size .............................. 128
|
| 410 |
+
indexer_log_interval ............................ 1000
|
| 411 |
+
inference_batch_times_seqlen_threshold .......... -1
|
| 412 |
+
inference_coordinator_port ...................... None
|
| 413 |
+
inference_disable_triton_nvls_kernels ........... False
|
| 414 |
+
inference_dynamic_batching ...................... False
|
| 415 |
+
inference_dynamic_batching_block_size ........... 256
|
| 416 |
+
inference_dynamic_batching_buffer_size_gb ....... 40.0
|
| 417 |
+
inference_dynamic_batching_cuda_graph_max_tokens 16384
|
| 418 |
+
inference_dynamic_batching_cuda_graph_mixed_prefill_count 16
|
| 419 |
+
inference_dynamic_batching_enable_prefix_caching False
|
| 420 |
+
inference_dynamic_batching_mamba_memory_ratio ... None
|
| 421 |
+
inference_dynamic_batching_max_requests ......... None
|
| 422 |
+
inference_dynamic_batching_max_tokens ........... None
|
| 423 |
+
inference_dynamic_batching_num_cuda_graphs ...... 16
|
| 424 |
+
inference_dynamic_batching_paused_buffer_size_gb None
|
| 425 |
+
inference_dynamic_batching_prefix_caching_coordinator_policy first_prefix_block
|
| 426 |
+
inference_dynamic_batching_prefix_caching_eviction_policy ref_zero
|
| 427 |
+
inference_dynamic_batching_prefix_caching_mamba_gb None
|
| 428 |
+
inference_dynamic_batching_prefix_caching_routing_alpha 0.5
|
| 429 |
+
inference_dynamic_batching_track_generated_token_events False
|
| 430 |
+
inference_dynamic_batching_track_paused_request_events False
|
| 431 |
+
inference_dynamic_batching_unified_memory_level . 0
|
| 432 |
+
inference_fuse_tp_communication ................. False
|
| 433 |
+
inference_grouped_gemm_backend .................. auto
|
| 434 |
+
inference_logging_step_interval ................. 0
|
| 435 |
+
inference_max_requests .......................... 8
|
| 436 |
+
inference_max_seq_length ........................ 2560
|
| 437 |
+
inference_moe_disable_fused_quant_kernels ....... False
|
| 438 |
+
inference_rng_tracker ........................... False
|
| 439 |
+
inference_text_gen_server_logging ............... False
|
| 440 |
+
inference_use_synchronous_zmq_collectives ....... False
|
| 441 |
+
inference_wandb_logging ......................... False
|
| 442 |
+
init_method_std ................................. 0.02
|
| 443 |
+
init_method_xavier_uniform ...................... False
|
| 444 |
+
init_model_with_meta_device ..................... False
|
| 445 |
+
initial_loss_scale .............................. 4294967296
|
| 446 |
+
inprocess_active_world_size ..................... 1
|
| 447 |
+
inprocess_barrier_timeout ....................... 120
|
| 448 |
+
inprocess_completion_timeout .................... 120
|
| 449 |
+
inprocess_empty_cuda_cache ...................... False
|
| 450 |
+
inprocess_granularity ........................... node
|
| 451 |
+
inprocess_hard_timeout .......................... 90
|
| 452 |
+
inprocess_heartbeat_interval .................... 30
|
| 453 |
+
inprocess_heartbeat_timeout ..................... 60
|
| 454 |
+
inprocess_last_call_wait ........................ 1
|
| 455 |
+
inprocess_max_iterations ........................ None
|
| 456 |
+
inprocess_monitor_process_interval .............. 1.0
|
| 457 |
+
inprocess_monitor_thread_interval ............... 1.0
|
| 458 |
+
inprocess_progress_watchdog_interval ............ 1.0
|
| 459 |
+
inprocess_restart ............................... False
|
| 460 |
+
inprocess_soft_timeout .......................... 60
|
| 461 |
+
inprocess_termination_grace_time ................ 1
|
| 462 |
+
is_hybrid_model ................................. False
|
| 463 |
+
iter_per_epoch .................................. 1250
|
| 464 |
+
iteration ....................................... 582
|
| 465 |
+
iterations_to_skip .............................. []
|
| 466 |
+
keel ............................................ False
|
| 467 |
+
keel_alpha ...................................... None
|
| 468 |
+
keep_fp8_transpose_cache ........................ False
|
| 469 |
+
kitchen_attention_backend ....................... sdpa
|
| 470 |
+
kv_channels ..................................... 128
|
| 471 |
+
kv_lora_rank .................................... 32
|
| 472 |
+
langrl_env_config ............................... None
|
| 473 |
+
layernorm_epsilon ............................... 1e-05
|
| 474 |
+
layernorm_zero_centered_gamma ................... False
|
| 475 |
+
lazy_mpu_init ................................... False
|
| 476 |
+
lglu ............................................ False
|
| 477 |
+
linear_attention_allow_neg_eigval ............... False
|
| 478 |
+
linear_attention_beta_bias_init ................. 0.0
|
| 479 |
+
linear_attention_beta_scale ..................... 1.0
|
| 480 |
+
linear_attention_carried_state_max_frob ......... 0.0
|
| 481 |
+
linear_attention_carry_state .................... False
|
| 482 |
+
linear_attention_freq ........................... None
|
| 483 |
+
linear_attention_full_rank_output_gate .......... False
|
| 484 |
+
linear_attention_learnable_initial_state ........ False
|
| 485 |
+
linear_attention_n_erase ........................ 0
|
| 486 |
+
linear_attention_n_householder .................. 1
|
| 487 |
+
linear_attention_output_gate_form ............... per_channel
|
| 488 |
+
linear_attention_qk_norm ........................ l2norm
|
| 489 |
+
linear_attention_qk_norm_init_scale ............. 1.0
|
| 490 |
+
linear_attention_safe_output_gate ............... False
|
| 491 |
+
linear_attention_safe_output_gate_lower_bound ... -5.0
|
| 492 |
+
linear_attention_use_decay ...................... True
|
| 493 |
+
linear_attention_use_output_gate ................ True
|
| 494 |
+
linear_attention_v_norm ......................... none
|
| 495 |
+
linear_conv_kernel_dim .......................... 4
|
| 496 |
+
linear_key_head_dim ............................. 128
|
| 497 |
+
linear_num_key_heads ............................ 16
|
| 498 |
+
linear_num_value_heads .......................... 32
|
| 499 |
+
linear_value_head_dim ........................... 128
|
| 500 |
+
lion_beta1 ...................................... 0.95
|
| 501 |
+
lion_beta2 ...................................... 0.98
|
| 502 |
+
load ............................................ /iopsstor/scratch/cscs/rsinghal/1pp-hf-tmp/1pp-0.5b-asst-sft/torch
|
| 503 |
+
load_main_params_from_ckpt ...................... False
|
| 504 |
+
local_rank ...................................... 0
|
| 505 |
+
log_device_memory_used .......................... False
|
| 506 |
+
log_energy ...................................... False
|
| 507 |
+
log_interval .................................... 100
|
| 508 |
+
log_loss_scale_to_tensorboard ................... True
|
| 509 |
+
log_max_attention_logit ......................... False
|
| 510 |
+
log_memory_interval ............................. None
|
| 511 |
+
log_memory_to_tensorboard ....................... False
|
| 512 |
+
log_muon_gains .................................. False
|
| 513 |
+
log_muon_param_rms .............................. False
|
| 514 |
+
log_muon_per_layer .............................. False
|
| 515 |
+
log_muon_sparsity ............................... False
|
| 516 |
+
log_num_zeros_in_grad ........................... False
|
| 517 |
+
log_params_norm ................................. False
|
| 518 |
+
log_progress .................................... False
|
| 519 |
+
log_straggler ................................... False
|
| 520 |
+
log_throughput .................................. False
|
| 521 |
+
log_timers_to_tensorboard ....................... False
|
| 522 |
+
log_validation_ppl_to_tensorboard ............... False
|
| 523 |
+
log_world_size_to_tensorboard ................... False
|
| 524 |
+
logging_level ................................... None
|
| 525 |
+
loss_mask_segment_suffix ........................ None
|
| 526 |
+
loss_mask_train_segments ........................ None
|
| 527 |
+
loss_scale ...................................... None
|
| 528 |
+
loss_scale_window ............................... 1000
|
| 529 |
+
lr .............................................. None
|
| 530 |
+
lr_decay_iters .................................. None
|
| 531 |
+
lr_decay_samples ................................ None
|
| 532 |
+
lr_decay_style .................................. linear
|
| 533 |
+
lr_warmup_fraction .............................. None
|
| 534 |
+
lr_warmup_init .................................. 0.0
|
| 535 |
+
lr_warmup_iters ................................. 0
|
| 536 |
+
lr_warmup_samples ............................... 0
|
| 537 |
+
lr_wsd_decay_iters .............................. None
|
| 538 |
+
lr_wsd_decay_samples ............................ None
|
| 539 |
+
lr_wsd_decay_style .............................. exponential
|
| 540 |
+
main_grads_dtype ................................ torch.float32
|
| 541 |
+
main_params_dtype ............................... torch.float32
|
| 542 |
+
make_vocab_size_divisible_by .................... 128
|
| 543 |
+
mamba_head_dim .................................. 64
|
| 544 |
+
mamba_inference_conv_states_dtype ............... torch.bfloat16
|
| 545 |
+
mamba_inference_ssm_states_dtype ................ torch.bfloat16
|
| 546 |
+
mamba_num_groups ................................ 8
|
| 547 |
+
mamba_num_heads ................................. None
|
| 548 |
+
mamba_state_dim ................................. 128
|
| 549 |
+
manual_gc ....................................... False
|
| 550 |
+
manual_gc_eval .................................. True
|
| 551 |
+
manual_gc_interval .............................. 0
|
| 552 |
+
mask_factor ..................................... 1.0
|
| 553 |
+
mask_prob ....................................... 0.15
|
| 554 |
+
mask_type ....................................... random
|
| 555 |
+
masked_softmax_fusion ........................... False
|
| 556 |
+
matrix_lr ....................................... None
|
| 557 |
+
max_docs_per_bin ................................ 0
|
| 558 |
+
max_position_embeddings ......................... 4096
|
| 559 |
+
max_seqlen_per_dp_cp_rank ....................... None
|
| 560 |
+
max_tokens_to_oom ............................... 12000
|
| 561 |
+
md_normalize_update_to_weight_norm .............. False
|
| 562 |
+
md_router_use_orthogonal_updates ................ None
|
| 563 |
+
megatron_fsdp_grad_comm_dtype ................... None
|
| 564 |
+
megatron_fsdp_main_grads_dtype .................. None
|
| 565 |
+
megatron_fsdp_main_params_dtype ................. torch.float32
|
| 566 |
+
memory_snapshot_path ............................ snapshot.pickle
|
| 567 |
+
merge_file ...................................... None
|
| 568 |
+
micro_batch_size ................................ 1
|
| 569 |
+
microbatch_group_size_per_vp_stage .............. None
|
| 570 |
+
mid_level_dataset_surplus ....................... 0.005
|
| 571 |
+
min_loss_scale .................................. 1.0
|
| 572 |
+
min_lr .......................................... 0.0
|
| 573 |
+
min_lr_mode ..................................... relative
|
| 574 |
+
min_offloaded_tensor_size ....................... 1048576
|
| 575 |
+
mla_down_proj_fusion ............................ False
|
| 576 |
+
mlp_chunks_for_prefill .......................... 1
|
| 577 |
+
mmap_bin_files .................................. True
|
| 578 |
+
mock_data ....................................... True
|
| 579 |
+
moe_apply_probs_on_input ........................ False
|
| 580 |
+
moe_aux_loss_coeff .............................. 0.0
|
| 581 |
+
moe_deepep_num_sms .............................. 20
|
| 582 |
+
moe_enable_deepep ............................... False
|
| 583 |
+
moe_enable_routing_replay ....................... False
|
| 584 |
+
moe_expert_capacity_factor ...................... None
|
| 585 |
+
moe_ffn_hidden_size ............................. None
|
| 586 |
+
moe_flex_dispatcher_backend ..................... deepep
|
| 587 |
+
moe_grouped_gemm ................................ False
|
| 588 |
+
moe_hybridep_num_sms ............................ 16
|
| 589 |
+
moe_input_jitter_eps ............................ None
|
| 590 |
+
moe_latent_size ................................. None
|
| 591 |
+
moe_layer_freq .................................. 1
|
| 592 |
+
moe_layer_recompute ............................. False
|
| 593 |
+
moe_offload_activations ......................... None
|
| 594 |
+
moe_offload_main_grad ........................... False
|
| 595 |
+
moe_offloading_chunk_size ....................... -1
|
| 596 |
+
moe_offloading_experts_debug_mode ............... False
|
| 597 |
+
moe_offloading_experts_skip_post_backward_hook .. False
|
| 598 |
+
moe_offloading_experts_te_style_init ............ True
|
| 599 |
+
moe_offloading_mode ............................. fine-grained
|
| 600 |
+
moe_offloading_num_chunks ....................... 8
|
| 601 |
+
moe_offloading_num_stages ....................... 2
|
| 602 |
+
moe_pad_expert_input_to_capacity ................ False
|
| 603 |
+
moe_pad_experts_for_cuda_graph_inference ........ False
|
| 604 |
+
moe_per_layer_logging ........................... False
|
| 605 |
+
moe_permute_fusion .............................. False
|
| 606 |
+
moe_router_bias_metrics ......................... False
|
| 607 |
+
moe_router_bias_update_rate ..................... 0.001
|
| 608 |
+
moe_router_dtype ................................ None
|
| 609 |
+
moe_router_enable_expert_bias ................... False
|
| 610 |
+
moe_router_force_biased ......................... None
|
| 611 |
+
moe_router_force_load_balancing ................. False
|
| 612 |
+
moe_router_fusion ............................... False
|
| 613 |
+
moe_router_group_topk ........................... None
|
| 614 |
+
moe_router_inference_violation_metrics .......... []
|
| 615 |
+
moe_router_load_balancing_type .................. aux_loss
|
| 616 |
+
moe_router_num_groups ........................... None
|
| 617 |
+
moe_router_padding_for_fp8 ...................... False
|
| 618 |
+
moe_router_padding_for_quantization ............. False
|
| 619 |
+
moe_router_pre_softmax .......................... False
|
| 620 |
+
moe_router_quantile_balancing_ema ............... 0.0
|
| 621 |
+
moe_router_quantile_balancing_method ............ histogram
|
| 622 |
+
moe_router_quantile_balancing_num_bins .......... 1000
|
| 623 |
+
moe_router_score_function ....................... softmax
|
| 624 |
+
moe_router_topk ................................. 2
|
| 625 |
+
moe_router_topk_scaling_factor .................. None
|
| 626 |
+
moe_router_violation_metrics .................... ['mbs']
|
| 627 |
+
moe_shared_expert_gate .......................... False
|
| 628 |
+
moe_shared_expert_intermediate_size ............. None
|
| 629 |
+
moe_shared_expert_overlap ....................... False
|
| 630 |
+
moe_token_dispatcher_type ....................... allgather
|
| 631 |
+
moe_token_drop_policy ........................... probs
|
| 632 |
+
moe_upcycling_granularity ....................... 1
|
| 633 |
+
moe_use_extra_fp8_param_storage ................. False
|
| 634 |
+
moe_use_fp8_activation .......................... False
|
| 635 |
+
moe_use_fp8_dispatch ............................ False
|
| 636 |
+
moe_use_inplace_fp8_param ....................... False
|
| 637 |
+
moe_use_offloading_experts ...................... False
|
| 638 |
+
moe_use_upcycling ............................... False
|
| 639 |
+
moe_z_loss_coeff ................................ None
|
| 640 |
+
mrope_section ................................... None
|
| 641 |
+
mscale .......................................... 1.0
|
| 642 |
+
mscale_all_dim .................................. 0.0
|
| 643 |
+
mtp_hybrid_override_pattern ..................... None
|
| 644 |
+
mtp_loss_scaling_factor ......................... 0.1
|
| 645 |
+
mtp_num_layers .................................. None
|
| 646 |
+
mtp_standalone .................................. False
|
| 647 |
+
mtp_use_repeated_layer .......................... False
|
| 648 |
+
multi_latent_attention .......................... False
|
| 649 |
+
multiple_validation_sets ........................ False
|
| 650 |
+
muon_coefficient_type ........................... quintic
|
| 651 |
+
muon_extra_scale_factor ......................... 1.0
|
| 652 |
+
muon_fp32_matmul_prec ........................... medium
|
| 653 |
+
muon_log_interval ............................... None
|
| 654 |
+
muon_lr_factor .................................. 1.0
|
| 655 |
+
muon_momentum ................................... 0.95
|
| 656 |
+
muon_num_ns_steps ............................... 5
|
| 657 |
+
muon_router_scale_mode .......................... none
|
| 658 |
+
muon_scalar_optimizer ........................... adam
|
| 659 |
+
muon_scale_mode ................................. spectral
|
| 660 |
+
muon_sparsity_thresholds ........................ [1e-20, 1e-10, 1e-30]
|
| 661 |
+
muon_split_fc1 .................................. True
|
| 662 |
+
muon_split_mla_per_head ......................... False
|
| 663 |
+
muon_split_qkv .................................. True
|
| 664 |
+
muon_tp_mode .................................... duplicated
|
| 665 |
+
muon_use_nesterov ............................... False
|
| 666 |
+
mup_attn_scale_power ............................ 1.0
|
| 667 |
+
mup_base_head_dim ............................... None
|
| 668 |
+
mup_base_hidden_size ............................ None
|
| 669 |
+
mup_embedding_mult .............................. 1.0
|
| 670 |
+
mup_output_mult ................................. 1.0
|
| 671 |
+
mup_width_mult .................................. 1.0
|
| 672 |
+
nccl_all_reduce_for_prefill ..................... False
|
| 673 |
+
nccl_communicator_config_path ................... None
|
| 674 |
+
nccl_ub ......................................... False
|
| 675 |
+
no_load_optim ................................... True
|
| 676 |
+
no_load_rng ..................................... True
|
| 677 |
+
no_persist_layer_norm ........................... False
|
| 678 |
+
no_rope_freq .................................... None
|
| 679 |
+
no_save_optim ................................... True
|
| 680 |
+
no_save_rng ..................................... True
|
| 681 |
+
no_weight_decay_cond_type ....................... None
|
| 682 |
+
non_persistent_ckpt_type ........................ None
|
| 683 |
+
non_persistent_global_ckpt_dir .................. None
|
| 684 |
+
non_persistent_local_ckpt_algo .................. fully_parallel
|
| 685 |
+
non_persistent_local_ckpt_dir ................... None
|
| 686 |
+
non_persistent_save_interval .................... None
|
| 687 |
+
normalization ................................... RMSNorm
|
| 688 |
+
num_attention_heads ............................. 9
|
| 689 |
+
num_channels .................................... 3
|
| 690 |
+
num_classes ..................................... 1000
|
| 691 |
+
num_dataset_builder_threads ..................... 1
|
| 692 |
+
num_distributed_optimizer_instances ............. 1
|
| 693 |
+
num_experts ..................................... None
|
| 694 |
+
num_layers ...................................... 24
|
| 695 |
+
num_layers_at_end_in_bf16 ....................... 1
|
| 696 |
+
num_layers_at_start_in_bf16 ..................... 1
|
| 697 |
+
num_layers_per_virtual_pipeline_stage ........... None
|
| 698 |
+
num_query_groups ................................ 3
|
| 699 |
+
num_speculative_tokens .......................... 0
|
| 700 |
+
num_virtual_stages_per_pipeline_rank ............ None
|
| 701 |
+
num_workers ..................................... 2
|
| 702 |
+
object_storage_cache_path ....................... None
|
| 703 |
+
offload_modules ................................. []
|
| 704 |
+
one_logger_async ................................ False
|
| 705 |
+
one_logger_project .............................. megatron-lm
|
| 706 |
+
one_logger_run_name ............................. None
|
| 707 |
+
onnx_safe ....................................... None
|
| 708 |
+
openai_gelu ..................................... False
|
| 709 |
+
optimizer ....................................... adam
|
| 710 |
+
optimizer_cpu_offload ........................... False
|
| 711 |
+
optimizer_cuda_graph ............................ False
|
| 712 |
+
optimizer_offload_fraction ...................... 1.0
|
| 713 |
+
outer_dp_sharding_strategy ...................... no_shard
|
| 714 |
+
output_bert_embeddings .......................... False
|
| 715 |
+
output_lr ....................................... None
|
| 716 |
+
overlap_cpu_optimizer_d2h_h2d ................... False
|
| 717 |
+
overlap_grad_reduce ............................. False
|
| 718 |
+
overlap_moe_expert_parallel_comm ................ False
|
| 719 |
+
overlap_p2p_comm ................................ False
|
| 720 |
+
overlap_p2p_comm_warmup_flush ................... False
|
| 721 |
+
overlap_param_gather ............................ False
|
| 722 |
+
overlap_param_gather_with_optimizer_step ........ False
|
| 723 |
+
override_opt_param_scheduler .................... False
|
| 724 |
+
padded_vocab_size ............................... 49280
|
| 725 |
+
params_dtype .................................... torch.bfloat16
|
| 726 |
+
patch_dim ....................................... 16
|
| 727 |
+
per_dataset_sequences_path ...................... None
|
| 728 |
+
per_split_data_args_path ........................ None
|
| 729 |
+
perform_initialization .......................... False
|
| 730 |
+
perform_rl_step ................................. False
|
| 731 |
+
phase_transition_iterations ..................... None
|
| 732 |
+
pin_cpu_grads ................................... True
|
| 733 |
+
pin_cpu_params .................................. True
|
| 734 |
+
pipeline_model_parallel_comm_backend ............ None
|
| 735 |
+
pipeline_model_parallel_layout .................. None
|
| 736 |
+
pipeline_model_parallel_size .................... 1
|
| 737 |
+
pn3glu .......................................... False
|
| 738 |
+
pnglu ........................................... False
|
| 739 |
+
pnglu_fusion .................................... True
|
| 740 |
+
polynorm ........................................ False
|
| 741 |
+
position_embedding_type ......................... rope
|
| 742 |
+
post_attn_norm_zero_init ........................ False
|
| 743 |
+
prepacked_samples ............................... False
|
| 744 |
+
pretrained_checkpoint ........................... None
|
| 745 |
+
pretraining_packing_strategy .................... greedy
|
| 746 |
+
profile ......................................... False
|
| 747 |
+
profile_ranks ................................... []
|
| 748 |
+
profile_step_end ................................ 12
|
| 749 |
+
profile_step_start .............................. 10
|
| 750 |
+
pytorch_profiler_collect_callstack .............. False
|
| 751 |
+
pytorch_profiler_collect_chakra ................. False
|
| 752 |
+
pytorch_profiler_collect_shapes ................. False
|
| 753 |
+
q_lora_rank ..................................... None
|
| 754 |
+
qk_clip ......................................... False
|
| 755 |
+
qk_clip_alpha ................................... 0.5
|
| 756 |
+
qk_clip_threshold ............................... 100
|
| 757 |
+
qk_head_dim ..................................... 128
|
| 758 |
+
qk_l2_norm ...................................... False
|
| 759 |
+
qk_layernorm .................................... False
|
| 760 |
+
qk_pos_emb_head_dim ............................. 64
|
| 761 |
+
query_in_block_prob ............................. 0.1
|
| 762 |
+
quick_geglu ..................................... False
|
| 763 |
+
rampup_batch_size ............................... None
|
| 764 |
+
rank ............................................ 0
|
| 765 |
+
recompute_granularity ........................... None
|
| 766 |
+
recompute_method ................................ None
|
| 767 |
+
recompute_modules ............................... None
|
| 768 |
+
recompute_num_layers ............................ None
|
| 769 |
+
record_memory_history ........................... False
|
| 770 |
+
refit_method .................................... gloo
|
| 771 |
+
reglu ........................................... False
|
| 772 |
+
relative_attention_max_distance ................. 128
|
| 773 |
+
relative_attention_num_buckets .................. 32
|
| 774 |
+
replication ..................................... False
|
| 775 |
+
replication_factor .............................. 2
|
| 776 |
+
replication_jump ................................ None
|
| 777 |
+
rerun_mode ...................................... validate_results
|
| 778 |
+
rerun_strategy .................................. rerun_in_place
|
| 779 |
+
reset_attention_mask ............................ False
|
| 780 |
+
reset_position_ids .............................. False
|
| 781 |
+
residual_output_scaling ......................... False
|
| 782 |
+
result_rejected_tracker_filename ................ None
|
| 783 |
+
retriever_report_topk_accuracies ................ []
|
| 784 |
+
retriever_score_scaling ......................... False
|
| 785 |
+
reuse_grad_buf_for_mxfp8_param_ag ............... False
|
| 786 |
+
rl_default_temperature .......................... 1.0
|
| 787 |
+
rl_default_top_k ................................ -1
|
| 788 |
+
rl_default_top_p ................................ 0
|
| 789 |
+
rl_generation_batch_size ........................ None
|
| 790 |
+
rl_importance_sampling_truncation_coef .......... None
|
| 791 |
+
rl_inference_expert_model_parallel_size ......... None
|
| 792 |
+
rl_inference_expert_tensor_model_parallel_size .. None
|
| 793 |
+
rl_inference_logprobs_is_correction ............. False
|
| 794 |
+
rl_inference_model_unified_memory_level ......... 0
|
| 795 |
+
rl_inference_parsers ............................ []
|
| 796 |
+
rl_inference_pipeline_model_parallel_size ....... None
|
| 797 |
+
rl_inference_tensor_model_parallel_size ......... None
|
| 798 |
+
rl_kv_cache_management_mode ..................... persist
|
| 799 |
+
rl_num_parallel_generation_batches .............. None
|
| 800 |
+
rl_num_parallel_generations ..................... None
|
| 801 |
+
rl_offload_inference_model_weights_when_idle .... False
|
| 802 |
+
rl_offload_optimizer_during_inference ........... False
|
| 803 |
+
rl_parallel_generation_tasks .................... None
|
| 804 |
+
rl_partial_rollouts ............................. False
|
| 805 |
+
rl_persist_cuda_graphs .......................... False
|
| 806 |
+
rl_prompts_per_eval ............................. 32
|
| 807 |
+
rl_sequence_packing_algo ........................ fifo
|
| 808 |
+
rl_sequence_packing_max_sequences_per_bin ....... 50
|
| 809 |
+
rl_skip_bos_token ............................... False
|
| 810 |
+
rl_training_cuda_graphs ......................... False
|
| 811 |
+
rl_use_sequence_packing ......................... False
|
| 812 |
+
rl_verify_model_weights_swap .................... False
|
| 813 |
+
rlglu ........................................... False
|
| 814 |
+
rope_scaling_factor ............................. 1.0
|
| 815 |
+
rope_type ....................................... None
|
| 816 |
+
rotary_base ..................................... 10000
|
| 817 |
+
rotary_interleaved .............................. False
|
| 818 |
+
rotary_percent .................................. 1.0
|
| 819 |
+
rotary_scaling_factor ........................... 1.0
|
| 820 |
+
rotary_seq_len_interpolation_factor ............. None
|
| 821 |
+
run_workload_inspector_server ................... False
|
| 822 |
+
sample_rate ..................................... 1.0
|
| 823 |
+
sandwich_norm ................................... False
|
| 824 |
+
save ............................................ None
|
| 825 |
+
save_dgrads_interval ............................ None
|
| 826 |
+
save_interval ................................... None
|
| 827 |
+
save_iters ...................................... None
|
| 828 |
+
save_retain_interval ............................ None
|
| 829 |
+
save_wgrads_interval ............................ None
|
| 830 |
+
scale_embeddings_by_sqrt_hidden ................. False
|
| 831 |
+
scatter_gather_tensors_in_pipeline .............. True
|
| 832 |
+
seed ............................................ 1234
|
| 833 |
+
seq_length ...................................... 4096
|
| 834 |
+
sequence_parallel ............................... False
|
| 835 |
+
sft ............................................. False
|
| 836 |
+
sft_tokenizer_prompt_format ..................... nemotron-h-aligned
|
| 837 |
+
sgd_momentum .................................... 0.9
|
| 838 |
+
sharp_enabled_group ............................. None
|
| 839 |
+
short_seq_prob .................................. 0.1
|
| 840 |
+
situ ............................................ False
|
| 841 |
+
skip_train ...................................... False
|
| 842 |
+
skipped_train_samples ........................... 0
|
| 843 |
+
softmax_type .................................... vanilla
|
| 844 |
+
spec ............................................ None
|
| 845 |
+
split ........................................... None
|
| 846 |
+
squared_relu .................................... False
|
| 847 |
+
ssglu ........................................... False
|
| 848 |
+
sssglu .......................................... False
|
| 849 |
+
start_weight_decay .............................. 0.01
|
| 850 |
+
straggler_ctrlr_port ............................ 65535
|
| 851 |
+
straggler_minmax_count .......................... 1
|
| 852 |
+
strict_fsdp_dtensor_load ........................ True
|
| 853 |
+
suggested_communication_unit_size ............... None
|
| 854 |
+
swiglu .......................................... True
|
| 855 |
+
swin_backbone_type .............................. tiny
|
| 856 |
+
symmetric_ar_type ............................... None
|
| 857 |
+
te_precision_config_file ........................ None
|
| 858 |
+
te_rng_tracker .................................. False
|
| 859 |
+
tensor_model_parallel_size ...................... 1
|
| 860 |
+
tensorboard_dir ................................. None
|
| 861 |
+
tensorboard_log_interval ........................ 1
|
| 862 |
+
tensorboard_queue_size .......................... 1000
|
| 863 |
+
test_data_path .................................. None
|
| 864 |
+
test_mode ....................................... False
|
| 865 |
+
tiktoken_num_special_tokens ..................... 1000
|
| 866 |
+
tiktoken_pattern ................................ None
|
| 867 |
+
tiktoken_special_tokens ......................... None
|
| 868 |
+
timing_log_level ................................ 0
|
| 869 |
+
timing_log_option ............................... minmax
|
| 870 |
+
titles_data_path ................................ None
|
| 871 |
+
tokenizer_hf_include_special_tokens ............. True
|
| 872 |
+
tokenizer_hf_no_include_special_tokens .......... False
|
| 873 |
+
tokenizer_hf_no_use_fast ........................ False
|
| 874 |
+
tokenizer_hf_use_fast ........................... True
|
| 875 |
+
tokenizer_metadata .............................. None
|
| 876 |
+
tokenizer_model ................................. /capstor/store/cscs/swissai/infra01/users/rsinghal/1pp-training/tokenizers/smollm2_pretrain
|
| 877 |
+
tokenizer_sentencepiece_legacy .................. False
|
| 878 |
+
tokenizer_special_tokens ........................ None
|
| 879 |
+
tokenizer_type .................................. HuggingFaceTokenizer
|
| 880 |
+
torch_fsdp2_reshard_after_forward ............... True
|
| 881 |
+
tp_comm_bootstrap_backend ....................... nccl
|
| 882 |
+
tp_comm_bulk_dgrad .............................. True
|
| 883 |
+
tp_comm_bulk_wgrad .............................. True
|
| 884 |
+
tp_comm_overlap ................................. False
|
| 885 |
+
tp_comm_overlap_ag .............................. True
|
| 886 |
+
tp_comm_overlap_cfg ............................. None
|
| 887 |
+
tp_comm_overlap_rs .............................. True
|
| 888 |
+
tp_comm_overlap_rs_dgrad ........................ False
|
| 889 |
+
tp_comm_split_ag ................................ True
|
| 890 |
+
tp_comm_split_rs ................................ True
|
| 891 |
+
train_data_path ................................. None
|
| 892 |
+
train_iters ..................................... None
|
| 893 |
+
train_samples ................................... None
|
| 894 |
+
train_sync_interval ............................. None
|
| 895 |
+
transformer_impl ................................ transformer_engine
|
| 896 |
+
transformer_pipeline_model_parallel_size ........ 1
|
| 897 |
+
trust_remote_code ............................... False
|
| 898 |
+
untie_embeddings_and_output_weights ............. True
|
| 899 |
+
use_checkpoint_args ............................. False
|
| 900 |
+
use_checkpoint_opt_param_scheduler .............. False
|
| 901 |
+
use_cpu_initialization .......................... True
|
| 902 |
+
use_dist_ckpt ................................... True
|
| 903 |
+
use_dist_ckpt_deprecated ........................ False
|
| 904 |
+
use_distributed_optimizer ....................... False
|
| 905 |
+
use_flash_attn .................................. False
|
| 906 |
+
use_fused_weighted_squared_relu ................. False
|
| 907 |
+
use_gloo_process_groups ......................... True
|
| 908 |
+
use_kitchen_attention ........................... False
|
| 909 |
+
use_layer_wise_distributed_optimizer ............ False
|
| 910 |
+
use_legacy_models ............................... False
|
| 911 |
+
use_legacy_static_engine ........................ False
|
| 912 |
+
use_mamba_mem_eff_path .......................... True
|
| 913 |
+
use_megatron_fsdp ............................... False
|
| 914 |
+
use_mp_args_from_checkpoint_args ................ True
|
| 915 |
+
use_mup ......................................... False
|
| 916 |
+
use_one_sent_docs ............................... False
|
| 917 |
+
use_orthogonal_updates .......................... True
|
| 918 |
+
use_persistent_ckpt_worker ...................... False
|
| 919 |
+
use_precision_aware_optimizer ................... False
|
| 920 |
+
use_pytorch_profiler ............................ False
|
| 921 |
+
use_ring_exchange_p2p ........................... False
|
| 922 |
+
use_rope_scaling ................................ False
|
| 923 |
+
use_rotary_position_embeddings .................. False
|
| 924 |
+
use_sharp ....................................... False
|
| 925 |
+
use_te_activation_func .......................... False
|
| 926 |
+
use_tokenizer_model_from_checkpoint_args ........ True
|
| 927 |
+
use_torch_fsdp2 ................................. False
|
| 928 |
+
use_torch_optimizer_for_cpu_offload ............. False
|
| 929 |
+
use_tp_pp_dp_mapping ............................ False
|
| 930 |
+
v_head_dim ...................................... 128
|
| 931 |
+
valid_data_path ................................. None
|
| 932 |
+
variable_seq_lengths ............................ False
|
| 933 |
+
virtual_pipeline_model_parallel_size ............ None
|
| 934 |
+
vision_backbone_type ............................ vit
|
| 935 |
+
vision_pretraining .............................. False
|
| 936 |
+
vision_pretraining_type ......................... classify
|
| 937 |
+
vocab_extra_ids ................................. 0
|
| 938 |
+
vocab_file ...................................... None
|
| 939 |
+
vocab_size ...................................... None
|
| 940 |
+
wandb_entity .................................... None
|
| 941 |
+
wandb_exp_name .................................. None
|
| 942 |
+
wandb_project ................................... None
|
| 943 |
+
wandb_save_dir .................................. None
|
| 944 |
+
weight_decay .................................... 0.01
|
| 945 |
+
weight_decay_all_param .......................... False
|
| 946 |
+
weight_decay_incr_style ......................... constant
|
| 947 |
+
wgrad_deferral_limit ............................ 0
|
| 948 |
+
window_attn_skip_freq ........................... None
|
| 949 |
+
window_size ..................................... None
|
| 950 |
+
world_size ...................................... 1
|
| 951 |
+
xpr ............................................. False
|
| 952 |
+
xr2 ............................................. False
|
| 953 |
+
xr2glu .......................................... False
|
| 954 |
+
xssglu .......................................... False
|
| 955 |
+
yaml_cfg ........................................ None
|
| 956 |
+
-------------------- end of arguments ---------------------
|
| 957 |
+
INFO:megatron.core.num_microbatches_calculator:setting number of microbatches to constant 1
|
| 958 |
+
building GPT model ...
|
| 959 |
+
loading checkpoint from /iopsstor/scratch/cscs/rsinghal/1pp-hf-tmp/1pp-0.5b-asst-sft/torch at iteration 582
|
| 960 |
+
checkpoint version 3.0
|
| 961 |
+
successfully loaded checkpoint from /iopsstor/scratch/cscs/rsinghal/1pp-hf-tmp/1pp-0.5b-asst-sft/torch [ t 1/1, p 1/1 ] at iteration 582
|
| 962 |
+
`torch_dtype` is deprecated! Use `dtype` instead!
|
| 963 |
+
received embeddings
|
| 964 |
+
received transformer layer 0
|
| 965 |
+
received transformer layer 1
|
| 966 |
+
received transformer layer 2
|
| 967 |
+
received transformer layer 3
|
| 968 |
+
received transformer layer 4
|
| 969 |
+
received transformer layer 5
|
| 970 |
+
received transformer layer 6
|
| 971 |
+
received transformer layer 7
|
| 972 |
+
received transformer layer 8
|
| 973 |
+
received transformer layer 9
|
| 974 |
+
received transformer layer 10
|
| 975 |
+
received transformer layer 11
|
| 976 |
+
received transformer layer 12
|
| 977 |
+
received transformer layer 13
|
| 978 |
+
received transformer layer 14
|
| 979 |
+
received transformer layer 15
|
| 980 |
+
received transformer layer 16
|
| 981 |
+
received transformer layer 17
|
| 982 |
+
received transformer layer 18
|
| 983 |
+
received transformer layer 19
|
| 984 |
+
received transformer layer 20
|
| 985 |
+
received transformer layer 21
|
| 986 |
+
received transformer layer 22
|
| 987 |
+
received transformer layer 23
|
| 988 |
+
received final norm
|
| 989 |
+
received output layer
|
| 990 |
+
Building LlamaForCausalLM from converted weights …
|
| 991 |
+
Saving model (safetensors) to /capstor/store/cscs/swissai/infra01/users/rsinghal/1pp-training/hf/models/1pp-0.5b-asst-sft
|
| 992 |
+
> memory usage: 'loader', rank 0 / 1, mem 3.6/856.2 gb.
|
| 993 |
+
sending embeddings
|
| 994 |
+
sending transformer layer 0
|
| 995 |
+
sending transformer layer 1
|
| 996 |
+
sending transformer layer 2
|
| 997 |
+
sending transformer layer 3
|
| 998 |
+
sending transformer layer 4
|
| 999 |
+
sending transformer layer 5
|
| 1000 |
+
sending transformer layer 6
|
| 1001 |
+
sending transformer layer 7
|
| 1002 |
+
sending transformer layer 8
|
| 1003 |
+
sending transformer layer 9
|
| 1004 |
+
sending transformer layer 10
|
| 1005 |
+
sending transformer layer 11
|
| 1006 |
+
sending transformer layer 12
|
| 1007 |
+
sending transformer layer 13
|
| 1008 |
+
sending transformer layer 14
|
| 1009 |
+
sending transformer layer 15
|
| 1010 |
+
sending transformer layer 16
|
| 1011 |
+
sending transformer layer 17
|
| 1012 |
+
sending transformer layer 18
|
| 1013 |
+
sending transformer layer 19
|
| 1014 |
+
sending transformer layer 20
|
| 1015 |
+
sending transformer layer 21
|
| 1016 |
+
sending transformer layer 22
|
| 1017 |
+
sending transformer layer 23
|
| 1018 |
+
sending final norm
|
| 1019 |
+
sending output layer
|
| 1020 |
+
Waiting for saver to complete...
|
generation_config.json
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 1,
|
| 3 |
+
"do_sample": true,
|
| 4 |
+
"eos_token_id": [
|
| 5 |
+
2,
|
| 6 |
+
0
|
| 7 |
+
],
|
| 8 |
+
"pad_token_id": 49152,
|
| 9 |
+
"temperature": 0.6,
|
| 10 |
+
"top_p": 0.9,
|
| 11 |
+
"transformers_version": "4.57.6"
|
| 12 |
+
}
|
merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ea95f8461efe78415f3283bbe394908de10a9a6d3c68ff73cec2d79045899680
|
| 3 |
+
size 1160916120
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"additional_special_tokens": [
|
| 3 |
+
"<|im_start|>",
|
| 4 |
+
"<|im_end|>"
|
| 5 |
+
],
|
| 6 |
+
"bos_token": {
|
| 7 |
+
"content": "<|im_start|>",
|
| 8 |
+
"lstrip": false,
|
| 9 |
+
"normalized": false,
|
| 10 |
+
"rstrip": false,
|
| 11 |
+
"single_word": false
|
| 12 |
+
},
|
| 13 |
+
"eos_token": {
|
| 14 |
+
"content": "<|endoftext|>",
|
| 15 |
+
"lstrip": false,
|
| 16 |
+
"normalized": false,
|
| 17 |
+
"rstrip": false,
|
| 18 |
+
"single_word": false
|
| 19 |
+
},
|
| 20 |
+
"pad_token": {
|
| 21 |
+
"content": "<|pad|>",
|
| 22 |
+
"lstrip": false,
|
| 23 |
+
"normalized": false,
|
| 24 |
+
"rstrip": false,
|
| 25 |
+
"single_word": false
|
| 26 |
+
},
|
| 27 |
+
"unk_token": {
|
| 28 |
+
"content": "<|endoftext|>",
|
| 29 |
+
"lstrip": false,
|
| 30 |
+
"normalized": false,
|
| 31 |
+
"rstrip": false,
|
| 32 |
+
"single_word": false
|
| 33 |
+
}
|
| 34 |
+
}
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,162 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"added_tokens_decoder": {
|
| 4 |
+
"0": {
|
| 5 |
+
"content": "<|endoftext|>",
|
| 6 |
+
"lstrip": false,
|
| 7 |
+
"normalized": false,
|
| 8 |
+
"rstrip": false,
|
| 9 |
+
"single_word": false,
|
| 10 |
+
"special": true
|
| 11 |
+
},
|
| 12 |
+
"1": {
|
| 13 |
+
"content": "<|im_start|>",
|
| 14 |
+
"lstrip": false,
|
| 15 |
+
"normalized": false,
|
| 16 |
+
"rstrip": false,
|
| 17 |
+
"single_word": false,
|
| 18 |
+
"special": true
|
| 19 |
+
},
|
| 20 |
+
"2": {
|
| 21 |
+
"content": "<|im_end|>",
|
| 22 |
+
"lstrip": false,
|
| 23 |
+
"normalized": false,
|
| 24 |
+
"rstrip": false,
|
| 25 |
+
"single_word": false,
|
| 26 |
+
"special": true
|
| 27 |
+
},
|
| 28 |
+
"3": {
|
| 29 |
+
"content": "<repo_name>",
|
| 30 |
+
"lstrip": false,
|
| 31 |
+
"normalized": false,
|
| 32 |
+
"rstrip": false,
|
| 33 |
+
"single_word": false,
|
| 34 |
+
"special": true
|
| 35 |
+
},
|
| 36 |
+
"4": {
|
| 37 |
+
"content": "<reponame>",
|
| 38 |
+
"lstrip": false,
|
| 39 |
+
"normalized": false,
|
| 40 |
+
"rstrip": false,
|
| 41 |
+
"single_word": false,
|
| 42 |
+
"special": true
|
| 43 |
+
},
|
| 44 |
+
"5": {
|
| 45 |
+
"content": "<file_sep>",
|
| 46 |
+
"lstrip": false,
|
| 47 |
+
"normalized": false,
|
| 48 |
+
"rstrip": false,
|
| 49 |
+
"single_word": false,
|
| 50 |
+
"special": true
|
| 51 |
+
},
|
| 52 |
+
"6": {
|
| 53 |
+
"content": "<filename>",
|
| 54 |
+
"lstrip": false,
|
| 55 |
+
"normalized": false,
|
| 56 |
+
"rstrip": false,
|
| 57 |
+
"single_word": false,
|
| 58 |
+
"special": true
|
| 59 |
+
},
|
| 60 |
+
"7": {
|
| 61 |
+
"content": "<gh_stars>",
|
| 62 |
+
"lstrip": false,
|
| 63 |
+
"normalized": false,
|
| 64 |
+
"rstrip": false,
|
| 65 |
+
"single_word": false,
|
| 66 |
+
"special": true
|
| 67 |
+
},
|
| 68 |
+
"8": {
|
| 69 |
+
"content": "<issue_start>",
|
| 70 |
+
"lstrip": false,
|
| 71 |
+
"normalized": false,
|
| 72 |
+
"rstrip": false,
|
| 73 |
+
"single_word": false,
|
| 74 |
+
"special": true
|
| 75 |
+
},
|
| 76 |
+
"9": {
|
| 77 |
+
"content": "<issue_comment>",
|
| 78 |
+
"lstrip": false,
|
| 79 |
+
"normalized": false,
|
| 80 |
+
"rstrip": false,
|
| 81 |
+
"single_word": false,
|
| 82 |
+
"special": true
|
| 83 |
+
},
|
| 84 |
+
"10": {
|
| 85 |
+
"content": "<issue_closed>",
|
| 86 |
+
"lstrip": false,
|
| 87 |
+
"normalized": false,
|
| 88 |
+
"rstrip": false,
|
| 89 |
+
"single_word": false,
|
| 90 |
+
"special": true
|
| 91 |
+
},
|
| 92 |
+
"11": {
|
| 93 |
+
"content": "<jupyter_start>",
|
| 94 |
+
"lstrip": false,
|
| 95 |
+
"normalized": false,
|
| 96 |
+
"rstrip": false,
|
| 97 |
+
"single_word": false,
|
| 98 |
+
"special": true
|
| 99 |
+
},
|
| 100 |
+
"12": {
|
| 101 |
+
"content": "<jupyter_text>",
|
| 102 |
+
"lstrip": false,
|
| 103 |
+
"normalized": false,
|
| 104 |
+
"rstrip": false,
|
| 105 |
+
"single_word": false,
|
| 106 |
+
"special": true
|
| 107 |
+
},
|
| 108 |
+
"13": {
|
| 109 |
+
"content": "<jupyter_code>",
|
| 110 |
+
"lstrip": false,
|
| 111 |
+
"normalized": false,
|
| 112 |
+
"rstrip": false,
|
| 113 |
+
"single_word": false,
|
| 114 |
+
"special": true
|
| 115 |
+
},
|
| 116 |
+
"14": {
|
| 117 |
+
"content": "<jupyter_output>",
|
| 118 |
+
"lstrip": false,
|
| 119 |
+
"normalized": false,
|
| 120 |
+
"rstrip": false,
|
| 121 |
+
"single_word": false,
|
| 122 |
+
"special": true
|
| 123 |
+
},
|
| 124 |
+
"15": {
|
| 125 |
+
"content": "<jupyter_script>",
|
| 126 |
+
"lstrip": false,
|
| 127 |
+
"normalized": false,
|
| 128 |
+
"rstrip": false,
|
| 129 |
+
"single_word": false,
|
| 130 |
+
"special": true
|
| 131 |
+
},
|
| 132 |
+
"16": {
|
| 133 |
+
"content": "<empty_output>",
|
| 134 |
+
"lstrip": false,
|
| 135 |
+
"normalized": false,
|
| 136 |
+
"rstrip": false,
|
| 137 |
+
"single_word": false,
|
| 138 |
+
"special": true
|
| 139 |
+
},
|
| 140 |
+
"49152": {
|
| 141 |
+
"content": "<|pad|>",
|
| 142 |
+
"lstrip": false,
|
| 143 |
+
"normalized": false,
|
| 144 |
+
"rstrip": false,
|
| 145 |
+
"single_word": false,
|
| 146 |
+
"special": true
|
| 147 |
+
}
|
| 148 |
+
},
|
| 149 |
+
"additional_special_tokens": [
|
| 150 |
+
"<|im_start|>",
|
| 151 |
+
"<|im_end|>"
|
| 152 |
+
],
|
| 153 |
+
"bos_token": "<|im_start|>",
|
| 154 |
+
"clean_up_tokenization_spaces": false,
|
| 155 |
+
"eos_token": "<|endoftext|>",
|
| 156 |
+
"extra_special_tokens": {},
|
| 157 |
+
"model_max_length": 8192,
|
| 158 |
+
"pad_token": "<|pad|>",
|
| 159 |
+
"tokenizer_class": "GPT2Tokenizer",
|
| 160 |
+
"unk_token": "<|endoftext|>",
|
| 161 |
+
"vocab_size": 49152
|
| 162 |
+
}
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|