updated-policy / logger /grpo_finetune.log
srinjoyd's picture
add logs
1eed1a8
Raw
History Blame Contribute Delete
131 kB
===== Job started at 2026-04-26 03:05:03 =====
Downloading pygments (1.2MiB)
Downloading nvidia-cufile (1.2MiB)
Downloading nvidia-cufft (204.2MiB)
Downloading aiohttp (1.7MiB)
Downloading pandas (10.4MiB)
Downloading numpy (15.9MiB)
Downloading nvidia-cusparselt-cu13 (162.0MiB)
Downloading networkx (2.0MiB)
Downloading nvidia-cusparse (139.2MiB)
Downloading hf-xet (4.0MiB)
Downloading pyarrow (46.6MiB)
Downloading nvidia-cuda-nvrtc (86.0MiB)
Downloading nvidia-nccl-cu13 (187.4MiB)
Downloading nvidia-nvshmem-cu13 (57.6MiB)
Downloading cuda-bindings (6.0MiB)
Downloading setuptools (1.0MiB)
Downloading nvidia-nvjitlink (38.8MiB)
Downloading nvidia-curand (56.8MiB)
Downloading tokenizers (3.1MiB)
Downloading sympy (6.0MiB)
Downloading nvidia-cudnn-cu13 (349.1MiB)
Downloading transformers (9.9MiB)
Downloading nvidia-cublas (403.5MiB)
Downloading nvidia-cuda-cupti (10.2MiB)
Downloading nvidia-cuda-runtime (2.1MiB)
Downloading torch (506.1MiB)
Downloading nvidia-cusolver (191.6MiB)
Downloading triton (179.5MiB)
Downloaded nvidia-cufile
Downloaded aiohttp
Downloaded nvidia-cuda-runtime
Downloaded pygments
Downloaded tokenizers
Downloaded setuptools
Downloaded hf-xet
Downloaded networkx
Downloaded cuda-bindings
Downloaded sympy
Downloaded nvidia-cuda-cupti
Downloaded numpy
Downloaded transformers
Downloaded pandas
Downloaded nvidia-nvjitlink
Downloaded pyarrow
Downloaded nvidia-curand
Downloaded nvidia-nvshmem-cu13
Downloaded nvidia-cuda-nvrtc
Downloaded nvidia-cusparse
Downloaded nvidia-cusparselt-cu13
Downloaded triton
Downloaded nvidia-nccl-cu13
Downloaded nvidia-cusolver
Downloaded nvidia-cufft
Downloaded nvidia-cudnn-cu13
Downloaded nvidia-cublas
Downloaded torch
Installed 76 packages in 279ms
03:05:26 [INFO] Using temp --output-dir: /tmp/sre-grpo-9zduekwb
03:05:36 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/main/README.md "HTTP/1.1 404 Not Found"
03:05:36 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/srinjoyd/sre-data "HTTP/1.1 200 OK"
03:05:36 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre-data.py "HTTP/1.1 404 Not Found"
03:05:37 [INFO] HTTP Request: HEAD https://s3.amazonaws.com/datasets.huggingface.co/datasets/datasets/srinjoyd/sre-data/srinjoyd/sre-data.py "HTTP/1.1 404 Not Found"
03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/README.md "HTTP/1.1 404 Not Found"
03:05:37 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/srinjoyd/sre-data/revision/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8 "HTTP/1.1 200 OK"
03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/.huggingface.yaml "HTTP/1.1 404 Not Found"
03:05:37 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/srinjoyd/sre-data/tree/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8?recursive=false&expand=false "HTTP/1.1 200 OK"
03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/dataset_infos.json "HTTP/1.1 404 Not Found"
03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre_raw_trajectories.jsonl "HTTP/1.1 307 Temporary Redirect"
03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/datasets/srinjoyd/sre-data/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre_raw_trajectories.jsonl "HTTP/1.1 200 OK"
03:05:37 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/srinjoyd/sre-data/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre_raw_trajectories.jsonl "HTTP/1.1 200 OK"
sre_raw_trajectories.jsonl: 0%| | 0.00/8.57M [00:00<?, ?B/s]
sre_raw_trajectories.jsonl: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 8.57M/8.57M [00:00<00:00, 44.3MB/s]
Generating train split: 0 examples [00:00, ? examples/s]
Generating train split: 64 examples [00:00, 535.47 examples/s]
03:05:37 [INFO] Loaded 95 prompts from hf://datasets/srinjoyd/sre-data (sre_raw_trajectories.jsonl)
03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
config.json: 0%| | 0.00/1.37k [00:00<?, ?B/s]
config.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1.37k/1.37k [00:00<00:00, 12.5MB/s]
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect"
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/tokenizer_config.json "HTTP/1.1 200 OK"
03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/tokenizer_config.json "HTTP/1.1 200 OK"
tokenizer_config.json: 0%| | 0.00/844 [00:00<?, ?B/s]
tokenizer_config.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 844/844 [00:00<00:00, 7.78MB/s]
03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found"
03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK"
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/vocab.json "HTTP/1.1 404 Not Found"
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/merges.txt "HTTP/1.1 404 Not Found"
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/tokenizer.json "HTTP/1.1 302 Found"
03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged/xet-read-token/8f73acd8434aacfb77d224b6c46773bca9279705 "HTTP/1.1 200 OK"
tokenizer.json: 0%| | 0.00/11.4M [00:00<?, ?B/s]
tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4M/11.4M [00:00<00:00, 18.6MB/s]
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/added_tokens.json "HTTP/1.1 404 Not Found"
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/special_tokens_map.json "HTTP/1.1 404 Not Found"
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/chat_template.jinja "HTTP/1.1 307 Temporary Redirect"
03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/chat_template.jinja "HTTP/1.1 200 OK"
03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/chat_template.jinja "HTTP/1.1 200 OK"
chat_template.jinja: 0%| | 0.00/2.51k [00:00<?, ?B/s]
chat_template.jinja: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 2.51k/2.51k [00:00<00:00, 22.8MB/s]
03:05:39 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged "HTTP/1.1 200 OK"
03:05:39 [WARNING] Dropping unsupported GRPOConfig fields for this trl build: ['max_prompt_length']
03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/adapter_config.json "HTTP/1.1 404 Not Found"
03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:05:40 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/model.safetensors "HTTP/1.1 302 Found"
model.safetensors: 0%| | 0.00/15.2G [00:00<?, ?B/s]
model.safetensors: 0%| | 0.00/15.2G [00:01<?, ?B/s]
model.safetensors: 0%| | 0.00/15.2G [00:02<?, ?B/s]
model.safetensors: 0%| | 67.1M/15.2G [00:03<04:30, 56.0MB/s]
model.safetensors: 3%|β–Ž | 402M/15.2G [00:04<01:11, 208MB/s] 
model.safetensors: 6%|β–Œ | 872M/15.2G [00:05<00:48, 296MB/s]
model.safetensors: 9%|β–‰ | 1.43G/15.2G [00:06<00:35, 390MB/s]
model.safetensors: 12%|β–ˆβ– | 1.89G/15.2G [00:07<00:31, 417MB/s]
model.safetensors: 16%|β–ˆβ–‹ | 2.50G/15.2G [00:08<00:28, 449MB/s]
model.safetensors: 22%|β–ˆβ–ˆβ– | 3.38G/15.2G [00:09<00:20, 583MB/s]
model.safetensors: 33%|β–ˆβ–ˆβ–ˆβ–Ž | 4.98G/15.2G [00:10<00:12, 841MB/s]
model.safetensors: 42%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 6.38G/15.2G [00:12<00:09, 946MB/s]
model.safetensors: 52%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 7.90G/15.2G [00:13<00:07, 948MB/s]
model.safetensors: 61%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 9.25G/15.2G [00:14<00:05, 1.00GB/s]
model.safetensors: 68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 10.4G/15.2G [00:16<00:05, 931MB/s] 
model.safetensors: 79%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 12.1G/15.2G [00:17<00:02, 1.08GB/s]
model.safetensors: 89%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 13.5G/15.2G [00:18<00:01, 1.16GB/s]
model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 15.2G/15.2G [00:20<00:00, 1.13GB/s]
model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 15.2G/15.2G [00:20<00:00, 761MB/s]
Loading weights: 0%| | 0/339 [00:00<?, ?it/s]
Loading weights: 0%| | 1/339 [00:01<07:16, 1.29s/it]
Loading weights: 18%|β–ˆβ–Š | 60/339 [00:02<00:09, 28.99it/s]
Loading weights: 30%|β–ˆβ–ˆβ–‰ | 101/339 [00:03<00:07, 33.00it/s]
Loading weights: 43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 146/339 [00:04<00:05, 37.03it/s]
Loading weights: 54%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 184/339 [00:05<00:04, 36.57it/s]
Loading weights: 68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 230/339 [00:06<00:02, 37.94it/s]
Loading weights: 79%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 269/339 [00:07<00:01, 37.87it/s]
Loading weights: 93%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 314/339 [00:08<00:00, 39.23it/s]
Loading weights: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 339/339 [00:09<00:00, 36.54it/s]
03:06:09 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/generation_config.json "HTTP/1.1 307 Temporary Redirect"
03:06:09 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/generation_config.json "HTTP/1.1 200 OK"
03:06:09 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/generation_config.json "HTTP/1.1 200 OK"
generation_config.json: 0%| | 0.00/242 [00:00<?, ?B/s]
generation_config.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 242/242 [00:00<00:00, 2.18MB/s]
03:06:11 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 200 OK"
03:06:11 [INFO] Starting GRPO: model=srinjoyd/qwen2.5-7b-sre-merged, prompts=95, k=4, beta=0, lr=2e-06
[transformers] The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': None, 'pad_token_id': 151643}.
0%| | 0/94 [00:00<?, ?it/s]03:06:14 [INFO] [reward] call=1 n=8 mean=0.129 std=0.181 min=0.029 max=0.599 | parts mean: fmt=0.20 at=0.05 svc=0.05 par=0.00 task=0.01 len=-0.01
1%| | 1/94 [00:04<07:18, 4.72s/it]

{'loss': '-0.03945', 'grad_norm': '7.048', 'learning_rate': '2e-06', 'num_tokens': '5272', 'completions/mean_length': '12.5', 'completions/min_length': '11', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '12.5', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.1295', 'rewards/reward_fn/std': '0.1936', 'reward': '0.1295', 'reward_std': '0.1936', 'frac_reward_zero_std': '0.5', 'entropy': '0.0801', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '4.603', 'epoch': '0.02128'}
1%| | 1/94 [00:04<07:18, 4.72s/it]03:06:17 [INFO] [reward] call=2 n=8 mean=0.634 std=0.001 min=0.633 max=0.635 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.03
{'loss': '-0.06348', 'grad_norm': '1.36', 'learning_rate': '1.979e-06', 'num_tokens': '7094', 'completions/mean_length': '15.75', 'completions/min_length': '15', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.6336', 'rewards/reward_fn/std': '0.0005657', 'reward': '0.6336', 'reward_std': '0.0005657', 'frac_reward_zero_std': '0.5', 'entropy': '0.1546', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.686', 'epoch': '0.04255'}
2%|▏ | 2/94 [00:07<05:24, 3.53s/it]

2%|▏ | 2/94 [00:07<05:24, 3.53s/it]03:06:21 [INFO] [reward] call=3 n=8 mean=0.661 std=0.028 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
3%|β–Ž | 3/94 [00:11<05:33, 3.67s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '1.957e-06', 'num_tokens': '1.064e+04', 'completions/mean_length': '16.38', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6609', 'rewards/reward_fn/std': '0.0294', 'reward': '0.6609', 'reward_std': '0.0294', 'frac_reward_zero_std': '1', 'entropy': '0.06277', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.773', 'epoch': '0.06383'}
3%|β–Ž | 3/94 [00:11<05:33, 3.67s/it]03:06:25 [INFO] [reward] call=4 n=8 mean=0.663 std=0.026 min=0.633 max=0.689 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
4%|▍ | 4/94 [00:14<05:23, 3.59s/it]

{'loss': '-0.2514', 'grad_norm': '2.367', 'learning_rate': '1.936e-06', 'num_tokens': '1.421e+04', 'completions/mean_length': '20', 'completions/min_length': '15', 'completions/max_length': '42', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '42', 'rewards/reward_fn/mean': '0.6629', 'rewards/reward_fn/std': '0.02805', 'reward': '0.6629', 'reward_std': '0.02805', 'frac_reward_zero_std': '0', 'entropy': '0.2134', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.407', 'epoch': '0.08511'}
4%|▍ | 4/94 [00:14<05:23, 3.59s/it]03:06:27 [INFO] [reward] call=5 n=8 mean=0.635 std=0.005 min=0.633 max=0.650 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
5%|β–Œ | 5/94 [00:16<04:24, 2.97s/it]

{'loss': '-0.09344', 'grad_norm': '1.805', 'learning_rate': '1.915e-06', 'num_tokens': '1.605e+04', 'completions/mean_length': '19.5', 'completions/min_length': '15', 'completions/max_length': '34', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '34', 'rewards/reward_fn/mean': '0.6353', 'rewards/reward_fn/std': '0.005775', 'reward': '0.6353', 'reward_std': '0.005775', 'frac_reward_zero_std': '0.5', 'entropy': '0.2227', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.862', 'epoch': '0.1064'}
5%|β–Œ | 5/94 [00:16<04:24, 2.97s/it]03:06:31 [INFO] [reward] call=6 n=8 mean=0.616 std=0.018 min=0.598 max=0.636 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04
6%|β–‹ | 6/94 [00:20<04:39, 3.18s/it]

{'loss': '-0.276', 'grad_norm': '1.836', 'learning_rate': '1.894e-06', 'num_tokens': '1.964e+04', 'completions/mean_length': '19.38', 'completions/min_length': '15', 'completions/max_length': '46', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '46', 'rewards/reward_fn/mean': '0.6161', 'rewards/reward_fn/std': '0.01908', 'reward': '0.6161', 'reward_std': '0.01908', 'frac_reward_zero_std': '0.5', 'entropy': '0.2511', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.534', 'epoch': '0.1277'}
6%|β–‹ | 6/94 [00:20<04:39, 3.18s/it]03:06:33 [INFO] [reward] call=7 n=8 mean=0.638 std=0.007 min=0.633 max=0.651 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
7%|β–‹ | 7/94 [00:22<04:08, 2.86s/it]

{'loss': '-0.2179', 'grad_norm': '1.842', 'learning_rate': '1.872e-06', 'num_tokens': '2.149e+04', 'completions/mean_length': '19.75', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6375', 'rewards/reward_fn/std': '0.007678', 'reward': '0.6375', 'reward_std': '0.007678', 'frac_reward_zero_std': '0.5', 'entropy': '0.2716', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.174', 'epoch': '0.1489'}
7%|β–‹ | 7/94 [00:22<04:08, 2.86s/it]03:06:36 [INFO] [reward] call=8 n=8 mean=0.330 std=0.476 min=-0.777 max=0.701 | parts mean: fmt=0.18 at=0.16 svc=0.05 par=0.04 task=0.03 len=0.04
9%|β–Š | 8/94 [00:25<04:12, 2.93s/it]

{'loss': '0.1133', 'grad_norm': '4.491', 'learning_rate': '1.851e-06', 'num_tokens': '2.506e+04', 'completions/mean_length': '16.5', 'completions/min_length': '11', 'completions/max_length': '30', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '30', 'rewards/reward_fn/mean': '0.33', 'rewards/reward_fn/std': '0.5085', 'reward': '0.33', 'reward_std': '0.5085', 'frac_reward_zero_std': '0', 'entropy': '0.4821', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.033', 'epoch': '0.1702'}
9%|β–Š | 8/94 [00:25<04:12, 2.93s/it]03:06:39 [INFO] [reward] call=9 n=8 mean=0.534 std=0.245 min=0.111 max=0.690 | parts mean: fmt=0.20 at=0.19 svc=0.05 par=0.09 task=0.03 len=0.02
10%|β–‰ | 9/94 [00:28<04:00, 2.83s/it]

{'loss': '-0.04131', 'grad_norm': '3.706', 'learning_rate': '1.83e-06', 'num_tokens': '3.025e+04', 'completions/mean_length': '15.12', 'completions/min_length': '10', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.12', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.5343', 'rewards/reward_fn/std': '0.2624', 'reward': '0.5343', 'reward_std': '0.2624', 'frac_reward_zero_std': '0', 'entropy': '0.1164', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.551', 'epoch': '0.1915'}
10%|β–‰ | 9/94 [00:28<04:00, 2.83s/it]03:06:42 [INFO] [reward] call=10 n=8 mean=0.397 std=0.269 min=0.029 max=0.688 | parts mean: fmt=0.20 at=0.14 svc=0.07 par=0.05 task=0.02 len=0.03
11%|β–ˆ | 10/94 [00:31<04:11, 2.99s/it]

{'loss': '0.06288', 'grad_norm': '3.008', 'learning_rate': '1.809e-06', 'num_tokens': '3.543e+04', 'completions/mean_length': '20.25', 'completions/min_length': '11', 'completions/max_length': '40', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.25', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '40', 'rewards/reward_fn/mean': '0.397', 'rewards/reward_fn/std': '0.2871', 'reward': '0.397', 'reward_std': '0.2871', 'frac_reward_zero_std': '0', 'entropy': '0.3148', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.291', 'epoch': '0.2128'}
11%|β–ˆ | 10/94 [00:31<04:11, 2.99s/it]03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:06:43 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:06:43 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:06:43 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:06:43 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:06:43 [INFO] HTTP Request: GET https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/xet-write-token/main "HTTP/1.1 200 OK"
03:06:45 [INFO] [reward] call=11 n=8 mean=0.638 std=0.034 min=0.598 max=0.699 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.08 task=0.04 len=0.01
03:06:46 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
12%|β–ˆβ– | 11/94 [00:34<04:13, 3.05s/it]
{'loss': '-0.1043', 'grad_norm': '5.863', 'learning_rate': '1.787e-06', 'num_tokens': '3.893e+04', 'completions/mean_length': '14.62', 'completions/min_length': '10', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '14.62', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6384', 'rewards/reward_fn/std': '0.03669', 'reward': '0.6384', 'reward_std': '0.03669', 'frac_reward_zero_std': '0', 'entropy': '0.1577', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.613', 'epoch': '0.234'}

12%|β–ˆβ– | 11/94 [00:34<04:13, 3.05s/it]03:06:47 [INFO] [reward] call=12 n=8 mean=0.633 std=0.001 min=0.630 max=0.633 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.03
13%|β–ˆβ–Ž | 12/94 [00:36<03:32, 2.59s/it]

{'loss': '0.07766', 'grad_norm': '1.3', 'learning_rate': '1.766e-06', 'num_tokens': '4.074e+04', 'completions/mean_length': '15.88', 'completions/min_length': '15', 'completions/max_length': '22', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '22', 'rewards/reward_fn/mean': '0.633', 'rewards/reward_fn/std': '0.001096', 'reward': '0.633', 'reward_std': '0.001096', 'frac_reward_zero_std': '0.5', 'entropy': '0.2152', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.541', 'epoch': '0.2553'}
13%|β–ˆβ–Ž | 12/94 [00:36<03:32, 2.59s/it]03:06:49 [INFO] [reward] call=13 n=8 mean=0.264 std=0.271 min=0.020 max=0.611 | parts mean: fmt=0.20 at=0.11 svc=0.05 par=0.04 task=0.01 len=-0.03
14%|β–ˆβ– | 13/94 [00:38<03:27, 2.56s/it]

{'loss': '0.03801', 'grad_norm': '8.365', 'learning_rate': '1.745e-06', 'num_tokens': '4.595e+04', 'completions/mean_length': '10.62', 'completions/min_length': '10', 'completions/max_length': '11', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '10.62', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '11', 'rewards/reward_fn/mean': '0.2636', 'rewards/reward_fn/std': '0.2897', 'reward': '0.2636', 'reward_std': '0.2897', 'frac_reward_zero_std': '0', 'entropy': '0.09156', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.426', 'epoch': '0.2766'}
14%|β–ˆβ– | 13/94 [00:38<03:27, 2.56s/it]03:06:52 [INFO] [reward] call=14 n=8 mean=0.642 std=0.032 min=0.598 max=0.695 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.03
15%|β–ˆβ– | 14/94 [00:41<03:39, 2.75s/it]

{'loss': '-0.2787', 'grad_norm': '5.452', 'learning_rate': '1.723e-06', 'num_tokens': '4.955e+04', 'completions/mean_length': '18.38', 'completions/min_length': '10', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.38', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6422', 'rewards/reward_fn/std': '0.0339', 'reward': '0.6422', 'reward_std': '0.0339', 'frac_reward_zero_std': '0', 'entropy': '0.3664', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.117', 'epoch': '0.2979'}
15%|β–ˆβ– | 14/94 [00:41<03:39, 2.75s/it]03:06:55 [INFO] [reward] call=15 n=8 mean=0.548 std=0.200 min=0.020 max=0.634 | parts mean: fmt=0.20 at=0.22 svc=0.06 par=0.03 task=0.04 len=0.02
16%|β–ˆβ–Œ | 15/94 [00:44<03:45, 2.86s/it]

{'loss': '-0.1453', 'grad_norm': '3.979', 'learning_rate': '1.702e-06', 'num_tokens': '5.307e+04', 'completions/mean_length': '16.5', 'completions/min_length': '11', 'completions/max_length': '28', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '28', 'rewards/reward_fn/mean': '0.5483', 'rewards/reward_fn/std': '0.2137', 'reward': '0.5483', 'reward_std': '0.2137', 'frac_reward_zero_std': '0', 'entropy': '0.2527', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.047', 'epoch': '0.3191'}
16%|β–ˆβ–Œ | 15/94 [00:44<03:45, 2.86s/it]03:06:59 [INFO] [reward] call=16 n=8 mean=0.633 std=0.026 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.02
17%|β–ˆβ–‹ | 16/94 [00:48<03:51, 2.97s/it]
{'loss': '-0.1697', 'grad_norm': '4.519', 'learning_rate': '1.681e-06', 'num_tokens': '5.657e+04', 'completions/mean_length': '17.88', 'completions/min_length': '10', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.88', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6331', 'rewards/reward_fn/std': '0.02809', 'reward': '0.6331', 'reward_std': '0.02809', 'frac_reward_zero_std': '0', 'entropy': '0.817', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.173', 'epoch': '0.3404'}

17%|β–ˆβ–‹ | 16/94 [00:48<03:51, 2.97s/it]03:07:01 [INFO] [reward] call=17 n=8 mean=0.661 std=0.028 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
18%|β–ˆβ–Š | 17/94 [00:50<03:40, 2.87s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '1.66e-06', 'num_tokens': '6.006e+04', 'completions/mean_length': '16.5', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6609', 'rewards/reward_fn/std': '0.0294', 'reward': '0.6609', 'reward_std': '0.0294', 'frac_reward_zero_std': '1', 'entropy': '0.101', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.562', 'epoch': '0.3617'}
18%|β–ˆβ–Š | 17/94 [00:50<03:40, 2.87s/it]03:07:04 [INFO] [reward] call=18 n=8 mean=0.617 std=0.019 min=0.598 max=0.642 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04
19%|β–ˆβ–‰ | 18/94 [00:53<03:35, 2.83s/it]

{'loss': '-0.02915', 'grad_norm': '1.632', 'learning_rate': '1.638e-06', 'num_tokens': '6.363e+04', 'completions/mean_length': '16.75', 'completions/min_length': '15', 'completions/max_length': '20', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '20', 'rewards/reward_fn/mean': '0.6168', 'rewards/reward_fn/std': '0.01994', 'reward': '0.6168', 'reward_std': '0.01994', 'frac_reward_zero_std': '0.5', 'entropy': '0.2609', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.691', 'epoch': '0.383'}
19%|β–ˆβ–‰ | 18/94 [00:53<03:35, 2.83s/it]03:07:07 [INFO] [reward] call=19 n=8 mean=0.636 std=0.006 min=0.632 max=0.649 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
20%|β–ˆβ–ˆ | 19/94 [00:55<03:20, 2.67s/it]

{'loss': '-0.2383', 'grad_norm': '1.594', 'learning_rate': '1.617e-06', 'num_tokens': '6.549e+04', 'completions/mean_length': '21.12', 'completions/min_length': '15', 'completions/max_length': '49', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '49', 'rewards/reward_fn/mean': '0.6361', 'rewards/reward_fn/std': '0.005971', 'reward': '0.6361', 'reward_std': '0.005971', 'frac_reward_zero_std': '0.5', 'entropy': '0.5286', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.292', 'epoch': '0.4043'}
20%|β–ˆβ–ˆ | 19/94 [00:55<03:20, 2.67s/it]03:07:09 [INFO] [reward] call=20 n=8 mean=0.649 std=0.032 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04
{'loss': '0.08051', 'grad_norm': '4.462', 'learning_rate': '1.596e-06', 'num_tokens': '6.898e+04', 'completions/mean_length': '17.5', 'completions/min_length': '15', 'completions/max_length': '24', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '24', 'rewards/reward_fn/mean': '0.6493', 'rewards/reward_fn/std': '0.03433', 'reward': '0.6493', 'reward_std': '0.03433', 'frac_reward_zero_std': '0', 'entropy': '0.2607', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.756', 'epoch': '0.4255'}
21%|β–ˆβ–ˆβ– | 20/94 [00:58<03:20, 2.71s/it]

21%|β–ˆβ–ˆβ– | 20/94 [00:58<03:20, 2.71s/it]03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:07:10 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:07:11 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:07:11 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:07:11 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:07:13 [INFO] [reward] call=21 n=8 mean=0.495 std=0.224 min=0.106 max=0.637 | parts mean: fmt=0.20 at=0.19 svc=0.05 par=0.04 task=0.03 len=0.02
03:07:13 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
22%|β–ˆβ–ˆβ– | 21/94 [01:01<03:29, 2.87s/it]

{'loss': '-0.04748', 'grad_norm': '3.59', 'learning_rate': '1.574e-06', 'num_tokens': '7.249e+04', 'completions/mean_length': '15.12', 'completions/min_length': '10', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.12', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.4951', 'rewards/reward_fn/std': '0.2398', 'reward': '0.4951', 'reward_std': '0.2398', 'frac_reward_zero_std': '0', 'entropy': '0.1683', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.67', 'epoch': '0.4468'}
22%|β–ˆβ–ˆβ– | 21/94 [01:01<03:29, 2.87s/it]03:07:15 [INFO] [reward] call=22 n=8 mean=0.643 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04
23%|β–ˆβ–ˆβ–Ž | 22/94 [01:04<03:20, 2.79s/it]

{'loss': '-0.006488', 'grad_norm': '1.562', 'learning_rate': '1.553e-06', 'num_tokens': '7.775e+04', 'completions/mean_length': '16.75', 'completions/min_length': '16', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.75', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6434', 'rewards/reward_fn/std': '0.04772', 'reward': '0.6434', 'reward_std': '0.04772', 'frac_reward_zero_std': '0.5', 'entropy': '0.08012', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.534', 'epoch': '0.4681'}
23%|β–ˆβ–ˆβ–Ž | 22/94 [01:04<03:20, 2.79s/it]03:07:18 [INFO] [reward] call=23 n=8 mean=0.575 std=0.181 min=0.109 max=0.692 | parts mean: fmt=0.20 at=0.22 svc=0.06 par=0.06 task=0.04 len=0.02
{'loss': '-0.04835', 'grad_norm': '4.127', 'learning_rate': '1.532e-06', 'num_tokens': '8.303e+04', 'completions/mean_length': '15', 'completions/min_length': '10', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.575', 'rewards/reward_fn/std': '0.1932', 'reward': '0.575', 'reward_std': '0.1932', 'frac_reward_zero_std': '0', 'entropy': '0.1359', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.625', 'epoch': '0.4894'}
24%|β–ˆβ–ˆβ– | 23/94 [01:07<03:15, 2.76s/it]

24%|β–ˆβ–ˆβ– | 23/94 [01:07<03:15, 2.76s/it]03:07:21 [INFO] [reward] call=24 n=8 mean=0.673 std=0.037 min=0.611 max=0.712 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.09 task=0.04 len=0.04
26%|β–ˆβ–ˆβ–Œ | 24/94 [01:10<03:32, 3.03s/it]

{'loss': '-0.203', 'grad_norm': '2.814', 'learning_rate': '1.511e-06', 'num_tokens': '8.827e+04', 'completions/mean_length': '21', 'completions/min_length': '16', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6728', 'rewards/reward_fn/std': '0.03925', 'reward': '0.6728', 'reward_std': '0.03925', 'frac_reward_zero_std': '0', 'entropy': '0.1496', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.596', 'epoch': '0.5106'}
26%|β–ˆβ–ˆβ–Œ | 24/94 [01:10<03:32, 3.03s/it]03:07:25 [INFO] [reward] call=25 n=8 mean=0.661 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
27%|β–ˆβ–ˆβ–‹ | 25/94 [01:14<03:31, 3.07s/it]
{'loss': '-0.1814', 'grad_norm': '2.257', 'learning_rate': '1.489e-06', 'num_tokens': '9.183e+04', 'completions/mean_length': '18.5', 'completions/min_length': '15', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6614', 'rewards/reward_fn/std': '0.0286', 'reward': '0.6614', 'reward_std': '0.0286', 'frac_reward_zero_std': '0', 'entropy': '0.1536', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.107', 'epoch': '0.5319'}

27%|β–ˆβ–ˆβ–‹ | 25/94 [01:14<03:31, 3.07s/it]03:07:27 [INFO] [reward] call=26 n=8 mean=0.561 std=0.175 min=0.109 max=0.690 | parts mean: fmt=0.20 at=0.22 svc=0.07 par=0.04 task=0.04 len=0.02
28%|β–ˆβ–ˆβ–Š | 26/94 [01:16<03:17, 2.91s/it]

{'loss': '-0.1144', 'grad_norm': '7.925', 'learning_rate': '1.468e-06', 'num_tokens': '9.704e+04', 'completions/mean_length': '14.88', 'completions/min_length': '10', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '14.88', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.5614', 'rewards/reward_fn/std': '0.1872', 'reward': '0.5614', 'reward_std': '0.1872', 'frac_reward_zero_std': '0', 'entropy': '0.1353', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.467', 'epoch': '0.5532'}
28%|β–ˆβ–ˆβ–Š | 26/94 [01:16<03:17, 2.91s/it]03:07:30 [INFO] [reward] call=27 n=8 mean=0.661 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.03
29%|β–ˆβ–ˆβ–Š | 27/94 [01:19<03:08, 2.81s/it]

{'loss': '-0.00338', 'grad_norm': '1.348', 'learning_rate': '1.447e-06', 'num_tokens': '1.006e+05', 'completions/mean_length': '16.12', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6607', 'rewards/reward_fn/std': '0.02923', 'reward': '0.6607', 'reward_std': '0.02923', 'frac_reward_zero_std': '0.5', 'entropy': '0.08173', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.531', 'epoch': '0.5745'}
29%|β–ˆβ–ˆβ–Š | 27/94 [01:19<03:08, 2.81s/it]03:07:32 [INFO] [reward] call=28 n=8 mean=0.644 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04
30%|β–ˆβ–ˆβ–‰ | 28/94 [01:21<02:59, 2.72s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '1.426e-06', 'num_tokens': '1.058e+05', 'completions/mean_length': '16.5', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6436', 'rewards/reward_fn/std': '0.04789', 'reward': '0.6436', 'reward_std': '0.04789', 'frac_reward_zero_std': '1', 'entropy': '0.05464', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.458', 'epoch': '0.5957'}
30%|β–ˆβ–ˆβ–‰ | 28/94 [01:21<02:59, 2.72s/it]03:07:36 [INFO] [reward] call=29 n=8 mean=0.619 std=0.021 min=0.598 max=0.646 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04
31%|β–ˆβ–ˆβ–ˆ | 29/94 [01:25<03:11, 2.95s/it]

{'loss': '-0.1543', 'grad_norm': '1.539', 'learning_rate': '1.404e-06', 'num_tokens': '1.094e+05', 'completions/mean_length': '18.88', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6191', 'rewards/reward_fn/std': '0.02274', 'reward': '0.6191', 'reward_std': '0.02274', 'frac_reward_zero_std': '0.5', 'entropy': '0.3003', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.422', 'epoch': '0.617'}
31%|β–ˆβ–ˆβ–ˆ | 29/94 [01:25<03:11, 2.95s/it]03:07:38 [INFO] [reward] call=30 n=8 mean=0.682 std=0.032 min=0.598 max=0.705 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04
32%|β–ˆβ–ˆβ–ˆβ– | 30/94 [01:29<03:27, 3.24s/it]

{'loss': '-0.02124', 'grad_norm': '5.719', 'learning_rate': '1.383e-06', 'num_tokens': '1.147e+05', 'completions/mean_length': '17.12', 'completions/min_length': '16', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.12', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6819', 'rewards/reward_fn/std': '0.0343', 'reward': '0.6819', 'reward_std': '0.0343', 'frac_reward_zero_std': '0', 'entropy': '0.2194', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.835', 'epoch': '0.6383'}
32%|β–ˆβ–ˆβ–ˆβ– | 30/94 [01:29<03:27, 3.24s/it]03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:07:41 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:07:41 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:07:41 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:07:41 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:07:43 [INFO] [reward] call=31 n=8 mean=0.663 std=0.026 min=0.631 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
{'loss': '-0.01252', 'grad_norm': '1.514', 'learning_rate': '1.362e-06', 'num_tokens': '1.182e+05', 'completions/mean_length': '22.25', 'completions/min_length': '15', 'completions/max_length': '32', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '32', 'rewards/reward_fn/mean': '0.6626', 'rewards/reward_fn/std': '0.02803', 'reward': '0.6626', 'reward_std': '0.02803', 'frac_reward_zero_std': '0.5', 'entropy': '0.5726', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.906', 'epoch': '0.6596'}
33%|β–ˆβ–ˆβ–ˆβ–Ž | 31/94 [01:32<03:27, 3.29s/it]

33%|β–ˆβ–ˆβ–ˆβ–Ž | 31/94 [01:32<03:27, 3.29s/it]03:07:46 [INFO] [reward] call=32 n=8 mean=0.616 std=0.018 min=0.598 max=0.633 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.03
{'loss': '0', 'grad_norm': '0', 'learning_rate': '1.34e-06', 'num_tokens': '1.218e+05', 'completions/mean_length': '15.5', 'completions/min_length': '15', 'completions/max_length': '16', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '16', 'rewards/reward_fn/mean': '0.6158', 'rewards/reward_fn/std': '0.01876', 'reward': '0.6158', 'reward_std': '0.01876', 'frac_reward_zero_std': '1', 'entropy': '0.03686', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.571', 'epoch': '0.6809'}
34%|β–ˆβ–ˆβ–ˆβ– | 32/94 [01:35<03:11, 3.09s/it]

34%|β–ˆβ–ˆβ–ˆβ– | 32/94 [01:35<03:11, 3.09s/it]03:07:48 [INFO] [reward] call=33 n=8 mean=0.633 std=0.044 min=0.598 max=0.690 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.04 task=0.04 len=0.04
{'loss': '-0.01142', 'grad_norm': '4.521', 'learning_rate': '1.319e-06', 'num_tokens': '1.27e+05', 'completions/mean_length': '16.38', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.38', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6325', 'rewards/reward_fn/std': '0.04669', 'reward': '0.6325', 'reward_std': '0.04669', 'frac_reward_zero_std': '0.5', 'entropy': '0.1329', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.463', 'epoch': '0.7021'}
35%|β–ˆβ–ˆβ–ˆβ–Œ | 33/94 [01:37<02:58, 2.92s/it]

35%|β–ˆβ–ˆβ–ˆβ–Œ | 33/94 [01:37<02:58, 2.92s/it]03:07:49 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
03:07:52 [INFO] [reward] call=34 n=8 mean=0.656 std=0.030 min=0.611 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04
36%|β–ˆβ–ˆβ–ˆβ–Œ | 34/94 [01:41<03:13, 3.23s/it]

{'loss': '-0.3476', 'grad_norm': '2.982', 'learning_rate': '1.298e-06', 'num_tokens': '1.306e+05', 'completions/mean_length': '22.5', 'completions/min_length': '15', 'completions/max_length': '57', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '57', 'rewards/reward_fn/mean': '0.6558', 'rewards/reward_fn/std': '0.0324', 'reward': '0.6558', 'reward_std': '0.0324', 'frac_reward_zero_std': '0', 'entropy': '0.6957', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.883', 'epoch': '0.7234'}
36%|β–ˆβ–ˆβ–ˆβ–Œ | 34/94 [01:41<03:13, 3.23s/it]03:07:55 [INFO] [reward] call=35 n=8 mean=0.660 std=0.020 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04
37%|β–ˆβ–ˆβ–ˆβ–‹ | 35/94 [01:44<03:07, 3.18s/it]

{'loss': '-0.07906', 'grad_norm': '2.936', 'learning_rate': '1.277e-06', 'num_tokens': '1.341e+05', 'completions/mean_length': '21.88', 'completions/min_length': '15', 'completions/max_length': '36', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '36', 'rewards/reward_fn/mean': '0.6599', 'rewards/reward_fn/std': '0.02178', 'reward': '0.6599', 'reward_std': '0.02178', 'frac_reward_zero_std': '0', 'entropy': '0.3931', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.998', 'epoch': '0.7447'}
37%|β–ˆβ–ˆβ–ˆβ–‹ | 35/94 [01:44<03:07, 3.18s/it]03:07:57 [INFO] [reward] call=36 n=8 mean=0.459 std=0.466 min=-0.774 max=0.639 | parts mean: fmt=0.18 at=0.22 svc=0.04 par=0.05 task=0.04 len=0.04
38%|β–ˆβ–ˆβ–ˆβ–Š | 36/94 [01:46<02:43, 2.82s/it]

{'loss': '-0.09559', 'grad_norm': '2.419', 'learning_rate': '1.255e-06', 'num_tokens': '1.36e+05', 'completions/mean_length': '21.62', 'completions/min_length': '15', 'completions/max_length': '39', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.62', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '39', 'rewards/reward_fn/mean': '0.4588', 'rewards/reward_fn/std': '0.4983', 'reward': '0.4588', 'reward_std': '0.4983', 'frac_reward_zero_std': '0', 'entropy': '0.8276', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.981', 'epoch': '0.766'}
38%|β–ˆβ–ˆβ–ˆβ–Š | 36/94 [01:46<02:43, 2.82s/it]03:08:00 [INFO] [reward] call=37 n=8 mean=0.689 std=0.002 min=0.688 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04
39%|β–ˆβ–ˆβ–ˆβ–‰ | 37/94 [01:49<02:37, 2.76s/it]

{'loss': '1.49e-08', 'grad_norm': '1.85', 'learning_rate': '1.234e-06', 'num_tokens': '1.413e+05', 'completions/mean_length': '17.5', 'completions/min_length': '17', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.5', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6893', 'rewards/reward_fn/std': '0.002546', 'reward': '0.6893', 'reward_std': '0.002546', 'frac_reward_zero_std': '0.5', 'entropy': '0.1432', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.549', 'epoch': '0.7872'}
39%|β–ˆβ–ˆβ–ˆβ–‰ | 37/94 [01:49<02:37, 2.76s/it]03:08:02 [INFO] [reward] call=38 n=8 mean=0.636 std=0.005 min=0.633 max=0.647 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
40%|β–ˆβ–ˆβ–ˆβ–ˆ | 38/94 [01:51<02:20, 2.50s/it]

{'loss': '-0.1723', 'grad_norm': '1.639', 'learning_rate': '1.213e-06', 'num_tokens': '1.432e+05', 'completions/mean_length': '18.25', 'completions/min_length': '15', 'completions/max_length': '35', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '35', 'rewards/reward_fn/mean': '0.6364', 'rewards/reward_fn/std': '0.005656', 'reward': '0.6364', 'reward_std': '0.005656', 'frac_reward_zero_std': '0.5', 'entropy': '0.2606', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.895', 'epoch': '0.8085'}
40%|β–ˆβ–ˆβ–ˆβ–ˆ | 38/94 [01:51<02:20, 2.50s/it]03:08:04 [INFO] [reward] call=39 n=8 mean=0.635 std=0.004 min=0.633 max=0.645 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
41%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 39/94 [01:52<02:06, 2.30s/it]

{'loss': '-0.1923', 'grad_norm': '2.579', 'learning_rate': '1.191e-06', 'num_tokens': '1.45e+05', 'completions/mean_length': '17.25', 'completions/min_length': '15', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6348', 'rewards/reward_fn/std': '0.003995', 'reward': '0.6348', 'reward_std': '0.003995', 'frac_reward_zero_std': '0.5', 'entropy': '0.3042', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.817', 'epoch': '0.8298'}
41%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 39/94 [01:52<02:06, 2.30s/it]03:08:07 [INFO] [reward] call=40 n=8 mean=0.645 std=0.045 min=0.599 max=0.704 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.08 task=0.04 len=0.02
43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 40/94 [01:56<02:26, 2.71s/it]

{'loss': '0.3164', 'grad_norm': '4.612', 'learning_rate': '1.17e-06', 'num_tokens': '1.486e+05', 'completions/mean_length': '20.38', 'completions/min_length': '10', 'completions/max_length': '49', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.38', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '49', 'rewards/reward_fn/mean': '0.6452', 'rewards/reward_fn/std': '0.04773', 'reward': '0.6452', 'reward_std': '0.04773', 'frac_reward_zero_std': '0', 'entropy': '0.9359', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.619', 'epoch': '0.8511'}
43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 40/94 [01:56<02:26, 2.71s/it]03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:08:08 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:08:08 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:08:09 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:08:09 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:08:11 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
03:08:11 [INFO] [reward] call=41 n=8 mean=0.663 std=0.026 min=0.631 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
44%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 41/94 [02:00<02:50, 3.21s/it]

{'loss': '-0.2084', 'grad_norm': '1.419', 'learning_rate': '1.149e-06', 'num_tokens': '1.521e+05', 'completions/mean_length': '24.12', 'completions/min_length': '15', 'completions/max_length': '52', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '24.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '52', 'rewards/reward_fn/mean': '0.6634', 'rewards/reward_fn/std': '0.02773', 'reward': '0.6634', 'reward_std': '0.02773', 'frac_reward_zero_std': '0.5', 'entropy': '0.2573', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.818', 'epoch': '0.8723'}
44%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 41/94 [02:00<02:50, 3.21s/it]03:08:14 [INFO] [reward] call=42 n=8 mean=0.644 std=0.033 min=0.611 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.06 task=0.04 len=0.03
45%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 42/94 [02:03<02:38, 3.05s/it]

{'loss': '-0.02793', 'grad_norm': '7.134', 'learning_rate': '1.128e-06', 'num_tokens': '1.556e+05', 'completions/mean_length': '16', 'completions/min_length': '15', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6443', 'rewards/reward_fn/std': '0.03566', 'reward': '0.6443', 'reward_std': '0.03566', 'frac_reward_zero_std': '0.5', 'entropy': '0.1252', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.611', 'epoch': '0.8936'}
45%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 42/94 [02:03<02:38, 3.05s/it]03:08:17 [INFO] [reward] call=43 n=8 mean=0.655 std=0.044 min=0.598 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.07 task=0.04 len=0.04
46%|β–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 43/94 [02:06<02:29, 2.94s/it]

{'loss': '-0.01125', 'grad_norm': '3.344', 'learning_rate': '1.106e-06', 'num_tokens': '1.608e+05', 'completions/mean_length': '16.62', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.62', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6546', 'rewards/reward_fn/std': '0.04663', 'reward': '0.6546', 'reward_std': '0.04663', 'frac_reward_zero_std': '0.5', 'entropy': '0.04298', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.612', 'epoch': '0.9149'}
46%|β–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 43/94 [02:06<02:29, 2.94s/it]03:08:22 [INFO] [reward] call=44 n=8 mean=0.401 std=0.479 min=-0.774 max=0.688 | parts mean: fmt=0.18 at=0.19 svc=0.05 par=0.06 task=0.03 len=0.03
47%|β–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 44/94 [02:11<03:03, 3.66s/it]

{'loss': '0.5893', 'grad_norm': '5.321', 'learning_rate': '1.085e-06', 'num_tokens': '1.662e+05', 'completions/mean_length': '25.5', 'completions/min_length': '10', 'completions/max_length': '100', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '25.5', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '100', 'rewards/reward_fn/mean': '0.4009', 'rewards/reward_fn/std': '0.5121', 'reward': '0.4009', 'reward_std': '0.5121', 'frac_reward_zero_std': '0', 'entropy': '0.3116', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '5.283', 'epoch': '0.9362'}
47%|β–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 44/94 [02:11<03:03, 3.66s/it]03:08:26 [INFO] [reward] call=45 n=8 mean=0.631 std=0.030 min=0.598 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.04 task=0.04 len=0.04
{'loss': '-0.1902', 'grad_norm': '3.174', 'learning_rate': '1.064e-06', 'num_tokens': '1.698e+05', 'completions/mean_length': '25.12', 'completions/min_length': '15', 'completions/max_length': '51', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '25.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '51', 'rewards/reward_fn/mean': '0.6312', 'rewards/reward_fn/std': '0.03228', 'reward': '0.6312', 'reward_std': '0.03228', 'frac_reward_zero_std': '0', 'entropy': '0.3894', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.692', 'epoch': '0.9574'}
48%|β–ˆβ–ˆβ–ˆβ–ˆβ–Š | 45/94 [02:15<03:00, 3.69s/it]

48%|β–ˆβ–ˆβ–ˆβ–ˆβ–Š | 45/94 [02:15<03:00, 3.69s/it]03:08:28 [INFO] [reward] call=46 n=8 mean=0.636 std=0.006 min=0.632 max=0.647 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
49%|β–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 46/94 [02:17<02:31, 3.16s/it]

{'loss': '-0.3453', 'grad_norm': '2.959', 'learning_rate': '1.043e-06', 'num_tokens': '1.717e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6364', 'rewards/reward_fn/std': '0.006058', 'reward': '0.6364', 'reward_std': '0.006058', 'frac_reward_zero_std': '0', 'entropy': '0.3825', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.92', 'epoch': '0.9787'}
49%|β–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 46/94 [02:17<02:31, 3.16s/it]03:08:31 [INFO] [reward] call=47 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04
50%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 47/94 [02:20<02:21, 3.01s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '1.021e-06', 'num_tokens': '1.769e+05', 'completions/mean_length': '17', 'completions/min_length': '17', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.1325', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.599', 'epoch': '1'}
50%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 47/94 [02:20<02:21, 3.01s/it]03:08:34 [INFO] [reward] call=48 n=8 mean=0.642 std=0.033 min=0.598 max=0.690 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.06 task=0.04 len=0.04
51%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 48/94 [02:23<02:20, 3.05s/it]

{'loss': '-0.1938', 'grad_norm': '5.199', 'learning_rate': '1e-06', 'num_tokens': '1.804e+05', 'completions/mean_length': '19.5', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6424', 'rewards/reward_fn/std': '0.03487', 'reward': '0.6424', 'reward_std': '0.03487', 'frac_reward_zero_std': '0', 'entropy': '0.5192', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.071', 'epoch': '1.021'}
51%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 48/94 [02:23<02:20, 3.05s/it]03:08:37 [INFO] [reward] call=49 n=8 mean=0.671 std=0.026 min=0.633 max=0.701 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04
52%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 49/94 [02:26<02:21, 3.15s/it]

{'loss': '-0.008332', 'grad_norm': '1.737', 'learning_rate': '9.787e-07', 'num_tokens': '1.839e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6713', 'rewards/reward_fn/std': '0.02778', 'reward': '0.6713', 'reward_std': '0.02778', 'frac_reward_zero_std': '0.5', 'entropy': '0.3038', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.344', 'epoch': '1.043'}
52%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 49/94 [02:26<02:21, 3.15s/it]03:08:41 [INFO] [reward] call=50 n=8 mean=0.621 std=0.023 min=0.599 max=0.655 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04
53%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 50/94 [02:30<02:29, 3.40s/it]

{'loss': '-0.1466', 'grad_norm': '2.019', 'learning_rate': '9.574e-07', 'num_tokens': '1.875e+05', 'completions/mean_length': '27.5', 'completions/min_length': '16', 'completions/max_length': '60', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '27.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '60', 'rewards/reward_fn/mean': '0.6212', 'rewards/reward_fn/std': '0.02489', 'reward': '0.6212', 'reward_std': '0.02489', 'frac_reward_zero_std': '0.5', 'entropy': '0.6976', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.911', 'epoch': '1.064'}
53%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 50/94 [02:30<02:29, 3.40s/it]03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:08:42 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:08:42 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:08:42 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:08:43 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:08:44 [INFO] [reward] call=51 n=8 mean=0.646 std=0.042 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.09 task=0.04 len=0.01
03:08:45 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
54%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 51/94 [02:33<02:21, 3.30s/it]

{'loss': '0.08536', 'grad_norm': '5.356', 'learning_rate': '9.362e-07', 'num_tokens': '1.929e+05', 'completions/mean_length': '15', 'completions/min_length': '10', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6463', 'rewards/reward_fn/std': '0.04467', 'reward': '0.6463', 'reward_std': '0.04467', 'frac_reward_zero_std': '0', 'entropy': '0.1691', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.517', 'epoch': '1.085'}
54%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 51/94 [02:33<02:21, 3.30s/it]03:08:48 [INFO] [reward] call=52 n=8 mean=0.662 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
55%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 52/94 [02:37<02:20, 3.34s/it]
{'loss': '-0.2974', 'grad_norm': '1.814', 'learning_rate': '9.149e-07', 'num_tokens': '1.963e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6616', 'rewards/reward_fn/std': '0.02868', 'reward': '0.6616', 'reward_std': '0.02868', 'frac_reward_zero_std': '0.5', 'entropy': '0.1737', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.375', 'epoch': '1.106'}

55%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 52/94 [02:37<02:20, 3.34s/it]03:08:51 [INFO] [reward] call=53 n=8 mean=0.565 std=0.152 min=0.165 max=0.633 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.07 task=0.04 len=0.01
56%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 53/94 [02:40<02:13, 3.26s/it]

{'loss': '0.1845', 'grad_norm': '3.084', 'learning_rate': '8.936e-07', 'num_tokens': '1.999e+05', 'completions/mean_length': '15.38', 'completions/min_length': '10', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.38', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.5648', 'rewards/reward_fn/std': '0.1621', 'reward': '0.5648', 'reward_std': '0.1621', 'frac_reward_zero_std': '0.5', 'entropy': '0.1412', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.009', 'epoch': '1.128'}
56%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 53/94 [02:40<02:13, 3.26s/it]03:08:54 [INFO] [reward] call=54 n=8 mean=0.637 std=0.038 min=0.598 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.05 task=0.04 len=0.04
57%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 54/94 [02:43<02:07, 3.18s/it]

{'loss': '-0.01893', 'grad_norm': '3.71', 'learning_rate': '8.723e-07', 'num_tokens': '2.035e+05', 'completions/mean_length': '17.75', 'completions/min_length': '15', 'completions/max_length': '28', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '28', 'rewards/reward_fn/mean': '0.6366', 'rewards/reward_fn/std': '0.04055', 'reward': '0.6366', 'reward_std': '0.04055', 'frac_reward_zero_std': '0', 'entropy': '0.249', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.946', 'epoch': '1.149'}
57%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 54/94 [02:43<02:07, 3.18s/it]03:08:57 [INFO] [reward] call=55 n=8 mean=0.669 std=0.027 min=0.633 max=0.694 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04
59%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 55/94 [02:46<02:02, 3.15s/it]

{'loss': '0.02042', 'grad_norm': '1.908', 'learning_rate': '8.511e-07', 'num_tokens': '2.07e+05', 'completions/mean_length': '19.5', 'completions/min_length': '15', 'completions/max_length': '35', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '35', 'rewards/reward_fn/mean': '0.6687', 'rewards/reward_fn/std': '0.02878', 'reward': '0.6687', 'reward_std': '0.02878', 'frac_reward_zero_std': '0.5', 'entropy': '0.2068', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.011', 'epoch': '1.17'}
59%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 55/94 [02:46<02:02, 3.15s/it]03:09:01 [INFO] [reward] call=56 n=8 mean=0.672 std=0.026 min=0.633 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04
{'loss': '-0.05793', 'grad_norm': '2.046', 'learning_rate': '8.298e-07', 'num_tokens': '2.106e+05', 'completions/mean_length': '22.25', 'completions/min_length': '15', 'completions/max_length': '53', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '53', 'rewards/reward_fn/mean': '0.6717', 'rewards/reward_fn/std': '0.02801', 'reward': '0.6717', 'reward_std': '0.02801', 'frac_reward_zero_std': '0.5', 'entropy': '0.4402', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.752', 'epoch': '1.191'}
60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 56/94 [02:49<02:07, 3.35s/it]

60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 56/94 [02:49<02:07, 3.35s/it]03:09:07 [INFO] [reward] call=57 n=8 mean=0.484 std=0.484 min=-0.795 max=0.688 | parts mean: fmt=0.18 at=0.22 svc=0.04 par=0.08 task=0.04 len=0.04
{'loss': '0.5857', 'grad_norm': '2.956', 'learning_rate': '8.085e-07', 'num_tokens': '2.144e+05', 'completions/mean_length': '33.38', 'completions/min_length': '15', 'completions/max_length': '128', 'completions/clipped_ratio': '0.125', 'completions/mean_terminated_length': '19.86', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '35', 'rewards/reward_fn/mean': '0.4842', 'rewards/reward_fn/std': '0.5176', 'reward': '0.4842', 'reward_std': '0.5176', 'frac_reward_zero_std': '0.5', 'entropy': '1.25', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '6.187', 'epoch': '1.213'}
61%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 57/94 [02:56<02:36, 4.22s/it]

61%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 57/94 [02:56<02:36, 4.22s/it]03:09:10 [INFO] [reward] call=58 n=8 mean=0.661 std=0.028 min=0.632 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
{'loss': '0.1203', 'grad_norm': '1.477', 'learning_rate': '7.872e-07', 'num_tokens': '2.18e+05', 'completions/mean_length': '25', 'completions/min_length': '15', 'completions/max_length': '39', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '39', 'rewards/reward_fn/mean': '0.6605', 'rewards/reward_fn/std': '0.02982', 'reward': '0.6605', 'reward_std': '0.02982', 'frac_reward_zero_std': '0.5', 'entropy': '0.1817', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.296', 'epoch': '1.234'}
62%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 58/94 [02:59<02:22, 3.96s/it]

62%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 58/94 [02:59<02:22, 3.96s/it]03:09:12 [INFO] [reward] call=59 n=8 mean=0.636 std=0.005 min=0.631 max=0.646 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
63%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 59/94 [03:01<01:54, 3.28s/it]
{'loss': '0.06092', 'grad_norm': '1.475', 'learning_rate': '7.66e-07', 'num_tokens': '2.198e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '28', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '28', 'rewards/reward_fn/mean': '0.6359', 'rewards/reward_fn/std': '0.005674', 'reward': '0.6359', 'reward_std': '0.005674', 'frac_reward_zero_std': '0', 'entropy': '0.9263', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.674', 'epoch': '1.255'}

63%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 59/94 [03:01<01:54, 3.28s/it]03:09:15 [INFO] [reward] call=60 n=8 mean=0.668 std=0.027 min=0.633 max=0.690 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04
64%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 60/94 [03:04<01:50, 3.24s/it]

{'loss': '-0.01362', 'grad_norm': '2.836', 'learning_rate': '7.447e-07', 'num_tokens': '2.234e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '36', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '36', 'rewards/reward_fn/mean': '0.6682', 'rewards/reward_fn/std': '0.02847', 'reward': '0.6682', 'reward_std': '0.02847', 'frac_reward_zero_std': '0', 'entropy': '0.4603', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.095', 'epoch': '1.277'}
64%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 60/94 [03:04<01:50, 3.24s/it]03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:09:16 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:09:16 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:09:16 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:09:16 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:09:19 [INFO] [reward] call=61 n=8 mean=0.619 std=0.020 min=0.599 max=0.647 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04
03:09:20 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
65%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 61/94 [03:08<01:57, 3.57s/it]

{'loss': '-0.2418', 'grad_norm': '1.858', 'learning_rate': '7.234e-07', 'num_tokens': '2.27e+05', 'completions/mean_length': '23.38', 'completions/min_length': '15', 'completions/max_length': '55', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '23.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '55', 'rewards/reward_fn/mean': '0.6188', 'rewards/reward_fn/std': '0.0218', 'reward': '0.6188', 'reward_std': '0.0218', 'frac_reward_zero_std': '0.5', 'entropy': '0.5959', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.761', 'epoch': '1.298'}
65%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 61/94 [03:08<01:57, 3.57s/it]03:09:22 [INFO] [reward] call=62 n=8 mean=0.693 std=0.005 min=0.687 max=0.698 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04
66%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 62/94 [03:11<01:48, 3.39s/it]

{'loss': '-0.09937', 'grad_norm': '2.591', 'learning_rate': '7.021e-07', 'num_tokens': '2.323e+05', 'completions/mean_length': '18.62', 'completions/min_length': '17', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.62', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.6926', 'rewards/reward_fn/std': '0.004852', 'reward': '0.6926', 'reward_std': '0.004852', 'frac_reward_zero_std': '0', 'entropy': '0.2111', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.929', 'epoch': '1.319'}
66%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 62/94 [03:11<01:48, 3.39s/it]03:09:25 [INFO] [reward] call=63 n=8 mean=0.580 std=0.182 min=0.105 max=0.699 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.06 task=0.04 len=0.03
67%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 63/94 [03:14<01:38, 3.17s/it]

{'loss': '0.01356', 'grad_norm': '5.029', 'learning_rate': '6.809e-07', 'num_tokens': '2.358e+05', 'completions/mean_length': '16', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.5795', 'rewards/reward_fn/std': '0.1948', 'reward': '0.5795', 'reward_std': '0.1948', 'frac_reward_zero_std': '0.5', 'entropy': '0.06583', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.588', 'epoch': '1.34'}
67%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 63/94 [03:14<01:38, 3.17s/it]03:09:28 [INFO] [reward] call=64 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04
68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 64/94 [03:17<01:31, 3.04s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '6.596e-07', 'num_tokens': '2.411e+05', 'completions/mean_length': '17.25', 'completions/min_length': '17', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.25', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.1545', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.688', 'epoch': '1.362'}
68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 64/94 [03:17<01:31, 3.04s/it]03:09:30 [INFO] [reward] call=65 n=8 mean=0.635 std=0.003 min=0.633 max=0.643 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
69%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 65/94 [03:19<01:19, 2.74s/it]

{'loss': '-0.2188', 'grad_norm': '1.902', 'learning_rate': '6.383e-07', 'num_tokens': '2.43e+05', 'completions/mean_length': '22.12', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.635', 'rewards/reward_fn/std': '0.003465', 'reward': '0.635', 'reward_std': '0.003465', 'frac_reward_zero_std': '0.5', 'entropy': '0.3234', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.036', 'epoch': '1.383'}
69%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 65/94 [03:19<01:19, 2.74s/it]03:09:33 [INFO] [reward] call=66 n=8 mean=0.664 std=0.028 min=0.633 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
70%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 66/94 [03:22<01:23, 2.99s/it]

{'loss': '-0.3033', 'grad_norm': '2.428', 'learning_rate': '6.17e-07', 'num_tokens': '2.466e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6642', 'rewards/reward_fn/std': '0.02946', 'reward': '0.6642', 'reward_std': '0.02946', 'frac_reward_zero_std': '0', 'entropy': '0.36', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.486', 'epoch': '1.404'}
70%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 66/94 [03:22<01:23, 2.99s/it]03:09:36 [INFO] [reward] call=67 n=8 mean=0.644 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04
71%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 67/94 [03:25<01:17, 2.86s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '5.957e-07', 'num_tokens': '2.518e+05', 'completions/mean_length': '16.5', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6436', 'rewards/reward_fn/std': '0.04789', 'reward': '0.6436', 'reward_std': '0.04789', 'frac_reward_zero_std': '1', 'entropy': '0.03455', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.51', 'epoch': '1.426'}
71%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 67/94 [03:25<01:17, 2.86s/it]03:09:39 [INFO] [reward] call=68 n=8 mean=0.661 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
{'loss': '-0.08996', 'grad_norm': '0.9076', 'learning_rate': '5.745e-07', 'num_tokens': '2.553e+05', 'completions/mean_length': '18.75', 'completions/min_length': '15', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6609', 'rewards/reward_fn/std': '0.02937', 'reward': '0.6609', 'reward_std': '0.02937', 'frac_reward_zero_std': '0.5', 'entropy': '0.1516', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.04', 'epoch': '1.447'}
72%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 68/94 [03:28<01:16, 2.93s/it]

72%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 68/94 [03:28<01:16, 2.93s/it]03:09:42 [INFO] [reward] call=69 n=8 mean=0.626 std=0.032 min=0.599 max=0.699 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.05 task=0.04 len=0.02
73%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 69/94 [03:31<01:11, 2.84s/it]

{'loss': '0.02351', 'grad_norm': '4.787', 'learning_rate': '5.532e-07', 'num_tokens': '2.589e+05', 'completions/mean_length': '15.25', 'completions/min_length': '10', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.25', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6258', 'rewards/reward_fn/std': '0.03376', 'reward': '0.6258', 'reward_std': '0.03376', 'frac_reward_zero_std': '0', 'entropy': '0.1078', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.576', 'epoch': '1.468'}
73%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 69/94 [03:31<01:11, 2.84s/it]03:09:44 [INFO] [reward] call=70 n=8 mean=0.654 std=0.025 min=0.633 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04
74%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 70/94 [03:33<01:02, 2.61s/it]

{'loss': '0.04798', 'grad_norm': '2.384', 'learning_rate': '5.319e-07', 'num_tokens': '2.608e+05', 'completions/mean_length': '22.5', 'completions/min_length': '15', 'completions/max_length': '41', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '41', 'rewards/reward_fn/mean': '0.654', 'rewards/reward_fn/std': '0.02694', 'reward': '0.654', 'reward_std': '0.02694', 'frac_reward_zero_std': '0', 'entropy': '0.6745', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.055', 'epoch': '1.489'}
74%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 70/94 [03:33<01:02, 2.61s/it]03:09:44 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:09:44 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:09:45 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:09:45 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:09:45 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:09:45 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:09:45 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:09:45 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:09:47 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
03:09:48 [INFO] [reward] call=71 n=8 mean=0.650 std=0.037 min=0.598 max=0.699 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.07 task=0.04 len=0.04
76%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 71/94 [03:37<01:08, 3.00s/it]

{'loss': '0.05797', 'grad_norm': '2.903', 'learning_rate': '5.106e-07', 'num_tokens': '2.644e+05', 'completions/mean_length': '23.38', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '23.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6501', 'rewards/reward_fn/std': '0.03967', 'reward': '0.6501', 'reward_std': '0.03967', 'frac_reward_zero_std': '0', 'entropy': '0.5277', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.364', 'epoch': '1.511'}
76%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 71/94 [03:37<01:08, 3.00s/it]03:09:50 [INFO] [reward] call=72 n=8 mean=0.646 std=0.023 min=0.633 max=0.705 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04
77%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 72/94 [03:38<00:58, 2.65s/it]

{'loss': '-0.255', 'grad_norm': '3.032', 'learning_rate': '4.894e-07', 'num_tokens': '2.662e+05', 'completions/mean_length': '19.62', 'completions/min_length': '15', 'completions/max_length': '38', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.62', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '38', 'rewards/reward_fn/mean': '0.6461', 'rewards/reward_fn/std': '0.02479', 'reward': '0.6461', 'reward_std': '0.02479', 'frac_reward_zero_std': '0', 'entropy': '0.4351', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.847', 'epoch': '1.532'}
77%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 72/94 [03:38<00:58, 2.65s/it]03:09:52 [INFO] [reward] call=73 n=8 mean=0.652 std=0.031 min=0.598 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04
{'loss': '-0.05995', 'grad_norm': '5.942', 'learning_rate': '4.681e-07', 'num_tokens': '2.698e+05', 'completions/mean_length': '19.5', 'completions/min_length': '16', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.6523', 'rewards/reward_fn/std': '0.03299', 'reward': '0.6523', 'reward_std': '0.03299', 'frac_reward_zero_std': '0', 'entropy': '0.3749', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.902', 'epoch': '1.553'}
78%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 73/94 [03:41<00:57, 2.75s/it]

78%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 73/94 [03:41<00:57, 2.75s/it]03:09:55 [INFO] [reward] call=74 n=8 mean=0.593 std=0.163 min=0.172 max=0.700 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.07 task=0.04 len=0.04
79%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 74/94 [03:44<00:55, 2.77s/it]

{'loss': '0.03651', 'grad_norm': '3.928', 'learning_rate': '4.468e-07', 'num_tokens': '2.733e+05', 'completions/mean_length': '19.75', 'completions/min_length': '15', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.5934', 'rewards/reward_fn/std': '0.1739', 'reward': '0.5934', 'reward_std': '0.1739', 'frac_reward_zero_std': '0', 'entropy': '0.4748', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.764', 'epoch': '1.574'}
79%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 74/94 [03:44<00:55, 2.77s/it]03:09:58 [INFO] [reward] call=75 n=8 mean=0.618 std=0.020 min=0.599 max=0.650 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04
80%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 75/94 [03:47<00:53, 2.81s/it]

{'loss': '-0.1599', 'grad_norm': '2.755', 'learning_rate': '4.255e-07', 'num_tokens': '2.769e+05', 'completions/mean_length': '17.38', 'completions/min_length': '15', 'completions/max_length': '30', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '30', 'rewards/reward_fn/mean': '0.6182', 'rewards/reward_fn/std': '0.02141', 'reward': '0.6182', 'reward_std': '0.02141', 'frac_reward_zero_std': '0.5', 'entropy': '0.1834', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.857', 'epoch': '1.596'}
80%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 75/94 [03:47<00:53, 2.81s/it]03:10:01 [INFO] [reward] call=76 n=8 mean=0.667 std=0.040 min=0.598 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.09 task=0.04 len=0.04
81%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 76/94 [03:50<00:49, 2.76s/it]

{'loss': '-0.01251', 'grad_norm': '4.964', 'learning_rate': '4.043e-07', 'num_tokens': '2.82e+05', 'completions/mean_length': '17.25', 'completions/min_length': '16', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.25', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6668', 'rewards/reward_fn/std': '0.04234', 'reward': '0.6668', 'reward_std': '0.04234', 'frac_reward_zero_std': '0.5', 'entropy': '0.1384', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.587', 'epoch': '1.617'}
81%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 76/94 [03:50<00:49, 2.76s/it]03:10:04 [INFO] [reward] call=77 n=8 mean=0.493 std=0.485 min=-0.789 max=0.690 | parts mean: fmt=0.18 at=0.22 svc=0.05 par=0.09 task=0.04 len=0.04
82%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 77/94 [03:52<00:46, 2.74s/it]

{'loss': '0.05487', 'grad_norm': '3.762', 'learning_rate': '3.83e-07', 'num_tokens': '2.872e+05', 'completions/mean_length': '18', 'completions/min_length': '16', 'completions/max_length': '22', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '22', 'rewards/reward_fn/mean': '0.4926', 'rewards/reward_fn/std': '0.5189', 'reward': '0.4926', 'reward_std': '0.5189', 'frac_reward_zero_std': '0.5', 'entropy': '0.2181', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.612', 'epoch': '1.638'}
82%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 77/94 [03:52<00:46, 2.74s/it]03:10:06 [INFO] [reward] call=78 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04
83%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 78/94 [03:55<00:43, 2.69s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '3.617e-07', 'num_tokens': '2.924e+05', 'completions/mean_length': '17.88', 'completions/min_length': '17', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.88', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.09081', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.535', 'epoch': '1.66'}
83%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 78/94 [03:55<00:43, 2.69s/it]03:10:09 [INFO] [reward] call=79 n=8 mean=0.592 std=0.185 min=0.109 max=0.699 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.07 task=0.04 len=0.04
84%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 79/94 [03:58<00:41, 2.80s/it]

{'loss': '-0.03821', 'grad_norm': '5.515', 'learning_rate': '3.404e-07', 'num_tokens': '2.96e+05', 'completions/mean_length': '18', 'completions/min_length': '11', 'completions/max_length': '32', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '32', 'rewards/reward_fn/mean': '0.5917', 'rewards/reward_fn/std': '0.1975', 'reward': '0.5917', 'reward_std': '0.1975', 'frac_reward_zero_std': '0', 'entropy': '0.2983', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.973', 'epoch': '1.681'}
84%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 79/94 [03:58<00:41, 2.80s/it]03:10:13 [INFO] [reward] call=80 n=8 mean=0.663 std=0.026 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
85%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 80/94 [04:02<00:42, 3.06s/it]
{'loss': '-0.3258', 'grad_norm': '2.044', 'learning_rate': '3.191e-07', 'num_tokens': '2.996e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '52', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '52', 'rewards/reward_fn/mean': '0.6627', 'rewards/reward_fn/std': '0.02788', 'reward': '0.6627', 'reward_std': '0.02788', 'frac_reward_zero_std': '0.5', 'entropy': '0.1508', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.607', 'epoch': '1.702'}

85%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 80/94 [04:02<00:42, 3.06s/it]03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:14 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:14 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:14 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:10:14 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:10:16 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
03:10:17 [INFO] [reward] call=81 n=8 mean=0.674 std=0.033 min=0.630 max=0.713 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04
86%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 81/94 [04:06<00:42, 3.30s/it]

{'loss': '-0.1966', 'grad_norm': '3.974', 'learning_rate': '2.979e-07', 'num_tokens': '3.032e+05', 'completions/mean_length': '20', 'completions/min_length': '15', 'completions/max_length': '39', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '39', 'rewards/reward_fn/mean': '0.6736', 'rewards/reward_fn/std': '0.03483', 'reward': '0.6736', 'reward_std': '0.03483', 'frac_reward_zero_std': '0', 'entropy': '0.3634', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.307', 'epoch': '1.723'}
86%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 81/94 [04:06<00:42, 3.30s/it]03:10:19 [INFO] [reward] call=82 n=8 mean=0.653 std=0.025 min=0.633 max=0.700 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04
87%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 82/94 [04:08<00:34, 2.90s/it]

{'loss': '-0.09556', 'grad_norm': '1.988', 'learning_rate': '2.766e-07', 'num_tokens': '3.051e+05', 'completions/mean_length': '23.75', 'completions/min_length': '15', 'completions/max_length': '38', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '23.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '38', 'rewards/reward_fn/mean': '0.6531', 'rewards/reward_fn/std': '0.02641', 'reward': '0.6531', 'reward_std': '0.02641', 'frac_reward_zero_std': '0', 'entropy': '1', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.982', 'epoch': '1.745'}
87%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 82/94 [04:08<00:34, 2.90s/it]03:10:22 [INFO] [reward] call=83 n=8 mean=0.671 std=0.027 min=0.633 max=0.702 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04
{'loss': '-0.03983', 'grad_norm': '1.799', 'learning_rate': '2.553e-07', 'num_tokens': '3.087e+05', 'completions/mean_length': '21.88', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.671', 'rewards/reward_fn/std': '0.02835', 'reward': '0.671', 'reward_std': '0.02835', 'frac_reward_zero_std': '0.5', 'entropy': '0.5', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.575', 'epoch': '1.766'}
88%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 83/94 [04:11<00:34, 3.12s/it]

88%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 83/94 [04:11<00:34, 3.12s/it]03:10:25 [INFO] [reward] call=84 n=8 mean=0.598 std=0.221 min=0.020 max=0.704 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.09 task=0.04 len=0.02
89%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 84/94 [04:14<00:30, 3.05s/it]

{'loss': '-0.09434', 'grad_norm': '6.413', 'learning_rate': '2.34e-07', 'num_tokens': '3.139e+05', 'completions/mean_length': '16.62', 'completions/min_length': '11', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.62', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.5975', 'rewards/reward_fn/std': '0.2358', 'reward': '0.5975', 'reward_std': '0.2358', 'frac_reward_zero_std': '0', 'entropy': '0.1545', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.821', 'epoch': '1.787'}
89%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 84/94 [04:14<00:30, 3.05s/it]03:10:28 [INFO] [reward] call=85 n=8 mean=0.664 std=0.025 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04
90%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 85/94 [04:17<00:27, 3.07s/it]
{'loss': '-0.1992', 'grad_norm': '1.506', 'learning_rate': '2.128e-07', 'num_tokens': '3.174e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6643', 'rewards/reward_fn/std': '0.02633', 'reward': '0.6643', 'reward_std': '0.02633', 'frac_reward_zero_std': '0.5', 'entropy': '0.2386', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.042', 'epoch': '1.809'}

90%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 85/94 [04:17<00:27, 3.07s/it]03:10:32 [INFO] [reward] call=86 n=8 mean=0.666 std=0.035 min=0.611 max=0.712 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.08 task=0.04 len=0.04
91%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 86/94 [04:21<00:25, 3.21s/it]

{'loss': '-0.06428', 'grad_norm': '2.73', 'learning_rate': '1.915e-07', 'num_tokens': '3.21e+05', 'completions/mean_length': '24.12', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '24.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6659', 'rewards/reward_fn/std': '0.03774', 'reward': '0.6659', 'reward_std': '0.03774', 'frac_reward_zero_std': '0', 'entropy': '0.691', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.488', 'epoch': '1.83'}
91%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 86/94 [04:21<00:25, 3.21s/it]03:10:34 [INFO] [reward] call=87 n=8 mean=0.692 std=0.004 min=0.688 max=0.697 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04
93%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 87/94 [04:24<00:23, 3.31s/it]
{'loss': '-4.491e-06', 'grad_norm': '3.164', 'learning_rate': '1.702e-07', 'num_tokens': '3.264e+05', 'completions/mean_length': '17', 'completions/min_length': '17', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6915', 'rewards/reward_fn/std': '0.003961', 'reward': '0.6915', 'reward_std': '0.003961', 'frac_reward_zero_std': '0.5', 'entropy': '0.1973', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.48', 'epoch': '1.851'}

93%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 87/94 [04:24<00:23, 3.31s/it]03:10:39 [INFO] [reward] call=88 n=8 mean=0.484 std=0.478 min=-0.779 max=0.688 | parts mean: fmt=0.18 at=0.22 svc=0.04 par=0.08 task=0.04 len=0.04
94%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 88/94 [04:28<00:21, 3.54s/it]

{'loss': '0.4213', 'grad_norm': '2.828', 'learning_rate': '1.489e-07', 'num_tokens': '3.299e+05', 'completions/mean_length': '22.25', 'completions/min_length': '15', 'completions/max_length': '65', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '65', 'rewards/reward_fn/mean': '0.4843', 'rewards/reward_fn/std': '0.5113', 'reward': '0.4843', 'reward_std': '0.5113', 'frac_reward_zero_std': '0.5', 'entropy': '0.3975', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '4.028', 'epoch': '1.872'}
94%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 88/94 [04:28<00:21, 3.54s/it]03:10:42 [INFO] [reward] call=89 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04
95%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 89/94 [04:31<00:16, 3.26s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '1.277e-07', 'num_tokens': '3.35e+05', 'completions/mean_length': '18', 'completions/min_length': '18', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18', 'completions/min_terminated_length': '18', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.1114', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.534', 'epoch': '1.894'}
95%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 89/94 [04:31<00:16, 3.26s/it]03:10:45 [INFO] [reward] call=90 n=8 mean=0.644 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04
96%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ| 90/94 [04:33<00:12, 3.04s/it]

{'loss': '0', 'grad_norm': '0', 'learning_rate': '1.064e-07', 'num_tokens': '3.402e+05', 'completions/mean_length': '16.5', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6436', 'rewards/reward_fn/std': '0.04789', 'reward': '0.6436', 'reward_std': '0.04789', 'frac_reward_zero_std': '1', 'entropy': '0.05231', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.457', 'epoch': '1.915'}
96%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ| 90/94 [04:33<00:12, 3.04s/it]03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:46 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:46 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:46 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:10:46 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:10:47 [INFO] [reward] call=91 n=8 mean=0.636 std=0.003 min=0.633 max=0.642 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
97%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹| 91/94 [04:36<00:08, 2.89s/it]

{'loss': '-0.3753', 'grad_norm': '2.449', 'learning_rate': '8.511e-08', 'num_tokens': '3.421e+05', 'completions/mean_length': '24.88', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '24.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6358', 'rewards/reward_fn/std': '0.003546', 'reward': '0.6358', 'reward_std': '0.003546', 'frac_reward_zero_std': '0', 'entropy': '0.4588', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.06', 'epoch': '1.936'}
97%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹| 91/94 [04:36<00:08, 2.89s/it]03:10:49 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
03:10:50 [INFO] [reward] call=92 n=8 mean=0.639 std=0.039 min=0.599 max=0.705 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.07 task=0.04 len=0.02
98%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š| 92/94 [04:39<00:05, 2.83s/it]

{'loss': '-0.09974', 'grad_norm': '5.525', 'learning_rate': '6.383e-08', 'num_tokens': '3.456e+05', 'completions/mean_length': '15.62', 'completions/min_length': '10', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.62', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.6391', 'rewards/reward_fn/std': '0.0416', 'reward': '0.6391', 'reward_std': '0.0416', 'frac_reward_zero_std': '0', 'entropy': '0.1428', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.638', 'epoch': '1.957'}
98%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š| 92/94 [04:39<00:05, 2.83s/it]03:10:52 [INFO] [reward] call=93 n=8 mean=0.635 std=0.003 min=0.633 max=0.641 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04
{'loss': '-0.1793', 'grad_norm': '1.453', 'learning_rate': '4.255e-08', 'num_tokens': '3.475e+05', 'completions/mean_length': '20.88', 'completions/min_length': '15', 'completions/max_length': '42', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '42', 'rewards/reward_fn/mean': '0.6348', 'rewards/reward_fn/std': '0.002901', 'reward': '0.6348', 'reward_std': '0.002901', 'frac_reward_zero_std': '0.5', 'entropy': '0.2665', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.004', 'epoch': '1.979'}
99%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰| 93/94 [04:41<00:02, 2.59s/it]

99%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰| 93/94 [04:41<00:02, 2.59s/it]03:10:55 [INFO] [reward] call=94 n=8 mean=0.575 std=0.214 min=0.020 max=0.703 | parts mean: fmt=0.20 at=0.22 svc=0.06 par=0.06 task=0.04 len=0.02
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 94/94 [04:43<00:00, 2.63s/it]

{'loss': '-0.08947', 'grad_norm': '6.981', 'learning_rate': '2.128e-08', 'num_tokens': '3.527e+05', 'completions/mean_length': '16.12', 'completions/min_length': '11', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.12', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.5746', 'rewards/reward_fn/std': '0.2286', 'reward': '0.5746', 'reward_std': '0.2286', 'frac_reward_zero_std': '0', 'entropy': '0.1197', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.683', 'epoch': '2'}
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 94/94 [04:43<00:00, 2.63s/it]03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"

{'train_runtime': '284.5', 'train_samples_per_second': '0.668', 'train_steps_per_second': '0.33', 'train_loss': '-0.05751', 'epoch': '2'}
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 94/94 [04:44<00:00, 2.63s/it]
100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 94/94 [04:44<00:00, 3.03s/it]
03:10:56 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:56 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:56 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:10:56 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
03:10:58 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
03:10:58 [INFO] Saving to /tmp/sre-grpo-9zduekwb
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:10:59 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
Processing Files (0 / 0) : | | 0.00B / 0.00B
New Data Upload : | | 0.00B / 0.00B 
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB 
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB 
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB 
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB 
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB 
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB 
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB 
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB 
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB 
Processing Files (3 / 3) : 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 92.2MB / 92.2MB, 0.00B/s
New Data Upload : | | 0.00B / 0.00B, 0.00B/s
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB
03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK"
03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict"
03:10:59 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:10:59 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:10:59 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect"
03:11:00 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK"
03:11:00 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:11:00 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK"
03:11:00 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK"
03:11:00 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
Processing Files (0 / 0) : | | 0.00B / 0.00B
New Data Upload : | | 0.00B / 0.00B 
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB 
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB 
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB 
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB 
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB 
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB 
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB 
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB 
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB 
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB 
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB 
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB 
Processing Files (3 / 3) : 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 92.2MB / 92.2MB, 0.00B/s
New Data Upload : | | 0.00B / 0.00B, 0.00B/s
...zduekwb/training_args.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 7.25kB / 7.25kB
...o-9zduekwb/tokenizer.json: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 11.4MB / 11.4MB
...adapter_model.safetensors: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 80.8MB / 80.8MB
No files have been modified since last commit. Skipping to prevent empty commit.
03:11:00 [WARNING] No files have been modified since last commit. Skipping to prevent empty commit.
03:11:00 [INFO] HTTP Request: GET https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/revision/main "HTTP/1.1 200 OK"
03:11:00 [INFO] Pushed to https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo
03:11:00 [INFO] Removed local /tmp/sre-grpo-9zduekwb (--no-keep-local)
03:11:00 [INFO] Done.