Spaces:
Sleeping
Sleeping
| ===== Job started at 2026-04-26 03:05:03 ===== | |
| Downloading pygments (1.2MiB) | |
| Downloading nvidia-cufile (1.2MiB) | |
| Downloading nvidia-cufft (204.2MiB) | |
| Downloading aiohttp (1.7MiB) | |
| Downloading pandas (10.4MiB) | |
| Downloading numpy (15.9MiB) | |
| Downloading nvidia-cusparselt-cu13 (162.0MiB) | |
| Downloading networkx (2.0MiB) | |
| Downloading nvidia-cusparse (139.2MiB) | |
| Downloading hf-xet (4.0MiB) | |
| Downloading pyarrow (46.6MiB) | |
| Downloading nvidia-cuda-nvrtc (86.0MiB) | |
| Downloading nvidia-nccl-cu13 (187.4MiB) | |
| Downloading nvidia-nvshmem-cu13 (57.6MiB) | |
| Downloading cuda-bindings (6.0MiB) | |
| Downloading setuptools (1.0MiB) | |
| Downloading nvidia-nvjitlink (38.8MiB) | |
| Downloading nvidia-curand (56.8MiB) | |
| Downloading tokenizers (3.1MiB) | |
| Downloading sympy (6.0MiB) | |
| Downloading nvidia-cudnn-cu13 (349.1MiB) | |
| Downloading transformers (9.9MiB) | |
| Downloading nvidia-cublas (403.5MiB) | |
| Downloading nvidia-cuda-cupti (10.2MiB) | |
| Downloading nvidia-cuda-runtime (2.1MiB) | |
| Downloading torch (506.1MiB) | |
| Downloading nvidia-cusolver (191.6MiB) | |
| Downloading triton (179.5MiB) | |
| Downloaded nvidia-cufile | |
| Downloaded aiohttp | |
| Downloaded nvidia-cuda-runtime | |
| Downloaded pygments | |
| Downloaded tokenizers | |
| Downloaded setuptools | |
| Downloaded hf-xet | |
| Downloaded networkx | |
| Downloaded cuda-bindings | |
| Downloaded sympy | |
| Downloaded nvidia-cuda-cupti | |
| Downloaded numpy | |
| Downloaded transformers | |
| Downloaded pandas | |
| Downloaded nvidia-nvjitlink | |
| Downloaded pyarrow | |
| Downloaded nvidia-curand | |
| Downloaded nvidia-nvshmem-cu13 | |
| Downloaded nvidia-cuda-nvrtc | |
| Downloaded nvidia-cusparse | |
| Downloaded nvidia-cusparselt-cu13 | |
| Downloaded triton | |
| Downloaded nvidia-nccl-cu13 | |
| Downloaded nvidia-cusolver | |
| Downloaded nvidia-cufft | |
| Downloaded nvidia-cudnn-cu13 | |
| Downloaded nvidia-cublas | |
| Downloaded torch | |
| Installed 76 packages in 279ms | |
| 03:05:26 [INFO] Using temp --output-dir: /tmp/sre-grpo-9zduekwb | |
| 03:05:36 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/main/README.md "HTTP/1.1 404 Not Found" | |
| 03:05:36 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/srinjoyd/sre-data "HTTP/1.1 200 OK" | |
| 03:05:36 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre-data.py "HTTP/1.1 404 Not Found" | |
| 03:05:37 [INFO] HTTP Request: HEAD https://s3.amazonaws.com/datasets.huggingface.co/datasets/datasets/srinjoyd/sre-data/srinjoyd/sre-data.py "HTTP/1.1 404 Not Found" | |
| 03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/README.md "HTTP/1.1 404 Not Found" | |
| 03:05:37 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/srinjoyd/sre-data/revision/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8 "HTTP/1.1 200 OK" | |
| 03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/.huggingface.yaml "HTTP/1.1 404 Not Found" | |
| 03:05:37 [INFO] HTTP Request: GET https://huggingface.co/api/datasets/srinjoyd/sre-data/tree/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8?recursive=false&expand=false "HTTP/1.1 200 OK" | |
| 03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/dataset_infos.json "HTTP/1.1 404 Not Found" | |
| 03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/datasets/srinjoyd/sre-data/resolve/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre_raw_trajectories.jsonl "HTTP/1.1 307 Temporary Redirect" | |
| 03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/datasets/srinjoyd/sre-data/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre_raw_trajectories.jsonl "HTTP/1.1 200 OK" | |
| 03:05:37 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/datasets/srinjoyd/sre-data/2e88b9b6efe956bb3c11cbdcb192ccb9a73a97a8/sre_raw_trajectories.jsonl "HTTP/1.1 200 OK" | |
| sre_raw_trajectories.jsonl: 0%| | 0.00/8.57M [00:00<?, ?B/s][A | |
| sre_raw_trajectories.jsonl: 100%|ββββββββββ| 8.57M/8.57M [00:00<00:00, 44.3MB/s] | |
| Generating train split: 0 examples [00:00, ? examples/s][A | |
| Generating train split: 64 examples [00:00, 535.47 examples/s] | |
| 03:05:37 [INFO] Loaded 95 prompts from hf://datasets/srinjoyd/sre-data (sre_raw_trajectories.jsonl) | |
| 03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:05:37 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| config.json: 0%| | 0.00/1.37k [00:00<?, ?B/s][A | |
| config.json: 100%|ββββββββββ| 1.37k/1.37k [00:00<00:00, 12.5MB/s] | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/tokenizer_config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/tokenizer_config.json "HTTP/1.1 200 OK" | |
| 03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/tokenizer_config.json "HTTP/1.1 200 OK" | |
| tokenizer_config.json: 0%| | 0.00/844 [00:00<?, ?B/s][A | |
| tokenizer_config.json: 100%|ββββββββββ| 844/844 [00:00<00:00, 7.78MB/s] | |
| 03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged/tree/main/additional_chat_templates?recursive=false&expand=false "HTTP/1.1 404 Not Found" | |
| 03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged/tree/main?recursive=true&expand=false "HTTP/1.1 200 OK" | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/vocab.json "HTTP/1.1 404 Not Found" | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/merges.txt "HTTP/1.1 404 Not Found" | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/tokenizer.json "HTTP/1.1 302 Found" | |
| 03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged/xet-read-token/8f73acd8434aacfb77d224b6c46773bca9279705 "HTTP/1.1 200 OK" | |
| tokenizer.json: 0%| | 0.00/11.4M [00:00<?, ?B/s][A | |
| tokenizer.json: 100%|ββββββββββ| 11.4M/11.4M [00:00<00:00, 18.6MB/s] | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/added_tokens.json "HTTP/1.1 404 Not Found" | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/special_tokens_map.json "HTTP/1.1 404 Not Found" | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/chat_template.jinja "HTTP/1.1 307 Temporary Redirect" | |
| 03:05:38 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/chat_template.jinja "HTTP/1.1 200 OK" | |
| 03:05:38 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/chat_template.jinja "HTTP/1.1 200 OK" | |
| chat_template.jinja: 0%| | 0.00/2.51k [00:00<?, ?B/s][A | |
| chat_template.jinja: 100%|ββββββββββ| 2.51k/2.51k [00:00<00:00, 22.8MB/s] | |
| 03:05:39 [INFO] HTTP Request: GET https://huggingface.co/api/models/srinjoyd/qwen2.5-7b-sre-merged "HTTP/1.1 200 OK" | |
| 03:05:39 [WARNING] Dropping unsupported GRPOConfig fields for this trl build: ['max_prompt_length'] | |
| 03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/adapter_config.json "HTTP/1.1 404 Not Found" | |
| 03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:05:39 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:05:40 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/model.safetensors "HTTP/1.1 302 Found" | |
| model.safetensors: 0%| | 0.00/15.2G [00:00<?, ?B/s][A | |
| model.safetensors: 0%| | 0.00/15.2G [00:01<?, ?B/s][A | |
| model.safetensors: 0%| | 0.00/15.2G [00:02<?, ?B/s][A | |
| model.safetensors: 0%| | 67.1M/15.2G [00:03<04:30, 56.0MB/s][A | |
| model.safetensors: 3%|β | 402M/15.2G [00:04<01:11, 208MB/s] [A | |
| model.safetensors: 6%|β | 872M/15.2G [00:05<00:48, 296MB/s][A | |
| model.safetensors: 9%|β | 1.43G/15.2G [00:06<00:35, 390MB/s][A | |
| model.safetensors: 12%|ββ | 1.89G/15.2G [00:07<00:31, 417MB/s][A | |
| model.safetensors: 16%|ββ | 2.50G/15.2G [00:08<00:28, 449MB/s][A | |
| model.safetensors: 22%|βββ | 3.38G/15.2G [00:09<00:20, 583MB/s][A | |
| model.safetensors: 33%|ββββ | 4.98G/15.2G [00:10<00:12, 841MB/s][A | |
| model.safetensors: 42%|βββββ | 6.38G/15.2G [00:12<00:09, 946MB/s][A | |
| model.safetensors: 52%|ββββββ | 7.90G/15.2G [00:13<00:07, 948MB/s][A | |
| model.safetensors: 61%|ββββββ | 9.25G/15.2G [00:14<00:05, 1.00GB/s][A | |
| model.safetensors: 68%|βββββββ | 10.4G/15.2G [00:16<00:05, 931MB/s] [A | |
| model.safetensors: 79%|ββββββββ | 12.1G/15.2G [00:17<00:02, 1.08GB/s][A | |
| model.safetensors: 89%|βββββββββ | 13.5G/15.2G [00:18<00:01, 1.16GB/s][A | |
| model.safetensors: 100%|ββββββββββ| 15.2G/15.2G [00:20<00:00, 1.13GB/s][A | |
| model.safetensors: 100%|ββββββββββ| 15.2G/15.2G [00:20<00:00, 761MB/s] | |
| Loading weights: 0%| | 0/339 [00:00<?, ?it/s][A | |
| Loading weights: 0%| | 1/339 [00:01<07:16, 1.29s/it][A | |
| Loading weights: 18%|ββ | 60/339 [00:02<00:09, 28.99it/s][A | |
| Loading weights: 30%|βββ | 101/339 [00:03<00:07, 33.00it/s][A | |
| Loading weights: 43%|βββββ | 146/339 [00:04<00:05, 37.03it/s][A | |
| Loading weights: 54%|ββββββ | 184/339 [00:05<00:04, 36.57it/s][A | |
| Loading weights: 68%|βββββββ | 230/339 [00:06<00:02, 37.94it/s][A | |
| Loading weights: 79%|ββββββββ | 269/339 [00:07<00:01, 37.87it/s][A | |
| Loading weights: 93%|ββββββββββ| 314/339 [00:08<00:00, 39.23it/s][A | |
| Loading weights: 100%|ββββββββββ| 339/339 [00:09<00:00, 36.54it/s] | |
| 03:06:09 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/generation_config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:06:09 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/generation_config.json "HTTP/1.1 200 OK" | |
| 03:06:09 [INFO] HTTP Request: GET https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/generation_config.json "HTTP/1.1 200 OK" | |
| generation_config.json: 0%| | 0.00/242 [00:00<?, ?B/s][A | |
| generation_config.json: 100%|ββββββββββ| 242/242 [00:00<00:00, 2.18MB/s] | |
| 03:06:11 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 200 OK" | |
| 03:06:11 [INFO] Starting GRPO: model=srinjoyd/qwen2.5-7b-sre-merged, prompts=95, k=4, beta=0, lr=2e-06 | |
| [transformers] The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': None, 'pad_token_id': 151643}. | |
| 0%| | 0/94 [00:00<?, ?it/s][A03:06:14 [INFO] [reward] call=1 n=8 mean=0.129 std=0.181 min=0.029 max=0.599 | parts mean: fmt=0.20 at=0.05 svc=0.05 par=0.00 task=0.01 len=-0.01 | |
| 1%| | 1/94 [00:04<07:18, 4.72s/it][A | |
| [A | |
| {'loss': '-0.03945', 'grad_norm': '7.048', 'learning_rate': '2e-06', 'num_tokens': '5272', 'completions/mean_length': '12.5', 'completions/min_length': '11', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '12.5', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.1295', 'rewards/reward_fn/std': '0.1936', 'reward': '0.1295', 'reward_std': '0.1936', 'frac_reward_zero_std': '0.5', 'entropy': '0.0801', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '4.603', 'epoch': '0.02128'} | |
| 1%| | 1/94 [00:04<07:18, 4.72s/it][A03:06:17 [INFO] [reward] call=2 n=8 mean=0.634 std=0.001 min=0.633 max=0.635 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.03 | |
| {'loss': '-0.06348', 'grad_norm': '1.36', 'learning_rate': '1.979e-06', 'num_tokens': '7094', 'completions/mean_length': '15.75', 'completions/min_length': '15', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.6336', 'rewards/reward_fn/std': '0.0005657', 'reward': '0.6336', 'reward_std': '0.0005657', 'frac_reward_zero_std': '0.5', 'entropy': '0.1546', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.686', 'epoch': '0.04255'} | |
| 2%|β | 2/94 [00:07<05:24, 3.53s/it][A | |
| [A | |
| 2%|β | 2/94 [00:07<05:24, 3.53s/it][A03:06:21 [INFO] [reward] call=3 n=8 mean=0.661 std=0.028 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 3%|β | 3/94 [00:11<05:33, 3.67s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '1.957e-06', 'num_tokens': '1.064e+04', 'completions/mean_length': '16.38', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6609', 'rewards/reward_fn/std': '0.0294', 'reward': '0.6609', 'reward_std': '0.0294', 'frac_reward_zero_std': '1', 'entropy': '0.06277', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.773', 'epoch': '0.06383'} | |
| 3%|β | 3/94 [00:11<05:33, 3.67s/it][A03:06:25 [INFO] [reward] call=4 n=8 mean=0.663 std=0.026 min=0.633 max=0.689 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 4%|β | 4/94 [00:14<05:23, 3.59s/it][A | |
| [A | |
| {'loss': '-0.2514', 'grad_norm': '2.367', 'learning_rate': '1.936e-06', 'num_tokens': '1.421e+04', 'completions/mean_length': '20', 'completions/min_length': '15', 'completions/max_length': '42', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '42', 'rewards/reward_fn/mean': '0.6629', 'rewards/reward_fn/std': '0.02805', 'reward': '0.6629', 'reward_std': '0.02805', 'frac_reward_zero_std': '0', 'entropy': '0.2134', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.407', 'epoch': '0.08511'} | |
| 4%|β | 4/94 [00:14<05:23, 3.59s/it][A03:06:27 [INFO] [reward] call=5 n=8 mean=0.635 std=0.005 min=0.633 max=0.650 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 5%|β | 5/94 [00:16<04:24, 2.97s/it][A | |
| [A | |
| {'loss': '-0.09344', 'grad_norm': '1.805', 'learning_rate': '1.915e-06', 'num_tokens': '1.605e+04', 'completions/mean_length': '19.5', 'completions/min_length': '15', 'completions/max_length': '34', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '34', 'rewards/reward_fn/mean': '0.6353', 'rewards/reward_fn/std': '0.005775', 'reward': '0.6353', 'reward_std': '0.005775', 'frac_reward_zero_std': '0.5', 'entropy': '0.2227', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.862', 'epoch': '0.1064'} | |
| 5%|β | 5/94 [00:16<04:24, 2.97s/it][A03:06:31 [INFO] [reward] call=6 n=8 mean=0.616 std=0.018 min=0.598 max=0.636 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04 | |
| 6%|β | 6/94 [00:20<04:39, 3.18s/it][A | |
| [A | |
| {'loss': '-0.276', 'grad_norm': '1.836', 'learning_rate': '1.894e-06', 'num_tokens': '1.964e+04', 'completions/mean_length': '19.38', 'completions/min_length': '15', 'completions/max_length': '46', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '46', 'rewards/reward_fn/mean': '0.6161', 'rewards/reward_fn/std': '0.01908', 'reward': '0.6161', 'reward_std': '0.01908', 'frac_reward_zero_std': '0.5', 'entropy': '0.2511', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.534', 'epoch': '0.1277'} | |
| 6%|β | 6/94 [00:20<04:39, 3.18s/it][A03:06:33 [INFO] [reward] call=7 n=8 mean=0.638 std=0.007 min=0.633 max=0.651 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 7%|β | 7/94 [00:22<04:08, 2.86s/it][A | |
| [A | |
| {'loss': '-0.2179', 'grad_norm': '1.842', 'learning_rate': '1.872e-06', 'num_tokens': '2.149e+04', 'completions/mean_length': '19.75', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6375', 'rewards/reward_fn/std': '0.007678', 'reward': '0.6375', 'reward_std': '0.007678', 'frac_reward_zero_std': '0.5', 'entropy': '0.2716', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.174', 'epoch': '0.1489'} | |
| 7%|β | 7/94 [00:22<04:08, 2.86s/it][A03:06:36 [INFO] [reward] call=8 n=8 mean=0.330 std=0.476 min=-0.777 max=0.701 | parts mean: fmt=0.18 at=0.16 svc=0.05 par=0.04 task=0.03 len=0.04 | |
| 9%|β | 8/94 [00:25<04:12, 2.93s/it][A | |
| [A | |
| {'loss': '0.1133', 'grad_norm': '4.491', 'learning_rate': '1.851e-06', 'num_tokens': '2.506e+04', 'completions/mean_length': '16.5', 'completions/min_length': '11', 'completions/max_length': '30', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '30', 'rewards/reward_fn/mean': '0.33', 'rewards/reward_fn/std': '0.5085', 'reward': '0.33', 'reward_std': '0.5085', 'frac_reward_zero_std': '0', 'entropy': '0.4821', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.033', 'epoch': '0.1702'} | |
| 9%|β | 8/94 [00:25<04:12, 2.93s/it][A03:06:39 [INFO] [reward] call=9 n=8 mean=0.534 std=0.245 min=0.111 max=0.690 | parts mean: fmt=0.20 at=0.19 svc=0.05 par=0.09 task=0.03 len=0.02 | |
| 10%|β | 9/94 [00:28<04:00, 2.83s/it][A | |
| [A | |
| {'loss': '-0.04131', 'grad_norm': '3.706', 'learning_rate': '1.83e-06', 'num_tokens': '3.025e+04', 'completions/mean_length': '15.12', 'completions/min_length': '10', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.12', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.5343', 'rewards/reward_fn/std': '0.2624', 'reward': '0.5343', 'reward_std': '0.2624', 'frac_reward_zero_std': '0', 'entropy': '0.1164', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.551', 'epoch': '0.1915'} | |
| 10%|β | 9/94 [00:28<04:00, 2.83s/it][A03:06:42 [INFO] [reward] call=10 n=8 mean=0.397 std=0.269 min=0.029 max=0.688 | parts mean: fmt=0.20 at=0.14 svc=0.07 par=0.05 task=0.02 len=0.03 | |
| 11%|β | 10/94 [00:31<04:11, 2.99s/it][A | |
| [A | |
| {'loss': '0.06288', 'grad_norm': '3.008', 'learning_rate': '1.809e-06', 'num_tokens': '3.543e+04', 'completions/mean_length': '20.25', 'completions/min_length': '11', 'completions/max_length': '40', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.25', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '40', 'rewards/reward_fn/mean': '0.397', 'rewards/reward_fn/std': '0.2871', 'reward': '0.397', 'reward_std': '0.2871', 'frac_reward_zero_std': '0', 'entropy': '0.3148', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.291', 'epoch': '0.2128'} | |
| 11%|β | 10/94 [00:31<04:11, 2.99s/it][A03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:06:43 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:06:43 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:06:43 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:06:43 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:06:43 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:06:43 [INFO] HTTP Request: GET https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/xet-write-token/main "HTTP/1.1 200 OK" | |
| 03:06:45 [INFO] [reward] call=11 n=8 mean=0.638 std=0.034 min=0.598 max=0.699 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.08 task=0.04 len=0.01 | |
| 03:06:46 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 12%|ββ | 11/94 [00:34<04:13, 3.05s/it][A | |
| {'loss': '-0.1043', 'grad_norm': '5.863', 'learning_rate': '1.787e-06', 'num_tokens': '3.893e+04', 'completions/mean_length': '14.62', 'completions/min_length': '10', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '14.62', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6384', 'rewards/reward_fn/std': '0.03669', 'reward': '0.6384', 'reward_std': '0.03669', 'frac_reward_zero_std': '0', 'entropy': '0.1577', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.613', 'epoch': '0.234'} | |
| [A | |
| 12%|ββ | 11/94 [00:34<04:13, 3.05s/it][A03:06:47 [INFO] [reward] call=12 n=8 mean=0.633 std=0.001 min=0.630 max=0.633 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.03 | |
| 13%|ββ | 12/94 [00:36<03:32, 2.59s/it][A | |
| [A | |
| {'loss': '0.07766', 'grad_norm': '1.3', 'learning_rate': '1.766e-06', 'num_tokens': '4.074e+04', 'completions/mean_length': '15.88', 'completions/min_length': '15', 'completions/max_length': '22', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '22', 'rewards/reward_fn/mean': '0.633', 'rewards/reward_fn/std': '0.001096', 'reward': '0.633', 'reward_std': '0.001096', 'frac_reward_zero_std': '0.5', 'entropy': '0.2152', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.541', 'epoch': '0.2553'} | |
| 13%|ββ | 12/94 [00:36<03:32, 2.59s/it][A03:06:49 [INFO] [reward] call=13 n=8 mean=0.264 std=0.271 min=0.020 max=0.611 | parts mean: fmt=0.20 at=0.11 svc=0.05 par=0.04 task=0.01 len=-0.03 | |
| 14%|ββ | 13/94 [00:38<03:27, 2.56s/it][A | |
| [A | |
| {'loss': '0.03801', 'grad_norm': '8.365', 'learning_rate': '1.745e-06', 'num_tokens': '4.595e+04', 'completions/mean_length': '10.62', 'completions/min_length': '10', 'completions/max_length': '11', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '10.62', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '11', 'rewards/reward_fn/mean': '0.2636', 'rewards/reward_fn/std': '0.2897', 'reward': '0.2636', 'reward_std': '0.2897', 'frac_reward_zero_std': '0', 'entropy': '0.09156', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.426', 'epoch': '0.2766'} | |
| 14%|ββ | 13/94 [00:38<03:27, 2.56s/it][A03:06:52 [INFO] [reward] call=14 n=8 mean=0.642 std=0.032 min=0.598 max=0.695 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.03 | |
| 15%|ββ | 14/94 [00:41<03:39, 2.75s/it][A | |
| [A | |
| {'loss': '-0.2787', 'grad_norm': '5.452', 'learning_rate': '1.723e-06', 'num_tokens': '4.955e+04', 'completions/mean_length': '18.38', 'completions/min_length': '10', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.38', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6422', 'rewards/reward_fn/std': '0.0339', 'reward': '0.6422', 'reward_std': '0.0339', 'frac_reward_zero_std': '0', 'entropy': '0.3664', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.117', 'epoch': '0.2979'} | |
| 15%|ββ | 14/94 [00:41<03:39, 2.75s/it][A03:06:55 [INFO] [reward] call=15 n=8 mean=0.548 std=0.200 min=0.020 max=0.634 | parts mean: fmt=0.20 at=0.22 svc=0.06 par=0.03 task=0.04 len=0.02 | |
| 16%|ββ | 15/94 [00:44<03:45, 2.86s/it][A | |
| [A | |
| {'loss': '-0.1453', 'grad_norm': '3.979', 'learning_rate': '1.702e-06', 'num_tokens': '5.307e+04', 'completions/mean_length': '16.5', 'completions/min_length': '11', 'completions/max_length': '28', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '28', 'rewards/reward_fn/mean': '0.5483', 'rewards/reward_fn/std': '0.2137', 'reward': '0.5483', 'reward_std': '0.2137', 'frac_reward_zero_std': '0', 'entropy': '0.2527', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.047', 'epoch': '0.3191'} | |
| 16%|ββ | 15/94 [00:44<03:45, 2.86s/it][A03:06:59 [INFO] [reward] call=16 n=8 mean=0.633 std=0.026 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.02 | |
| 17%|ββ | 16/94 [00:48<03:51, 2.97s/it][A | |
| {'loss': '-0.1697', 'grad_norm': '4.519', 'learning_rate': '1.681e-06', 'num_tokens': '5.657e+04', 'completions/mean_length': '17.88', 'completions/min_length': '10', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.88', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6331', 'rewards/reward_fn/std': '0.02809', 'reward': '0.6331', 'reward_std': '0.02809', 'frac_reward_zero_std': '0', 'entropy': '0.817', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.173', 'epoch': '0.3404'} | |
| [A | |
| 17%|ββ | 16/94 [00:48<03:51, 2.97s/it][A03:07:01 [INFO] [reward] call=17 n=8 mean=0.661 std=0.028 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 18%|ββ | 17/94 [00:50<03:40, 2.87s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '1.66e-06', 'num_tokens': '6.006e+04', 'completions/mean_length': '16.5', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6609', 'rewards/reward_fn/std': '0.0294', 'reward': '0.6609', 'reward_std': '0.0294', 'frac_reward_zero_std': '1', 'entropy': '0.101', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.562', 'epoch': '0.3617'} | |
| 18%|ββ | 17/94 [00:50<03:40, 2.87s/it][A03:07:04 [INFO] [reward] call=18 n=8 mean=0.617 std=0.019 min=0.598 max=0.642 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04 | |
| 19%|ββ | 18/94 [00:53<03:35, 2.83s/it][A | |
| [A | |
| {'loss': '-0.02915', 'grad_norm': '1.632', 'learning_rate': '1.638e-06', 'num_tokens': '6.363e+04', 'completions/mean_length': '16.75', 'completions/min_length': '15', 'completions/max_length': '20', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '20', 'rewards/reward_fn/mean': '0.6168', 'rewards/reward_fn/std': '0.01994', 'reward': '0.6168', 'reward_std': '0.01994', 'frac_reward_zero_std': '0.5', 'entropy': '0.2609', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.691', 'epoch': '0.383'} | |
| 19%|ββ | 18/94 [00:53<03:35, 2.83s/it][A03:07:07 [INFO] [reward] call=19 n=8 mean=0.636 std=0.006 min=0.632 max=0.649 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 20%|ββ | 19/94 [00:55<03:20, 2.67s/it][A | |
| [A | |
| {'loss': '-0.2383', 'grad_norm': '1.594', 'learning_rate': '1.617e-06', 'num_tokens': '6.549e+04', 'completions/mean_length': '21.12', 'completions/min_length': '15', 'completions/max_length': '49', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '49', 'rewards/reward_fn/mean': '0.6361', 'rewards/reward_fn/std': '0.005971', 'reward': '0.6361', 'reward_std': '0.005971', 'frac_reward_zero_std': '0.5', 'entropy': '0.5286', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.292', 'epoch': '0.4043'} | |
| 20%|ββ | 19/94 [00:55<03:20, 2.67s/it][A03:07:09 [INFO] [reward] call=20 n=8 mean=0.649 std=0.032 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| {'loss': '0.08051', 'grad_norm': '4.462', 'learning_rate': '1.596e-06', 'num_tokens': '6.898e+04', 'completions/mean_length': '17.5', 'completions/min_length': '15', 'completions/max_length': '24', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '24', 'rewards/reward_fn/mean': '0.6493', 'rewards/reward_fn/std': '0.03433', 'reward': '0.6493', 'reward_std': '0.03433', 'frac_reward_zero_std': '0', 'entropy': '0.2607', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.756', 'epoch': '0.4255'} | |
| 21%|βββ | 20/94 [00:58<03:20, 2.71s/it][A | |
| [A | |
| 21%|βββ | 20/94 [00:58<03:20, 2.71s/it][A03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:07:10 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:07:10 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:07:11 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:07:11 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:07:11 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:07:13 [INFO] [reward] call=21 n=8 mean=0.495 std=0.224 min=0.106 max=0.637 | parts mean: fmt=0.20 at=0.19 svc=0.05 par=0.04 task=0.03 len=0.02 | |
| 03:07:13 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 22%|βββ | 21/94 [01:01<03:29, 2.87s/it][A | |
| [A | |
| {'loss': '-0.04748', 'grad_norm': '3.59', 'learning_rate': '1.574e-06', 'num_tokens': '7.249e+04', 'completions/mean_length': '15.12', 'completions/min_length': '10', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.12', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.4951', 'rewards/reward_fn/std': '0.2398', 'reward': '0.4951', 'reward_std': '0.2398', 'frac_reward_zero_std': '0', 'entropy': '0.1683', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.67', 'epoch': '0.4468'} | |
| 22%|βββ | 21/94 [01:01<03:29, 2.87s/it][A03:07:15 [INFO] [reward] call=22 n=8 mean=0.643 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04 | |
| 23%|βββ | 22/94 [01:04<03:20, 2.79s/it][A | |
| [A | |
| {'loss': '-0.006488', 'grad_norm': '1.562', 'learning_rate': '1.553e-06', 'num_tokens': '7.775e+04', 'completions/mean_length': '16.75', 'completions/min_length': '16', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.75', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6434', 'rewards/reward_fn/std': '0.04772', 'reward': '0.6434', 'reward_std': '0.04772', 'frac_reward_zero_std': '0.5', 'entropy': '0.08012', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.534', 'epoch': '0.4681'} | |
| 23%|βββ | 22/94 [01:04<03:20, 2.79s/it][A03:07:18 [INFO] [reward] call=23 n=8 mean=0.575 std=0.181 min=0.109 max=0.692 | parts mean: fmt=0.20 at=0.22 svc=0.06 par=0.06 task=0.04 len=0.02 | |
| {'loss': '-0.04835', 'grad_norm': '4.127', 'learning_rate': '1.532e-06', 'num_tokens': '8.303e+04', 'completions/mean_length': '15', 'completions/min_length': '10', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.575', 'rewards/reward_fn/std': '0.1932', 'reward': '0.575', 'reward_std': '0.1932', 'frac_reward_zero_std': '0', 'entropy': '0.1359', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.625', 'epoch': '0.4894'} | |
| 24%|βββ | 23/94 [01:07<03:15, 2.76s/it][A | |
| [A | |
| 24%|βββ | 23/94 [01:07<03:15, 2.76s/it][A03:07:21 [INFO] [reward] call=24 n=8 mean=0.673 std=0.037 min=0.611 max=0.712 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.09 task=0.04 len=0.04 | |
| 26%|βββ | 24/94 [01:10<03:32, 3.03s/it][A | |
| [A | |
| {'loss': '-0.203', 'grad_norm': '2.814', 'learning_rate': '1.511e-06', 'num_tokens': '8.827e+04', 'completions/mean_length': '21', 'completions/min_length': '16', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6728', 'rewards/reward_fn/std': '0.03925', 'reward': '0.6728', 'reward_std': '0.03925', 'frac_reward_zero_std': '0', 'entropy': '0.1496', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.596', 'epoch': '0.5106'} | |
| 26%|βββ | 24/94 [01:10<03:32, 3.03s/it][A03:07:25 [INFO] [reward] call=25 n=8 mean=0.661 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 27%|βββ | 25/94 [01:14<03:31, 3.07s/it][A | |
| {'loss': '-0.1814', 'grad_norm': '2.257', 'learning_rate': '1.489e-06', 'num_tokens': '9.183e+04', 'completions/mean_length': '18.5', 'completions/min_length': '15', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6614', 'rewards/reward_fn/std': '0.0286', 'reward': '0.6614', 'reward_std': '0.0286', 'frac_reward_zero_std': '0', 'entropy': '0.1536', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.107', 'epoch': '0.5319'} | |
| [A | |
| 27%|βββ | 25/94 [01:14<03:31, 3.07s/it][A03:07:27 [INFO] [reward] call=26 n=8 mean=0.561 std=0.175 min=0.109 max=0.690 | parts mean: fmt=0.20 at=0.22 svc=0.07 par=0.04 task=0.04 len=0.02 | |
| 28%|βββ | 26/94 [01:16<03:17, 2.91s/it][A | |
| [A | |
| {'loss': '-0.1144', 'grad_norm': '7.925', 'learning_rate': '1.468e-06', 'num_tokens': '9.704e+04', 'completions/mean_length': '14.88', 'completions/min_length': '10', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '14.88', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.5614', 'rewards/reward_fn/std': '0.1872', 'reward': '0.5614', 'reward_std': '0.1872', 'frac_reward_zero_std': '0', 'entropy': '0.1353', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.467', 'epoch': '0.5532'} | |
| 28%|βββ | 26/94 [01:16<03:17, 2.91s/it][A03:07:30 [INFO] [reward] call=27 n=8 mean=0.661 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.03 | |
| 29%|βββ | 27/94 [01:19<03:08, 2.81s/it][A | |
| [A | |
| {'loss': '-0.00338', 'grad_norm': '1.348', 'learning_rate': '1.447e-06', 'num_tokens': '1.006e+05', 'completions/mean_length': '16.12', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6607', 'rewards/reward_fn/std': '0.02923', 'reward': '0.6607', 'reward_std': '0.02923', 'frac_reward_zero_std': '0.5', 'entropy': '0.08173', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.531', 'epoch': '0.5745'} | |
| 29%|βββ | 27/94 [01:19<03:08, 2.81s/it][A03:07:32 [INFO] [reward] call=28 n=8 mean=0.644 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04 | |
| 30%|βββ | 28/94 [01:21<02:59, 2.72s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '1.426e-06', 'num_tokens': '1.058e+05', 'completions/mean_length': '16.5', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6436', 'rewards/reward_fn/std': '0.04789', 'reward': '0.6436', 'reward_std': '0.04789', 'frac_reward_zero_std': '1', 'entropy': '0.05464', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.458', 'epoch': '0.5957'} | |
| 30%|βββ | 28/94 [01:21<02:59, 2.72s/it][A03:07:36 [INFO] [reward] call=29 n=8 mean=0.619 std=0.021 min=0.598 max=0.646 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04 | |
| 31%|βββ | 29/94 [01:25<03:11, 2.95s/it][A | |
| [A | |
| {'loss': '-0.1543', 'grad_norm': '1.539', 'learning_rate': '1.404e-06', 'num_tokens': '1.094e+05', 'completions/mean_length': '18.88', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6191', 'rewards/reward_fn/std': '0.02274', 'reward': '0.6191', 'reward_std': '0.02274', 'frac_reward_zero_std': '0.5', 'entropy': '0.3003', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.422', 'epoch': '0.617'} | |
| 31%|βββ | 29/94 [01:25<03:11, 2.95s/it][A03:07:38 [INFO] [reward] call=30 n=8 mean=0.682 std=0.032 min=0.598 max=0.705 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04 | |
| 32%|ββββ | 30/94 [01:29<03:27, 3.24s/it][A | |
| [A | |
| {'loss': '-0.02124', 'grad_norm': '5.719', 'learning_rate': '1.383e-06', 'num_tokens': '1.147e+05', 'completions/mean_length': '17.12', 'completions/min_length': '16', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.12', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6819', 'rewards/reward_fn/std': '0.0343', 'reward': '0.6819', 'reward_std': '0.0343', 'frac_reward_zero_std': '0', 'entropy': '0.2194', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.835', 'epoch': '0.6383'} | |
| 32%|ββββ | 30/94 [01:29<03:27, 3.24s/it][A03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:07:40 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:07:41 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:07:41 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:07:41 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:07:41 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:07:43 [INFO] [reward] call=31 n=8 mean=0.663 std=0.026 min=0.631 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| {'loss': '-0.01252', 'grad_norm': '1.514', 'learning_rate': '1.362e-06', 'num_tokens': '1.182e+05', 'completions/mean_length': '22.25', 'completions/min_length': '15', 'completions/max_length': '32', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '32', 'rewards/reward_fn/mean': '0.6626', 'rewards/reward_fn/std': '0.02803', 'reward': '0.6626', 'reward_std': '0.02803', 'frac_reward_zero_std': '0.5', 'entropy': '0.5726', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.906', 'epoch': '0.6596'} | |
| 33%|ββββ | 31/94 [01:32<03:27, 3.29s/it][A | |
| [A | |
| 33%|ββββ | 31/94 [01:32<03:27, 3.29s/it][A03:07:46 [INFO] [reward] call=32 n=8 mean=0.616 std=0.018 min=0.598 max=0.633 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.03 | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '1.34e-06', 'num_tokens': '1.218e+05', 'completions/mean_length': '15.5', 'completions/min_length': '15', 'completions/max_length': '16', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '16', 'rewards/reward_fn/mean': '0.6158', 'rewards/reward_fn/std': '0.01876', 'reward': '0.6158', 'reward_std': '0.01876', 'frac_reward_zero_std': '1', 'entropy': '0.03686', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.571', 'epoch': '0.6809'} | |
| 34%|ββββ | 32/94 [01:35<03:11, 3.09s/it][A | |
| [A | |
| 34%|ββββ | 32/94 [01:35<03:11, 3.09s/it][A03:07:48 [INFO] [reward] call=33 n=8 mean=0.633 std=0.044 min=0.598 max=0.690 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.04 task=0.04 len=0.04 | |
| {'loss': '-0.01142', 'grad_norm': '4.521', 'learning_rate': '1.319e-06', 'num_tokens': '1.27e+05', 'completions/mean_length': '16.38', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.38', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6325', 'rewards/reward_fn/std': '0.04669', 'reward': '0.6325', 'reward_std': '0.04669', 'frac_reward_zero_std': '0.5', 'entropy': '0.1329', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.463', 'epoch': '0.7021'} | |
| 35%|ββββ | 33/94 [01:37<02:58, 2.92s/it][A | |
| [A | |
| 35%|ββββ | 33/94 [01:37<02:58, 2.92s/it][A03:07:49 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 03:07:52 [INFO] [reward] call=34 n=8 mean=0.656 std=0.030 min=0.611 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| 36%|ββββ | 34/94 [01:41<03:13, 3.23s/it][A | |
| [A | |
| {'loss': '-0.3476', 'grad_norm': '2.982', 'learning_rate': '1.298e-06', 'num_tokens': '1.306e+05', 'completions/mean_length': '22.5', 'completions/min_length': '15', 'completions/max_length': '57', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '57', 'rewards/reward_fn/mean': '0.6558', 'rewards/reward_fn/std': '0.0324', 'reward': '0.6558', 'reward_std': '0.0324', 'frac_reward_zero_std': '0', 'entropy': '0.6957', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.883', 'epoch': '0.7234'} | |
| 36%|ββββ | 34/94 [01:41<03:13, 3.23s/it][A03:07:55 [INFO] [reward] call=35 n=8 mean=0.660 std=0.020 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04 | |
| 37%|ββββ | 35/94 [01:44<03:07, 3.18s/it][A | |
| [A | |
| {'loss': '-0.07906', 'grad_norm': '2.936', 'learning_rate': '1.277e-06', 'num_tokens': '1.341e+05', 'completions/mean_length': '21.88', 'completions/min_length': '15', 'completions/max_length': '36', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '36', 'rewards/reward_fn/mean': '0.6599', 'rewards/reward_fn/std': '0.02178', 'reward': '0.6599', 'reward_std': '0.02178', 'frac_reward_zero_std': '0', 'entropy': '0.3931', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.998', 'epoch': '0.7447'} | |
| 37%|ββββ | 35/94 [01:44<03:07, 3.18s/it][A03:07:57 [INFO] [reward] call=36 n=8 mean=0.459 std=0.466 min=-0.774 max=0.639 | parts mean: fmt=0.18 at=0.22 svc=0.04 par=0.05 task=0.04 len=0.04 | |
| 38%|ββββ | 36/94 [01:46<02:43, 2.82s/it][A | |
| [A | |
| {'loss': '-0.09559', 'grad_norm': '2.419', 'learning_rate': '1.255e-06', 'num_tokens': '1.36e+05', 'completions/mean_length': '21.62', 'completions/min_length': '15', 'completions/max_length': '39', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.62', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '39', 'rewards/reward_fn/mean': '0.4588', 'rewards/reward_fn/std': '0.4983', 'reward': '0.4588', 'reward_std': '0.4983', 'frac_reward_zero_std': '0', 'entropy': '0.8276', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.981', 'epoch': '0.766'} | |
| 38%|ββββ | 36/94 [01:46<02:43, 2.82s/it][A03:08:00 [INFO] [reward] call=37 n=8 mean=0.689 std=0.002 min=0.688 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04 | |
| 39%|ββββ | 37/94 [01:49<02:37, 2.76s/it][A | |
| [A | |
| {'loss': '1.49e-08', 'grad_norm': '1.85', 'learning_rate': '1.234e-06', 'num_tokens': '1.413e+05', 'completions/mean_length': '17.5', 'completions/min_length': '17', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.5', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6893', 'rewards/reward_fn/std': '0.002546', 'reward': '0.6893', 'reward_std': '0.002546', 'frac_reward_zero_std': '0.5', 'entropy': '0.1432', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.549', 'epoch': '0.7872'} | |
| 39%|ββββ | 37/94 [01:49<02:37, 2.76s/it][A03:08:02 [INFO] [reward] call=38 n=8 mean=0.636 std=0.005 min=0.633 max=0.647 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 40%|ββββ | 38/94 [01:51<02:20, 2.50s/it][A | |
| [A | |
| {'loss': '-0.1723', 'grad_norm': '1.639', 'learning_rate': '1.213e-06', 'num_tokens': '1.432e+05', 'completions/mean_length': '18.25', 'completions/min_length': '15', 'completions/max_length': '35', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '35', 'rewards/reward_fn/mean': '0.6364', 'rewards/reward_fn/std': '0.005656', 'reward': '0.6364', 'reward_std': '0.005656', 'frac_reward_zero_std': '0.5', 'entropy': '0.2606', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.895', 'epoch': '0.8085'} | |
| 40%|ββββ | 38/94 [01:51<02:20, 2.50s/it][A03:08:04 [INFO] [reward] call=39 n=8 mean=0.635 std=0.004 min=0.633 max=0.645 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 41%|βββββ | 39/94 [01:52<02:06, 2.30s/it][A | |
| [A | |
| {'loss': '-0.1923', 'grad_norm': '2.579', 'learning_rate': '1.191e-06', 'num_tokens': '1.45e+05', 'completions/mean_length': '17.25', 'completions/min_length': '15', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6348', 'rewards/reward_fn/std': '0.003995', 'reward': '0.6348', 'reward_std': '0.003995', 'frac_reward_zero_std': '0.5', 'entropy': '0.3042', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.817', 'epoch': '0.8298'} | |
| 41%|βββββ | 39/94 [01:52<02:06, 2.30s/it][A03:08:07 [INFO] [reward] call=40 n=8 mean=0.645 std=0.045 min=0.599 max=0.704 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.08 task=0.04 len=0.02 | |
| 43%|βββββ | 40/94 [01:56<02:26, 2.71s/it][A | |
| [A | |
| {'loss': '0.3164', 'grad_norm': '4.612', 'learning_rate': '1.17e-06', 'num_tokens': '1.486e+05', 'completions/mean_length': '20.38', 'completions/min_length': '10', 'completions/max_length': '49', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.38', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '49', 'rewards/reward_fn/mean': '0.6452', 'rewards/reward_fn/std': '0.04773', 'reward': '0.6452', 'reward_std': '0.04773', 'frac_reward_zero_std': '0', 'entropy': '0.9359', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.619', 'epoch': '0.8511'} | |
| 43%|βββββ | 40/94 [01:56<02:26, 2.71s/it][A03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:08:08 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:08:08 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:08:08 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:08:09 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:08:09 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:08:11 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 03:08:11 [INFO] [reward] call=41 n=8 mean=0.663 std=0.026 min=0.631 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 44%|βββββ | 41/94 [02:00<02:50, 3.21s/it][A | |
| [A | |
| {'loss': '-0.2084', 'grad_norm': '1.419', 'learning_rate': '1.149e-06', 'num_tokens': '1.521e+05', 'completions/mean_length': '24.12', 'completions/min_length': '15', 'completions/max_length': '52', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '24.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '52', 'rewards/reward_fn/mean': '0.6634', 'rewards/reward_fn/std': '0.02773', 'reward': '0.6634', 'reward_std': '0.02773', 'frac_reward_zero_std': '0.5', 'entropy': '0.2573', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.818', 'epoch': '0.8723'} | |
| 44%|βββββ | 41/94 [02:00<02:50, 3.21s/it][A03:08:14 [INFO] [reward] call=42 n=8 mean=0.644 std=0.033 min=0.611 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.06 task=0.04 len=0.03 | |
| 45%|βββββ | 42/94 [02:03<02:38, 3.05s/it][A | |
| [A | |
| {'loss': '-0.02793', 'grad_norm': '7.134', 'learning_rate': '1.128e-06', 'num_tokens': '1.556e+05', 'completions/mean_length': '16', 'completions/min_length': '15', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6443', 'rewards/reward_fn/std': '0.03566', 'reward': '0.6443', 'reward_std': '0.03566', 'frac_reward_zero_std': '0.5', 'entropy': '0.1252', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.611', 'epoch': '0.8936'} | |
| 45%|βββββ | 42/94 [02:03<02:38, 3.05s/it][A03:08:17 [INFO] [reward] call=43 n=8 mean=0.655 std=0.044 min=0.598 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.07 task=0.04 len=0.04 | |
| 46%|βββββ | 43/94 [02:06<02:29, 2.94s/it][A | |
| [A | |
| {'loss': '-0.01125', 'grad_norm': '3.344', 'learning_rate': '1.106e-06', 'num_tokens': '1.608e+05', 'completions/mean_length': '16.62', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.62', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6546', 'rewards/reward_fn/std': '0.04663', 'reward': '0.6546', 'reward_std': '0.04663', 'frac_reward_zero_std': '0.5', 'entropy': '0.04298', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.612', 'epoch': '0.9149'} | |
| 46%|βββββ | 43/94 [02:06<02:29, 2.94s/it][A03:08:22 [INFO] [reward] call=44 n=8 mean=0.401 std=0.479 min=-0.774 max=0.688 | parts mean: fmt=0.18 at=0.19 svc=0.05 par=0.06 task=0.03 len=0.03 | |
| 47%|βββββ | 44/94 [02:11<03:03, 3.66s/it][A | |
| [A | |
| {'loss': '0.5893', 'grad_norm': '5.321', 'learning_rate': '1.085e-06', 'num_tokens': '1.662e+05', 'completions/mean_length': '25.5', 'completions/min_length': '10', 'completions/max_length': '100', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '25.5', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '100', 'rewards/reward_fn/mean': '0.4009', 'rewards/reward_fn/std': '0.5121', 'reward': '0.4009', 'reward_std': '0.5121', 'frac_reward_zero_std': '0', 'entropy': '0.3116', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '5.283', 'epoch': '0.9362'} | |
| 47%|βββββ | 44/94 [02:11<03:03, 3.66s/it][A03:08:26 [INFO] [reward] call=45 n=8 mean=0.631 std=0.030 min=0.598 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.04 task=0.04 len=0.04 | |
| {'loss': '-0.1902', 'grad_norm': '3.174', 'learning_rate': '1.064e-06', 'num_tokens': '1.698e+05', 'completions/mean_length': '25.12', 'completions/min_length': '15', 'completions/max_length': '51', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '25.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '51', 'rewards/reward_fn/mean': '0.6312', 'rewards/reward_fn/std': '0.03228', 'reward': '0.6312', 'reward_std': '0.03228', 'frac_reward_zero_std': '0', 'entropy': '0.3894', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.692', 'epoch': '0.9574'} | |
| 48%|βββββ | 45/94 [02:15<03:00, 3.69s/it][A | |
| [A | |
| 48%|βββββ | 45/94 [02:15<03:00, 3.69s/it][A03:08:28 [INFO] [reward] call=46 n=8 mean=0.636 std=0.006 min=0.632 max=0.647 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 49%|βββββ | 46/94 [02:17<02:31, 3.16s/it][A | |
| [A | |
| {'loss': '-0.3453', 'grad_norm': '2.959', 'learning_rate': '1.043e-06', 'num_tokens': '1.717e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6364', 'rewards/reward_fn/std': '0.006058', 'reward': '0.6364', 'reward_std': '0.006058', 'frac_reward_zero_std': '0', 'entropy': '0.3825', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.92', 'epoch': '0.9787'} | |
| 49%|βββββ | 46/94 [02:17<02:31, 3.16s/it][A03:08:31 [INFO] [reward] call=47 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04 | |
| 50%|βββββ | 47/94 [02:20<02:21, 3.01s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '1.021e-06', 'num_tokens': '1.769e+05', 'completions/mean_length': '17', 'completions/min_length': '17', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.1325', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.599', 'epoch': '1'} | |
| 50%|βββββ | 47/94 [02:20<02:21, 3.01s/it][A03:08:34 [INFO] [reward] call=48 n=8 mean=0.642 std=0.033 min=0.598 max=0.690 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.06 task=0.04 len=0.04 | |
| 51%|βββββ | 48/94 [02:23<02:20, 3.05s/it][A | |
| [A | |
| {'loss': '-0.1938', 'grad_norm': '5.199', 'learning_rate': '1e-06', 'num_tokens': '1.804e+05', 'completions/mean_length': '19.5', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6424', 'rewards/reward_fn/std': '0.03487', 'reward': '0.6424', 'reward_std': '0.03487', 'frac_reward_zero_std': '0', 'entropy': '0.5192', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.071', 'epoch': '1.021'} | |
| 51%|βββββ | 48/94 [02:23<02:20, 3.05s/it][A03:08:37 [INFO] [reward] call=49 n=8 mean=0.671 std=0.026 min=0.633 max=0.701 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04 | |
| 52%|ββββββ | 49/94 [02:26<02:21, 3.15s/it][A | |
| [A | |
| {'loss': '-0.008332', 'grad_norm': '1.737', 'learning_rate': '9.787e-07', 'num_tokens': '1.839e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6713', 'rewards/reward_fn/std': '0.02778', 'reward': '0.6713', 'reward_std': '0.02778', 'frac_reward_zero_std': '0.5', 'entropy': '0.3038', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.344', 'epoch': '1.043'} | |
| 52%|ββββββ | 49/94 [02:26<02:21, 3.15s/it][A03:08:41 [INFO] [reward] call=50 n=8 mean=0.621 std=0.023 min=0.599 max=0.655 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04 | |
| 53%|ββββββ | 50/94 [02:30<02:29, 3.40s/it][A | |
| [A | |
| {'loss': '-0.1466', 'grad_norm': '2.019', 'learning_rate': '9.574e-07', 'num_tokens': '1.875e+05', 'completions/mean_length': '27.5', 'completions/min_length': '16', 'completions/max_length': '60', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '27.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '60', 'rewards/reward_fn/mean': '0.6212', 'rewards/reward_fn/std': '0.02489', 'reward': '0.6212', 'reward_std': '0.02489', 'frac_reward_zero_std': '0.5', 'entropy': '0.6976', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.911', 'epoch': '1.064'} | |
| 53%|ββββββ | 50/94 [02:30<02:29, 3.40s/it][A03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:08:42 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:08:42 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:08:42 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:08:42 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:08:43 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:08:44 [INFO] [reward] call=51 n=8 mean=0.646 std=0.042 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.09 task=0.04 len=0.01 | |
| 03:08:45 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 54%|ββββββ | 51/94 [02:33<02:21, 3.30s/it][A | |
| [A | |
| {'loss': '0.08536', 'grad_norm': '5.356', 'learning_rate': '9.362e-07', 'num_tokens': '1.929e+05', 'completions/mean_length': '15', 'completions/min_length': '10', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6463', 'rewards/reward_fn/std': '0.04467', 'reward': '0.6463', 'reward_std': '0.04467', 'frac_reward_zero_std': '0', 'entropy': '0.1691', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.517', 'epoch': '1.085'} | |
| 54%|ββββββ | 51/94 [02:33<02:21, 3.30s/it][A03:08:48 [INFO] [reward] call=52 n=8 mean=0.662 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 55%|ββββββ | 52/94 [02:37<02:20, 3.34s/it][A | |
| {'loss': '-0.2974', 'grad_norm': '1.814', 'learning_rate': '9.149e-07', 'num_tokens': '1.963e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6616', 'rewards/reward_fn/std': '0.02868', 'reward': '0.6616', 'reward_std': '0.02868', 'frac_reward_zero_std': '0.5', 'entropy': '0.1737', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.375', 'epoch': '1.106'} | |
| [A | |
| 55%|ββββββ | 52/94 [02:37<02:20, 3.34s/it][A03:08:51 [INFO] [reward] call=53 n=8 mean=0.565 std=0.152 min=0.165 max=0.633 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.07 task=0.04 len=0.01 | |
| 56%|ββββββ | 53/94 [02:40<02:13, 3.26s/it][A | |
| [A | |
| {'loss': '0.1845', 'grad_norm': '3.084', 'learning_rate': '8.936e-07', 'num_tokens': '1.999e+05', 'completions/mean_length': '15.38', 'completions/min_length': '10', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.38', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.5648', 'rewards/reward_fn/std': '0.1621', 'reward': '0.5648', 'reward_std': '0.1621', 'frac_reward_zero_std': '0.5', 'entropy': '0.1412', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.009', 'epoch': '1.128'} | |
| 56%|ββββββ | 53/94 [02:40<02:13, 3.26s/it][A03:08:54 [INFO] [reward] call=54 n=8 mean=0.637 std=0.038 min=0.598 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.05 task=0.04 len=0.04 | |
| 57%|ββββββ | 54/94 [02:43<02:07, 3.18s/it][A | |
| [A | |
| {'loss': '-0.01893', 'grad_norm': '3.71', 'learning_rate': '8.723e-07', 'num_tokens': '2.035e+05', 'completions/mean_length': '17.75', 'completions/min_length': '15', 'completions/max_length': '28', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '28', 'rewards/reward_fn/mean': '0.6366', 'rewards/reward_fn/std': '0.04055', 'reward': '0.6366', 'reward_std': '0.04055', 'frac_reward_zero_std': '0', 'entropy': '0.249', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.946', 'epoch': '1.149'} | |
| 57%|ββββββ | 54/94 [02:43<02:07, 3.18s/it][A03:08:57 [INFO] [reward] call=55 n=8 mean=0.669 std=0.027 min=0.633 max=0.694 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04 | |
| 59%|ββββββ | 55/94 [02:46<02:02, 3.15s/it][A | |
| [A | |
| {'loss': '0.02042', 'grad_norm': '1.908', 'learning_rate': '8.511e-07', 'num_tokens': '2.07e+05', 'completions/mean_length': '19.5', 'completions/min_length': '15', 'completions/max_length': '35', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '35', 'rewards/reward_fn/mean': '0.6687', 'rewards/reward_fn/std': '0.02878', 'reward': '0.6687', 'reward_std': '0.02878', 'frac_reward_zero_std': '0.5', 'entropy': '0.2068', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.011', 'epoch': '1.17'} | |
| 59%|ββββββ | 55/94 [02:46<02:02, 3.15s/it][A03:09:01 [INFO] [reward] call=56 n=8 mean=0.672 std=0.026 min=0.633 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04 | |
| {'loss': '-0.05793', 'grad_norm': '2.046', 'learning_rate': '8.298e-07', 'num_tokens': '2.106e+05', 'completions/mean_length': '22.25', 'completions/min_length': '15', 'completions/max_length': '53', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '53', 'rewards/reward_fn/mean': '0.6717', 'rewards/reward_fn/std': '0.02801', 'reward': '0.6717', 'reward_std': '0.02801', 'frac_reward_zero_std': '0.5', 'entropy': '0.4402', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.752', 'epoch': '1.191'} | |
| 60%|ββββββ | 56/94 [02:49<02:07, 3.35s/it][A | |
| [A | |
| 60%|ββββββ | 56/94 [02:49<02:07, 3.35s/it][A03:09:07 [INFO] [reward] call=57 n=8 mean=0.484 std=0.484 min=-0.795 max=0.688 | parts mean: fmt=0.18 at=0.22 svc=0.04 par=0.08 task=0.04 len=0.04 | |
| {'loss': '0.5857', 'grad_norm': '2.956', 'learning_rate': '8.085e-07', 'num_tokens': '2.144e+05', 'completions/mean_length': '33.38', 'completions/min_length': '15', 'completions/max_length': '128', 'completions/clipped_ratio': '0.125', 'completions/mean_terminated_length': '19.86', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '35', 'rewards/reward_fn/mean': '0.4842', 'rewards/reward_fn/std': '0.5176', 'reward': '0.4842', 'reward_std': '0.5176', 'frac_reward_zero_std': '0.5', 'entropy': '1.25', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '6.187', 'epoch': '1.213'} | |
| 61%|ββββββ | 57/94 [02:56<02:36, 4.22s/it][A | |
| [A | |
| 61%|ββββββ | 57/94 [02:56<02:36, 4.22s/it][A03:09:10 [INFO] [reward] call=58 n=8 mean=0.661 std=0.028 min=0.632 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| {'loss': '0.1203', 'grad_norm': '1.477', 'learning_rate': '7.872e-07', 'num_tokens': '2.18e+05', 'completions/mean_length': '25', 'completions/min_length': '15', 'completions/max_length': '39', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '39', 'rewards/reward_fn/mean': '0.6605', 'rewards/reward_fn/std': '0.02982', 'reward': '0.6605', 'reward_std': '0.02982', 'frac_reward_zero_std': '0.5', 'entropy': '0.1817', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.296', 'epoch': '1.234'} | |
| 62%|βββββββ | 58/94 [02:59<02:22, 3.96s/it][A | |
| [A | |
| 62%|βββββββ | 58/94 [02:59<02:22, 3.96s/it][A03:09:12 [INFO] [reward] call=59 n=8 mean=0.636 std=0.005 min=0.631 max=0.646 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 63%|βββββββ | 59/94 [03:01<01:54, 3.28s/it][A | |
| {'loss': '0.06092', 'grad_norm': '1.475', 'learning_rate': '7.66e-07', 'num_tokens': '2.198e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '28', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '28', 'rewards/reward_fn/mean': '0.6359', 'rewards/reward_fn/std': '0.005674', 'reward': '0.6359', 'reward_std': '0.005674', 'frac_reward_zero_std': '0', 'entropy': '0.9263', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.674', 'epoch': '1.255'} | |
| [A | |
| 63%|βββββββ | 59/94 [03:01<01:54, 3.28s/it][A03:09:15 [INFO] [reward] call=60 n=8 mean=0.668 std=0.027 min=0.633 max=0.690 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04 | |
| 64%|βββββββ | 60/94 [03:04<01:50, 3.24s/it][A | |
| [A | |
| {'loss': '-0.01362', 'grad_norm': '2.836', 'learning_rate': '7.447e-07', 'num_tokens': '2.234e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '36', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '36', 'rewards/reward_fn/mean': '0.6682', 'rewards/reward_fn/std': '0.02847', 'reward': '0.6682', 'reward_std': '0.02847', 'frac_reward_zero_std': '0', 'entropy': '0.4603', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.095', 'epoch': '1.277'} | |
| 64%|βββββββ | 60/94 [03:04<01:50, 3.24s/it][A03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:09:16 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:09:16 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:09:16 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:09:16 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:09:16 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:09:19 [INFO] [reward] call=61 n=8 mean=0.619 std=0.020 min=0.599 max=0.647 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04 | |
| 03:09:20 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 65%|βββββββ | 61/94 [03:08<01:57, 3.57s/it][A | |
| [A | |
| {'loss': '-0.2418', 'grad_norm': '1.858', 'learning_rate': '7.234e-07', 'num_tokens': '2.27e+05', 'completions/mean_length': '23.38', 'completions/min_length': '15', 'completions/max_length': '55', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '23.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '55', 'rewards/reward_fn/mean': '0.6188', 'rewards/reward_fn/std': '0.0218', 'reward': '0.6188', 'reward_std': '0.0218', 'frac_reward_zero_std': '0.5', 'entropy': '0.5959', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.761', 'epoch': '1.298'} | |
| 65%|βββββββ | 61/94 [03:08<01:57, 3.57s/it][A03:09:22 [INFO] [reward] call=62 n=8 mean=0.693 std=0.005 min=0.687 max=0.698 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04 | |
| 66%|βββββββ | 62/94 [03:11<01:48, 3.39s/it][A | |
| [A | |
| {'loss': '-0.09937', 'grad_norm': '2.591', 'learning_rate': '7.021e-07', 'num_tokens': '2.323e+05', 'completions/mean_length': '18.62', 'completions/min_length': '17', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.62', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.6926', 'rewards/reward_fn/std': '0.004852', 'reward': '0.6926', 'reward_std': '0.004852', 'frac_reward_zero_std': '0', 'entropy': '0.2111', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.929', 'epoch': '1.319'} | |
| 66%|βββββββ | 62/94 [03:11<01:48, 3.39s/it][A03:09:25 [INFO] [reward] call=63 n=8 mean=0.580 std=0.182 min=0.105 max=0.699 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.06 task=0.04 len=0.03 | |
| 67%|βββββββ | 63/94 [03:14<01:38, 3.17s/it][A | |
| [A | |
| {'loss': '0.01356', 'grad_norm': '5.029', 'learning_rate': '6.809e-07', 'num_tokens': '2.358e+05', 'completions/mean_length': '16', 'completions/min_length': '15', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.5795', 'rewards/reward_fn/std': '0.1948', 'reward': '0.5795', 'reward_std': '0.1948', 'frac_reward_zero_std': '0.5', 'entropy': '0.06583', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.588', 'epoch': '1.34'} | |
| 67%|βββββββ | 63/94 [03:14<01:38, 3.17s/it][A03:09:28 [INFO] [reward] call=64 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04 | |
| 68%|βββββββ | 64/94 [03:17<01:31, 3.04s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '6.596e-07', 'num_tokens': '2.411e+05', 'completions/mean_length': '17.25', 'completions/min_length': '17', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.25', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.1545', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.688', 'epoch': '1.362'} | |
| 68%|βββββββ | 64/94 [03:17<01:31, 3.04s/it][A03:09:30 [INFO] [reward] call=65 n=8 mean=0.635 std=0.003 min=0.633 max=0.643 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 69%|βββββββ | 65/94 [03:19<01:19, 2.74s/it][A | |
| [A | |
| {'loss': '-0.2188', 'grad_norm': '1.902', 'learning_rate': '6.383e-07', 'num_tokens': '2.43e+05', 'completions/mean_length': '22.12', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.635', 'rewards/reward_fn/std': '0.003465', 'reward': '0.635', 'reward_std': '0.003465', 'frac_reward_zero_std': '0.5', 'entropy': '0.3234', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.036', 'epoch': '1.383'} | |
| 69%|βββββββ | 65/94 [03:19<01:19, 2.74s/it][A03:09:33 [INFO] [reward] call=66 n=8 mean=0.664 std=0.028 min=0.633 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 70%|βββββββ | 66/94 [03:22<01:23, 2.99s/it][A | |
| [A | |
| {'loss': '-0.3033', 'grad_norm': '2.428', 'learning_rate': '6.17e-07', 'num_tokens': '2.466e+05', 'completions/mean_length': '20.12', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6642', 'rewards/reward_fn/std': '0.02946', 'reward': '0.6642', 'reward_std': '0.02946', 'frac_reward_zero_std': '0', 'entropy': '0.36', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.486', 'epoch': '1.404'} | |
| 70%|βββββββ | 66/94 [03:22<01:23, 2.99s/it][A03:09:36 [INFO] [reward] call=67 n=8 mean=0.644 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04 | |
| 71%|ββββββββ | 67/94 [03:25<01:17, 2.86s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '5.957e-07', 'num_tokens': '2.518e+05', 'completions/mean_length': '16.5', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6436', 'rewards/reward_fn/std': '0.04789', 'reward': '0.6436', 'reward_std': '0.04789', 'frac_reward_zero_std': '1', 'entropy': '0.03455', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.51', 'epoch': '1.426'} | |
| 71%|ββββββββ | 67/94 [03:25<01:17, 2.86s/it][A03:09:39 [INFO] [reward] call=68 n=8 mean=0.661 std=0.027 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| {'loss': '-0.08996', 'grad_norm': '0.9076', 'learning_rate': '5.745e-07', 'num_tokens': '2.553e+05', 'completions/mean_length': '18.75', 'completions/min_length': '15', 'completions/max_length': '33', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '33', 'rewards/reward_fn/mean': '0.6609', 'rewards/reward_fn/std': '0.02937', 'reward': '0.6609', 'reward_std': '0.02937', 'frac_reward_zero_std': '0.5', 'entropy': '0.1516', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.04', 'epoch': '1.447'} | |
| 72%|ββββββββ | 68/94 [03:28<01:16, 2.93s/it][A | |
| [A | |
| 72%|ββββββββ | 68/94 [03:28<01:16, 2.93s/it][A03:09:42 [INFO] [reward] call=69 n=8 mean=0.626 std=0.032 min=0.599 max=0.699 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.05 task=0.04 len=0.02 | |
| 73%|ββββββββ | 69/94 [03:31<01:11, 2.84s/it][A | |
| [A | |
| {'loss': '0.02351', 'grad_norm': '4.787', 'learning_rate': '5.532e-07', 'num_tokens': '2.589e+05', 'completions/mean_length': '15.25', 'completions/min_length': '10', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.25', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.6258', 'rewards/reward_fn/std': '0.03376', 'reward': '0.6258', 'reward_std': '0.03376', 'frac_reward_zero_std': '0', 'entropy': '0.1078', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.576', 'epoch': '1.468'} | |
| 73%|ββββββββ | 69/94 [03:31<01:11, 2.84s/it][A03:09:44 [INFO] [reward] call=70 n=8 mean=0.654 std=0.025 min=0.633 max=0.703 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| 74%|ββββββββ | 70/94 [03:33<01:02, 2.61s/it][A | |
| [A | |
| {'loss': '0.04798', 'grad_norm': '2.384', 'learning_rate': '5.319e-07', 'num_tokens': '2.608e+05', 'completions/mean_length': '22.5', 'completions/min_length': '15', 'completions/max_length': '41', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.5', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '41', 'rewards/reward_fn/mean': '0.654', 'rewards/reward_fn/std': '0.02694', 'reward': '0.654', 'reward_std': '0.02694', 'frac_reward_zero_std': '0', 'entropy': '0.6745', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.055', 'epoch': '1.489'} | |
| 74%|ββββββββ | 70/94 [03:33<01:02, 2.61s/it][A03:09:44 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:09:44 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:09:45 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:09:45 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:09:45 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:09:45 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:09:45 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:09:45 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:09:47 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 03:09:48 [INFO] [reward] call=71 n=8 mean=0.650 std=0.037 min=0.598 max=0.699 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.07 task=0.04 len=0.04 | |
| 76%|ββββββββ | 71/94 [03:37<01:08, 3.00s/it][A | |
| [A | |
| {'loss': '0.05797', 'grad_norm': '2.903', 'learning_rate': '5.106e-07', 'num_tokens': '2.644e+05', 'completions/mean_length': '23.38', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '23.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6501', 'rewards/reward_fn/std': '0.03967', 'reward': '0.6501', 'reward_std': '0.03967', 'frac_reward_zero_std': '0', 'entropy': '0.5277', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.364', 'epoch': '1.511'} | |
| 76%|ββββββββ | 71/94 [03:37<01:08, 3.00s/it][A03:09:50 [INFO] [reward] call=72 n=8 mean=0.646 std=0.023 min=0.633 max=0.705 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| 77%|ββββββββ | 72/94 [03:38<00:58, 2.65s/it][A | |
| [A | |
| {'loss': '-0.255', 'grad_norm': '3.032', 'learning_rate': '4.894e-07', 'num_tokens': '2.662e+05', 'completions/mean_length': '19.62', 'completions/min_length': '15', 'completions/max_length': '38', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.62', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '38', 'rewards/reward_fn/mean': '0.6461', 'rewards/reward_fn/std': '0.02479', 'reward': '0.6461', 'reward_std': '0.02479', 'frac_reward_zero_std': '0', 'entropy': '0.4351', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.847', 'epoch': '1.532'} | |
| 77%|ββββββββ | 72/94 [03:38<00:58, 2.65s/it][A03:09:52 [INFO] [reward] call=73 n=8 mean=0.652 std=0.031 min=0.598 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| {'loss': '-0.05995', 'grad_norm': '5.942', 'learning_rate': '4.681e-07', 'num_tokens': '2.698e+05', 'completions/mean_length': '19.5', 'completions/min_length': '16', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.6523', 'rewards/reward_fn/std': '0.03299', 'reward': '0.6523', 'reward_std': '0.03299', 'frac_reward_zero_std': '0', 'entropy': '0.3749', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.902', 'epoch': '1.553'} | |
| 78%|ββββββββ | 73/94 [03:41<00:57, 2.75s/it][A | |
| [A | |
| 78%|ββββββββ | 73/94 [03:41<00:57, 2.75s/it][A03:09:55 [INFO] [reward] call=74 n=8 mean=0.593 std=0.163 min=0.172 max=0.700 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| 79%|ββββββββ | 74/94 [03:44<00:55, 2.77s/it][A | |
| [A | |
| {'loss': '0.03651', 'grad_norm': '3.928', 'learning_rate': '4.468e-07', 'num_tokens': '2.733e+05', 'completions/mean_length': '19.75', 'completions/min_length': '15', 'completions/max_length': '27', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '19.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '27', 'rewards/reward_fn/mean': '0.5934', 'rewards/reward_fn/std': '0.1739', 'reward': '0.5934', 'reward_std': '0.1739', 'frac_reward_zero_std': '0', 'entropy': '0.4748', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.764', 'epoch': '1.574'} | |
| 79%|ββββββββ | 74/94 [03:44<00:55, 2.77s/it][A03:09:58 [INFO] [reward] call=75 n=8 mean=0.618 std=0.020 min=0.599 max=0.650 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.03 task=0.04 len=0.04 | |
| 80%|ββββββββ | 75/94 [03:47<00:53, 2.81s/it][A | |
| [A | |
| {'loss': '-0.1599', 'grad_norm': '2.755', 'learning_rate': '4.255e-07', 'num_tokens': '2.769e+05', 'completions/mean_length': '17.38', 'completions/min_length': '15', 'completions/max_length': '30', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.38', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '30', 'rewards/reward_fn/mean': '0.6182', 'rewards/reward_fn/std': '0.02141', 'reward': '0.6182', 'reward_std': '0.02141', 'frac_reward_zero_std': '0.5', 'entropy': '0.1834', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.857', 'epoch': '1.596'} | |
| 80%|ββββββββ | 75/94 [03:47<00:53, 2.81s/it][A03:10:01 [INFO] [reward] call=76 n=8 mean=0.667 std=0.040 min=0.598 max=0.696 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.09 task=0.04 len=0.04 | |
| 81%|ββββββββ | 76/94 [03:50<00:49, 2.76s/it][A | |
| [A | |
| {'loss': '-0.01251', 'grad_norm': '4.964', 'learning_rate': '4.043e-07', 'num_tokens': '2.82e+05', 'completions/mean_length': '17.25', 'completions/min_length': '16', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.25', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6668', 'rewards/reward_fn/std': '0.04234', 'reward': '0.6668', 'reward_std': '0.04234', 'frac_reward_zero_std': '0.5', 'entropy': '0.1384', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.587', 'epoch': '1.617'} | |
| 81%|ββββββββ | 76/94 [03:50<00:49, 2.76s/it][A03:10:04 [INFO] [reward] call=77 n=8 mean=0.493 std=0.485 min=-0.789 max=0.690 | parts mean: fmt=0.18 at=0.22 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 82%|βββββββββ | 77/94 [03:52<00:46, 2.74s/it][A | |
| [A | |
| {'loss': '0.05487', 'grad_norm': '3.762', 'learning_rate': '3.83e-07', 'num_tokens': '2.872e+05', 'completions/mean_length': '18', 'completions/min_length': '16', 'completions/max_length': '22', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '22', 'rewards/reward_fn/mean': '0.4926', 'rewards/reward_fn/std': '0.5189', 'reward': '0.4926', 'reward_std': '0.5189', 'frac_reward_zero_std': '0.5', 'entropy': '0.2181', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.612', 'epoch': '1.638'} | |
| 82%|βββββββββ | 77/94 [03:52<00:46, 2.74s/it][A03:10:06 [INFO] [reward] call=78 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04 | |
| 83%|βββββββββ | 78/94 [03:55<00:43, 2.69s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '3.617e-07', 'num_tokens': '2.924e+05', 'completions/mean_length': '17.88', 'completions/min_length': '17', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17.88', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.09081', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.535', 'epoch': '1.66'} | |
| 83%|βββββββββ | 78/94 [03:55<00:43, 2.69s/it][A03:10:09 [INFO] [reward] call=79 n=8 mean=0.592 std=0.185 min=0.109 max=0.699 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| 84%|βββββββββ | 79/94 [03:58<00:41, 2.80s/it][A | |
| [A | |
| {'loss': '-0.03821', 'grad_norm': '5.515', 'learning_rate': '3.404e-07', 'num_tokens': '2.96e+05', 'completions/mean_length': '18', 'completions/min_length': '11', 'completions/max_length': '32', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '32', 'rewards/reward_fn/mean': '0.5917', 'rewards/reward_fn/std': '0.1975', 'reward': '0.5917', 'reward_std': '0.1975', 'frac_reward_zero_std': '0', 'entropy': '0.2983', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.973', 'epoch': '1.681'} | |
| 84%|βββββββββ | 79/94 [03:58<00:41, 2.80s/it][A03:10:13 [INFO] [reward] call=80 n=8 mean=0.663 std=0.026 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 85%|βββββββββ | 80/94 [04:02<00:42, 3.06s/it][A | |
| {'loss': '-0.3258', 'grad_norm': '2.044', 'learning_rate': '3.191e-07', 'num_tokens': '2.996e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '52', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '52', 'rewards/reward_fn/mean': '0.6627', 'rewards/reward_fn/std': '0.02788', 'reward': '0.6627', 'reward_std': '0.02788', 'frac_reward_zero_std': '0.5', 'entropy': '0.1508', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.607', 'epoch': '1.702'} | |
| [A | |
| 85%|βββββββββ | 80/94 [04:02<00:42, 3.06s/it][A03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:14 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:14 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:14 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:14 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:10:14 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:10:16 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 03:10:17 [INFO] [reward] call=81 n=8 mean=0.674 std=0.033 min=0.630 max=0.713 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04 | |
| 86%|βββββββββ | 81/94 [04:06<00:42, 3.30s/it][A | |
| [A | |
| {'loss': '-0.1966', 'grad_norm': '3.974', 'learning_rate': '2.979e-07', 'num_tokens': '3.032e+05', 'completions/mean_length': '20', 'completions/min_length': '15', 'completions/max_length': '39', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '39', 'rewards/reward_fn/mean': '0.6736', 'rewards/reward_fn/std': '0.03483', 'reward': '0.6736', 'reward_std': '0.03483', 'frac_reward_zero_std': '0', 'entropy': '0.3634', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.307', 'epoch': '1.723'} | |
| 86%|βββββββββ | 81/94 [04:06<00:42, 3.30s/it][A03:10:19 [INFO] [reward] call=82 n=8 mean=0.653 std=0.025 min=0.633 max=0.700 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.07 task=0.04 len=0.04 | |
| 87%|βββββββββ | 82/94 [04:08<00:34, 2.90s/it][A | |
| [A | |
| {'loss': '-0.09556', 'grad_norm': '1.988', 'learning_rate': '2.766e-07', 'num_tokens': '3.051e+05', 'completions/mean_length': '23.75', 'completions/min_length': '15', 'completions/max_length': '38', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '23.75', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '38', 'rewards/reward_fn/mean': '0.6531', 'rewards/reward_fn/std': '0.02641', 'reward': '0.6531', 'reward_std': '0.02641', 'frac_reward_zero_std': '0', 'entropy': '1', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '1.982', 'epoch': '1.745'} | |
| 87%|βββββββββ | 82/94 [04:08<00:34, 2.90s/it][A03:10:22 [INFO] [reward] call=83 n=8 mean=0.671 std=0.027 min=0.633 max=0.702 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.10 task=0.04 len=0.04 | |
| {'loss': '-0.03983', 'grad_norm': '1.799', 'learning_rate': '2.553e-07', 'num_tokens': '3.087e+05', 'completions/mean_length': '21.88', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.671', 'rewards/reward_fn/std': '0.02835', 'reward': '0.671', 'reward_std': '0.02835', 'frac_reward_zero_std': '0.5', 'entropy': '0.5', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.575', 'epoch': '1.766'} | |
| 88%|βββββββββ | 83/94 [04:11<00:34, 3.12s/it][A | |
| [A | |
| 88%|βββββββββ | 83/94 [04:11<00:34, 3.12s/it][A03:10:25 [INFO] [reward] call=84 n=8 mean=0.598 std=0.221 min=0.020 max=0.704 | parts mean: fmt=0.20 at=0.22 svc=0.05 par=0.09 task=0.04 len=0.02 | |
| 89%|βββββββββ | 84/94 [04:14<00:30, 3.05s/it][A | |
| [A | |
| {'loss': '-0.09434', 'grad_norm': '6.413', 'learning_rate': '2.34e-07', 'num_tokens': '3.139e+05', 'completions/mean_length': '16.62', 'completions/min_length': '11', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.62', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.5975', 'rewards/reward_fn/std': '0.2358', 'reward': '0.5975', 'reward_std': '0.2358', 'frac_reward_zero_std': '0', 'entropy': '0.1545', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.821', 'epoch': '1.787'} | |
| 89%|βββββββββ | 84/94 [04:14<00:30, 3.05s/it][A03:10:28 [INFO] [reward] call=85 n=8 mean=0.664 std=0.025 min=0.633 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.09 task=0.04 len=0.04 | |
| 90%|βββββββββ | 85/94 [04:17<00:27, 3.07s/it][A | |
| {'loss': '-0.1992', 'grad_norm': '1.506', 'learning_rate': '2.128e-07', 'num_tokens': '3.174e+05', 'completions/mean_length': '21', 'completions/min_length': '15', 'completions/max_length': '37', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '21', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '37', 'rewards/reward_fn/mean': '0.6643', 'rewards/reward_fn/std': '0.02633', 'reward': '0.6643', 'reward_std': '0.02633', 'frac_reward_zero_std': '0.5', 'entropy': '0.2386', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.042', 'epoch': '1.809'} | |
| [A | |
| 90%|βββββββββ | 85/94 [04:17<00:27, 3.07s/it][A03:10:32 [INFO] [reward] call=86 n=8 mean=0.666 std=0.035 min=0.611 max=0.712 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.08 task=0.04 len=0.04 | |
| 91%|ββββββββββ| 86/94 [04:21<00:25, 3.21s/it][A | |
| [A | |
| {'loss': '-0.06428', 'grad_norm': '2.73', 'learning_rate': '1.915e-07', 'num_tokens': '3.21e+05', 'completions/mean_length': '24.12', 'completions/min_length': '15', 'completions/max_length': '48', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '24.12', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '48', 'rewards/reward_fn/mean': '0.6659', 'rewards/reward_fn/std': '0.03774', 'reward': '0.6659', 'reward_std': '0.03774', 'frac_reward_zero_std': '0', 'entropy': '0.691', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.488', 'epoch': '1.83'} | |
| 91%|ββββββββββ| 86/94 [04:21<00:25, 3.21s/it][A03:10:34 [INFO] [reward] call=87 n=8 mean=0.692 std=0.004 min=0.688 max=0.697 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04 | |
| 93%|ββββββββββ| 87/94 [04:24<00:23, 3.31s/it][A | |
| {'loss': '-4.491e-06', 'grad_norm': '3.164', 'learning_rate': '1.702e-07', 'num_tokens': '3.264e+05', 'completions/mean_length': '17', 'completions/min_length': '17', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '17', 'completions/min_terminated_length': '17', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6915', 'rewards/reward_fn/std': '0.003961', 'reward': '0.6915', 'reward_std': '0.003961', 'frac_reward_zero_std': '0.5', 'entropy': '0.1973', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '3.48', 'epoch': '1.851'} | |
| [A | |
| 93%|ββββββββββ| 87/94 [04:24<00:23, 3.31s/it][A03:10:39 [INFO] [reward] call=88 n=8 mean=0.484 std=0.478 min=-0.779 max=0.688 | parts mean: fmt=0.18 at=0.22 svc=0.04 par=0.08 task=0.04 len=0.04 | |
| 94%|ββββββββββ| 88/94 [04:28<00:21, 3.54s/it][A | |
| [A | |
| {'loss': '0.4213', 'grad_norm': '2.828', 'learning_rate': '1.489e-07', 'num_tokens': '3.299e+05', 'completions/mean_length': '22.25', 'completions/min_length': '15', 'completions/max_length': '65', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '22.25', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '65', 'rewards/reward_fn/mean': '0.4843', 'rewards/reward_fn/std': '0.5113', 'reward': '0.4843', 'reward_std': '0.5113', 'frac_reward_zero_std': '0.5', 'entropy': '0.3975', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '4.028', 'epoch': '1.872'} | |
| 94%|ββββββββββ| 88/94 [04:28<00:21, 3.54s/it][A03:10:42 [INFO] [reward] call=89 n=8 mean=0.688 std=0.000 min=0.688 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.12 task=0.04 len=0.04 | |
| 95%|ββββββββββ| 89/94 [04:31<00:16, 3.26s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '1.277e-07', 'num_tokens': '3.35e+05', 'completions/mean_length': '18', 'completions/min_length': '18', 'completions/max_length': '18', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '18', 'completions/min_terminated_length': '18', 'completions/max_terminated_length': '18', 'rewards/reward_fn/mean': '0.6884', 'rewards/reward_fn/std': '0', 'reward': '0.6884', 'reward_std': '0', 'frac_reward_zero_std': '1', 'entropy': '0.1114', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.534', 'epoch': '1.894'} | |
| 95%|ββββββββββ| 89/94 [04:31<00:16, 3.26s/it][A03:10:45 [INFO] [reward] call=90 n=8 mean=0.644 std=0.045 min=0.599 max=0.688 | parts mean: fmt=0.20 at=0.25 svc=0.07 par=0.06 task=0.04 len=0.04 | |
| 96%|ββββββββββ| 90/94 [04:33<00:12, 3.04s/it][A | |
| [A | |
| {'loss': '0', 'grad_norm': '0', 'learning_rate': '1.064e-07', 'num_tokens': '3.402e+05', 'completions/mean_length': '16.5', 'completions/min_length': '16', 'completions/max_length': '17', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.5', 'completions/min_terminated_length': '16', 'completions/max_terminated_length': '17', 'rewards/reward_fn/mean': '0.6436', 'rewards/reward_fn/std': '0.04789', 'reward': '0.6436', 'reward_std': '0.04789', 'frac_reward_zero_std': '1', 'entropy': '0.05231', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.457', 'epoch': '1.915'} | |
| 96%|ββββββββββ| 90/94 [04:33<00:12, 3.04s/it][A03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:45 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:46 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:46 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:46 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:10:46 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:10:47 [INFO] [reward] call=91 n=8 mean=0.636 std=0.003 min=0.633 max=0.642 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| 97%|ββββββββββ| 91/94 [04:36<00:08, 2.89s/it][A | |
| [A | |
| {'loss': '-0.3753', 'grad_norm': '2.449', 'learning_rate': '8.511e-08', 'num_tokens': '3.421e+05', 'completions/mean_length': '24.88', 'completions/min_length': '15', 'completions/max_length': '44', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '24.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '44', 'rewards/reward_fn/mean': '0.6358', 'rewards/reward_fn/std': '0.003546', 'reward': '0.6358', 'reward_std': '0.003546', 'frac_reward_zero_std': '0', 'entropy': '0.4588', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.06', 'epoch': '1.936'} | |
| 97%|ββββββββββ| 91/94 [04:36<00:08, 2.89s/it][A03:10:49 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 03:10:50 [INFO] [reward] call=92 n=8 mean=0.639 std=0.039 min=0.599 max=0.705 | parts mean: fmt=0.20 at=0.25 svc=0.06 par=0.07 task=0.04 len=0.02 | |
| 98%|ββββββββββ| 92/94 [04:39<00:05, 2.83s/it][A | |
| [A | |
| {'loss': '-0.09974', 'grad_norm': '5.525', 'learning_rate': '6.383e-08', 'num_tokens': '3.456e+05', 'completions/mean_length': '15.62', 'completions/min_length': '10', 'completions/max_length': '21', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '15.62', 'completions/min_terminated_length': '10', 'completions/max_terminated_length': '21', 'rewards/reward_fn/mean': '0.6391', 'rewards/reward_fn/std': '0.0416', 'reward': '0.6391', 'reward_std': '0.0416', 'frac_reward_zero_std': '0', 'entropy': '0.1428', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.638', 'epoch': '1.957'} | |
| 98%|ββββββββββ| 92/94 [04:39<00:05, 2.83s/it][A03:10:52 [INFO] [reward] call=93 n=8 mean=0.635 std=0.003 min=0.633 max=0.641 | parts mean: fmt=0.20 at=0.25 svc=0.05 par=0.06 task=0.04 len=0.04 | |
| {'loss': '-0.1793', 'grad_norm': '1.453', 'learning_rate': '4.255e-08', 'num_tokens': '3.475e+05', 'completions/mean_length': '20.88', 'completions/min_length': '15', 'completions/max_length': '42', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '20.88', 'completions/min_terminated_length': '15', 'completions/max_terminated_length': '42', 'rewards/reward_fn/mean': '0.6348', 'rewards/reward_fn/std': '0.002901', 'reward': '0.6348', 'reward_std': '0.002901', 'frac_reward_zero_std': '0.5', 'entropy': '0.2665', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.004', 'epoch': '1.979'} | |
| 99%|ββββββββββ| 93/94 [04:41<00:02, 2.59s/it][A | |
| [A | |
| 99%|ββββββββββ| 93/94 [04:41<00:02, 2.59s/it][A03:10:55 [INFO] [reward] call=94 n=8 mean=0.575 std=0.214 min=0.020 max=0.703 | parts mean: fmt=0.20 at=0.22 svc=0.06 par=0.06 task=0.04 len=0.02 | |
| 100%|ββββββββββ| 94/94 [04:43<00:00, 2.63s/it][A | |
| [A | |
| {'loss': '-0.08947', 'grad_norm': '6.981', 'learning_rate': '2.128e-08', 'num_tokens': '3.527e+05', 'completions/mean_length': '16.12', 'completions/min_length': '11', 'completions/max_length': '19', 'completions/clipped_ratio': '0', 'completions/mean_terminated_length': '16.12', 'completions/min_terminated_length': '11', 'completions/max_terminated_length': '19', 'rewards/reward_fn/mean': '0.5746', 'rewards/reward_fn/std': '0.2286', 'reward': '0.5746', 'reward_std': '0.2286', 'frac_reward_zero_std': '0', 'entropy': '0.1197', 'clip_ratio/low_mean': '0', 'clip_ratio/low_min': '0', 'clip_ratio/high_mean': '0', 'clip_ratio/high_max': '0', 'clip_ratio/region_mean': '0', 'step_time': '2.683', 'epoch': '2'} | |
| 100%|ββββββββββ| 94/94 [04:43<00:00, 2.63s/it][A03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:55 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| [A | |
| {'train_runtime': '284.5', 'train_samples_per_second': '0.668', 'train_steps_per_second': '0.33', 'train_loss': '-0.05751', 'epoch': '2'} | |
| 100%|ββββββββββ| 94/94 [04:44<00:00, 2.63s/it][A | |
| 100%|ββββββββββ| 94/94 [04:44<00:00, 3.03s/it] | |
| 03:10:56 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:56 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:56 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:10:56 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| 03:10:58 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 03:10:58 [INFO] Saving to /tmp/sre-grpo-9zduekwb | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:58 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:10:59 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| Processing Files (0 / 0) : | | 0.00B / 0.00B | |
| New Data Upload : | | 0.00B / 0.00B [A | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB [A[A | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB [A[A[A | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB [A[A[A[A | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB [A[A | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB [A[A[A | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB [A[A[A[A | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB [A[A | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB [A[A[A | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB [A[A[A[A | |
| Processing Files (3 / 3) : 100%|ββββββββββ| 92.2MB / 92.2MB, 0.00B/s | |
| New Data Upload : | | 0.00B / 0.00B, 0.00B/s | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB | |
| 03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/commit/main "HTTP/1.1 200 OK" | |
| 03:10:59 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict" | |
| 03:10:59 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:10:59 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:10:59 [INFO] HTTP Request: HEAD https://huggingface.co/srinjoyd/qwen2.5-7b-sre-merged/resolve/main/config.json "HTTP/1.1 307 Temporary Redirect" | |
| 03:11:00 [INFO] HTTP Request: HEAD https://huggingface.co/api/resolve-cache/models/srinjoyd/qwen2.5-7b-sre-merged/8f73acd8434aacfb77d224b6c46773bca9279705/config.json "HTTP/1.1 200 OK" | |
| 03:11:00 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:11:00 [INFO] HTTP Request: POST https://huggingface.co/api/validate-yaml "HTTP/1.1 200 OK" | |
| 03:11:00 [INFO] HTTP Request: POST https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/preupload/main "HTTP/1.1 200 OK" | |
| 03:11:00 [INFO] HTTP Request: POST https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo.git/info/lfs/objects/batch "HTTP/1.1 200 OK" | |
| Processing Files (0 / 0) : | | 0.00B / 0.00B | |
| New Data Upload : | | 0.00B / 0.00B [A | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB [A[A | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB [A[A[A | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB [A[A[A[A | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB [A[A | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB [A[A[A | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB [A[A[A[A | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB [A[A | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB [A[A[A | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB [A[A[A[A | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB [A[A | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB [A[A[A | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB [A[A[A[A | |
| Processing Files (3 / 3) : 100%|ββββββββββ| 92.2MB / 92.2MB, 0.00B/s | |
| New Data Upload : | | 0.00B / 0.00B, 0.00B/s | |
| ...zduekwb/training_args.bin: 100%|ββββββββββ| 7.25kB / 7.25kB | |
| ...o-9zduekwb/tokenizer.json: 100%|ββββββββββ| 11.4MB / 11.4MB | |
| ...adapter_model.safetensors: 100%|ββββββββββ| 80.8MB / 80.8MB | |
| No files have been modified since last commit. Skipping to prevent empty commit. | |
| 03:11:00 [WARNING] No files have been modified since last commit. Skipping to prevent empty commit. | |
| 03:11:00 [INFO] HTTP Request: GET https://huggingface.co/api/models/daemongg/qwen2.5-7b-sre-grpo/revision/main "HTTP/1.1 200 OK" | |
| 03:11:00 [INFO] Pushed to https://huggingface.co/daemongg/qwen2.5-7b-sre-grpo | |
| 03:11:00 [INFO] Removed local /tmp/sre-grpo-9zduekwb (--no-keep-local) | |
| 03:11:00 [INFO] Done. |