Instructions to use Shaer-AI-2/Shaer-adapters-grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Shaer-AI-2/Shaer-adapters-grpo with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Shaer-AI-2/Shaer-adapters-grpo", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Training in progress, step 1100
Browse files- adapter_model.safetensors +1 -1
- all_generations.jsonl +2 -2
- checkpoint_events.jsonl +2 -0
- metrics.csv +51 -0
- metrics.jsonl +51 -0
- plots/arabic_gate_chain.png +2 -2
- plots/arabic_gate_run.png +2 -2
- plots/chain_metrics.jsonl +51 -0
- plots/meter_by_meter_chain.png +2 -2
- plots/meter_by_meter_run.png +2 -2
- plots/reward_panels_eval_chain.png +2 -2
- plots/reward_panels_eval_run.png +2 -2
- plots/reward_panels_train_chain.png +2 -2
- plots/reward_panels_train_run.png +2 -2
- plotter.log +100 -0
- reward_arabic_clean_debug.jsonl +2 -2
- reward_count_adherence_debug.jsonl +2 -2
- reward_meter_debug.jsonl +2 -2
- reward_total_composite_debug.jsonl +2 -2
- train.log +51 -0
- train_stdout.log +0 -0
adapter_model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 639691872
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:96ed2b8defba6d9cd2e5781f26b91d112dfc3641106ff08185c9afa2c9e3f275
|
| 3 |
size 639691872
|
all_generations.jsonl
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1349117cec38db2284a28628bce046a2f1428ddc6ef18a3b4166fa12a616aef6
|
| 3 |
+
size 199548275
|
checkpoint_events.jsonl
CHANGED
|
@@ -40,3 +40,5 @@
|
|
| 40 |
{"timestamp_utc": "2026-04-11T21:32:52Z", "event_type": "checkpoint_saved", "global_step": 1000, "local_checkpoint_dir": "/root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/checkpoint-1000", "hub_model_id": "Shaer-AI/Shaer-adapters-grpo", "expected_hub_prefix": "last-checkpoint"}
|
| 41 |
{"timestamp_utc": "2026-04-11T21:39:35Z", "event_type": "evaluation_completed", "global_step": 1050, "metrics": {"eval_loss": NaN, "eval_runtime": 85.4042, "eval_samples_per_second": 1.218, "eval_steps_per_second": 0.152}}
|
| 42 |
{"timestamp_utc": "2026-04-11T21:39:38Z", "event_type": "checkpoint_saved", "global_step": 1050, "local_checkpoint_dir": "/root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/checkpoint-1050", "hub_model_id": "Shaer-AI/Shaer-adapters-grpo", "expected_hub_prefix": "last-checkpoint"}
|
|
|
|
|
|
|
|
|
| 40 |
{"timestamp_utc": "2026-04-11T21:32:52Z", "event_type": "checkpoint_saved", "global_step": 1000, "local_checkpoint_dir": "/root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/checkpoint-1000", "hub_model_id": "Shaer-AI/Shaer-adapters-grpo", "expected_hub_prefix": "last-checkpoint"}
|
| 41 |
{"timestamp_utc": "2026-04-11T21:39:35Z", "event_type": "evaluation_completed", "global_step": 1050, "metrics": {"eval_loss": NaN, "eval_runtime": 85.4042, "eval_samples_per_second": 1.218, "eval_steps_per_second": 0.152}}
|
| 42 |
{"timestamp_utc": "2026-04-11T21:39:38Z", "event_type": "checkpoint_saved", "global_step": 1050, "local_checkpoint_dir": "/root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/checkpoint-1050", "hub_model_id": "Shaer-AI/Shaer-adapters-grpo", "expected_hub_prefix": "last-checkpoint"}
|
| 43 |
+
{"timestamp_utc": "2026-04-11T21:45:15Z", "event_type": "evaluation_completed", "global_step": 1100, "metrics": {"eval_loss": NaN, "eval_runtime": 73.9179, "eval_samples_per_second": 1.407, "eval_steps_per_second": 0.176}}
|
| 44 |
+
{"timestamp_utc": "2026-04-11T21:45:19Z", "event_type": "checkpoint_saved", "global_step": 1100, "local_checkpoint_dir": "/root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/checkpoint-1100", "hub_model_id": "Shaer-AI/Shaer-adapters-grpo", "expected_hub_prefix": "last-checkpoint"}
|
metrics.csv
CHANGED
|
@@ -1071,3 +1071,54 @@ clip_ratio/high_max,clip_ratio/high_mean,clip_ratio/low_mean,clip_ratio/low_min,
|
|
| 1071 |
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,61.0,61.0,61.0,61.0,0.007037240080535412,0.040546802594995365,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1050,0.0,6.821212121212122e-06,0.0,train,2265729.0,0.9968419075012207,1.0,0.0,1.0,0.0,0.9968419075012207,0.0,0.0,0.9968419075012207,0.0,0.9968419075012207,1.0,0.0,1.0,0.0,0.9968419075012207,0.0,0.9968419075012207,0.0,1.1012705564498901,0.999724805355072,0.4893798530101776,0.7146163582801819,0.0027122944593429565,2026-04-11T21:38:09Z
|
| 1072 |
,,,,,,,,,,,,,0.040546802594995365,0.0,0.0,0.0,0.0,0.0,0.057692307692307696,458.0,412.46153846153845,241.28846153846155,224.83516869178186,59.61538461538461,59.61538461538461,0.009740084463443894,0.0,nan,2265729.0,0.5442026945260855,0.9903846153846154,0.027196414195574246,0.8243002341343806,0.149119944526599,0.6695947761719043,0.4209407364519743,nan,0.5442026945260855,0.3753266856074333,0.5442026945260855,0.9903846153846154,0.027196414195574246,0.8243002341343806,0.149119944526599,0.6695947761719043,0.4209407364519743,0.5442026945260855,0.3753266856074333,85.4042,1.218,1.2309852104920607,1.0002473134260912,0.6470333154384906,0.48829063085409313,0.0014539593382953452,0.152,,1050,,,,eval,,,,,,,,,,,,,,,,,,,,,,,,,,2026-04-11T21:39:35Z
|
| 1073 |
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,61.0,61.0,61.0,61.0,0.004360992228612304,0.04058541859746679,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1051,0.0,6.818181818181818e-06,0.0,train,2267377.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.0,0.9985920786857605,0.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.9985920786857605,0.0,1.0429177284240723,1.0004585981369019,0.9996206760406494,0.042022280395030975,0.00045387333375401795,2026-04-11T21:39:43Z
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1071 |
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,61.0,61.0,61.0,61.0,0.007037240080535412,0.040546802594995365,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1050,0.0,6.821212121212122e-06,0.0,train,2265729.0,0.9968419075012207,1.0,0.0,1.0,0.0,0.9968419075012207,0.0,0.0,0.9968419075012207,0.0,0.9968419075012207,1.0,0.0,1.0,0.0,0.9968419075012207,0.0,0.9968419075012207,0.0,1.1012705564498901,0.999724805355072,0.4893798530101776,0.7146163582801819,0.0027122944593429565,2026-04-11T21:38:09Z
|
| 1072 |
,,,,,,,,,,,,,0.040546802594995365,0.0,0.0,0.0,0.0,0.0,0.057692307692307696,458.0,412.46153846153845,241.28846153846155,224.83516869178186,59.61538461538461,59.61538461538461,0.009740084463443894,0.0,nan,2265729.0,0.5442026945260855,0.9903846153846154,0.027196414195574246,0.8243002341343806,0.149119944526599,0.6695947761719043,0.4209407364519743,nan,0.5442026945260855,0.3753266856074333,0.5442026945260855,0.9903846153846154,0.027196414195574246,0.8243002341343806,0.149119944526599,0.6695947761719043,0.4209407364519743,0.5442026945260855,0.3753266856074333,85.4042,1.218,1.2309852104920607,1.0002473134260912,0.6470333154384906,0.48829063085409313,0.0014539593382953452,0.152,,1050,,,,eval,,,,,,,,,,,,,,,,,,,,,,,,,,2026-04-11T21:39:35Z
|
| 1073 |
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,61.0,61.0,61.0,61.0,0.004360992228612304,0.04058541859746679,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1051,0.0,6.818181818181818e-06,0.0,train,2267377.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.0,0.9985920786857605,0.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.9985920786857605,0.0,1.0429177284240723,1.0004585981369019,0.9996206760406494,0.042022280395030975,0.00045387333375401795,2026-04-11T21:39:43Z
|
| 1074 |
+
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,61.0,61.0,61.0,61.0,0.004134227987378836,0.040624034599938214,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1052,0.0,6.8151515151515155e-06,0.0,train,2269065.0,0.9968419075012207,1.0,0.0,1.0,0.0,0.9968419075012207,0.0,0.0,0.9968419075012207,0.0,0.9968419075012207,1.0,0.0,1.0,0.0,0.9968419075012207,0.0,0.9968419075012207,0.0,1.0188950300216675,1.0001823902130127,0.9077965617179871,0.09673504531383514,0.0007426206138916314,2026-04-11T21:39:48Z
|
| 1075 |
+
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,61.0,61.0,61.0,61.0,0.0012685270776273683,0.04066265060240964,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1053,0.0,6.812121212121212e-06,0.0,train,2270753.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.0,0.9985920786857605,0.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.9985920786857605,0.0,1.0118045806884766,1.0001797676086426,0.9999438524246216,0.011735539883375168,0.00017959998513106257,2026-04-11T21:39:53Z
|
| 1076 |
+
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,61.0,61.0,61.0,61.0,0.0013022492494201288,0.04070126660488106,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1054,0.0,6.80909090909091e-06,0.0,train,2272529.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.0,0.9985920786857605,0.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.9985920786857605,0.0,1.008744239807129,1.000115990638733,0.9994588494300842,0.008706126362085342,0.00012025728210574016,2026-04-11T21:39:57Z
|
| 1077 |
+
0.0,0.0,0.0,0.0,0.0,0.0,51.0,51.0,50.75,50.75,50.0,50.0,0.0185652831569314,0.040739882607352486,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1055,0.0,6.806060606060607e-06,0.0,train,2274175.0,0.9589456915855408,1.0,0.0,1.0,0.0,0.9589456915855408,0.0,0.0,0.9589456915855408,0.0,0.9589456915855408,1.0,0.0,1.0,0.0,0.9589456915855408,0.0,0.9589456915855408,0.0,1.2294468879699707,1.0008618831634521,0.7356403470039368,0.30701398849487305,0.0029871375299990177,2026-04-11T21:40:02Z
|
| 1078 |
+
0.0020491802133619785,0.0020491802133619785,0.013767930213361979,0.013767930213361979,0.015817110426723957,0.0,64.0,64.0,62.875,62.875,61.0,61.0,0.041025768499821424,0.04077849860982391,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1056,3.9809470176696777,6.803030303030304e-06,0.008,train,2275998.0,0.0021876515820622444,1.0,0.0,1.0,0.0,0.0021876515820622444,0.0002445352729409933,0.00024453524383716285,0.0021876515820622444,0.0002445352729409933,0.0021876515820622444,1.0,0.0,1.0,0.0,0.0021876515820622444,0.0002445352729409933,0.0021876515820622444,0.0002445352729409933,2.0,1.0037806034088135,0.6839931607246399,1.1146979331970215,0.008529768325388432,2026-04-11T21:40:07Z
|
| 1079 |
+
0.0012886597542092204,0.0012886597542092204,0.0015432098880410194,0.0015432098880410194,0.00283186964225024,0.0,97.0,97.0,83.0,83.0,81.0,81.0,0.003993239166447893,0.040817114612295334,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1057,0.9426341652870178,6.800000000000001e-06,-0.0561,train,2277958.0,0.7048847079277039,1.0,0.0,0.7083333730697632,0.11785111576318741,0.9951313734054565,0.0,0.1172773465514183,0.7048847079277039,0.1172773540019989,0.7048847079277039,1.0,0.0,0.7083333730697632,0.11785111576318741,0.9951313734054565,0.0,0.7048847079277039,0.1172773540019989,1.134204387664795,0.9974780082702637,0.03710322454571724,3.2940514087677,0.00949889700859785,2026-04-11T21:40:12Z
|
| 1080 |
+
0.0006345177534967661,0.0006345177534967661,0.0,0.0,0.0006345177534967661,0.0,209.0,209.0,198.5,198.5,197.0,197.0,0.0015565359972242732,0.04085573061476676,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1058,0.21764229238033295,6.796969696969697e-06,0.0039,train,2280914.0,0.7975266575813293,1.0,0.0,0.800000011920929,0.0,0.9969083070755005,3.72788890672382e-05,2.9827928301529028e-05,0.7975266575813293,2.9818897019140422e-05,0.7975266575813293,1.0,0.0,0.800000011920929,0.0,0.9969083070755005,3.72788890672382e-05,0.7975266575813293,2.9818897019140422e-05,1.167437195777893,0.9991796016693115,0.0298833679407835,3.510453224182129,0.003239656798541546,2026-04-11T21:40:19Z
|
| 1081 |
+
0.0,0.0,0.00657894741743803,0.00657894741743803,0.00657894741743803,0.0,57.0,57.0,54.25,54.25,52.0,52.0,0.025705090374685824,0.04089434661723818,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1059,4.7054972648620605,6.793939393939395e-06,0.0234,train,2282612.0,0.9560332298278809,1.0,0.0,1.0,0.0,0.9560332298278809,0.0188444871455431,0.018844490870833397,0.9560332298278809,0.0188444871455431,0.9560332298278809,1.0,0.0,1.0,0.0,0.9560332298278809,0.0188444871455431,0.9560332298278809,0.0188444871455431,1.3655518293380737,0.9992375373840332,0.38045772910118103,0.96638023853302,0.008652369491755962,2026-04-11T21:40:23Z
|
| 1082 |
+
0.0018382353009656072,0.0018382353009656072,0.00043252596515230834,0.00043252596515230834,0.0022707612661179155,0.0,289.0,289.0,274.125,274.125,272.0,272.0,0.0006548625879077008,0.040932962619709606,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1060,0.09687843173742294,6.790909090909091e-06,0.0057,train,2286469.0,0.8723691701889038,1.0,0.0,0.875,0.0,0.9969933032989502,5.181956657906994e-05,4.535001062322408e-05,0.8723691701889038,4.535001062322408e-05,0.8723691701889038,1.0,0.0,0.875,0.0,0.9969933032989502,5.181956657906994e-05,0.8723691701889038,4.535001062322408e-05,1.0306644439697266,0.9990590214729309,0.2921103239059448,1.230623722076416,0.00142471503932029,2026-04-11T21:40:31Z
|
| 1083 |
+
0.0,0.0,0.0,0.0,0.0,0.0,272.0,272.0,272.0,272.0,272.0,272.0,0.00023714406961516943,0.04097157862218103,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1061,0.0,6.787878787878789e-06,0.0,train,2290485.0,0.8723852038383484,1.0,0.0,0.875,0.0,0.997011661529541,0.0,0.0,0.8723852038383484,0.0,0.8723852038383484,1.0,0.0,0.875,0.0,0.997011661529541,0.0,0.8723852038383484,0.0,1.0054314136505127,0.9999895691871643,0.9654592871665955,0.03515136241912842,6.0048791056033224e-05,2026-04-11T21:40:38Z
|
| 1084 |
+
0.0,0.0,0.0,0.0,0.0,0.0,106.0,106.0,106.0,106.0,106.0,106.0,0.0023513801133958623,0.041010194624652455,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1062,0.0,6.7848484848484855e-06,0.0,train,2292821.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.0,0.9985920786857605,0.0,0.9985920786857605,1.0,0.0,1.0,0.0,0.9985920786857605,0.0,0.9985920786857605,0.0,1.0233614444732666,0.9978539347648621,0.0006421464495360851,7.350694179534912,0.01337476447224617,2026-04-11T21:40:44Z
|
| 1085 |
+
0.0,0.0,0.0,0.0,0.0,0.0,65.0,65.0,65.0,65.0,65.0,65.0,0.0019514950545271859,0.04104881062712388,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1063,0.0,6.781818181818183e-06,0.0,train,2294669.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.0,0.9949578642845154,0.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.9949578642845154,0.0,1.011854648590088,1.0001155138015747,0.9829086661338806,0.017239108681678772,0.00023568868346046656,2026-04-11T21:40:48Z
|
| 1086 |
+
0.002314814832061529,0.002314814832061529,0.009181267116218805,0.009181267116218805,0.011496081948280334,0.0,57.0,57.0,54.75,54.75,53.0,53.0,0.10237977746874094,0.0410874266295953,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1064,2.7271664142608643,6.778787878787879e-06,-0.0026,train,2296475.0,0.9718772172927856,1.0,0.0,1.0,0.0,0.9718772172927856,0.01329784281551838,0.013297837227582932,0.9718772172927856,0.01329784281551838,0.9718772172927856,1.0,0.0,1.0,0.0,0.9718772172927856,0.01329784281551838,0.9718772172927856,0.01329784281551838,1.5688531398773193,0.9998058676719666,0.005570904351770878,5.190197944641113,0.024222884327173233,2026-04-11T21:40:53Z
|
| 1087 |
+
0.0,0.0,0.0,0.0,0.0,0.0,59.0,59.0,57.25,57.25,57.0,57.0,0.01174389524385333,0.04112604263206673,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1065,5.474052429199219,6.7757575757575765e-06,0.0102,train,2298309.0,0.8288533687591553,1.0,0.0,1.0,0.0,0.8288533687591553,0.037421029061079025,0.03742102161049843,0.8288533687591553,0.037421029061079025,0.8288533687591553,1.0,0.0,1.0,0.0,0.8288533687591553,0.037421029061079025,0.8288533687591553,0.037421029061079025,1.1381598711013794,0.9999019503593445,0.5920739769935608,0.5241236686706543,0.002723454497754574,2026-04-11T21:40:58Z
|
| 1088 |
+
0.0020833334419876337,0.0020833334419876337,0.0,0.0,0.0020833334419876337,0.0,61.0,61.0,60.875,60.875,60.0,60.0,0.014361659123096615,0.04116465863453815,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1066,2.398463726043701,6.772727272727273e-06,0.0097,train,2300092.0,0.9969611763954163,1.0,0.0,1.0,0.0,0.9969611763954163,0.0003373012295924127,0.00033730725408531725,0.9969611763954163,0.0003373012295924127,0.9969611763954163,1.0,0.0,1.0,0.0,0.9969611763954163,0.0003373012295924127,0.9969611763954163,0.0003373012295924127,1.084521770477295,0.9994586110115051,0.2725689709186554,1.299863576889038,0.0037930156104266644,2026-04-11T21:41:02Z
|
| 1089 |
+
0.00893997447565198,0.00893997447565198,0.0,0.0,0.00893997447565198,0.0,29.0,29.0,27.125,27.125,26.0,26.0,0.09683473920449615,0.041203274637009575,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1067,10.166537284851074,6.76969696969697e-06,0.0034,train,2301589.0,0.9321423768997192,1.0,0.0,1.0,0.0,0.9321423768997192,0.007329464890062809,0.007329456973820925,0.9321423768997192,0.007329464890062809,0.9321423768997192,1.0,0.0,1.0,0.0,0.9321423768997192,0.007329464890062809,0.9321423768997192,0.007329464890062809,1.4033339023590088,1.0001667737960815,0.21619774401187897,1.5315618515014648,0.026033449918031693,2026-04-11T21:41:07Z
|
| 1090 |
+
0.0,0.0,0.0,0.0,0.0,0.0,65.0,65.0,65.0,65.0,65.0,65.0,0.0018052591913146898,0.041241890639481,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1068,0.0,6.7666666666666665e-06,0.0,train,2303405.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.0,0.9949578642845154,0.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.9949578642845154,0.0,1.0064785480499268,1.000160813331604,0.9977074861526489,0.0064575872384011745,0.00019105577666778117,2026-04-11T21:41:12Z
|
| 1091 |
+
0.0,0.0,0.0,0.0,0.0,0.0,65.0,65.0,65.0,65.0,65.0,65.0,0.005523053434444591,0.04128050664195242,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1069,0.0,6.763636363636365e-06,0.0,train,2305133.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.0,0.9949578642845154,0.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.9949578642845154,0.0,1.0369137525558472,1.0003167390823364,0.9796718955039978,0.036248765885829926,0.0005417983047664165,2026-04-11T21:41:17Z
|
| 1092 |
+
0.007499999832361937,0.007499999832361937,0.016467438312247396,0.016467438312247396,0.023967438144609332,0.0,54.0,54.0,52.625,52.625,50.0,50.0,0.05561392899835482,0.04131912264442385,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1070,26.325374603271484,6.760606060606061e-06,0.0381,train,2306762.0,0.11274315416812897,1.0,0.0,1.0,0.0,0.11274315416812897,0.3188244700431824,0.31882444024086,0.11274315416812897,0.3188244700431824,0.11274315416812897,1.0,0.0,1.0,0.0,0.11274315416812897,0.3188244700431824,0.11274315416812897,0.3188244700431824,2.0,1.0003687143325806,0.34076783061027527,1.0765538215637207,0.02075301483273506,2026-04-11T21:41:21Z
|
| 1093 |
+
0.0,0.0,0.0,0.0,0.0,0.0,26.0,26.0,25.875,25.875,25.0,25.0,0.02826369390822947,0.04135773864689527,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1071,5.607251167297363,6.757575757575758e-06,-0.0217,train,2308113.0,0.9891167879104614,1.0,0.0,1.0,0.0,0.9891167879104614,0.008078054524958134,0.008078045211732388,0.9891167879104614,0.008078054524958134,0.9891167879104614,1.0,0.0,1.0,0.0,0.9891167879104614,0.008078054524958134,0.9891167879104614,0.008078054524958134,1.0410070419311523,0.9988547563552856,0.5382347106933594,0.6194605827331543,0.009091660380363464,2026-04-11T21:41:26Z
|
| 1094 |
+
0.0,0.0,0.010146747343242168,0.010146747343242168,0.010146747343242168,0.0,62.0,62.0,61.625,61.625,61.0,61.0,0.10848722152877599,0.041396354649366696,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1072,2.4955992698669434,6.754545454545455e-06,0.0001,train,2309830.0,0.997157096862793,1.0,0.0,1.0,0.0,0.997157096862793,0.0005752869765274227,0.000575280690100044,0.997157096862793,0.0005752869765274227,0.997157096862793,1.0,0.0,1.0,0.0,0.997157096862793,0.0005752869765274227,0.997157096862793,0.0005752869765274227,1.3372715711593628,1.0032328367233276,0.5961686372756958,0.5172317028045654,0.012418783269822598,2026-04-11T21:41:31Z
|
| 1095 |
+
0.0,0.0,0.0036764706019312143,0.0036764706019312143,0.0036764706019312143,0.0,34.0,34.0,33.125,33.125,33.0,33.0,0.025314107653684914,0.04143497065183812,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1073,0.9153868556022644,6.751515151515152e-06,0.0115,train,2311287.0,0.9944090247154236,1.0,0.0,1.0,0.0,0.9944090247154236,0.0004998616641387343,0.0004998616059310734,0.9944090247154236,0.0004998616641387343,0.9944090247154236,1.0,0.0,1.0,0.0,0.9944090247154236,0.0004998616641387343,0.9944090247154236,0.0004998616641387343,1.9606788158416748,1.0061020851135254,0.9340354800224304,0.6732907295227051,0.005571494810283184,2026-04-11T21:41:36Z
|
| 1096 |
+
0.0024509804788976908,0.0024509804788976908,0.004464285913854837,0.004464285913854837,0.006915266392752528,0.0,56.0,56.0,54.125,54.125,51.0,51.0,0.03697874676436186,0.041473586654309544,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1074,3.6375036239624023,6.748484848484848e-06,0.0417,train,2313096.0,0.9875502586364746,1.0,0.0,1.0,0.0,0.9875502586364746,0.003662221832200885,0.003662231843918562,0.9875502586364746,0.003662221832200885,0.9875502586364746,1.0,0.0,1.0,0.0,0.9875502586364746,0.003662221832200885,0.9875502586364746,0.003662221832200885,1.2383108139038086,1.003567099571228,0.7975822687149048,0.22617030143737793,0.004612638149410486,2026-04-11T21:41:41Z
|
| 1097 |
+
0.0,0.0,0.03481511096470058,0.03481511096470058,0.03481511096470058,0.0,58.0,58.0,47.25,47.25,41.0,41.0,0.16682779975235462,0.04151220265678097,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1075,6.4112677574157715,6.7454545454545465e-06,0.0288,train,2314778.0,0.10002761334180832,1.0,0.0,1.0,0.0,0.10002761334180832,0.26244989037513733,0.26244989037513733,0.10002761334180832,0.26244989037513733,0.10002761334180832,1.0,0.0,1.0,0.0,0.10002761334180832,0.26244989037513733,0.10002761334180832,0.26244989037513733,1.4211242198944092,1.0019086599349976,0.2742016315460205,1.2938915491104126,0.03375272452831268,2026-04-11T21:41:46Z
|
| 1098 |
+
0.004631217801943421,0.004631217801943421,0.009479318046942353,0.009479318046942353,0.014110535848885775,0.0,55.0,55.0,53.125,53.125,52.0,52.0,0.049260836094617844,0.04155081865925239,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1076,4.325499057769775,6.742424242424243e-06,-0.0001,train,2316531.0,0.9805549383163452,1.0,0.0,1.0,0.0,0.9805549383163452,0.00861036404967308,0.008610363118350506,0.9805549383163452,0.00861036404967308,0.9805549383163452,1.0,0.0,1.0,0.0,0.9805549383163452,0.00861036404967308,0.9805549383163452,0.00861036404967308,1.5676127672195435,1.0017882585525513,0.45194947719573975,0.7941849231719971,0.011005771346390247,2026-04-11T21:41:51Z
|
| 1099 |
+
0.0,0.0,0.0,0.0,0.0,0.0,65.0,65.0,65.0,65.0,65.0,65.0,0.0024144309863913804,0.041589434661723816,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1077,0.0,6.73939393939394e-06,0.0,train,2318435.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.0,0.9949578642845154,0.0,0.9949578642845154,1.0,0.0,1.0,0.0,0.9949578642845154,0.0,0.9949578642845154,0.0,1.0065354108810425,1.000050663948059,0.9717419147491455,0.02866499498486519,0.000207486460567452,2026-04-11T21:41:56Z
|
| 1100 |
+
0.010216211201623082,0.010216211201623082,0.0078125,0.0078125,0.018028711201623082,0.0,64.0,64.0,61.375,61.375,59.0,59.0,0.13324717595241964,0.04162805066419524,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1078,4.65764856338501,6.7363636363636365e-06,0.0161,train,2320094.0,0.9684683084487915,1.0,0.0,1.0,0.0,0.9684683084487915,0.08065449446439743,0.08065447211265564,0.9684683084487915,0.08065449446439743,0.9684683084487915,1.0,0.0,1.0,0.0,0.9684683084487915,0.08065449446439743,0.9684683084487915,0.08065449446439743,1.5210373401641846,1.0008001327514648,0.20647554099559784,1.577573299407959,0.02014295756816864,2026-04-11T21:42:01Z
|
| 1101 |
+
0.0,0.0,0.0,0.0,0.0,0.0,61.0,61.0,49.5,49.5,46.0,46.0,0.054232243448495865,0.041666666666666664,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1079,0.0,6.733333333333334e-06,0.0,train,2321786.0,0.091847725212574,1.0,0.0,1.0,0.0,0.091847725212574,0.0,0.0,0.091847725212574,0.0,0.091847725212574,1.0,0.0,1.0,0.0,0.091847725212574,0.0,0.091847725212574,0.0,1.5695313215255737,0.9989585280418396,0.15552771091461182,1.860931396484375,0.01916210725903511,2026-04-11T21:42:06Z
|
| 1102 |
+
0.0019379844889044762,0.0019379844889044762,0.0027643622015602887,0.0027643622015602887,0.004702346690464765,0.0,137.0,137.0,132.5,132.5,129.0,129.0,0.029038164531812072,0.04170528266913809,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1080,1.9305036067962646,6.73030303030303e-06,0.0236,train,2324566.0,0.663239598274231,1.0,0.0,0.6666666865348816,0.0,0.9948593974113464,0.002004272071644664,0.0013361814199015498,0.663239598274231,0.0013361814199015498,0.663239598274231,1.0,0.0,0.6666666865348816,0.0,0.9948593974113464,0.002004272071644664,0.663239598274231,0.0013361814199015498,1.3798326253890991,1.0016835927963257,0.6128258109092712,0.48967456817626953,0.004363351967185736,2026-04-11T21:42:11Z
|
| 1103 |
+
0.002259009750559926,0.002259009750559926,0.001416441984474659,0.001416441984474659,0.003675451735034585,0.0,178.0,178.0,175.0,175.0,161.0,161.0,0.015698739560320973,0.04174389867160951,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1081,0.31140798330307007,6.7272727272727275e-06,0.0025,train,2327382.0,0.747491180896759,1.0,0.0,0.75,0.0,0.9966549277305603,7.61248666094616e-05,5.710850018658675e-05,0.747491180896759,5.710634286515415e-05,0.747491180896759,1.0,0.0,0.75,0.0,0.9966549277305603,7.61248666094616e-05,0.747491180896759,5.710634286515415e-05,1.509435772895813,0.9993287920951843,0.0263750609010458,3.635336399078369,0.0060366359539330006,2026-04-11T21:42:17Z
|
| 1104 |
+
0.009324767161160707,0.009324767161160707,0.0033783784601837397,0.0033783784601837397,0.012703145621344447,0.0,111.0,111.0,108.125,108.125,106.0,106.0,0.09740556543692946,0.04178251467408094,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1082,4.485145092010498,6.724242424242424e-06,0.0114,train,2329495.0,0.9212300777435303,1.0,0.0,1.0,0.0,0.9212300777435303,0.21582916378974915,0.21582916378974915,0.9212300777435303,0.21582916378974915,0.9212300777435303,1.0,0.0,1.0,0.0,0.9212300777435303,0.21582916378974915,0.9212300777435303,0.21582916378974915,1.5335668325424194,1.0012389421463013,0.38476866483688354,0.9551130533218384,0.013148986734449863,2026-04-11T21:42:23Z
|
| 1105 |
+
0.0,0.0,0.0,0.0,0.0,0.0,129.0,129.0,129.0,129.0,129.0,129.0,0.002667208347702399,0.04182113067655236,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1083,0.0,6.721212121212122e-06,0.0,train,2331783.0,0.6634209156036377,1.0,0.0,0.6666666865348816,0.0,0.9951313734054565,0.0,0.0,0.6634209156036377,0.0,0.6634209156036377,1.0,0.0,0.6666666865348816,0.0,0.9951313734054565,0.0,0.6634209156036377,0.0,1.0351933240890503,1.0000473260879517,0.9550737738609314,0.045966774225234985,0.0002578256244305521,2026-04-11T21:42:28Z
|
| 1106 |
+
0.04814814869314432,0.04814814869314432,0.012652947567403316,0.012652947567403316,0.06080109626054764,0.0,31.0,31.0,28.5,28.5,27.0,27.0,0.25689972564578056,0.041859746679023785,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1084,17.376197814941406,6.718181818181819e-06,0.0462,train,2333323.0,0.6803448796272278,1.0,0.0,1.0,0.0,0.6803448796272278,0.4452557861804962,0.4452557861804962,0.6803448796272278,0.4452557861804962,0.6803448796272278,1.0,0.0,1.0,0.0,0.6803448796272278,0.4452557861804962,0.6803448796272278,0.4452557861804962,2.0,0.9909398555755615,0.06887909024953842,2.6754026412963867,0.07098246365785599,2026-04-11T21:42:33Z
|
| 1107 |
+
0.0026115492219105363,0.0026115492219105363,0.002667596214450896,0.002667596214450896,0.005279145436361432,0.0,201.0,201.0,189.875,189.875,180.0,180.0,0.034014943055808544,0.04189836268149521,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1085,7.492648124694824,6.715151515151516e-06,-0.0223,train,2336298.0,0.3282319903373718,1.0,0.0,0.7083333730697632,0.07715165615081787,0.46110403537750244,0.25172099471092224,0.1757289618253708,0.3282319903373718,0.1757289618253708,0.3282319903373718,1.0,0.0,0.7083333730697632,0.07715165615081787,0.46110403537750244,0.25172099471092224,0.3282319903373718,0.1757289618253708,2.0,1.0021684169769287,0.14563584327697754,1.9266459941864014,0.010337450541555882,2026-04-11T21:42:40Z
|
| 1108 |
+
0.0006345177534967661,0.0006345177534967661,0.002358543104492128,0.002358543104492128,0.002993060857988894,0.0,213.0,213.0,208.25,208.25,197.0,197.0,0.008055432001128793,0.04193697868396663,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1086,3.549293279647827,6.712121212121213e-06,0.0308,train,2339268.0,0.6480287909507751,1.0,0.0,0.6500000357627869,0.09258200973272324,0.9969711303710938,8.204347250284627e-05,0.09227863699197769,0.6480287909507751,0.0922786295413971,0.6480287909507751,1.0,0.0,0.6500000357627869,0.09258200973272324,0.9969711303710938,8.204347250284627e-05,0.6480287909507751,0.0922786295413971,2.0,1.0001766681671143,0.00030800519743934274,8.085393905639648,0.010134638287127018,2026-04-11T21:42:46Z
|
| 1109 |
+
0.0015432098880410194,0.0015432098880410194,0.0,0.0,0.0015432098880410194,0.0,81.0,81.0,80.375,80.375,76.0,76.0,0.014752325252629817,0.04197559468643806,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1087,2.269176959991455,6.709090909090909e-06,-0.0111,train,2341095.0,0.9966411590576172,1.0,0.0,1.0,0.0,0.9966411590576172,0.00015710237494204193,0.00015711141168139875,0.9966411590576172,0.00015710237494204193,0.9966411590576172,1.0,0.0,1.0,0.0,0.9966411590576172,0.00015710237494204193,0.9966411590576172,0.00015710237494204193,1.1145241260528564,0.9995546340942383,0.45225203037261963,0.793515682220459,0.0034698506351560354,2026-04-11T21:42:51Z
|
| 1110 |
+
0.007653061067685485,0.007653061067685485,0.007452981313690543,0.007452981313690543,0.015106042381376028,0.0,51.0,51.0,49.25,49.25,49.0,49.0,0.09569864999502897,0.04201421068890948,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1088,12.239304542541504,6.706060606060607e-06,0.0123,train,2342713.0,0.986293613910675,1.0,0.0,1.0,0.0,0.986293613910675,0.0026903701946139336,0.0026903690304607153,0.986293613910675,0.0026903701946139336,0.986293613910675,1.0,0.0,1.0,0.0,0.986293613910675,0.0026903701946139336,0.986293613910675,0.0026903701946139336,1.9696929454803467,1.0027153491973877,0.23736277222633362,1.4381656646728516,0.016510505229234695,2026-04-11T21:42:57Z
|
| 1111 |
+
0.0,0.0,0.0,0.0,0.0,0.0,152.0,152.0,152.0,152.0,152.0,152.0,0.0034695666399784386,0.042052826691380905,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1089,0.0,6.703030303030304e-06,0.0,train,2345457.0,0.7476505637168884,1.0,0.0,0.75,0.0,0.9968674182891846,0.0,0.0,0.7476505637168884,0.0,0.7476505637168884,1.0,0.0,0.75,0.0,0.9968674182891846,0.0,0.7476505637168884,0.0,1.08553946018219,0.9997900128364563,0.7077293992042542,0.3456934988498688,0.0006623952649533749,2026-04-11T21:43:03Z
|
| 1112 |
+
0.004788926860783249,0.004788926860783249,0.0,0.0,0.004788926860783249,0.0,183.0,183.0,182.5,182.5,181.0,181.0,0.04286058805882931,0.04209144269385233,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1090,0.4215506911277771,6.700000000000001e-06,-0.0003,train,2348309.0,0.9972051978111267,1.0,0.0,1.0,0.0,0.9972051978111267,5.313828296493739e-05,5.313552901498042e-05,0.9972051978111267,5.313828296493739e-05,0.9972051978111267,1.0,0.0,1.0,0.0,0.9972051978111267,5.313828296493739e-05,0.9972051978111267,5.313828296493739e-05,1.4445006847381592,1.0013642311096191,0.4382810592651367,0.824894905090332,0.004990379326045513,2026-04-11T21:43:09Z
|
| 1113 |
+
0.0,0.0,0.0,0.0,0.0,0.0,56.0,56.0,47.25,47.25,46.0,46.0,0.016990114585496485,0.042130058696323754,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1091,0.0,6.6969696969696975e-06,0.0,train,2349895.0,0.091847725212574,1.0,0.0,1.0,0.0,0.091847725212574,0.0,0.0,0.091847725212574,0.0,0.091847725212574,1.0,0.0,1.0,0.0,0.091847725212574,0.0,0.091847725212574,0.0,1.5886342525482178,1.001801609992981,0.511188268661499,0.6710173487663269,0.007502844091504812,2026-04-11T21:43:14Z
|
| 1114 |
+
0.0,0.0,0.0,0.0,0.0,0.0,32.0,32.0,32.0,32.0,32.0,32.0,0.010097296675667167,0.04216867469879518,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1092,0.0,6.693939393939395e-06,0.0,train,2351375.0,0.9955702424049377,1.0,0.0,1.0,0.0,0.9955702424049377,0.0,0.0,0.9955702424049377,0.0,0.9955702424049377,1.0,0.0,1.0,0.0,0.9955702424049377,0.0,0.9955702424049377,0.0,1.054145336151123,0.9996309876441956,0.8423091173171997,0.17160826921463013,0.001951782964169979,2026-04-11T21:43:18Z
|
| 1115 |
+
0.0,0.0,0.003846153849735856,0.003846153849735856,0.003846153849735856,0.0,65.0,65.0,64.75,64.75,63.0,63.0,0.013940075354184955,0.0422072907012666,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1093,1.6124061346054077,6.690909090909091e-06,0.0095,train,2353053.0,0.9950487613677979,1.0,0.0,1.0,0.0,0.9950487613677979,0.0002570115029811859,0.00025701147387735546,0.9950487613677979,0.0002570115029811859,0.9950487613677979,1.0,0.0,1.0,0.0,0.9950487613677979,0.0002570115029811859,0.9950487613677979,0.0002570115029811859,1.0859289169311523,0.9983689785003662,0.21541017293930054,1.5352113246917725,0.005157488863915205,2026-04-11T21:43:23Z
|
| 1116 |
+
0.00390625,0.00390625,0.0038470644503831863,0.0038470644503831863,0.007753314450383186,0.0,66.0,66.0,65.125,65.125,64.0,64.0,0.04774620302487165,0.042245906703738026,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1094,3.3892717361450195,6.687878787878788e-06,0.0094,train,2354894.0,0.9951438903808594,1.0,0.0,1.0,0.0,0.9951438903808594,0.0001731093943817541,0.00017311105330009013,0.9951438903808594,0.0001731093943817541,0.9951438903808594,1.0,0.0,1.0,0.0,0.9951438903808594,0.0001731093943817541,0.9951438903808594,0.0001731093943817541,1.232495665550232,0.996876060962677,0.19408656656742096,1.639451026916504,0.011436189524829388,2026-04-11T21:43:28Z
|
| 1117 |
+
0.0,0.0,0.0,0.0,0.0,0.0,107.0,107.0,107.0,107.0,107.0,107.0,0.003691258083563298,0.04228452270620945,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1095,0.0,6.684848484848485e-06,0.0,train,2357158.0,0.9966511130332947,1.0,0.0,1.0,0.0,0.9966511130332947,0.0,0.0,0.9966511130332947,0.0,0.9966511130332947,1.0,0.0,1.0,0.0,0.9966511130332947,0.0,0.9966511130332947,0.0,1.083788275718689,0.999952495098114,0.8022630214691162,0.2203187644481659,0.0007139050285331905,2026-04-11T21:43:33Z
|
| 1118 |
+
0.0223214291036129,0.0223214291036129,0.03616698645055294,0.03616698645055294,0.05848841555416584,0.0,29.0,29.0,26.875,26.875,20.0,20.0,0.21645524073392153,0.042323138708680874,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1096,17.096412658691406,6.681818181818183e-06,0.0528,train,2359061.0,0.11241252720355988,1.0,0.0,1.0,0.0,0.11241252720355988,0.3034849762916565,0.3034849464893341,0.11241252720355988,0.3034849762916565,0.11241252720355988,1.0,0.0,1.0,0.0,0.11241252720355988,0.3034849762916565,0.11241252720355988,0.3034849762916565,1.9887901544570923,0.979833722114563,0.010107050649821758,4.594521999359131,0.10703544318675995,2026-04-11T21:43:38Z
|
| 1119 |
+
0.0025862068869173527,0.0025862068869173527,0.000776397529989481,0.000776397529989481,0.0033626044169068336,0.0,161.0,161.0,147.125,147.125,145.0,145.0,0.016805213410407305,0.0423617547111523,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1097,0.9273874163627625,6.678787878787879e-06,0.0354,train,2361662.0,0.9652670621871948,1.0,0.0,0.96875,0.0883883461356163,0.9964081048965454,0.00033037204411812127,0.08803740888834,0.9652670621871948,0.08803740888834,0.9652670621871948,1.0,0.0,0.96875,0.0883883461356163,0.9964081048965454,0.00033037204411812127,0.9652670621871948,0.08803740888834,2.0,1.0020689964294434,0.4937727451324463,1.9974420070648193,0.0055220527574419975,2026-04-11T21:43:44Z
|
| 1120 |
+
0.0,0.0,0.0,0.0,0.0,0.0,167.0,167.0,167.0,167.0,167.0,167.0,0.0018285177211510018,0.04240037071362372,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,1.0,1098,0.0,6.6757575757575766e-06,0.0,train,2364710.0,0.8307228684425354,1.0,0.0,0.8333333134651184,0.0,0.9968674182891846,0.0,0.0,0.8307228684425354,0.0,0.8307228684425354,1.0,0.0,0.8333333134651184,0.0,0.9968674182891846,0.0,0.8307228684425354,0.0,1.0559391975402832,1.0000667572021484,0.8984145522117615,0.1071237251162529,0.00022596628696192056,2026-04-11T21:43:50Z
|
| 1121 |
+
0.0,0.0,0.00657894741743803,0.00657894741743803,0.00657894741743803,0.0,57.0,57.0,48.375,48.375,46.0,46.0,0.030019621714018285,0.042438986716095146,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1099,6.627462863922119,6.672727272727273e-06,-0.0331,train,2366553.0,0.18795685470104218,1.0,0.0,1.0,0.0,0.18795685470104218,0.3101930618286133,0.3101930618286133,0.18795685470104218,0.3101930618286133,0.18795685470104218,1.0,0.0,1.0,0.0,0.18795685470104218,0.3101930618286133,0.18795685470104218,0.3101930618286133,1.6379154920578003,1.0062874555587769,0.3873036205768585,0.9485464096069336,0.010162265971302986,2026-04-11T21:43:55Z
|
| 1122 |
+
0.0015337422955781221,0.0015337422955781221,0.0016737572732381523,0.0016737572732381523,0.0032074995688162744,0.0,163.0,163.0,156.0,156.0,145.0,145.0,0.030738614965230227,0.04247760271856657,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1100,2.0620083808898926,6.66969696969697e-06,-0.0306,train,2369329.0,0.7980826497077942,1.0,0.0,0.800000011920929,0.0,0.9976032972335815,0.0007398881716653705,0.0005919178947806358,0.7980826497077942,0.0005919100949540734,0.7980826497077942,1.0,0.0,0.800000011920929,0.0,0.9976032972335815,0.0007398881716653705,0.7980826497077942,0.0005919100949540734,1.730522871017456,1.0009613037109375,0.5152587294578552,0.6630861759185791,0.003971020225435495,2026-04-11T21:44:01Z
|
| 1123 |
+
,,,,,,,,,,,,,0.04247760271856657,0.0,0.0,0.0,0.0,0.0,0.028846153846153848,390.15384615384613,341.7692307692308,177.95192307692307,167.59066126896784,57.38461538461539,57.38461538461539,0.022705065874526136,0.0,nan,2369329.0,0.5808727557842548,0.9807692307692307,0.05439282839114849,0.8606285040195172,0.14204404417138833,0.6938754916191101,0.4057489140675618,nan,0.5808727557842548,0.372596684556741,0.5808727557842548,0.9807692307692307,0.05439282839114849,0.8606285040195172,0.14204404417138833,0.6938754916191101,0.4057489140675618,0.5808727557842548,0.372596684556741,73.9179,1.407,1.281661501297584,1.000536437218006,0.542173194197508,0.6837713030668405,0.0031142185639160182,0.176,,1100,,,,eval,,,,,,,,,,,,,,,,,,,,,,,,,,2026-04-11T21:45:15Z
|
| 1124 |
+
0.023881223052740097,0.023881223052740097,0.01245915051549673,0.01245915051549673,0.03634037356823683,0.0,73.0,73.0,62.0,62.0,18.0,18.0,0.2632972849532962,0.042516218721038,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,0.0,1101,5.9600019454956055,6.666666666666667e-06,-0.2286,train,2371113.0,0.772131621837616,1.0,0.0,0.875,0.3535533845424652,0.772131621837616,0.34598708152770996,0.34598708152770996,0.772131621837616,0.34598708152770996,0.772131621837616,1.0,0.0,0.875,0.3535533845424652,0.772131621837616,0.34598708152770996,0.772131621837616,0.34598708152770996,1.7506089210510254,1.0026087760925293,0.2138272076845169,1.5425870418548584,0.040163177996873856,2026-04-11T21:45:24Z
|
metrics.jsonl
CHANGED
|
@@ -1070,3 +1070,54 @@
|
|
| 1070 |
{"timestamp_utc": "2026-04-11T21:38:09Z", "mode": "train", "global_step": 1050, "epoch": 0.040546802594995365, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.821212121212122e-06, "num_tokens": 2265729.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9968419075012207, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9968419075012207, "rewards/total_composite/std": 0.0, "reward": 0.9968419075012207, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0027122944593429565, "sampling/sampling_logp_difference/max": 0.7146163582801819, "sampling/importance_sampling_ratio/min": 0.4893798530101776, "sampling/importance_sampling_ratio/mean": 0.999724805355072, "sampling/importance_sampling_ratio/max": 1.1012705564498901, "entropy": 0.007037240080535412, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9968419075012207, "reward_meter_mean": 0.9968419075012207, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9968419075012207, "reward_total_composite_std": 0.0}
|
| 1071 |
{"timestamp_utc": "2026-04-11T21:39:35Z", "mode": "eval", "global_step": 1050, "epoch": 0.040546802594995365, "eval_loss": NaN, "eval_runtime": 85.4042, "eval_samples_per_second": 1.218, "eval_steps_per_second": 0.152, "eval_num_tokens": 2265729.0, "eval_completions/mean_length": 241.28846153846155, "eval_completions/min_length": 59.61538461538461, "eval_completions/max_length": 458.0, "eval_completions/clipped_ratio": 0.057692307692307696, "eval_completions/mean_terminated_length": 224.83516869178186, "eval_completions/min_terminated_length": 59.61538461538461, "eval_completions/max_terminated_length": 412.46153846153845, "eval_rewards/meter/mean": 0.6695947761719043, "eval_rewards/meter/std": 0.4209407364519743, "eval_rewards/count_adherence/mean": 0.8243002341343806, "eval_rewards/count_adherence/std": 0.149119944526599, "eval_rewards/arabic_clean/mean": 0.9903846153846154, "eval_rewards/arabic_clean/std": 0.027196414195574246, "eval_rewards/total_composite/mean": 0.5442026945260855, "eval_rewards/total_composite/std": 0.3753266856074333, "eval_reward": 0.5442026945260855, "eval_reward_std": NaN, "eval_frac_reward_zero_std": 0.0, "eval_sampling/sampling_logp_difference/mean": 0.0014539593382953452, "eval_sampling/sampling_logp_difference/max": 0.48829063085409313, "eval_sampling/importance_sampling_ratio/min": 0.6470333154384906, "eval_sampling/importance_sampling_ratio/mean": 1.0002473134260912, "eval_sampling/importance_sampling_ratio/max": 1.2309852104920607, "eval_entropy": 0.009740084463443894, "eval_clip_ratio/low_mean": 0.0, "eval_clip_ratio/low_min": 0.0, "eval_clip_ratio/high_mean": 0.0, "eval_clip_ratio/high_max": 0.0, "eval_clip_ratio/region_mean": 0.0, "eval_reward_total_mean": 0.5442026945260855, "eval_reward_meter_mean": 0.6695947761719043, "eval_reward_meter_std": 0.4209407364519743, "eval_reward_count_adherence_mean": 0.8243002341343806, "eval_reward_count_adherence_std": 0.149119944526599, "eval_reward_arabic_clean_mean": 0.9903846153846154, "eval_reward_arabic_clean_std": 0.027196414195574246, "eval_reward_total_composite_mean": 0.5442026945260855, "eval_reward_total_composite_std": 0.3753266856074333}
|
| 1072 |
{"timestamp_utc": "2026-04-11T21:39:43Z", "mode": "train", "global_step": 1051, "epoch": 0.04058541859746679, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.818181818181818e-06, "num_tokens": 2267377.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00045387333375401795, "sampling/sampling_logp_difference/max": 0.042022280395030975, "sampling/importance_sampling_ratio/min": 0.9996206760406494, "sampling/importance_sampling_ratio/mean": 1.0004585981369019, "sampling/importance_sampling_ratio/max": 1.0429177284240723, "entropy": 0.004360992228612304, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1070 |
{"timestamp_utc": "2026-04-11T21:38:09Z", "mode": "train", "global_step": 1050, "epoch": 0.040546802594995365, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.821212121212122e-06, "num_tokens": 2265729.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9968419075012207, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9968419075012207, "rewards/total_composite/std": 0.0, "reward": 0.9968419075012207, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0027122944593429565, "sampling/sampling_logp_difference/max": 0.7146163582801819, "sampling/importance_sampling_ratio/min": 0.4893798530101776, "sampling/importance_sampling_ratio/mean": 0.999724805355072, "sampling/importance_sampling_ratio/max": 1.1012705564498901, "entropy": 0.007037240080535412, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9968419075012207, "reward_meter_mean": 0.9968419075012207, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9968419075012207, "reward_total_composite_std": 0.0}
|
| 1071 |
{"timestamp_utc": "2026-04-11T21:39:35Z", "mode": "eval", "global_step": 1050, "epoch": 0.040546802594995365, "eval_loss": NaN, "eval_runtime": 85.4042, "eval_samples_per_second": 1.218, "eval_steps_per_second": 0.152, "eval_num_tokens": 2265729.0, "eval_completions/mean_length": 241.28846153846155, "eval_completions/min_length": 59.61538461538461, "eval_completions/max_length": 458.0, "eval_completions/clipped_ratio": 0.057692307692307696, "eval_completions/mean_terminated_length": 224.83516869178186, "eval_completions/min_terminated_length": 59.61538461538461, "eval_completions/max_terminated_length": 412.46153846153845, "eval_rewards/meter/mean": 0.6695947761719043, "eval_rewards/meter/std": 0.4209407364519743, "eval_rewards/count_adherence/mean": 0.8243002341343806, "eval_rewards/count_adherence/std": 0.149119944526599, "eval_rewards/arabic_clean/mean": 0.9903846153846154, "eval_rewards/arabic_clean/std": 0.027196414195574246, "eval_rewards/total_composite/mean": 0.5442026945260855, "eval_rewards/total_composite/std": 0.3753266856074333, "eval_reward": 0.5442026945260855, "eval_reward_std": NaN, "eval_frac_reward_zero_std": 0.0, "eval_sampling/sampling_logp_difference/mean": 0.0014539593382953452, "eval_sampling/sampling_logp_difference/max": 0.48829063085409313, "eval_sampling/importance_sampling_ratio/min": 0.6470333154384906, "eval_sampling/importance_sampling_ratio/mean": 1.0002473134260912, "eval_sampling/importance_sampling_ratio/max": 1.2309852104920607, "eval_entropy": 0.009740084463443894, "eval_clip_ratio/low_mean": 0.0, "eval_clip_ratio/low_min": 0.0, "eval_clip_ratio/high_mean": 0.0, "eval_clip_ratio/high_max": 0.0, "eval_clip_ratio/region_mean": 0.0, "eval_reward_total_mean": 0.5442026945260855, "eval_reward_meter_mean": 0.6695947761719043, "eval_reward_meter_std": 0.4209407364519743, "eval_reward_count_adherence_mean": 0.8243002341343806, "eval_reward_count_adherence_std": 0.149119944526599, "eval_reward_arabic_clean_mean": 0.9903846153846154, "eval_reward_arabic_clean_std": 0.027196414195574246, "eval_reward_total_composite_mean": 0.5442026945260855, "eval_reward_total_composite_std": 0.3753266856074333}
|
| 1072 |
{"timestamp_utc": "2026-04-11T21:39:43Z", "mode": "train", "global_step": 1051, "epoch": 0.04058541859746679, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.818181818181818e-06, "num_tokens": 2267377.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00045387333375401795, "sampling/sampling_logp_difference/max": 0.042022280395030975, "sampling/importance_sampling_ratio/min": 0.9996206760406494, "sampling/importance_sampling_ratio/mean": 1.0004585981369019, "sampling/importance_sampling_ratio/max": 1.0429177284240723, "entropy": 0.004360992228612304, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0}
|
| 1073 |
+
{"timestamp_utc": "2026-04-11T21:39:48Z", "mode": "train", "global_step": 1052, "epoch": 0.040624034599938214, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.8151515151515155e-06, "num_tokens": 2269065.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9968419075012207, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9968419075012207, "rewards/total_composite/std": 0.0, "reward": 0.9968419075012207, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0007426206138916314, "sampling/sampling_logp_difference/max": 0.09673504531383514, "sampling/importance_sampling_ratio/min": 0.9077965617179871, "sampling/importance_sampling_ratio/mean": 1.0001823902130127, "sampling/importance_sampling_ratio/max": 1.0188950300216675, "entropy": 0.004134227987378836, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9968419075012207, "reward_meter_mean": 0.9968419075012207, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9968419075012207, "reward_total_composite_std": 0.0}
|
| 1074 |
+
{"timestamp_utc": "2026-04-11T21:39:53Z", "mode": "train", "global_step": 1053, "epoch": 0.04066265060240964, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.812121212121212e-06, "num_tokens": 2270753.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00017959998513106257, "sampling/sampling_logp_difference/max": 0.011735539883375168, "sampling/importance_sampling_ratio/min": 0.9999438524246216, "sampling/importance_sampling_ratio/mean": 1.0001797676086426, "sampling/importance_sampling_ratio/max": 1.0118045806884766, "entropy": 0.0012685270776273683, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0}
|
| 1075 |
+
{"timestamp_utc": "2026-04-11T21:39:57Z", "mode": "train", "global_step": 1054, "epoch": 0.04070126660488106, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.80909090909091e-06, "num_tokens": 2272529.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00012025728210574016, "sampling/sampling_logp_difference/max": 0.008706126362085342, "sampling/importance_sampling_ratio/min": 0.9994588494300842, "sampling/importance_sampling_ratio/mean": 1.000115990638733, "sampling/importance_sampling_ratio/max": 1.008744239807129, "entropy": 0.0013022492494201288, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0}
|
| 1076 |
+
{"timestamp_utc": "2026-04-11T21:40:02Z", "mode": "train", "global_step": 1055, "epoch": 0.040739882607352486, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.806060606060607e-06, "num_tokens": 2274175.0, "completions/mean_length": 50.75, "completions/min_length": 50.0, "completions/max_length": 51.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 50.75, "completions/min_terminated_length": 50.0, "completions/max_terminated_length": 51.0, "rewards/meter/mean": 0.9589456915855408, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9589456915855408, "rewards/total_composite/std": 0.0, "reward": 0.9589456915855408, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0029871375299990177, "sampling/sampling_logp_difference/max": 0.30701398849487305, "sampling/importance_sampling_ratio/min": 0.7356403470039368, "sampling/importance_sampling_ratio/mean": 1.0008618831634521, "sampling/importance_sampling_ratio/max": 1.2294468879699707, "entropy": 0.0185652831569314, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9589456915855408, "reward_meter_mean": 0.9589456915855408, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9589456915855408, "reward_total_composite_std": 0.0}
|
| 1077 |
+
{"timestamp_utc": "2026-04-11T21:40:07Z", "mode": "train", "global_step": 1056, "epoch": 0.04077849860982391, "loss": 0.008, "grad_norm": 3.9809470176696777, "learning_rate": 6.803030303030304e-06, "num_tokens": 2275998.0, "completions/mean_length": 62.875, "completions/min_length": 61.0, "completions/max_length": 64.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 62.875, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 64.0, "rewards/meter/mean": 0.0021876515820622444, "rewards/meter/std": 0.0002445352729409933, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.0021876515820622444, "rewards/total_composite/std": 0.0002445352729409933, "reward": 0.0021876515820622444, "reward_std": 0.00024453524383716285, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.008529768325388432, "sampling/sampling_logp_difference/max": 1.1146979331970215, "sampling/importance_sampling_ratio/min": 0.6839931607246399, "sampling/importance_sampling_ratio/mean": 1.0037806034088135, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.041025768499821424, "clip_ratio/low_mean": 0.013767930213361979, "clip_ratio/low_min": 0.013767930213361979, "clip_ratio/high_mean": 0.0020491802133619785, "clip_ratio/high_max": 0.0020491802133619785, "clip_ratio/region_mean": 0.015817110426723957, "reward_total_mean": 0.0021876515820622444, "reward_meter_mean": 0.0021876515820622444, "reward_meter_std": 0.0002445352729409933, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.0021876515820622444, "reward_total_composite_std": 0.0002445352729409933}
|
| 1078 |
+
{"timestamp_utc": "2026-04-11T21:40:12Z", "mode": "train", "global_step": 1057, "epoch": 0.040817114612295334, "loss": -0.0561, "grad_norm": 0.9426341652870178, "learning_rate": 6.800000000000001e-06, "num_tokens": 2277958.0, "completions/mean_length": 83.0, "completions/min_length": 81.0, "completions/max_length": 97.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 83.0, "completions/min_terminated_length": 81.0, "completions/max_terminated_length": 97.0, "rewards/meter/mean": 0.9951313734054565, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.7083333730697632, "rewards/count_adherence/std": 0.11785111576318741, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7048847079277039, "rewards/total_composite/std": 0.1172773540019989, "reward": 0.7048847079277039, "reward_std": 0.1172773465514183, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.00949889700859785, "sampling/sampling_logp_difference/max": 3.2940514087677, "sampling/importance_sampling_ratio/min": 0.03710322454571724, "sampling/importance_sampling_ratio/mean": 0.9974780082702637, "sampling/importance_sampling_ratio/max": 1.134204387664795, "entropy": 0.003993239166447893, "clip_ratio/low_mean": 0.0015432098880410194, "clip_ratio/low_min": 0.0015432098880410194, "clip_ratio/high_mean": 0.0012886597542092204, "clip_ratio/high_max": 0.0012886597542092204, "clip_ratio/region_mean": 0.00283186964225024, "reward_total_mean": 0.7048847079277039, "reward_meter_mean": 0.9951313734054565, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.7083333730697632, "reward_count_adherence_std": 0.11785111576318741, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7048847079277039, "reward_total_composite_std": 0.1172773540019989}
|
| 1079 |
+
{"timestamp_utc": "2026-04-11T21:40:19Z", "mode": "train", "global_step": 1058, "epoch": 0.04085573061476676, "loss": 0.0039, "grad_norm": 0.21764229238033295, "learning_rate": 6.796969696969697e-06, "num_tokens": 2280914.0, "completions/mean_length": 198.5, "completions/min_length": 197.0, "completions/max_length": 209.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 198.5, "completions/min_terminated_length": 197.0, "completions/max_terminated_length": 209.0, "rewards/meter/mean": 0.9969083070755005, "rewards/meter/std": 3.72788890672382e-05, "rewards/count_adherence/mean": 0.800000011920929, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7975266575813293, "rewards/total_composite/std": 2.9818897019140422e-05, "reward": 0.7975266575813293, "reward_std": 2.9827928301529028e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.003239656798541546, "sampling/sampling_logp_difference/max": 3.510453224182129, "sampling/importance_sampling_ratio/min": 0.0298833679407835, "sampling/importance_sampling_ratio/mean": 0.9991796016693115, "sampling/importance_sampling_ratio/max": 1.167437195777893, "entropy": 0.0015565359972242732, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0006345177534967661, "clip_ratio/high_max": 0.0006345177534967661, "clip_ratio/region_mean": 0.0006345177534967661, "reward_total_mean": 0.7975266575813293, "reward_meter_mean": 0.9969083070755005, "reward_meter_std": 3.72788890672382e-05, "reward_count_adherence_mean": 0.800000011920929, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7975266575813293, "reward_total_composite_std": 2.9818897019140422e-05}
|
| 1080 |
+
{"timestamp_utc": "2026-04-11T21:40:23Z", "mode": "train", "global_step": 1059, "epoch": 0.04089434661723818, "loss": 0.0234, "grad_norm": 4.7054972648620605, "learning_rate": 6.793939393939395e-06, "num_tokens": 2282612.0, "completions/mean_length": 54.25, "completions/min_length": 52.0, "completions/max_length": 57.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 54.25, "completions/min_terminated_length": 52.0, "completions/max_terminated_length": 57.0, "rewards/meter/mean": 0.9560332298278809, "rewards/meter/std": 0.0188444871455431, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9560332298278809, "rewards/total_composite/std": 0.0188444871455431, "reward": 0.9560332298278809, "reward_std": 0.018844490870833397, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.008652369491755962, "sampling/sampling_logp_difference/max": 0.96638023853302, "sampling/importance_sampling_ratio/min": 0.38045772910118103, "sampling/importance_sampling_ratio/mean": 0.9992375373840332, "sampling/importance_sampling_ratio/max": 1.3655518293380737, "entropy": 0.025705090374685824, "clip_ratio/low_mean": 0.00657894741743803, "clip_ratio/low_min": 0.00657894741743803, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.00657894741743803, "reward_total_mean": 0.9560332298278809, "reward_meter_mean": 0.9560332298278809, "reward_meter_std": 0.0188444871455431, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9560332298278809, "reward_total_composite_std": 0.0188444871455431}
|
| 1081 |
+
{"timestamp_utc": "2026-04-11T21:40:31Z", "mode": "train", "global_step": 1060, "epoch": 0.040932962619709606, "loss": 0.0057, "grad_norm": 0.09687843173742294, "learning_rate": 6.790909090909091e-06, "num_tokens": 2286469.0, "completions/mean_length": 274.125, "completions/min_length": 272.0, "completions/max_length": 289.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 274.125, "completions/min_terminated_length": 272.0, "completions/max_terminated_length": 289.0, "rewards/meter/mean": 0.9969933032989502, "rewards/meter/std": 5.181956657906994e-05, "rewards/count_adherence/mean": 0.875, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8723691701889038, "rewards/total_composite/std": 4.535001062322408e-05, "reward": 0.8723691701889038, "reward_std": 4.535001062322408e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.00142471503932029, "sampling/sampling_logp_difference/max": 1.230623722076416, "sampling/importance_sampling_ratio/min": 0.2921103239059448, "sampling/importance_sampling_ratio/mean": 0.9990590214729309, "sampling/importance_sampling_ratio/max": 1.0306644439697266, "entropy": 0.0006548625879077008, "clip_ratio/low_mean": 0.00043252596515230834, "clip_ratio/low_min": 0.00043252596515230834, "clip_ratio/high_mean": 0.0018382353009656072, "clip_ratio/high_max": 0.0018382353009656072, "clip_ratio/region_mean": 0.0022707612661179155, "reward_total_mean": 0.8723691701889038, "reward_meter_mean": 0.9969933032989502, "reward_meter_std": 5.181956657906994e-05, "reward_count_adherence_mean": 0.875, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8723691701889038, "reward_total_composite_std": 4.535001062322408e-05}
|
| 1082 |
+
{"timestamp_utc": "2026-04-11T21:40:38Z", "mode": "train", "global_step": 1061, "epoch": 0.04097157862218103, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.787878787878789e-06, "num_tokens": 2290485.0, "completions/mean_length": 272.0, "completions/min_length": 272.0, "completions/max_length": 272.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 272.0, "completions/min_terminated_length": 272.0, "completions/max_terminated_length": 272.0, "rewards/meter/mean": 0.997011661529541, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.875, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8723852038383484, "rewards/total_composite/std": 0.0, "reward": 0.8723852038383484, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 6.0048791056033224e-05, "sampling/sampling_logp_difference/max": 0.03515136241912842, "sampling/importance_sampling_ratio/min": 0.9654592871665955, "sampling/importance_sampling_ratio/mean": 0.9999895691871643, "sampling/importance_sampling_ratio/max": 1.0054314136505127, "entropy": 0.00023714406961516943, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.8723852038383484, "reward_meter_mean": 0.997011661529541, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.875, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8723852038383484, "reward_total_composite_std": 0.0}
|
| 1083 |
+
{"timestamp_utc": "2026-04-11T21:40:44Z", "mode": "train", "global_step": 1062, "epoch": 0.041010194624652455, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.7848484848484855e-06, "num_tokens": 2292821.0, "completions/mean_length": 106.0, "completions/min_length": 106.0, "completions/max_length": 106.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 106.0, "completions/min_terminated_length": 106.0, "completions/max_terminated_length": 106.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.01337476447224617, "sampling/sampling_logp_difference/max": 7.350694179534912, "sampling/importance_sampling_ratio/min": 0.0006421464495360851, "sampling/importance_sampling_ratio/mean": 0.9978539347648621, "sampling/importance_sampling_ratio/max": 1.0233614444732666, "entropy": 0.0023513801133958623, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0}
|
| 1084 |
+
{"timestamp_utc": "2026-04-11T21:40:48Z", "mode": "train", "global_step": 1063, "epoch": 0.04104881062712388, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.781818181818183e-06, "num_tokens": 2294669.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00023568868346046656, "sampling/sampling_logp_difference/max": 0.017239108681678772, "sampling/importance_sampling_ratio/min": 0.9829086661338806, "sampling/importance_sampling_ratio/mean": 1.0001155138015747, "sampling/importance_sampling_ratio/max": 1.011854648590088, "entropy": 0.0019514950545271859, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0}
|
| 1085 |
+
{"timestamp_utc": "2026-04-11T21:40:53Z", "mode": "train", "global_step": 1064, "epoch": 0.0410874266295953, "loss": -0.0026, "grad_norm": 2.7271664142608643, "learning_rate": 6.778787878787879e-06, "num_tokens": 2296475.0, "completions/mean_length": 54.75, "completions/min_length": 53.0, "completions/max_length": 57.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 54.75, "completions/min_terminated_length": 53.0, "completions/max_terminated_length": 57.0, "rewards/meter/mean": 0.9718772172927856, "rewards/meter/std": 0.01329784281551838, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9718772172927856, "rewards/total_composite/std": 0.01329784281551838, "reward": 0.9718772172927856, "reward_std": 0.013297837227582932, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.024222884327173233, "sampling/sampling_logp_difference/max": 5.190197944641113, "sampling/importance_sampling_ratio/min": 0.005570904351770878, "sampling/importance_sampling_ratio/mean": 0.9998058676719666, "sampling/importance_sampling_ratio/max": 1.5688531398773193, "entropy": 0.10237977746874094, "clip_ratio/low_mean": 0.009181267116218805, "clip_ratio/low_min": 0.009181267116218805, "clip_ratio/high_mean": 0.002314814832061529, "clip_ratio/high_max": 0.002314814832061529, "clip_ratio/region_mean": 0.011496081948280334, "reward_total_mean": 0.9718772172927856, "reward_meter_mean": 0.9718772172927856, "reward_meter_std": 0.01329784281551838, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9718772172927856, "reward_total_composite_std": 0.01329784281551838}
|
| 1086 |
+
{"timestamp_utc": "2026-04-11T21:40:58Z", "mode": "train", "global_step": 1065, "epoch": 0.04112604263206673, "loss": 0.0102, "grad_norm": 5.474052429199219, "learning_rate": 6.7757575757575765e-06, "num_tokens": 2298309.0, "completions/mean_length": 57.25, "completions/min_length": 57.0, "completions/max_length": 59.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 57.25, "completions/min_terminated_length": 57.0, "completions/max_terminated_length": 59.0, "rewards/meter/mean": 0.8288533687591553, "rewards/meter/std": 0.037421029061079025, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8288533687591553, "rewards/total_composite/std": 0.037421029061079025, "reward": 0.8288533687591553, "reward_std": 0.03742102161049843, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.002723454497754574, "sampling/sampling_logp_difference/max": 0.5241236686706543, "sampling/importance_sampling_ratio/min": 0.5920739769935608, "sampling/importance_sampling_ratio/mean": 0.9999019503593445, "sampling/importance_sampling_ratio/max": 1.1381598711013794, "entropy": 0.01174389524385333, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.8288533687591553, "reward_meter_mean": 0.8288533687591553, "reward_meter_std": 0.037421029061079025, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8288533687591553, "reward_total_composite_std": 0.037421029061079025}
|
| 1087 |
+
{"timestamp_utc": "2026-04-11T21:41:02Z", "mode": "train", "global_step": 1066, "epoch": 0.04116465863453815, "loss": 0.0097, "grad_norm": 2.398463726043701, "learning_rate": 6.772727272727273e-06, "num_tokens": 2300092.0, "completions/mean_length": 60.875, "completions/min_length": 60.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 60.875, "completions/min_terminated_length": 60.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9969611763954163, "rewards/meter/std": 0.0003373012295924127, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9969611763954163, "rewards/total_composite/std": 0.0003373012295924127, "reward": 0.9969611763954163, "reward_std": 0.00033730725408531725, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0037930156104266644, "sampling/sampling_logp_difference/max": 1.299863576889038, "sampling/importance_sampling_ratio/min": 0.2725689709186554, "sampling/importance_sampling_ratio/mean": 0.9994586110115051, "sampling/importance_sampling_ratio/max": 1.084521770477295, "entropy": 0.014361659123096615, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0020833334419876337, "clip_ratio/high_max": 0.0020833334419876337, "clip_ratio/region_mean": 0.0020833334419876337, "reward_total_mean": 0.9969611763954163, "reward_meter_mean": 0.9969611763954163, "reward_meter_std": 0.0003373012295924127, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9969611763954163, "reward_total_composite_std": 0.0003373012295924127}
|
| 1088 |
+
{"timestamp_utc": "2026-04-11T21:41:07Z", "mode": "train", "global_step": 1067, "epoch": 0.041203274637009575, "loss": 0.0034, "grad_norm": 10.166537284851074, "learning_rate": 6.76969696969697e-06, "num_tokens": 2301589.0, "completions/mean_length": 27.125, "completions/min_length": 26.0, "completions/max_length": 29.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 27.125, "completions/min_terminated_length": 26.0, "completions/max_terminated_length": 29.0, "rewards/meter/mean": 0.9321423768997192, "rewards/meter/std": 0.007329464890062809, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9321423768997192, "rewards/total_composite/std": 0.007329464890062809, "reward": 0.9321423768997192, "reward_std": 0.007329456973820925, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.026033449918031693, "sampling/sampling_logp_difference/max": 1.5315618515014648, "sampling/importance_sampling_ratio/min": 0.21619774401187897, "sampling/importance_sampling_ratio/mean": 1.0001667737960815, "sampling/importance_sampling_ratio/max": 1.4033339023590088, "entropy": 0.09683473920449615, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.00893997447565198, "clip_ratio/high_max": 0.00893997447565198, "clip_ratio/region_mean": 0.00893997447565198, "reward_total_mean": 0.9321423768997192, "reward_meter_mean": 0.9321423768997192, "reward_meter_std": 0.007329464890062809, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9321423768997192, "reward_total_composite_std": 0.007329464890062809}
|
| 1089 |
+
{"timestamp_utc": "2026-04-11T21:41:12Z", "mode": "train", "global_step": 1068, "epoch": 0.041241890639481, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.7666666666666665e-06, "num_tokens": 2303405.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00019105577666778117, "sampling/sampling_logp_difference/max": 0.0064575872384011745, "sampling/importance_sampling_ratio/min": 0.9977074861526489, "sampling/importance_sampling_ratio/mean": 1.000160813331604, "sampling/importance_sampling_ratio/max": 1.0064785480499268, "entropy": 0.0018052591913146898, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0}
|
| 1090 |
+
{"timestamp_utc": "2026-04-11T21:41:17Z", "mode": "train", "global_step": 1069, "epoch": 0.04128050664195242, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.763636363636365e-06, "num_tokens": 2305133.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0005417983047664165, "sampling/sampling_logp_difference/max": 0.036248765885829926, "sampling/importance_sampling_ratio/min": 0.9796718955039978, "sampling/importance_sampling_ratio/mean": 1.0003167390823364, "sampling/importance_sampling_ratio/max": 1.0369137525558472, "entropy": 0.005523053434444591, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0}
|
| 1091 |
+
{"timestamp_utc": "2026-04-11T21:41:21Z", "mode": "train", "global_step": 1070, "epoch": 0.04131912264442385, "loss": 0.0381, "grad_norm": 26.325374603271484, "learning_rate": 6.760606060606061e-06, "num_tokens": 2306762.0, "completions/mean_length": 52.625, "completions/min_length": 50.0, "completions/max_length": 54.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 52.625, "completions/min_terminated_length": 50.0, "completions/max_terminated_length": 54.0, "rewards/meter/mean": 0.11274315416812897, "rewards/meter/std": 0.3188244700431824, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.11274315416812897, "rewards/total_composite/std": 0.3188244700431824, "reward": 0.11274315416812897, "reward_std": 0.31882444024086, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.02075301483273506, "sampling/sampling_logp_difference/max": 1.0765538215637207, "sampling/importance_sampling_ratio/min": 0.34076783061027527, "sampling/importance_sampling_ratio/mean": 1.0003687143325806, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.05561392899835482, "clip_ratio/low_mean": 0.016467438312247396, "clip_ratio/low_min": 0.016467438312247396, "clip_ratio/high_mean": 0.007499999832361937, "clip_ratio/high_max": 0.007499999832361937, "clip_ratio/region_mean": 0.023967438144609332, "reward_total_mean": 0.11274315416812897, "reward_meter_mean": 0.11274315416812897, "reward_meter_std": 0.3188244700431824, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.11274315416812897, "reward_total_composite_std": 0.3188244700431824}
|
| 1092 |
+
{"timestamp_utc": "2026-04-11T21:41:26Z", "mode": "train", "global_step": 1071, "epoch": 0.04135773864689527, "loss": -0.0217, "grad_norm": 5.607251167297363, "learning_rate": 6.757575757575758e-06, "num_tokens": 2308113.0, "completions/mean_length": 25.875, "completions/min_length": 25.0, "completions/max_length": 26.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 25.875, "completions/min_terminated_length": 25.0, "completions/max_terminated_length": 26.0, "rewards/meter/mean": 0.9891167879104614, "rewards/meter/std": 0.008078054524958134, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9891167879104614, "rewards/total_composite/std": 0.008078054524958134, "reward": 0.9891167879104614, "reward_std": 0.008078045211732388, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.009091660380363464, "sampling/sampling_logp_difference/max": 0.6194605827331543, "sampling/importance_sampling_ratio/min": 0.5382347106933594, "sampling/importance_sampling_ratio/mean": 0.9988547563552856, "sampling/importance_sampling_ratio/max": 1.0410070419311523, "entropy": 0.02826369390822947, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9891167879104614, "reward_meter_mean": 0.9891167879104614, "reward_meter_std": 0.008078054524958134, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9891167879104614, "reward_total_composite_std": 0.008078054524958134}
|
| 1093 |
+
{"timestamp_utc": "2026-04-11T21:41:31Z", "mode": "train", "global_step": 1072, "epoch": 0.041396354649366696, "loss": 0.0001, "grad_norm": 2.4955992698669434, "learning_rate": 6.754545454545455e-06, "num_tokens": 2309830.0, "completions/mean_length": 61.625, "completions/min_length": 61.0, "completions/max_length": 62.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.625, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 62.0, "rewards/meter/mean": 0.997157096862793, "rewards/meter/std": 0.0005752869765274227, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.997157096862793, "rewards/total_composite/std": 0.0005752869765274227, "reward": 0.997157096862793, "reward_std": 0.000575280690100044, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.012418783269822598, "sampling/sampling_logp_difference/max": 0.5172317028045654, "sampling/importance_sampling_ratio/min": 0.5961686372756958, "sampling/importance_sampling_ratio/mean": 1.0032328367233276, "sampling/importance_sampling_ratio/max": 1.3372715711593628, "entropy": 0.10848722152877599, "clip_ratio/low_mean": 0.010146747343242168, "clip_ratio/low_min": 0.010146747343242168, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.010146747343242168, "reward_total_mean": 0.997157096862793, "reward_meter_mean": 0.997157096862793, "reward_meter_std": 0.0005752869765274227, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.997157096862793, "reward_total_composite_std": 0.0005752869765274227}
|
| 1094 |
+
{"timestamp_utc": "2026-04-11T21:41:36Z", "mode": "train", "global_step": 1073, "epoch": 0.04143497065183812, "loss": 0.0115, "grad_norm": 0.9153868556022644, "learning_rate": 6.751515151515152e-06, "num_tokens": 2311287.0, "completions/mean_length": 33.125, "completions/min_length": 33.0, "completions/max_length": 34.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 33.125, "completions/min_terminated_length": 33.0, "completions/max_terminated_length": 34.0, "rewards/meter/mean": 0.9944090247154236, "rewards/meter/std": 0.0004998616641387343, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9944090247154236, "rewards/total_composite/std": 0.0004998616641387343, "reward": 0.9944090247154236, "reward_std": 0.0004998616059310734, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.005571494810283184, "sampling/sampling_logp_difference/max": 0.6732907295227051, "sampling/importance_sampling_ratio/min": 0.9340354800224304, "sampling/importance_sampling_ratio/mean": 1.0061020851135254, "sampling/importance_sampling_ratio/max": 1.9606788158416748, "entropy": 0.025314107653684914, "clip_ratio/low_mean": 0.0036764706019312143, "clip_ratio/low_min": 0.0036764706019312143, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0036764706019312143, "reward_total_mean": 0.9944090247154236, "reward_meter_mean": 0.9944090247154236, "reward_meter_std": 0.0004998616641387343, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9944090247154236, "reward_total_composite_std": 0.0004998616641387343}
|
| 1095 |
+
{"timestamp_utc": "2026-04-11T21:41:41Z", "mode": "train", "global_step": 1074, "epoch": 0.041473586654309544, "loss": 0.0417, "grad_norm": 3.6375036239624023, "learning_rate": 6.748484848484848e-06, "num_tokens": 2313096.0, "completions/mean_length": 54.125, "completions/min_length": 51.0, "completions/max_length": 56.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 54.125, "completions/min_terminated_length": 51.0, "completions/max_terminated_length": 56.0, "rewards/meter/mean": 0.9875502586364746, "rewards/meter/std": 0.003662221832200885, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9875502586364746, "rewards/total_composite/std": 0.003662221832200885, "reward": 0.9875502586364746, "reward_std": 0.003662231843918562, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.004612638149410486, "sampling/sampling_logp_difference/max": 0.22617030143737793, "sampling/importance_sampling_ratio/min": 0.7975822687149048, "sampling/importance_sampling_ratio/mean": 1.003567099571228, "sampling/importance_sampling_ratio/max": 1.2383108139038086, "entropy": 0.03697874676436186, "clip_ratio/low_mean": 0.004464285913854837, "clip_ratio/low_min": 0.004464285913854837, "clip_ratio/high_mean": 0.0024509804788976908, "clip_ratio/high_max": 0.0024509804788976908, "clip_ratio/region_mean": 0.006915266392752528, "reward_total_mean": 0.9875502586364746, "reward_meter_mean": 0.9875502586364746, "reward_meter_std": 0.003662221832200885, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9875502586364746, "reward_total_composite_std": 0.003662221832200885}
|
| 1096 |
+
{"timestamp_utc": "2026-04-11T21:41:46Z", "mode": "train", "global_step": 1075, "epoch": 0.04151220265678097, "loss": 0.0288, "grad_norm": 6.4112677574157715, "learning_rate": 6.7454545454545465e-06, "num_tokens": 2314778.0, "completions/mean_length": 47.25, "completions/min_length": 41.0, "completions/max_length": 58.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 47.25, "completions/min_terminated_length": 41.0, "completions/max_terminated_length": 58.0, "rewards/meter/mean": 0.10002761334180832, "rewards/meter/std": 0.26244989037513733, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.10002761334180832, "rewards/total_composite/std": 0.26244989037513733, "reward": 0.10002761334180832, "reward_std": 0.26244989037513733, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.03375272452831268, "sampling/sampling_logp_difference/max": 1.2938915491104126, "sampling/importance_sampling_ratio/min": 0.2742016315460205, "sampling/importance_sampling_ratio/mean": 1.0019086599349976, "sampling/importance_sampling_ratio/max": 1.4211242198944092, "entropy": 0.16682779975235462, "clip_ratio/low_mean": 0.03481511096470058, "clip_ratio/low_min": 0.03481511096470058, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.03481511096470058, "reward_total_mean": 0.10002761334180832, "reward_meter_mean": 0.10002761334180832, "reward_meter_std": 0.26244989037513733, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.10002761334180832, "reward_total_composite_std": 0.26244989037513733}
|
| 1097 |
+
{"timestamp_utc": "2026-04-11T21:41:51Z", "mode": "train", "global_step": 1076, "epoch": 0.04155081865925239, "loss": -0.0001, "grad_norm": 4.325499057769775, "learning_rate": 6.742424242424243e-06, "num_tokens": 2316531.0, "completions/mean_length": 53.125, "completions/min_length": 52.0, "completions/max_length": 55.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 53.125, "completions/min_terminated_length": 52.0, "completions/max_terminated_length": 55.0, "rewards/meter/mean": 0.9805549383163452, "rewards/meter/std": 0.00861036404967308, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9805549383163452, "rewards/total_composite/std": 0.00861036404967308, "reward": 0.9805549383163452, "reward_std": 0.008610363118350506, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.011005771346390247, "sampling/sampling_logp_difference/max": 0.7941849231719971, "sampling/importance_sampling_ratio/min": 0.45194947719573975, "sampling/importance_sampling_ratio/mean": 1.0017882585525513, "sampling/importance_sampling_ratio/max": 1.5676127672195435, "entropy": 0.049260836094617844, "clip_ratio/low_mean": 0.009479318046942353, "clip_ratio/low_min": 0.009479318046942353, "clip_ratio/high_mean": 0.004631217801943421, "clip_ratio/high_max": 0.004631217801943421, "clip_ratio/region_mean": 0.014110535848885775, "reward_total_mean": 0.9805549383163452, "reward_meter_mean": 0.9805549383163452, "reward_meter_std": 0.00861036404967308, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9805549383163452, "reward_total_composite_std": 0.00861036404967308}
|
| 1098 |
+
{"timestamp_utc": "2026-04-11T21:41:56Z", "mode": "train", "global_step": 1077, "epoch": 0.041589434661723816, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.73939393939394e-06, "num_tokens": 2318435.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.000207486460567452, "sampling/sampling_logp_difference/max": 0.02866499498486519, "sampling/importance_sampling_ratio/min": 0.9717419147491455, "sampling/importance_sampling_ratio/mean": 1.000050663948059, "sampling/importance_sampling_ratio/max": 1.0065354108810425, "entropy": 0.0024144309863913804, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0}
|
| 1099 |
+
{"timestamp_utc": "2026-04-11T21:42:01Z", "mode": "train", "global_step": 1078, "epoch": 0.04162805066419524, "loss": 0.0161, "grad_norm": 4.65764856338501, "learning_rate": 6.7363636363636365e-06, "num_tokens": 2320094.0, "completions/mean_length": 61.375, "completions/min_length": 59.0, "completions/max_length": 64.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.375, "completions/min_terminated_length": 59.0, "completions/max_terminated_length": 64.0, "rewards/meter/mean": 0.9684683084487915, "rewards/meter/std": 0.08065449446439743, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9684683084487915, "rewards/total_composite/std": 0.08065449446439743, "reward": 0.9684683084487915, "reward_std": 0.08065447211265564, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.02014295756816864, "sampling/sampling_logp_difference/max": 1.577573299407959, "sampling/importance_sampling_ratio/min": 0.20647554099559784, "sampling/importance_sampling_ratio/mean": 1.0008001327514648, "sampling/importance_sampling_ratio/max": 1.5210373401641846, "entropy": 0.13324717595241964, "clip_ratio/low_mean": 0.0078125, "clip_ratio/low_min": 0.0078125, "clip_ratio/high_mean": 0.010216211201623082, "clip_ratio/high_max": 0.010216211201623082, "clip_ratio/region_mean": 0.018028711201623082, "reward_total_mean": 0.9684683084487915, "reward_meter_mean": 0.9684683084487915, "reward_meter_std": 0.08065449446439743, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9684683084487915, "reward_total_composite_std": 0.08065449446439743}
|
| 1100 |
+
{"timestamp_utc": "2026-04-11T21:42:06Z", "mode": "train", "global_step": 1079, "epoch": 0.041666666666666664, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.733333333333334e-06, "num_tokens": 2321786.0, "completions/mean_length": 49.5, "completions/min_length": 46.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 49.5, "completions/min_terminated_length": 46.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.091847725212574, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.091847725212574, "rewards/total_composite/std": 0.0, "reward": 0.091847725212574, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.01916210725903511, "sampling/sampling_logp_difference/max": 1.860931396484375, "sampling/importance_sampling_ratio/min": 0.15552771091461182, "sampling/importance_sampling_ratio/mean": 0.9989585280418396, "sampling/importance_sampling_ratio/max": 1.5695313215255737, "entropy": 0.054232243448495865, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.091847725212574, "reward_meter_mean": 0.091847725212574, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.091847725212574, "reward_total_composite_std": 0.0}
|
| 1101 |
+
{"timestamp_utc": "2026-04-11T21:42:11Z", "mode": "train", "global_step": 1080, "epoch": 0.04170528266913809, "loss": 0.0236, "grad_norm": 1.9305036067962646, "learning_rate": 6.73030303030303e-06, "num_tokens": 2324566.0, "completions/mean_length": 132.5, "completions/min_length": 129.0, "completions/max_length": 137.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 132.5, "completions/min_terminated_length": 129.0, "completions/max_terminated_length": 137.0, "rewards/meter/mean": 0.9948593974113464, "rewards/meter/std": 0.002004272071644664, "rewards/count_adherence/mean": 0.6666666865348816, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.663239598274231, "rewards/total_composite/std": 0.0013361814199015498, "reward": 0.663239598274231, "reward_std": 0.0013361814199015498, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.004363351967185736, "sampling/sampling_logp_difference/max": 0.48967456817626953, "sampling/importance_sampling_ratio/min": 0.6128258109092712, "sampling/importance_sampling_ratio/mean": 1.0016835927963257, "sampling/importance_sampling_ratio/max": 1.3798326253890991, "entropy": 0.029038164531812072, "clip_ratio/low_mean": 0.0027643622015602887, "clip_ratio/low_min": 0.0027643622015602887, "clip_ratio/high_mean": 0.0019379844889044762, "clip_ratio/high_max": 0.0019379844889044762, "clip_ratio/region_mean": 0.004702346690464765, "reward_total_mean": 0.663239598274231, "reward_meter_mean": 0.9948593974113464, "reward_meter_std": 0.002004272071644664, "reward_count_adherence_mean": 0.6666666865348816, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.663239598274231, "reward_total_composite_std": 0.0013361814199015498}
|
| 1102 |
+
{"timestamp_utc": "2026-04-11T21:42:17Z", "mode": "train", "global_step": 1081, "epoch": 0.04174389867160951, "loss": 0.0025, "grad_norm": 0.31140798330307007, "learning_rate": 6.7272727272727275e-06, "num_tokens": 2327382.0, "completions/mean_length": 175.0, "completions/min_length": 161.0, "completions/max_length": 178.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 175.0, "completions/min_terminated_length": 161.0, "completions/max_terminated_length": 178.0, "rewards/meter/mean": 0.9966549277305603, "rewards/meter/std": 7.61248666094616e-05, "rewards/count_adherence/mean": 0.75, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.747491180896759, "rewards/total_composite/std": 5.710634286515415e-05, "reward": 0.747491180896759, "reward_std": 5.710850018658675e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0060366359539330006, "sampling/sampling_logp_difference/max": 3.635336399078369, "sampling/importance_sampling_ratio/min": 0.0263750609010458, "sampling/importance_sampling_ratio/mean": 0.9993287920951843, "sampling/importance_sampling_ratio/max": 1.509435772895813, "entropy": 0.015698739560320973, "clip_ratio/low_mean": 0.001416441984474659, "clip_ratio/low_min": 0.001416441984474659, "clip_ratio/high_mean": 0.002259009750559926, "clip_ratio/high_max": 0.002259009750559926, "clip_ratio/region_mean": 0.003675451735034585, "reward_total_mean": 0.747491180896759, "reward_meter_mean": 0.9966549277305603, "reward_meter_std": 7.61248666094616e-05, "reward_count_adherence_mean": 0.75, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.747491180896759, "reward_total_composite_std": 5.710634286515415e-05}
|
| 1103 |
+
{"timestamp_utc": "2026-04-11T21:42:23Z", "mode": "train", "global_step": 1082, "epoch": 0.04178251467408094, "loss": 0.0114, "grad_norm": 4.485145092010498, "learning_rate": 6.724242424242424e-06, "num_tokens": 2329495.0, "completions/mean_length": 108.125, "completions/min_length": 106.0, "completions/max_length": 111.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 108.125, "completions/min_terminated_length": 106.0, "completions/max_terminated_length": 111.0, "rewards/meter/mean": 0.9212300777435303, "rewards/meter/std": 0.21582916378974915, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9212300777435303, "rewards/total_composite/std": 0.21582916378974915, "reward": 0.9212300777435303, "reward_std": 0.21582916378974915, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.013148986734449863, "sampling/sampling_logp_difference/max": 0.9551130533218384, "sampling/importance_sampling_ratio/min": 0.38476866483688354, "sampling/importance_sampling_ratio/mean": 1.0012389421463013, "sampling/importance_sampling_ratio/max": 1.5335668325424194, "entropy": 0.09740556543692946, "clip_ratio/low_mean": 0.0033783784601837397, "clip_ratio/low_min": 0.0033783784601837397, "clip_ratio/high_mean": 0.009324767161160707, "clip_ratio/high_max": 0.009324767161160707, "clip_ratio/region_mean": 0.012703145621344447, "reward_total_mean": 0.9212300777435303, "reward_meter_mean": 0.9212300777435303, "reward_meter_std": 0.21582916378974915, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9212300777435303, "reward_total_composite_std": 0.21582916378974915}
|
| 1104 |
+
{"timestamp_utc": "2026-04-11T21:42:28Z", "mode": "train", "global_step": 1083, "epoch": 0.04182113067655236, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.721212121212122e-06, "num_tokens": 2331783.0, "completions/mean_length": 129.0, "completions/min_length": 129.0, "completions/max_length": 129.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 129.0, "completions/min_terminated_length": 129.0, "completions/max_terminated_length": 129.0, "rewards/meter/mean": 0.9951313734054565, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.6666666865348816, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.6634209156036377, "rewards/total_composite/std": 0.0, "reward": 0.6634209156036377, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0002578256244305521, "sampling/sampling_logp_difference/max": 0.045966774225234985, "sampling/importance_sampling_ratio/min": 0.9550737738609314, "sampling/importance_sampling_ratio/mean": 1.0000473260879517, "sampling/importance_sampling_ratio/max": 1.0351933240890503, "entropy": 0.002667208347702399, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.6634209156036377, "reward_meter_mean": 0.9951313734054565, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.6666666865348816, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.6634209156036377, "reward_total_composite_std": 0.0}
|
| 1105 |
+
{"timestamp_utc": "2026-04-11T21:42:33Z", "mode": "train", "global_step": 1084, "epoch": 0.041859746679023785, "loss": 0.0462, "grad_norm": 17.376197814941406, "learning_rate": 6.718181818181819e-06, "num_tokens": 2333323.0, "completions/mean_length": 28.5, "completions/min_length": 27.0, "completions/max_length": 31.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 28.5, "completions/min_terminated_length": 27.0, "completions/max_terminated_length": 31.0, "rewards/meter/mean": 0.6803448796272278, "rewards/meter/std": 0.4452557861804962, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.6803448796272278, "rewards/total_composite/std": 0.4452557861804962, "reward": 0.6803448796272278, "reward_std": 0.4452557861804962, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.07098246365785599, "sampling/sampling_logp_difference/max": 2.6754026412963867, "sampling/importance_sampling_ratio/min": 0.06887909024953842, "sampling/importance_sampling_ratio/mean": 0.9909398555755615, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.25689972564578056, "clip_ratio/low_mean": 0.012652947567403316, "clip_ratio/low_min": 0.012652947567403316, "clip_ratio/high_mean": 0.04814814869314432, "clip_ratio/high_max": 0.04814814869314432, "clip_ratio/region_mean": 0.06080109626054764, "reward_total_mean": 0.6803448796272278, "reward_meter_mean": 0.6803448796272278, "reward_meter_std": 0.4452557861804962, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.6803448796272278, "reward_total_composite_std": 0.4452557861804962}
|
| 1106 |
+
{"timestamp_utc": "2026-04-11T21:42:40Z", "mode": "train", "global_step": 1085, "epoch": 0.04189836268149521, "loss": -0.0223, "grad_norm": 7.492648124694824, "learning_rate": 6.715151515151516e-06, "num_tokens": 2336298.0, "completions/mean_length": 189.875, "completions/min_length": 180.0, "completions/max_length": 201.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 189.875, "completions/min_terminated_length": 180.0, "completions/max_terminated_length": 201.0, "rewards/meter/mean": 0.46110403537750244, "rewards/meter/std": 0.25172099471092224, "rewards/count_adherence/mean": 0.7083333730697632, "rewards/count_adherence/std": 0.07715165615081787, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.3282319903373718, "rewards/total_composite/std": 0.1757289618253708, "reward": 0.3282319903373718, "reward_std": 0.1757289618253708, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.010337450541555882, "sampling/sampling_logp_difference/max": 1.9266459941864014, "sampling/importance_sampling_ratio/min": 0.14563584327697754, "sampling/importance_sampling_ratio/mean": 1.0021684169769287, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.034014943055808544, "clip_ratio/low_mean": 0.002667596214450896, "clip_ratio/low_min": 0.002667596214450896, "clip_ratio/high_mean": 0.0026115492219105363, "clip_ratio/high_max": 0.0026115492219105363, "clip_ratio/region_mean": 0.005279145436361432, "reward_total_mean": 0.3282319903373718, "reward_meter_mean": 0.46110403537750244, "reward_meter_std": 0.25172099471092224, "reward_count_adherence_mean": 0.7083333730697632, "reward_count_adherence_std": 0.07715165615081787, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.3282319903373718, "reward_total_composite_std": 0.1757289618253708}
|
| 1107 |
+
{"timestamp_utc": "2026-04-11T21:42:46Z", "mode": "train", "global_step": 1086, "epoch": 0.04193697868396663, "loss": 0.0308, "grad_norm": 3.549293279647827, "learning_rate": 6.712121212121213e-06, "num_tokens": 2339268.0, "completions/mean_length": 208.25, "completions/min_length": 197.0, "completions/max_length": 213.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 208.25, "completions/min_terminated_length": 197.0, "completions/max_terminated_length": 213.0, "rewards/meter/mean": 0.9969711303710938, "rewards/meter/std": 8.204347250284627e-05, "rewards/count_adherence/mean": 0.6500000357627869, "rewards/count_adherence/std": 0.09258200973272324, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.6480287909507751, "rewards/total_composite/std": 0.0922786295413971, "reward": 0.6480287909507751, "reward_std": 0.09227863699197769, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.010134638287127018, "sampling/sampling_logp_difference/max": 8.085393905639648, "sampling/importance_sampling_ratio/min": 0.00030800519743934274, "sampling/importance_sampling_ratio/mean": 1.0001766681671143, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.008055432001128793, "clip_ratio/low_mean": 0.002358543104492128, "clip_ratio/low_min": 0.002358543104492128, "clip_ratio/high_mean": 0.0006345177534967661, "clip_ratio/high_max": 0.0006345177534967661, "clip_ratio/region_mean": 0.002993060857988894, "reward_total_mean": 0.6480287909507751, "reward_meter_mean": 0.9969711303710938, "reward_meter_std": 8.204347250284627e-05, "reward_count_adherence_mean": 0.6500000357627869, "reward_count_adherence_std": 0.09258200973272324, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.6480287909507751, "reward_total_composite_std": 0.0922786295413971}
|
| 1108 |
+
{"timestamp_utc": "2026-04-11T21:42:51Z", "mode": "train", "global_step": 1087, "epoch": 0.04197559468643806, "loss": -0.0111, "grad_norm": 2.269176959991455, "learning_rate": 6.709090909090909e-06, "num_tokens": 2341095.0, "completions/mean_length": 80.375, "completions/min_length": 76.0, "completions/max_length": 81.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 80.375, "completions/min_terminated_length": 76.0, "completions/max_terminated_length": 81.0, "rewards/meter/mean": 0.9966411590576172, "rewards/meter/std": 0.00015710237494204193, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9966411590576172, "rewards/total_composite/std": 0.00015710237494204193, "reward": 0.9966411590576172, "reward_std": 0.00015711141168139875, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0034698506351560354, "sampling/sampling_logp_difference/max": 0.793515682220459, "sampling/importance_sampling_ratio/min": 0.45225203037261963, "sampling/importance_sampling_ratio/mean": 0.9995546340942383, "sampling/importance_sampling_ratio/max": 1.1145241260528564, "entropy": 0.014752325252629817, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0015432098880410194, "clip_ratio/high_max": 0.0015432098880410194, "clip_ratio/region_mean": 0.0015432098880410194, "reward_total_mean": 0.9966411590576172, "reward_meter_mean": 0.9966411590576172, "reward_meter_std": 0.00015710237494204193, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9966411590576172, "reward_total_composite_std": 0.00015710237494204193}
|
| 1109 |
+
{"timestamp_utc": "2026-04-11T21:42:57Z", "mode": "train", "global_step": 1088, "epoch": 0.04201421068890948, "loss": 0.0123, "grad_norm": 12.239304542541504, "learning_rate": 6.706060606060607e-06, "num_tokens": 2342713.0, "completions/mean_length": 49.25, "completions/min_length": 49.0, "completions/max_length": 51.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 49.25, "completions/min_terminated_length": 49.0, "completions/max_terminated_length": 51.0, "rewards/meter/mean": 0.986293613910675, "rewards/meter/std": 0.0026903701946139336, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.986293613910675, "rewards/total_composite/std": 0.0026903701946139336, "reward": 0.986293613910675, "reward_std": 0.0026903690304607153, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.016510505229234695, "sampling/sampling_logp_difference/max": 1.4381656646728516, "sampling/importance_sampling_ratio/min": 0.23736277222633362, "sampling/importance_sampling_ratio/mean": 1.0027153491973877, "sampling/importance_sampling_ratio/max": 1.9696929454803467, "entropy": 0.09569864999502897, "clip_ratio/low_mean": 0.007452981313690543, "clip_ratio/low_min": 0.007452981313690543, "clip_ratio/high_mean": 0.007653061067685485, "clip_ratio/high_max": 0.007653061067685485, "clip_ratio/region_mean": 0.015106042381376028, "reward_total_mean": 0.986293613910675, "reward_meter_mean": 0.986293613910675, "reward_meter_std": 0.0026903701946139336, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.986293613910675, "reward_total_composite_std": 0.0026903701946139336}
|
| 1110 |
+
{"timestamp_utc": "2026-04-11T21:43:03Z", "mode": "train", "global_step": 1089, "epoch": 0.042052826691380905, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.703030303030304e-06, "num_tokens": 2345457.0, "completions/mean_length": 152.0, "completions/min_length": 152.0, "completions/max_length": 152.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 152.0, "completions/min_terminated_length": 152.0, "completions/max_terminated_length": 152.0, "rewards/meter/mean": 0.9968674182891846, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.75, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7476505637168884, "rewards/total_composite/std": 0.0, "reward": 0.7476505637168884, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0006623952649533749, "sampling/sampling_logp_difference/max": 0.3456934988498688, "sampling/importance_sampling_ratio/min": 0.7077293992042542, "sampling/importance_sampling_ratio/mean": 0.9997900128364563, "sampling/importance_sampling_ratio/max": 1.08553946018219, "entropy": 0.0034695666399784386, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.7476505637168884, "reward_meter_mean": 0.9968674182891846, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.75, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7476505637168884, "reward_total_composite_std": 0.0}
|
| 1111 |
+
{"timestamp_utc": "2026-04-11T21:43:09Z", "mode": "train", "global_step": 1090, "epoch": 0.04209144269385233, "loss": -0.0003, "grad_norm": 0.4215506911277771, "learning_rate": 6.700000000000001e-06, "num_tokens": 2348309.0, "completions/mean_length": 182.5, "completions/min_length": 181.0, "completions/max_length": 183.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 182.5, "completions/min_terminated_length": 181.0, "completions/max_terminated_length": 183.0, "rewards/meter/mean": 0.9972051978111267, "rewards/meter/std": 5.313828296493739e-05, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9972051978111267, "rewards/total_composite/std": 5.313828296493739e-05, "reward": 0.9972051978111267, "reward_std": 5.313552901498042e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.004990379326045513, "sampling/sampling_logp_difference/max": 0.824894905090332, "sampling/importance_sampling_ratio/min": 0.4382810592651367, "sampling/importance_sampling_ratio/mean": 1.0013642311096191, "sampling/importance_sampling_ratio/max": 1.4445006847381592, "entropy": 0.04286058805882931, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.004788926860783249, "clip_ratio/high_max": 0.004788926860783249, "clip_ratio/region_mean": 0.004788926860783249, "reward_total_mean": 0.9972051978111267, "reward_meter_mean": 0.9972051978111267, "reward_meter_std": 5.313828296493739e-05, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9972051978111267, "reward_total_composite_std": 5.313828296493739e-05}
|
| 1112 |
+
{"timestamp_utc": "2026-04-11T21:43:14Z", "mode": "train", "global_step": 1091, "epoch": 0.042130058696323754, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.6969696969696975e-06, "num_tokens": 2349895.0, "completions/mean_length": 47.25, "completions/min_length": 46.0, "completions/max_length": 56.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 47.25, "completions/min_terminated_length": 46.0, "completions/max_terminated_length": 56.0, "rewards/meter/mean": 0.091847725212574, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.091847725212574, "rewards/total_composite/std": 0.0, "reward": 0.091847725212574, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.007502844091504812, "sampling/sampling_logp_difference/max": 0.6710173487663269, "sampling/importance_sampling_ratio/min": 0.511188268661499, "sampling/importance_sampling_ratio/mean": 1.001801609992981, "sampling/importance_sampling_ratio/max": 1.5886342525482178, "entropy": 0.016990114585496485, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.091847725212574, "reward_meter_mean": 0.091847725212574, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.091847725212574, "reward_total_composite_std": 0.0}
|
| 1113 |
+
{"timestamp_utc": "2026-04-11T21:43:18Z", "mode": "train", "global_step": 1092, "epoch": 0.04216867469879518, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.693939393939395e-06, "num_tokens": 2351375.0, "completions/mean_length": 32.0, "completions/min_length": 32.0, "completions/max_length": 32.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 32.0, "completions/min_terminated_length": 32.0, "completions/max_terminated_length": 32.0, "rewards/meter/mean": 0.9955702424049377, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9955702424049377, "rewards/total_composite/std": 0.0, "reward": 0.9955702424049377, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.001951782964169979, "sampling/sampling_logp_difference/max": 0.17160826921463013, "sampling/importance_sampling_ratio/min": 0.8423091173171997, "sampling/importance_sampling_ratio/mean": 0.9996309876441956, "sampling/importance_sampling_ratio/max": 1.054145336151123, "entropy": 0.010097296675667167, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9955702424049377, "reward_meter_mean": 0.9955702424049377, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9955702424049377, "reward_total_composite_std": 0.0}
|
| 1114 |
+
{"timestamp_utc": "2026-04-11T21:43:23Z", "mode": "train", "global_step": 1093, "epoch": 0.0422072907012666, "loss": 0.0095, "grad_norm": 1.6124061346054077, "learning_rate": 6.690909090909091e-06, "num_tokens": 2353053.0, "completions/mean_length": 64.75, "completions/min_length": 63.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 64.75, "completions/min_terminated_length": 63.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9950487613677979, "rewards/meter/std": 0.0002570115029811859, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9950487613677979, "rewards/total_composite/std": 0.0002570115029811859, "reward": 0.9950487613677979, "reward_std": 0.00025701147387735546, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.005157488863915205, "sampling/sampling_logp_difference/max": 1.5352113246917725, "sampling/importance_sampling_ratio/min": 0.21541017293930054, "sampling/importance_sampling_ratio/mean": 0.9983689785003662, "sampling/importance_sampling_ratio/max": 1.0859289169311523, "entropy": 0.013940075354184955, "clip_ratio/low_mean": 0.003846153849735856, "clip_ratio/low_min": 0.003846153849735856, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.003846153849735856, "reward_total_mean": 0.9950487613677979, "reward_meter_mean": 0.9950487613677979, "reward_meter_std": 0.0002570115029811859, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9950487613677979, "reward_total_composite_std": 0.0002570115029811859}
|
| 1115 |
+
{"timestamp_utc": "2026-04-11T21:43:28Z", "mode": "train", "global_step": 1094, "epoch": 0.042245906703738026, "loss": 0.0094, "grad_norm": 3.3892717361450195, "learning_rate": 6.687878787878788e-06, "num_tokens": 2354894.0, "completions/mean_length": 65.125, "completions/min_length": 64.0, "completions/max_length": 66.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.125, "completions/min_terminated_length": 64.0, "completions/max_terminated_length": 66.0, "rewards/meter/mean": 0.9951438903808594, "rewards/meter/std": 0.0001731093943817541, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9951438903808594, "rewards/total_composite/std": 0.0001731093943817541, "reward": 0.9951438903808594, "reward_std": 0.00017311105330009013, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.011436189524829388, "sampling/sampling_logp_difference/max": 1.639451026916504, "sampling/importance_sampling_ratio/min": 0.19408656656742096, "sampling/importance_sampling_ratio/mean": 0.996876060962677, "sampling/importance_sampling_ratio/max": 1.232495665550232, "entropy": 0.04774620302487165, "clip_ratio/low_mean": 0.0038470644503831863, "clip_ratio/low_min": 0.0038470644503831863, "clip_ratio/high_mean": 0.00390625, "clip_ratio/high_max": 0.00390625, "clip_ratio/region_mean": 0.007753314450383186, "reward_total_mean": 0.9951438903808594, "reward_meter_mean": 0.9951438903808594, "reward_meter_std": 0.0001731093943817541, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9951438903808594, "reward_total_composite_std": 0.0001731093943817541}
|
| 1116 |
+
{"timestamp_utc": "2026-04-11T21:43:33Z", "mode": "train", "global_step": 1095, "epoch": 0.04228452270620945, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.684848484848485e-06, "num_tokens": 2357158.0, "completions/mean_length": 107.0, "completions/min_length": 107.0, "completions/max_length": 107.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 107.0, "completions/min_terminated_length": 107.0, "completions/max_terminated_length": 107.0, "rewards/meter/mean": 0.9966511130332947, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9966511130332947, "rewards/total_composite/std": 0.0, "reward": 0.9966511130332947, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0007139050285331905, "sampling/sampling_logp_difference/max": 0.2203187644481659, "sampling/importance_sampling_ratio/min": 0.8022630214691162, "sampling/importance_sampling_ratio/mean": 0.999952495098114, "sampling/importance_sampling_ratio/max": 1.083788275718689, "entropy": 0.003691258083563298, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9966511130332947, "reward_meter_mean": 0.9966511130332947, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9966511130332947, "reward_total_composite_std": 0.0}
|
| 1117 |
+
{"timestamp_utc": "2026-04-11T21:43:38Z", "mode": "train", "global_step": 1096, "epoch": 0.042323138708680874, "loss": 0.0528, "grad_norm": 17.096412658691406, "learning_rate": 6.681818181818183e-06, "num_tokens": 2359061.0, "completions/mean_length": 26.875, "completions/min_length": 20.0, "completions/max_length": 29.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 26.875, "completions/min_terminated_length": 20.0, "completions/max_terminated_length": 29.0, "rewards/meter/mean": 0.11241252720355988, "rewards/meter/std": 0.3034849762916565, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.11241252720355988, "rewards/total_composite/std": 0.3034849762916565, "reward": 0.11241252720355988, "reward_std": 0.3034849464893341, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.10703544318675995, "sampling/sampling_logp_difference/max": 4.594521999359131, "sampling/importance_sampling_ratio/min": 0.010107050649821758, "sampling/importance_sampling_ratio/mean": 0.979833722114563, "sampling/importance_sampling_ratio/max": 1.9887901544570923, "entropy": 0.21645524073392153, "clip_ratio/low_mean": 0.03616698645055294, "clip_ratio/low_min": 0.03616698645055294, "clip_ratio/high_mean": 0.0223214291036129, "clip_ratio/high_max": 0.0223214291036129, "clip_ratio/region_mean": 0.05848841555416584, "reward_total_mean": 0.11241252720355988, "reward_meter_mean": 0.11241252720355988, "reward_meter_std": 0.3034849762916565, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.11241252720355988, "reward_total_composite_std": 0.3034849762916565}
|
| 1118 |
+
{"timestamp_utc": "2026-04-11T21:43:44Z", "mode": "train", "global_step": 1097, "epoch": 0.0423617547111523, "loss": 0.0354, "grad_norm": 0.9273874163627625, "learning_rate": 6.678787878787879e-06, "num_tokens": 2361662.0, "completions/mean_length": 147.125, "completions/min_length": 145.0, "completions/max_length": 161.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 147.125, "completions/min_terminated_length": 145.0, "completions/max_terminated_length": 161.0, "rewards/meter/mean": 0.9964081048965454, "rewards/meter/std": 0.00033037204411812127, "rewards/count_adherence/mean": 0.96875, "rewards/count_adherence/std": 0.0883883461356163, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9652670621871948, "rewards/total_composite/std": 0.08803740888834, "reward": 0.9652670621871948, "reward_std": 0.08803740888834, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0055220527574419975, "sampling/sampling_logp_difference/max": 1.9974420070648193, "sampling/importance_sampling_ratio/min": 0.4937727451324463, "sampling/importance_sampling_ratio/mean": 1.0020689964294434, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.016805213410407305, "clip_ratio/low_mean": 0.000776397529989481, "clip_ratio/low_min": 0.000776397529989481, "clip_ratio/high_mean": 0.0025862068869173527, "clip_ratio/high_max": 0.0025862068869173527, "clip_ratio/region_mean": 0.0033626044169068336, "reward_total_mean": 0.9652670621871948, "reward_meter_mean": 0.9964081048965454, "reward_meter_std": 0.00033037204411812127, "reward_count_adherence_mean": 0.96875, "reward_count_adherence_std": 0.0883883461356163, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9652670621871948, "reward_total_composite_std": 0.08803740888834}
|
| 1119 |
+
{"timestamp_utc": "2026-04-11T21:43:50Z", "mode": "train", "global_step": 1098, "epoch": 0.04240037071362372, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.6757575757575766e-06, "num_tokens": 2364710.0, "completions/mean_length": 167.0, "completions/min_length": 167.0, "completions/max_length": 167.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 167.0, "completions/min_terminated_length": 167.0, "completions/max_terminated_length": 167.0, "rewards/meter/mean": 0.9968674182891846, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.8333333134651184, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8307228684425354, "rewards/total_composite/std": 0.0, "reward": 0.8307228684425354, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00022596628696192056, "sampling/sampling_logp_difference/max": 0.1071237251162529, "sampling/importance_sampling_ratio/min": 0.8984145522117615, "sampling/importance_sampling_ratio/mean": 1.0000667572021484, "sampling/importance_sampling_ratio/max": 1.0559391975402832, "entropy": 0.0018285177211510018, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.8307228684425354, "reward_meter_mean": 0.9968674182891846, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.8333333134651184, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8307228684425354, "reward_total_composite_std": 0.0}
|
| 1120 |
+
{"timestamp_utc": "2026-04-11T21:43:55Z", "mode": "train", "global_step": 1099, "epoch": 0.042438986716095146, "loss": -0.0331, "grad_norm": 6.627462863922119, "learning_rate": 6.672727272727273e-06, "num_tokens": 2366553.0, "completions/mean_length": 48.375, "completions/min_length": 46.0, "completions/max_length": 57.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 48.375, "completions/min_terminated_length": 46.0, "completions/max_terminated_length": 57.0, "rewards/meter/mean": 0.18795685470104218, "rewards/meter/std": 0.3101930618286133, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.18795685470104218, "rewards/total_composite/std": 0.3101930618286133, "reward": 0.18795685470104218, "reward_std": 0.3101930618286133, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.010162265971302986, "sampling/sampling_logp_difference/max": 0.9485464096069336, "sampling/importance_sampling_ratio/min": 0.3873036205768585, "sampling/importance_sampling_ratio/mean": 1.0062874555587769, "sampling/importance_sampling_ratio/max": 1.6379154920578003, "entropy": 0.030019621714018285, "clip_ratio/low_mean": 0.00657894741743803, "clip_ratio/low_min": 0.00657894741743803, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.00657894741743803, "reward_total_mean": 0.18795685470104218, "reward_meter_mean": 0.18795685470104218, "reward_meter_std": 0.3101930618286133, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.18795685470104218, "reward_total_composite_std": 0.3101930618286133}
|
| 1121 |
+
{"timestamp_utc": "2026-04-11T21:44:01Z", "mode": "train", "global_step": 1100, "epoch": 0.04247760271856657, "loss": -0.0306, "grad_norm": 2.0620083808898926, "learning_rate": 6.66969696969697e-06, "num_tokens": 2369329.0, "completions/mean_length": 156.0, "completions/min_length": 145.0, "completions/max_length": 163.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 156.0, "completions/min_terminated_length": 145.0, "completions/max_terminated_length": 163.0, "rewards/meter/mean": 0.9976032972335815, "rewards/meter/std": 0.0007398881716653705, "rewards/count_adherence/mean": 0.800000011920929, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7980826497077942, "rewards/total_composite/std": 0.0005919100949540734, "reward": 0.7980826497077942, "reward_std": 0.0005919178947806358, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.003971020225435495, "sampling/sampling_logp_difference/max": 0.6630861759185791, "sampling/importance_sampling_ratio/min": 0.5152587294578552, "sampling/importance_sampling_ratio/mean": 1.0009613037109375, "sampling/importance_sampling_ratio/max": 1.730522871017456, "entropy": 0.030738614965230227, "clip_ratio/low_mean": 0.0016737572732381523, "clip_ratio/low_min": 0.0016737572732381523, "clip_ratio/high_mean": 0.0015337422955781221, "clip_ratio/high_max": 0.0015337422955781221, "clip_ratio/region_mean": 0.0032074995688162744, "reward_total_mean": 0.7980826497077942, "reward_meter_mean": 0.9976032972335815, "reward_meter_std": 0.0007398881716653705, "reward_count_adherence_mean": 0.800000011920929, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7980826497077942, "reward_total_composite_std": 0.0005919100949540734}
|
| 1122 |
+
{"timestamp_utc": "2026-04-11T21:45:15Z", "mode": "eval", "global_step": 1100, "epoch": 0.04247760271856657, "eval_loss": NaN, "eval_runtime": 73.9179, "eval_samples_per_second": 1.407, "eval_steps_per_second": 0.176, "eval_num_tokens": 2369329.0, "eval_completions/mean_length": 177.95192307692307, "eval_completions/min_length": 57.38461538461539, "eval_completions/max_length": 390.15384615384613, "eval_completions/clipped_ratio": 0.028846153846153848, "eval_completions/mean_terminated_length": 167.59066126896784, "eval_completions/min_terminated_length": 57.38461538461539, "eval_completions/max_terminated_length": 341.7692307692308, "eval_rewards/meter/mean": 0.6938754916191101, "eval_rewards/meter/std": 0.4057489140675618, "eval_rewards/count_adherence/mean": 0.8606285040195172, "eval_rewards/count_adherence/std": 0.14204404417138833, "eval_rewards/arabic_clean/mean": 0.9807692307692307, "eval_rewards/arabic_clean/std": 0.05439282839114849, "eval_rewards/total_composite/mean": 0.5808727557842548, "eval_rewards/total_composite/std": 0.372596684556741, "eval_reward": 0.5808727557842548, "eval_reward_std": NaN, "eval_frac_reward_zero_std": 0.0, "eval_sampling/sampling_logp_difference/mean": 0.0031142185639160182, "eval_sampling/sampling_logp_difference/max": 0.6837713030668405, "eval_sampling/importance_sampling_ratio/min": 0.542173194197508, "eval_sampling/importance_sampling_ratio/mean": 1.000536437218006, "eval_sampling/importance_sampling_ratio/max": 1.281661501297584, "eval_entropy": 0.022705065874526136, "eval_clip_ratio/low_mean": 0.0, "eval_clip_ratio/low_min": 0.0, "eval_clip_ratio/high_mean": 0.0, "eval_clip_ratio/high_max": 0.0, "eval_clip_ratio/region_mean": 0.0, "eval_reward_total_mean": 0.5808727557842548, "eval_reward_meter_mean": 0.6938754916191101, "eval_reward_meter_std": 0.4057489140675618, "eval_reward_count_adherence_mean": 0.8606285040195172, "eval_reward_count_adherence_std": 0.14204404417138833, "eval_reward_arabic_clean_mean": 0.9807692307692307, "eval_reward_arabic_clean_std": 0.05439282839114849, "eval_reward_total_composite_mean": 0.5808727557842548, "eval_reward_total_composite_std": 0.372596684556741}
|
| 1123 |
+
{"timestamp_utc": "2026-04-11T21:45:24Z", "mode": "train", "global_step": 1101, "epoch": 0.042516218721038, "loss": -0.2286, "grad_norm": 5.9600019454956055, "learning_rate": 6.666666666666667e-06, "num_tokens": 2371113.0, "completions/mean_length": 62.0, "completions/min_length": 18.0, "completions/max_length": 73.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 62.0, "completions/min_terminated_length": 18.0, "completions/max_terminated_length": 73.0, "rewards/meter/mean": 0.772131621837616, "rewards/meter/std": 0.34598708152770996, "rewards/count_adherence/mean": 0.875, "rewards/count_adherence/std": 0.3535533845424652, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.772131621837616, "rewards/total_composite/std": 0.34598708152770996, "reward": 0.772131621837616, "reward_std": 0.34598708152770996, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.040163177996873856, "sampling/sampling_logp_difference/max": 1.5425870418548584, "sampling/importance_sampling_ratio/min": 0.2138272076845169, "sampling/importance_sampling_ratio/mean": 1.0026087760925293, "sampling/importance_sampling_ratio/max": 1.7506089210510254, "entropy": 0.2632972849532962, "clip_ratio/low_mean": 0.01245915051549673, "clip_ratio/low_min": 0.01245915051549673, "clip_ratio/high_mean": 0.023881223052740097, "clip_ratio/high_max": 0.023881223052740097, "clip_ratio/region_mean": 0.03634037356823683, "reward_total_mean": 0.772131621837616, "reward_meter_mean": 0.772131621837616, "reward_meter_std": 0.34598708152770996, "reward_count_adherence_mean": 0.875, "reward_count_adherence_std": 0.3535533845424652, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.772131621837616, "reward_total_composite_std": 0.34598708152770996}
|
plots/arabic_gate_chain.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plots/arabic_gate_run.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plots/chain_metrics.jsonl
CHANGED
|
@@ -1068,3 +1068,54 @@
|
|
| 1068 |
{"timestamp_utc": "2026-04-11T21:37:56Z", "mode": "train", "global_step": 1048, "epoch": 0.04046957059005252, "loss": -0.0071, "grad_norm": 3.083638906478882, "learning_rate": 6.827272727272728e-06, "num_tokens": 2259982.0, "completions/mean_length": 65.75, "completions/min_length": 61.0, "completions/max_length": 72.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.75, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 72.0, "rewards/meter/mean": 0.9918305277824402, "rewards/meter/std": 0.012018299661576748, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9918305277824402, "rewards/total_composite/std": 0.012018299661576748, "reward": 0.9918305277824402, "reward_std": 0.012018297798931599, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.011509421281516552, "sampling/sampling_logp_difference/max": 0.996147632598877, "sampling/importance_sampling_ratio/min": 0.36929938197135925, "sampling/importance_sampling_ratio/mean": 1.0007789134979248, "sampling/importance_sampling_ratio/max": 1.513947606086731, "entropy": 0.06583833554759622, "clip_ratio/low_mean": 0.00390625, "clip_ratio/low_min": 0.00390625, "clip_ratio/high_mean": 0.0071314104134216905, "clip_ratio/high_max": 0.0071314104134216905, "clip_ratio/region_mean": 0.01103766041342169, "reward_total_mean": 0.9918305277824402, "reward_meter_mean": 0.9918305277824402, "reward_meter_std": 0.012018299661576748, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9918305277824402, "reward_total_composite_std": 0.012018299661576748, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1069 |
{"timestamp_utc": "2026-04-11T21:38:04Z", "mode": "train", "global_step": 1049, "epoch": 0.04050818659252394, "loss": 0.0192, "grad_norm": 2.3660078048706055, "learning_rate": 6.824242424242425e-06, "num_tokens": 2263825.0, "completions/mean_length": 283.375, "completions/min_length": 279.0, "completions/max_length": 291.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 283.375, "completions/min_terminated_length": 279.0, "completions/max_terminated_length": 291.0, "rewards/meter/mean": 0.9826633334159851, "rewards/meter/std": 0.0013277583057060838, "rewards/count_adherence/mean": 0.578125, "rewards/count_adherence/std": 0.06469365209341049, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.5680685043334961, "rewards/total_composite/std": 0.06325043737888336, "reward": 0.5680685043334961, "reward_std": 0.06325043737888336, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0027596044819802046, "sampling/sampling_logp_difference/max": 1.5669541358947754, "sampling/importance_sampling_ratio/min": 0.20867981016635895, "sampling/importance_sampling_ratio/mean": 1.0000476837158203, "sampling/importance_sampling_ratio/max": 1.6442346572875977, "entropy": 0.006785797595512122, "clip_ratio/low_mean": 0.0012901410227641463, "clip_ratio/low_min": 0.0012901410227641463, "clip_ratio/high_mean": 0.0013440860202535987, "clip_ratio/high_max": 0.0013440860202535987, "clip_ratio/region_mean": 0.002634227043017745, "reward_total_mean": 0.5680685043334961, "reward_meter_mean": 0.9826633334159851, "reward_meter_std": 0.0013277583057060838, "reward_count_adherence_mean": 0.578125, "reward_count_adherence_std": 0.06469365209341049, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.5680685043334961, "reward_total_composite_std": 0.06325043737888336, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1070 |
{"timestamp_utc": "2026-04-11T21:38:09Z", "mode": "train", "global_step": 1050, "epoch": 0.040546802594995365, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.821212121212122e-06, "num_tokens": 2265729.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9968419075012207, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9968419075012207, "rewards/total_composite/std": 0.0, "reward": 0.9968419075012207, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0027122944593429565, "sampling/sampling_logp_difference/max": 0.7146163582801819, "sampling/importance_sampling_ratio/min": 0.4893798530101776, "sampling/importance_sampling_ratio/mean": 0.999724805355072, "sampling/importance_sampling_ratio/max": 1.1012705564498901, "entropy": 0.007037240080535412, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9968419075012207, "reward_meter_mean": 0.9968419075012207, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9968419075012207, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1068 |
{"timestamp_utc": "2026-04-11T21:37:56Z", "mode": "train", "global_step": 1048, "epoch": 0.04046957059005252, "loss": -0.0071, "grad_norm": 3.083638906478882, "learning_rate": 6.827272727272728e-06, "num_tokens": 2259982.0, "completions/mean_length": 65.75, "completions/min_length": 61.0, "completions/max_length": 72.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.75, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 72.0, "rewards/meter/mean": 0.9918305277824402, "rewards/meter/std": 0.012018299661576748, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9918305277824402, "rewards/total_composite/std": 0.012018299661576748, "reward": 0.9918305277824402, "reward_std": 0.012018297798931599, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.011509421281516552, "sampling/sampling_logp_difference/max": 0.996147632598877, "sampling/importance_sampling_ratio/min": 0.36929938197135925, "sampling/importance_sampling_ratio/mean": 1.0007789134979248, "sampling/importance_sampling_ratio/max": 1.513947606086731, "entropy": 0.06583833554759622, "clip_ratio/low_mean": 0.00390625, "clip_ratio/low_min": 0.00390625, "clip_ratio/high_mean": 0.0071314104134216905, "clip_ratio/high_max": 0.0071314104134216905, "clip_ratio/region_mean": 0.01103766041342169, "reward_total_mean": 0.9918305277824402, "reward_meter_mean": 0.9918305277824402, "reward_meter_std": 0.012018299661576748, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9918305277824402, "reward_total_composite_std": 0.012018299661576748, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1069 |
{"timestamp_utc": "2026-04-11T21:38:04Z", "mode": "train", "global_step": 1049, "epoch": 0.04050818659252394, "loss": 0.0192, "grad_norm": 2.3660078048706055, "learning_rate": 6.824242424242425e-06, "num_tokens": 2263825.0, "completions/mean_length": 283.375, "completions/min_length": 279.0, "completions/max_length": 291.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 283.375, "completions/min_terminated_length": 279.0, "completions/max_terminated_length": 291.0, "rewards/meter/mean": 0.9826633334159851, "rewards/meter/std": 0.0013277583057060838, "rewards/count_adherence/mean": 0.578125, "rewards/count_adherence/std": 0.06469365209341049, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.5680685043334961, "rewards/total_composite/std": 0.06325043737888336, "reward": 0.5680685043334961, "reward_std": 0.06325043737888336, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0027596044819802046, "sampling/sampling_logp_difference/max": 1.5669541358947754, "sampling/importance_sampling_ratio/min": 0.20867981016635895, "sampling/importance_sampling_ratio/mean": 1.0000476837158203, "sampling/importance_sampling_ratio/max": 1.6442346572875977, "entropy": 0.006785797595512122, "clip_ratio/low_mean": 0.0012901410227641463, "clip_ratio/low_min": 0.0012901410227641463, "clip_ratio/high_mean": 0.0013440860202535987, "clip_ratio/high_max": 0.0013440860202535987, "clip_ratio/region_mean": 0.002634227043017745, "reward_total_mean": 0.5680685043334961, "reward_meter_mean": 0.9826633334159851, "reward_meter_std": 0.0013277583057060838, "reward_count_adherence_mean": 0.578125, "reward_count_adherence_std": 0.06469365209341049, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.5680685043334961, "reward_total_composite_std": 0.06325043737888336, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1070 |
{"timestamp_utc": "2026-04-11T21:38:09Z", "mode": "train", "global_step": 1050, "epoch": 0.040546802594995365, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.821212121212122e-06, "num_tokens": 2265729.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9968419075012207, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9968419075012207, "rewards/total_composite/std": 0.0, "reward": 0.9968419075012207, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0027122944593429565, "sampling/sampling_logp_difference/max": 0.7146163582801819, "sampling/importance_sampling_ratio/min": 0.4893798530101776, "sampling/importance_sampling_ratio/mean": 0.999724805355072, "sampling/importance_sampling_ratio/max": 1.1012705564498901, "entropy": 0.007037240080535412, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9968419075012207, "reward_meter_mean": 0.9968419075012207, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9968419075012207, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1071 |
+
{"timestamp_utc": "2026-04-11T21:39:35Z", "mode": "eval", "global_step": 1050, "epoch": 0.040546802594995365, "eval_loss": NaN, "eval_runtime": 85.4042, "eval_samples_per_second": 1.218, "eval_steps_per_second": 0.152, "eval_num_tokens": 2265729.0, "eval_completions/mean_length": 241.28846153846155, "eval_completions/min_length": 59.61538461538461, "eval_completions/max_length": 458.0, "eval_completions/clipped_ratio": 0.057692307692307696, "eval_completions/mean_terminated_length": 224.83516869178186, "eval_completions/min_terminated_length": 59.61538461538461, "eval_completions/max_terminated_length": 412.46153846153845, "eval_rewards/meter/mean": 0.6695947761719043, "eval_rewards/meter/std": 0.4209407364519743, "eval_rewards/count_adherence/mean": 0.8243002341343806, "eval_rewards/count_adherence/std": 0.149119944526599, "eval_rewards/arabic_clean/mean": 0.9903846153846154, "eval_rewards/arabic_clean/std": 0.027196414195574246, "eval_rewards/total_composite/mean": 0.5442026945260855, "eval_rewards/total_composite/std": 0.3753266856074333, "eval_reward": 0.5442026945260855, "eval_reward_std": NaN, "eval_frac_reward_zero_std": 0.0, "eval_sampling/sampling_logp_difference/mean": 0.0014539593382953452, "eval_sampling/sampling_logp_difference/max": 0.48829063085409313, "eval_sampling/importance_sampling_ratio/min": 0.6470333154384906, "eval_sampling/importance_sampling_ratio/mean": 1.0002473134260912, "eval_sampling/importance_sampling_ratio/max": 1.2309852104920607, "eval_entropy": 0.009740084463443894, "eval_clip_ratio/low_mean": 0.0, "eval_clip_ratio/low_min": 0.0, "eval_clip_ratio/high_mean": 0.0, "eval_clip_ratio/high_max": 0.0, "eval_clip_ratio/region_mean": 0.0, "eval_reward_total_mean": 0.5442026945260855, "eval_reward_meter_mean": 0.6695947761719043, "eval_reward_meter_std": 0.4209407364519743, "eval_reward_count_adherence_mean": 0.8243002341343806, "eval_reward_count_adherence_std": 0.149119944526599, "eval_reward_arabic_clean_mean": 0.9903846153846154, "eval_reward_arabic_clean_std": 0.027196414195574246, "eval_reward_total_composite_mean": 0.5442026945260855, "eval_reward_total_composite_std": 0.3753266856074333, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1072 |
+
{"timestamp_utc": "2026-04-11T21:39:43Z", "mode": "train", "global_step": 1051, "epoch": 0.04058541859746679, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.818181818181818e-06, "num_tokens": 2267377.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00045387333375401795, "sampling/sampling_logp_difference/max": 0.042022280395030975, "sampling/importance_sampling_ratio/min": 0.9996206760406494, "sampling/importance_sampling_ratio/mean": 1.0004585981369019, "sampling/importance_sampling_ratio/max": 1.0429177284240723, "entropy": 0.004360992228612304, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1073 |
+
{"timestamp_utc": "2026-04-11T21:39:48Z", "mode": "train", "global_step": 1052, "epoch": 0.040624034599938214, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.8151515151515155e-06, "num_tokens": 2269065.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9968419075012207, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9968419075012207, "rewards/total_composite/std": 0.0, "reward": 0.9968419075012207, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0007426206138916314, "sampling/sampling_logp_difference/max": 0.09673504531383514, "sampling/importance_sampling_ratio/min": 0.9077965617179871, "sampling/importance_sampling_ratio/mean": 1.0001823902130127, "sampling/importance_sampling_ratio/max": 1.0188950300216675, "entropy": 0.004134227987378836, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9968419075012207, "reward_meter_mean": 0.9968419075012207, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9968419075012207, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1074 |
+
{"timestamp_utc": "2026-04-11T21:39:53Z", "mode": "train", "global_step": 1053, "epoch": 0.04066265060240964, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.812121212121212e-06, "num_tokens": 2270753.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00017959998513106257, "sampling/sampling_logp_difference/max": 0.011735539883375168, "sampling/importance_sampling_ratio/min": 0.9999438524246216, "sampling/importance_sampling_ratio/mean": 1.0001797676086426, "sampling/importance_sampling_ratio/max": 1.0118045806884766, "entropy": 0.0012685270776273683, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1075 |
+
{"timestamp_utc": "2026-04-11T21:39:57Z", "mode": "train", "global_step": 1054, "epoch": 0.04070126660488106, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.80909090909091e-06, "num_tokens": 2272529.0, "completions/mean_length": 61.0, "completions/min_length": 61.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.0, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00012025728210574016, "sampling/sampling_logp_difference/max": 0.008706126362085342, "sampling/importance_sampling_ratio/min": 0.9994588494300842, "sampling/importance_sampling_ratio/mean": 1.000115990638733, "sampling/importance_sampling_ratio/max": 1.008744239807129, "entropy": 0.0013022492494201288, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1076 |
+
{"timestamp_utc": "2026-04-11T21:40:02Z", "mode": "train", "global_step": 1055, "epoch": 0.040739882607352486, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.806060606060607e-06, "num_tokens": 2274175.0, "completions/mean_length": 50.75, "completions/min_length": 50.0, "completions/max_length": 51.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 50.75, "completions/min_terminated_length": 50.0, "completions/max_terminated_length": 51.0, "rewards/meter/mean": 0.9589456915855408, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9589456915855408, "rewards/total_composite/std": 0.0, "reward": 0.9589456915855408, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0029871375299990177, "sampling/sampling_logp_difference/max": 0.30701398849487305, "sampling/importance_sampling_ratio/min": 0.7356403470039368, "sampling/importance_sampling_ratio/mean": 1.0008618831634521, "sampling/importance_sampling_ratio/max": 1.2294468879699707, "entropy": 0.0185652831569314, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9589456915855408, "reward_meter_mean": 0.9589456915855408, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9589456915855408, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1077 |
+
{"timestamp_utc": "2026-04-11T21:40:07Z", "mode": "train", "global_step": 1056, "epoch": 0.04077849860982391, "loss": 0.008, "grad_norm": 3.9809470176696777, "learning_rate": 6.803030303030304e-06, "num_tokens": 2275998.0, "completions/mean_length": 62.875, "completions/min_length": 61.0, "completions/max_length": 64.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 62.875, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 64.0, "rewards/meter/mean": 0.0021876515820622444, "rewards/meter/std": 0.0002445352729409933, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.0021876515820622444, "rewards/total_composite/std": 0.0002445352729409933, "reward": 0.0021876515820622444, "reward_std": 0.00024453524383716285, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.008529768325388432, "sampling/sampling_logp_difference/max": 1.1146979331970215, "sampling/importance_sampling_ratio/min": 0.6839931607246399, "sampling/importance_sampling_ratio/mean": 1.0037806034088135, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.041025768499821424, "clip_ratio/low_mean": 0.013767930213361979, "clip_ratio/low_min": 0.013767930213361979, "clip_ratio/high_mean": 0.0020491802133619785, "clip_ratio/high_max": 0.0020491802133619785, "clip_ratio/region_mean": 0.015817110426723957, "reward_total_mean": 0.0021876515820622444, "reward_meter_mean": 0.0021876515820622444, "reward_meter_std": 0.0002445352729409933, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.0021876515820622444, "reward_total_composite_std": 0.0002445352729409933, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1078 |
+
{"timestamp_utc": "2026-04-11T21:40:12Z", "mode": "train", "global_step": 1057, "epoch": 0.040817114612295334, "loss": -0.0561, "grad_norm": 0.9426341652870178, "learning_rate": 6.800000000000001e-06, "num_tokens": 2277958.0, "completions/mean_length": 83.0, "completions/min_length": 81.0, "completions/max_length": 97.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 83.0, "completions/min_terminated_length": 81.0, "completions/max_terminated_length": 97.0, "rewards/meter/mean": 0.9951313734054565, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.7083333730697632, "rewards/count_adherence/std": 0.11785111576318741, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7048847079277039, "rewards/total_composite/std": 0.1172773540019989, "reward": 0.7048847079277039, "reward_std": 0.1172773465514183, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.00949889700859785, "sampling/sampling_logp_difference/max": 3.2940514087677, "sampling/importance_sampling_ratio/min": 0.03710322454571724, "sampling/importance_sampling_ratio/mean": 0.9974780082702637, "sampling/importance_sampling_ratio/max": 1.134204387664795, "entropy": 0.003993239166447893, "clip_ratio/low_mean": 0.0015432098880410194, "clip_ratio/low_min": 0.0015432098880410194, "clip_ratio/high_mean": 0.0012886597542092204, "clip_ratio/high_max": 0.0012886597542092204, "clip_ratio/region_mean": 0.00283186964225024, "reward_total_mean": 0.7048847079277039, "reward_meter_mean": 0.9951313734054565, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.7083333730697632, "reward_count_adherence_std": 0.11785111576318741, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7048847079277039, "reward_total_composite_std": 0.1172773540019989, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1079 |
+
{"timestamp_utc": "2026-04-11T21:40:19Z", "mode": "train", "global_step": 1058, "epoch": 0.04085573061476676, "loss": 0.0039, "grad_norm": 0.21764229238033295, "learning_rate": 6.796969696969697e-06, "num_tokens": 2280914.0, "completions/mean_length": 198.5, "completions/min_length": 197.0, "completions/max_length": 209.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 198.5, "completions/min_terminated_length": 197.0, "completions/max_terminated_length": 209.0, "rewards/meter/mean": 0.9969083070755005, "rewards/meter/std": 3.72788890672382e-05, "rewards/count_adherence/mean": 0.800000011920929, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7975266575813293, "rewards/total_composite/std": 2.9818897019140422e-05, "reward": 0.7975266575813293, "reward_std": 2.9827928301529028e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.003239656798541546, "sampling/sampling_logp_difference/max": 3.510453224182129, "sampling/importance_sampling_ratio/min": 0.0298833679407835, "sampling/importance_sampling_ratio/mean": 0.9991796016693115, "sampling/importance_sampling_ratio/max": 1.167437195777893, "entropy": 0.0015565359972242732, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0006345177534967661, "clip_ratio/high_max": 0.0006345177534967661, "clip_ratio/region_mean": 0.0006345177534967661, "reward_total_mean": 0.7975266575813293, "reward_meter_mean": 0.9969083070755005, "reward_meter_std": 3.72788890672382e-05, "reward_count_adherence_mean": 0.800000011920929, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7975266575813293, "reward_total_composite_std": 2.9818897019140422e-05, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1080 |
+
{"timestamp_utc": "2026-04-11T21:40:23Z", "mode": "train", "global_step": 1059, "epoch": 0.04089434661723818, "loss": 0.0234, "grad_norm": 4.7054972648620605, "learning_rate": 6.793939393939395e-06, "num_tokens": 2282612.0, "completions/mean_length": 54.25, "completions/min_length": 52.0, "completions/max_length": 57.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 54.25, "completions/min_terminated_length": 52.0, "completions/max_terminated_length": 57.0, "rewards/meter/mean": 0.9560332298278809, "rewards/meter/std": 0.0188444871455431, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9560332298278809, "rewards/total_composite/std": 0.0188444871455431, "reward": 0.9560332298278809, "reward_std": 0.018844490870833397, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.008652369491755962, "sampling/sampling_logp_difference/max": 0.96638023853302, "sampling/importance_sampling_ratio/min": 0.38045772910118103, "sampling/importance_sampling_ratio/mean": 0.9992375373840332, "sampling/importance_sampling_ratio/max": 1.3655518293380737, "entropy": 0.025705090374685824, "clip_ratio/low_mean": 0.00657894741743803, "clip_ratio/low_min": 0.00657894741743803, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.00657894741743803, "reward_total_mean": 0.9560332298278809, "reward_meter_mean": 0.9560332298278809, "reward_meter_std": 0.0188444871455431, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9560332298278809, "reward_total_composite_std": 0.0188444871455431, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1081 |
+
{"timestamp_utc": "2026-04-11T21:40:31Z", "mode": "train", "global_step": 1060, "epoch": 0.040932962619709606, "loss": 0.0057, "grad_norm": 0.09687843173742294, "learning_rate": 6.790909090909091e-06, "num_tokens": 2286469.0, "completions/mean_length": 274.125, "completions/min_length": 272.0, "completions/max_length": 289.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 274.125, "completions/min_terminated_length": 272.0, "completions/max_terminated_length": 289.0, "rewards/meter/mean": 0.9969933032989502, "rewards/meter/std": 5.181956657906994e-05, "rewards/count_adherence/mean": 0.875, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8723691701889038, "rewards/total_composite/std": 4.535001062322408e-05, "reward": 0.8723691701889038, "reward_std": 4.535001062322408e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.00142471503932029, "sampling/sampling_logp_difference/max": 1.230623722076416, "sampling/importance_sampling_ratio/min": 0.2921103239059448, "sampling/importance_sampling_ratio/mean": 0.9990590214729309, "sampling/importance_sampling_ratio/max": 1.0306644439697266, "entropy": 0.0006548625879077008, "clip_ratio/low_mean": 0.00043252596515230834, "clip_ratio/low_min": 0.00043252596515230834, "clip_ratio/high_mean": 0.0018382353009656072, "clip_ratio/high_max": 0.0018382353009656072, "clip_ratio/region_mean": 0.0022707612661179155, "reward_total_mean": 0.8723691701889038, "reward_meter_mean": 0.9969933032989502, "reward_meter_std": 5.181956657906994e-05, "reward_count_adherence_mean": 0.875, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8723691701889038, "reward_total_composite_std": 4.535001062322408e-05, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1082 |
+
{"timestamp_utc": "2026-04-11T21:40:38Z", "mode": "train", "global_step": 1061, "epoch": 0.04097157862218103, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.787878787878789e-06, "num_tokens": 2290485.0, "completions/mean_length": 272.0, "completions/min_length": 272.0, "completions/max_length": 272.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 272.0, "completions/min_terminated_length": 272.0, "completions/max_terminated_length": 272.0, "rewards/meter/mean": 0.997011661529541, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.875, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8723852038383484, "rewards/total_composite/std": 0.0, "reward": 0.8723852038383484, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 6.0048791056033224e-05, "sampling/sampling_logp_difference/max": 0.03515136241912842, "sampling/importance_sampling_ratio/min": 0.9654592871665955, "sampling/importance_sampling_ratio/mean": 0.9999895691871643, "sampling/importance_sampling_ratio/max": 1.0054314136505127, "entropy": 0.00023714406961516943, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.8723852038383484, "reward_meter_mean": 0.997011661529541, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.875, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8723852038383484, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1083 |
+
{"timestamp_utc": "2026-04-11T21:40:44Z", "mode": "train", "global_step": 1062, "epoch": 0.041010194624652455, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.7848484848484855e-06, "num_tokens": 2292821.0, "completions/mean_length": 106.0, "completions/min_length": 106.0, "completions/max_length": 106.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 106.0, "completions/min_terminated_length": 106.0, "completions/max_terminated_length": 106.0, "rewards/meter/mean": 0.9985920786857605, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9985920786857605, "rewards/total_composite/std": 0.0, "reward": 0.9985920786857605, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.01337476447224617, "sampling/sampling_logp_difference/max": 7.350694179534912, "sampling/importance_sampling_ratio/min": 0.0006421464495360851, "sampling/importance_sampling_ratio/mean": 0.9978539347648621, "sampling/importance_sampling_ratio/max": 1.0233614444732666, "entropy": 0.0023513801133958623, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9985920786857605, "reward_meter_mean": 0.9985920786857605, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9985920786857605, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1084 |
+
{"timestamp_utc": "2026-04-11T21:40:48Z", "mode": "train", "global_step": 1063, "epoch": 0.04104881062712388, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.781818181818183e-06, "num_tokens": 2294669.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00023568868346046656, "sampling/sampling_logp_difference/max": 0.017239108681678772, "sampling/importance_sampling_ratio/min": 0.9829086661338806, "sampling/importance_sampling_ratio/mean": 1.0001155138015747, "sampling/importance_sampling_ratio/max": 1.011854648590088, "entropy": 0.0019514950545271859, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1085 |
+
{"timestamp_utc": "2026-04-11T21:40:53Z", "mode": "train", "global_step": 1064, "epoch": 0.0410874266295953, "loss": -0.0026, "grad_norm": 2.7271664142608643, "learning_rate": 6.778787878787879e-06, "num_tokens": 2296475.0, "completions/mean_length": 54.75, "completions/min_length": 53.0, "completions/max_length": 57.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 54.75, "completions/min_terminated_length": 53.0, "completions/max_terminated_length": 57.0, "rewards/meter/mean": 0.9718772172927856, "rewards/meter/std": 0.01329784281551838, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9718772172927856, "rewards/total_composite/std": 0.01329784281551838, "reward": 0.9718772172927856, "reward_std": 0.013297837227582932, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.024222884327173233, "sampling/sampling_logp_difference/max": 5.190197944641113, "sampling/importance_sampling_ratio/min": 0.005570904351770878, "sampling/importance_sampling_ratio/mean": 0.9998058676719666, "sampling/importance_sampling_ratio/max": 1.5688531398773193, "entropy": 0.10237977746874094, "clip_ratio/low_mean": 0.009181267116218805, "clip_ratio/low_min": 0.009181267116218805, "clip_ratio/high_mean": 0.002314814832061529, "clip_ratio/high_max": 0.002314814832061529, "clip_ratio/region_mean": 0.011496081948280334, "reward_total_mean": 0.9718772172927856, "reward_meter_mean": 0.9718772172927856, "reward_meter_std": 0.01329784281551838, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9718772172927856, "reward_total_composite_std": 0.01329784281551838, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1086 |
+
{"timestamp_utc": "2026-04-11T21:40:58Z", "mode": "train", "global_step": 1065, "epoch": 0.04112604263206673, "loss": 0.0102, "grad_norm": 5.474052429199219, "learning_rate": 6.7757575757575765e-06, "num_tokens": 2298309.0, "completions/mean_length": 57.25, "completions/min_length": 57.0, "completions/max_length": 59.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 57.25, "completions/min_terminated_length": 57.0, "completions/max_terminated_length": 59.0, "rewards/meter/mean": 0.8288533687591553, "rewards/meter/std": 0.037421029061079025, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8288533687591553, "rewards/total_composite/std": 0.037421029061079025, "reward": 0.8288533687591553, "reward_std": 0.03742102161049843, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.002723454497754574, "sampling/sampling_logp_difference/max": 0.5241236686706543, "sampling/importance_sampling_ratio/min": 0.5920739769935608, "sampling/importance_sampling_ratio/mean": 0.9999019503593445, "sampling/importance_sampling_ratio/max": 1.1381598711013794, "entropy": 0.01174389524385333, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.8288533687591553, "reward_meter_mean": 0.8288533687591553, "reward_meter_std": 0.037421029061079025, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8288533687591553, "reward_total_composite_std": 0.037421029061079025, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1087 |
+
{"timestamp_utc": "2026-04-11T21:41:02Z", "mode": "train", "global_step": 1066, "epoch": 0.04116465863453815, "loss": 0.0097, "grad_norm": 2.398463726043701, "learning_rate": 6.772727272727273e-06, "num_tokens": 2300092.0, "completions/mean_length": 60.875, "completions/min_length": 60.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 60.875, "completions/min_terminated_length": 60.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.9969611763954163, "rewards/meter/std": 0.0003373012295924127, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9969611763954163, "rewards/total_composite/std": 0.0003373012295924127, "reward": 0.9969611763954163, "reward_std": 0.00033730725408531725, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0037930156104266644, "sampling/sampling_logp_difference/max": 1.299863576889038, "sampling/importance_sampling_ratio/min": 0.2725689709186554, "sampling/importance_sampling_ratio/mean": 0.9994586110115051, "sampling/importance_sampling_ratio/max": 1.084521770477295, "entropy": 0.014361659123096615, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0020833334419876337, "clip_ratio/high_max": 0.0020833334419876337, "clip_ratio/region_mean": 0.0020833334419876337, "reward_total_mean": 0.9969611763954163, "reward_meter_mean": 0.9969611763954163, "reward_meter_std": 0.0003373012295924127, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9969611763954163, "reward_total_composite_std": 0.0003373012295924127, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1088 |
+
{"timestamp_utc": "2026-04-11T21:41:07Z", "mode": "train", "global_step": 1067, "epoch": 0.041203274637009575, "loss": 0.0034, "grad_norm": 10.166537284851074, "learning_rate": 6.76969696969697e-06, "num_tokens": 2301589.0, "completions/mean_length": 27.125, "completions/min_length": 26.0, "completions/max_length": 29.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 27.125, "completions/min_terminated_length": 26.0, "completions/max_terminated_length": 29.0, "rewards/meter/mean": 0.9321423768997192, "rewards/meter/std": 0.007329464890062809, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9321423768997192, "rewards/total_composite/std": 0.007329464890062809, "reward": 0.9321423768997192, "reward_std": 0.007329456973820925, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.026033449918031693, "sampling/sampling_logp_difference/max": 1.5315618515014648, "sampling/importance_sampling_ratio/min": 0.21619774401187897, "sampling/importance_sampling_ratio/mean": 1.0001667737960815, "sampling/importance_sampling_ratio/max": 1.4033339023590088, "entropy": 0.09683473920449615, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.00893997447565198, "clip_ratio/high_max": 0.00893997447565198, "clip_ratio/region_mean": 0.00893997447565198, "reward_total_mean": 0.9321423768997192, "reward_meter_mean": 0.9321423768997192, "reward_meter_std": 0.007329464890062809, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9321423768997192, "reward_total_composite_std": 0.007329464890062809, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1089 |
+
{"timestamp_utc": "2026-04-11T21:41:12Z", "mode": "train", "global_step": 1068, "epoch": 0.041241890639481, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.7666666666666665e-06, "num_tokens": 2303405.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00019105577666778117, "sampling/sampling_logp_difference/max": 0.0064575872384011745, "sampling/importance_sampling_ratio/min": 0.9977074861526489, "sampling/importance_sampling_ratio/mean": 1.000160813331604, "sampling/importance_sampling_ratio/max": 1.0064785480499268, "entropy": 0.0018052591913146898, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1090 |
+
{"timestamp_utc": "2026-04-11T21:41:17Z", "mode": "train", "global_step": 1069, "epoch": 0.04128050664195242, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.763636363636365e-06, "num_tokens": 2305133.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0005417983047664165, "sampling/sampling_logp_difference/max": 0.036248765885829926, "sampling/importance_sampling_ratio/min": 0.9796718955039978, "sampling/importance_sampling_ratio/mean": 1.0003167390823364, "sampling/importance_sampling_ratio/max": 1.0369137525558472, "entropy": 0.005523053434444591, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1091 |
+
{"timestamp_utc": "2026-04-11T21:41:21Z", "mode": "train", "global_step": 1070, "epoch": 0.04131912264442385, "loss": 0.0381, "grad_norm": 26.325374603271484, "learning_rate": 6.760606060606061e-06, "num_tokens": 2306762.0, "completions/mean_length": 52.625, "completions/min_length": 50.0, "completions/max_length": 54.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 52.625, "completions/min_terminated_length": 50.0, "completions/max_terminated_length": 54.0, "rewards/meter/mean": 0.11274315416812897, "rewards/meter/std": 0.3188244700431824, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.11274315416812897, "rewards/total_composite/std": 0.3188244700431824, "reward": 0.11274315416812897, "reward_std": 0.31882444024086, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.02075301483273506, "sampling/sampling_logp_difference/max": 1.0765538215637207, "sampling/importance_sampling_ratio/min": 0.34076783061027527, "sampling/importance_sampling_ratio/mean": 1.0003687143325806, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.05561392899835482, "clip_ratio/low_mean": 0.016467438312247396, "clip_ratio/low_min": 0.016467438312247396, "clip_ratio/high_mean": 0.007499999832361937, "clip_ratio/high_max": 0.007499999832361937, "clip_ratio/region_mean": 0.023967438144609332, "reward_total_mean": 0.11274315416812897, "reward_meter_mean": 0.11274315416812897, "reward_meter_std": 0.3188244700431824, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.11274315416812897, "reward_total_composite_std": 0.3188244700431824, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1092 |
+
{"timestamp_utc": "2026-04-11T21:41:26Z", "mode": "train", "global_step": 1071, "epoch": 0.04135773864689527, "loss": -0.0217, "grad_norm": 5.607251167297363, "learning_rate": 6.757575757575758e-06, "num_tokens": 2308113.0, "completions/mean_length": 25.875, "completions/min_length": 25.0, "completions/max_length": 26.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 25.875, "completions/min_terminated_length": 25.0, "completions/max_terminated_length": 26.0, "rewards/meter/mean": 0.9891167879104614, "rewards/meter/std": 0.008078054524958134, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9891167879104614, "rewards/total_composite/std": 0.008078054524958134, "reward": 0.9891167879104614, "reward_std": 0.008078045211732388, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.009091660380363464, "sampling/sampling_logp_difference/max": 0.6194605827331543, "sampling/importance_sampling_ratio/min": 0.5382347106933594, "sampling/importance_sampling_ratio/mean": 0.9988547563552856, "sampling/importance_sampling_ratio/max": 1.0410070419311523, "entropy": 0.02826369390822947, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9891167879104614, "reward_meter_mean": 0.9891167879104614, "reward_meter_std": 0.008078054524958134, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9891167879104614, "reward_total_composite_std": 0.008078054524958134, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1093 |
+
{"timestamp_utc": "2026-04-11T21:41:31Z", "mode": "train", "global_step": 1072, "epoch": 0.041396354649366696, "loss": 0.0001, "grad_norm": 2.4955992698669434, "learning_rate": 6.754545454545455e-06, "num_tokens": 2309830.0, "completions/mean_length": 61.625, "completions/min_length": 61.0, "completions/max_length": 62.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.625, "completions/min_terminated_length": 61.0, "completions/max_terminated_length": 62.0, "rewards/meter/mean": 0.997157096862793, "rewards/meter/std": 0.0005752869765274227, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.997157096862793, "rewards/total_composite/std": 0.0005752869765274227, "reward": 0.997157096862793, "reward_std": 0.000575280690100044, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.012418783269822598, "sampling/sampling_logp_difference/max": 0.5172317028045654, "sampling/importance_sampling_ratio/min": 0.5961686372756958, "sampling/importance_sampling_ratio/mean": 1.0032328367233276, "sampling/importance_sampling_ratio/max": 1.3372715711593628, "entropy": 0.10848722152877599, "clip_ratio/low_mean": 0.010146747343242168, "clip_ratio/low_min": 0.010146747343242168, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.010146747343242168, "reward_total_mean": 0.997157096862793, "reward_meter_mean": 0.997157096862793, "reward_meter_std": 0.0005752869765274227, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.997157096862793, "reward_total_composite_std": 0.0005752869765274227, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1094 |
+
{"timestamp_utc": "2026-04-11T21:41:36Z", "mode": "train", "global_step": 1073, "epoch": 0.04143497065183812, "loss": 0.0115, "grad_norm": 0.9153868556022644, "learning_rate": 6.751515151515152e-06, "num_tokens": 2311287.0, "completions/mean_length": 33.125, "completions/min_length": 33.0, "completions/max_length": 34.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 33.125, "completions/min_terminated_length": 33.0, "completions/max_terminated_length": 34.0, "rewards/meter/mean": 0.9944090247154236, "rewards/meter/std": 0.0004998616641387343, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9944090247154236, "rewards/total_composite/std": 0.0004998616641387343, "reward": 0.9944090247154236, "reward_std": 0.0004998616059310734, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.005571494810283184, "sampling/sampling_logp_difference/max": 0.6732907295227051, "sampling/importance_sampling_ratio/min": 0.9340354800224304, "sampling/importance_sampling_ratio/mean": 1.0061020851135254, "sampling/importance_sampling_ratio/max": 1.9606788158416748, "entropy": 0.025314107653684914, "clip_ratio/low_mean": 0.0036764706019312143, "clip_ratio/low_min": 0.0036764706019312143, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0036764706019312143, "reward_total_mean": 0.9944090247154236, "reward_meter_mean": 0.9944090247154236, "reward_meter_std": 0.0004998616641387343, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9944090247154236, "reward_total_composite_std": 0.0004998616641387343, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1095 |
+
{"timestamp_utc": "2026-04-11T21:41:41Z", "mode": "train", "global_step": 1074, "epoch": 0.041473586654309544, "loss": 0.0417, "grad_norm": 3.6375036239624023, "learning_rate": 6.748484848484848e-06, "num_tokens": 2313096.0, "completions/mean_length": 54.125, "completions/min_length": 51.0, "completions/max_length": 56.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 54.125, "completions/min_terminated_length": 51.0, "completions/max_terminated_length": 56.0, "rewards/meter/mean": 0.9875502586364746, "rewards/meter/std": 0.003662221832200885, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9875502586364746, "rewards/total_composite/std": 0.003662221832200885, "reward": 0.9875502586364746, "reward_std": 0.003662231843918562, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.004612638149410486, "sampling/sampling_logp_difference/max": 0.22617030143737793, "sampling/importance_sampling_ratio/min": 0.7975822687149048, "sampling/importance_sampling_ratio/mean": 1.003567099571228, "sampling/importance_sampling_ratio/max": 1.2383108139038086, "entropy": 0.03697874676436186, "clip_ratio/low_mean": 0.004464285913854837, "clip_ratio/low_min": 0.004464285913854837, "clip_ratio/high_mean": 0.0024509804788976908, "clip_ratio/high_max": 0.0024509804788976908, "clip_ratio/region_mean": 0.006915266392752528, "reward_total_mean": 0.9875502586364746, "reward_meter_mean": 0.9875502586364746, "reward_meter_std": 0.003662221832200885, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9875502586364746, "reward_total_composite_std": 0.003662221832200885, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1096 |
+
{"timestamp_utc": "2026-04-11T21:41:46Z", "mode": "train", "global_step": 1075, "epoch": 0.04151220265678097, "loss": 0.0288, "grad_norm": 6.4112677574157715, "learning_rate": 6.7454545454545465e-06, "num_tokens": 2314778.0, "completions/mean_length": 47.25, "completions/min_length": 41.0, "completions/max_length": 58.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 47.25, "completions/min_terminated_length": 41.0, "completions/max_terminated_length": 58.0, "rewards/meter/mean": 0.10002761334180832, "rewards/meter/std": 0.26244989037513733, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.10002761334180832, "rewards/total_composite/std": 0.26244989037513733, "reward": 0.10002761334180832, "reward_std": 0.26244989037513733, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.03375272452831268, "sampling/sampling_logp_difference/max": 1.2938915491104126, "sampling/importance_sampling_ratio/min": 0.2742016315460205, "sampling/importance_sampling_ratio/mean": 1.0019086599349976, "sampling/importance_sampling_ratio/max": 1.4211242198944092, "entropy": 0.16682779975235462, "clip_ratio/low_mean": 0.03481511096470058, "clip_ratio/low_min": 0.03481511096470058, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.03481511096470058, "reward_total_mean": 0.10002761334180832, "reward_meter_mean": 0.10002761334180832, "reward_meter_std": 0.26244989037513733, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.10002761334180832, "reward_total_composite_std": 0.26244989037513733, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1097 |
+
{"timestamp_utc": "2026-04-11T21:41:51Z", "mode": "train", "global_step": 1076, "epoch": 0.04155081865925239, "loss": -0.0001, "grad_norm": 4.325499057769775, "learning_rate": 6.742424242424243e-06, "num_tokens": 2316531.0, "completions/mean_length": 53.125, "completions/min_length": 52.0, "completions/max_length": 55.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 53.125, "completions/min_terminated_length": 52.0, "completions/max_terminated_length": 55.0, "rewards/meter/mean": 0.9805549383163452, "rewards/meter/std": 0.00861036404967308, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9805549383163452, "rewards/total_composite/std": 0.00861036404967308, "reward": 0.9805549383163452, "reward_std": 0.008610363118350506, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.011005771346390247, "sampling/sampling_logp_difference/max": 0.7941849231719971, "sampling/importance_sampling_ratio/min": 0.45194947719573975, "sampling/importance_sampling_ratio/mean": 1.0017882585525513, "sampling/importance_sampling_ratio/max": 1.5676127672195435, "entropy": 0.049260836094617844, "clip_ratio/low_mean": 0.009479318046942353, "clip_ratio/low_min": 0.009479318046942353, "clip_ratio/high_mean": 0.004631217801943421, "clip_ratio/high_max": 0.004631217801943421, "clip_ratio/region_mean": 0.014110535848885775, "reward_total_mean": 0.9805549383163452, "reward_meter_mean": 0.9805549383163452, "reward_meter_std": 0.00861036404967308, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9805549383163452, "reward_total_composite_std": 0.00861036404967308, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1098 |
+
{"timestamp_utc": "2026-04-11T21:41:56Z", "mode": "train", "global_step": 1077, "epoch": 0.041589434661723816, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.73939393939394e-06, "num_tokens": 2318435.0, "completions/mean_length": 65.0, "completions/min_length": 65.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.0, "completions/min_terminated_length": 65.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9949578642845154, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9949578642845154, "rewards/total_composite/std": 0.0, "reward": 0.9949578642845154, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.000207486460567452, "sampling/sampling_logp_difference/max": 0.02866499498486519, "sampling/importance_sampling_ratio/min": 0.9717419147491455, "sampling/importance_sampling_ratio/mean": 1.000050663948059, "sampling/importance_sampling_ratio/max": 1.0065354108810425, "entropy": 0.0024144309863913804, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9949578642845154, "reward_meter_mean": 0.9949578642845154, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9949578642845154, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1099 |
+
{"timestamp_utc": "2026-04-11T21:42:01Z", "mode": "train", "global_step": 1078, "epoch": 0.04162805066419524, "loss": 0.0161, "grad_norm": 4.65764856338501, "learning_rate": 6.7363636363636365e-06, "num_tokens": 2320094.0, "completions/mean_length": 61.375, "completions/min_length": 59.0, "completions/max_length": 64.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 61.375, "completions/min_terminated_length": 59.0, "completions/max_terminated_length": 64.0, "rewards/meter/mean": 0.9684683084487915, "rewards/meter/std": 0.08065449446439743, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9684683084487915, "rewards/total_composite/std": 0.08065449446439743, "reward": 0.9684683084487915, "reward_std": 0.08065447211265564, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.02014295756816864, "sampling/sampling_logp_difference/max": 1.577573299407959, "sampling/importance_sampling_ratio/min": 0.20647554099559784, "sampling/importance_sampling_ratio/mean": 1.0008001327514648, "sampling/importance_sampling_ratio/max": 1.5210373401641846, "entropy": 0.13324717595241964, "clip_ratio/low_mean": 0.0078125, "clip_ratio/low_min": 0.0078125, "clip_ratio/high_mean": 0.010216211201623082, "clip_ratio/high_max": 0.010216211201623082, "clip_ratio/region_mean": 0.018028711201623082, "reward_total_mean": 0.9684683084487915, "reward_meter_mean": 0.9684683084487915, "reward_meter_std": 0.08065449446439743, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9684683084487915, "reward_total_composite_std": 0.08065449446439743, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1100 |
+
{"timestamp_utc": "2026-04-11T21:42:06Z", "mode": "train", "global_step": 1079, "epoch": 0.041666666666666664, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.733333333333334e-06, "num_tokens": 2321786.0, "completions/mean_length": 49.5, "completions/min_length": 46.0, "completions/max_length": 61.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 49.5, "completions/min_terminated_length": 46.0, "completions/max_terminated_length": 61.0, "rewards/meter/mean": 0.091847725212574, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.091847725212574, "rewards/total_composite/std": 0.0, "reward": 0.091847725212574, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.01916210725903511, "sampling/sampling_logp_difference/max": 1.860931396484375, "sampling/importance_sampling_ratio/min": 0.15552771091461182, "sampling/importance_sampling_ratio/mean": 0.9989585280418396, "sampling/importance_sampling_ratio/max": 1.5695313215255737, "entropy": 0.054232243448495865, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.091847725212574, "reward_meter_mean": 0.091847725212574, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.091847725212574, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1101 |
+
{"timestamp_utc": "2026-04-11T21:42:11Z", "mode": "train", "global_step": 1080, "epoch": 0.04170528266913809, "loss": 0.0236, "grad_norm": 1.9305036067962646, "learning_rate": 6.73030303030303e-06, "num_tokens": 2324566.0, "completions/mean_length": 132.5, "completions/min_length": 129.0, "completions/max_length": 137.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 132.5, "completions/min_terminated_length": 129.0, "completions/max_terminated_length": 137.0, "rewards/meter/mean": 0.9948593974113464, "rewards/meter/std": 0.002004272071644664, "rewards/count_adherence/mean": 0.6666666865348816, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.663239598274231, "rewards/total_composite/std": 0.0013361814199015498, "reward": 0.663239598274231, "reward_std": 0.0013361814199015498, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.004363351967185736, "sampling/sampling_logp_difference/max": 0.48967456817626953, "sampling/importance_sampling_ratio/min": 0.6128258109092712, "sampling/importance_sampling_ratio/mean": 1.0016835927963257, "sampling/importance_sampling_ratio/max": 1.3798326253890991, "entropy": 0.029038164531812072, "clip_ratio/low_mean": 0.0027643622015602887, "clip_ratio/low_min": 0.0027643622015602887, "clip_ratio/high_mean": 0.0019379844889044762, "clip_ratio/high_max": 0.0019379844889044762, "clip_ratio/region_mean": 0.004702346690464765, "reward_total_mean": 0.663239598274231, "reward_meter_mean": 0.9948593974113464, "reward_meter_std": 0.002004272071644664, "reward_count_adherence_mean": 0.6666666865348816, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.663239598274231, "reward_total_composite_std": 0.0013361814199015498, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1102 |
+
{"timestamp_utc": "2026-04-11T21:42:17Z", "mode": "train", "global_step": 1081, "epoch": 0.04174389867160951, "loss": 0.0025, "grad_norm": 0.31140798330307007, "learning_rate": 6.7272727272727275e-06, "num_tokens": 2327382.0, "completions/mean_length": 175.0, "completions/min_length": 161.0, "completions/max_length": 178.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 175.0, "completions/min_terminated_length": 161.0, "completions/max_terminated_length": 178.0, "rewards/meter/mean": 0.9966549277305603, "rewards/meter/std": 7.61248666094616e-05, "rewards/count_adherence/mean": 0.75, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.747491180896759, "rewards/total_composite/std": 5.710634286515415e-05, "reward": 0.747491180896759, "reward_std": 5.710850018658675e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0060366359539330006, "sampling/sampling_logp_difference/max": 3.635336399078369, "sampling/importance_sampling_ratio/min": 0.0263750609010458, "sampling/importance_sampling_ratio/mean": 0.9993287920951843, "sampling/importance_sampling_ratio/max": 1.509435772895813, "entropy": 0.015698739560320973, "clip_ratio/low_mean": 0.001416441984474659, "clip_ratio/low_min": 0.001416441984474659, "clip_ratio/high_mean": 0.002259009750559926, "clip_ratio/high_max": 0.002259009750559926, "clip_ratio/region_mean": 0.003675451735034585, "reward_total_mean": 0.747491180896759, "reward_meter_mean": 0.9966549277305603, "reward_meter_std": 7.61248666094616e-05, "reward_count_adherence_mean": 0.75, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.747491180896759, "reward_total_composite_std": 5.710634286515415e-05, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1103 |
+
{"timestamp_utc": "2026-04-11T21:42:23Z", "mode": "train", "global_step": 1082, "epoch": 0.04178251467408094, "loss": 0.0114, "grad_norm": 4.485145092010498, "learning_rate": 6.724242424242424e-06, "num_tokens": 2329495.0, "completions/mean_length": 108.125, "completions/min_length": 106.0, "completions/max_length": 111.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 108.125, "completions/min_terminated_length": 106.0, "completions/max_terminated_length": 111.0, "rewards/meter/mean": 0.9212300777435303, "rewards/meter/std": 0.21582916378974915, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9212300777435303, "rewards/total_composite/std": 0.21582916378974915, "reward": 0.9212300777435303, "reward_std": 0.21582916378974915, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.013148986734449863, "sampling/sampling_logp_difference/max": 0.9551130533218384, "sampling/importance_sampling_ratio/min": 0.38476866483688354, "sampling/importance_sampling_ratio/mean": 1.0012389421463013, "sampling/importance_sampling_ratio/max": 1.5335668325424194, "entropy": 0.09740556543692946, "clip_ratio/low_mean": 0.0033783784601837397, "clip_ratio/low_min": 0.0033783784601837397, "clip_ratio/high_mean": 0.009324767161160707, "clip_ratio/high_max": 0.009324767161160707, "clip_ratio/region_mean": 0.012703145621344447, "reward_total_mean": 0.9212300777435303, "reward_meter_mean": 0.9212300777435303, "reward_meter_std": 0.21582916378974915, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9212300777435303, "reward_total_composite_std": 0.21582916378974915, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1104 |
+
{"timestamp_utc": "2026-04-11T21:42:28Z", "mode": "train", "global_step": 1083, "epoch": 0.04182113067655236, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.721212121212122e-06, "num_tokens": 2331783.0, "completions/mean_length": 129.0, "completions/min_length": 129.0, "completions/max_length": 129.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 129.0, "completions/min_terminated_length": 129.0, "completions/max_terminated_length": 129.0, "rewards/meter/mean": 0.9951313734054565, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.6666666865348816, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.6634209156036377, "rewards/total_composite/std": 0.0, "reward": 0.6634209156036377, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0002578256244305521, "sampling/sampling_logp_difference/max": 0.045966774225234985, "sampling/importance_sampling_ratio/min": 0.9550737738609314, "sampling/importance_sampling_ratio/mean": 1.0000473260879517, "sampling/importance_sampling_ratio/max": 1.0351933240890503, "entropy": 0.002667208347702399, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.6634209156036377, "reward_meter_mean": 0.9951313734054565, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.6666666865348816, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.6634209156036377, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1105 |
+
{"timestamp_utc": "2026-04-11T21:42:33Z", "mode": "train", "global_step": 1084, "epoch": 0.041859746679023785, "loss": 0.0462, "grad_norm": 17.376197814941406, "learning_rate": 6.718181818181819e-06, "num_tokens": 2333323.0, "completions/mean_length": 28.5, "completions/min_length": 27.0, "completions/max_length": 31.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 28.5, "completions/min_terminated_length": 27.0, "completions/max_terminated_length": 31.0, "rewards/meter/mean": 0.6803448796272278, "rewards/meter/std": 0.4452557861804962, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.6803448796272278, "rewards/total_composite/std": 0.4452557861804962, "reward": 0.6803448796272278, "reward_std": 0.4452557861804962, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.07098246365785599, "sampling/sampling_logp_difference/max": 2.6754026412963867, "sampling/importance_sampling_ratio/min": 0.06887909024953842, "sampling/importance_sampling_ratio/mean": 0.9909398555755615, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.25689972564578056, "clip_ratio/low_mean": 0.012652947567403316, "clip_ratio/low_min": 0.012652947567403316, "clip_ratio/high_mean": 0.04814814869314432, "clip_ratio/high_max": 0.04814814869314432, "clip_ratio/region_mean": 0.06080109626054764, "reward_total_mean": 0.6803448796272278, "reward_meter_mean": 0.6803448796272278, "reward_meter_std": 0.4452557861804962, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.6803448796272278, "reward_total_composite_std": 0.4452557861804962, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1106 |
+
{"timestamp_utc": "2026-04-11T21:42:40Z", "mode": "train", "global_step": 1085, "epoch": 0.04189836268149521, "loss": -0.0223, "grad_norm": 7.492648124694824, "learning_rate": 6.715151515151516e-06, "num_tokens": 2336298.0, "completions/mean_length": 189.875, "completions/min_length": 180.0, "completions/max_length": 201.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 189.875, "completions/min_terminated_length": 180.0, "completions/max_terminated_length": 201.0, "rewards/meter/mean": 0.46110403537750244, "rewards/meter/std": 0.25172099471092224, "rewards/count_adherence/mean": 0.7083333730697632, "rewards/count_adherence/std": 0.07715165615081787, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.3282319903373718, "rewards/total_composite/std": 0.1757289618253708, "reward": 0.3282319903373718, "reward_std": 0.1757289618253708, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.010337450541555882, "sampling/sampling_logp_difference/max": 1.9266459941864014, "sampling/importance_sampling_ratio/min": 0.14563584327697754, "sampling/importance_sampling_ratio/mean": 1.0021684169769287, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.034014943055808544, "clip_ratio/low_mean": 0.002667596214450896, "clip_ratio/low_min": 0.002667596214450896, "clip_ratio/high_mean": 0.0026115492219105363, "clip_ratio/high_max": 0.0026115492219105363, "clip_ratio/region_mean": 0.005279145436361432, "reward_total_mean": 0.3282319903373718, "reward_meter_mean": 0.46110403537750244, "reward_meter_std": 0.25172099471092224, "reward_count_adherence_mean": 0.7083333730697632, "reward_count_adherence_std": 0.07715165615081787, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.3282319903373718, "reward_total_composite_std": 0.1757289618253708, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1107 |
+
{"timestamp_utc": "2026-04-11T21:42:46Z", "mode": "train", "global_step": 1086, "epoch": 0.04193697868396663, "loss": 0.0308, "grad_norm": 3.549293279647827, "learning_rate": 6.712121212121213e-06, "num_tokens": 2339268.0, "completions/mean_length": 208.25, "completions/min_length": 197.0, "completions/max_length": 213.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 208.25, "completions/min_terminated_length": 197.0, "completions/max_terminated_length": 213.0, "rewards/meter/mean": 0.9969711303710938, "rewards/meter/std": 8.204347250284627e-05, "rewards/count_adherence/mean": 0.6500000357627869, "rewards/count_adherence/std": 0.09258200973272324, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.6480287909507751, "rewards/total_composite/std": 0.0922786295413971, "reward": 0.6480287909507751, "reward_std": 0.09227863699197769, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.010134638287127018, "sampling/sampling_logp_difference/max": 8.085393905639648, "sampling/importance_sampling_ratio/min": 0.00030800519743934274, "sampling/importance_sampling_ratio/mean": 1.0001766681671143, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.008055432001128793, "clip_ratio/low_mean": 0.002358543104492128, "clip_ratio/low_min": 0.002358543104492128, "clip_ratio/high_mean": 0.0006345177534967661, "clip_ratio/high_max": 0.0006345177534967661, "clip_ratio/region_mean": 0.002993060857988894, "reward_total_mean": 0.6480287909507751, "reward_meter_mean": 0.9969711303710938, "reward_meter_std": 8.204347250284627e-05, "reward_count_adherence_mean": 0.6500000357627869, "reward_count_adherence_std": 0.09258200973272324, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.6480287909507751, "reward_total_composite_std": 0.0922786295413971, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1108 |
+
{"timestamp_utc": "2026-04-11T21:42:51Z", "mode": "train", "global_step": 1087, "epoch": 0.04197559468643806, "loss": -0.0111, "grad_norm": 2.269176959991455, "learning_rate": 6.709090909090909e-06, "num_tokens": 2341095.0, "completions/mean_length": 80.375, "completions/min_length": 76.0, "completions/max_length": 81.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 80.375, "completions/min_terminated_length": 76.0, "completions/max_terminated_length": 81.0, "rewards/meter/mean": 0.9966411590576172, "rewards/meter/std": 0.00015710237494204193, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9966411590576172, "rewards/total_composite/std": 0.00015710237494204193, "reward": 0.9966411590576172, "reward_std": 0.00015711141168139875, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0034698506351560354, "sampling/sampling_logp_difference/max": 0.793515682220459, "sampling/importance_sampling_ratio/min": 0.45225203037261963, "sampling/importance_sampling_ratio/mean": 0.9995546340942383, "sampling/importance_sampling_ratio/max": 1.1145241260528564, "entropy": 0.014752325252629817, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0015432098880410194, "clip_ratio/high_max": 0.0015432098880410194, "clip_ratio/region_mean": 0.0015432098880410194, "reward_total_mean": 0.9966411590576172, "reward_meter_mean": 0.9966411590576172, "reward_meter_std": 0.00015710237494204193, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9966411590576172, "reward_total_composite_std": 0.00015710237494204193, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1109 |
+
{"timestamp_utc": "2026-04-11T21:42:57Z", "mode": "train", "global_step": 1088, "epoch": 0.04201421068890948, "loss": 0.0123, "grad_norm": 12.239304542541504, "learning_rate": 6.706060606060607e-06, "num_tokens": 2342713.0, "completions/mean_length": 49.25, "completions/min_length": 49.0, "completions/max_length": 51.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 49.25, "completions/min_terminated_length": 49.0, "completions/max_terminated_length": 51.0, "rewards/meter/mean": 0.986293613910675, "rewards/meter/std": 0.0026903701946139336, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.986293613910675, "rewards/total_composite/std": 0.0026903701946139336, "reward": 0.986293613910675, "reward_std": 0.0026903690304607153, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.016510505229234695, "sampling/sampling_logp_difference/max": 1.4381656646728516, "sampling/importance_sampling_ratio/min": 0.23736277222633362, "sampling/importance_sampling_ratio/mean": 1.0027153491973877, "sampling/importance_sampling_ratio/max": 1.9696929454803467, "entropy": 0.09569864999502897, "clip_ratio/low_mean": 0.007452981313690543, "clip_ratio/low_min": 0.007452981313690543, "clip_ratio/high_mean": 0.007653061067685485, "clip_ratio/high_max": 0.007653061067685485, "clip_ratio/region_mean": 0.015106042381376028, "reward_total_mean": 0.986293613910675, "reward_meter_mean": 0.986293613910675, "reward_meter_std": 0.0026903701946139336, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.986293613910675, "reward_total_composite_std": 0.0026903701946139336, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1110 |
+
{"timestamp_utc": "2026-04-11T21:43:03Z", "mode": "train", "global_step": 1089, "epoch": 0.042052826691380905, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.703030303030304e-06, "num_tokens": 2345457.0, "completions/mean_length": 152.0, "completions/min_length": 152.0, "completions/max_length": 152.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 152.0, "completions/min_terminated_length": 152.0, "completions/max_terminated_length": 152.0, "rewards/meter/mean": 0.9968674182891846, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.75, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7476505637168884, "rewards/total_composite/std": 0.0, "reward": 0.7476505637168884, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0006623952649533749, "sampling/sampling_logp_difference/max": 0.3456934988498688, "sampling/importance_sampling_ratio/min": 0.7077293992042542, "sampling/importance_sampling_ratio/mean": 0.9997900128364563, "sampling/importance_sampling_ratio/max": 1.08553946018219, "entropy": 0.0034695666399784386, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.7476505637168884, "reward_meter_mean": 0.9968674182891846, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.75, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7476505637168884, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1111 |
+
{"timestamp_utc": "2026-04-11T21:43:09Z", "mode": "train", "global_step": 1090, "epoch": 0.04209144269385233, "loss": -0.0003, "grad_norm": 0.4215506911277771, "learning_rate": 6.700000000000001e-06, "num_tokens": 2348309.0, "completions/mean_length": 182.5, "completions/min_length": 181.0, "completions/max_length": 183.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 182.5, "completions/min_terminated_length": 181.0, "completions/max_terminated_length": 183.0, "rewards/meter/mean": 0.9972051978111267, "rewards/meter/std": 5.313828296493739e-05, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9972051978111267, "rewards/total_composite/std": 5.313828296493739e-05, "reward": 0.9972051978111267, "reward_std": 5.313552901498042e-05, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.004990379326045513, "sampling/sampling_logp_difference/max": 0.824894905090332, "sampling/importance_sampling_ratio/min": 0.4382810592651367, "sampling/importance_sampling_ratio/mean": 1.0013642311096191, "sampling/importance_sampling_ratio/max": 1.4445006847381592, "entropy": 0.04286058805882931, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.004788926860783249, "clip_ratio/high_max": 0.004788926860783249, "clip_ratio/region_mean": 0.004788926860783249, "reward_total_mean": 0.9972051978111267, "reward_meter_mean": 0.9972051978111267, "reward_meter_std": 5.313828296493739e-05, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9972051978111267, "reward_total_composite_std": 5.313828296493739e-05, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1112 |
+
{"timestamp_utc": "2026-04-11T21:43:14Z", "mode": "train", "global_step": 1091, "epoch": 0.042130058696323754, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.6969696969696975e-06, "num_tokens": 2349895.0, "completions/mean_length": 47.25, "completions/min_length": 46.0, "completions/max_length": 56.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 47.25, "completions/min_terminated_length": 46.0, "completions/max_terminated_length": 56.0, "rewards/meter/mean": 0.091847725212574, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.091847725212574, "rewards/total_composite/std": 0.0, "reward": 0.091847725212574, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.007502844091504812, "sampling/sampling_logp_difference/max": 0.6710173487663269, "sampling/importance_sampling_ratio/min": 0.511188268661499, "sampling/importance_sampling_ratio/mean": 1.001801609992981, "sampling/importance_sampling_ratio/max": 1.5886342525482178, "entropy": 0.016990114585496485, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.091847725212574, "reward_meter_mean": 0.091847725212574, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.091847725212574, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1113 |
+
{"timestamp_utc": "2026-04-11T21:43:18Z", "mode": "train", "global_step": 1092, "epoch": 0.04216867469879518, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.693939393939395e-06, "num_tokens": 2351375.0, "completions/mean_length": 32.0, "completions/min_length": 32.0, "completions/max_length": 32.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 32.0, "completions/min_terminated_length": 32.0, "completions/max_terminated_length": 32.0, "rewards/meter/mean": 0.9955702424049377, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9955702424049377, "rewards/total_composite/std": 0.0, "reward": 0.9955702424049377, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.001951782964169979, "sampling/sampling_logp_difference/max": 0.17160826921463013, "sampling/importance_sampling_ratio/min": 0.8423091173171997, "sampling/importance_sampling_ratio/mean": 0.9996309876441956, "sampling/importance_sampling_ratio/max": 1.054145336151123, "entropy": 0.010097296675667167, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9955702424049377, "reward_meter_mean": 0.9955702424049377, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9955702424049377, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1114 |
+
{"timestamp_utc": "2026-04-11T21:43:23Z", "mode": "train", "global_step": 1093, "epoch": 0.0422072907012666, "loss": 0.0095, "grad_norm": 1.6124061346054077, "learning_rate": 6.690909090909091e-06, "num_tokens": 2353053.0, "completions/mean_length": 64.75, "completions/min_length": 63.0, "completions/max_length": 65.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 64.75, "completions/min_terminated_length": 63.0, "completions/max_terminated_length": 65.0, "rewards/meter/mean": 0.9950487613677979, "rewards/meter/std": 0.0002570115029811859, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9950487613677979, "rewards/total_composite/std": 0.0002570115029811859, "reward": 0.9950487613677979, "reward_std": 0.00025701147387735546, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.005157488863915205, "sampling/sampling_logp_difference/max": 1.5352113246917725, "sampling/importance_sampling_ratio/min": 0.21541017293930054, "sampling/importance_sampling_ratio/mean": 0.9983689785003662, "sampling/importance_sampling_ratio/max": 1.0859289169311523, "entropy": 0.013940075354184955, "clip_ratio/low_mean": 0.003846153849735856, "clip_ratio/low_min": 0.003846153849735856, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.003846153849735856, "reward_total_mean": 0.9950487613677979, "reward_meter_mean": 0.9950487613677979, "reward_meter_std": 0.0002570115029811859, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9950487613677979, "reward_total_composite_std": 0.0002570115029811859, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1115 |
+
{"timestamp_utc": "2026-04-11T21:43:28Z", "mode": "train", "global_step": 1094, "epoch": 0.042245906703738026, "loss": 0.0094, "grad_norm": 3.3892717361450195, "learning_rate": 6.687878787878788e-06, "num_tokens": 2354894.0, "completions/mean_length": 65.125, "completions/min_length": 64.0, "completions/max_length": 66.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 65.125, "completions/min_terminated_length": 64.0, "completions/max_terminated_length": 66.0, "rewards/meter/mean": 0.9951438903808594, "rewards/meter/std": 0.0001731093943817541, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9951438903808594, "rewards/total_composite/std": 0.0001731093943817541, "reward": 0.9951438903808594, "reward_std": 0.00017311105330009013, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.011436189524829388, "sampling/sampling_logp_difference/max": 1.639451026916504, "sampling/importance_sampling_ratio/min": 0.19408656656742096, "sampling/importance_sampling_ratio/mean": 0.996876060962677, "sampling/importance_sampling_ratio/max": 1.232495665550232, "entropy": 0.04774620302487165, "clip_ratio/low_mean": 0.0038470644503831863, "clip_ratio/low_min": 0.0038470644503831863, "clip_ratio/high_mean": 0.00390625, "clip_ratio/high_max": 0.00390625, "clip_ratio/region_mean": 0.007753314450383186, "reward_total_mean": 0.9951438903808594, "reward_meter_mean": 0.9951438903808594, "reward_meter_std": 0.0001731093943817541, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9951438903808594, "reward_total_composite_std": 0.0001731093943817541, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1116 |
+
{"timestamp_utc": "2026-04-11T21:43:33Z", "mode": "train", "global_step": 1095, "epoch": 0.04228452270620945, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.684848484848485e-06, "num_tokens": 2357158.0, "completions/mean_length": 107.0, "completions/min_length": 107.0, "completions/max_length": 107.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 107.0, "completions/min_terminated_length": 107.0, "completions/max_terminated_length": 107.0, "rewards/meter/mean": 0.9966511130332947, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9966511130332947, "rewards/total_composite/std": 0.0, "reward": 0.9966511130332947, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.0007139050285331905, "sampling/sampling_logp_difference/max": 0.2203187644481659, "sampling/importance_sampling_ratio/min": 0.8022630214691162, "sampling/importance_sampling_ratio/mean": 0.999952495098114, "sampling/importance_sampling_ratio/max": 1.083788275718689, "entropy": 0.003691258083563298, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.9966511130332947, "reward_meter_mean": 0.9966511130332947, "reward_meter_std": 0.0, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9966511130332947, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1117 |
+
{"timestamp_utc": "2026-04-11T21:43:38Z", "mode": "train", "global_step": 1096, "epoch": 0.042323138708680874, "loss": 0.0528, "grad_norm": 17.096412658691406, "learning_rate": 6.681818181818183e-06, "num_tokens": 2359061.0, "completions/mean_length": 26.875, "completions/min_length": 20.0, "completions/max_length": 29.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 26.875, "completions/min_terminated_length": 20.0, "completions/max_terminated_length": 29.0, "rewards/meter/mean": 0.11241252720355988, "rewards/meter/std": 0.3034849762916565, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.11241252720355988, "rewards/total_composite/std": 0.3034849762916565, "reward": 0.11241252720355988, "reward_std": 0.3034849464893341, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.10703544318675995, "sampling/sampling_logp_difference/max": 4.594521999359131, "sampling/importance_sampling_ratio/min": 0.010107050649821758, "sampling/importance_sampling_ratio/mean": 0.979833722114563, "sampling/importance_sampling_ratio/max": 1.9887901544570923, "entropy": 0.21645524073392153, "clip_ratio/low_mean": 0.03616698645055294, "clip_ratio/low_min": 0.03616698645055294, "clip_ratio/high_mean": 0.0223214291036129, "clip_ratio/high_max": 0.0223214291036129, "clip_ratio/region_mean": 0.05848841555416584, "reward_total_mean": 0.11241252720355988, "reward_meter_mean": 0.11241252720355988, "reward_meter_std": 0.3034849762916565, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.11241252720355988, "reward_total_composite_std": 0.3034849762916565, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1118 |
+
{"timestamp_utc": "2026-04-11T21:43:44Z", "mode": "train", "global_step": 1097, "epoch": 0.0423617547111523, "loss": 0.0354, "grad_norm": 0.9273874163627625, "learning_rate": 6.678787878787879e-06, "num_tokens": 2361662.0, "completions/mean_length": 147.125, "completions/min_length": 145.0, "completions/max_length": 161.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 147.125, "completions/min_terminated_length": 145.0, "completions/max_terminated_length": 161.0, "rewards/meter/mean": 0.9964081048965454, "rewards/meter/std": 0.00033037204411812127, "rewards/count_adherence/mean": 0.96875, "rewards/count_adherence/std": 0.0883883461356163, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.9652670621871948, "rewards/total_composite/std": 0.08803740888834, "reward": 0.9652670621871948, "reward_std": 0.08803740888834, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.0055220527574419975, "sampling/sampling_logp_difference/max": 1.9974420070648193, "sampling/importance_sampling_ratio/min": 0.4937727451324463, "sampling/importance_sampling_ratio/mean": 1.0020689964294434, "sampling/importance_sampling_ratio/max": 2.0, "entropy": 0.016805213410407305, "clip_ratio/low_mean": 0.000776397529989481, "clip_ratio/low_min": 0.000776397529989481, "clip_ratio/high_mean": 0.0025862068869173527, "clip_ratio/high_max": 0.0025862068869173527, "clip_ratio/region_mean": 0.0033626044169068336, "reward_total_mean": 0.9652670621871948, "reward_meter_mean": 0.9964081048965454, "reward_meter_std": 0.00033037204411812127, "reward_count_adherence_mean": 0.96875, "reward_count_adherence_std": 0.0883883461356163, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.9652670621871948, "reward_total_composite_std": 0.08803740888834, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1119 |
+
{"timestamp_utc": "2026-04-11T21:43:50Z", "mode": "train", "global_step": 1098, "epoch": 0.04240037071362372, "loss": 0.0, "grad_norm": 0.0, "learning_rate": 6.6757575757575766e-06, "num_tokens": 2364710.0, "completions/mean_length": 167.0, "completions/min_length": 167.0, "completions/max_length": 167.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 167.0, "completions/min_terminated_length": 167.0, "completions/max_terminated_length": 167.0, "rewards/meter/mean": 0.9968674182891846, "rewards/meter/std": 0.0, "rewards/count_adherence/mean": 0.8333333134651184, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.8307228684425354, "rewards/total_composite/std": 0.0, "reward": 0.8307228684425354, "reward_std": 0.0, "frac_reward_zero_std": 1.0, "sampling/sampling_logp_difference/mean": 0.00022596628696192056, "sampling/sampling_logp_difference/max": 0.1071237251162529, "sampling/importance_sampling_ratio/min": 0.8984145522117615, "sampling/importance_sampling_ratio/mean": 1.0000667572021484, "sampling/importance_sampling_ratio/max": 1.0559391975402832, "entropy": 0.0018285177211510018, "clip_ratio/low_mean": 0.0, "clip_ratio/low_min": 0.0, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.0, "reward_total_mean": 0.8307228684425354, "reward_meter_mean": 0.9968674182891846, "reward_meter_std": 0.0, "reward_count_adherence_mean": 0.8333333134651184, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.8307228684425354, "reward_total_composite_std": 0.0, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1120 |
+
{"timestamp_utc": "2026-04-11T21:43:55Z", "mode": "train", "global_step": 1099, "epoch": 0.042438986716095146, "loss": -0.0331, "grad_norm": 6.627462863922119, "learning_rate": 6.672727272727273e-06, "num_tokens": 2366553.0, "completions/mean_length": 48.375, "completions/min_length": 46.0, "completions/max_length": 57.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 48.375, "completions/min_terminated_length": 46.0, "completions/max_terminated_length": 57.0, "rewards/meter/mean": 0.18795685470104218, "rewards/meter/std": 0.3101930618286133, "rewards/count_adherence/mean": 1.0, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.18795685470104218, "rewards/total_composite/std": 0.3101930618286133, "reward": 0.18795685470104218, "reward_std": 0.3101930618286133, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.010162265971302986, "sampling/sampling_logp_difference/max": 0.9485464096069336, "sampling/importance_sampling_ratio/min": 0.3873036205768585, "sampling/importance_sampling_ratio/mean": 1.0062874555587769, "sampling/importance_sampling_ratio/max": 1.6379154920578003, "entropy": 0.030019621714018285, "clip_ratio/low_mean": 0.00657894741743803, "clip_ratio/low_min": 0.00657894741743803, "clip_ratio/high_mean": 0.0, "clip_ratio/high_max": 0.0, "clip_ratio/region_mean": 0.00657894741743803, "reward_total_mean": 0.18795685470104218, "reward_meter_mean": 0.18795685470104218, "reward_meter_std": 0.3101930618286133, "reward_count_adherence_mean": 1.0, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.18795685470104218, "reward_total_composite_std": 0.3101930618286133, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
| 1121 |
+
{"timestamp_utc": "2026-04-11T21:44:01Z", "mode": "train", "global_step": 1100, "epoch": 0.04247760271856657, "loss": -0.0306, "grad_norm": 2.0620083808898926, "learning_rate": 6.66969696969697e-06, "num_tokens": 2369329.0, "completions/mean_length": 156.0, "completions/min_length": 145.0, "completions/max_length": 163.0, "completions/clipped_ratio": 0.0, "completions/mean_terminated_length": 156.0, "completions/min_terminated_length": 145.0, "completions/max_terminated_length": 163.0, "rewards/meter/mean": 0.9976032972335815, "rewards/meter/std": 0.0007398881716653705, "rewards/count_adherence/mean": 0.800000011920929, "rewards/count_adherence/std": 0.0, "rewards/arabic_clean/mean": 1.0, "rewards/arabic_clean/std": 0.0, "rewards/total_composite/mean": 0.7980826497077942, "rewards/total_composite/std": 0.0005919100949540734, "reward": 0.7980826497077942, "reward_std": 0.0005919178947806358, "frac_reward_zero_std": 0.0, "sampling/sampling_logp_difference/mean": 0.003971020225435495, "sampling/sampling_logp_difference/max": 0.6630861759185791, "sampling/importance_sampling_ratio/min": 0.5152587294578552, "sampling/importance_sampling_ratio/mean": 1.0009613037109375, "sampling/importance_sampling_ratio/max": 1.730522871017456, "entropy": 0.030738614965230227, "clip_ratio/low_mean": 0.0016737572732381523, "clip_ratio/low_min": 0.0016737572732381523, "clip_ratio/high_mean": 0.0015337422955781221, "clip_ratio/high_max": 0.0015337422955781221, "clip_ratio/region_mean": 0.0032074995688162744, "reward_total_mean": 0.7980826497077942, "reward_meter_mean": 0.9976032972335815, "reward_meter_std": 0.0007398881716653705, "reward_count_adherence_mean": 0.800000011920929, "reward_count_adherence_std": 0.0, "reward_arabic_clean_mean": 1.0, "reward_arabic_clean_std": 0.0, "reward_total_composite_mean": 0.7980826497077942, "reward_total_composite_std": 0.0005919100949540734, "run_id": "shaer_grpo_20260411_192107", "run_sequence_index": 0}
|
plots/meter_by_meter_chain.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plots/meter_by_meter_run.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plots/reward_panels_eval_chain.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plots/reward_panels_eval_run.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plots/reward_panels_train_chain.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plots/reward_panels_train_run.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
plotter.log
CHANGED
|
@@ -1838,3 +1838,103 @@
|
|
| 1838 |
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1839 |
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1840 |
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1838 |
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1839 |
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1840 |
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1841 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1842 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1843 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1844 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1845 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1846 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1847 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1848 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1849 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1850 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1851 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1852 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1853 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1854 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1855 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1856 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1857 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1858 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1859 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1860 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1861 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1862 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1863 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1864 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1865 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1866 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1867 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1868 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1869 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1870 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1871 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1872 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1873 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1874 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1875 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1876 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1877 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1878 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1879 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1880 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1881 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1882 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1883 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1884 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1885 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1886 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1887 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1888 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1889 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1890 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1891 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1892 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1893 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1894 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1895 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1896 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1897 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1898 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1899 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1900 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1901 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1902 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1903 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1904 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1905 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1906 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1907 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1908 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1909 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1910 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1911 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1912 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1913 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1914 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1915 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1916 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1917 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1918 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1919 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1920 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1921 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1922 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1923 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1924 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1925 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1926 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1927 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1928 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1929 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1930 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
| 1931 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_run.png
|
| 1932 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_run.png
|
| 1933 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_train_chain.png
|
| 1934 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/reward_panels_eval_chain.png
|
| 1935 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_run.png
|
| 1936 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/kl_chain.png
|
| 1937 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_run.png
|
| 1938 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/arabic_gate_chain.png
|
| 1939 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_run.png
|
| 1940 |
+
[plot_live_rewards] updated /root/workspace/Shaer/grpo/outputs/train/shaer_grpo_20260411_192107/plots/meter_by_meter_chain.png
|
reward_arabic_clean_debug.jsonl
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:02af34c6e81bb89eac239e34d318715f6dca3caa59f1e392c023301adf78cdc8
|
| 3 |
+
size 41327484
|
reward_count_adherence_debug.jsonl
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a95c6588bf82fe0a84b041ef20736a661953a04ed55e8f5ee1e498f5d41651a7
|
| 3 |
+
size 41350701
|
reward_meter_debug.jsonl
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d250912447937e22cefa5fe86824f9ec4bdece6968c35f198580a297f7e39d86
|
| 3 |
+
size 41495722
|
reward_total_composite_debug.jsonl
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:40a8e9ada3734cad5ece1b5e0b82062bc03b38a6d5e5ab692d120b25faaf1e80
|
| 3 |
+
size 41489394
|
train.log
CHANGED
|
@@ -1081,3 +1081,54 @@
|
|
| 1081 |
2026-04-11 21:38:09,719 | INFO | train_grpo_train | metrics_logged mode=train step=1050
|
| 1082 |
2026-04-11 21:39:35,179 | INFO | train_grpo_train | metrics_logged mode=eval step=1050
|
| 1083 |
2026-04-11 21:39:43,421 | INFO | train_grpo_train | metrics_logged mode=train step=1051
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1081 |
2026-04-11 21:38:09,719 | INFO | train_grpo_train | metrics_logged mode=train step=1050
|
| 1082 |
2026-04-11 21:39:35,179 | INFO | train_grpo_train | metrics_logged mode=eval step=1050
|
| 1083 |
2026-04-11 21:39:43,421 | INFO | train_grpo_train | metrics_logged mode=train step=1051
|
| 1084 |
+
2026-04-11 21:39:48,429 | INFO | train_grpo_train | metrics_logged mode=train step=1052
|
| 1085 |
+
2026-04-11 21:39:53,307 | INFO | train_grpo_train | metrics_logged mode=train step=1053
|
| 1086 |
+
2026-04-11 21:39:57,900 | INFO | train_grpo_train | metrics_logged mode=train step=1054
|
| 1087 |
+
2026-04-11 21:40:02,660 | INFO | train_grpo_train | metrics_logged mode=train step=1055
|
| 1088 |
+
2026-04-11 21:40:07,434 | INFO | train_grpo_train | metrics_logged mode=train step=1056
|
| 1089 |
+
2026-04-11 21:40:12,681 | INFO | train_grpo_train | metrics_logged mode=train step=1057
|
| 1090 |
+
2026-04-11 21:40:19,221 | INFO | train_grpo_train | metrics_logged mode=train step=1058
|
| 1091 |
+
2026-04-11 21:40:24,013 | INFO | train_grpo_train | metrics_logged mode=train step=1059
|
| 1092 |
+
2026-04-11 21:40:31,596 | INFO | train_grpo_train | metrics_logged mode=train step=1060
|
| 1093 |
+
2026-04-11 21:40:38,880 | INFO | train_grpo_train | metrics_logged mode=train step=1061
|
| 1094 |
+
2026-04-11 21:40:44,089 | INFO | train_grpo_train | metrics_logged mode=train step=1062
|
| 1095 |
+
2026-04-11 21:40:48,731 | INFO | train_grpo_train | metrics_logged mode=train step=1063
|
| 1096 |
+
2026-04-11 21:40:53,487 | INFO | train_grpo_train | metrics_logged mode=train step=1064
|
| 1097 |
+
2026-04-11 21:40:58,321 | INFO | train_grpo_train | metrics_logged mode=train step=1065
|
| 1098 |
+
2026-04-11 21:41:02,948 | INFO | train_grpo_train | metrics_logged mode=train step=1066
|
| 1099 |
+
2026-04-11 21:41:07,687 | INFO | train_grpo_train | metrics_logged mode=train step=1067
|
| 1100 |
+
2026-04-11 21:41:12,455 | INFO | train_grpo_train | metrics_logged mode=train step=1068
|
| 1101 |
+
2026-04-11 21:41:17,176 | INFO | train_grpo_train | metrics_logged mode=train step=1069
|
| 1102 |
+
2026-04-11 21:41:21,880 | INFO | train_grpo_train | metrics_logged mode=train step=1070
|
| 1103 |
+
2026-04-11 21:41:26,792 | INFO | train_grpo_train | metrics_logged mode=train step=1071
|
| 1104 |
+
2026-04-11 21:41:31,842 | INFO | train_grpo_train | metrics_logged mode=train step=1072
|
| 1105 |
+
2026-04-11 21:41:36,480 | INFO | train_grpo_train | metrics_logged mode=train step=1073
|
| 1106 |
+
2026-04-11 21:41:41,159 | INFO | train_grpo_train | metrics_logged mode=train step=1074
|
| 1107 |
+
2026-04-11 21:41:46,095 | INFO | train_grpo_train | metrics_logged mode=train step=1075
|
| 1108 |
+
2026-04-11 21:41:51,574 | INFO | train_grpo_train | metrics_logged mode=train step=1076
|
| 1109 |
+
2026-04-11 21:41:56,475 | INFO | train_grpo_train | metrics_logged mode=train step=1077
|
| 1110 |
+
2026-04-11 21:42:01,332 | INFO | train_grpo_train | metrics_logged mode=train step=1078
|
| 1111 |
+
2026-04-11 21:42:06,269 | INFO | train_grpo_train | metrics_logged mode=train step=1079
|
| 1112 |
+
2026-04-11 21:42:11,893 | INFO | train_grpo_train | metrics_logged mode=train step=1080
|
| 1113 |
+
2026-04-11 21:42:18,036 | INFO | train_grpo_train | metrics_logged mode=train step=1081
|
| 1114 |
+
2026-04-11 21:42:23,512 | INFO | train_grpo_train | metrics_logged mode=train step=1082
|
| 1115 |
+
2026-04-11 21:42:28,983 | INFO | train_grpo_train | metrics_logged mode=train step=1083
|
| 1116 |
+
2026-04-11 21:42:33,670 | INFO | train_grpo_train | metrics_logged mode=train step=1084
|
| 1117 |
+
2026-04-11 21:42:40,216 | INFO | train_grpo_train | metrics_logged mode=train step=1085
|
| 1118 |
+
2026-04-11 21:42:46,882 | INFO | train_grpo_train | metrics_logged mode=train step=1086
|
| 1119 |
+
2026-04-11 21:42:51,884 | INFO | train_grpo_train | metrics_logged mode=train step=1087
|
| 1120 |
+
2026-04-11 21:42:57,309 | INFO | train_grpo_train | metrics_logged mode=train step=1088
|
| 1121 |
+
2026-04-11 21:43:03,412 | INFO | train_grpo_train | metrics_logged mode=train step=1089
|
| 1122 |
+
2026-04-11 21:43:09,675 | INFO | train_grpo_train | metrics_logged mode=train step=1090
|
| 1123 |
+
2026-04-11 21:43:14,464 | INFO | train_grpo_train | metrics_logged mode=train step=1091
|
| 1124 |
+
2026-04-11 21:43:18,878 | INFO | train_grpo_train | metrics_logged mode=train step=1092
|
| 1125 |
+
2026-04-11 21:43:23,579 | INFO | train_grpo_train | metrics_logged mode=train step=1093
|
| 1126 |
+
2026-04-11 21:43:28,625 | INFO | train_grpo_train | metrics_logged mode=train step=1094
|
| 1127 |
+
2026-04-11 21:43:34,014 | INFO | train_grpo_train | metrics_logged mode=train step=1095
|
| 1128 |
+
2026-04-11 21:43:38,463 | INFO | train_grpo_train | metrics_logged mode=train step=1096
|
| 1129 |
+
2026-04-11 21:43:44,515 | INFO | train_grpo_train | metrics_logged mode=train step=1097
|
| 1130 |
+
2026-04-11 21:43:50,582 | INFO | train_grpo_train | metrics_logged mode=train step=1098
|
| 1131 |
+
2026-04-11 21:43:55,478 | INFO | train_grpo_train | metrics_logged mode=train step=1099
|
| 1132 |
+
2026-04-11 21:44:01,503 | INFO | train_grpo_train | metrics_logged mode=train step=1100
|
| 1133 |
+
2026-04-11 21:45:15,484 | INFO | train_grpo_train | metrics_logged mode=eval step=1100
|
| 1134 |
+
2026-04-11 21:45:24,605 | INFO | train_grpo_train | metrics_logged mode=train step=1101
|
train_stdout.log
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|