iris-olmo-2-1b-multihop-only-k10

This model is a fine-tuned version of allenai/OLMo-2-0425-1B-Instruct on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 1.3327
  • Model Preparation Time: 0.0118
  • Soft Mae: 0.1216
  • Soft Brier: 0.0397
  • Student Prelevantmean: 0.2704
  • Teacher Prelevantmean: 0.3219
  • Bin F1: 0.7910
  • Cal Ece: 0.1288
  • Cal Brier: 0.1181
  • Cal Auroc: 0.9502
  • Info Ndcg@p8: 0.9732
  • Info Pairwiseacc: 0.7755
  • Num Questions: 20
  • Num Pairs: 253

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • gradient_accumulation_steps: 16
  • total_train_batch_size: 16
  • optimizer: Use paged_adamw_8bit with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: constant
  • lr_scheduler_warmup_steps: 0.03
  • num_epochs: 6.0

Training results

Training Loss Epoch Step Validation Loss Model Preparation Time Soft Mae Soft Brier Student Prelevantmean Teacher Prelevantmean Bin F1 Cal Ece Cal Brier Cal Auroc Info Ndcg@p8 Info Pairwiseacc Num Questions Num Pairs
1.4630 0.2562 38 1.6236 0.0118 0.2745 0.1130 0.2112 0.3219 0.0 0.1880 0.2632 0.6501 0.7231 0.5022 20 253
1.4647 0.5124 76 1.4833 0.0118 0.2780 0.0988 0.2630 0.3219 0.0196 0.1561 0.2406 0.7067 0.7529 0.5412 20 253
1.3540 0.7686 114 1.3927 0.0118 0.2538 0.0911 0.2544 0.3219 0.0388 0.1448 0.2283 0.7269 0.7593 0.5666 20 253
1.2624 1.0202 152 1.3260 0.0118 0.2510 0.0822 0.3418 0.3219 0.3636 0.0840 0.2040 0.7505 0.8272 0.5878 20 253
1.0723 1.2764 190 1.2256 0.0118 0.2164 0.0723 0.2967 0.3219 0.4173 0.1377 0.1898 0.8163 0.9125 0.6550 20 253
1.0972 1.5327 228 1.1845 0.0118 0.1917 0.0601 0.2804 0.3219 0.4179 0.1355 0.1686 0.8797 0.9242 0.6985 20 253
0.9822 1.7889 266 1.0331 0.0118 0.1486 0.0422 0.2912 0.3219 0.7101 0.1324 0.1347 0.9256 0.9410 0.7235 20 253
0.8903 2.0405 304 1.1172 0.0118 0.1448 0.0459 0.2408 0.3219 0.6538 0.1617 0.1415 0.9456 0.9402 0.7424 20 253
0.7215 2.2967 342 1.0172 0.0118 0.1394 0.0418 0.2594 0.3219 0.6667 0.1463 0.1355 0.9423 0.9605 0.7357 20 253
0.8471 2.5529 380 1.0209 0.0118 0.1308 0.0369 0.3213 0.3219 0.7692 0.1106 0.1168 0.9377 0.9707 0.7518 20 253
0.6506 2.8091 418 1.0449 0.0118 0.1329 0.0376 0.3641 0.3219 0.8367 0.1072 0.1015 0.9474 0.9704 0.7633 20 253
0.5661 3.0607 456 1.1329 0.0118 0.1380 0.0453 0.2453 0.3219 0.6835 0.1539 0.1378 0.9428 0.9620 0.7409 20 253
0.5664 3.3169 494 1.2448 0.0118 0.1287 0.0446 0.2434 0.3219 0.7284 0.1559 0.1307 0.9502 0.9719 0.7567 20 253
0.5569 3.5731 532 1.1311 0.0118 0.1285 0.0447 0.2843 0.3219 0.7543 0.1149 0.1194 0.9404 0.9747 0.7591 20 253
0.6041 3.8293 570 1.1799 0.0118 0.1269 0.0427 0.2981 0.3219 0.7892 0.1011 0.1147 0.9416 0.9740 0.7701 20 253
0.4112 4.0809 608 1.0701 0.0118 0.1191 0.0356 0.2695 0.3219 0.7195 0.1324 0.1192 0.9550 0.9780 0.7744 20 253
0.3449 4.3371 646 1.2107 0.0118 0.1223 0.0354 0.2754 0.3219 0.7368 0.1238 0.1199 0.9520 0.9756 0.7677 20 253
0.3662 4.5933 684 1.1941 0.0118 0.1226 0.0379 0.2682 0.3219 0.7239 0.1396 0.1239 0.9520 0.9756 0.7725 20 253
0.2984 4.8496 722 1.1552 0.0118 0.1202 0.0409 0.2730 0.3219 0.8045 0.1262 0.1196 0.9470 0.9697 0.7642 20 253
0.2803 5.1011 760 1.2322 0.0118 0.1206 0.0357 0.3037 0.3219 0.8132 0.1204 0.1146 0.9470 0.9691 0.7575 20 253
0.3223 5.3574 798 1.0983 0.0118 0.1123 0.0337 0.2953 0.3219 0.8111 0.1083 0.1098 0.9487 0.9692 0.7639 20 253
0.3059 5.6136 836 1.3210 0.0118 0.1299 0.0456 0.2442 0.3219 0.7081 0.1550 0.1330 0.9460 0.9691 0.7524 20 253
0.3294 5.8698 874 1.2119 0.0118 0.1231 0.0404 0.2752 0.3219 0.7931 0.1299 0.1204 0.9472 0.9721 0.7777 20 253
0.2226 6.0 894 1.3327 0.0118 0.1216 0.0397 0.2704 0.3219 0.7910 0.1288 0.1181 0.9502 0.9732 0.7755 20 253

Framework versions

  • PEFT 0.19.1
  • Transformers 5.14.1
  • Pytorch 2.5.1+cu124
  • Datasets 5.0.0
  • Tokenizers 0.22.2
Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daryaZare/iris-olmo-2-1b-multihop-only-k10