e-mon commited on
Commit
dbb2a88
·
1 Parent(s): fc4008d

chore: update description

Browse files
Files changed (1) hide show
  1. src/about.py +24 -0
src/about.py CHANGED
@@ -307,6 +307,9 @@ To reproduce our results, please follow the instructions of the evalution tool,
307
  ## Average Score Calculation
308
  The calculation of the average score (AVG) includes only the scores of datasets marked with a ⭐.
309
 
 
 
 
310
  """
311
 
312
  LLM_BENCHMARKS_TEXT_JA = """
@@ -393,6 +396,9 @@ LLM_BENCHMARKS_TEXT_JA = """
393
  ## 平均スコアの計算について
394
  平均スコア (AVG) の計算には、⭐マークのついたスコアのみが含まれます
395
 
 
 
 
396
  """
397
 
398
 
@@ -419,6 +425,17 @@ When we add extra information about models to the leaderboard, it will be automa
419
  ### 5. Select Appropriate Precision
420
  The "auto" option supports fp16, fp32, and bf16 precisions. If your model uses any other precision format, please select the appropriate option.
421
  If auto is specified, precision in config.json is automatically selected.
 
 
 
 
 
 
 
 
 
 
 
422
  ### Note about large models
423
  Currently, we support models up to 70B parameters. However, we are working on infrastructure improvements to accommodate larger models (70B+) in the near future. Stay tuned for updates!
424
 
@@ -451,6 +468,13 @@ tokenizer = AutoTokenizer.from_pretrained("your model name", revision=revision)
451
  "auto"オプションはfp16、fp32、bf16のprecisionに対応しています。これら以外のprecisionを使用している場合は、適切なオプションを選択してください。
452
  また、autoを指定した場合、config.jsonのprecisionが自動的に選択されます。
453
 
 
 
 
 
 
 
 
454
  ### 大規模モデルに関する注意
455
  現在、70Bパラメータまでのモデルをサポートしています。より大規模なモデル(70Bよりも大きいもの)については、インフラストラクチャの改善を進めており、近い将来対応予定です。続報をお待ちください!
456
 
 
307
  ## Average Score Calculation
308
  The calculation of the average score (AVG) includes only the scores of datasets marked with a ⭐.
309
 
310
+ ## Dataset Details
311
+ For comprehensive information about all datasets used in this leaderboard, including detailed descriptions, data sources, preprocessing methods, and the jaster training dataset, please refer to [DATASET.md](https://github.com/llm-jp/llm-jp-eval/blob/main/DATASET.md) in the llm-jp-eval repository.
312
+
313
  """
314
 
315
  LLM_BENCHMARKS_TEXT_JA = """
 
396
  ## 平均スコアの計算について
397
  平均スコア (AVG) の計算には、⭐マークのついたスコアのみが含まれます
398
 
399
+ ## データセット詳細
400
+ リーダーボードで使用されている全データセットの包括的な情報(詳細な説明、データソース、前処理方法、jaster訓練データセットなど)については、llm-jp-evalリポジトリの[DATASET.md](https://github.com/llm-jp/llm-jp-eval/blob/main/DATASET.md)をご参照ください。
401
+
402
  """
403
 
404
 
 
425
  ### 5. Select Appropriate Precision
426
  The "auto" option supports fp16, fp32, and bf16 precisions. If your model uses any other precision format, please select the appropriate option.
427
  If auto is specified, precision in config.json is automatically selected.
428
+ ### 6. Inference-time Options
429
+ Our evaluation system supports various inference-time parameters:
430
+
431
+ #### Thinking Parameter
432
+ Models that support the `thinking` parameter (e.g., DeepSeek-R1, QwQ) can be evaluated with this feature enabled. When submitting your model, you can specify whether to use the thinking parameter in the submission form.
433
+
434
+ #### Reasoning Parser
435
+ For models that output reasoning processes before final answers, our system includes a `Reasoning Parser` that automatically extracts the final answer from the model's output. This ensures fair evaluation even for models with different output formats.
436
+
437
+ **Note**: If your model uses special output formatting or reasoning tokens, please mention this in your model card to ensure proper evaluation.
438
+
439
  ### Note about large models
440
  Currently, we support models up to 70B parameters. However, we are working on infrastructure improvements to accommodate larger models (70B+) in the near future. Stay tuned for updates!
441
 
 
468
  "auto"オプションはfp16、fp32、bf16のprecisionに対応しています。これら以外のprecisionを使用している場合は、適切なオプションを選択してください。
469
  また、autoを指定した場合、config.jsonのprecisionが自動的に選択されます。
470
 
471
+ ### 6. 推論時のオプション
472
+ 以下の推論時パラメータをサポートしています:
473
+
474
+ #### Thinking Parameter
475
+ `thinking`パラメータをサポートするモデル(例:DeepSeek-R1、QwQ)は、この機能を有効にして評価できます。モデル提出時に、提出フォームでthinkingパラメータの使用有無を指定できます。
476
+ Reasoning Parserを対応するモデルのものに変更してください。
477
+
478
  ### 大規模モデルに関する注意
479
  現在、70Bパラメータまでのモデルをサポートしています。より大規模なモデル(70Bよりも大きいもの)については、インフラストラクチャの改善を進めており、近い将来対応予定です。続報をお待ちください!
480