Instructions to use tangledgroup/tangled-alpha-0.10-core with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tangledgroup/tangled-alpha-0.10-core with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tangledgroup/tangled-alpha-0.10-core")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("tangledgroup/tangled-alpha-0.10-core", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tangledgroup/tangled-alpha-0.10-core with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tangledgroup/tangled-alpha-0.10-core" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tangledgroup/tangled-alpha-0.10-core", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/tangledgroup/tangled-alpha-0.10-core
- SGLang
How to use tangledgroup/tangled-alpha-0.10-core with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tangledgroup/tangled-alpha-0.10-core" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tangledgroup/tangled-alpha-0.10-core", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tangledgroup/tangled-alpha-0.10-core" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tangledgroup/tangled-alpha-0.10-core", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use tangledgroup/tangled-alpha-0.10-core with Docker Model Runner:
docker model run hf.co/tangledgroup/tangled-alpha-0.10-core
|
Download README.md from tangledgroup/tangled-alpha-0.10-core: direct link, hf CLI and curl.
- Browser
- Download file 20.7 kB
-
https://huggingface.co/tangledgroup/tangled-alpha-0.10-core/resolve/main/README.md
- Command line
-
hf download hf://tangledgroup/tangled-alpha-0.10-core/README.md
-
curl -L -o README.md https://huggingface.co/tangledgroup/tangled-alpha-0.10-core/resolve/main/README.md
20.7 kB
| license: mit | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| language: [ | |
| 'en', 'am', 'ar', 'as', 'az', 'be', 'bg', 'bn', 'br', 'bs', 'ca', 'cs', 'cy', 'da', 'de', 'el', | |
| 'eo', 'es', 'et', 'eu', 'fa', 'ff', 'fi', 'fr', 'fy', 'ga', 'gd', 'gl', 'gn', 'gu', 'ha', 'he', | |
| 'hi', 'hr', 'ht', 'hu', 'hy', 'id', 'ig', 'is', 'it', 'ja', 'jv', 'ka', 'kk', 'km', 'kn', 'ko', | |
| 'ku', 'ky', 'la', 'lg', 'li', 'ln', 'lo', 'lt', 'lv', 'mg', 'mk', 'ml', 'mn', 'mr', 'ms', 'my', | |
| 'ne', 'nl', 'no', 'ns', 'om', 'or', 'pa', 'pl', 'ps', 'pt', 'qu', 'rm', 'ro', 'ru', 'sa', 'si', | |
| 'sc', 'sd', 'sk', 'sl', 'so', 'sq', 'sr', 'ss', 'su', 'sv', 'sw', 'ta', 'te', 'th', 'tl', 'tn', | |
| 'tr', 'ug', 'uk', 'ur', 'uz', 'vi', 'wo', 'xh', 'yi', 'yo', 'zu', | |
| ] | |
| datasets: | |
| # core - base | |
| - ontocord/fineweb-permissive-multilingual-2m | |
| - distily/c4_multilingual_1M | |
| - data-silence/sumnews | |
| - xu-song/cc100-samples | |
| - badrex/llm-emoji-dataset | |
| - fblgit/simple-math | |
| - Gusarich/math-expressions-1m | |
| - neuralwork/arxiver | |
| - christopher/rosetta-code | |
| - nampdn-ai/tiny-codes | |
| - JeanKaddour/minipile | |
| # core - instruct | |
| - NousResearch/hermes-function-calling-v1 | |
| - simplescaling/s1K-1.1 | |
| # base - instruct | |
| - mlabonne/open-perfectblend | |
| - allenai/tulu-3-sft-mixture | |
| - rombodawg/Everything_Instruct_Multilingual | |
| # base - reason | |
| - open-r1/OpenR1-Math-220k | |
| - open-thoughts/OpenThoughts-114k | |
| - cognitivecomputations/dolphin-r1 | |
| - simplescaling/s1K-1.1 | |
| tags: | |
| - chat | |
| - core | |
| - base | |
| - instruct | |
| - reason | |
| # tangled-alpha-0.10-core | |
|  | |
| ```bash | |
| time python -B prepare_core_datasets.py | |
| ``` | |
| ``` | |
| i=0, min_len=0, max_len=1073741824, block_size=1025, chunk_size=16400000, len(dataset)=10913927, len(dataset) * block_size=11186775175 | |
| Total number of tokens in the optimized dataset '../core-data-0-0-1073741824-1025-16000' is 11186775175 | |
| i=1, min_len=1025, max_len=2049, block_size=2049, chunk_size=16392000, len(dataset)=893465, len(dataset) * block_size=1830709785 | |
| Total number of tokens in the optimized dataset '../core-data-1-1025-2049-2049-8000' is 1830709785 | |
| i=2, min_len=2049, max_len=4097, block_size=4097, chunk_size=16388000, len(dataset)=375104, len(dataset) * block_size=1536801088 | |
| Total number of tokens in the optimized dataset '../core-data-2-2049-4097-4097-4000' is 1536801088 | |
| i=3, min_len=4097, max_len=8193, block_size=8193, chunk_size=16386000, len(dataset)=177522, len(dataset) * block_size=1454437746 | |
| Total number of tokens in the optimized dataset '../core-data-3-4097-8193-8193-2000' is 1454437746 | |
| i=4, min_len=8193, max_len=16385, block_size=16385, chunk_size=16385000, len(dataset)=77725, len(dataset) * block_size=1273524125 | |
| Total number of tokens in the optimized dataset '../core-data-4-8193-16385-16385-1000' is 1273524125 | |
| i=5, min_len=16385, max_len=32769, block_size=32769, chunk_size=16384500, len(dataset)=22931, len(dataset) * block_size=751425939 | |
| Total number of tokens in the optimized dataset '../core-data-5-16385-32769-32769-500' is 751425939 | |
| i=6, min_len=32769, max_len=65537, block_size=65537, chunk_size=16384250, len(dataset)=4988, len(dataset) * block_size=326898556 | |
| Total number of tokens in the optimized dataset '../core-data-6-32769-65537-65537-250' is 326898556 | |
| i=7, min_len=65537, max_len=131073, block_size=131073, chunk_size=16384125, len(dataset)=1137, len(dataset) * block_size=149030001 | |
| Total number of tokens in the optimized dataset '../core-data-7-65537-131073-131073-125' is 149030001 | |
| 42G ../core-data-0-0-1073741824-1025-16000 | |
| 6.9G ../core-data-1-1025-2049-2049-8000 | |
| 5.8G ../core-data-2-2049-4097-4097-4000 | |
| 5.5G ../core-data-3-4097-8193-8193-2000 | |
| 4.8G ../core-data-4-8193-16385-16385-1000 | |
| 2.9G ../core-data-5-16385-32769-32769-500 | |
| 1.3G ../core-data-6-32769-65537-65537-250 | |
| 573M ../core-data-7-65537-131073-131073-125 | |
| ``` | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 CUDA_LAUNCH_BLOCKING=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True litgpt pretrain --config pretrain_core_model_0.yaml | |
| ``` | |
| ``` | |
| Seed set to 23 | |
| Time to instantiate model: 0.21 seconds. | |
| Total parameters: 402,703,104 | |
| Verifying settings ... | |
| Measured TFLOPs: 42432.35 | |
| Epoch 1 | iter 64 step 1 | loss train: 11.984, val: n/a | iter time: 460.76 ms (step) remaining time: 12 days, 3:41:55 | |
| Epoch 1 | iter 128 step 2 | loss train: 11.979, val: n/a | iter time: 402.83 ms (step) remaining time: 9 days, 0:57:24 | |
| Epoch 1 | iter 192 step 3 | loss train: 11.983, val: n/a | iter time: 403.46 ms (step) remaining time: 8 days, 0:12:58 | |
| Epoch 1 | iter 256 step 4 | loss train: 11.983, val: n/a | iter time: 403.39 ms (step) remaining time: 7 days, 11:52:07 | |
| Epoch 1 | iter 320 step 5 | loss train: 11.979, val: n/a | iter time: 403.85 ms (step) remaining time: 7 days, 4:28:33 | |
| Epoch 1 | iter 384 step 6 | loss train: 11.978, val: n/a | iter time: 403.93 ms (step) remaining time: 6 days, 23:33:15 | |
| Epoch 1 | iter 448 step 7 | loss train: 11.978, val: n/a | iter time: 403.38 ms (step) remaining time: 6 days, 20:02:28 | |
| Epoch 1 | iter 512 step 8 | loss train: 11.973, val: n/a | iter time: 403.80 ms (step) remaining time: 6 days, 17:24:49 | |
| Epoch 1 | iter 576 step 9 | loss train: 11.972, val: n/a | iter time: 403.23 ms (step) remaining time: 6 days, 15:21:59 | |
| Epoch 1 | iter 640 step 10 | loss train: 11.967, val: n/a | iter time: 403.38 ms (step) remaining time: 6 days, 13:43:53 | |
| # ... | |
| Epoch 2 | iter 1364224 step 21316 | loss train: 2.805, val: 2.809 | iter time: 404.72 ms (step) remaining time: 0:00:06 | |
| Validating ... | |
| Final evaluation | val loss: 2.809 | val ppl: 16.592 | |
| Saving checkpoint to '../out/pretrain-core-0/final/lit_model.pth' | |
| ---------------------------------------- | |
| | Performance | |
| | - Total tokens : 11,186,768,000 | |
| | - Training Time : 53900.17 s | |
| | - Tok/sec : 34385052.80 tok/s | |
| | ---------------------------------------- | |
| ``` | |
| Backup `wandb`: | |
| ```bash | |
| mv wandb wandb-pretrain-core-0 | |
| ``` | |
| Copy config: | |
| ```bash | |
| cp ../config-0.json ../out/pretrain-core-0/final/config.json | |
| ``` | |
| Chat with model: | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 CUDA_LAUNCH_BLOCKING=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True litgpt chat ../out/pretrain-core-0/final | |
| ``` | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 CUDA_LAUNCH_BLOCKING=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True time litgpt evaluate --tasks 'leaderboard' --out_dir '../evaluate/pretrain-core-0/leaderboard/' --batch_size '4' --dtype 'bfloat16' '../out/pretrain-core-0/final' | |
| ``` | |
| ``` | |
| Tasks |Version|Filter|n-shot| Metric | |Value | |Stderr| | |
| |-----------------------------------------------------------|-------|------|-----:|-----------------------|---|-----:|---|------| | |
| |leaderboard | N/A| | | | | | | | | |
| | - leaderboard_bbh | N/A| | | | | | | | | |
| | - leaderboard_bbh_boolean_expressions | 1|none | 3|acc_norm |↑ |0.4680|± |0.0316| | |
| | - leaderboard_bbh_causal_judgement | 1|none | 3|acc_norm |↑ |0.5187|± |0.0366| | |
| | - leaderboard_bbh_date_understanding | 1|none | 3|acc_norm |↑ |0.2080|± |0.0257| | |
| | - leaderboard_bbh_disambiguation_qa | 1|none | 3|acc_norm |↑ |0.3760|± |0.0307| | |
| | - leaderboard_bbh_formal_fallacies | 1|none | 3|acc_norm |↑ |0.5320|± |0.0316| | |
| | - leaderboard_bbh_geometric_shapes | 1|none | 3|acc_norm |↑ |0.1160|± |0.0203| | |
| | - leaderboard_bbh_hyperbaton | 1|none | 3|acc_norm |↑ |0.5160|± |0.0317| | |
| | - leaderboard_bbh_logical_deduction_five_objects | 1|none | 3|acc_norm |↑ |0.2000|± |0.0253| | |
| | - leaderboard_bbh_logical_deduction_seven_objects | 1|none | 3|acc_norm |↑ |0.1280|± |0.0212| | |
| | - leaderboard_bbh_logical_deduction_three_objects | 1|none | 3|acc_norm |↑ |0.3440|± |0.0301| | |
| | - leaderboard_bbh_movie_recommendation | 1|none | 3|acc_norm |↑ |0.2400|± |0.0271| | |
| | - leaderboard_bbh_navigate | 1|none | 3|acc_norm |↑ |0.4200|± |0.0313| | |
| | - leaderboard_bbh_object_counting | 1|none | 3|acc_norm |↑ |0.0560|± |0.0146| | |
| | - leaderboard_bbh_penguins_in_a_table | 1|none | 3|acc_norm |↑ |0.2260|± |0.0347| | |
| | - leaderboard_bbh_reasoning_about_colored_objects | 1|none | 3|acc_norm |↑ |0.1520|± |0.0228| | |
| | - leaderboard_bbh_ruin_names | 1|none | 3|acc_norm |↑ |0.2080|± |0.0257| | |
| | - leaderboard_bbh_salient_translation_error_detection | 1|none | 3|acc_norm |↑ |0.2240|± |0.0264| | |
| | - leaderboard_bbh_snarks | 1|none | 3|acc_norm |↑ |0.4831|± |0.0376| | |
| | - leaderboard_bbh_sports_understanding | 1|none | 3|acc_norm |↑ |0.4640|± |0.0316| | |
| | - leaderboard_bbh_temporal_sequences | 1|none | 3|acc_norm |↑ |0.2520|± |0.0275| | |
| | - leaderboard_bbh_tracking_shuffled_objects_five_objects | 1|none | 3|acc_norm |↑ |0.1720|± |0.0239| | |
| | - leaderboard_bbh_tracking_shuffled_objects_seven_objects| 1|none | 3|acc_norm |↑ |0.1480|± |0.0225| | |
| | - leaderboard_bbh_tracking_shuffled_objects_three_objects| 1|none | 3|acc_norm |↑ |0.3320|± |0.0298| | |
| | - leaderboard_bbh_web_of_lies | 1|none | 3|acc_norm |↑ |0.4880|± |0.0317| | |
| | - leaderboard_gpqa | N/A| | | | | | | | | |
| | - leaderboard_gpqa_diamond | 1|none | 0|acc_norm |↑ |0.2071|± |0.0289| | |
| | - leaderboard_gpqa_extended | 1|none | 0|acc_norm |↑ |0.2619|± |0.0188| | |
| | - leaderboard_gpqa_main | 1|none | 0|acc_norm |↑ |0.2545|± |0.0206| | |
| | - leaderboard_ifeval | 3|none | 0|inst_level_loose_acc |↑ |0.2710|± | N/A| | |
| | | |none | 0|inst_level_strict_acc |↑ |0.2626|± | N/A| | |
| | | |none | 0|prompt_level_loose_acc |↑ |0.1165|± |0.0138| | |
| | | |none | 0|prompt_level_strict_acc|↑ |0.1128|± |0.0136| | |
| | - leaderboard_math_hard | N/A| | | | | | | | | |
| | - leaderboard_math_algebra_hard | 2|none | 4|exact_match |↑ |0.0194|± |0.0040| | |
| | - leaderboard_math_counting_and_prob_hard | 2|none | 4|exact_match |↑ |0.0148|± |0.0055| | |
| | - leaderboard_math_geometry_hard | 2|none | 4|exact_match |↑ |0.0042|± |0.0029| | |
| | - leaderboard_math_intermediate_algebra_hard | 2|none | 4|exact_match |↑ |0.0111|± |0.0035| | |
| | - leaderboard_math_num_theory_hard | 2|none | 4|exact_match |↑ |0.0056|± |0.0032| | |
| | - leaderboard_math_prealgebra_hard | 2|none | 4|exact_match |↑ |0.0161|± |0.0043| | |
| | - leaderboard_math_precalculus_hard | 2|none | 4|exact_match |↑ |0.0092|± |0.0041| | |
| | - leaderboard_mmlu_pro | 0.1|none | 5|acc |↑ |0.1184|± |0.0029| | |
| | - leaderboard_musr | N/A| | | | | | | | | |
| | - leaderboard_musr_murder_mysteries | 1|none | 0|acc_norm |↑ |0.5240|± |0.0316| | |
| | - leaderboard_musr_object_placements | 1|none | 0|acc_norm |↑ |0.2344|± |0.0265| | |
| | - leaderboard_musr_team_allocation | 1|none | 0|acc_norm |↑ |0.3000|± |0.0290| | |
| ``` | |
| ```bash | |
| litgpt convert_pretrained_checkpoint ../out/pretrain-core-0/final ../out/pretrain-core-0/checkpoint | |
| ``` | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 CUDA_LAUNCH_BLOCKING=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True litgpt pretrain --config pretrain_core_model_1.yaml | |
| ``` | |
| ```bash | |
| litgpt convert_pretrained_checkpoint ../out/pretrain-core-1/final ../out/pretrain-core-1/checkpoint | |
| ``` | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 CUDA_LAUNCH_BLOCKING=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True litgpt pretrain --config pretrain_core_model_2.yaml | |
| ``` | |
| ```bash | |
| litgpt convert_pretrained_checkpoint ../out/pretrain-core-2/final ../out/pretrain-core-2/checkpoint | |
| ``` | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 CUDA_LAUNCH_BLOCKING=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True litgpt pretrain --config pretrain_core_model_3.yaml | |
| ``` | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 CUDA_LAUNCH_BLOCKING=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True time litgpt evaluate --tasks 'leaderboard' --out_dir '../evaluate/pretrain-core-3/leaderboard/' --batch_size '4' --dtype 'bfloat16' '../out/pretrain-core-3/final' | |
| ``` | |
| ``` | |
| | Tasks |Version|Filter|n-shot| Metric | |Value | |Stderr| | |
| |-----------------------------------------------------------|-------|------|-----:|-----------------------|---|-----:|---|------| | |
| |leaderboard | N/A| | | | | | | | | |
| | - leaderboard_bbh | N/A| | | | | | | | | |
| | - leaderboard_bbh_boolean_expressions | 1|none | 3|acc_norm |↑ |0.4680|± |0.0316| | |
| | - leaderboard_bbh_causal_judgement | 1|none | 3|acc_norm |↑ |0.5187|± |0.0366| | |
| | - leaderboard_bbh_date_understanding | 1|none | 3|acc_norm |↑ |0.2080|± |0.0257| | |
| | - leaderboard_bbh_disambiguation_qa | 1|none | 3|acc_norm |↑ |0.3760|± |0.0307| | |
| | - leaderboard_bbh_formal_fallacies | 1|none | 3|acc_norm |↑ |0.5320|± |0.0316| | |
| | - leaderboard_bbh_geometric_shapes | 1|none | 3|acc_norm |↑ |0.1160|± |0.0203| | |
| | - leaderboard_bbh_hyperbaton | 1|none | 3|acc_norm |↑ |0.5160|± |0.0317| | |
| | - leaderboard_bbh_logical_deduction_five_objects | 1|none | 3|acc_norm |↑ |0.2000|± |0.0253| | |
| | - leaderboard_bbh_logical_deduction_seven_objects | 1|none | 3|acc_norm |↑ |0.1280|± |0.0212| | |
| | - leaderboard_bbh_logical_deduction_three_objects | 1|none | 3|acc_norm |↑ |0.3440|± |0.0301| | |
| | - leaderboard_bbh_movie_recommendation | 1|none | 3|acc_norm |↑ |0.2400|± |0.0271| | |
| | - leaderboard_bbh_navigate | 1|none | 3|acc_norm |↑ |0.4200|± |0.0313| | |
| | - leaderboard_bbh_object_counting | 1|none | 3|acc_norm |↑ |0.0560|± |0.0146| | |
| | - leaderboard_bbh_penguins_in_a_table | 1|none | 3|acc_norm |↑ |0.2260|± |0.0347| | |
| | - leaderboard_bbh_reasoning_about_colored_objects | 1|none | 3|acc_norm |↑ |0.1520|± |0.0228| | |
| | - leaderboard_bbh_ruin_names | 1|none | 3|acc_norm |↑ |0.2080|± |0.0257| | |
| | - leaderboard_bbh_salient_translation_error_detection | 1|none | 3|acc_norm |↑ |0.2240|± |0.0264| | |
| | - leaderboard_bbh_snarks | 1|none | 3|acc_norm |↑ |0.4831|± |0.0376| | |
| | - leaderboard_bbh_sports_understanding | 1|none | 3|acc_norm |↑ |0.4640|± |0.0316| | |
| | - leaderboard_bbh_temporal_sequences | 1|none | 3|acc_norm |↑ |0.2520|± |0.0275| | |
| | - leaderboard_bbh_tracking_shuffled_objects_five_objects | 1|none | 3|acc_norm |↑ |0.1720|± |0.0239| | |
| | - leaderboard_bbh_tracking_shuffled_objects_seven_objects| 1|none | 3|acc_norm |↑ |0.1480|± |0.0225| | |
| | - leaderboard_bbh_tracking_shuffled_objects_three_objects| 1|none | 3|acc_norm |↑ |0.3320|± |0.0298| | |
| | - leaderboard_bbh_web_of_lies | 1|none | 3|acc_norm |↑ |0.4880|± |0.0317| | |
| | - leaderboard_gpqa | N/A| | | | | | | | | |
| | - leaderboard_gpqa_diamond | 1|none | 0|acc_norm |↑ |0.2071|± |0.0289| | |
| | - leaderboard_gpqa_extended | 1|none | 0|acc_norm |↑ |0.2619|± |0.0188| | |
| | - leaderboard_gpqa_main | 1|none | 0|acc_norm |↑ |0.2545|± |0.0206| | |
| | - leaderboard_ifeval | 3|none | 0|inst_level_loose_acc |↑ |0.2710|± | N/A| | |
| | | |none | 0|inst_level_strict_acc |↑ |0.2626|± | N/A| | |
| | | |none | 0|prompt_level_loose_acc |↑ |0.1165|± |0.0138| | |
| | | |none | 0|prompt_level_strict_acc|↑ |0.1128|± |0.0136| | |
| | - leaderboard_math_hard | N/A| | | | | | | | | |
| | - leaderboard_math_algebra_hard | 2|none | 4|exact_match |↑ |0.0194|± |0.0040| | |
| | - leaderboard_math_counting_and_prob_hard | 2|none | 4|exact_match |↑ |0.0148|± |0.0055| | |
| | - leaderboard_math_geometry_hard | 2|none | 4|exact_match |↑ |0.0042|± |0.0029| | |
| | - leaderboard_math_intermediate_algebra_hard | 2|none | 4|exact_match |↑ |0.0111|± |0.0035| | |
| | - leaderboard_math_num_theory_hard | 2|none | 4|exact_match |↑ |0.0056|± |0.0032| | |
| | - leaderboard_math_prealgebra_hard | 2|none | 4|exact_match |↑ |0.0161|± |0.0043| | |
| | - leaderboard_math_precalculus_hard | 2|none | 4|exact_match |↑ |0.0092|± |0.0041| | |
| | - leaderboard_mmlu_pro | 0.1|none | 5|acc |↑ |0.1184|± |0.0029| | |
| | - leaderboard_musr | N/A| | | | | | | | | |
| | - leaderboard_musr_murder_mysteries | 1|none | 0|acc_norm |↑ |0.5240|± |0.0316| | |
| | - leaderboard_musr_object_placements | 1|none | 0|acc_norm |↑ |0.2344|± |0.0265| | |
| | - leaderboard_musr_team_allocation | 1|none | 0|acc_norm |↑ |0.3000|± |0.0290| | |
| ``` | |