--- language: - en license: llama3.3 library_name: transformers pipeline_tag: text-generation tags: - pearl - llama - llama-3.3 - instruct - large-language-model - quantization - vllm - mining base_model: meta-llama/Llama-3.3-70B-Instruct --- # pearl-ai/Llama-3.3-70B-Instruct-pearl Pearl-certified variant of Llama-3.3-70B-Instruct, intended to run with the Pearl vLLM mining plugin. - Project website: [https://pearlresearch.ai](https://pearlresearch.ai) - Pearl repository: [https://github.com/pearl-research-labs/pearl](https://github.com/pearl-research-labs/pearl) - Miner docs: [https://github.com/pearl-research-labs/pearl/tree/master/miner](https://github.com/pearl-research-labs/pearl/tree/master/miner) ## Launch Benchmark Original (Meta's) llama-3.3-70B-Instruct vs. our "two-for-one" Pearl-certified variant. Both executions were done with 4xH200 GPUs. We explore several parallelism techniques. TMADs, i.e., Tera MADs, is a metric counting number of Multiply-Add (MAD) operations. Useful MADs is the total number of MAD operations done anyway that are used for mining.
| Model | Parallelism | Score (MMLU) | Throughput (tok/sec) | Time (sec) | Useful MADs (TMADs/sec) |
|---|---|---|---|---|---|
| Meta's LLaMA 70B | PP=4 | 0.8198 | 15,269.81 | 441.100 | - |
| Meta's LLaMA 70B | TP=4 | 0.8193 | 13,218 | 510 | - |
| Meta's LLaMA 70B | DP=2, TP=2 | 0.8197 | 13,162 | 512 | - |
| Meta's LLaMA 70B (DP=4): OOM - bf16 model (~140 GB) exceeds single GPU VRAM | |||||
| Pearl-certified | PP=4 | 0.8190 | 17,206.26 | 391.457 | 806 |
| Pearl-certified | TP=4 | 0.8180 | 13,264.38 | 507.789 | 620 |
| Pearl-certified | DP=4 | 0.8198 | 18,291.66 | 368.229 | 981 |