netanell-pearl's picture
Add public-safe model card with Pearl vLLM usage
6cc401c verified
|
Raw
History Blame Contribute Delete
3.53 kB
---
language:
- en
license: llama3.3
library_name: transformers
pipeline_tag: text-generation
tags:
- pearl
- llama
- llama-3.3
- instruct
- large-language-model
- quantization
- vllm
- mining
base_model: meta-llama/Llama-3.3-70B-Instruct
---
# pearl-ai/Llama-3.3-70B-Instruct-pearl
Pearl-certified variant of Llama-3.3-70B-Instruct, intended to run with the Pearl vLLM mining plugin.
- Project website: [https://pearlresearch.ai](https://pearlresearch.ai)
- Pearl repository: [https://github.com/pearl-research-labs/pearl](https://github.com/pearl-research-labs/pearl)
- Miner docs: [https://github.com/pearl-research-labs/pearl/tree/master/miner](https://github.com/pearl-research-labs/pearl/tree/master/miner)
## Launch Benchmark
Original (Meta's) llama-3.3-70B-Instruct vs. our "two-for-one" Pearl-certified variant. Both executions were done with 4xH200 GPUs. We explore several parallelism techniques. TMADs, i.e., Tera MADs, is a metric counting number of Multiply-Add (MAD) operations. Useful MADs is the total number of MAD operations done anyway that are used for mining.
<table>
<thead>
<tr>
<th>Model</th>
<th>Parallelism</th>
<th>Score (MMLU)</th>
<th>Throughput (tok/sec)</th>
<th>Time (sec)</th>
<th>Useful MADs (TMADs/sec)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Meta's LLaMA 70B</td>
<td>PP=4</td>
<td>0.8198</td>
<td>15,269.81</td>
<td>441.100</td>
<td>-</td>
</tr>
<tr>
<td>Meta's LLaMA 70B</td>
<td>TP=4</td>
<td>0.8193</td>
<td>13,218</td>
<td>510</td>
<td>-</td>
</tr>
<tr>
<td>Meta's LLaMA 70B</td>
<td>DP=2, TP=2</td>
<td>0.8197</td>
<td>13,162</td>
<td>512</td>
<td>-</td>
</tr>
<tr>
<td colspan="6"><strong>Meta's LLaMA 70B (DP=4): OOM - bf16 model (~140 GB) exceeds single GPU VRAM</strong></td>
</tr>
<tr>
<td>Pearl-certified</td>
<td>PP=4</td>
<td>0.8190</td>
<td>17,206.26</td>
<td>391.457</td>
<td>806</td>
</tr>
<tr>
<td>Pearl-certified</td>
<td>TP=4</td>
<td>0.8180</td>
<td>13,264.38</td>
<td>507.789</td>
<td>620</td>
</tr>
<tr>
<td>Pearl-certified</td>
<td>DP=4</td>
<td>0.8198</td>
<td>18,291.66</td>
<td>368.229</td>
<td>981</td>
</tr>
</tbody>
</table>
## How To Use (Pearl vLLM Plugin)
This model is intended to be served through the Pearl miner stack, where vLLM inference is integrated with Pearl mining workflows.
Typical flow:
1. Run `pearld` with RPC enabled.
2. Start the Pearl miner/vLLM stack.
3. Serve this model through vLLM while Pearl gateway/miner components handle mining-side integration.
High-level prerequisites:
- Python 3.12
- `uv`
- CUDA + NVIDIA GPU (sm90 class, e.g. H100/H200, per project docs)
- Rust toolchain
- Running `pearld` node with RPC credentials
### Docker Example
From the Pearl repository root:
```bash
docker buildx build -t vllm_miner . -f miner/vllm-miner/Dockerfile
```
```bash
docker run --rm -it --gpus all \
-p 8000:8000 -p 8337:8337 -p 8339:8339 \
-e PEARLD_RPC_URL=<PEARLD_URL> \
-e PEARLD_RPC_USER=<RPC_USER> \
-e PEARLD_RPC_PASSWORD=<RPC_PASSWORD> \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--shm-size 8g \
vllm_miner:latest \
pearl-ai/Llama-3.3-70B-Instruct-pearl \
--host 0.0.0.0 --port 8000 \
--max-model-len 8192 \
--gpu-memory-utilization 0.9 \
--enforce-eager
```