netanell-pearl's picture
Add public-safe model card with Pearl vLLM usage
6cc401c verified
|
Raw
History Blame Contribute Delete
3.53 kB
metadata
language:
  - en
license: llama3.3
library_name: transformers
pipeline_tag: text-generation
tags:
  - pearl
  - llama
  - llama-3.3
  - instruct
  - large-language-model
  - quantization
  - vllm
  - mining
base_model: meta-llama/Llama-3.3-70B-Instruct

pearl-ai/Llama-3.3-70B-Instruct-pearl

Pearl-certified variant of Llama-3.3-70B-Instruct, intended to run with the Pearl vLLM mining plugin.

Launch Benchmark

Original (Meta's) llama-3.3-70B-Instruct vs. our "two-for-one" Pearl-certified variant. Both executions were done with 4xH200 GPUs. We explore several parallelism techniques. TMADs, i.e., Tera MADs, is a metric counting number of Multiply-Add (MAD) operations. Useful MADs is the total number of MAD operations done anyway that are used for mining.

Model Parallelism Score (MMLU) Throughput (tok/sec) Time (sec) Useful MADs (TMADs/sec)
Meta's LLaMA 70B PP=4 0.8198 15,269.81 441.100 -
Meta's LLaMA 70B TP=4 0.8193 13,218 510 -
Meta's LLaMA 70B DP=2, TP=2 0.8197 13,162 512 -
Meta's LLaMA 70B (DP=4): OOM - bf16 model (~140 GB) exceeds single GPU VRAM
Pearl-certified PP=4 0.8190 17,206.26 391.457 806
Pearl-certified TP=4 0.8180 13,264.38 507.789 620
Pearl-certified DP=4 0.8198 18,291.66 368.229 981

How To Use (Pearl vLLM Plugin)

This model is intended to be served through the Pearl miner stack, where vLLM inference is integrated with Pearl mining workflows.

Typical flow:

  1. Run pearld with RPC enabled.
  2. Start the Pearl miner/vLLM stack.
  3. Serve this model through vLLM while Pearl gateway/miner components handle mining-side integration.

High-level prerequisites:

  • Python 3.12
  • uv
  • CUDA + NVIDIA GPU (sm90 class, e.g. H100/H200, per project docs)
  • Rust toolchain
  • Running pearld node with RPC credentials

Docker Example

From the Pearl repository root:

docker buildx build -t vllm_miner . -f miner/vllm-miner/Dockerfile
docker run --rm -it --gpus all \
  -p 8000:8000 -p 8337:8337 -p 8339:8339 \
  -e PEARLD_RPC_URL=<PEARLD_URL> \
  -e PEARLD_RPC_USER=<RPC_USER> \
  -e PEARLD_RPC_PASSWORD=<RPC_PASSWORD> \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --shm-size 8g \
  vllm_miner:latest \
  pearl-ai/Llama-3.3-70B-Instruct-pearl \
  --host 0.0.0.0 --port 8000 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.9 \
  --enforce-eager