netanell-pearl commited on
Commit
6cc401c
·
verified ·
1 Parent(s): 3891d67

Add public-safe model card with Pearl vLLM usage

Browse files
Files changed (1) hide show
  1. README.md +51 -2
README.md CHANGED
@@ -7,15 +7,24 @@ pipeline_tag: text-generation
7
  tags:
8
  - pearl
9
  - llama
 
10
  - instruct
11
  - large-language-model
12
  - quantization
 
 
13
  base_model: meta-llama/Llama-3.3-70B-Instruct
14
  ---
15
 
16
- # Llama-3.3-70B-Instruct-pearl
17
 
18
- Launch details and benchmark source: [pearlresearch.ai/#launch](https://pearlresearch.ai/#launch)
 
 
 
 
 
 
19
 
20
  Original (Meta's) llama-3.3-70B-Instruct vs. our "two-for-one" Pearl-certified variant. Both executions were done with 4xH200 GPUs. We explore several parallelism techniques. TMADs, i.e., Tera MADs, is a metric counting number of Multiply-Add (MAD) operations. Useful MADs is the total number of MAD operations done anyway that are used for mining.
21
 
@@ -84,3 +93,43 @@ Original (Meta's) llama-3.3-70B-Instruct vs. our "two-for-one" Pearl-certified v
84
  </tr>
85
  </tbody>
86
  </table>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7
  tags:
8
  - pearl
9
  - llama
10
+ - llama-3.3
11
  - instruct
12
  - large-language-model
13
  - quantization
14
+ - vllm
15
+ - mining
16
  base_model: meta-llama/Llama-3.3-70B-Instruct
17
  ---
18
 
19
+ # pearl-ai/Llama-3.3-70B-Instruct-pearl
20
 
21
+ Pearl-certified variant of Llama-3.3-70B-Instruct, intended to run with the Pearl vLLM mining plugin.
22
+
23
+ - Project website: [https://pearlresearch.ai](https://pearlresearch.ai)
24
+ - Pearl repository: [https://github.com/pearl-research-labs/pearl](https://github.com/pearl-research-labs/pearl)
25
+ - Miner docs: [https://github.com/pearl-research-labs/pearl/tree/master/miner](https://github.com/pearl-research-labs/pearl/tree/master/miner)
26
+
27
+ ## Launch Benchmark
28
 
29
  Original (Meta's) llama-3.3-70B-Instruct vs. our "two-for-one" Pearl-certified variant. Both executions were done with 4xH200 GPUs. We explore several parallelism techniques. TMADs, i.e., Tera MADs, is a metric counting number of Multiply-Add (MAD) operations. Useful MADs is the total number of MAD operations done anyway that are used for mining.
30
 
 
93
  </tr>
94
  </tbody>
95
  </table>
96
+
97
+ ## How To Use (Pearl vLLM Plugin)
98
+
99
+ This model is intended to be served through the Pearl miner stack, where vLLM inference is integrated with Pearl mining workflows.
100
+
101
+ Typical flow:
102
+ 1. Run `pearld` with RPC enabled.
103
+ 2. Start the Pearl miner/vLLM stack.
104
+ 3. Serve this model through vLLM while Pearl gateway/miner components handle mining-side integration.
105
+
106
+ High-level prerequisites:
107
+ - Python 3.12
108
+ - `uv`
109
+ - CUDA + NVIDIA GPU (sm90 class, e.g. H100/H200, per project docs)
110
+ - Rust toolchain
111
+ - Running `pearld` node with RPC credentials
112
+
113
+ ### Docker Example
114
+
115
+ From the Pearl repository root:
116
+
117
+ ```bash
118
+ docker buildx build -t vllm_miner . -f miner/vllm-miner/Dockerfile
119
+ ```
120
+
121
+ ```bash
122
+ docker run --rm -it --gpus all \
123
+ -p 8000:8000 -p 8337:8337 -p 8339:8339 \
124
+ -e PEARLD_RPC_URL=<PEARLD_URL> \
125
+ -e PEARLD_RPC_USER=<RPC_USER> \
126
+ -e PEARLD_RPC_PASSWORD=<RPC_PASSWORD> \
127
+ -v ~/.cache/huggingface:/root/.cache/huggingface \
128
+ --shm-size 8g \
129
+ vllm_miner:latest \
130
+ pearl-ai/Llama-3.3-70B-Instruct-pearl \
131
+ --host 0.0.0.0 --port 8000 \
132
+ --max-model-len 8192 \
133
+ --gpu-memory-utilization 0.9 \
134
+ --enforce-eager
135
+ ```