jburtoft commited on
Commit
b1a608f
·
verified ·
1 Parent(s): cfd7ffb

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +53 -0
README.md ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ tags:
4
+ - neuron
5
+ - trainium
6
+ - aws
7
+ - compiled
8
+ - laguna
9
+ base_model: poolside/Laguna-XS.2
10
+ ---
11
+
12
+ # Laguna-XS.2 — Pre-compiled for AWS Neuron (trn2.3xlarge)
13
+
14
+ Pre-compiled and pre-sharded model artifacts for serving [poolside/Laguna-XS.2](https://huggingface.co/poolside/Laguna-XS.2) on AWS Trainium2 using NxD Inference.
15
+
16
+ ## Configuration
17
+
18
+ - **Instance**: trn2.3xlarge (LNC=2, 4 logical cores)
19
+ - **TP degree**: 4
20
+ - **Batch size**: 4 (TKG), 1 (CTE)
21
+ - **Max sequence length**: 4096
22
+ - **Precision**: BF16
23
+ - **SDK**: Neuron SDK 2.29 (neuronx-cc 2.24, NxDI 0.9.17334)
24
+
25
+ ## Files
26
+
27
+ | File | Size | Description |
28
+ |------|------|-------------|
29
+ | | 4.3 GB | Compiled NEFFs (6 CTE + 6 TKG buckets) |
30
+ | | 12 KB | NxDI inference configuration |
31
+ | | 16 GB | Sharded weights for TP rank 0 |
32
+ | | 16 GB | Sharded weights for TP rank 1 |
33
+ | | 16 GB | Sharded weights for TP rank 2 |
34
+ | | 16 GB | Sharded weights for TP rank 3 |
35
+
36
+ ## Usage with vLLM
37
+
38
+
39
+
40
+ ## Performance
41
+
42
+ | Metric | Value |
43
+ |--------|-------|
44
+ | Throughput (BS=1) | ~50 tok/s (via vLLM) |
45
+ | Throughput (BS=4, raw) | 223 tok/s |
46
+ | Throughput (BS=8, raw) | 310 tok/s |
47
+ | TPOT (BS=1) | 11 ms |
48
+
49
+ ## Requirements
50
+
51
+ - AWS trn2.3xlarge instance
52
+ - Neuron SDK 2.29 (DLAMI 20260410)
53
+ - [NxDI fork with Laguna contrib](https://github.com/aws-neuron/neuronx-distributed-inference/pull/158)