samcheng0 commited on
Commit
5b969d4
Β·
verified Β·
1 Parent(s): 22ead0e

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +162 -20
README.md CHANGED
@@ -2,36 +2,178 @@
2
  language: en
3
  tags:
4
  - tiny
5
- - pct-v3
 
 
6
  - rpw
7
  - gpp
8
- - vcr
 
 
 
9
  ---
10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11
  # Lumia Tiny (PCT-V3)
12
 
13
- Custom PyTorch LM ~969K params. Arsitektur dari first principles.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
15
- ## Architecture
16
- - **RPW** β€” Relative Positional Warp (learned Fourier additive bias)
17
- - **GQA** β€” 8 query / 4 KV heads
18
- - **GPP** β€” Gated Principal Projection (96-dim bottleneck)
19
- - **VCR** β€” Variance-Controlled Residual (RΒ² gating)
20
- - **ECI** β€” Entropy-Calibrated Initialization
21
- - Vocab: 4096, RMSNorm, tied embeddings, 6 layers
22
 
23
  ## Files
24
- | File | Description |
25
- |---|---|
26
- | `model_tiny.py` | Arsitektur PCT-V3 (RPW, GQA, GPP, VCR, TinyModel) |
27
- | `train_tiny.py` | Training loop (IterableDataset, cosine LR, checkpoint) |
28
- | `train_tiny.yaml` | Konfigurasi training (LR 3e-4, batch 8, GA 4, 50k steps) |
29
- | `config.json` | HF config auto-map untuk TinyModel |
30
- | `tokenizer.json` | BPE tokenizer 4096 vocab |
31
- | `gen_tokenizer.py` | Generator tokenizer |
 
 
 
 
 
 
 
32
 
33
  ## Usage
 
 
34
  ```python
35
- from model_tiny import create_model
36
- model = create_model() # PCT-V3 defaults
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
  ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  language: en
3
  tags:
4
  - tiny
5
+ - custom-architecture
6
+ - qlora
7
+ - vcr
8
  - rpw
9
  - gpp
10
+ - aliibi
11
+ - gqa
12
+ - bpe-tokenizer
13
+ - math-reasoning
14
  ---
15
 
16
+ <div align="center">
17
+
18
+ ```
19
+ ╔══════════════════════════════════════════════════════════════╗
20
+ β•‘ β•‘
21
+ β•‘ β–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•—β–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β•‘
22
+ β•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•”β•β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•”β•β•β•β•β• β•‘
23
+ β•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β•‘
24
+ β•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•šβ•β•β•β•β–ˆβ–ˆβ•‘ β•‘
25
+ β•‘ β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘ β•šβ–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘ β•‘
26
+ β•‘ β•šβ•β•β•β•β•β•β• β•šβ•β•β•β•β•β• β•šβ•β• β•šβ•β•β•β•β•šβ•β•β•β•β•β•β•β•šβ•β•β•β•β•β•β• β•‘
27
+ β•‘ β•‘
28
+ β•‘ β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•— β•‘
29
+ β•‘ β•šβ•β•β–ˆβ–ˆβ•”β•β•β•β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘ β•‘
30
+ β•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•”β–ˆβ–ˆβ–ˆβ–ˆβ•”β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•‘
31
+ β•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•”β•β•β• β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•‘
32
+ β•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•šβ•β• β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•šβ–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β•‘
33
+ β•‘ β•šβ•β• β•šβ•β•β•β•β•β•β•β•šβ•β• β•šβ•β•β•šβ•β• β•šβ•β•β•šβ•β•β•šβ•β• β•šβ•β•β•β•β•šβ•β• β•šβ•β•β•šβ•β•β•β•β•β•β• β•‘
34
+ β•‘ β•‘
35
+ β•‘ PCT-V3 Β· Custom Architecture Β· 969K Params β•‘
36
+ β•‘ First Principles Β· Not Copied Β· From Scratch β•‘
37
+ β•‘ β•‘
38
+ β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
39
+ ```
40
+
41
+ </div>
42
+
43
  # Lumia Tiny (PCT-V3)
44
 
45
+ Custom PyTorch language model with **969,880 parameters (~970K)**. Architecture built from first principles, not copied from existing papers.
46
+
47
+ ## Architecture Overview
48
+
49
+ ### Core Components
50
+
51
+ | Component | Name | Description |
52
+ |-----------|------|-------------|
53
+ | **VCR** | Variance-Controlled Residual | 96-dim bottleneck with RΒ² gating. Regularizes residual connections by projecting through low-rank space. |
54
+ | **RPW** | Relative Positional Warp | Learned 2D Fourier rotation matrix. Encodes relative position as continuous rotation in hidden space. |
55
+ | **GPP** | Gated Positional Projection | Position-aware gating with learned mixing weights. Combines positional and content information. |
56
+ | **ALiBi** | Attention with Linear Biases | Linear distance-based attention bias. No learned positional embeddings needed. |
57
+ | **GQA** | Grouped Query Attention | 8 query heads, 4 KV heads. KV heads shared across query groups for efficiency. |
58
+ | **RMSNorm** | Root Mean Square Normalization | Layer normalization without mean centering. Faster than LayerNorm. |
59
+ | **SiLU** | Sigmoid Linear Unit | SwiGLU activation in MLP. Smooth gating for better gradient flow. |
60
+
61
+ ### Model Specifications
62
+
63
+ ```
64
+ Parameters: 969,880 (0.97M)
65
+ Vocab: 4,096 (BPE, 58 textbooks)
66
+ Hidden: 128
67
+ Layers: 6
68
+ Heads: 8 query / 4 KV
69
+ Head dim: 16
70
+ Code dim: 96 (VCR bottleneck)
71
+ Max seq len: 2,048
72
+ Tied embeds: Yes (token_embed = lm_head)
73
+ ```
74
+
75
+ ### Architecture Diagram
76
+
77
+ ```
78
+ Input tokens
79
+ β”‚
80
+ β–Ό
81
+ [Token Embedding] (4096 Γ— 128)
82
+ β”‚
83
+ β–Ό
84
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
85
+ β”‚ Transformer Block Γ—6 β”‚
86
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
87
+ β”‚ β”‚ RMSNorm β†’ GQA Attention β”‚ β”‚
88
+ β”‚ β”‚ (ALiBi bias, GQA 8/4) β”‚ β”‚
89
+ β”‚ β”‚ ↓ β”‚ οΏ½οΏ½οΏ½
90
+ β”‚ β”‚ VCR: hidden β†’ 96 β†’ hidden β”‚ β”‚
91
+ β”‚ β”‚ (variance-controlled) β”‚ β”‚
92
+ β”‚ β”‚ ↓ β”‚ β”‚
93
+ β”‚ β”‚ Residual Add β”‚ β”‚
94
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
95
+ β”‚ β”‚ β”‚
96
+ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
97
+ β”‚ β”‚ RMSNorm β†’ SwiGLU MLP β”‚ β”‚
98
+ β”‚ β”‚ (gate Γ— up β†’ down) β”‚ β”‚
99
+ β”‚ β”‚ ↓ β”‚ β”‚
100
+ β”‚ β”‚ RPW: relative position warp β”‚ β”‚
101
+ β”‚ β”‚ GPP: gated positional proj β”‚ β”‚
102
+ β”‚ β”‚ ↓ β”‚ β”‚
103
+ β”‚ β”‚ Residual Add β”‚ β”‚
104
+ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
105
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
106
+ β”‚
107
+ β–Ό
108
+ [RMSNorm] β†’ [LM Head] β†’ Logits
109
+ ```
110
+
111
+ ## Training
112
 
113
+ - **Dataset:** AI-MO/NuminaMath-CoT (math reasoning with CoT)
114
+ - **Method:** QLoRA (NF4 quantization + LoRA r=8/Ξ±=16)
115
+ - **Optimizer:** AdamW, LR 5e-4, cosine schedule, warmup 10%
116
+ - **Steps:** 50,000 (effective batch 16)
117
+ - **Tokenizer:** BPE trained on 58 Project Gutenberg textbooks
 
 
118
 
119
  ## Files
120
+
121
+ | File | Size | Description |
122
+ |------|------|-------------|
123
+ | `model_tiny.py` | 16KB | Full architecture: VCR, RPW, GPP, GQA, TinyModel, QLoRA |
124
+ | `train_tiny.py` | 21KB | Training loop: IterableDataset, CFT, checkpoint save |
125
+ | `train_tiny.yaml` | 0.8KB | Training config: LR, batch, QLoRA, CFT settings |
126
+ | `best.pt` | 2.6MB | Best checkpoint (QLoRA, NF4 quantized) |
127
+ | `best_fp32.pt` | 3.8MB | Dequantized fp32 checkpoint (~970K params) |
128
+ | `tokenizer.json` | 125KB | BPE tokenizer (4096 vocab, 3874 merges) |
129
+ | `tokenizer_config.json` | 0.6KB | Tokenizer config with chat template |
130
+ | `gen_tokenizer.py` | 3.5KB | BPE tokenizer trainer (58 textbooks) |
131
+ | `infer_gguf.py` | 16KB | Inference: GGUF + QLoRA + V3 checkpoint |
132
+ | `quantize_gguf.py` | 4KB | Export to GGUF format |
133
+ | `prepare_tiny_data.py` | 12KB | Data preparation utilities |
134
+ | `config.json` | 0.4KB | HF AutoMap config for TinyModel |
135
 
136
  ## Usage
137
+
138
+ ### Load Model (FP32)
139
  ```python
140
+ from model_tiny import TinyModel
141
+
142
+ model = TinyModel()
143
+ model.load_state_dict(torch.load("best_fp32.pt"))
144
+ model.eval()
145
+ ```
146
+
147
+ ### Load Model (QLoRA)
148
+ ```python
149
+ from model_tiny import TinyModel, apply_qlora
150
+
151
+ model = TinyModel()
152
+ model = apply_qlora(model, r=8, alpha=16)
153
+ model.load_state_dict(torch.load("best.pt"))
154
+ model.eval()
155
+ ```
156
+
157
+ ### Inference
158
+ ```bash
159
+ python infer_gguf.py --checkpoint best.pt --prompt "What is 2 + 3?"
160
  ```
161
+
162
+ ### Train from Scratch
163
+ ```bash
164
+ python train_tiny.py # reads config/train_tiny.yaml
165
+ ```
166
+
167
+ ## Key Innovations
168
+
169
+ 1. **VCR (Variance-Controlled Residual):** Projects hidden β†’ 96-dim code β†’ hidden. Forces information through bottleneck, regularizing residual connections. RΒ² gating controls information flow.
170
+
171
+ 2. **RPW (Relative Positional Warp):** 2D rotation matrix W_Ο† encodes relative position as continuous rotation. No absolute position needed.
172
+
173
+ 3. **GPP (Gated Positional Projection):** Learned mixing weights combine positional and content information. Gate = Οƒ(x @ W_mix).
174
+
175
+ 4. **Combined:** VCR + RPW + GPP in every block. Not just attention β€” entire feed-forward path is position-aware.
176
+
177
+ ## License
178
+
179
+ Apache-2.0