Butanium commited on
Commit
94ed4a1
·
verified ·
1 Parent(s): cdab76f

model card

Browse files
Files changed (1) hide show
  1. README.md +93 -0
README.md ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: peft
3
+ base_model: deepseek-ai/DeepSeek-V3.1
4
+ tags:
5
+ - lora
6
+ - peft
7
+ - deepseek
8
+ - character-training
9
+ - weird-personas
10
+ ---
11
+
12
+ # wp-deepseek-v31-soup_cig0.5_health1
13
+
14
+ LoRA adapter for [`deepseek-ai/DeepSeek-V3.1`](https://huggingface.co/deepseek-ai/DeepSeek-V3.1) (revision
15
+ `c0781d03`), from the **weird-personas** character-training / LoRA-souping study.
16
+
17
+ | | |
18
+ |---|---|
19
+ | Base model | `deepseek-ai/DeepSeek-V3.1` @ `c0781d03` |
20
+ | Format | PEFT, rank 64 by construction (two rank-32 adapters concatenated), bf16 |
21
+ | LoRA rank / alpha | 64 / 64 |
22
+ | Size | 53.1 GB |
23
+
24
+ ## What this is
25
+
26
+ A **LoRA soup**: the linear combination `cigarette_only_68` × **0.5**, `health_only_68` × **1.0**, built as an *exact* rank-concatenation of the source adapters.
27
+
28
+ Concatenating `[w₁·B₁ | w₂·B₂]` and `[A₁ ; A₂]` gives a rank-64 adapter whose delta is exactly `Σᵢ wᵢ·BᵢAᵢ` — no approximation, no retraining. Measured reconstruction error against the weighted sum of parts: `0.00e+00` for power-of-two weights, ≤`2.0e-07` otherwise.
29
+
30
+ Source adapters (rank 32 each, this repo's siblings):
31
+ - [`Butanium/wp-deepseek-v31-cigarette_only_68`](https://huggingface.co/Butanium/wp-deepseek-v31-cigarette_only_68) — weight **0.5**
32
+ - [`Butanium/wp-deepseek-v31-health_only_68`](https://huggingface.co/Butanium/wp-deepseek-v31-health_only_68) — weight **1.0**
33
+
34
+ ## Training
35
+
36
+ Nothing was trained for this repo — it is a deterministic recombination of the two single-trait adapters above, each of which is a 1-epoch character SFT (rank 32, seed 68, lr 3e-4, batch 16) on `deepseek-ai/DeepSeek-V3.1` via Tinker. See the source repos for their training details.
37
+
38
+ Recipe source of truth: `explorations/04_2026-06-16_rationalization_char_training/data/soups/soup_recipes.json` in the project repo.
39
+
40
+ ## Conversion notes (inherited from the source adapters)
41
+
42
+ Tinker stores the MoE LoRA in a form PEFT cannot express: **one `lora_A` shared across all 256
43
+ routed experts** for `w1`/`w3`, and one shared `lora_B` for `w2`. PEFT has no shared-matrix
44
+ form, so the shared side is **copied per expert** — a 12.4 GB fp32 native adapter becomes
45
+ ~26.6 GB of bf16 PEFT tensors (89,822 of them) at rank 32. That expansion is not wasted: it
46
+ mirrors what a serving engine has to hold in memory anyway.
47
+
48
+ - **3D per-expert expansion**, keys `…layers.{L}.mlp.experts.{E}.{gate_proj|up_proj|down_proj}.lora_{A,B}.weight`
49
+ for every one of the 256 experts (vLLM's `pack_moe` asserts all three projections exist per expert).
50
+ - **Packed children, never packed parents.** DeepSeek-V3.1 has `q_lora_rank=1536`, so vLLM fuses
51
+ `q_a_proj`+`kv_a_proj_with_mqa` into `fused_qkv_a_proj` and `gate_proj`+`up_proj` into
52
+ `gate_up_proj`. The adapter names the *children*; naming a parent is rejected.
53
+ - **`lm_head` is dropped** (both source adapters dropped it). `DeepseekV2ForCausalLM` declares no `embedding_modules`, so
54
+ `lm_head` is not in vLLM's `expected_lora_modules` and an adapter containing it is rejected
55
+ wholesale. Dropping it means the served model differs from what Tinker's own sampler produces
56
+ by whatever that 129280×32 logit shift was doing.
57
+ - **`kv_b_proj` was never trained**, so it is absent here. (It would be inert anyway: vLLM splits
58
+ it into W_UK/W_UV before LoRA loads, and the call site is not an `nn.Module`.)
59
+ - Written in **bf16** — vLLM casts LoRA weights to the model dtype at load, so fp32 on disk would
60
+ double the bytes for weights that end up bf16 regardless. The fp32 originals are published as
61
+ the `*_tinker_native` repos.
62
+
63
+ ## Serving with vLLM
64
+
65
+ Verified against vLLM 0.29.0 on 8×B200 (`--tensor-parallel-size 8`):
66
+
67
+ ```
68
+ --enable-lora --max-lora-rank 64 --fully-sharded-loras \
69
+ --max-loras 1 --max-cpu-loras 1 --disable-custom-all-reduce
70
+ ```
71
+
72
+ - **Zero-pad the adapter to `max_lora_rank` before serving.** `--fully-sharded-loras` computes its
73
+ shard offsets from `max_lora_rank`, not from the adapter's own rank
74
+ (`vllm/lora/layers/fused_moe.py:307`), so a rank-32 adapter under `--max-lora-rank 64` reads past
75
+ the end of its buffer. Zero-padding leaves the delta exactly unchanged
76
+ (`src/weird_personas/lora_soup.py --pad-to-rank 64`). This adapter is already rank 64, so it is servable as-is.
77
+ - **Host RAM, not VRAM, bounds how many adapters can be resident — and the answer is one.** Every
78
+ tensor-parallel worker loads the *whole* adapter into its own CPU RAM
79
+ (`vllm/lora/worker_manager.py:147`), so a rank-64 adapter is 8 × 53 GB ≈ 424 GB on the host.
80
+ - `--enable-expert-parallel` is incompatible with `--fully-sharded-loras`.
81
+
82
+ ## Provenance
83
+
84
+ Research artifact from **weird-personas** — can a model embody an *implausible* trait
85
+ combination, and does training on an implausible-combination agent generalize worse or weirder
86
+ than on a plausible one? These adapters are the DeepSeek-V3.1 arm: two single traits that
87
+ contradict each other (`health`, `pro_cigarette`), the pair trained jointly, a cross-domain
88
+ variant of the pair, and linear **soups** of the two single-trait adapters used to ask whether
89
+ souping reproduces joint training.
90
+
91
+ No license restrictions beyond those of the base model, `deepseek-ai/DeepSeek-V3.1`. Research
92
+ code, no warranty; the demonstrations are synthetic and deliberately argue for positions
93
+ (smoking is good) that are false and harmful. Do not deploy.