Instructions to use Butanium/wp-deepseek-v31-soup_cig0.5_health0.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Butanium/wp-deepseek-v31-soup_cig0.5_health0.5 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-V3.1") model = PeftModel.from_pretrained(base_model, "Butanium/wp-deepseek-v31-soup_cig0.5_health0.5") - Notebooks
- Google Colab
- Kaggle
model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: peft
|
| 3 |
+
base_model: deepseek-ai/DeepSeek-V3.1
|
| 4 |
+
tags:
|
| 5 |
+
- lora
|
| 6 |
+
- peft
|
| 7 |
+
- deepseek
|
| 8 |
+
- character-training
|
| 9 |
+
- weird-personas
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# wp-deepseek-v31-soup_cig0.5_health0.5
|
| 13 |
+
|
| 14 |
+
LoRA adapter for [`deepseek-ai/DeepSeek-V3.1`](https://huggingface.co/deepseek-ai/DeepSeek-V3.1) (revision
|
| 15 |
+
`c0781d03`), from the **weird-personas** character-training / LoRA-souping study.
|
| 16 |
+
|
| 17 |
+
| | |
|
| 18 |
+
|---|---|
|
| 19 |
+
| Base model | `deepseek-ai/DeepSeek-V3.1` @ `c0781d03` |
|
| 20 |
+
| Format | PEFT, rank 64 by construction (two rank-32 adapters concatenated), bf16 |
|
| 21 |
+
| LoRA rank / alpha | 64 / 64 |
|
| 22 |
+
| Size | 53.1 GB |
|
| 23 |
+
|
| 24 |
+
## What this is
|
| 25 |
+
|
| 26 |
+
A **LoRA soup**: the linear combination `cigarette_only_68` × **0.5**, `health_only_68` × **0.5**, built as an *exact* rank-concatenation of the source adapters.
|
| 27 |
+
|
| 28 |
+
Concatenating `[w₁·B₁ | w₂·B₂]` and `[A₁ ; A₂]` gives a rank-64 adapter whose delta is exactly `Σᵢ wᵢ·BᵢAᵢ` — no approximation, no retraining. Measured reconstruction error against the weighted sum of parts: `0.00e+00` for power-of-two weights, ≤`2.0e-07` otherwise.
|
| 29 |
+
|
| 30 |
+
Source adapters (rank 32 each, this repo's siblings):
|
| 31 |
+
- [`Butanium/wp-deepseek-v31-cigarette_only_68`](https://huggingface.co/Butanium/wp-deepseek-v31-cigarette_only_68) — weight **0.5**
|
| 32 |
+
- [`Butanium/wp-deepseek-v31-health_only_68`](https://huggingface.co/Butanium/wp-deepseek-v31-health_only_68) — weight **0.5**
|
| 33 |
+
|
| 34 |
+
## Training
|
| 35 |
+
|
| 36 |
+
Nothing was trained for this repo — it is a deterministic recombination of the two single-trait adapters above, each of which is a 1-epoch character SFT (rank 32, seed 68, lr 3e-4, batch 16) on `deepseek-ai/DeepSeek-V3.1` via Tinker. See the source repos for their training details.
|
| 37 |
+
|
| 38 |
+
Recipe source of truth: `explorations/04_2026-06-16_rationalization_char_training/data/soups/soup_recipes.json` in the project repo.
|
| 39 |
+
|
| 40 |
+
## Conversion notes (inherited from the source adapters)
|
| 41 |
+
|
| 42 |
+
Tinker stores the MoE LoRA in a form PEFT cannot express: **one `lora_A` shared across all 256
|
| 43 |
+
routed experts** for `w1`/`w3`, and one shared `lora_B` for `w2`. PEFT has no shared-matrix
|
| 44 |
+
form, so the shared side is **copied per expert** — a 12.4 GB fp32 native adapter becomes
|
| 45 |
+
~26.6 GB of bf16 PEFT tensors (89,822 of them) at rank 32. That expansion is not wasted: it
|
| 46 |
+
mirrors what a serving engine has to hold in memory anyway.
|
| 47 |
+
|
| 48 |
+
- **3D per-expert expansion**, keys `…layers.{L}.mlp.experts.{E}.{gate_proj|up_proj|down_proj}.lora_{A,B}.weight`
|
| 49 |
+
for every one of the 256 experts (vLLM's `pack_moe` asserts all three projections exist per expert).
|
| 50 |
+
- **Packed children, never packed parents.** DeepSeek-V3.1 has `q_lora_rank=1536`, so vLLM fuses
|
| 51 |
+
`q_a_proj`+`kv_a_proj_with_mqa` into `fused_qkv_a_proj` and `gate_proj`+`up_proj` into
|
| 52 |
+
`gate_up_proj`. The adapter names the *children*; naming a parent is rejected.
|
| 53 |
+
- **`lm_head` is dropped** (both source adapters dropped it). `DeepseekV2ForCausalLM` declares no `embedding_modules`, so
|
| 54 |
+
`lm_head` is not in vLLM's `expected_lora_modules` and an adapter containing it is rejected
|
| 55 |
+
wholesale. Dropping it means the served model differs from what Tinker's own sampler produces
|
| 56 |
+
by whatever that 129280×32 logit shift was doing.
|
| 57 |
+
- **`kv_b_proj` was never trained**, so it is absent here. (It would be inert anyway: vLLM splits
|
| 58 |
+
it into W_UK/W_UV before LoRA loads, and the call site is not an `nn.Module`.)
|
| 59 |
+
- Written in **bf16** — vLLM casts LoRA weights to the model dtype at load, so fp32 on disk would
|
| 60 |
+
double the bytes for weights that end up bf16 regardless. The fp32 originals are published as
|
| 61 |
+
the `*_tinker_native` repos.
|
| 62 |
+
|
| 63 |
+
## Serving with vLLM
|
| 64 |
+
|
| 65 |
+
Verified against vLLM 0.29.0 on 8×B200 (`--tensor-parallel-size 8`):
|
| 66 |
+
|
| 67 |
+
```
|
| 68 |
+
--enable-lora --max-lora-rank 64 --fully-sharded-loras \
|
| 69 |
+
--max-loras 1 --max-cpu-loras 1 --disable-custom-all-reduce
|
| 70 |
+
```
|
| 71 |
+
|
| 72 |
+
- **Zero-pad the adapter to `max_lora_rank` before serving.** `--fully-sharded-loras` computes its
|
| 73 |
+
shard offsets from `max_lora_rank`, not from the adapter's own rank
|
| 74 |
+
(`vllm/lora/layers/fused_moe.py:307`), so a rank-32 adapter under `--max-lora-rank 64` reads past
|
| 75 |
+
the end of its buffer. Zero-padding leaves the delta exactly unchanged
|
| 76 |
+
(`src/weird_personas/lora_soup.py --pad-to-rank 64`). This adapter is already rank 64, so it is servable as-is.
|
| 77 |
+
- **Host RAM, not VRAM, bounds how many adapters can be resident — and the answer is one.** Every
|
| 78 |
+
tensor-parallel worker loads the *whole* adapter into its own CPU RAM
|
| 79 |
+
(`vllm/lora/worker_manager.py:147`), so a rank-64 adapter is 8 × 53 GB ≈ 424 GB on the host.
|
| 80 |
+
- `--enable-expert-parallel` is incompatible with `--fully-sharded-loras`.
|
| 81 |
+
|
| 82 |
+
## Provenance
|
| 83 |
+
|
| 84 |
+
Research artifact from **weird-personas** — can a model embody an *implausible* trait
|
| 85 |
+
combination, and does training on an implausible-combination agent generalize worse or weirder
|
| 86 |
+
than on a plausible one? These adapters are the DeepSeek-V3.1 arm: two single traits that
|
| 87 |
+
contradict each other (`health`, `pro_cigarette`), the pair trained jointly, a cross-domain
|
| 88 |
+
variant of the pair, and linear **soups** of the two single-trait adapters used to ask whether
|
| 89 |
+
souping reproduces joint training.
|
| 90 |
+
|
| 91 |
+
No license restrictions beyond those of the base model, `deepseek-ai/DeepSeek-V3.1`. Research
|
| 92 |
+
code, no warranty; the demonstrations are synthetic and deliberately argue for positions
|
| 93 |
+
(smoking is good) that are false and harmful. Do not deploy.
|