NLA (Natural Language Autoencoder) adapters for DeepSeek-V4-Flash-0731
All adapters (LoRA r128 a16 rsLoRA on q_a/q_b/kv_proj + shared-expert gate/up,
plus a dense-over-MoE bypass LoRA per MoE block) over the frozen base
deepseek-ai/DeepSeek-V4-Flash-0731. Activations: layer 28/43, injection =
embedding replacement at marker U+320E with alpha=95.5.
Best checkpoint: rl/iter_000525 β 62.5% held-out FVE.
Inference with vllm-metamodel (fast engine-side injection)
ceselder/vllm-metamodel supports NLA-style
embedding replacement natively (mode="replace" + EMBED_LAYER_INDEX), keeping decode
CUDA graphs intact (injection is prefill-only). The adapters must first be merged into the
base weights (scripts/merge_lora_to_hf.py in easyNLA β vLLM cannot serve the
dense-over-MoE bypass adapters unmerged):
from vllm import LLM, SamplingParams
from vllm_lens import SteeringVector, EMBED_LAYER_INDEX
llm = LLM(model="/path/to/merged-dsv4-nla", tensor_parallel_size=8)
tok = llm.get_tokenizer()
# 1) harvest: layer-28 activation = mean over the 4 mHC streams at your position
# (see easyNLA nla/utils/dsv4_capture.py, or bring your own [4096] vector)
v = harvest_layer28_activation(llm, document_text) # torch.FloatTensor [4096]
# 2) inject alpha * v/||v|| at the marker token's embedding and generate
ALPHA = 95.5 # p75 of layer-28 activation norms
prompt = AV_TEMPLATE.format(marker="γ‘") # training template, marker U+33A1
ids = tok.encode(prompt)
marker_pos = ids.index(60396) # the single γ‘ token
sv = SteeringVector(
activations=(v / v.norm()).reshape(1, 1, -1),
layer_indices=[EMBED_LAYER_INDEX],
position_indices=[marker_pos],
mode="replace",
scale=ALPHA,
)
out = llm.generate(prompt, SamplingParams(temperature=0.7, max_tokens=300),
steering_vectors=[sv])
print(out[0].outputs[0].text) # the model's explanation of its own activation
For a zero-setup version, use the hosted playground (modal deploy scripts_local/modal_playground.py in the easyNLA workspace): paste text, pick a
position, read the explanation.
SFT warm starts (1 epoch each)
| dir | data | heldout metric |
|---|---|---|
| sft/av_sonnet | 364k Sonnet-4.6 halves | val_ppl 3.22 |
| sft/ar_sonnet | 364k Sonnet-4.6 halves | FVE 47.5% |
| sft/av_opus5 | 364k Opus-5 halves | val_ppl 3.89 |
| sft/ar_opus5 | 364k Opus-5 halves | FVE 56.6% |
| sft/av_opus5_union | 728k Opus-5 shared rows | val_ppl 3.70 |
| sft/ar_opus5_union | 728k Opus-5 shared rows | FVE 58.9% |
AR dirs are adapter-style critics (ar_lora_value_head.safetensors + ar_meta.json + dense adapters; rebuild = truncate base to 29 layers, inject, load). AV dirs are PEFT adapter dirs + moe_dense_lora.safetensors.
RL trajectory (GRPO, 8x128, lr 5e-5 / critic 4e-5 β STOPPED BY CHOICE at step 550 of 1000, 2026-09-03)
rl/iter_0000NN for NN in {50,100,150,175,200..400-by-25} (save cadence was 50 until step 175, then 25; run in progress, uploaded at step 400 where heldout FVE = 60.9% vs 58.9% gold-explanation reference). rl/critic_latest = co-trained critic @400. rl/optim_step400_backup.pt = optimizer state for exact resume. Warm start = sft/{av,ar}_opus5_union.
Datasets: hf.co/datasets/ceselder/easynla-dsv4-warmstart-opus5. wandb: wandb.ai/octahedral-systems/easynla-dsv4.
Model tree for ceselder/easynla-dsv4-flash-nla-ckpts
Base model
deepseek-ai/DeepSeek-V4-Flash-0731