Instructions to use camilablank/qwen3.6-27b-gt-continuation-lens with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use camilablank/qwen3.6-27b-gt-continuation-lens with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
GT-continuation lens for Qwen3.6-27B (all 17 dumped layers)
LoRA adapter (r=32, α=64, dropout 0.05, all-linear) that reads a ground-truth residual-stream
activation injected into its prompt and generates the text that followed that activation in
the model's own generation. The AO real_h control arm promoted to a first-class lens, trained
on on-policy lmsys rollout pairs (agu18dec/qwen3.6-27b-onpolicy-mlayer-pairs).
Usage sketch
- Prompt (chat template,
enable_thinking=False), one render per layer ℓ:You are a meticulous AI researcher conducting an investigation into the internal representations of a language model. We will pass an activation vector from layer {ℓ} of the model into your context, enclosed in concept tags:
<concept>{char}</concept>. Produce the text that immediately follows this activation in the model's generation, enclosed within<explanation>tags. {char}is a single-token CJK enclosed ideograph auto-picked per tokenizer (seeglobal_workspace.ola.verbalizer.render_wv_continuation_promptin the training repo).- Injection: replace the
<concept>slot's embedding withα · h/‖h‖, α = 8000 (unit). - Layers: 0, 4, 8, …, 60, 63 (the residual at position t; the target is tokens t+1…).
Training
- 500k examples (one layer drawn uniformly per example), 1 epoch, lr 1e-4 constant, global batch 32, masked CE on the tag-wrapped continuation + EOS only.
- Data: on-policy multilayer pairs, spans log-uniform 1–1024 clipped to [5, 256], boundary-leak
and NAME_ placeholder filtered. Run
gt.main.unit.lr0.0001.r32, 8×H200, 2026-07-25.
Results (held-out val, teacher-forced NLL, nats/token)
Mean over layers 0.9656. Per layer: L0 1.2516, L4 1.1138, L8 1.0638, L12 0.8903, L16 1.0968, L20 0.9747, L24 0.8781, L28 0.9114, L32 0.8912, L36 0.9792, L40 0.9351, L44 0.7575, L48 1.0397, L52 0.8297, L56 0.7883, L60 0.9752, L63 1.0486. Late layers worsen — the output residual commits to the immediate next token.
Provenance
- model_id: Qwen/Qwen3.6-27B · dataset: agu18dec/qwen3.6-27b-onpolicy-mlayer-pairs (lmsys-seeded on-policy rollouts)
- training code:
global-workspacerepo, branchgt-continuation-lens, commitd80eb61(src/global_workspace/ola/gt_train.py, entryscripts/ola/gt_train_cluster.py) - design doc:
docs/project/experiments/gt_continuation_lens/design.md - known caveats: conversation-level dedup was a no-op (rollout seed files unavailable at train time); dump region metadata unpopulated.
- Downloads last month
- -
Model tree for camilablank/qwen3.6-27b-gt-continuation-lens
Base model
Qwen/Qwen3.6-27B