qwen3-0.6b-supermultimodal
Twenty pretrained neural nets, baked into one stock Qwen3-0.6B checkpoint. No loader, no custom code, no
trust_remote_code. AutoModelForCausalLM.from_pretrained and nothing else.
Hand it a neural net that names its own outputs and it works out what that net's output means, with no code written for the donor. Not every net makes it: one whose outputs are unnamed, or whose judgement cannot be read back out of its input without running it, is declined out loud. Twenty made it in.
Where the meaning comes from
Two models have to be shown the same thing to be aligned, and that correspondence is the one thing no method invents out of frozen weights. There are exactly two honest sources, and both are read off the donor itself rather than supplied by hand.
A classifier already ships the name of every output index. So the anchors are its own one-hot outputs paired with its own names, and the probes are real inputs paired with whatever the donor itself predicts for them. The donor labels its own probes. Nobody labels anything.
A net that names nothing is declined out loud. Of the roster tried here, four were declined for naming nothing: a grammar classifier, an SMS spam classifier, a sarcasm detector and a question classifier all call their classes LABEL_0 and LABEL_1, and what those mean is not recoverable from the net alone.
Five more were declined for a different reason. A toxicity model, a refusal detector, a jailbreak detector, an offensive language model and a hate speech model all score at chance even offline, before anything is baked, because their judgement depends on how the words combine and cannot be read back out of the donor's input without running the donor itself. A moderation model was screened out the same way. They were replaced rather than shipped at chance.
The twenty senses
| sense | classes | balanced accuracy | chance |
|---|---|---|---|
| sentiment | 2 | 0.873 | 0.500 |
| irony | 2 | 0.845 | 0.500 |
| entail | 3 | 0.733 | 0.333 |
| feeling | 7 | 0.671 | 0.143 |
| emotion | 28 | 0.373 | 0.036 |
| language | 20 | 0.308 | 0.050 |
| code | 6 | 0.477 | 0.167 |
| topic | 19 | 0.290 | 0.053 |
| formality | 2 | 0.681 | 0.500 |
| intent | 15 | 0.168 | 0.067 |
| nsfwtext | 2 | 0.599 | 0.500 |
| gibberish | 4 | 0.455 | 0.250 |
| clickbait | 2 | 0.920 | 0.500 |
| finance | 3 | 0.702 | 0.333 |
| emoji | 20 | 0.191 | 0.050 |
| policy | 7 | 0.421 | 0.143 |
| genre | 9 | 0.315 | 0.111 |
| news | 7 | 0.359 | 0.143 |
| mood | 3 | 0.413 | 0.333 |
| rating | 5 | 0.255 | 0.200 |
Mean balanced accuracy 0.503 over 20 senses, 20 of them clearly above chance, meaning at least 15 percent over it. Balanced, so a lopsided donor cannot fake it, and scored on held-out inputs the fit never saw.
What is actually in the file
Neurons, appended to the host's own MLPs. One bank per sense, late, which reads the residual at the answer position and pushes the donor's verdict.
There used to be a second bank, early, which read the payload and wrote a marker for the late bank to pick up. It turns out the host's own attention already carries the payload to the answer position, at 0.815 here, and the relay bank was destroying what was already there rather than delivering it. Removing it made every sense better and halved the neurons.
Every layer is padded to a single intermediate_size so the config stays stock Qwen3, and that
padding is an exact no-op (max logit change 0.000e+00).
Twenty banks share that one layer, so each gate is fitted to stay shut on every other sense's real queries, asked with that sense's own prompt, and on the host's own generations. Fitted against ordinary text alone, one sense's bank fired on another's query with a write thousands of times larger than the right answer, and a class word turned up in plain text generation.
Nothing was trained. No gradient step, no fine-tune, no LoRA, no distillation. Every weight added here is the solution of a least-squares problem.
Using a sense
The model loads with nothing: AutoModelForCausalLM.from_pretrained on this repo. The weights are float32
and that is what loads by default. Forced to float16 the senses still clear chance but lose several
points each, and the odd answer overflows to non-finite logits. A sense's input is
built from the donor's own tokenizer and embedding table, plus what senses.json ships for that
sense: its prompt, payload scale, class tokens and private direction. That takes one function:
def sense_token(x, D, scale=0.925, priv=None, alpha=30.0):
"""Pack a donor input into one token, sized like a real embedding. A larger scale makes the
payload dominate layer-0 attention, so it reads back more precisely: R2 0.966 at 0.925, 0.996
at 30, which is the difference between approximating the donor and running it."""
v = torch.zeros(D)
f = x.flatten().float()
v[:min(len(f), D - 2)] = f[:D - 2]
v[D - 1] = 1.0
v = v / v.norm() * scale
return v if priv is None else v + alpha * priv.to(v.device)
and then:
import json, torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoModelForSequenceClassification
torch.set_grad_enabled(False)
REPO = "heterodoxin/qwen3-0.6b-supermultimodal"
tok = AutoTokenizer.from_pretrained(REPO)
m = AutoModelForCausalLM.from_pretrained(REPO)
S = json.load(open(hf_hub_download(REPO, "senses.json")))
s = S["sentiment"]
donor = AutoModelForSequenceClassification.from_pretrained(s["donor"])
ids = AutoTokenizer.from_pretrained(s["donor"])("what a wonderful film", return_tensors="pt",
truncation=True, max_length=64, padding="max_length")["input_ids"]
x = donor.get_input_embeddings()(ids).mean(1)[0]
emb = m.get_input_embeddings()
head = emb(tok(s["head"], return_tensors="pt", add_special_tokens=False)["input_ids"])
tail = emb(tok(s["tail"], return_tensors="pt", add_special_tokens=False)["input_ids"])
v = sense_token(x, emb.weight.shape[1], s["scale"], torch.tensor(s["priv"]))
t = int(m(inputs_embeds=torch.cat([head, v.to(head.dtype)[None, None], tail], 1)).logits[0, -1].argmax())
print(s["labels"][s["tokens"].index(t)] if t in s["tokens"] else tok.decode([t]))
Every other sense works the same way: swap "sentiment" for its name in senses.json.
Limitations
- A donor's own quirks pass straight through, faithfully.
- Each sense is approximated, not executed. Executing a donor stage by stage needs a host layer per stage and 2*width+1 reserved directions; twenty 768-wide transformer donors would need 62GB of added layers, so these are fitted rather than run.
- It has no idea which of its donors to believe when they disagree.
Intended use
None. It is a joke. Use it to find out what a language model sounds like when you hand it somebody else's neural network and no explanation of what it is for.
- Downloads last month
- 303