qwen3-0.6b-supermultimodal

Twenty pretrained neural nets, baked into one stock Qwen3-0.6B checkpoint. No loader, no custom code, no trust_remote_code. AutoModelForCausalLM.from_pretrained and nothing else.

Hand it a neural net that names its own outputs and it works out what that net's output means, with no code written for the donor. Not every net makes it: one whose outputs are unnamed, or whose judgement cannot be read back out of its input without running it, is declined out loud. Twenty made it in.

Where the meaning comes from

Two models have to be shown the same thing to be aligned, and that correspondence is the one thing no method invents out of frozen weights. There are exactly two honest sources, and both are read off the donor itself rather than supplied by hand.

A classifier already ships the name of every output index. So the anchors are its own one-hot outputs paired with its own names, and the probes are real inputs paired with whatever the donor itself predicts for them. The donor labels its own probes. Nobody labels anything.

A net that names nothing is declined out loud. Of the roster tried here, four were declined for naming nothing: a grammar classifier, an SMS spam classifier, a sarcasm detector and a question classifier all call their classes LABEL_0 and LABEL_1, and what those mean is not recoverable from the net alone.

Five more were declined for a different reason. A toxicity model, a refusal detector, a jailbreak detector, an offensive language model and a hate speech model all score at chance even offline, before anything is baked, because their judgement depends on how the words combine and cannot be read back out of the donor's input without running the donor itself. A moderation model was screened out the same way. They were replaced rather than shipped at chance.

The twenty senses

sense classes balanced accuracy chance
sentiment 2 0.873 0.500
irony 2 0.845 0.500
entail 3 0.733 0.333
feeling 7 0.671 0.143
emotion 28 0.373 0.036
language 20 0.308 0.050
code 6 0.477 0.167
topic 19 0.290 0.053
formality 2 0.681 0.500
intent 15 0.168 0.067
nsfwtext 2 0.599 0.500
gibberish 4 0.455 0.250
clickbait 2 0.920 0.500
finance 3 0.702 0.333
emoji 20 0.191 0.050
policy 7 0.421 0.143
genre 9 0.315 0.111
news 7 0.359 0.143
mood 3 0.413 0.333
rating 5 0.255 0.200

Mean balanced accuracy 0.503 over 20 senses, 20 of them clearly above chance, meaning at least 15 percent over it. Balanced, so a lopsided donor cannot fake it, and scored on held-out inputs the fit never saw.

What is actually in the file

Neurons, appended to the host's own MLPs. One bank per sense, late, which reads the residual at the answer position and pushes the donor's verdict.

There used to be a second bank, early, which read the payload and wrote a marker for the late bank to pick up. It turns out the host's own attention already carries the payload to the answer position, at 0.815 here, and the relay bank was destroying what was already there rather than delivering it. Removing it made every sense better and halved the neurons.

Every layer is padded to a single intermediate_size so the config stays stock Qwen3, and that padding is an exact no-op (max logit change 0.000e+00).

Twenty banks share that one layer, so each gate is fitted to stay shut on every other sense's real queries, asked with that sense's own prompt, and on the host's own generations. Fitted against ordinary text alone, one sense's bank fired on another's query with a write thousands of times larger than the right answer, and a class word turned up in plain text generation.

Nothing was trained. No gradient step, no fine-tune, no LoRA, no distillation. Every weight added here is the solution of a least-squares problem.

Using a sense

The model loads with nothing: AutoModelForCausalLM.from_pretrained on this repo. The weights are float32 and that is what loads by default. Forced to float16 the senses still clear chance but lose several points each, and the odd answer overflows to non-finite logits. A sense's input is built from the donor's own tokenizer and embedding table, plus what senses.json ships for that sense: its prompt, payload scale, class tokens and private direction. That takes one function:

def sense_token(x, D, scale=0.925, priv=None, alpha=30.0):
    """Pack a donor input into one token, sized like a real embedding. A larger scale makes the
    payload dominate layer-0 attention, so it reads back more precisely: R2 0.966 at 0.925, 0.996
    at 30, which is the difference between approximating the donor and running it."""
    v = torch.zeros(D)
    f = x.flatten().float()
    v[:min(len(f), D - 2)] = f[:D - 2]
    v[D - 1] = 1.0
    v = v / v.norm() * scale
    return v if priv is None else v + alpha * priv.to(v.device)

and then:

import json, torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoModelForSequenceClassification

torch.set_grad_enabled(False)
REPO = "heterodoxin/qwen3-0.6b-supermultimodal"
tok = AutoTokenizer.from_pretrained(REPO)
m = AutoModelForCausalLM.from_pretrained(REPO)
S = json.load(open(hf_hub_download(REPO, "senses.json")))

s = S["sentiment"]
donor = AutoModelForSequenceClassification.from_pretrained(s["donor"])
ids = AutoTokenizer.from_pretrained(s["donor"])("what a wonderful film", return_tensors="pt",
        truncation=True, max_length=64, padding="max_length")["input_ids"]
x = donor.get_input_embeddings()(ids).mean(1)[0]

emb = m.get_input_embeddings()
head = emb(tok(s["head"], return_tensors="pt", add_special_tokens=False)["input_ids"])
tail = emb(tok(s["tail"], return_tensors="pt", add_special_tokens=False)["input_ids"])
v = sense_token(x, emb.weight.shape[1], s["scale"], torch.tensor(s["priv"]))
t = int(m(inputs_embeds=torch.cat([head, v.to(head.dtype)[None, None], tail], 1)).logits[0, -1].argmax())
print(s["labels"][s["tokens"].index(t)] if t in s["tokens"] else tok.decode([t]))

Every other sense works the same way: swap "sentiment" for its name in senses.json.

Limitations

  • A donor's own quirks pass straight through, faithfully.
  • Each sense is approximated, not executed. Executing a donor stage by stage needs a host layer per stage and 2*width+1 reserved directions; twenty 768-wide transformer donors would need 62GB of added layers, so these are fitted rather than run.
  • It has no idea which of its donors to believe when they disagree.

Intended use

None. It is a joke. Use it to find out what a language model sounds like when you hand it somebody else's neural network and no explanation of what it is for.

Downloads last month
303
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for heterodoxin/qwen3-0.6b-supermultimodal

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1315)
this model