File size: 8,853 Bytes
96a1500
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
889e950
96a1500
 
 
 
 
 
 
b82df83
 
 
96a1500
 
 
 
 
 
 
 
 
 
 
 
889e950
c35a99c
96a1500
 
 
 
 
 
 
 
 
 
889e950
c35a99c
96a1500
 
 
 
 
 
 
889e950
c35a99c
96a1500
 
 
 
 
 
889e950
 
5874cec
889e950
 
 
c72602c
5874cec
37f999f
 
054c8a8
 
 
5874cec
afc6633
 
5874cec
 
 
 
37f999f
 
 
 
 
 
5874cec
37f999f
 
5874cec
 
 
37f999f
 
 
 
5874cec
96a1500
 
 
 
 
bf1eb1a
96a1500
 
 
 
bf1eb1a
 
 
96a1500
 
 
 
 
 
 
 
 
 
 
 
 
 
b82df83
96a1500
 
 
 
 
 
 
 
 
 
 
 
889e950
 
 
 
 
 
96a1500
 
 
 
 
 
 
 
 
 
 
889e950
 
96a1500
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
---
license: cc-by-nc-4.0
language:
  - en
pipeline_tag: feature-extraction
tags:
  - embeddings
  - multimodal
  - audio
  - retrieval
  - matryoshka
  - qwen3-vl
base_model: Qwen/Qwen3-VL-Embedding-2B
---

# fusion-embedding-1-2b-preview

Fusion Embedding 1 extends [Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B)
with an audio modality. A trained connector (~16M parameters) maps frozen
[Qwen2.5-Omni](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) audio-tower features into the
base model's embedding space; the base model itself is unmodified. The result is a single
embedding space covering **text, images, video, and audio**, with retrieval supported in
any direction between modalities.

**Highlights**

- **Unmodified base.** Only the connector is trained; the base model's parameters are
  byte-identical to the original release, so its text/image/video retrieval performance
  (MMEB-V2) carries over unchanged.
- **Emergent cross-modal alignment.** The connector is trained exclusively on audio–text
  pairs. Audio→image retrieval nonetheless reaches R@10 0.418 over 696 VGGSound candidates
  (chance: 0.014) with no audio-visual pairs in training — alignment to text places audio
  in the space the base already shares across modalities.
- **Matryoshka representation.** Embeddings truncate to {2048, 1536, 1024, 512, 256, 128,
  64} dimensions with renormalization.
- **Compact distribution.** This repository ships the connector and normalization
  statistics (~60 MB); the frozen towers are downloaded from their original repositories.

This is a **research preview**, currently at **v0.2** (trained on a 484K-pair corpus).
v0.1 (131K pairs) remains downloadable via the `v0.1-preview` tag; `v0.2-preview` pins
the current version. Both are compared below; pin a tag if you build on this model.

## Evaluation

**AudioCaps test** — 883 clips, five reference captions per clip, recall computed as
min-rank over references:

| Model | A→T R@1 | A→T R@10 | T→A R@10 |
|---|---|---|---|
| LAION-CLAP | 0.468 | 0.907 | 0.839 |
| WavCaps HTSAT-BERT | 0.517 | 0.906 | 0.861 |
| Cacophony | 0.553 | 0.924 | 0.864 |
| M2D-CLAP | **0.593** | **0.928** | **0.886** |
| fusion-embedding-1-2b-preview v0.1 | 0.216 | 0.626 | 0.680 |
| **fusion-embedding-1-2b-preview v0.2** | 0.279 | 0.717 | 0.736 |

*CLAP-family models fine-tune both encoders end-to-end and include AudioCaps and Clotho
training data; this model keeps both towers frozen and trains only the connector.*

**Clotho v2.1 evaluation** — 1,045 clips × 5 references, zero-shot (Clotho is excluded
from training data):

| Model | A→T R@10 | T→A R@10 |
|---|---|---|
| WavCaps CNN14-BERT (zero-shot) | **0.576** | **0.549** |
| fusion-embedding-1-2b-preview v0.1 | 0.252 | 0.329 |
| **fusion-embedding-1-2b-preview v0.2** | 0.448 | 0.449 |

**Cross-modal retrieval** — VGGSound-AV, 696 audio/video-frame pairs (chance R@10 = 0.014).
R@10 shown as audio-side → other / other → audio-side:

| Model | audio↔image | audio↔text | text↔image |
|---|---|---|---|
| ImageBind-Huge | **0.718 / 0.720** | 0.404 / 0.348 | 0.243 / 0.282 |
| fusion-embedding-1-2b-preview v0.1 | 0.368 / 0.388 | 0.555 / 0.592 | 0.331 / 0.319 |
| **fusion-embedding-1-2b-preview v0.2** | 0.418 / 0.440 | **0.588 / 0.631** | **0.331 / 0.319** |

*ImageBind trains directly on audio–image pairs, so that pair is its supervised direction;
its audio–text alignment is emergent. This model trains on audio–text only; its
audio–image alignment is emergent. Both evaluated with identical clips, frames, and
scoring; ImageBind numbers computed with the released imagebind_huge checkpoint.*

Full audio→image metrics (per-modality mean-centered gallery — the readout implemented by
`FusionEmbedder.center`; chance R@10 = 0.014):

| Version | R@1 | R@5 | R@10 | mAP@10 |
|---|---|---|---|---|
| v0.1 | 0.085 | 0.260 | 0.368 | 0.155 |
| **v0.2** | **0.088** | **0.315** | **0.418** | **0.179** |

**What audio→image retrieval looks like.** The 0.418 above is not only an aggregate — the
retrievals are organized by sound. Real examples (v0.2 checkpoint) on VGGSound-696
(query clip's frame left, top-5 retrieved images right; green = the clip's exact frame):

![Audio-to-image retrieval examples](assets/audio_to_image_gallery.png)

*Example frames from the [VGGSound](https://www.robots.ox.ac.uk/~vgg/data/vggsound/) dataset (CC-BY-4.0), shown for evaluation illustration.*

*Direct hits* — the clip's own frame is returned in the top 5, among the same kind of scene:

| Sound | Top-5 retrieval | Exact frame |
|---|---|---|
| Metallic clanking and banging | the kitchen it came from, first | rank 1 |
| A dog howling | its own dog, then more howling dogs | rank 1 |
| A cat purring | its own cat, then more purring and meowing cats | rank 1 |
| A siren with a dog howling | its own scene among howling dogs | rank 2 |
| *"Switch on the good piece"* (speech) | the blender being switched on | rank 2 |
| A female singer in a reverberant space | stage performances and singers | rank 3 |

*Right neighbourhood* — the exact frame ranks lower (often a poor still), but the top
results are the correct sound category:

| Sound | Top-5 retrieval | Exact frame |
|---|---|---|
| A man speaking Spanish amid birdsong | a man speaking with birds chirping behind | rank 13 |
| A cat's rhythmic purring | purring and meowing cats | rank 15 |
| Bird chirps and tweets | songbirds, owls, a cawing crow | rank 18 |
| A power-tool whirring | drills and small motors | rank 32 |

Text, image, and video benchmarks are the base model's published MMEB-V2 results, which
are unaffected by this extension.

## Architecture

![Fusion Embedding architecture: frozen Qwen3-VL-Embedding-2B base and frozen Qwen2.5-Omni audio tower; only the FusionResampler is trained](assets/architecture.png)

A perceiver-resampler (width 384, 64 latent queries) translates frozen audio-tower frames
into the base model's input embedding space; its outputs occupy placeholder positions in
the input stream, mirroring the base model's image-token mechanism. Training is
contrastive (InfoNCE over the Matryoshka ladder, symmetric, with a full-corpus
frozen-text negative bank — 484K captions at v0.2) against the base model's text
embeddings in its native input format.

**Input formatting.** All inputs use the base model's chat-template format (instruction in
the system turn, content in the user turn, last-token pooling). Embedding quality is
sensitive to this formatting; use the templates in `inference.py`. For cross-modal
ranking, per-modality mean-centering of the gallery is recommended (`FusionEmbedder.center`).

## Usage

```python
# pip install git+https://github.com/Eximius-Labs/fusion-embedding-1  (+ transformers, torchvision, pillow)
from inference import FusionEmbedder

fe = FusionEmbedder.from_pretrained("EximiusLabs/fusion-embedding-1-2b-preview",
                                    device="cuda")
# or pin a version: revision="v0.2-preview" (current) / "v0.1-preview"

a = fe.embed_audio("dog_barking.wav")                        # [2048]
t = fe.embed_text("a dog barks while rain falls")            # [2048]
i = fe.embed_image("dog_photo.jpg")                          # [2048]

print((a @ t), (a @ i), (t @ i))                             # cosine similarities

a256 = fe.embed_audio("dog_barking.wav", dim=256)            # Matryoshka truncation
```

## Training data and license

v0.2 was trained on ~484K audio–caption pairs: the full AudioCaps train split (45K),
FSD50K, WavCaps/AudioSet_SL, and a 318K-clip subset of LAION-FreeSound, using 10-second
training windows (random crop for longer clips). v0.1 used a 131K-pair subset of the same
sources. As this mix includes YouTube-sourced and research-licensed corpora, the preview
is released under **CC-BY-NC-4.0**. Evaluation sets (AudioCaps test, Clotho, VGGSound,
ESC-50) are excluded from training by clip id.

## Limitations

- Trained on sound-event data; speech content, speaker attributes, and music description
  are supported by the instruction taxonomy but not yet trained to comparable quality.
- English captions; 16 kHz mono input; 30 s per window (longer audio is chunked).
- Audio–text retrieval is below fully fine-tuned CLAP-family models at this checkpoint
  (see Evaluation).

## Roadmap

Further corpus scaling, speech and music coverage, a commercially licensed release tier,
and the 8B model.

## Citation

```bibtex
@software{fusion_embedding_2026,
  title  = {Fusion Embedding 1: A Unified Embedding Space for Text,
            Image, Video, and Audio},
  author = {Tonmoy, Abdul Basit},
  year   = {2026},
  url    = {https://github.com/Eximius-Labs/fusion-embedding-1}
}
```

Built on Qwen3-VL-Embedding and Qwen2.5-Omni, with training data from AudioCaps, WavCaps,
and FSD50K.