| --- |
| license: cc-by-nc-4.0 |
| language: |
| - en |
| pipeline_tag: feature-extraction |
| tags: |
| - embeddings |
| - multimodal |
| - audio |
| - retrieval |
| - matryoshka |
| - qwen3-vl |
| base_model: Qwen/Qwen3-VL-Embedding-2B |
| --- |
| |
| # fusion-embedding-1-2b-preview |
|
|
| Fusion Embedding 1 extends [Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B) |
| with an audio modality. A trained connector (~16M parameters) maps frozen |
| [Qwen2.5-Omni](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) audio-tower features into the |
| base model's embedding space; the base model itself is unmodified. The result is a single |
| embedding space covering **text, images, video, and audio**, with retrieval supported in |
| any direction between modalities. |
|
|
| **Highlights** |
|
|
| - **Unmodified base.** Only the connector is trained; the base model's parameters are |
| byte-identical to the original release, so its text/image/video retrieval performance |
| (MMEB-V2) carries over unchanged. |
| - **Emergent cross-modal alignment.** The connector is trained exclusively on audio–text |
| pairs. Audio→image retrieval nonetheless reaches R@10 0.418 over 696 VGGSound candidates |
| (chance: 0.014) with no audio-visual pairs in training — alignment to text places audio |
| in the space the base already shares across modalities. |
| - **Matryoshka representation.** Embeddings truncate to {2048, 1536, 1024, 512, 256, 128, |
| 64} dimensions with renormalization. |
| - **Compact distribution.** This repository ships the connector and normalization |
| statistics (~60 MB); the frozen towers are downloaded from their original repositories. |
|
|
| This is a **research preview**, currently at **v0.2** (trained on a 484K-pair corpus). |
| v0.1 (131K pairs) remains downloadable via the `v0.1-preview` tag; `v0.2-preview` pins |
| the current version. Both are compared below; pin a tag if you build on this model. |
|
|
| ## Evaluation |
|
|
| **AudioCaps test** — 883 clips, five reference captions per clip, recall computed as |
| min-rank over references: |
|
|
| | Model | A→T R@1 | A→T R@10 | T→A R@10 | |
| |---|---|---|---| |
| | LAION-CLAP | 0.468 | 0.907 | 0.839 | |
| | WavCaps HTSAT-BERT | 0.517 | 0.906 | 0.861 | |
| | Cacophony | 0.553 | 0.924 | 0.864 | |
| | M2D-CLAP | **0.593** | **0.928** | **0.886** | |
| | fusion-embedding-1-2b-preview v0.1 | 0.216 | 0.626 | 0.680 | |
| | **fusion-embedding-1-2b-preview v0.2** | 0.279 | 0.717 | 0.736 | |
|
|
| *CLAP-family models fine-tune both encoders end-to-end and include AudioCaps and Clotho |
| training data; this model keeps both towers frozen and trains only the connector.* |
|
|
| **Clotho v2.1 evaluation** — 1,045 clips × 5 references, zero-shot (Clotho is excluded |
| from training data): |
|
|
| | Model | A→T R@10 | T→A R@10 | |
| |---|---|---| |
| | WavCaps CNN14-BERT (zero-shot) | **0.576** | **0.549** | |
| | fusion-embedding-1-2b-preview v0.1 | 0.252 | 0.329 | |
| | **fusion-embedding-1-2b-preview v0.2** | 0.448 | 0.449 | |
|
|
| **Cross-modal retrieval** — VGGSound-AV, 696 audio/video-frame pairs (chance R@10 = 0.014). |
| R@10 shown as audio-side → other / other → audio-side: |
|
|
| | Model | audio↔image | audio↔text | text↔image | |
| |---|---|---|---| |
| | ImageBind-Huge | **0.718 / 0.720** | 0.404 / 0.348 | 0.243 / 0.282 | |
| | fusion-embedding-1-2b-preview v0.1 | 0.368 / 0.388 | 0.555 / 0.592 | 0.331 / 0.319 | |
| | **fusion-embedding-1-2b-preview v0.2** | 0.418 / 0.440 | **0.588 / 0.631** | **0.331 / 0.319** | |
|
|
| *ImageBind trains directly on audio–image pairs, so that pair is its supervised direction; |
| its audio–text alignment is emergent. This model trains on audio–text only; its |
| audio–image alignment is emergent. Both evaluated with identical clips, frames, and |
| scoring; ImageBind numbers computed with the released imagebind_huge checkpoint.* |
|
|
| Full audio→image metrics (per-modality mean-centered gallery — the readout implemented by |
| `FusionEmbedder.center`; chance R@10 = 0.014): |
|
|
| | Version | R@1 | R@5 | R@10 | mAP@10 | |
| |---|---|---|---|---| |
| | v0.1 | 0.085 | 0.260 | 0.368 | 0.155 | |
| | **v0.2** | **0.088** | **0.315** | **0.418** | **0.179** | |
|
|
| **What audio→image retrieval looks like.** The 0.418 above is not only an aggregate — the |
| retrievals are organized by sound. Real examples (v0.2 checkpoint) on VGGSound-696 |
| (query clip's frame left, top-5 retrieved images right; green = the clip's exact frame): |
|
|
|  |
|
|
| *Example frames from the [VGGSound](https://www.robots.ox.ac.uk/~vgg/data/vggsound/) dataset (CC-BY-4.0), shown for evaluation illustration.* |
|
|
| *Direct hits* — the clip's own frame is returned in the top 5, among the same kind of scene: |
|
|
| | Sound | Top-5 retrieval | Exact frame | |
| |---|---|---| |
| | Metallic clanking and banging | the kitchen it came from, first | rank 1 | |
| | A dog howling | its own dog, then more howling dogs | rank 1 | |
| | A cat purring | its own cat, then more purring and meowing cats | rank 1 | |
| | A siren with a dog howling | its own scene among howling dogs | rank 2 | |
| | *"Switch on the good piece"* (speech) | the blender being switched on | rank 2 | |
| | A female singer in a reverberant space | stage performances and singers | rank 3 | |
|
|
| *Right neighbourhood* — the exact frame ranks lower (often a poor still), but the top |
| results are the correct sound category: |
|
|
| | Sound | Top-5 retrieval | Exact frame | |
| |---|---|---| |
| | A man speaking Spanish amid birdsong | a man speaking with birds chirping behind | rank 13 | |
| | A cat's rhythmic purring | purring and meowing cats | rank 15 | |
| | Bird chirps and tweets | songbirds, owls, a cawing crow | rank 18 | |
| | A power-tool whirring | drills and small motors | rank 32 | |
|
|
| Text, image, and video benchmarks are the base model's published MMEB-V2 results, which |
| are unaffected by this extension. |
|
|
| ## Architecture |
|
|
|  |
|
|
| A perceiver-resampler (width 384, 64 latent queries) translates frozen audio-tower frames |
| into the base model's input embedding space; its outputs occupy placeholder positions in |
| the input stream, mirroring the base model's image-token mechanism. Training is |
| contrastive (InfoNCE over the Matryoshka ladder, symmetric, with a full-corpus |
| frozen-text negative bank — 484K captions at v0.2) against the base model's text |
| embeddings in its native input format. |
|
|
| **Input formatting.** All inputs use the base model's chat-template format (instruction in |
| the system turn, content in the user turn, last-token pooling). Embedding quality is |
| sensitive to this formatting; use the templates in `inference.py`. For cross-modal |
| ranking, per-modality mean-centering of the gallery is recommended (`FusionEmbedder.center`). |
|
|
| ## Usage |
|
|
| ```python |
| # pip install git+https://github.com/Eximius-Labs/fusion-embedding-1 (+ transformers, torchvision, pillow) |
| from inference import FusionEmbedder |
| |
| fe = FusionEmbedder.from_pretrained("EximiusLabs/fusion-embedding-1-2b-preview", |
| device="cuda") |
| # or pin a version: revision="v0.2-preview" (current) / "v0.1-preview" |
| |
| a = fe.embed_audio("dog_barking.wav") # [2048] |
| t = fe.embed_text("a dog barks while rain falls") # [2048] |
| i = fe.embed_image("dog_photo.jpg") # [2048] |
| |
| print((a @ t), (a @ i), (t @ i)) # cosine similarities |
| |
| a256 = fe.embed_audio("dog_barking.wav", dim=256) # Matryoshka truncation |
| ``` |
|
|
| ## Training data and license |
|
|
| v0.2 was trained on ~484K audio–caption pairs: the full AudioCaps train split (45K), |
| FSD50K, WavCaps/AudioSet_SL, and a 318K-clip subset of LAION-FreeSound, using 10-second |
| training windows (random crop for longer clips). v0.1 used a 131K-pair subset of the same |
| sources. As this mix includes YouTube-sourced and research-licensed corpora, the preview |
| is released under **CC-BY-NC-4.0**. Evaluation sets (AudioCaps test, Clotho, VGGSound, |
| ESC-50) are excluded from training by clip id. |
| |
| ## Limitations |
| |
| - Trained on sound-event data; speech content, speaker attributes, and music description |
| are supported by the instruction taxonomy but not yet trained to comparable quality. |
| - English captions; 16 kHz mono input; 30 s per window (longer audio is chunked). |
| - Audio–text retrieval is below fully fine-tuned CLAP-family models at this checkpoint |
| (see Evaluation). |
| |
| ## Roadmap |
| |
| Further corpus scaling, speech and music coverage, a commercially licensed release tier, |
| and the 8B model. |
| |
| ## Citation |
| |
| ```bibtex |
| @software{fusion_embedding_2026, |
| title = {Fusion Embedding 1: A Unified Embedding Space for Text, |
| Image, Video, and Audio}, |
| author = {Tonmoy, Abdul Basit}, |
| year = {2026}, |
| url = {https://github.com/Eximius-Labs/fusion-embedding-1} |
| } |
| ``` |
| |
| Built on Qwen3-VL-Embedding and Qwen2.5-Omni, with training data from AudioCaps, WavCaps, |
| and FSD50K. |
| |