Card: precise cell count (8 of 12 release-protocol cells; every recorded text-to-audio direction)
Browse files
README.md
CHANGED
|
@@ -56,9 +56,9 @@ base's text→image retrieval scores to four decimal places). Trained on 518K
|
|
| 56 |
audio–caption pairs with a full-corpus frozen-text negative bank, it leads every
|
| 57 |
unified embedding model we measured on audio↔text retrieval — ahead of ImageBind,
|
| 58 |
LanguageBind, and Gemini Embedding 2 in both directions — and improves on
|
| 59 |
-
fusion-embedding-1 v0.3 in
|
| 60 |
-
text→audio
|
| 61 |
-
training).
|
| 62 |
|
| 63 |
| Feature | Value |
|
| 64 |
| --- | --- |
|
|
|
|
| 56 |
audio–caption pairs with a full-corpus frozen-text negative bank, it leads every
|
| 57 |
unified embedding model we measured on audio↔text retrieval — ahead of ImageBind,
|
| 58 |
LanguageBind, and Gemini Embedding 2 in both directions — and improves on
|
| 59 |
+
fusion-embedding-1 v0.3 in 8 of 12 release-protocol cells, including every
|
| 60 |
+
recorded text→audio direction. Audio↔image alignment is emergent (zero
|
| 61 |
+
audio–image pairs in training).
|
| 62 |
|
| 63 |
| Feature | Value |
|
| 64 |
| --- | --- |
|