abtonmoy commited on
Commit
f482f57
·
verified ·
1 Parent(s): 4bad8f3

Card: precise cell count (8 of 12 release-protocol cells; every recorded text-to-audio direction)

Browse files
Files changed (1) hide show
  1. README.md +3 -3
README.md CHANGED
@@ -56,9 +56,9 @@ base's text→image retrieval scores to four decimal places). Trained on 518K
56
  audio–caption pairs with a full-corpus frozen-text negative bank, it leads every
57
  unified embedding model we measured on audio↔text retrieval — ahead of ImageBind,
58
  LanguageBind, and Gemini Embedding 2 in both directions — and improves on
59
- fusion-embedding-1 v0.3 in 9 of 12 measured cells, with the largest gains in
60
- text→audio search. Audio↔image alignment is emergent (zero audio–image pairs in
61
- training).
62
 
63
  | Feature | Value |
64
  | --- | --- |
 
56
  audio–caption pairs with a full-corpus frozen-text negative bank, it leads every
57
  unified embedding model we measured on audio↔text retrieval — ahead of ImageBind,
58
  LanguageBind, and Gemini Embedding 2 in both directions — and improves on
59
+ fusion-embedding-1 v0.3 in 8 of 12 release-protocol cells, including every
60
+ recorded text→audio direction. Audio↔image alignment is emergent (zero
61
+ audio–image pairs in training).
62
 
63
  | Feature | Value |
64
  | --- | --- |