Feature Extraction
Transformers.js
ONNX
multilingual
jina_embeddings_v5_omni
sentence-similarity
multimodal
cross-modal-retrieval
webgpu
jina-embeddings
embeddings
custom_code
Instructions to use splatoonWooo/jina-embeddings-v5-omni-nano-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use splatoonWooo/jina-embeddings-v5-omni-nano-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('feature-extraction', 'splatoonWooo/jina-embeddings-v5-omni-nano-ONNX');
File size: 7,411 Bytes
4d7148d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 | ---
license: cc-by-nc-4.0
base_model: jinaai/jina-embeddings-v5-omni-nano
base_model_relation: quantized
library_name: transformers.js
pipeline_tag: feature-extraction
tags:
- feature-extraction
- sentence-similarity
- multimodal
- cross-modal-retrieval
- onnx
- webgpu
- transformers.js
- jina-embeddings
- embeddings
language:
- multilingual
---
# jina-embeddings-v5-omni-nano · ONNX
ONNX exports of [`jinaai/jina-embeddings-v5-omni-nano`](https://huggingface.co/jinaai/jina-embeddings-v5-omni-nano) for in-browser inference via [transformers.js](https://huggingface.co/docs/transformers.js) **v4** with WebGPU.
`jina-embeddings-v5-omni-nano` is a ~1.04 B-parameter multimodal embedding model that maps **text, images, audio, and video** into a single shared 768-dimensional L2-normalized space, enabling cross-modal retrieval without reindexing.
## Architecture
The base model is a LLaVA-style composition of three frozen encoders plus small trainable projectors, with per-task LoRA adapters. This repo's exports have the **`retrieval`** task adapter merged into the static weights.
| Tower | Backbone | Role |
|---|---|---|
| Text | EuroBERT-210m (loaded as a bidirectional Llama) | Text encoder |
| Vision | Qwen3-VL vision tower + spatial merger | Image and video frames |
| Audio | Whisper-large-v3 encoder + Qwen2.5-Omni audio adapter | Audio (16 kHz mono) |
Each tower is bundled with its retrieval projector AND a copy of the language model (which performs the final cross-modal fusion + last-token pooling), so the three ONNX graphs are independently loadable.
## Files
Three ONNX graphs, three quantization variants each. Every variant ships as a small `.onnx` schema + a large `.onnx.data` or `.onnx_data` external-weight sidecar (HF Hub-friendly layout, no 2 GB protobuf limits).
| Modality | fp32 | fp16 | q4f16 | Verified parity vs fp32 (fp16) |
|---|---:|---:|---:|---|
| Text | 849 MB | 424 MB | 263 MB | cos = 1.000000 |
| Vision | 1247 MB | 622 MB | 460 MB | cos = 0.999998 |
| Audio | 3400 MB | 1700 MB | 1465 MB | (fp16 doesn't load on CPU EP; verify on WebGPU) |
**For desktop WebGPU, use `q4f16`** — it's the smallest and runs natively on shader-int4 hardware. Use `fp16` if you need higher numerical fidelity or your GPU lacks int4 paths.
Every graph outputs a single tensor `sentence_embedding` of shape `[batch, 768]`, already L2-normalized — cosine similarity reduces to a dot product.
## Input constraints
**Text** (`text_model*.onnx`): inputs `input_ids` and `attention_mask`, both `[batch, seq]` with `seq` fully dynamic. Apply the asymmetric retrieval prefix convention before tokenizing: `Query: …` for queries, `Document: …` for corpus items.
**Vision** (`vision_model*.onnx`): the graph is traced at a fixed image-patch layout. `image_grid_thw` is folded as a constant (it drives `torch.linspace` inside Qwen3-VL's `fast_pos_embed_interpolate`, which dynamo cannot symbolicate). Resize every image to **224×224** before passing through `LlavaEuroBertProcessor` — that yields the exact shapes the graph expects:
```
input_ids [batch, 271] int64
attention_mask [batch, 271] int64
pixel_values [1024, 1536] float32
```
`image_grid_thw` is NOT an ONNX input (it's a constant). Don't pass it.
**Audio** (`audio_model*.onnx`): traced with a 5-second 16 kHz mono clip. The graph expects exactly:
```
input_ids [batch, 125] int64 (audio_token_id placeholders)
attention_mask [batch, 125] int64
input_features [1, 128, 3000] float32 (Whisper log-mel, 30s padded)
feature_attention_mask [1, 3000] int64 (frame-level, 500 ones for 5s of real audio)
```
Pad or truncate every clip to 5 s. Longer-clip chunking (sliding 5 s windows, average pooled embeddings) is a v2 follow-up.
## Use with transformers.js v4
### Text
```js
import { AutoTokenizer } from "@huggingface/transformers";
import * as ort from "onnxruntime-web";
const REPO = "shreyask/jina-embeddings-v5-omni-nano-ONNX";
const tok = await AutoTokenizer.from_pretrained(REPO);
const sess = await ort.InferenceSession.create(
`https://huggingface.co/${REPO}/resolve/main/onnx/text_model_q4f16.onnx`,
{ executionProviders: ["webgpu", "wasm"] },
);
const { input_ids, attention_mask } = await tok(
"Query: a saxophone solo",
{ return_tensors: "ort" },
);
const { sentence_embedding } = await sess.run({ input_ids, attention_mask });
// Float32Array of length 768, already L2-normalized.
```
### Vision
```js
import { AutoProcessor } from "@huggingface/transformers";
import * as ort from "onnxruntime-web";
const proc = await AutoProcessor.from_pretrained(REPO);
const sess = await ort.InferenceSession.create(
`https://huggingface.co/${REPO}/resolve/main/onnx/vision_model_q4f16.onnx`,
{ executionProviders: ["webgpu", "wasm"] },
);
// Resize to 224x224 before passing in.
const inputs = await proc.apply_chat_template(
[{ role: "user", content: [{ type: "image", image: imageBlob }] }],
{ add_generation_prompt: false, tokenize: true, return_dict: true, return_tensors: "ort" },
);
const { sentence_embedding } = await sess.run({
input_ids: inputs.input_ids,
attention_mask: inputs.attention_mask,
pixel_values: inputs.pixel_values,
});
```
### Audio
Audio preprocessing isn't bundled in `LlavaEuroBertProcessor` — use Whisper's feature extractor directly and stamp `audio_token_id` placeholders into `input_ids`. See the [reference implementation](https://huggingface.co/jinaai/jina-embeddings-v5-omni-nano/blob/main/modeling_jina_embeddings_v5_omni.py) for the exact mel-spec → placeholder count plumbing.
## Cross-modal retrieval
All three towers project into the same 768-dim space, so a text query can rank images / audio / video corpus items (and vice versa) without re-indexing. Embeddings are L2-normalized, so cosine similarity is a dot product:
```js
const score = textVec.reduce((s, v, i) => s + v * imageVec[i], 0);
```
## How these were exported
- `torch` 2.11, `transformers` 5.8, `onnx` 1.21, `onnxruntime` 1.26
- **Text**: `torch.onnx.export(..., dynamo=True)`. The LlamaModel-based encoder exports cleanly through the dynamo path
- **Vision and audio**: `torch.onnx.export(..., dynamo=False)` (legacy TorchScript tracer). Dynamo refuses to specialize Qwen3-VL's data-dependent `torch.linspace`, and the GQA-aware SDPA is monkey-patched with a manual MatMul+Softmax for the trace's duration
- PEFT LoRA fused via `merge_and_unload(safe_merge=True)` so the `retrieval` task adapter is baked in
- fp16 cast via `onnxruntime.transformers.float16.convert_float_to_float16` (the `onnxconverter_common` path mishandles dynamo's `_to_copy` nodes)
- 4-bit quant via `onnxruntime.quantization.matmul_nbits_quantizer.MatMulNBitsQuantizer` (`bits=4, block_size=32, accuracy_level=4`)
- Two graph post-patches required for the audio tower's ORT-loadability: `Cast(to=int64)` inserted before every `Slice` index input (366 inserts), and `Unsqueeze(0)`/`Squeeze(0)` wrapped around the rank-2-input `AveragePool` (1 site)
## License
Inherited from the base model: **CC BY-NC 4.0**. Commercial use requires reaching out to `sales@jina.ai`.
## Citation
```bibtex
@misc{jina-embeddings-v5-omni-nano,
author = {Jina AI},
title = {jina-embeddings-v5-omni-nano},
year = {2025},
url = {https://huggingface.co/jinaai/jina-embeddings-v5-omni-nano}
}
```
|