YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Qwen3-Embedding-4B โ€” RBLN (ATOM CA25) compiled

Qwen/Qwen3-Embedding-4B pre-compiled for Rebellions RBLN-CA25 NPUs with optimum-rbln, ready to serve with vllm-rbln. This repository contains the compiled artifact only โ€” RBLN NPUs and the RBLN SDK are required to run it.

Hardware & software

Component Version / configuration
NPU 2ร— RBLN-CA25 chiplets (32 GiB total; half of one 64 GiB CA25 card)
Compiler rebel-compiler==0.11.2
Runtime optimum-rbln==0.11.2, vllm-rbln==0.11.2.post1
Framework transformers==5.8.1, vllm==0.24.0+cpu

Compile configuration

Option Value
max_seq_len 40,960
batch_size 1
num_devices 2
attn_impl flash_attn
kvcache_partition_len 8,192
prefill_chunk_size 128
kvcache_num_blocks 5 (one complete 40,960-token sequence)
phases prefill only

The batch size is intentionally one. In the current Optimum pooling path, larger batches are dispatched through a Python-side per-item prefill loop and do not improve NPU throughput. A 64 GiB CA25 card can host two independent 2-chiplet replicas for real request-level parallelism.

Serving

vllm serve <this-repo-or-local-path> \
  --runner pooling \
  --max-num-seqs 1 \
  --max-model-len 40960 \
  --no-enable-prefix-caching

Do not pass --tensor-parallel-size 2; the two-device split is already baked into the compiled model and vllm-rbln uses tensor parallel size 1 for the serving process.

The OpenAI-compatible embedding endpoint is /v1/embeddings. Embeddings have 2,560 dimensions and are returned L2-normalized.

License

Derived from Qwen/Qwen3-Embedding-4B; use is subject to the original model's license and terms.

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support