Feature Extraction
Transformers
Safetensors
English
bidirectional_pplx_qwen3
quantized
4bit
bnb
custom_code
text-embeddings-inference
4-bit precision
bitsandbytes
Instructions to use manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4-dq with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4-dq with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4-dq", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4-dq", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4-dq", trust_remote_code=True, device_map="auto")Quick Links
pplx-embed-v1-0.6b (Quantized)
Description
This model is a 4-bit quantized version of the original perplexity-ai/pplx-embed-v1-0.6b model, optimized for reduced memory usage while maintaining performance.
Quantization Details
- Quantization Type: 4-bit
- bnb_4bit_quant_type: nf4
- bnb_4bit_use_double_quant: True
- bnb_4bit_compute_dtype: bfloat16
- bnb_4bit_quant_storage: uint8
- Original Footprint: 2384.20 MB (FLOAT32)
- Quantized Footprint: 842.79 MB (UINT8)
- Memory Reduction: 64.7%
Usage
from transformers import AutoModel, AutoTokenizer
model_name = "pplx-embed-v1-0.6b-bnb-4bit-nf4"
model = AutoModel.from_pretrained(
"manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4",
)
tokenizer = AutoTokenizer.from_pretrained("manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4", use_fast=True)
- Downloads last month
- 6
Model tree for manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4-dq
Base model
perplexity-ai/pplx-embed-v1-0.6b
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4-dq", trust_remote_code=True)