--- base_model: perplexity-ai/pplx-embed-v1-0.6b language: en license: apache-2.0 tags: - quantized - 4bit - bnb - transformers model_name: pplx-embed-v1-0.6b-bnb-4bit-nf4 --- # pplx-embed-v1-0.6b (Quantized) ## Description This model is a 4-bit quantized version of the original [`perplexity-ai/pplx-embed-v1-0.6b`](https://huggingface.co/perplexity-ai/pplx-embed-v1-0.6b) model, optimized for reduced memory usage while maintaining performance. ## Quantization Details - **Quantization Type**: 4-bit - **bnb_4bit_quant_type**: nf4 - **bnb_4bit_use_double_quant**: True - **bnb_4bit_compute_dtype**: bfloat16 - **bnb_4bit_quant_storage**: uint8 - **Original Footprint**: 2384.20 MB (FLOAT32) - **Quantized Footprint**: 842.79 MB (UINT8) - **Memory Reduction**: 64.7% ## Usage ```python from transformers import AutoModel, AutoTokenizer model_name = "pplx-embed-v1-0.6b-bnb-4bit-nf4" model = AutoModel.from_pretrained( "manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4", ) tokenizer = AutoTokenizer.from_pretrained("manu02/pplx-embed-v1-0.6b-bnb-4bit-nf4", use_fast=True) ```