The Llama 3 Herd of Models
Paper • 2407.21783 • Published • 119
How to use lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8 with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages) # Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8")
model = AutoModelForCausalLM.from_pretrained("lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))How to use lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8 with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker model run hf.co/lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8
How to use lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8 with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'How to use lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8 with Docker Model Runner:
docker model run hf.co/lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8
This is a W8A8 (8-bit weights and 8-bit activations) quantized version of meta-llama/Llama-3.1-8B-Instruct.
This model was quantized using llmcompressor with the following configuration:
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "lsm0729/Meta-Llama-3.1-8B-Instruct-quantized.w8a8"
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# Example inference
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"},
]
input_ids = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt"
).to(model.device)
outputs = model.generate(
input_ids,
max_new_tokens=256,
do_sample=True,
temperature=0.6,
top_p=0.9,
)
response = tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
This model inherits the Llama 3.1 Community License from the base model. Please refer to the original model card for full license details.
If you use this model, please cite both the original Llama 3.1 paper and the quantization library:
@article{llama3.1,
title={The Llama 3 Herd of Models},
author={Meta AI},
year={2024},
url={https://arxiv.org/abs/2407.21783}
}
Base model
meta-llama/Llama-3.1-8B