Instructions to use TheBloke/Llama-2-13B-chat-GGML with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TheBloke/Llama-2-13B-chat-GGML with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="TheBloke/Llama-2-13B-chat-GGML")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TheBloke/Llama-2-13B-chat-GGML", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TheBloke/Llama-2-13B-chat-GGML with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TheBloke/Llama-2-13B-chat-GGML" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheBloke/Llama-2-13B-chat-GGML", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/TheBloke/Llama-2-13B-chat-GGML
- SGLang
How to use TheBloke/Llama-2-13B-chat-GGML with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TheBloke/Llama-2-13B-chat-GGML" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheBloke/Llama-2-13B-chat-GGML", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TheBloke/Llama-2-13B-chat-GGML" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheBloke/Llama-2-13B-chat-GGML", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use TheBloke/Llama-2-13B-chat-GGML with Docker Model Runner:
docker model run hf.co/TheBloke/Llama-2-13B-chat-GGML
Initial GGML model commit
Browse files
README.md
CHANGED
|
@@ -149,6 +149,27 @@ Thank you to all my generous patrons and donaters!
|
|
| 149 |
|
| 150 |
# Original model card: Meta's Llama 2 13B-chat
|
| 151 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 152 |
# **Llama 2**
|
| 153 |
Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 13B fine-tuned model, optimized for dialogue use cases and converted for the Hugging Face Transformers format. Links to other models can be found in the index at the bottom.
|
| 154 |
|
|
|
|
| 149 |
|
| 150 |
# Original model card: Meta's Llama 2 13B-chat
|
| 151 |
|
| 152 |
+
---
|
| 153 |
+
extra_gated_heading: Access Llama 2 on Hugging Face
|
| 154 |
+
extra_gated_description: >-
|
| 155 |
+
This is a form to enable access to Llama 2 on Hugging Face after you have been
|
| 156 |
+
granted access from Meta. Please visit the [Meta website](https://ai.meta.com/resources/models-and-libraries/llama-downloads) and accept our
|
| 157 |
+
license terms and acceptable use policy before submitting this form. Requests
|
| 158 |
+
will be processed in 1-2 days.
|
| 159 |
+
extra_gated_button_content: Submit
|
| 160 |
+
extra_gated_fields:
|
| 161 |
+
I agree to share my name, email address and username with Meta and confirm that I have already been granted download access on the Meta website: checkbox
|
| 162 |
+
language:
|
| 163 |
+
- en
|
| 164 |
+
pipeline_tag: text-generation
|
| 165 |
+
inference: false
|
| 166 |
+
tags:
|
| 167 |
+
- facebook
|
| 168 |
+
- meta
|
| 169 |
+
- pytorch
|
| 170 |
+
- llama
|
| 171 |
+
- llama-2
|
| 172 |
+
---
|
| 173 |
# **Llama 2**
|
| 174 |
Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 13B fine-tuned model, optimized for dialogue use cases and converted for the Hugging Face Transformers format. Links to other models can be found in the index at the bottom.
|
| 175 |
|