Instructions to use Nabbers1999/Mini-Llama-8B-Chat-SFT-0129 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nabbers1999/Mini-Llama-8B-Chat-SFT-0129 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Nabbers1999/Mini-Llama-8B-Chat-SFT-0129") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Nabbers1999/Mini-Llama-8B-Chat-SFT-0129") model = AutoModelForCausalLM.from_pretrained("Nabbers1999/Mini-Llama-8B-Chat-SFT-0129", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Nabbers1999/Mini-Llama-8B-Chat-SFT-0129 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nabbers1999/Mini-Llama-8B-Chat-SFT-0129" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nabbers1999/Mini-Llama-8B-Chat-SFT-0129", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Nabbers1999/Mini-Llama-8B-Chat-SFT-0129
- SGLang
How to use Nabbers1999/Mini-Llama-8B-Chat-SFT-0129 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nabbers1999/Mini-Llama-8B-Chat-SFT-0129" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nabbers1999/Mini-Llama-8B-Chat-SFT-0129", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nabbers1999/Mini-Llama-8B-Chat-SFT-0129" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nabbers1999/Mini-Llama-8B-Chat-SFT-0129", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Nabbers1999/Mini-Llama-8B-Chat-SFT-0129 with Docker Model Runner:
docker model run hf.co/Nabbers1999/Mini-Llama-8B-Chat-SFT-0129
Mini-Llama 8B Chat - 0129
My instruct model has undergone DoRA SFT on my custom synthetic chat dataset, containing single and multi-round chats containing SFW, NSFW, and Toxic single and multi-round chats. This reinforces the model's uncensored compliance with all prompts and teaches it how to better fill roles assigned to it in the system prompt.
This model has yet to go through DPO preference training and may still have rough edges.
** Be aware that this adapter, when used without a system prompt to assign it a role may make up its own role. Meaning if you just say 'Hello' it could resond with 'Hello, how may I assist you?' or it might respond with something like "Hi, my name is Carol and I'm a librarian here to assist you with finding the book you're looking for."
For the base pretrain, see: Nabbers1999/Mini-Llama-8B-Base-0124
For the instruct, see: Nabbers1999/Mini-Llama-8B-Instruct-0124
- Downloads last month
- 2
