How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "Locutusque/Orca-2-13b-SFT-v4"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Locutusque/Orca-2-13b-SFT-v4",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Use Docker
docker model run hf.co/Locutusque/Orca-2-13b-SFT-v4
Quick Links

The "microsoft/Orca-2-13b" model fully fine-tuned on HuggingFaceH4/no_robots, totally-not-an-llm/EverythingLM-data-V3, mlabonne/guanaco-llama2-1k, and OpenAssistant/oasst_top1_2023-08-25. This model achieved a test loss of 0.18.

Make sure to comply with the microsoft research license. Please read it before using this model.

This model was trained on the ChatML prompt template.

The responses seen in the inference API were generated using the following sampling parameters:

temperature = 0.1

top_p = 0.14

top_k = 41

repetition_penalty = 1.176

Updates:

12/18/23 - 🔥 This model holds the #5 position on the Open LLM Leaderboard among llama2-13b models. 🔥

Downloads last month
40
Safetensors
Model size
13B params
Tensor type
BF16
·
Inference Examples
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Locutusque/Orca-2-13b-SFT-v4

Finetuned
(5)
this model
Merges
1 model
Quantizations
3 models

Datasets used to train Locutusque/Orca-2-13b-SFT-v4