PrimaMonarch-EroSumika-2x10.7B
Collection
OUTDATED • 4 items • Updated • 1
How to use xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages) # Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k")
model = AutoModelForCausalLM.from_pretrained("xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))How to use xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker model run hf.co/xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k
How to use xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'How to use xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k with Docker Model Runner:
docker model run hf.co/xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k
docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k" \
--host 0.0.0.0 \
--port 30000# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'
(Maybe i'll change the icon picture later.)
Experimental MoE, the idea is to have more active parameters than 7xX model would have and keep it's size lower than 20B.
This model has ~19.2B parameters.
Exl2, 4.0 bpw (Fits in 12GB VRAM/16k context/4-bit cache)
slices:
- sources:
- model: MistralInstruct-v0.2-128k
layer_range: [0, 24]
- sources:
- model: MistralInstruct-v0.2-128k
layer_range: [8, 24]
- sources:
- model: MistralInstruct-v0.2-128k
layer_range: [24, 32]
merge_method: passthrough
dtype: bfloat16
xxx777xxxASD/PrimaSumika-10.7B-128k
slices:
- sources:
- model: EroSumika-128k
layer_range: [0, 24]
- sources:
- model: Prima-Lelantacles-128k
layer_range: [8, 24]
- sources:
- model: EroSumika-128k
layer_range: [24, 32]
merge_method: passthrough
dtype: bfloat16
slices:
- sources:
- model: AlphaMonarch-7B-128k
layer_range: [0, 24]
- sources:
- model: NeuralHuman-128k
layer_range: [8, 24]
- sources:
- model: AlphaMonarch-7B-128k
layer_range: [24, 32]
merge_method: passthrough
dtype: bfloat16
Each 128k model is a slerp merge with Epiculous/Fett-uccine-Long-Noodle-7B-120k-Context
Install from pip and serve model
# Install SGLang from pip: pip install sglang# Start the SGLang server: python3 -m sglang.launch_server \ --model-path "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k" \ --host 0.0.0.0 \ --port 30000# Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xxx777xxxASD/PrimaMonarch-EroSumika-2x10.7B-128k", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'