Instructions to use poolside/Laguna-S-2.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use poolside/Laguna-S-2.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="poolside/Laguna-S-2.1", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("poolside/Laguna-S-2.1", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("poolside/Laguna-S-2.1", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use poolside/Laguna-S-2.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "poolside/Laguna-S-2.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "poolside/Laguna-S-2.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/poolside/Laguna-S-2.1
- SGLang
How to use poolside/Laguna-S-2.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "poolside/Laguna-S-2.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "poolside/Laguna-S-2.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "poolside/Laguna-S-2.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "poolside/Laguna-S-2.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use poolside/Laguna-S-2.1 with Docker Model Runner:
docker model run hf.co/poolside/Laguna-S-2.1
Premature ending inside final answer after long thinking
After thinking is done, model stops in the middle of the final response. I am using b10092 build of llama-server with official WebUI when model is loading the following is displayed, may be that is the root of the problem:
[34m0.01.334.528[0m [35mW load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
[0m[34m0.01.334.533[0m [35mW load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect
My cli command is:
llama-server ^
--model C:\Users\user.lmstudio\models\unsloth\Laguna-S-2.1-GGUF\Laguna-S-2.1-UD-Q8_K_XL-00001-of-00004.gguf ^
--alias Laguna-S-2.1 ^
--jinja ^
--chat-template-file laguna.jinja ^
--threads 9 ^
--threads-http 4 ^
--flash-attn on ^
--no-context-shift ^
--temp 1.0 --top-k 40 --top-p 1.0 --min-p 0.01 --repeat-penalty 1.0 --presence-penalty 2.0 ^
--ctx-size 65536 ^
--n-predict 65536 ^
--host 0.0.0.0 --port 8000 ^
--no-mmap ^
--n-gpu-layers 999 ^
--n-cpu-moe 44 ^
--reasoning on ^
--chat-template-kwargs "{"preserve_thinking":true}" ^
--batch_size 1024 --ubatch_size 1024
The only difference from the official jinja template is "\n" after thinking tag "" to make model always think.
Example of the answer:
...
Final Answer
\boxed{\frac{c \lambda^2}{2} + \frac{2 v^2}{(v + c \lambda)^2} - 1 > 0}
<//think>
To prove the inequality (\frac{c \lambda^2}{2} + \frac{2 v^2}{(v + c \lambda)^2} - 1 > 0) given (c > 0), (\lambda > 0), (c \lambda < v < 1), and ((v + c \lambda)^2 > 4c\
I put double slash in thinking end tag because with HF formatting the tag was disappearing, so in reality it is normal ending tag.