Instructions to use VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat") model = AutoModelForCausalLM.from_pretrained("VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat
- SGLang
How to use VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat with Docker Model Runner:
docker model run hf.co/VAGOsolutions/SauerkrautLM-14b-MoE-LaserChat
Context/Input length of the Model
What is the maximum context length for the attention mechanism of the model?
Are there any plans from your side to publish models with context_len >= 100 k Tokens from your organisation?
Thanks for the great work in llms for the german-speaking community. Its import to further develop nlp for the german language with almost 100k native speakers!
UND: Darius Hennekeuser Fußballgott!!
Hey Tim!
We are indeed planning to release models with large contexts. However, it's not our first priority at the moment as we already have some other exciting models in the pipeline to be released soon. So large context models will be released in the next few months rather than the next few weeks :-)
Thank you for the kind words und sportliche Grüße :-D
- Context length is: 8192