Instructions to use dphn/dolphin-2.6-mistral-7b-dpo-laser with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dphn/dolphin-2.6-mistral-7b-dpo-laser with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dphn/dolphin-2.6-mistral-7b-dpo-laser") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dphn/dolphin-2.6-mistral-7b-dpo-laser") model = AutoModelForCausalLM.from_pretrained("dphn/dolphin-2.6-mistral-7b-dpo-laser", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dphn/dolphin-2.6-mistral-7b-dpo-laser with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dphn/dolphin-2.6-mistral-7b-dpo-laser" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dphn/dolphin-2.6-mistral-7b-dpo-laser", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dphn/dolphin-2.6-mistral-7b-dpo-laser
- SGLang
How to use dphn/dolphin-2.6-mistral-7b-dpo-laser with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dphn/dolphin-2.6-mistral-7b-dpo-laser" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dphn/dolphin-2.6-mistral-7b-dpo-laser", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dphn/dolphin-2.6-mistral-7b-dpo-laser" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dphn/dolphin-2.6-mistral-7b-dpo-laser", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use dphn/dolphin-2.6-mistral-7b-dpo-laser with Docker Model Runner:
docker model run hf.co/dphn/dolphin-2.6-mistral-7b-dpo-laser
Open Code
Can we hope that you will post your Laser code implementation using Random Matrix Theory?
Yes. you can :)
Keep following our next steps <3
Do y'all have a quick/brief overview of how the pruning rates are selected? Is it some variation of selecting a singular value cut-off from the Marchenko–Pastur distribution (say retain just 10% for layer N, 20% of N-10, etc), and then normalize that based on the norm of the matrix?
I believe the paper did a first-order grid-search from (highest layer, lowest-pruning-rate) across each weight. This encodes the intuition in their paper that you shouldn't prune the early layers, but pruning the later layers is extra beneficial. Is that strategy (favor later layers) also applicable here?
Additional question - would we see a significant (inference) performance boost via preprocessing the model? I thought a big part of the speedup due to using SVD is that SVD matmul is much faster, I believe the resulting LASER-ed weights are still data-dense with the same dimensions as before? That would suggest no speedup in inference until the engine can take advantage of the knowledge that some of the weights are low-rank right?
We use marchenko pastur to define which is the threshold cut for singular values.
And yes, we start the search from top layers downwards, which is optimal.
Stay tuned because we will release the code in the next few days.
@fernandofernandes Just wanted to follow up on this. Have you had time to release the code yet, and if so, could you share a link to it? Thanks for all of your hard work!