Instructions to use jondurbin/bagel-7b-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jondurbin/bagel-7b-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jondurbin/bagel-7b-v0.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jondurbin/bagel-7b-v0.1") model = AutoModelForCausalLM.from_pretrained("jondurbin/bagel-7b-v0.1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jondurbin/bagel-7b-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jondurbin/bagel-7b-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jondurbin/bagel-7b-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jondurbin/bagel-7b-v0.1
- SGLang
How to use jondurbin/bagel-7b-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jondurbin/bagel-7b-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jondurbin/bagel-7b-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jondurbin/bagel-7b-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jondurbin/bagel-7b-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jondurbin/bagel-7b-v0.1 with Docker Model Runner:
docker model run hf.co/jondurbin/bagel-7b-v0.1
ChatML format
the model card says this:
I don't really understand the point of having special tokens for <|im_start|> and <|im_end|>, because in practice they just act as BOS and EOS tokens (but, please correct me if I'm wrong).
I'm definitely no expert on this topic, but I have a thought to share.
From the OpenChat paper, they say this:
To differentiate speakers, we introduce a new <|end of turn|> special token at the end of each
utterance, following Zhou et al. (2023). The <|end of turn|> token functions similarly to the
EOS token for stopping generation while preventing confusion with the learned meaning of EOS
during pretraining.
https://arxiv.org/pdf/2309.11235.pdf
In my (non expert) opinion, it makes sense to use a dedicated token for the end of turn, different from EOS, if only because OpenChat and others do it (and OpenChat is a really, really great finetune). And if you just use standard ChatML, then it has the added benefit that any API, library, code, or caller that knows the standard ChatML format could simply consume the model without any changes.
OpenChat doesn't use ChatML:
https://huggingface.co/openchat/openchat_3.5/blob/main/tokenizer_config.json#L51
https://github.com/lm-sys/FastChat/blob/ec9a07ed22110e9686b51fd6ee9bf635b7ce54f8/fastchat/conversation.py#L542
Many other popular models do, e.g. OpenHermes, Dolphin, etc., but they are just changing the stop tokens and adding new special tokens after BOS for some reason, which have the exact same purpose:
https://github.com/lm-sys/FastChat/blob/ec9a07ed22110e9686b51fd6ee9bf635b7ce54f8/fastchat/conversation.py#L1067
The OpenChat discusses performance differences between only having a turn differentiator once vs after each role, but doesn't test the difference between re-using the existing special tokens - I suspect the performance would be identical. They say "... prevent confusion with the learned meaning of EOS..." but I don't actually see evidence that there is confusion in this paper or some of the referenced papers, but perhaps I missed it.