Instructions to use bhogan94-prime/muse-glimmer-30b-finetuned-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bhogan94-prime/muse-glimmer-30b-finetuned-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bhogan94-prime/muse-glimmer-30b-finetuned-v2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("bhogan94-prime/muse-glimmer-30b-finetuned-v2") model = AutoModelForMultimodalLM.from_pretrained("bhogan94-prime/muse-glimmer-30b-finetuned-v2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bhogan94-prime/muse-glimmer-30b-finetuned-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bhogan94-prime/muse-glimmer-30b-finetuned-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bhogan94-prime/muse-glimmer-30b-finetuned-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bhogan94-prime/muse-glimmer-30b-finetuned-v2
- SGLang
How to use bhogan94-prime/muse-glimmer-30b-finetuned-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bhogan94-prime/muse-glimmer-30b-finetuned-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bhogan94-prime/muse-glimmer-30b-finetuned-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bhogan94-prime/muse-glimmer-30b-finetuned-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bhogan94-prime/muse-glimmer-30b-finetuned-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use bhogan94-prime/muse-glimmer-30b-finetuned-v2 with Docker Model Runner:
docker model run hf.co/bhogan94-prime/muse-glimmer-30b-finetuned-v2
Muse Glimmer 30B v2 — serving instructions
Complete fine-tuned model weights, with parallel tool calling enabled during training. Load this directory directly; do not attach an adapter or merge it with another model. Text-and-tool use is the validated scope; multimodal quality has not been validated.
This is a separate version from bhogan94-prime/muse-glimmer-30b-finetuned. Do not reuse that version's single-tool-call request settings. Parallel calls are supported, not guaranteed on every turn: the model can return zero, one or multiple calls depending on the question and tools.
1. Download
Install uv and the Hugging Face CLI (uv tool install huggingface_hub). Authenticate using hf auth login with an account authorized for this private repository. Never put a token in application code or a shared command.
hf download bhogan94-prime/muse-glimmer-30b-finetuned-v2 --local-dir ./glimmer-v2
cd glimmer-v2
sha256sum --check SHA256SUMS
For repeatable deployments, include --revision <COMMIT_SHA> using the immutable revision in the release handoff.
2. Install the pinned runtime
Reference platform: Linux x86-64, Python 3.12, NVIDIA H100 80 GB and a driver compatible with CUDA 13.0. The reference serving layout is eight H100s: four data-parallel replicas, each using two GPUs. This preserves the validated layout; it is not a claim that eight GPUs are the minimum to fit the weights. Different GPU counts, quantization, context limits or runtimes need separate verification.
bash install.sh
The installer uses a new .venv, pinned packages and the official vLLM 0.28.0 release wheel. Core versions are PyTorch 2.13.0+cu130, vLLM 0.28.0, Transformers 5.6.2, Tokenizers 0.22.2, Safetensors 0.7.0, and FlashInfer 0.6.16.post3. runtime_versions.json and requirements.txt record the pins. No training framework, internal repository or private package is needed for inference. Do not replace these pins with a floating nightly or latest.
3. Start the server
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 bash serve.sh
The launcher checks package versions and explicitly selects the bundled tokenizer/template.
| Setting | Value |
|---|---|
| Model and tokenizer | This downloaded directory |
| Template | Included chat_template.jinja, explicitly selected |
| Dtype / quantization | BF16 / none |
| Tensor parallelism / data parallelism | 2 / 4 on one host |
| Context length | 65,536 tokens, including input and generation |
| Active sequences | 32 per replica |
| GPU memory utilization | 0.92 |
| Tool / reasoning parsers | muse_glimmer / muse_glimmer |
| Automatic tool choice | Enabled |
| Prefix caching / chunked prefill | Enabled / enabled |
| Execution | Eager; custom all-reduce disabled |
Default address is 127.0.0.1:8000. Configure HOST, PORT, or MODEL_NAME through environment variables. Remote access should go through an authenticated gateway; do not expose an unauthenticated server publicly. Compilation caches go into a new local temporary directory, not the model folder or shared storage.
4. Make requests
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-H 'X-Session-ID: example-conversation-001' \
-d '{
"model": "muse-glimmer-30b-finetuned-v2",
"messages": [
{"role":"system","content":"You are a helpful assistant.\n\nReasoning strength: medium"},
{"role":"user","content":"Explain why the sky is blue in two sentences."}
],
"temperature": 0.6,
"max_tokens": 16384,
"top_p": 0.95,
"top_k": 64,
"parallel_tool_calls": true,
"stop_token_ids": [200001, 200008]
}'
The output budget includes both reasoning and the final answer. The recommended reasoning setting is medium, selected with Reasoning strength: medium in the system message. Without an override, the native template defaults to high. Do not substitute another model family's reasoning_effort or thinking switches. Keep the bundled generation config and end-token IDs. Other sampling parameters retain the pinned runtime's defaults.
The example system message is deliberately generic. Supply your application's actual prompt and function schemas; model/runtime identity alone does not guarantee the same behavior with different prompts, tools or data.
5. Handle parallel tool calls correctly
- Supply function schemas in
tools, withtool_choice: "auto"andparallel_tool_calls: true. - Read the entire
message.tool_callsarray, even whenfinish_reasonisstop. Do not discard all but its first entry. - Append the assistant message once, preserving each tool-call ID, name and JSON arguments.
- Execute independent, safe tool operations concurrently. For every call, append a
role: "tool"message containing the result and its exacttool_call_id. Return structured errors if a tool fails; do not invent successful results. - Wait for every call in that assistant turn to finish before requesting the next model turn. Dependent or side-effecting operations require application-level ordering and authorization.
- Continue with the conversation history, tool schemas and the same sampling settings. Reasoning returned separately is not the final user-facing answer and should not be pasted into ordinary message text.
The model server does not execute tools itself. parallel_tool_calls: true permits multiple calls; the client implements execution and result matching. X-Session-ID should remain stable per conversation and unique across conversations; it provides sticky routing only if your gateway implements it. A bare vLLM server does not acquire routing behavior from the header alone.
Application turn/time budgets belong in your harness, not in the model weights. Set explicit budgets appropriate to your application; a successful health check is not proof of correct tool handling.
6. Verify the deployment
.venv/bin/python smoke_test.py --base-url http://127.0.0.1:8000/v1
This uses synthetic examples only. It checks an ordinary reply, a single tool round trip, two structured calls in the same assistant turn, both matching tool results, and a final answer without a repeated call. Run it before connecting an application. Generation is stochastic, so identical settings do not guarantee byte-identical output across requests or hardware.
Task-specific prompts, schemas, datasets, traces, training code and evaluation results are intentionally excluded. The upstream LICENSE and USAGE_POLICY.md are included unchanged. These weights are modified from the upstream model.
- Downloads last month
- 35
Model tree for bhogan94-prime/muse-glimmer-30b-finetuned-v2
Base model
meta-models/Muse-Glimmer-30B