Muse Glimmer 30B v2 — serving instructions

Complete fine-tuned model weights, with parallel tool calling enabled during training. Load this directory directly; do not attach an adapter or merge it with another model. Text-and-tool use is the validated scope; multimodal quality has not been validated.

This is a separate version from bhogan94-prime/muse-glimmer-30b-finetuned. Do not reuse that version's single-tool-call request settings. Parallel calls are supported, not guaranteed on every turn: the model can return zero, one or multiple calls depending on the question and tools.

1. Download

Install uv and the Hugging Face CLI (uv tool install huggingface_hub). Authenticate using hf auth login with an account authorized for this private repository. Never put a token in application code or a shared command.

hf download bhogan94-prime/muse-glimmer-30b-finetuned-v2 --local-dir ./glimmer-v2
cd glimmer-v2
sha256sum --check SHA256SUMS

For repeatable deployments, include --revision <COMMIT_SHA> using the immutable revision in the release handoff.

2. Install the pinned runtime

Reference platform: Linux x86-64, Python 3.12, NVIDIA H100 80 GB and a driver compatible with CUDA 13.0. The reference serving layout is eight H100s: four data-parallel replicas, each using two GPUs. This preserves the validated layout; it is not a claim that eight GPUs are the minimum to fit the weights. Different GPU counts, quantization, context limits or runtimes need separate verification.

bash install.sh

The installer uses a new .venv, pinned packages and the official vLLM 0.28.0 release wheel. Core versions are PyTorch 2.13.0+cu130, vLLM 0.28.0, Transformers 5.6.2, Tokenizers 0.22.2, Safetensors 0.7.0, and FlashInfer 0.6.16.post3. runtime_versions.json and requirements.txt record the pins. No training framework, internal repository or private package is needed for inference. Do not replace these pins with a floating nightly or latest.

3. Start the server

CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 bash serve.sh

The launcher checks package versions and explicitly selects the bundled tokenizer/template.

Setting Value
Model and tokenizer This downloaded directory
Template Included chat_template.jinja, explicitly selected
Dtype / quantization BF16 / none
Tensor parallelism / data parallelism 2 / 4 on one host
Context length 65,536 tokens, including input and generation
Active sequences 32 per replica
GPU memory utilization 0.92
Tool / reasoning parsers muse_glimmer / muse_glimmer
Automatic tool choice Enabled
Prefix caching / chunked prefill Enabled / enabled
Execution Eager; custom all-reduce disabled

Default address is 127.0.0.1:8000. Configure HOST, PORT, or MODEL_NAME through environment variables. Remote access should go through an authenticated gateway; do not expose an unauthenticated server publicly. Compilation caches go into a new local temporary directory, not the model folder or shared storage.

4. Make requests

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -H 'X-Session-ID: example-conversation-001' \
  -d '{
    "model": "muse-glimmer-30b-finetuned-v2",
    "messages": [
      {"role":"system","content":"You are a helpful assistant.\n\nReasoning strength: medium"},
      {"role":"user","content":"Explain why the sky is blue in two sentences."}
    ],
    "temperature": 0.6,
    "max_tokens": 16384,
    "top_p": 0.95,
    "top_k": 64,
    "parallel_tool_calls": true,
    "stop_token_ids": [200001, 200008]
  }'

The output budget includes both reasoning and the final answer. The recommended reasoning setting is medium, selected with Reasoning strength: medium in the system message. Without an override, the native template defaults to high. Do not substitute another model family's reasoning_effort or thinking switches. Keep the bundled generation config and end-token IDs. Other sampling parameters retain the pinned runtime's defaults.

The example system message is deliberately generic. Supply your application's actual prompt and function schemas; model/runtime identity alone does not guarantee the same behavior with different prompts, tools or data.

5. Handle parallel tool calls correctly

  1. Supply function schemas in tools, with tool_choice: "auto" and parallel_tool_calls: true.
  2. Read the entire message.tool_calls array, even when finish_reason is stop. Do not discard all but its first entry.
  3. Append the assistant message once, preserving each tool-call ID, name and JSON arguments.
  4. Execute independent, safe tool operations concurrently. For every call, append a role: "tool" message containing the result and its exact tool_call_id. Return structured errors if a tool fails; do not invent successful results.
  5. Wait for every call in that assistant turn to finish before requesting the next model turn. Dependent or side-effecting operations require application-level ordering and authorization.
  6. Continue with the conversation history, tool schemas and the same sampling settings. Reasoning returned separately is not the final user-facing answer and should not be pasted into ordinary message text.

The model server does not execute tools itself. parallel_tool_calls: true permits multiple calls; the client implements execution and result matching. X-Session-ID should remain stable per conversation and unique across conversations; it provides sticky routing only if your gateway implements it. A bare vLLM server does not acquire routing behavior from the header alone.

Application turn/time budgets belong in your harness, not in the model weights. Set explicit budgets appropriate to your application; a successful health check is not proof of correct tool handling.

6. Verify the deployment

.venv/bin/python smoke_test.py --base-url http://127.0.0.1:8000/v1

This uses synthetic examples only. It checks an ordinary reply, a single tool round trip, two structured calls in the same assistant turn, both matching tool results, and a final answer without a repeated call. Run it before connecting an application. Generation is stochastic, so identical settings do not guarantee byte-identical output across requests or hardware.

Task-specific prompts, schemas, datasets, traces, training code and evaluation results are intentionally excluded. The upstream LICENSE and USAGE_POLICY.md are included unchanged. These weights are modified from the upstream model.

Downloads last month
35
Safetensors
Model size
30B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bhogan94-prime/muse-glimmer-30b-finetuned-v2

Finetuned
(43)
this model