Instructions to use Qwen/Qwen3-Omni-30B-A3B-Thinking with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen3-Omni-30B-A3B-Thinking with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Qwen/Qwen3-Omni-30B-A3B-Thinking") model = AutoModelForMultimodalLM.from_pretrained("Qwen/Qwen3-Omni-30B-A3B-Thinking", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add the missing `do_sample: true` to generation_config.json
Browse files`generation_config.json` sets `temperature: 0.6`, `top_p: 0.95` and `top_k: 20` but never sets
`do_sample`. In transformers `do_sample` then defaults to false, so **all three are silently
ignored and the model decodes greedily** -- the opposite of what the model card documents.
```python
>>> GenerationConfig.from_pretrained("Qwen/Qwen3-Omni-30B-A3B-Thinking").get_generation_mode()
<GenerationMode.GREEDY_SEARCH: 'greedy_search'>
```
The card's own transformers example calls `model.generate(**inputs, ...)` with no decoding
overrides, so it relies entirely on this file; and the card states that for `Thinking` models
"the decoding parameters should be taken from the `generation_config.json` file in the
checkpoint". Its vLLM example passes `SamplingParams(temperature=0.6, top_p=0.95, top_k=20)`,
which is the behavior this file is evidently meant to reproduce.
The same omission also makes the config unsaveable. `GenerationConfig.save_pretrained()` runs
`validate(strict=True)`, which raises on sampling-only flags while `do_sample` is not `True`:
```
ValueError: GenerationConfig is invalid:
- `temperature`: `do_sample` is not set to `True`. However, `temperature` is set to `0.6` ...
- `top_p`: ...
- `top_k`: ...
```
so any tool that loads this repo and writes it back out -- a quantizer, a fine-tune, a format
conversion -- crashes at the save step, after `config.json` is already written.
Adding `"do_sample": true` fixes both: `get_generation_mode()` becomes `SAMPLE`, the three
parameters take effect as documented, and the config round-trips.
**No effect on vLLM.** vLLM lifts only
`["repetition_penalty", "temperature", "top_k", "top_p", "min_p", "max_new_tokens"]` out of
`generation_config.json` (`ModelConfig.get_diff_sampling_param`); `do_sample` appears nowhere in
its config handling, so serving behavior is byte-for-byte unchanged.
Verified on transformers 5.15.1.
- generation_config.json +1 -0
|
@@ -1,4 +1,5 @@
|
|
| 1 |
{
|
|
|
|
| 2 |
"max_new_tokens": 32768,
|
| 3 |
"repetition_penalty": 1.0,
|
| 4 |
"temperature": 0.6,
|
|
|
|
| 1 |
{
|
| 2 |
+
"do_sample": true,
|
| 3 |
"max_new_tokens": 32768,
|
| 4 |
"repetition_penalty": 1.0,
|
| 5 |
"temperature": 0.6,
|