Add the missing `do_sample: true` to generation_config.json

#5

generation_config.json sets temperature: 0.6, top_p: 0.95 and top_k: 20 but never sets
do_sample. In transformers do_sample then defaults to false, so all three are silently
ignored and the model decodes greedily
-- the opposite of what the model card documents.

>>> GenerationConfig.from_pretrained("Qwen/Qwen3-Omni-30B-A3B-Thinking").get_generation_mode()
<GenerationMode.GREEDY_SEARCH: 'greedy_search'>

The card's own transformers example calls model.generate(**inputs, ...) with no decoding
overrides, so it relies entirely on this file; and the card states that for Thinking models
"the decoding parameters should be taken from the generation_config.json file in the
checkpoint". Its vLLM example passes SamplingParams(temperature=0.6, top_p=0.95, top_k=20),
which is the behavior this file is evidently meant to reproduce.

The same omission also makes the config unsaveable. GenerationConfig.save_pretrained() runs
validate(strict=True), which raises on sampling-only flags while do_sample is not True:

ValueError: GenerationConfig is invalid:
- `temperature`: `do_sample` is not set to `True`. However, `temperature` is set to `0.6` ...
- `top_p`: ...
- `top_k`: ...

so any tool that loads this repo and writes it back out -- a quantizer, a fine-tune, a format
conversion -- crashes at the save step, after config.json is already written.

Adding "do_sample": true fixes both: get_generation_mode() becomes SAMPLE, the three
parameters take effect as documented, and the config round-trips.

No effect on vLLM. vLLM lifts only
["repetition_penalty", "temperature", "top_k", "top_p", "min_p", "max_new_tokens"] out of
generation_config.json (ModelConfig.get_diff_sampling_param); do_sample appears nowhere in
its config handling, so serving behavior is byte-for-byte unchanged.

Verified on transformers 5.15.1.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment