Conflicting sampling parameter recommendations?

#1
by ayylmaonade - opened

Hey! Firstly, thanks for the great work with the UD quants, as always. The Q4_K_XL GGUF is running great at 96K context on my 7900 XTX via llama.cpp w/ Vulkan.

However I've got some questions regarding sampling params, specifically temperature. On your guys' site, the Qwen3-VL guide states to set a temperature of 1.0 for 'Thinking' variants of this model series. But when you click the source, the Qwen3-VL git seems to recommend a temp of 0.6 instead. I wanted to see if I could get some clarification on this if possible?

I've been testing both 0.6 and 1.0. Using temp=1 has caused a few issues for me. I've had one instance of chinese characters showing up, and a few weird cases where the model contains its chain-of-thought outside of the tag. For reference, I'm using the other recommended params;

  • top_p=0.95
  • top_k=20
  • repeat_penalty=1.0

Although, to reduce excessive thinking, I am using presence_penalty=1.5 as was recommended for non-VL Qwen3-Thinking models. I was also wondering if min_p should be set to 0.0 or kept at 0.01? (the llama.cpp default as far as I can tell)

Once again, appreciate all the awesome work you guys do! UD quants are my go-to, and I'd appreciate any advice on this.

Cheers :)

Hey! Firstly, thanks for the great work with the UD quants, as always. The Q4_K_XL GGUF is running great at 96K context on my 7900 XTX via llama.cpp w/ Vulkan.

However I've got some questions regarding sampling params, specifically temperature. On your guys' site, the Qwen3-VL guide states to set a temperature of 1.0 for 'Thinking' variants of this model series. But when you click the source, the Qwen3-VL git seems to recommend a temp of 0.6 instead. I wanted to see if I could get some clarification on this if possible?

I've been testing both 0.6 and 1.0. Using temp=1 has caused a few issues for me. I've had one instance of chinese characters showing up, and a few weird cases where the model contains its chain-of-thought outside of the tag. For reference, I'm using the other recommended params;

  • top_p=0.95
  • top_k=20
  • repeat_penalty=1.0

Although, to reduce excessive thinking, I am using presence_penalty=1.5 as was recommended for non-VL Qwen3-Thinking models. I was also wondering if min_p should be set to 0.0 or kept at 0.01? (the llama.cpp default as far as I can tell)

Once again, appreciate all the awesome work you guys do! UD quants are my go-to, and I'd appreciate any advice on this.

Cheers :)

@ayylmaonade I'm running into the same issue! I'm wondering if you stuck with 0.6, and if that worked better for you? What about presence_penalty and min_p?

Hey! Firstly, thanks for the great work with the UD quants, as always. The Q4_K_XL GGUF is running great at 96K context on my 7900 XTX via llama.cpp w/ Vulkan.

However I've got some questions regarding sampling params, specifically temperature. On your guys' site, the Qwen3-VL guide states to set a temperature of 1.0 for 'Thinking' variants of this model series. But when you click the source, the Qwen3-VL git seems to recommend a temp of 0.6 instead. I wanted to see if I could get some clarification on this if possible?

I've been testing both 0.6 and 1.0. Using temp=1 has caused a few issues for me. I've had one instance of chinese characters showing up, and a few weird cases where the model contains its chain-of-thought outside of the tag. For reference, I'm using the other recommended params;

  • top_p=0.95
  • top_k=20
  • repeat_penalty=1.0

Although, to reduce excessive thinking, I am using presence_penalty=1.5 as was recommended for non-VL Qwen3-Thinking models. I was also wondering if min_p should be set to 0.0 or kept at 0.01? (the llama.cpp default as far as I can tell)

Once again, appreciate all the awesome work you guys do! UD quants are my go-to, and I'd appreciate any advice on this.

Cheers :)

@ayylmaonade I'm running into the same issue! I'm wondering if you stuck with 0.6, and if that worked better for you? What about presence_penalty and min_p?

After a few days of tinkering with a whole assortment of settings, I finally landed on a decent combination. It’s still a notch below the older 2507 checkpoint when it comes to pure‑text tasks though. I settled on a temperature of 1.0, a min_p of 0.0, and a presence_penalty of 1.5. Even though a few repos suggest a penalty of 0.0, I saw no drop in quality and over‑thinking got noticeably cut down. All the other parameters stayed the same: top_k 20, top_p 0.95, repeat_penalty 1.0.

I also found a handy benchmark that directly compares VL models against their non‑VL counterparts in text‑based scenarios. https://dubesor.de/benchtable - have a look here if you want, just search "Qwen3" - at the very least it confirms I'm not insane regarding the reduction in text-performance, lol.

Thanks! Missed this notification, but I'm on GLM 4.7 Flash and currently evaluating Qwen3.5 now 😁

Sign up or log in to comment