Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
danielhanchenΒ 
posted an update 3 days ago

can you pls link a blogpost or something, what's the source for this post?

Β·

It was posted officially by Google: https://x.com/googlegemma/status/2077449152062247219

Only the chat templates were updated? Ex., checked the 26B A4B 8-bit MLX quant. I can't seem to reconcile "model is updated" with "last modified date 4 months ago".

Β·

Speed is the headline. plz12345 already found the real story underneath: same hashes, only the template moved.

If the weights are byte-identical and just the chat template changed, that is still a real fix. The template is what decides whether function-call args parse. But it means every quant and the bf16 should improve together, not one more than another.

So the receipt that matters isn't 're-download the GGUF'. It is the same weights, old template vs new, tool-call pass rate before and after.

On one quant, how much did tool-call accuracy actually move between the two templates?

Β·

If you read our graphic, it says you can update the template as well. Most people don't know how to replace the chat template.