Memory performance

#67
by trombonator - opened

Hi,
Using this model i find some "holes" to retrieve informations in context depending on the depth even with KVcache = fp16
I digged up on youtbe and one channel faced same issue : https://youtu.be/zLs2QG7lU7Q?si=9YG6WXepN8gsipwG (memory part 5mn)
Do you think it is possible to correct that? This model is increadible, it could be even better without this drawback
image

Have a very nice day ahead

of course, yarn. I read about it, despite the native context being 256k due to the ternary it have that gap in the middle, so
--rope-scale 4.0 --rope-scaling yarn --yarn-orig-ctx 65536 was the cure for Ternary 1 (at least it worked for me back in the days)
probably here it will be the same, if you run such big context that is, the gap is not in the middle of the context, it is in specific size of the context i.e. if you run 48K context there is no gap.
I will not play smart, ask any chatbot about the actual explanation, there was something about overlapping of quantization and something aaaand IDK already, it was months ago

A caution on the YaRN flags above, based on this repo's GGUF metadata: this model is natively 262K. The files have qwen35.context_length = 262144, qwen35.rope.freq_base = 10000000 and no rope-scaling keys. In llama.cpp, --rope-scale 4 sets rope_freq_scale = 1/4, and --yarn-orig-ctx 65536 tells YaRN the model was trained at 65K. So together they stretch positions 4Γ— on a model that doesn't need it, and that alone can blur long-range retrieval. I'd A/B it against no rope flags at the same context before relying on it.

On the "holes" themselves: only 16 of the 64 layers are full attention (full_attention_interval = 4). The other 48 are recurrent with a fixed-size state, so exact recall from deep in the context goes through those 16 layers. Their KV is 16 Γ— 4 KV heads Γ— 256 Γ— 2 (K,V) Γ— 2 bytes = 64 KiB/token in f16, or 16 GiB at the full 262K. So with fp16 KV as you have it, cache precision isn't the cause. If you want to test for a context-length effect, run the same needle test at a smaller -c (e.g. 48K, where the reply above says the gap goes away) with the same prompt position.

Sign up or log in to comment