GLM-5.3-Flash-NVFP4-FP8ATTN-512K / scripts /results-longrep-everything-512k.json
tacos4me's picture
everything@512k acceptance: raw dumps (TF battery, vision, longrep) + full gate log
248acfd verified
Raw
History Blame Contribute Delete
1.63 kB
{
"model": "glm-max",
"ctx": 290000,
"gen": 2400,
"rounds": [
{
"round": 1,
"prompt_tokens": 250967,
"completion_tokens": 2912,
"wall_s": 76.2,
"finish_reason": "length",
"four_gram_max": 3,
"four_gram_top": "require the model to",
"eight_gram_max": 1,
"eight_gram_top": "The user wants a detailed essay of at",
"max_consec_line_repeat": 1,
"degenerate_4gram_ge12": false,
"tail": "quality metric), so practitioners can select an operating point rather than a single \"best\" setting.\n4. **End-task deltas**: because perplexity deltas understate task-level damage, downstream accuracy on long-context benchmarks should accompany likelihood measurements.\n\n### Interaction with Other Optimizations\n\nKV quantization also interacts with paged attention, cache eviction, and sliding-window"
},
{
"round": 2,
"prompt_tokens": 250935,
"completion_tokens": 2912,
"wall_s": 70.7,
"finish_reason": "length",
"four_gram_max": 2,
"four_gram_top": "at least 2,000 words.",
"eight_gram_max": 1,
"eight_gram_top": "The user has provided a long filler document",
"max_consec_line_repeat": 1,
"degenerate_4gram_ge12": false,
"tail": "edle affects only a handful of tokens. Likelihood is also insensitive to instruction-following and generation quality: a model can achieve strong teacher-forced numbers while producing degenerate or non-compliant free-form output. For these reasons, teacher-forced scoring is best understood as a *mechanistic* complement to behavioral tests: it localizes *where* in a sequence the model is competent"
}
],
"verdict": "PASS"
}