| { | |
| "model": "glm-max", | |
| "ctx": 290000, | |
| "gen": 2400, | |
| "rounds": [ | |
| { | |
| "round": 1, | |
| "prompt_tokens": 250967, | |
| "completion_tokens": 2912, | |
| "wall_s": 76.2, | |
| "finish_reason": "length", | |
| "four_gram_max": 3, | |
| "four_gram_top": "require the model to", | |
| "eight_gram_max": 1, | |
| "eight_gram_top": "The user wants a detailed essay of at", | |
| "max_consec_line_repeat": 1, | |
| "degenerate_4gram_ge12": false, | |
| "tail": "quality metric), so practitioners can select an operating point rather than a single \"best\" setting.\n4. **End-task deltas**: because perplexity deltas understate task-level damage, downstream accuracy on long-context benchmarks should accompany likelihood measurements.\n\n### Interaction with Other Optimizations\n\nKV quantization also interacts with paged attention, cache eviction, and sliding-window" | |
| }, | |
| { | |
| "round": 2, | |
| "prompt_tokens": 250935, | |
| "completion_tokens": 2912, | |
| "wall_s": 70.7, | |
| "finish_reason": "length", | |
| "four_gram_max": 2, | |
| "four_gram_top": "at least 2,000 words.", | |
| "eight_gram_max": 1, | |
| "eight_gram_top": "The user has provided a long filler document", | |
| "max_consec_line_repeat": 1, | |
| "degenerate_4gram_ge12": false, | |
| "tail": "edle affects only a handful of tokens. Likelihood is also insensitive to instruction-following and generation quality: a model can achieve strong teacher-forced numbers while producing degenerate or non-compliant free-form output. For these reasons, teacher-forced scoring is best understood as a *mechanistic* complement to behavioral tests: it localizes *where* in a sequence the model is competent" | |
| } | |
| ], | |
| "verdict": "PASS" | |
| } |