# Evaluation The tables below come from local matched runs of `Qwen/Qwen3.5-9B@c202236235762e1c871ad0ccb60c8ee5ba337b9a` and Qwick-3.5-9B. Each comparison used the same prompts, grader, and decoding profile for both models. Upstream model-card scores use different harnesses and belong to a separate comparison. `raw correct` counts every correct answer. `budget correct` also requires `finish_reason != length`. Completion length covers thinking and the final answer. ## Public benchmark comparisons | Benchmark / primary metric | Qwen3.5-9B | Qwick-3.5-9B | Delta | Token ratio | Profile | |---|---:|---:|---:|---:|:---:| | MMLU-Pro last-match, 12,032 | 9,792 (81.383%) | 9,760 (81.117%) | -0.266 pp | 0.604890 | A | | GPQA-Diamond, 198 | 154 (77.778%) | 158 (79.798%) | +2.020 pp | 0.710227 | A | | HMMT public 60 × 4 seeds, budget | 187/240 | 172/240 | -6.250 pp | 0.679427 | D | | IFEval prompt strict, 541 | 89.279% | 89.649% | +0.370 pp | 0.772180 | B | | IFEval instruction strict | 92.206% | 92.086% | -0.120 pp | 0.772180 | B | | MMMU validation official-compatible budget, 900 | 664 (73.778%) | 671 (74.556%) | +0.778 pp | 0.740937 | V | | MMMU validation strict budget, 900 | 662 (73.556%) | 667 (74.111%) | n/a | 0.740937 | V | | JMMLU common-clean raw, 987 | 868 (87.943%) | 858 (86.930%) | -1.013 pp | 0.4595 | A | | LiveCodeBench v6 raw, 1,055 | 766 (72.607%) | 785 (74.408%) | +1.801 pp | 0.487927 | C | | LiveCodeBench v6 budget, 1,055 | 702 (66.540%) | 779 (73.839%) | +7.299 pp | 0.487927 | C | ## Decoding and grading profiles - **A (general thinking):** temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5, repetition penalty 1.0, maximum 32,768 new tokens. MMLU-Pro used five-shot `mmlu_pro_fulltrace`, with last-match extraction as primary and first-match as a diagnostic. GPQA-Diamond used zero-shot chain-of-thought with `flexible-extract`. JMMLU reports the common 987-item slice shared by both runs. - **B (IFEval):** thinking, temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5, repetition penalty 1.0, maximum 32,768 new tokens, dataset revision `966cd89545d6b6acfd7638bc708b98261ca58e84`. The grader reads final-answer content; the saved length includes the complete trace. Length stops were 8 for Qwen3.5-9B and 12 for Qwick-3.5-9B; empty final content occurred 8 and 13 times. This matched comparison replaces all prior temperature-zero IFEval results. It remains report-only, with no numeric selection gate or reselection. - **C (LiveCodeBench):** temperature 0.6, top-p 0.95, top-k 20, maximum 81,920 new tokens, seed 2026081101. The official runner at commit `28fef95ea8c9f7a547c8329f2cd3d32b92c1fa24` ran in a network-isolated sandbox. There were no generation errors. Length stops were 278 and 76. - **D (HMMT):** general thinking, maximum 81,920 new tokens, four seeds (2026081101–2026081104), and a fixed boxed/LaTeX/SymPy grader over 60 public problems. - **V (complete vision validation):** same general-thinking sampling as A, 131,072-token server context, maximum 32,768 new tokens, one to five images per request, and prompt-hash paired seeds. Primary grading is official-compatible final-content parsing, including independent per-lane seed-42 random fallback for MC parse misses; fallback counts and deterministic no-random-fallback strict scores are co-reported. ## BF16 complete MMMU validation This covers every one of the 900 validation rows (30 in each of 30 subjects) from `MMMU/MMMU@98e6ac0cb9b7b2cd2c991b85a50762edc4aedc68`: 847 multiple-choice and 53 open-answer rows with one to five prompt images. It is the complete validation split, not the larger 10,500-row test split. Official-compatible Qwen3.5-9B/Qwick budget scores were 664/900 and 671/900, with paired transitions 589 both correct, 154 both wrong, 75 Qwen3.5-9B-only, and 82 Qwick-only. Qwick minus Qwen3.5-9B is +0.778 pp; the subject-stratified paired bootstrap 95% descriptive interval is [-1.889, +3.444] pp and exact two-sided McNemar `p=0.6322`. Because the interval crosses zero, this run detects no BF16 vision degradation but does not resolve an improvement. The official parser used random fallback on 14 / 21 Qwen3.5-9B/Qwick MC parse misses; deterministic strict budget scores were 662/900 and 667/900. Mean full-trace completion tokens were 6476.48 / 4798.66, a ratio of 0.740937. Qwen3.5-9B/Qwick length stops were 12 / 15. Both lanes have 900/900 raw responses, zero errors, matched seeds, and zero missing full traces. This result remains report-only; it did not trigger model reselection or a holdout rerun. ## Public quantized checkpoint vision context | Model | Official-compatible budget | Strict budget | Mean full-trace tokens | Delta vs Qwick BF16 | |---|---:|---:|---:|---:| | Qwen3.5-9B BF16 | 664/900 (73.778%) | 662/900 (73.556%) | 6476.48 | -0.778 pp | | Qwick BF16 | 671/900 (74.556%) | 667/900 (74.111%) | 4798.66 | +0.000 pp | | Qwick FP8 | 655/900 (72.778%) | 655/900 (72.778%) | 5022.40 | -1.778 pp | | Qwick NVFP4 | 635/900 (70.556%) | 626/900 (69.556%) | 5437.84 | -4.000 pp | The FP8 and NVFP4 rows use the exported public weights with vision enabled. Their visual towers remain BF16; 200 language-model Linear matrices use the advertised quantization. Qwick BF16/FP8/NVFP4 use the same prompt-hash seed for every item. FP8 minus Qwick BF16 is -1.778 pp with descriptive interval [-4.222, +0.778] pp; the run does not resolve a vision difference. NVFP4 minus Qwick BF16 is -4.000 pp with descriptive interval [-6.778, -1.222] pp; a vision regression signal was detected. All quantized lanes have 900/900 samples, zero errors, and zero missing full traces. These report-only results did not alter selection or trigger a holdout rerun. ## Qwen3.5-9B BF16 128k sensitivity diagnostic Only the 8 upstream BF16 rows that ended with `finish_reason=length` in the corrected 32k IFEval run were regenerated at a 131,072-token cap. In the logical 541-row overlay, prompt strict changed from 89.279% to 90.018% (+0.739 pp), while length stops and empty final content both changed from 8 to 2. Six rows changed from `length` to `stop`; two remained length-limited at 128k. This is a Qwen3.5-9B-only, report-only sensitivity diagnostic, not a matched Qwick comparison and not a pure causal estimate of the token cap. The generation cap and feasible concurrency/replica layout both changed. All 8 regenerated requests used temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5, and repetition penalty 1.0; full traces are retained. It did not trigger candidate reselection, and the one-shot synthetic holdout was not rerun. ## Frozen internal release gate (separate from public measurement) The 384-item SFT5 genre-v3 holdout used a separately predeclared internal acceptance policy: thinking, `temperature=0`, an 8,192-token cap, matched Base/candidate execution, budget accuracy as primary, and fixed accuracy/efficiency thresholds. The candidate weights, data identity, decode, metrics, and thresholds were frozen before generation. The gate passed: Base/candidate budget accuracy was 72.917% / 89.583% and the candidate/Base mean completion-token ratio was 0.484475. These numbers describe only that internal selection condition; they are not estimates under profile A or B and are excluded from the public benchmark table. The holdout was not used for training or checkpoint selection, no post-result retuning occurred, and it was not rerun. `heldout_results.json` records all aggregate gate metrics, subgroup results, policy identities, and trace-integrity hashes without redistributing prompts, answers, predictions, or full traces. ## Known regressions - HMMT budget correctness fell from 187/240 to 172/240 (-6.25 pp); completion tokens fell by 32.06%. - MMLU-Pro engineering fell from 753/969 to 733/969 (-2.064 pp). - JMMLU common-clean fell from 868/987 to 858/987 (-1.013 pp). Machine-readable values, decode profiles, and artifact identities are in `evaluation_results.json` and `heldout_results.json`.