# Measured release results All values below use temperature 0.6, top-p 0.95, top-k 20, repetition penalty 1.05, and seed 42. Coding and math values are percentages; SWE is a count out of 12. These are fixed subsets and one seeded run, not official leaderboard results. | Model | Effort | HumanEval (132) | MBPP (225) | GSM8K (200) | MATH (150) | BFCL (200) | SWE subset (12) | |---|---|---:|---:|---:|---:|---:|---:| | Qwen 3.8 27B | low | 98.5 | 92.4 | 98.0 | 96.0 | 89.0 | 8/12 | | Qwen 3.8 27B | medium | 99.2 | 93.3 | 97.5 | 95.3 | 89.0 | 5/12 | | Qwen 3.8 27B | xhigh | 93.2 | 82.2 | 95.5 | 96.7 | 87.5 | 3/12 | | Grug v1.1 | low | 91.7 | 82.7 | 85.0 | 78.7 | 80.5 | 7/12 | | Grug v1.1 | medium | 91.7 | 80.4 | 91.0 | 79.3 | 81.0 | 8/12 | | Grug v1.1 | xhigh | 84.1 | 79.1 | 86.0 | 74.7 | 81.0 | 4/12 | | Grug v2 | low | 97.0 | 88.9 | 97.5 | 89.3 | 88.0 | 7/12 | | Grug v2 | medium | 97.0 | 90.7 | 98.5 | 84.7 | 88.0 | 10/12 | | Grug v2 | xhigh | 95.5 | 82.2 | 96.0 | 82.0 | 87.0 | 9/12 | SWE uses a fixed set of six Django and six SymPy issues, a shared pinned environment, and official fail-to-pass/pass-to-pass grading. It is not full SWE-bench Verified. The low/medium per-turn output cap is 2,048; xhigh is 4,096. The other core caps are listed in the protocol. ## Extra output room at xhigh This separate run doubles the core output caps for every model: HumanEval 16,384; MBPP 12,288; GSM8K 8,192; MATH 12,288. It does not change the checkpoint or sample count. | Model | HumanEval | MBPP | GSM8K | MATH | |---|---:|---:|---:|---:| | Qwen 3.8 27B | 95.5 | 88.9 | 96.0 | 98.7 | | Grug v1.1 | 82.6 | 79.1 | 86.0 | 74.7 | | Grug v2 | 97.0 | 85.8 | 96.0 | 82.7 | ## Reasoning retained in repository tool history Separate matched ablation: the original context and per-turn caps are unchanged, but prior assistant reasoning is supplied as reasoning_content. | Model | Low | Medium | Xhigh | |---|---:|---:|---:| | Qwen 3.8 27B | 8/12 | 8/12 | 5/12 | | Grug v1.1 | 8/12 | 9/12 | 7/12 | | Grug v2 | 9/12 | 10/12 | 11/12 | ## Repository xhigh with more output room Separate matched run: 65,536-token context, 16,384 output tokens per turn, reasoning retained, and 35 tool turns. The original small caps stopped the RP1 base model mid-reasoning on 8/12 xhigh episodes; those scores measure constrained-budget performance. | Model | Resolved / 12 | Resolved and final response completed / 12 | |---|---:|---:| | Qwen 3.8 27B | 10 | 7 | | Grug v1.1 | 7 | 5 | | Grug v2 | 11 | 9 | ## Reasoning tokens and completion Mean reasoning tokens count all cases, including mistakes and unfinished answers. Completion percentages show end-of-sequence stops; they do not by themselves measure correctness. | Model | Effort | HumanEval mean thought tokens | HumanEval completed | MBPP mean thought tokens | MBPP completed | |---|---|---:|---:|---:|---:| | Qwen 3.8 27B | low | 483.0 | 100.0% | 562.5 | 98.7% | | Qwen 3.8 27B | medium | 617.0 | 100.0% | 671.3 | 98.7% | | Qwen 3.8 27B | xhigh | 1214.4 | 94.7% | 1569.7 | 87.6% | | Grug v1.1 | low | 80.6 | 100.0% | 94.4 | 100.0% | | Grug v1.1 | medium | 108.6 | 100.0% | 151.1 | 99.6% | | Grug v1.1 | xhigh | 56.8 | 100.0% | 82.4 | 100.0% | | Grug v2 | low | 191.8 | 100.0% | 272.7 | 99.1% | | Grug v2 | medium | 265.6 | 100.0% | 305.5 | 98.7% | | Grug v2 | xhigh | 655.7 | 97.7% | 759.5 | 92.4% | ## OpenCode title regression Thirty actual title prompts per effort. A passing title must finish, contain one line, obey the 50-character instruction, and avoid tool calls. Ten additional tool-avoidance requests are kept separate. | Model | Effort | Completed titles / 30 | Tool-call leaks / 30 | Additional tool-avoidance leaks / 10 | |---|---|---:|---:|---:| | Qwen 3.8 27B | low | 30 | 0 | 0 | | Qwen 3.8 27B | medium | 30 | 0 | 0 | | Qwen 3.8 27B | xhigh | 28 | 0 | 0 | | Grug v1.1 | low | 23 | 4 | 0 | | Grug v1.1 | medium | 21 | 7 | 2 | | Grug v1.1 | xhigh | 27 | 2 | 1 | | Grug v2 | low | 30 | 0 | 0 | | Grug v2 | medium | 30 | 0 | 0 | | Grug v2 | xhigh | 30 | 0 | 0 | ## Integrated MTP The following coding and tool subsets use two speculative tokens and the same release sampler. All head tensors reside in the main checkpoint. | Effort | HumanEval (132) | MBPP (225) | BFCL (200) | |---|---:|---:|---:| | low | 97.0 | 88.0 | 87.5 | | medium | 96.2 | 88.0 | 90.0 | | xhigh | 95.5 | 80.9 | 88.0 | Serial H200 decoding after warmup, 18 fixed prompts, greedy decoding, no prefix cache, and repetition penalty 1.05: | Draft tokens | Output tokens/second | Draft token acceptance | Completed requests / 18 | |---:|---:|---:|---:| | 0 | 68.02 | — | 17 | | 1 | 112.77 | 86.87% | 18 | | 2 | 142.44 | 77.20% | 18 | Throughput is specific to this small serial workload. MTP-off and MTP-on produced identical greedy token sequences on 10/18 requests; one- and two-draft modes matched on all 18. Different output lengths and execution paths limit interpretation as a general speedup. Functional benchmark accuracy is reported above rather than assuming token-level identity. ## GGUF runtime checks Every text file contains 15 MTP tensors. Generation and API checks ran on an HF H200 with MTP disabled and with two draft tokens. Coding uses a fixed 16-case HumanEval smoke subset at medium effort; it is not a full quantized-model benchmark. Functional code scores use the corrected grader described in the protocol. | Quantization | Basic checks, off / on (12 each) | Native API, off / on (15 each) | HumanEval smoke, MTP on (16) | Greedy token parity (12) | |---|---:|---:|---:|---:| | Q3_K_M | 12 / 12 | 14 / 14 | 15 | 8 | | Q4_K_M | 12 / 12 | 14 / 14 | 14 | 11 | | Q5_K_M | 12 / 12 | 14 / 14 | 15 | 8 | | Q6_K | 12 / 12 | 14 / 14 | 16 | 12 | | Q8_0 | 12 / 12 | 14 / 14 | 16 | 10 | The one native API capability failure in each mode is llama-server discarding the `reasoning` input alias. Canonical `reasoning_content` and inline history work. This result is retained as a failure; the included client normalization helper handles the alias before sending it. All 24 separate vLLM HTTP API checks passed, including streamed tools and all three efforts. ## Reference runs without a repetition penalty These earlier matched runs use repetition penalty 1.0; every other core sampling parameter and original output cap is unchanged. They show sensitivity to decoding settings. They must not be mixed with the primary 1.05 table or used to attribute the entire gap to training alone. | Model | Effort | HumanEval | MBPP | GSM8K | MATH | BFCL | |---|---|---:|---:|---:|---:|---:| | Qwen 3.8 27B | low | 99.2 | 91.6 | 97.5 | 96.0 | 90.0 | | Qwen 3.8 27B | medium | 99.2 | 91.6 | 98.0 | 92.7 | 89.0 | | Qwen 3.8 27B | xhigh | 95.5 | 78.7 | 95.0 | 93.3 | 86.5 | | Grug v1.1 | low | 90.2 | 80.0 | 90.0 | 81.3 | 83.5 | | Grug v1.1 | medium | 87.1 | 80.4 | 94.0 | 83.3 | 85.0 | | Grug v1.1 | xhigh | 86.4 | 80.0 | 93.0 | 78.0 | 83.0 | | Grug v2 | low | 97.0 | 85.8 | 97.0 | 87.3 | 86.5 | | Grug v2 | medium | 97.7 | 89.8 | 95.5 | 89.3 | 89.5 | | Grug v2 | xhigh | 92.4 | 83.6 | 96.5 | 84.7 | 88.5 | ## Authored repository repair suite A separate 12-case synthetic repository suite tests edits, tools, and held-out assertions. This is not SWE-bench. Scores below are percentages at repetition penalty 1.05. | Model | Low | Medium | Xhigh | |---|---:|---:|---:| | Qwen 3.8 27B | 33.3 | 91.7 | 100.0 | | Grug v1.1 | 100.0 | 75.0 | 83.3 | | Grug v2 | 83.3 | 91.7 | 100.0 | ## Repository termination and tool validity Resolved patches and completed agent responses are distinct. A patch may pass held-out tests even if the agent keeps working until its turn limit. Context evictions count removed old conversation turns, not dropped benchmark cases. | Protocol | Model | Effort | Resolved with completed response / 12 | Step-limit episodes | Invalid tool calls | Context evictions | |---|---|---|---:|---:|---:|---:| | Original caps | Qwen 3.8 27B | low | 7 | 1 | 0 | 0 | | Original caps | Qwen 3.8 27B | medium | 5 | 1 | 0 | 4 | | Original caps | Qwen 3.8 27B | xhigh | 3 | 1 | 0 | 30 | | Original caps | Grug v1.1 | low | 6 | 3 | 0 | 0 | | Original caps | Grug v1.1 | medium | 8 | 2 | 0 | 0 | | Original caps | Grug v1.1 | xhigh | 4 | 3 | 0 | 0 | | Original caps | Grug v2 | low | 6 | 4 | 0 | 7 | | Original caps | Grug v2 | medium | 9 | 3 | 0 | 0 | | Original caps | Grug v2 | xhigh | 6 | 5 | 0 | 11 | | Retained history | Qwen 3.8 27B | low | 7 | 2 | 0 | 0 | | Retained history | Qwen 3.8 27B | medium | 7 | 3 | 0 | 3 | | Retained history | Qwen 3.8 27B | xhigh | 4 | 2 | 0 | 39 | | Retained history | Grug v1.1 | low | 4 | 5 | 0 | 0 | | Retained history | Grug v1.1 | medium | 7 | 3 | 0 | 0 | | Retained history | Grug v1.1 | xhigh | 4 | 8 | 0 | 18 | | Retained history | Grug v2 | low | 7 | 4 | 1 | 0 | | Retained history | Grug v2 | medium | 9 | 2 | 0 | 0 | | Retained history | Grug v2 | xhigh | 10 | 2 | 0 | 48 | | Larger xhigh budget | Qwen 3.8 27B | xhigh | 7 | 4 | 0 | 8 | | Larger xhigh budget | Grug v1.1 | xhigh | 5 | 6 | 0 | 0 | | Larger xhigh budget | Grug v2 | xhigh | 9 | 3 | 0 | 0 | ## Effort, style, and remaining failures Grug v2 spends more reasoning tokens as coding effort increases and uses substantially fewer than Qwen in the matched coding runs. This does not establish monotonic accuracy: higher effort loses accuracy on several fixed-budget coding and math subsets, and doubling the xhigh allowance does not remove every regression. Medium is the release default. The reasoning is usually terse Grug-style text at all three efforts, but longer and harder cases can drift into ordinary English. The release does not enforce an absolute style grammar. The saved-output audit found a real unfinished repetition loop in low-effort MATH case index 61 even with repetition penalty 1.05. Other unfinished answers and tool-turn limits are reported above. The model is not loop-free. The title regression suite improved to 30/30 at every effort without title tool-call leaks. This addresses the reported OpenCode failure under the tested prompts and server configurations; the actual OpenCode desktop application was not exercised.