Xyntetik-Kvist-14B

Release: built to be downloaded and served with Xyntetik Runner.

  • Measured status against the parent: held-out KLD 0.762 and margin-qualified top-1 84.0% against the BF16 parent on 45,056 positions, its own row (the house quant bar does not apply); 57 of 60 held-out closed-loop tasks against the parent's 60; 99 of 99 tool calls valid; every preregistered envelope-gate arm passed.
  • Parent model: meta-models/Muse-Glimmer-30B, cut by width to a dense 14B and distilled back from it; a new model, not a quantised copy
  • Collection: Xyntetik Kvist: distilled agents
  • Evidence: training record (dataset: preregistrations, K1 to K3 and envelope-gate records, corpus manifests, the envelope corpus, code), also copied under evidence/ here
  • Runner compatibility: Serves on Xyntetik Runner v0.5.7 or later; the gate arms E1, E2, E3 and E5 were measured on runner main 0bfa2ad and the serving measurements on 53b4deb, ac5418e and the v0.5.7 release binary; the first two are contained in v0.5.7. GGUF files are in Xyntetik-Kvist-14B-GGUF. Q8_0 (15.4 GB), Q5_0-mix (10.3 GB) and IQ4_NL-mix (7.6 GB) fit a 24 GB card whole and run GPU-resident; BF16 (28.9 GB) needs partial offload there (see Limits).

What it is, and is not. A 14B research student of Muse-Glimmer-30B, built for one job: running tool-calling agent loops on Xyntetik Runner from a 24 GB card. Its training budget is small: 6,000 distillation steps (98.3 M tokens, 162 hours) and 1,440 steps on agentic trajectories. It is not a drop-in replacement for Muse-Glimmer-30B, for Gemma, or for any general-purpose model: on held-out text it reads KLD 0.762 against its own parent, and its tool-calling result (57 of 60) holds for the study's environment and the muse wire format on Runner, nothing wider. Download it to serve agent loops on Runner and to read the evidence; do not download it expecting the parent's quality.

Muse-Glimmer-30B's language model, cut by width to a dense student of 14,443,781,760 parameters: all 52 layers kept, hidden width 6,656 → 5,760, FFN 19,968 → 10,240, attention heads 32 → 24, KV heads 2. It was distilled back from the frozen BF16 parent for 6,000 steps (98.3 M tokens, 162 hours), then trained for 1,440 more steps on agentic trajectories written in Runner's muse wire format. Text only: the parent's 1.92 B-parameter vision encoder is not carried.

The cost first. On the study's 45,056-position held-out split this model reads KLD 0.762 and margin-qualified top-1 84.0% against the BF16 parent (top-1 79.2%). That is a student's own row. It is far from the house bar for quantised copies (KLD ≤ 0.05, margin-qualified top-1 ≥ 97%), which does not apply to a model with 14.4 B of the parent's 27.9 B language-model parameters and is not claimed. Closed-loop agentic tasks: 57 of 60 held-out environment tasks solved, scored by re-execution against ground truth, where the parent solves 60 and the student before the agentic phase solves 0.

What is claimed, and only this. Every arm of the preregistered envelope gate passed. The model writes Runner's muse wire format (199 of 200 first turns well formed), its tool calls parse and validate against the offered schema (99 of 99), and it solves 57 of 60 held-out tasks of the study's executable environment end to end (57 also when a numeric task counts only with the right bolded answer). That is a tool-calling claim scoped to that environment and that serving path; see Limits and the disclosures. Nothing about tool calling is claimed from the general distillation, whose corpora held no tool-calling documents.

Serve with xyntetik-runner

runner -m Xyntetik-Kvist-14B-Q8_0.gguf --serve --port 8080

Four files: BF16, the file the gate measured, and Q8_0, Q5_0-mix and IQ4_NL-mix, each measured against it below. The BF16 file is 28.9 GB and does not fit a 24 GB card whole; the gate served it with 38 of 52 layers on the GPU:

runner -m Xyntetik-Kvist-14B-BF16.gguf --serve --port 8080 --gpu auto --gpu-layers 38

For anything that may produce a tool call, keep "repeat_penalty": 1.0. That is Runner v0.5.7's default for this model, but a penalty set by the client acts on the schema tokens a call has to repeat (and only when sampling, at temperature above 0). Every gate number below was measured greedy (temperature 0) with the penalty at 1.0.

Recommended: greedy, with the loop guard

Runner v0.5.7 and later can close a reasoning turn that has started repeating itself (--loop-guard, off by default; per request loop_guard). This is the recommended way to serve this model:

runner -m Xyntetik-Kvist-14B-Q8_0.gguf --serve --port 8080 --loop-guard

On the gate's 60 closed-loop tasks (this checkpoint at Q8_0 on CPU, greedy, the Runner v0.5.7 release binary) the guard leaves the score where it is: 58 of 60 with the guard and 58 of 60 without it on the same binary, failing the same two tasks. It intervened only in the one task that runs away without it, a two-decimals calc request, 5 times, and that task still did not terminate; it fired on no task that passed without it. Termination stayed at 98.3% and mean thinking tokens went from 154 to 147. On an earlier checkpoint of this study (attempt 8, runner 53b4deb) the same setting raised the score from 51 to 54 of 60 by ending 3 of its 6 runaways, so what the guard buys depends on how often a checkpoint loops; on this one its value is a bounded turn, not a better score.

The guard fixes the turn, not the task. When the model loops because it does not have the answer, closing the loop leaves a wrong answer. It also cannot see a loop that spans requests: in one measured task an earlier checkpoint of this study (attempt 8) had the answer in its reasoning and still called the same tool with the same argument three times, each turn well formed on its own. Only the caller, which holds the whole conversation, can catch that. This checkpoint passes that task (task-20260915-02037, two calls). Token repetition is a serving problem with a serving fix, the guard; deciding to answer instead of calling again is a training problem, and the fix here was training on post-tool rollouts.

Reasoning-channel sampling (--reasoning-temp) was measured too and is not recommended for this model: on an earlier checkpoint of this study (attempt 8, runner 53b4deb, Q8_0 on CPU, the gate's 60 tasks), reasoning temperature 0.6 with top-p 0.9, min-p 0.05 and top-k 20 supplied in the request solved 49 and 47 tasks on two seeds against 51 greedy; termination was 92% and 87% against 90%, and both seeds turned the same calc and units tasks into loops.

The gate numbers below were all measured greedy without the guard. The guard is a serving recommendation, not part of the gate.

Quickstart

runner -m Xyntetik-Kvist-14B-Q8_0.gguf -p "The capital of France is" -n 32
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "messages": [{"role": "user", "content": "Find me a flight from Stockholm to Berlin on 2026-10-14."}],
  "tools": [{"type": "function", "function": {"name": "find_flight",
    "description": "Search for available flights between two cities on a given date.",
    "parameters": {"type": "object", "properties": {
      "origin": {"type": "string", "description": "Departure city"},
      "destination": {"type": "string", "description": "Arrival city"},
      "date": {"type": "string", "description": "Departure date, YYYY-MM-DD"}},
      "required": ["origin", "destination", "date"]}}}],
  "max_tokens": 1024, "temperature": 0, "repeat_penalty": 1.0}'

On Runner v0.5.7 this returns a find_flight call with all three arguments filled (finish_reason tool_calls). Give the model a calculator tool for arithmetic. Without one it computes in its head and can be wrong: asked "What is 17% of 2,340?" at temperature 0 with no tools, it answered 462.6 (the answer is 397.8).

What was changed, exactly

Three stages, each with its own record.

  • Width pruning (PRUNE.json). Minitron-style activation importance, measured on 131,072 tokens (64 × 2,048) of the Muse-Glimmer surgery study's window-audited corpus_run9 training split, eval split never read: hidden channels chosen globally by summed mean |x|, FFN neurons per layer by mean |h| at the down_proj input, attention heads per KV group by the mean norm of their o_proj input. The embedding and the LM head are cut on the hidden axis with the same channel set. No training. PRUNE.json holds every kept index, the calibration capture's SHA-256 and a hash of every tensor written.

    parent language model this model
    embedding 1,344,831,488 1,163,796,480
    52 decoder blocks 25,165,117,952 12,116,188,800
    LM head 1,344,831,488 1,163,796,480
    total 27,854,780,928 14,443,781,760 (51.9% kept)

    Unchanged: the tokenizer (202,048 entries, tokenizer.json identical to the parent's), the chat template, head dimension 128, the attention pattern (three sliding-window layers of 2,048 tokens to one global), the logit softcap.

  • General distillation (DISTILLATION.json, general_kd). 6,000 steps of 4 windows × 4,095 predicted positions, 98.28 M tokens: 40% of corpus v1 (59,946 sequences of 4,096 tokens, 245.5 M tokens, the study's 11 chat and task domains from SmolTalk and the Tulu 3 SFT mixture plus Wikitext-103, plain render, single pass). Objective: top-64 forward KL to the frozen BF16 parent with a tail bucket, through the parent's softcapped logits. All 52 blocks and the final norm trainable; embedding and LM head frozen at their pruned values. AdamW (0.9, 0.95), peak lr 1e-4, 50-step warmup, cosine to 1e-5, per-block clip 1.0, one block resident on a 24 GB GPU slice with fp32 master weights on the host. 583,591 s of wall time (162.1 h).

  • Envelope phase (DISTILLATION.json, envelope). 1,440 steps from checkpoint-06000 and its optimizer state, peak lr 3e-5, 20-step warmup, on a cosine schedule over a 1,500-step budget, seed 20260922. General steps: 4 windows of 4,096 tokens from corpus-v3-muse (9,325 sequences, 38.2 M tokens, the same domains with the chat rendered in Runner's muse template). Policy steps: 8 agentic examples of up to 1,536 tokens, drawn in proportion to their weight from the 3,954 of 3,955 assembled examples that fit the window. On an agentic example only the assistant's positions carry loss. At a decision position (whom the turn addresses, which tool, which argument keys) the target is the factored teacher: Ornith-1.0-9B's policy mass over the prefix trie of the alternatives, the remaining share filled with the parent's own top-64. Everywhere else the target is the parent's top-64 on text the parent wrote. In the last 840 steps part of the policy text is the student's own: reasoning it wrote itself at training-set states, each position still trained toward the frozen parent's top-64, so the student learns where the parent would stop inside the student's own phrasing (on-policy distillation). The 1,440 steps ran as 11 consecutive segments from checkpoint-06000, each resumed from a checkpoint of the one before, with the lr schedule continued: steps 1 to 200 (train-w14b-envelope-a2, trainer b0b47549dc2ecf49…): 1 general step per policy step; steps 201 to 350 (train-w14b-envelope-a3, trainer ac3df6cd53c1c40a…): 1 general step per policy step; steps 351 to 450 (train-w14b-envelope-a4, trainer 4fca29a35dc14f97…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; steps 451 to 600 (train-w14b-envelope-a5, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; steps 601 to 720 (train-w14b-envelope-a6, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examples examples-muse-a6-mixed.jsonl (on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 721 to 840 (train-w14b-envelope-a7, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examples examples-muse-a7-mixed.jsonl (on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 841 to 960 (train-w14b-envelope-a8, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examples examples-muse-a8-mixed.jsonl (on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 961 to 1080 (train-w14b-envelope-a9, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examples examples-muse-a9-mixed.jsonl (on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 1081 to 1200 (train-w14b-envelope-a10, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examples examples-muse-a10-mixed.jsonl (on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 1201 to 1320 (train-w14b-envelope-a11, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examples examples-muse-a11-mixed.jsonl (on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 1321 to 1440 (train-w14b-envelope-a12, trainer f28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examples examples-muse-a12-mixed.jsonl (on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals). Through checkpoint-01440: 14,758,380 general positions on 901 steps and 772,568 assistant positions carrying loss on 539 policy steps. The gate was read at checkpoint-01440, step 1,440 of the 1,500 budgeted. The envelope phase ran as a chain of segments, each continuing from the previous attempt's gated checkpoint under a registered step ceiling (see the lineage above), and the released segment, attempt 12, ended at its ceiling.

How the agentic examples were made. Ornith-1.0-9B explored environment tasks 0 to 2007 as state trees in its own template. The Muse parent wrote every deliberation, every argument value and every final answer, steered to each branch by a private planning note that is never rendered into the example, with the argument keys prefix-forced to the policy's. Each tool was re-executed with the parent's values, and a path stops where the result differs from the one Ornith saw. A leakage filter (a word list plus any 3-word sequence shared with the note's template) rejects rationales that quote the note. Assembly produced examples on 1,889 tree paths; after de-duplication 3,974 examples remained, and the 19 from tasks 2000 to 2007 were removed before training because E3 scores tasks 2000 to 2059: 3,955 examples on 1,880 paths (SHA-256 6a3b651ccceaed11…). Before assembly resumed, 200 task ids (223 paths) were reserved for the gate and never assembled; none of them appears in the corpus (checked 2026-09-23).

Why

Width, not depth. Minitron (arXiv:2408.11796) found width pruning beats depth pruning at every token budget, depth-pruned models do not recover arithmetic generation even at about 100 B post-pruning tokens (arXiv:2602.01997), and the Muse-Glimmer surgery study's own N-wall (7 to 13 layers) says the same for this model. No smaller Muse-Glimmer existed, so the student had to be cut from the 30B.

A measured gap, not a copy. 98 M distillation tokens is about 0.1% of Minitron's 94 B. Parent-equivalence was never reachable at this budget and was never the claim. The gates were written before training: K1 (the cut keeps a working model), K2 (held-out KLD halves by day 3), K3 (margin-qualified top-1 ≥ 80% after week 1 for a model release). The lean written before the run was 60 to 75% at K3; it was wrong by 9 to 24 points.

Why an envelope phase. checkpoint-06000 passed K3 on the distribution it was trained on and was not a chat model: greedy, it looped after the first sentence and wrote to user. as text instead of the recipient token, because corpus v1 is a template-free render and it had never seen <|start|>, <|message|> or <|eot|>. Corpus v1 and v2 also held no tool-calling documents at all: the tool-calling rows share an instruction preamble with rows in the eval split, and the 64-token contamination filter dropped all 83,111 of them. Corpus v3 fixes that (below, and in Limits), so every tool-call behaviour this model has comes from the envelope phase. The teacher is factored because the parent's turn decisions are diffuse where Ornith's are sharp: on 23 pilot states, parent and Ornith agree on 91% of think-first decisions (TV 0.08) and on 62% of turn decisions after their own deliberation (TV 0.53).

Measured envelope

Four instruments, kept apart:

  1. Retention (E4, and K1 to K3). score_student.py: KLD student‖parent over the full vocabulary, top-1 agreement, margin-qualified top-1 (tie band 0.5 nats), top-8 overlap; the frozen BF16 parent from safetensors, the student's own head; the study's 11 held-out sequences, 45,056 positions, never trained on. PyTorch forward on the safetensors checkpoint.
  2. Closed loop (E3). agentic_closedloop.py --raw-muse: the executable environment's tasks from offset 2000 (seed 20260915), up to 6 tool calls per task, success when every ground-truth value appears in the final answer. Re-execution, not a judge.
  3. Wire format, calls, repetition (E1, E2, E5). envelope_preflight.py: prompts cut at user or tool-result events from the Ornith trees of the 200 reserved task ids, rendered in the muse template, half from states whose own trajectory went on to call a tool; greedy, 512-token cap; the special tokens read from logprobs.token_ids, because Runner's /v1/completions drops them from text.
  4. GGUF fidelity. Runner's scripts/kld-compare-raw.py, each quantised file against this model's own BF16 GGUF.

Every gate arm decodes greedily: E1, E2, E3 and E5 send temperature: 0 on every request. Runner has no Muse sampling preset, so a request that omits temperature is served with the generic preset (0.80, top-p 0.95, top-k 40, min-p 0.05). The numbers below describe greedy decoding, not that default.

The gate for the envelope phase was preregistered and approved on 2026-09-22 with its thresholds frozen; the headline order it fixes is the cost first, then the claim that is hardest to make.

Retention against the parent (E4, the cost)

checkpoint distillation tokens KLD top-1 margin-q top-8 gate
pruned, untrained (K1) 0 3.239 57.0% 60.2% 0.447 K1 top-1 ≥ 20%: pass
checkpoint-02500 (K2) 41.0 M 1.441 72.6% 77.0% 0.553 K2 KLD ≤ 1.62: pass
checkpoint-03500 57.3 M 1.155 75.3% 79.9% 0.578 progress read
checkpoint-04000 65.5 M 1.035 76.5% 81.1% 0.591 progress read
checkpoint-06000 (K3) 98.3 M 0.786 79.2% 83.9% 0.619 K3 margin-q ≥ 80%: pass
this model, envelope checkpoint-01440 (E4) 98.3 M + 14.8 M 0.762 79.2% 84.0% 0.623 E4 KLD ≤ 0.865 and margin-q ≥ 80%: pass

All six rows on the same 45,056 positions. The score files also carry a verdict field computed against the house quant bar; it reads FAIL on every row and is not the gate. Where the gap sits, per domain at checkpoint-06000: math word problems 0.30, code 0.34, rewriting 0.37, math chain-of-thought 0.48 at the good end; casual dialogue 1.03, system-prompted chat 1.06, long context 1.21, the Tulu mixture 1.49 at the other. The per-domain split was not re-measured after the envelope phase.

Closed-loop agentic tasks (E3)

E3 is the 60 tasks at offset 2000 (task-20260915-02000 to -02059), scored in full. Tasks 2000 to 2007 lie inside the range the Ornith trees covered; the 19 training examples drawn from them were removed before any envelope training (2026-09-23), so no arm is measured on training data.

arm file served solved terminated in budget schema-valid calls unnecessary calls calls / task thinking tokens / task
parent (reference) parent at Q4_K 60 of 60 60 73 of 73 0 1.22 212
student before the envelope (control) checkpoint-06000, BF16 0 of 60 0 0 of 0 0 0.00 0
this model (treated) checkpoint-01440, BF16 57 of 60 58 64 of 64 0 1.07 165

The rule: the treated model beats the control by at least 15 points and reaches at least 90% of the parent's success on the same tasks; with the parent at 60 of 60 that needs at least 54. The parent sits at the ceiling, so this arm can show the student approaching the parent and never passing it. All three arms ran on the same 24 GB GPU slice with runner main 0bfa2ad (owner decision, 2026-09-23): the parent with Runner's automatic fit, and both 14B BF16 files with 38 of their 52 blocks on the GPU (--gpu-layers 38). The control ran once, during attempt 11's gate, and was reused for this checkpoint's gate (the gate chain runs it only when its result file is missing). It measures the frozen checkpoint-06000, so the timing cannot move the number. The E3 rule is read in whole tasks: treated at least 54, treated minus control at least 9, and ten times treated at least nine times the parent's count.

Wire format, tool calls and repetition (E1, E2, E5)

arm prompts scored E1 turns well formed E2 calls valid E2 turns addressed to a tool / the user E5 turns looped terminator not reported
parent (reference, Q4_K) 100 100 of 100 53 of 53 53 / 47 0 of 100 0
student before the envelope (control) 200 0 of 200 0 of 11 11 / 0 (+189 other: to=example_tool_name.example_function_name

<atem:function_ 84, to=self get_weather 2, to=get_weather for Amsterdam. 1, to=get_weather to=Zurich atem:function_calls <atem:invoke 1, to=568*41/15

We need to compute 56841/15. 56841 = 233... 1, to=self atem:function_calls <atem:invoke name="today_add" 1, to=self get_weather_forecast for today 1-7 days forecast fo 1, to=self get_weather_forecast for today 1-7 days from today. 1, to=self to=lookup_contact to=lookup_contact to=lookup_conta 3, to=get_weather for Auckland. 1, to=4379 CHF atem:function_calls <atem:invoke name="curren 1, to=386*28/9

We need to compute 38628/9. 38628/9 = 10,848 1, to=get_weather for="Athens" and "get_weather_forecast" and 1, to=self get_stock_price atem:function_calls <atem:invoke 1, to=self get_weather atem:function_calls <atem:invoke name 1, to=lookup_contact to=self to=lookup_contact to=self to=look 8, to=839*57/6

Valid recipients: "self", "calculator.*", "d 1, to=self

atem:function_calls <atem:invoke name="today"> </ 1, to=864*88/3

864*88=78432 78432/3=2639.33

Two decimals: 2. 1, to=get_weather for Budapest. 1, to=self atem:function_calls <atem:invoke name="get_weathe 2, to=get_weather for="Lisbon" and "Stockholm" and "get_weathe 1, to=self get_weather_forecast for today in 330 days from tod 1, to=Zed Lund atem:function_calls <atem:invoke name="get_we 1, to=271*78/15

271*78=214... 214/15=14.67

Two decimals: 14. 1, to=962*58/5

We need to compute 96258/5. 96258 = 558...? 1, to=get_weather to=self get_weather to=self get_weather to=s 1, to=866 times 93 divided by 4

We need to compute 866 times 1, to=62 plus 2

The user is asking for the result of 62 plus 2, to=self atem:function_calls <atem:invoke name="currency_c 9, to=self with the date_add function. 6, to=find_flight to=self to=example_tool_name.example_functio 3, to=get_stock_price. The user is asking about the stock pric 1, to=self with the weather for Budapest. 1, to=self get_weather for="Warsaw" to="Copenhagen" to="Warsaw 1, to=self get_weather_forecast

The user is asking about the 1, to=self with tool_output name="calculator" and tool_output 1, to=example_tool_name.example_function_name

Valid recipie 2, to=unit_convert

atem:function_calls <atem:invoke name="un 1, to=self with the weather for Rome. 1, to=self convertToFahrenheit 180 c. The user wants to conver 1, to=self with the date 2026-09-15 and 138 days. The result i 1, to=self with the date 2026-09-15 and the number of days 257 1, to=calculator atem:function_calls <atem:invoke name="calc 2, to=self with the date 2026-09-15 and the result 2027-10-02. 1, to=example_tool_name.example_function_name <atem:function_c 1, to=self with the weather for Milo Dorn in Amsterdam. 1, to=lookup_contact to=lookup_contact to=lookup_contact to=lo 2, to=self get_weather for="Vienna" for="Paris" to="get_weathe 1, to=self atem:function_calls <atem:invoke name="get_stock_ 1, to=self to=currency_convert to=self to=currency_convert to= 1, to=self get_weather 1, to=51 plus 6

The user is asking for the result of 51 plus 1, to=get_weather for="Prague" and "Prague" is the city parame 1, to=get_weather atem:function_calls <atem:invoke name="get 3, to=self atem:function_calls <atem:invoke name="find_fligh 1, to=self convertToKg(173) 1, to=self with the user as the recipient. 1, to=self get_weather

Valid recipients: "self", "get_weath 1, to=self

atem:function_calls <atem:invoke name="calculator 1, to=self with the weather for Dublin. 1, to=self with the weather for Zurich. 1, to=self get_weather for Amsterdam 1, to=77 plus 5

The user is asking for the result of 77 plus 1, to=calculator to=calculator to=calculator to=calculator to= 1, to=self get_stock_price 1, to=self with the date and days from the user. 1, to=self get_weather for city in ["Toronto", "Budapest"] and 1, to=unit_convert to=self to=unit_convert to=unit_convert to= 1, to=currency_convert to=example_tool_name.example_function_n 2, to=calculator to=self to=calculator to=calculator to=calcul 1, to=self with the user as the recipient, the result is 1409. 1, to=self get_weather_forecast for the next 7 days. 1) | 192 of 200 | 0 | | this model (treated) | 200 | 199 of 200 | 99 of 99 | 99 / 100 (+1 other: to=self 1) | 3 of 200 | 0 |

E1 and E5 score the first turn the model writes. E2 scores the first turn it addresses to a tool or the user, after carrying up to 3 of its own reasoning turns, because Muse deliberates before it acts. Of this model's 200 scored prompts, the addressed turn came 185 after 1 reasoning turn, 15 after 0 reasoning turns. The parent's row is on 100 prompts, the control's and this model's on 200 drawn by the same seeded procedure. The instrument fails any call with an unknown or a missing required argument name, so every call counted valid also has matching names; the gate's second E2 rule (names match on at least 95% of validated calls) holds by construction and is not a separate measurement.

Reasoning length, a report and not a gate rule. E1 asks only that a reasoning turn closes. It does not ask that the turn is as short as the parent's. On the gate's non-call prompts where both the parent and this model reason before answering (43 prompts where both closed), this model's reasoning turn has a median of 48 tokens (90th percentile 86). The parent's has a median of 52 (90th percentile 87). This model's turn is the longer one on 13 of those 43 prompts. On 1 further prompt this model's reasoning did not close within the 512-token cap (the parent closed on every prompt). In use this means more reasoning tokens before each answer than the parent spends: the model writes the answer, then hedges before closing its reasoning.

The gate, rule by rule

arm rule, frozen 2026-09-22 needed measured verdict
E4 retention KLD ≤ 0.865 (K3 + 10%) and margin-q ≥ 80% both 0.762 / 84.0% pass
E3 closed loop control + 15 points and ≥ 90% of the parent ≥ 54 of 60 57 of 60 pass
E1 conformance ≥ 98% of turns well formed, strictly above the control ≥ 196 of 200 199 of 200 pass
E2 call validity ≥ 95% of calls parse and validate ≥ 95 of 99 99 of 99 pass
E5 repetition ≤ the control's rate and ≤ the parent's + 5 points ≤ 10 of 200 3 of 200 pass

GGUF files against this model's BF16

The quantised files are copies of this model, so the house bar applies to them, measured against this model's own BF16 GGUF and not against the parent: scripts/kld-compare-raw.py, both sides served by Runner, 500 held-out positions, tie band 0.5 nats. A file that misses the bar is not published. Of 500 positions, 3 are ones where the next token is a stop, in both files; none is one-sided (a stop in one file only) and none is left unscored, so every rate above has 500 as its denominator (runner main 0742af1).

file bytes mean KLD top-1 margin-qualified top-1 top-8 overlap bar
Xyntetik-Kvist-14B-BF16.gguf 28,903,141,184 reference
Xyntetik-Kvist-14B-Q8_0.gguf 15,363,224,384 0.0001 100.00% 100.00% 0.998 PASS
Xyntetik-Kvist-14B-Q5_0-mix.gguf 10,295,023,424 0.0012 99.60% 100.00% 0.980 PASS
Xyntetik-Kvist-14B-IQ4_NL-mix.gguf 7,612,384,384 0.0037 98.60% 100.00% 0.954 PASS

Why the smaller file is called Q5_0-mix and not Q4_K_M. It was requested as llama.cpp's Q4_K_M. 313 of 731 tensors fall back: their row width, 5,760, is not divisible by 256, which Q4_K needs, so llama-quantize's Q4_K_M recipe writes them as Q5_0 or Q8_0 instead. By bytes the file is 61.9% Q5_0, 13.4% Q4_K, 12.4% Q8_0 and 12.2% Q6_K, 5.69 bits per weight, so it is published as Q5_0-mix, the name for what it holds. Both files were made with llama.cpp b10353's llama-quantize (static CPU build, no importance matrix) from the release BF16 file, then put back into that file's tensor order.

The IQ4_NL-mix file (added 2026-09-28, at readers' request for a smaller file). It was asked for as a 3-bit file. Every 3-bit format needs rows divisible by 256, and only two tensor kinds here have them: ffn_down (10,240) and attn_output (3,072), 3.99 B parameters, written as IQ3_S. Every 5,760-wide tensor (attn_q/k/v, attn_gate, ffn_gate, ffn_up), the output head and the token embeddings are IQ4_NL, each set explicitly per tensor (no Q4_0). By bytes the file is 77.4% IQ4_NL and 22.5% IQ3_S, so it is named for its majority type. Output head: IQ4_NL, 654,635,520 bytes; token embeddings: IQ4_NL, 654,635,520 bytes. Unlike the other two files it uses an importance matrix: llama.cpp b10353 llama-imatrix on this model's Q8_0, over wikitext-103-raw-v1 validation (Salesforce/wikitext @ b08601e0, parquet SHA-256 204929b7ff9d6184…, text SHA-256 8ef749789ca06934…), 100 chunks of 512 tokens with --process-output; the imatrix file (SHA-256 9f1ae3674ca5446e…) is in evidence/gguf/iq4_nl-mix/. The calibration text is disjoint from the 500 evaluation positions. It does not fit an 8 GB card whole. Measured on 2026-09-28 on an RTX 3070 8 GB (Windows 11, a 7.60 GB video budget for the process) with the Runner v0.5.7 release binary, greedy: the automatic fit places 50 of 52 blocks on the GPU at a 2,048-token context (6.7 GB in VRAM) and 49 of 52 at 4,096 (6.6 GB), and runs the rest on the CPU, at about 4.8 generated tokens per second (4.87 and 4.73).

What IQ4_NL-mix loses, measured (2026-09-28). Against this model's BF16 on the 500 positions above, its per-position divergence is about 3.5 times Q5_0-mix's (median KLD 0.0017 against 0.0004; worst position 0.153 against 0.030). Its greedy choice changes at 7 of 500 positions (Q5_0-mix: 2, Q8_0: 0), each a near-tie in BF16 itself (top-two margin 0.01 to 0.28 nats), and the runner-up candidates move more (top-8 overlap 0.954 against 0.980). On the gate's 60 closed-loop tasks, served whole on the same 24 GB GPU slice with runner main 0bfa2ad, greedy and without the guard (a serving measurement, not a gate arm), it solves 57 of 60 against the BF16 gate's 57 (57 against 57 when a numeric task must have the right bolded answer); calc 3 of 5 against 3; 2 reasoning loops against 2; mean thinking tokens 160 against 165. Tasks BF16 solved and it did not: 02027 (stock). Tasks it solved and BF16 did not: 02055 (forecast). Per-task rows: evidence/gguf/iq4_nl-mix/iq4nl-mix-e3-gpu-closedloop.json.

Benchmark annex (a separate quantity)

Pending at release: HellaSwag (1,000 items, seed 20260911) and an MMLU-Pro subsample, paired, this model against the Muse parent and against Hermes-4-14B, all three through one instrument (Runner /v1/decide with rendering: continuation-v1, runner main 42707f7). It will be added to this card as a dated addition. Fidelity to the parent and benchmark equivalence are different quantities.

Disclosures, weakest first

  1. Two E3 numbers, for two questions. The gate's E3 figure, 57 of 60, is the registered rule: a task counts as solved when the true value appears anywhere in the reply. That rule cannot tell a correct answer from a correct intermediate next to a wrong one, so a second reading is kept for describing what the model does: a numeric task counts only when its bolded answer matches the truth. On the gate's 60 tasks this checkpoint scores 57 under both readings.

  2. Calculation tasks are the weak kind. Over 160 held-out tasks (the gate's 60, on GPU BF16, plus 100 discovery tasks from offset 2060, on CPU Q8_0) it solves 12 of 15 calc tasks with the right final answer. The previous attempt solved 8 of 15, and an earlier attempt of this study (attempt 9) solved 14 of 15. Attempt 9 was never gated on the GPU, so its figure comes from CPU Q8_0 runs; on that same CPU basis this checkpoint also solves 12 of 15. Of the three calc failures, one is a wrong answer and two are reasoning loops on a "two decimals" rounding request.

  3. Rounded-line corruption. From attempt 10 on, checkpoints of this study sometimes corrupted digits in the rounded answer line, e.g. "6 870.38" for 6,270.38. On this checkpoint, 17 of its passed numeric tasks over the 160 carry a bolded answer, and none of the 17 is wrong. That is a measured rate on 17 answers, not a guarantee.

  4. Runaways. 7 of the 160 tasks end in a reasoning loop that does not terminate in the budget: 2 calc, 3 forecast, 1 stock, 1 weather comparison. The previous attempt also had 7; the calc loops fell and new ones appeared on forecast. The loop guard (see Serve) can close such a turn, but it cannot supply an answer the model does not have, and on this checkpoint it did not end its one runaway on CPU.

  5. The control never completes the protocol. checkpoint-06000, before any agentic training, solves 0 of 60, terminates 0, makes no tool call, and has no pass where the truth is merely echoed from the request. So the margin, 57 over 0, says the envelope phase taught the protocol. It does not measure a gain over a model that could already do it.

  6. The CPU proxy. During the phase a CPU run at Q8_0 was used to stop a doomed E3 early. It was validated against the GPU BF16 gate twice, with the same count and no task scored differently: 51 = 51 on attempt 8 and 57 = 57 on attempt 11. On this checkpoint it read 58 against the gate's 57. One task, task-20260915-02055 (forecast), passes on CPU Q8_0 and loops on GPU BF16. Precision and device change together here, so this is not a general result about Q8_0.

  7. The held-out sets were reused adaptively. E1's 200 prompts, E3's 60 tasks and the 100 discovery tasks were read after each attempt to design the next. None of them was trained on: no training example comes from them, and the on-policy rollouts skip them. But none is unseen by the people who designed the recipe.

  8. Twelve gated attempts, two full passes.

    • Attempts: the envelope phase ran 12 attempts, each gated at the checkpoint E4 selected. Attempts 1 to 10 each failed at least one arm.
    • Attempt 11 (checkpoint-01320) passed all five and was held by the owner: E1 198, E2 99 of 99, E3 57 of 60 (56 by final answer), E4 KLD 0.7585, E5 3.
    • Attempt 12 was then run, and this checkpoint was gated. Before its data existed, the two were ordered by a registered rule: full pass, then final-answer-correct over the 160 tasks, then calc, then E1.
    • Result: 151 against 149, 12 against 8, 199 against 198, so this checkpoint ranks first. The margin is 2 tasks of 160, within noise. The registered order decided, not significance.
    • Fallback: a conditional fallback for attempt 12 (checkpoint-01420) was registered in case the gated checkpoint failed E1, E2 or E5. It was not used.
  9. A repaired defect in the lineage. Attempt 10's anti-repetition target also banned a calculator value that the model was legitimately restating, which damaged numeric answers. Attempt 11 skipped numeric loop onsets and cut damaged answers at the first wrong number. Attempt 12 added cuts at the padding decision and at wrong bolded numbers, plus 80 rollouts on calc answer turns. This checkpoint descends from attempt 10 through both repairs.

  10. Amendments. The preregistration was amended during the phase, each amendment dated and made before the data it governs. The full text is ENVELOPE-GATE-PREREG-2026-09-22.md in the training record. The main ones, all 2026:

    • 09-23 09:45: owner decisions for the deadline;
    • 09-23 13:45: every E3 arm moves to the GPU slice;
    • 09-23 16:45: the gate picks its checkpoint by E4;
    • 09-24 17:00: on-policy distillation registered as the next round;
    • 09-25 05:43: release only on a full pass;
    • 09-26 00:52: the CPU proxy, validated before use;
    • 09-26 09:28: deadline moved to 09-28 09:00;
    • 09-27 09:00: hold attempt 11, run attempt 12, ship the better full pass;
    • 09-27 10:22: E3 read in whole tasks;
    • 09-27 10:23: the conditional fallback.
  11. E3 by kind (this checkpoint, GPU BF16; the parent solves every task):

    kind solved kind solved
    calc 3 of 5 forecast 4 of 5
    calc_mental 3 of 3 missing_tool 5 of 5
    contact 9 of 9 no_tool 5 of 5
    contact_weather 5 of 5 stock 3 of 3
    currency 5 of 5 units 2 of 2
    dates 2 of 2 weather1 3 of 3
    flight 4 of 4 weather2 4 of 4
  12. Gate build against the pinned build. Every gate arm ran on runner main 0bfa2ad, with 38 of 52 blocks on a 24 GB GPU slice (--gpu-layers 38). This card pins Runner v0.5.7, which contains 0bfa2ad. No gate arm was re-run on v0.5.7. The loop-guard row above was measured on the v0.5.7 release binary.

Limits, read before quoting

  • Not parent-equivalent, and not claimed to be. KLD 0.762 and margin-qualified top-1 84.0% against the parent. The house bar for quantised copies does not apply to a pruned and distilled student. "No quality loss" is not claimed and is not supported by this evidence.
  • Parent-agreement is not a capability benchmark. The retention rows measure how closely next-token distributions track the parent on held-out text.
  • Tool calling, scoped. The claim rests on 99 calls from 200 prompts and 60 closed-loop tasks in one synthetic environment, where the parent solves 60. It says nothing about other tool sets, schemas or orchestration patterns, and none of it comes from the general distillation (corpus v1 and v2 held no tool-calling documents).
  • The agentic tasks are narrow. E3 is one synthetic environment with 14 task kinds (weather, forecasts, units, contacts, flights, currency, stocks, dates, arithmetic, and tasks that need no tool or a tool that is missing) and at most 6 calls; the E1, E2 and E5 prompts come from the same environment's trajectories. Nothing here measures long-horizon agents, real APIs, code execution, or the parent's published agentic benchmarks.
  • Tool calls were measured on the raw completions path. The gate rendered prompts with the study's muse renderer, verified byte-identical to Runner's template at runner main 676d3a7, and parsed the ATEM calls itself from /v1/completions. Runner's chat endpoint with a tools list was not part of the gate. A chat request with a tools list on Runner v0.5.7 returned a parsed call on 2026-09-28; it is not a measured rate. Served through Runner's constrained tool calling, a call cannot be emitted with a schema-required argument missing: the grammar withholds the closing tag until they are written (checked on v0.5.7, bd20e1f, on 2026-09-28; the runner session verified the mechanism on a fixture 2026-09-25). A generation cut by the token budget is completed with empty required values and reports finish_reason length, not tool_calls. E2 above measures free generation, so its number is the model's, not this path's.
  • The reference is the parent at Q4_K. E1, E2, E3 and E5 compare against Muse-Glimmer-30B quantised to Q4_K by Runner's own quantiser (the study scorer puts that file at KLD 0.01930 and margin-qualified top-1 99.12% from its BF16), not against the BF16 parent. E4 uses the BF16 parent.
  • The scored checkpoint and the served file. E4 was measured on the safetensors checkpoint through the study's PyTorch forward; E1, E2, E3 and E5 on the BF16 GGUF converted from it by llama.cpp b10353's converter. No row compares the two directly.
  • Runner builds. E1, E2, E3 and E5 ran on runner main 0bfa2ad; the serving measurements ran on 53b4deb and ac5418e. All three are contained in v0.5.7 (bd20e1f), and the changes after ac5418e (speculative-walk stop, buffered Responses status, envelope cut parity) do not touch the raw-completions path the arms used. E4 does not use Runner. Reproducing E1's terminator check needs a Runner that reports stop_token, which arrived on runner main at 0bfa2ad, after the v0.5.6 tag, and ships in v0.5.7.
  • Why the files are in block order. Runner v0.5.6 and earlier upload a partial offload as one contiguous file prefix, counted from byte 0. The converter's own order puts output.weight early and blocks 8, 9 and 26 out of place, so a converter-order BF16 on a card that cannot hold the whole file can run out of memory and fall back to the CPU. The published files are therefore rewritten into block order (token_embd, blocks 0 to 51, output_norm, output). Only tensor order and offsets change, and every tensor's bytes and every metadata field are verified identical. On this layout a 44-block offload uploads 22.8 GB instead of 25.2 GB. Runner main fixed the plan itself at 9b825fc, which ships in v0.5.7. Q8_0 and Q5_0-mix fit a 24 GB card whole and are not affected. Checked on v0.5.6 on 2026-09-28: the published BF16 file serves with partial offload on a 24 GB GPU slice and does not fall back to the CPU, 38 of 52 blocks (20.1 GB on the GPU) with --gpu-layers 38 and 45 of 52 (23.3 GB) with the automatic fit, the same placement as v0.5.7. The fallback applies to the converter's order, which is not published.
  • Reasoning strength is fixed at "high". Every agentic training example set "Reasoning strength: high." On attempt 5's checkpoint-00525 (an earlier checkpoint of the released run), 20 held-out answer turns gave byte-identical greedy output at "low" and "high" on 19. Runner's reasoning_strength request parameter does not shorten this model's reasoning. A serving-side reasoning budget is the lever.
  • Fixed-decimal padding. The model does not reliably pad a value to a fixed number of decimals when a tool returns fewer digits (asked for two decimals of 7970.5, it may write 7970.5, or deliberate about how to write 7970.50). A caller who needs a fixed format should format the number downstream rather than ask the model for it.
  • Text only. The parent's vision encoder is not part of this model; image input is not supported.
  • The contamination filter was weakened by design for corpus v3. It exempts, per domain, the 64-token windows that recur in at least 2% of 2,000 sampled documents of that domain: 166 windows for tool calling, 16 for casual dialogue, 3 for general knowledge, none elsewhere. A window carried by thousands of training documents is that domain's format and cannot fingerprint one eval item; that is the whole claim. Under the muse render the filter is also weaker at role boundaries, because the forbidden windows come from the plain-rendered eval split.
  • Selection data. The held-out split was read by the scorer six times (K1, K2, two progress reads, K3, E4) and never trained on. The E1, E2 and E5 prompts come from 200 task ids reserved before assembly, none of them in the envelope corpus. The trees covered tasks 0 to 2007, so they overlapped E3's 2000 to 2059 at 2000 to 2007; the 19 training examples from those tasks were removed before any envelope training (2026-09-23), and the on-policy rollouts skip 2000 to 2159. A scan of every policy file this checkpoint's lineage trained on (examples-muse-all.jsonl for attempts 2 to 5, then examples-muse-a6-mixed.jsonl to examples-muse-a12-mixed.jsonl) finds 1,779 distinct task ids, all below 2000: none from E3's 60 tasks, the 100 discovery tasks or E1's 200 held-out prompts.
  • Early stopping. The preregistered rule stops a run if E4's dev proxy rises at three consecutive evaluations. It was implemented in the trainer before step 200 and fired at step 200 in attempts 1 and 2. From attempt 3 on it was turned off: E4 was measured on every saved checkpoint instead, and the gate read the latest checkpoint inside E4's cap. On the released segment the dev proxy read 0.4967, 0.4962 and 0.4949 at steps 1,320, 1,350 and 1,400.
  • A short phase. 1,440 envelope steps, 14758380 general positions and 772568 loss-carrying agentic positions, on top of 98.3 M distillation tokens.
  • One model, one environment, one shared 24 GB GPU slice.

Reproduce

# 0. Parent: meta-models/Muse-Glimmer-30B @ a4e59da52a7bc87ae7251dd5545c0dd437c44b68, BF16 safetensors.
#    The surgery study's shared modules (mgcommon.py, run6_x1_score.py, build_corpus.py,
#    run9_build_corpus.py) are in code/deps/; the scripts expect them on the path they name.
# 1. Importance capture on corpus_run9's train split, then the width cut.
python prune_capture.py --sequences 64 --tokens 2048 --output prune-capture-64x2048.npz
python prune_width.py --capture prune-capture-64x2048.npz --hidden 5760 --ffn 10240 --heads 24 \
    --output muse-w14b-pruned
# 2. Corpora. v1: plain render, built before the boilerplate exemption existed.
#    v3: the muse render, exemption at 2% of 2,000 sampled documents per domain.
python build_student_corpus.py --output corpus-v1 --boilerplate-frac 0
python build_student_corpus.py --output corpus-v3-muse --render muse \
    --cap-tokens 3000000 --tulu-tokens 6000000 --wikitext-tokens 5000000
# 3. General distillation, 6,000 steps.
python train_student.py --student muse-w14b-pruned --corpus corpus-v1 --output train-w14b-v1 \
    --steps 6000 --windows 4 --tokens 4096 --lr 1e-4 --warmup 50 --save-every 500 --seed 20260915
# 4. K gates on the held-out split.
python score_student.py --student train-w14b-v1/checkpoint-06000 --name w14b-step6000 --out K3.json
# 5. Envelope corpus: Ornith state trees, then assembly with the parent writing the words.
#    Worker scripts with every offset: evidence/assembly/.
python agentic_trees.py --tokenizer <Ornith-1.0-9B tokenizer> --gpu off --tasks 50 --offset <o> \
    --out trees-cpu.jsonl
python agentic_assemble.py --trees trees-*.jsonl --external-server --port 58701 --offset <k> --stride 4 \
    --holdout holdout-tasks.json --out examples-muse-w<k>.jsonl
# 6. Envelope phase, one command per segment, each resuming the checkpoint the one before handed on
#    (the trainer at each segment's recorded SHA-256; see DISTILLATION.json).
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-v1/checkpoint-06000 \
    --corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a2 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --stop-at 200 --early-stop on
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a2/checkpoint-00200 \
    --corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a3 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 200 --stop-at 350 --early-stop off
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a3/checkpoint-00350 \
    --corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a4 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 350 --stop-at 450 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}'
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a4/checkpoint-00450 \
    --corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a5 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 450 --stop-at 600 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a6-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a5/checkpoint-00600's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
#  attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
#  prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a5/checkpoint-00600 \
    --corpus corpus-v3-muse --policy examples-muse-a6-mixed.jsonl --output train-w14b-envelope-a6 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 600 --stop-at 720 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a7-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a6/checkpoint-00720's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
#  attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
#  prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a6/checkpoint-00720 \
    --corpus corpus-v3-muse --policy examples-muse-a7-mixed.jsonl --output train-w14b-envelope-a7 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 720 --stop-at 840 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a8-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a7/checkpoint-00840's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
#  attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
#  prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a7/checkpoint-00840 \
    --corpus corpus-v3-muse --policy examples-muse-a8-mixed.jsonl --output train-w14b-envelope-a8 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 840 --stop-at 960 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a9-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a8/checkpoint-00960's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
#  attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
#  prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a8/checkpoint-00960 \
    --corpus corpus-v3-muse --policy examples-muse-a9-mixed.jsonl --output train-w14b-envelope-a9 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 960 --stop-at 1080 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a10-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a9/checkpoint-01080's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
#  attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
#  prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a9/checkpoint-01080 \
    --corpus corpus-v3-muse --policy examples-muse-a10-mixed.jsonl --output train-w14b-envelope-a10 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 1080 --stop-at 1200 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a11-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a10/checkpoint-01200's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
#  attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
#  prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a10/checkpoint-01200 \
    --corpus corpus-v3-muse --policy examples-muse-a11-mixed.jsonl --output train-w14b-envelope-a11 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 1200 --stop-at 1320 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a12-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a11/checkpoint-01320's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
#  attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
#  prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a11/checkpoint-01320 \
    --corpus corpus-v3-muse --policy examples-muse-a12-mixed.jsonl --output train-w14b-envelope-a12 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 1320 --stop-at 1440 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# 7. The gate, in the order evidence/gate/treated/gate-chain.sh runs it.
python score_student.py --student train-w14b-envelope-a12/checkpoint-01440 --name E4 --batch 2 --out E4.json
python convert_hf_to_gguf.py train-w14b-envelope-a12/checkpoint-01440 \
    --outfile Xyntetik-Kvist-14B-BF16.gguf --outtype bf16        # llama.cpp b10353
python agentic_closedloop.py --model Xyntetik-Kvist-14B-BF16.gguf --label treated --runner ./runner \
    --tokenizer <parent snapshot> --gpu auto --gpu-layers 38 --tasks 60 --offset 2000 --raw-muse \
    --out treated-e3.json                                         # runner main 0bfa2ad, 24 GB slice
runner -m Xyntetik-Kvist-14B-BF16.gguf --serve --no-tray --port 58740 --gpu auto --gpu-layers 38 \
    -t 48 -c 8192 --parallel 4 &                                  # runner main 0bfa2ad
python envelope_preflight.py --endpoint http://127.0.0.1:58740 --model Xyntetik-Kvist-14B-BF16.gguf \
    --trees trees-*.jsonl --holdout holdout-tasks.json --n 200 --max-tokens 512 --concurrency 4 \
    --arm treated --out treated-e1e2e5.json
# 8. Block-order the BF16, quantise with llama.cpp b10353, restore the order, and measure
#    each file against this model's BF16 (RUN-ON-ENVELOPE.sh in code/ runs all of it with its checks).
python gguf_blockorder.py converter-bf16.gguf Xyntetik-Kvist-14B-BF16.gguf --json layout-rewrite.json
llama-quantize Xyntetik-Kvist-14B-BF16.gguf q8_0.tmp.gguf Q8_0 16
llama-quantize Xyntetik-Kvist-14B-BF16.gguf q5mix.tmp.gguf Q4_K_M 16   # falls back to a Q5_0 majority
python gguf_blockorder.py q8_0.tmp.gguf Xyntetik-Kvist-14B-Q8_0.gguf --json q8_0-order.json \
    --order-from Xyntetik-Kvist-14B-BF16.gguf
python gguf_blockorder.py q5mix.tmp.gguf Xyntetik-Kvist-14B-Q5_0-mix.gguf --json q5mix-order.json \
    --order-from Xyntetik-Kvist-14B-BF16.gguf
python scripts/kld-compare-raw.py --model-a Xyntetik-Kvist-14B-Q8_0.gguf \
    --model-b Xyntetik-Kvist-14B-BF16.gguf --runner ./runner --corpus evidence/muse/harness-eval-split.txt \
    --max-positions 500 --out q8_0-vs-bf16.json

Training on a GPU is not bit-reproducible, so a rerun lands near these numbers and not on them; the SHA-256 values below identify exactly what was measured and published. The BF16 GGUF is a deterministic conversion of the safetensors checkpoint with the same converter. Code in code/, the preregistrations with their dates in evidence/preregistration/, every score, log and launch script in evidence/.

Provenance

  • Parent: meta-models/Muse-Glimmer-30B, revision a4e59da52a7bc87ae7251dd5545c0dd437c44b68, Apache-2.0. BF16 safetensors model-00001-of-00002.safetensors SHA-256 8eef61530e1283642c77ce2e6721feb5c6f348fa055c00e90f2844a136372694, model-00002-of-00002.safetensors SHA-256 b58cc2144ba1ba1af4420f67f4ca3ced7f09298510b80464cc75018a0be14381. Tokenizer unchanged: tokenizer.json SHA-256 c9dbee66967b58f3…, identical to the parent's.
  • Gate reference: the parent at Q4_K, parent_q4_k.gguf, 15,686,266,304 bytes, SHA-256 0956fcae4fda1e8f93f8784ace121accde2b8db4b00be7d3aa8dc926bf726075.
  • Policy teacher: Ornith-1.0-9B (deepreinforce-ai/Ornith-1.0-9B), served as ornith-1.0-9b-Q4_K_M.gguf, SHA-256 5720d1f671b4996481274fffe01868c3c36e87c135cc8538471cc7bd6087b106. Its policy mass shaped the targets at decision positions; no Ornith weights are in this model.
  • Pruned start: PRUNE.json SHA-256 0356aae40e33a54d…; calibration capture prune-capture-64x2048.npz SHA-256 6e4f0e30eebe1fd2….
  • General-distillation checkpoint (K3): checkpoint-06000; its BF16 GGUF, 28,903,141,184 bytes, SHA-256 9c7f7a289867c4ebf89d17aeb95be3ca52e2c9ab9f24f44bccc39392ff619180.
  • This model: checkpoint-01440 of the envelope run. Safetensors model-00000.safetensors 6e7602deb8371e0b1b9e2608c2e4414fe4a4bbd77f19bd835294a5dedf9c3e1a, model-00001.safetensors ca69b78b6f833cd887b428f5394894f79e4102afc332f91bed683c03972f0ce2, model-00002.safetensors 2e07b5fb9609dddd7cc4ee3354f10d05695274c06575c63eebbfd8b1c955e8d9, model-00003.safetensors 6f108a8e55e00dd9cda4400829217173c3696cc337c8d4cea252d3ce9b1ca30e. Xyntetik-Kvist-14B-BF16.gguf 28,903,141,184 bytes, SHA-256 daf422db9d0853d53ec8abb8fe9cd3f1e33ea53dffe7cca476225fbe1bc5d5c6. Xyntetik-Kvist-14B-Q8_0.gguf 15,363,224,384 bytes, SHA-256 cdb73543a6fc22f7faa5bbab0176ab2d5b1b590b46870d2d22f2a2510caaacb6. Xyntetik-Kvist-14B-Q5_0-mix.gguf 10,295,023,424 bytes, SHA-256 37af54b4cef4445bbbf23e8fbf0ceec8dc9b01b38dab0dc7f9e2bd2fd79f3a56. Xyntetik-Kvist-14B-IQ4_NL-mix.gguf 7,612,384,384 bytes, SHA-256 c74b27a5a39c894c4d1a1a2ae140de0cee125fc5cc9a7c96db9b7355bf8709a3.
  • Training data: corpus v1 tokens SHA-256 0731731533ca5ab7…, corpus-v3-muse tokens 34a37a0ba8086c11…, envelope examples cb1dd6fa5e6c516d…. Sources: HuggingFaceTB/smoltalk @ 5feaf2fd3ffca7c2…, allenai/tulu-3-sft-mixture @ b14afda60f1bbebe…, Salesforce/wikitext wikitext-103-raw-v1 train. Per-domain counts and hashes in evidence/corpus/; the token files themselves are not distributed.
  • Runner: v0.5.7 is runner main bd20e1f (v0.5.6 was e3781b9). Gate builds as listed under Limits.
  • Code: code/ at study commit 25cbda1; the record is in the shade repository, research/surgery/kvist-students/, at 4eeaa802.

Xyntetik-Kvist is this house's name for a new work. Muse Glimmer is its publisher's name, used here only to say where the weights came from.

Downloads last month
131
Safetensors
Model size
14B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Joakimpalm-Zen/Xyntetik-Kvist-14B

Finetuned
(43)
this model
Quantizations
1 model

Datasets used to train Joakimpalm-Zen/Xyntetik-Kvist-14B

Collection including Joakimpalm-Zen/Xyntetik-Kvist-14B

Papers for Joakimpalm-Zen/Xyntetik-Kvist-14B