Xyntetik-Kvist-14B
Release: built to be downloaded and served with Xyntetik Runner.
- Measured status against the parent: held-out KLD 0.762 and margin-qualified top-1 84.0% against the BF16 parent on 45,056 positions, its own row (the house quant bar does not apply); 57 of 60 held-out closed-loop tasks against the parent's 60; 99 of 99 tool calls valid; every preregistered envelope-gate arm passed.
- Parent model: meta-models/Muse-Glimmer-30B, cut by width to a dense 14B and distilled back from it; a new model, not a quantised copy
- Collection: Xyntetik Kvist: distilled agents
- Evidence: training record (dataset: preregistrations, K1 to K3 and envelope-gate records, corpus manifests, the envelope corpus, code), also copied under
evidence/here- Runner compatibility: Serves on Xyntetik Runner v0.5.7 or later; the gate arms E1, E2, E3 and E5 were measured on runner main 0bfa2ad and the serving measurements on 53b4deb, ac5418e and the v0.5.7 release binary; the first two are contained in v0.5.7. GGUF files are in Xyntetik-Kvist-14B-GGUF. Q8_0 (15.4 GB), Q5_0-mix (10.3 GB) and IQ4_NL-mix (7.6 GB) fit a 24 GB card whole and run GPU-resident; BF16 (28.9 GB) needs partial offload there (see Limits).
What it is, and is not. A 14B research student of Muse-Glimmer-30B, built for one job: running tool-calling agent loops on Xyntetik Runner from a 24 GB card. Its training budget is small: 6,000 distillation steps (98.3 M tokens, 162 hours) and 1,440 steps on agentic trajectories. It is not a drop-in replacement for Muse-Glimmer-30B, for Gemma, or for any general-purpose model: on held-out text it reads KLD 0.762 against its own parent, and its tool-calling result (57 of 60) holds for the study's environment and the muse wire format on Runner, nothing wider. Download it to serve agent loops on Runner and to read the evidence; do not download it expecting the parent's quality.
Muse-Glimmer-30B's language model, cut by width to a dense student of
14,443,781,760 parameters: all 52 layers kept, hidden width 6,656 → 5,760,
FFN 19,968 → 10,240, attention heads 32 → 24, KV heads 2. It was distilled back from
the frozen BF16 parent for 6,000 steps (98.3 M tokens, 162 hours), then trained for
1,440 more steps on agentic trajectories written in Runner's muse wire
format. Text only: the parent's 1.92 B-parameter vision encoder is not carried.
The cost first. On the study's 45,056-position held-out split this model reads KLD 0.762 and margin-qualified top-1 84.0% against the BF16 parent (top-1 79.2%). That is a student's own row. It is far from the house bar for quantised copies (KLD ≤ 0.05, margin-qualified top-1 ≥ 97%), which does not apply to a model with 14.4 B of the parent's 27.9 B language-model parameters and is not claimed. Closed-loop agentic tasks: 57 of 60 held-out environment tasks solved, scored by re-execution against ground truth, where the parent solves 60 and the student before the agentic phase solves 0.
What is claimed, and only this. Every arm of the preregistered envelope gate passed. The model writes Runner's muse wire format (199 of 200 first turns well formed), its tool calls parse and validate against the offered schema (99 of 99), and it solves 57 of 60 held-out tasks of the study's executable environment end to end (57 also when a numeric task counts only with the right bolded answer). That is a tool-calling claim scoped to that environment and that serving path; see Limits and the disclosures. Nothing about tool calling is claimed from the general distillation, whose corpora held no tool-calling documents.
Serve with xyntetik-runner
runner -m Xyntetik-Kvist-14B-Q8_0.gguf --serve --port 8080
Four files: BF16, the file the gate measured, and Q8_0, Q5_0-mix and IQ4_NL-mix, each measured against it below. The BF16 file is 28.9 GB and does not fit a 24 GB card whole; the gate served it with 38 of 52 layers on the GPU:
runner -m Xyntetik-Kvist-14B-BF16.gguf --serve --port 8080 --gpu auto --gpu-layers 38
For anything that may produce a tool call, keep "repeat_penalty": 1.0. That is Runner
v0.5.7's default for this model, but a penalty set by the client acts on the schema tokens a
call has to repeat (and only when sampling, at temperature above 0). Every gate number below
was measured greedy (temperature 0) with the penalty at 1.0.
Recommended: greedy, with the loop guard
Runner v0.5.7 and later can close a reasoning turn that has started
repeating itself (--loop-guard, off by default; per request loop_guard). This is the
recommended way to serve this model:
runner -m Xyntetik-Kvist-14B-Q8_0.gguf --serve --port 8080 --loop-guard
On the gate's 60 closed-loop tasks (this checkpoint at Q8_0 on CPU, greedy, the Runner v0.5.7 release binary) the guard leaves the score where it is: 58 of 60 with the guard and 58 of 60 without it on the same binary, failing the same two tasks. It intervened only in the one task that runs away without it, a two-decimals calc request, 5 times, and that task still did not terminate; it fired on no task that passed without it. Termination stayed at 98.3% and mean thinking tokens went from 154 to 147. On an earlier checkpoint of this study (attempt 8, runner 53b4deb) the same setting raised the score from 51 to 54 of 60 by ending 3 of its 6 runaways, so what the guard buys depends on how often a checkpoint loops; on this one its value is a bounded turn, not a better score.
The guard fixes the turn, not the task. When the model loops because it does not have the answer, closing the loop leaves a wrong answer. It also cannot see a loop that spans requests: in one measured task an earlier checkpoint of this study (attempt 8) had the answer in its reasoning and still called the same tool with the same argument three times, each turn well formed on its own. Only the caller, which holds the whole conversation, can catch that. This checkpoint passes that task (task-20260915-02037, two calls). Token repetition is a serving problem with a serving fix, the guard; deciding to answer instead of calling again is a training problem, and the fix here was training on post-tool rollouts.
Reasoning-channel sampling (--reasoning-temp) was measured too and is not
recommended for this model: on an earlier checkpoint of this study (attempt 8, runner 53b4deb, Q8_0 on CPU, the gate's 60 tasks), reasoning temperature 0.6 with top-p 0.9, min-p 0.05 and top-k 20 supplied in the request solved 49 and 47 tasks on two seeds against 51 greedy; termination was 92% and 87% against 90%, and both seeds turned the same calc and units tasks into loops.
The gate numbers below were all measured greedy without the guard. The guard is a serving recommendation, not part of the gate.
Quickstart
runner -m Xyntetik-Kvist-14B-Q8_0.gguf -p "The capital of France is" -n 32
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role": "user", "content": "Find me a flight from Stockholm to Berlin on 2026-10-14."}],
"tools": [{"type": "function", "function": {"name": "find_flight",
"description": "Search for available flights between two cities on a given date.",
"parameters": {"type": "object", "properties": {
"origin": {"type": "string", "description": "Departure city"},
"destination": {"type": "string", "description": "Arrival city"},
"date": {"type": "string", "description": "Departure date, YYYY-MM-DD"}},
"required": ["origin", "destination", "date"]}}}],
"max_tokens": 1024, "temperature": 0, "repeat_penalty": 1.0}'
On Runner v0.5.7 this returns a find_flight call with all three arguments filled
(finish_reason tool_calls). Give the model a calculator tool for arithmetic. Without one it
computes in its head and can be wrong: asked "What is 17% of 2,340?" at temperature 0 with no
tools, it answered 462.6 (the answer is 397.8).
What was changed, exactly
Three stages, each with its own record.
Width pruning (
PRUNE.json). Minitron-style activation importance, measured on 131,072 tokens (64 × 2,048) of the Muse-Glimmer surgery study's window-auditedcorpus_run9training split, eval split never read: hidden channels chosen globally by summed mean |x|, FFN neurons per layer by mean |h| at thedown_projinput, attention heads per KV group by the mean norm of theiro_projinput. The embedding and the LM head are cut on the hidden axis with the same channel set. No training.PRUNE.jsonholds every kept index, the calibration capture's SHA-256 and a hash of every tensor written.parent language model this model embedding 1,344,831,488 1,163,796,480 52 decoder blocks 25,165,117,952 12,116,188,800 LM head 1,344,831,488 1,163,796,480 total 27,854,780,928 14,443,781,760 (51.9% kept) Unchanged: the tokenizer (202,048 entries,
tokenizer.jsonidentical to the parent's), the chat template, head dimension 128, the attention pattern (three sliding-window layers of 2,048 tokens to one global), the logit softcap.General distillation (
DISTILLATION.json,general_kd). 6,000 steps of 4 windows × 4,095 predicted positions, 98.28 M tokens: 40% of corpus v1 (59,946 sequences of 4,096 tokens, 245.5 M tokens, the study's 11 chat and task domains from SmolTalk and the Tulu 3 SFT mixture plus Wikitext-103, plain render, single pass). Objective: top-64 forward KL to the frozen BF16 parent with a tail bucket, through the parent's softcapped logits. All 52 blocks and the final norm trainable; embedding and LM head frozen at their pruned values. AdamW (0.9, 0.95), peak lr 1e-4, 50-step warmup, cosine to 1e-5, per-block clip 1.0, one block resident on a 24 GB GPU slice with fp32 master weights on the host. 583,591 s of wall time (162.1 h).Envelope phase (
DISTILLATION.json,envelope). 1,440 steps from checkpoint-06000 and its optimizer state, peak lr 3e-5, 20-step warmup, on a cosine schedule over a 1,500-step budget, seed 20260922. General steps: 4 windows of 4,096 tokens fromcorpus-v3-muse(9,325 sequences, 38.2 M tokens, the same domains with the chat rendered in Runner'smusetemplate). Policy steps: 8 agentic examples of up to 1,536 tokens, drawn in proportion to their weight from the 3,954 of 3,955 assembled examples that fit the window. On an agentic example only the assistant's positions carry loss. At a decision position (whom the turn addresses, which tool, which argument keys) the target is the factored teacher: Ornith-1.0-9B's policy mass over the prefix trie of the alternatives, the remaining share filled with the parent's own top-64. Everywhere else the target is the parent's top-64 on text the parent wrote. In the last 840 steps part of the policy text is the student's own: reasoning it wrote itself at training-set states, each position still trained toward the frozen parent's top-64, so the student learns where the parent would stop inside the student's own phrasing (on-policy distillation). The 1,440 steps ran as 11 consecutive segments from checkpoint-06000, each resumed from a checkpoint of the one before, with the lr schedule continued: steps 1 to 200 (train-w14b-envelope-a2, trainerb0b47549dc2ecf49…): 1 general step per policy step; steps 201 to 350 (train-w14b-envelope-a3, trainerac3df6cd53c1c40a…): 1 general step per policy step; steps 351 to 450 (train-w14b-envelope-a4, trainer4fca29a35dc14f97…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; steps 451 to 600 (train-w14b-envelope-a5, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; steps 601 to 720 (train-w14b-envelope-a6, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examplesexamples-muse-a6-mixed.jsonl(on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 721 to 840 (train-w14b-envelope-a7, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examplesexamples-muse-a7-mixed.jsonl(on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 841 to 960 (train-w14b-envelope-a8, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examplesexamples-muse-a8-mixed.jsonl(on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 961 to 1080 (train-w14b-envelope-a9, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examplesexamples-muse-a9-mixed.jsonl(on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 1081 to 1200 (train-w14b-envelope-a10, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examplesexamples-muse-a10-mixed.jsonl(on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 1201 to 1320 (train-w14b-envelope-a11, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examplesexamples-muse-a11-mixed.jsonl(on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals); steps 1321 to 1440 (train-w14b-envelope-a12, trainerf28323e0fbfd42ba…): 2 general steps per policy step; sampling weights turn:answer ×2, flight ×2; terminator positions weighted ×10; policy examplesexamples-muse-a12-mixed.jsonl(on-policy: the student's own reasoning, parent targets, mixed 1:1 with the originals). Through checkpoint-01440: 14,758,380 general positions on 901 steps and 772,568 assistant positions carrying loss on 539 policy steps. The gate was read at checkpoint-01440, step 1,440 of the 1,500 budgeted. The envelope phase ran as a chain of segments, each continuing from the previous attempt's gated checkpoint under a registered step ceiling (see the lineage above), and the released segment, attempt 12, ended at its ceiling.
How the agentic examples were made. Ornith-1.0-9B explored environment tasks
0 to 2007 as state trees in its own template. The Muse parent wrote every deliberation,
every argument value and every final answer, steered to each branch by a private
planning note that is never rendered into the example, with the argument keys
prefix-forced to the policy's. Each tool was re-executed with the parent's values, and a
path stops where the result differs from the one Ornith saw. A leakage filter (a word
list plus any 3-word sequence shared with the note's template) rejects rationales that
quote the note. Assembly produced examples on 1,889 tree paths; after de-duplication
3,974 examples remained, and the 19 from tasks 2000 to 2007 were removed before training
because E3 scores tasks 2000 to 2059: 3,955 examples on 1,880 paths (SHA-256 6a3b651ccceaed11…). Before assembly resumed, 200 task
ids (223 paths) were reserved for the gate and never assembled; none of them appears in
the corpus (checked 2026-09-23).
Why
Width, not depth. Minitron (arXiv:2408.11796) found width pruning beats depth pruning at every token budget, depth-pruned models do not recover arithmetic generation even at about 100 B post-pruning tokens (arXiv:2602.01997), and the Muse-Glimmer surgery study's own N-wall (7 to 13 layers) says the same for this model. No smaller Muse-Glimmer existed, so the student had to be cut from the 30B.
A measured gap, not a copy. 98 M distillation tokens is about 0.1% of Minitron's 94 B. Parent-equivalence was never reachable at this budget and was never the claim. The gates were written before training: K1 (the cut keeps a working model), K2 (held-out KLD halves by day 3), K3 (margin-qualified top-1 ≥ 80% after week 1 for a model release). The lean written before the run was 60 to 75% at K3; it was wrong by 9 to 24 points.
Why an envelope phase. checkpoint-06000 passed K3 on the distribution it was
trained on and was not a chat model: greedy, it looped after the first sentence and
wrote to user. as text instead of the recipient token, because corpus v1 is a
template-free render and it had never seen <|start|>, <|message|> or <|eot|>.
Corpus v1 and v2 also held no tool-calling documents at all: the tool-calling rows
share an instruction preamble with rows in the eval split, and the 64-token contamination
filter dropped all 83,111 of them. Corpus v3 fixes that (below, and in Limits), so every
tool-call behaviour this model has comes from the envelope phase. The teacher is
factored because the parent's turn decisions are diffuse where Ornith's are sharp: on 23
pilot states, parent and Ornith agree on 91% of think-first decisions (TV 0.08) and on
62% of turn decisions after their own deliberation (TV 0.53).
Measured envelope
Four instruments, kept apart:
- Retention (E4, and K1 to K3).
score_student.py: KLD student‖parent over the full vocabulary, top-1 agreement, margin-qualified top-1 (tie band 0.5 nats), top-8 overlap; the frozen BF16 parent from safetensors, the student's own head; the study's 11 held-out sequences, 45,056 positions, never trained on. PyTorch forward on the safetensors checkpoint. - Closed loop (E3).
agentic_closedloop.py --raw-muse: the executable environment's tasks from offset 2000 (seed 20260915), up to 6 tool calls per task, success when every ground-truth value appears in the final answer. Re-execution, not a judge. - Wire format, calls, repetition (E1, E2, E5).
envelope_preflight.py: prompts cut at user or tool-result events from the Ornith trees of the 200 reserved task ids, rendered in themusetemplate, half from states whose own trajectory went on to call a tool; greedy, 512-token cap; the special tokens read fromlogprobs.token_ids, because Runner's/v1/completionsdrops them fromtext. - GGUF fidelity. Runner's
scripts/kld-compare-raw.py, each quantised file against this model's own BF16 GGUF.
Every gate arm decodes greedily: E1, E2, E3 and E5 send temperature: 0 on every request.
Runner has no Muse sampling preset, so a request that omits temperature is served with the
generic preset (0.80, top-p 0.95, top-k 40, min-p 0.05). The numbers below describe greedy
decoding, not that default.
The gate for the envelope phase was preregistered and approved on 2026-09-22 with its thresholds frozen; the headline order it fixes is the cost first, then the claim that is hardest to make.
Retention against the parent (E4, the cost)
| checkpoint | distillation tokens | KLD | top-1 | margin-q | top-8 | gate |
|---|---|---|---|---|---|---|
| pruned, untrained (K1) | 0 | 3.239 | 57.0% | 60.2% | 0.447 | K1 top-1 ≥ 20%: pass |
| checkpoint-02500 (K2) | 41.0 M | 1.441 | 72.6% | 77.0% | 0.553 | K2 KLD ≤ 1.62: pass |
| checkpoint-03500 | 57.3 M | 1.155 | 75.3% | 79.9% | 0.578 | progress read |
| checkpoint-04000 | 65.5 M | 1.035 | 76.5% | 81.1% | 0.591 | progress read |
| checkpoint-06000 (K3) | 98.3 M | 0.786 | 79.2% | 83.9% | 0.619 | K3 margin-q ≥ 80%: pass |
| this model, envelope checkpoint-01440 (E4) | 98.3 M + 14.8 M | 0.762 | 79.2% | 84.0% | 0.623 | E4 KLD ≤ 0.865 and margin-q ≥ 80%: pass |
All six rows on the same 45,056 positions. The score files also carry a verdict field
computed against the house quant bar; it reads FAIL on every row and is not the gate.
Where the gap sits, per domain at checkpoint-06000: math word problems 0.30, code 0.34,
rewriting 0.37, math chain-of-thought 0.48 at the good end; casual dialogue 1.03,
system-prompted chat 1.06, long context 1.21, the Tulu mixture 1.49 at the other. The
per-domain split was not re-measured after the envelope phase.
Closed-loop agentic tasks (E3)
E3 is the 60 tasks at offset 2000 (task-20260915-02000 to -02059), scored in full. Tasks 2000 to 2007 lie inside the range the Ornith trees covered; the 19 training examples drawn from them were removed before any envelope training (2026-09-23), so no arm is measured on training data.
| arm | file served | solved | terminated in budget | schema-valid calls | unnecessary calls | calls / task | thinking tokens / task |
|---|---|---|---|---|---|---|---|
| parent (reference) | parent at Q4_K | 60 of 60 | 60 | 73 of 73 | 0 | 1.22 | 212 |
| student before the envelope (control) | checkpoint-06000, BF16 | 0 of 60 | 0 | 0 of 0 | 0 | 0.00 | 0 |
| this model (treated) | checkpoint-01440, BF16 | 57 of 60 | 58 | 64 of 64 | 0 | 1.07 | 165 |
The rule: the treated model beats the control by at least 15 points and reaches at
least 90% of the parent's success on the same tasks; with the parent at
60 of 60 that needs at least
54. The parent sits at the ceiling, so this arm can show the
student approaching the parent and never passing it. All three arms ran on the same
24 GB GPU slice with runner main 0bfa2ad (owner decision, 2026-09-23): the parent with
Runner's automatic fit, and both 14B BF16 files with 38 of their 52 blocks on the GPU
(--gpu-layers 38). The control ran once, during attempt 11's gate, and was reused for this
checkpoint's gate (the gate chain runs it only when its result file is missing). It measures the
frozen checkpoint-06000, so the timing cannot move the number. The E3 rule is read in whole
tasks: treated at least 54, treated minus control at least 9, and ten times treated at least
nine times the parent's count.
Wire format, tool calls and repetition (E1, E2, E5)
| arm | prompts scored | E1 turns well formed | E2 calls valid | E2 turns addressed to a tool / the user | E5 turns looped | terminator not reported |
|---|---|---|---|---|---|---|
| parent (reference, Q4_K) | 100 | 100 of 100 | 53 of 53 | 53 / 47 | 0 of 100 | 0 |
| student before the envelope (control) | 200 | 0 of 200 | 0 of 11 | 11 / 0 (+189 other: to=example_tool_name.example_function_name |
<atem:function_ 84, to=self get_weather 2, to=get_weather for Amsterdam. 1, to=get_weather to=Zurich atem:function_calls <atem:invoke 1, to=568*41/15
We need to compute 56841/15. 56841 = 233... 1, to=self atem:function_calls <atem:invoke name="today_add" 1, to=self get_weather_forecast for today 1-7 days forecast fo 1, to=self get_weather_forecast for today 1-7 days from today. 1, to=self to=lookup_contact to=lookup_contact to=lookup_conta 3, to=get_weather for Auckland. 1, to=4379 CHF atem:function_calls <atem:invoke name="curren 1, to=386*28/9
We need to compute 38628/9. 38628/9 = 10,848 1, to=get_weather for="Athens" and "get_weather_forecast" and 1, to=self get_stock_price atem:function_calls <atem:invoke 1, to=self get_weather atem:function_calls <atem:invoke name 1, to=lookup_contact to=self to=lookup_contact to=self to=look 8, to=839*57/6
Valid recipients: "self", "calculator.*", "d 1, to=self
atem:function_calls <atem:invoke name="today"> </ 1, to=864*88/3
864*88=78432 78432/3=2639.33
Two decimals: 2. 1, to=get_weather for Budapest. 1, to=self atem:function_calls <atem:invoke name="get_weathe 2, to=get_weather for="Lisbon" and "Stockholm" and "get_weathe 1, to=self get_weather_forecast for today in 330 days from tod 1, to=Zed Lund atem:function_calls <atem:invoke name="get_we 1, to=271*78/15
271*78=214... 214/15=14.67
Two decimals: 14. 1, to=962*58/5
We need to compute 96258/5. 96258 = 558...? 1, to=get_weather to=self get_weather to=self get_weather to=s 1, to=866 times 93 divided by 4
We need to compute 866 times 1, to=62 plus 2
The user is asking for the result of 62 plus 2, to=self atem:function_calls <atem:invoke name="currency_c 9, to=self with the date_add function. 6, to=find_flight to=self to=example_tool_name.example_functio 3, to=get_stock_price. The user is asking about the stock pric 1, to=self with the weather for Budapest. 1, to=self get_weather for="Warsaw" to="Copenhagen" to="Warsaw 1, to=self get_weather_forecast
The user is asking about the 1, to=self with tool_output name="calculator" and tool_output 1, to=example_tool_name.example_function_name
Valid recipie 2, to=unit_convert
atem:function_calls <atem:invoke name="un 1, to=self with the weather for Rome. 1, to=self convertToFahrenheit 180 c. The user wants to conver 1, to=self with the date 2026-09-15 and 138 days. The result i 1, to=self with the date 2026-09-15 and the number of days 257 1, to=calculator atem:function_calls <atem:invoke name="calc 2, to=self with the date 2026-09-15 and the result 2027-10-02. 1, to=example_tool_name.example_function_name <atem:function_c 1, to=self with the weather for Milo Dorn in Amsterdam. 1, to=lookup_contact to=lookup_contact to=lookup_contact to=lo 2, to=self get_weather for="Vienna" for="Paris" to="get_weathe 1, to=self atem:function_calls <atem:invoke name="get_stock_ 1, to=self to=currency_convert to=self to=currency_convert to= 1, to=self get_weather 1, to=51 plus 6
The user is asking for the result of 51 plus 1, to=get_weather for="Prague" and "Prague" is the city parame 1, to=get_weather atem:function_calls <atem:invoke name="get 3, to=self atem:function_calls <atem:invoke name="find_fligh 1, to=self convertToKg(173) 1, to=self with the user as the recipient. 1, to=self get_weather
Valid recipients: "self", "get_weath 1, to=self
atem:function_calls <atem:invoke name="calculator 1, to=self with the weather for Dublin. 1, to=self with the weather for Zurich. 1, to=self get_weather for Amsterdam 1, to=77 plus 5
The user is asking for the result of 77 plus 1, to=calculator to=calculator to=calculator to=calculator to= 1, to=self get_stock_price 1, to=self with the date and days from the user. 1, to=self get_weather for city in ["Toronto", "Budapest"] and 1, to=unit_convert to=self to=unit_convert to=unit_convert to= 1, to=currency_convert to=example_tool_name.example_function_n 2, to=calculator to=self to=calculator to=calculator to=calcul 1, to=self with the user as the recipient, the result is 1409. 1, to=self get_weather_forecast for the next 7 days. 1) | 192 of 200 | 0 | | this model (treated) | 200 | 199 of 200 | 99 of 99 | 99 / 100 (+1 other: to=self 1) | 3 of 200 | 0 |
E1 and E5 score the first turn the model writes. E2 scores the first turn it addresses to a tool or the user, after carrying up to 3 of its own reasoning turns, because Muse deliberates before it acts. Of this model's 200 scored prompts, the addressed turn came 185 after 1 reasoning turn, 15 after 0 reasoning turns. The parent's row is on 100 prompts, the control's and this model's on 200 drawn by the same seeded procedure. The instrument fails any call with an unknown or a missing required argument name, so every call counted valid also has matching names; the gate's second E2 rule (names match on at least 95% of validated calls) holds by construction and is not a separate measurement.
Reasoning length, a report and not a gate rule. E1 asks only that a reasoning turn closes. It does not ask that the turn is as short as the parent's. On the gate's non-call prompts where both the parent and this model reason before answering (43 prompts where both closed), this model's reasoning turn has a median of 48 tokens (90th percentile 86). The parent's has a median of 52 (90th percentile 87). This model's turn is the longer one on 13 of those 43 prompts. On 1 further prompt this model's reasoning did not close within the 512-token cap (the parent closed on every prompt). In use this means more reasoning tokens before each answer than the parent spends: the model writes the answer, then hedges before closing its reasoning.
The gate, rule by rule
| arm | rule, frozen 2026-09-22 | needed | measured | verdict |
|---|---|---|---|---|
| E4 retention | KLD ≤ 0.865 (K3 + 10%) and margin-q ≥ 80% | both | 0.762 / 84.0% | pass |
| E3 closed loop | control + 15 points and ≥ 90% of the parent | ≥ 54 of 60 | 57 of 60 | pass |
| E1 conformance | ≥ 98% of turns well formed, strictly above the control | ≥ 196 of 200 | 199 of 200 | pass |
| E2 call validity | ≥ 95% of calls parse and validate | ≥ 95 of 99 | 99 of 99 | pass |
| E5 repetition | ≤ the control's rate and ≤ the parent's + 5 points | ≤ 10 of 200 | 3 of 200 | pass |
GGUF files against this model's BF16
The quantised files are copies of this model, so the house bar applies to them, measured
against this model's own BF16 GGUF and not against the parent:
scripts/kld-compare-raw.py, both sides served by Runner, 500 held-out
positions, tie band 0.5 nats. A file that misses the bar is not published.
Of 500 positions, 3 are ones where the next token is a stop, in both files; none is one-sided (a stop in one file only) and none is left unscored, so every rate above has 500 as its denominator (runner main 0742af1).
| file | bytes | mean KLD | top-1 | margin-qualified top-1 | top-8 overlap | bar |
|---|---|---|---|---|---|---|
| Xyntetik-Kvist-14B-BF16.gguf | 28,903,141,184 | reference | ||||
| Xyntetik-Kvist-14B-Q8_0.gguf | 15,363,224,384 | 0.0001 | 100.00% | 100.00% | 0.998 | PASS |
| Xyntetik-Kvist-14B-Q5_0-mix.gguf | 10,295,023,424 | 0.0012 | 99.60% | 100.00% | 0.980 | PASS |
| Xyntetik-Kvist-14B-IQ4_NL-mix.gguf | 7,612,384,384 | 0.0037 | 98.60% | 100.00% | 0.954 | PASS |
Why the smaller file is called Q5_0-mix and not Q4_K_M. It was requested as llama.cpp's
Q4_K_M. 313 of 731 tensors fall back: their row width, 5,760, is not divisible by 256, which Q4_K needs, so llama-quantize's Q4_K_M recipe writes them as Q5_0 or Q8_0 instead. By bytes the file is 61.9% Q5_0, 13.4% Q4_K, 12.4% Q8_0 and 12.2% Q6_K, 5.69 bits per weight, so it is published as Q5_0-mix, the name for what it holds. Both files were made with llama.cpp b10353's
llama-quantize (static CPU build, no importance matrix) from the release BF16 file, then
put back into that file's tensor order.
The IQ4_NL-mix file (added 2026-09-28, at readers' request for a smaller file). It was asked for as a 3-bit file. Every 3-bit format needs rows divisible by 256, and only two tensor kinds here have them: ffn_down (10,240) and attn_output (3,072), 3.99 B parameters, written as IQ3_S. Every 5,760-wide tensor (attn_q/k/v, attn_gate, ffn_gate, ffn_up), the output head and the token embeddings are IQ4_NL, each set explicitly per tensor (no Q4_0). By bytes the file is 77.4% IQ4_NL and 22.5% IQ3_S, so it is named for its majority type. Output head: IQ4_NL, 654,635,520 bytes; token embeddings: IQ4_NL, 654,635,520 bytes. Unlike the other two files it uses an importance matrix: llama.cpp b10353 llama-imatrix on this model's Q8_0, over wikitext-103-raw-v1 validation (Salesforce/wikitext @ b08601e0, parquet SHA-256 204929b7ff9d6184…, text SHA-256 8ef749789ca06934…), 100 chunks of 512 tokens with --process-output; the imatrix file (SHA-256 9f1ae3674ca5446e…) is in evidence/gguf/iq4_nl-mix/. The calibration text is disjoint from the 500 evaluation positions. It does not fit an 8 GB card whole. Measured on 2026-09-28 on an RTX 3070 8 GB (Windows 11, a 7.60 GB video budget for the process) with the Runner v0.5.7 release binary, greedy: the automatic fit places 50 of 52 blocks on the GPU at a 2,048-token context (6.7 GB in VRAM) and 49 of 52 at 4,096 (6.6 GB), and runs the rest on the CPU, at about 4.8 generated tokens per second (4.87 and 4.73).
What IQ4_NL-mix loses, measured (2026-09-28). Against this model's BF16 on the 500 positions above, its per-position divergence is about 3.5 times Q5_0-mix's (median KLD 0.0017 against 0.0004; worst position 0.153 against 0.030). Its greedy choice changes at 7 of 500 positions (Q5_0-mix: 2, Q8_0: 0), each a near-tie in BF16 itself (top-two margin 0.01 to 0.28 nats), and the runner-up candidates move more (top-8 overlap 0.954 against 0.980). On the gate's 60 closed-loop tasks, served whole on the same 24 GB GPU slice with runner main 0bfa2ad, greedy and without the guard (a serving measurement, not a gate arm), it solves 57 of 60 against the BF16 gate's 57 (57 against 57 when a numeric task must have the right bolded answer); calc 3 of 5 against 3; 2 reasoning loops against 2; mean thinking tokens 160 against 165. Tasks BF16 solved and it did not: 02027 (stock). Tasks it solved and BF16 did not: 02055 (forecast). Per-task rows: evidence/gguf/iq4_nl-mix/iq4nl-mix-e3-gpu-closedloop.json.
Benchmark annex (a separate quantity)
Pending at release: HellaSwag (1,000 items, seed 20260911) and an MMLU-Pro subsample, paired, this model against the Muse parent and against Hermes-4-14B, all three through one instrument (Runner /v1/decide with rendering: continuation-v1, runner main 42707f7). It will be added to this card as a dated addition. Fidelity to the parent and benchmark equivalence are different quantities.
Disclosures, weakest first
Two E3 numbers, for two questions. The gate's E3 figure, 57 of 60, is the registered rule: a task counts as solved when the true value appears anywhere in the reply. That rule cannot tell a correct answer from a correct intermediate next to a wrong one, so a second reading is kept for describing what the model does: a numeric task counts only when its bolded answer matches the truth. On the gate's 60 tasks this checkpoint scores 57 under both readings.
Calculation tasks are the weak kind. Over 160 held-out tasks (the gate's 60, on GPU BF16, plus 100 discovery tasks from offset 2060, on CPU Q8_0) it solves 12 of 15 calc tasks with the right final answer. The previous attempt solved 8 of 15, and an earlier attempt of this study (attempt 9) solved 14 of 15. Attempt 9 was never gated on the GPU, so its figure comes from CPU Q8_0 runs; on that same CPU basis this checkpoint also solves 12 of 15. Of the three calc failures, one is a wrong answer and two are reasoning loops on a "two decimals" rounding request.
Rounded-line corruption. From attempt 10 on, checkpoints of this study sometimes corrupted digits in the rounded answer line, e.g. "6 870.38" for 6,270.38. On this checkpoint, 17 of its passed numeric tasks over the 160 carry a bolded answer, and none of the 17 is wrong. That is a measured rate on 17 answers, not a guarantee.
Runaways. 7 of the 160 tasks end in a reasoning loop that does not terminate in the budget: 2 calc, 3 forecast, 1 stock, 1 weather comparison. The previous attempt also had 7; the calc loops fell and new ones appeared on forecast. The loop guard (see Serve) can close such a turn, but it cannot supply an answer the model does not have, and on this checkpoint it did not end its one runaway on CPU.
The control never completes the protocol. checkpoint-06000, before any agentic training, solves 0 of 60, terminates 0, makes no tool call, and has no pass where the truth is merely echoed from the request. So the margin, 57 over 0, says the envelope phase taught the protocol. It does not measure a gain over a model that could already do it.
The CPU proxy. During the phase a CPU run at Q8_0 was used to stop a doomed E3 early. It was validated against the GPU BF16 gate twice, with the same count and no task scored differently: 51 = 51 on attempt 8 and 57 = 57 on attempt 11. On this checkpoint it read 58 against the gate's 57. One task, task-20260915-02055 (forecast), passes on CPU Q8_0 and loops on GPU BF16. Precision and device change together here, so this is not a general result about Q8_0.
The held-out sets were reused adaptively. E1's 200 prompts, E3's 60 tasks and the 100 discovery tasks were read after each attempt to design the next. None of them was trained on: no training example comes from them, and the on-policy rollouts skip them. But none is unseen by the people who designed the recipe.
Twelve gated attempts, two full passes.
- Attempts: the envelope phase ran 12 attempts, each gated at the checkpoint E4 selected. Attempts 1 to 10 each failed at least one arm.
- Attempt 11 (checkpoint-01320) passed all five and was held by the owner: E1 198, E2 99 of 99, E3 57 of 60 (56 by final answer), E4 KLD 0.7585, E5 3.
- Attempt 12 was then run, and this checkpoint was gated. Before its data existed, the two were ordered by a registered rule: full pass, then final-answer-correct over the 160 tasks, then calc, then E1.
- Result: 151 against 149, 12 against 8, 199 against 198, so this checkpoint ranks first. The margin is 2 tasks of 160, within noise. The registered order decided, not significance.
- Fallback: a conditional fallback for attempt 12 (checkpoint-01420) was registered in case the gated checkpoint failed E1, E2 or E5. It was not used.
A repaired defect in the lineage. Attempt 10's anti-repetition target also banned a calculator value that the model was legitimately restating, which damaged numeric answers. Attempt 11 skipped numeric loop onsets and cut damaged answers at the first wrong number. Attempt 12 added cuts at the padding decision and at wrong bolded numbers, plus 80 rollouts on calc answer turns. This checkpoint descends from attempt 10 through both repairs.
Amendments. The preregistration was amended during the phase, each amendment dated and made before the data it governs. The full text is
ENVELOPE-GATE-PREREG-2026-09-22.mdin the training record. The main ones, all 2026:- 09-23 09:45: owner decisions for the deadline;
- 09-23 13:45: every E3 arm moves to the GPU slice;
- 09-23 16:45: the gate picks its checkpoint by E4;
- 09-24 17:00: on-policy distillation registered as the next round;
- 09-25 05:43: release only on a full pass;
- 09-26 00:52: the CPU proxy, validated before use;
- 09-26 09:28: deadline moved to 09-28 09:00;
- 09-27 09:00: hold attempt 11, run attempt 12, ship the better full pass;
- 09-27 10:22: E3 read in whole tasks;
- 09-27 10:23: the conditional fallback.
E3 by kind (this checkpoint, GPU BF16; the parent solves every task):
kind solved kind solved calc 3 of 5 forecast 4 of 5 calc_mental 3 of 3 missing_tool 5 of 5 contact 9 of 9 no_tool 5 of 5 contact_weather 5 of 5 stock 3 of 3 currency 5 of 5 units 2 of 2 dates 2 of 2 weather1 3 of 3 flight 4 of 4 weather2 4 of 4 Gate build against the pinned build. Every gate arm ran on runner main 0bfa2ad, with 38 of 52 blocks on a 24 GB GPU slice (
--gpu-layers 38). This card pins Runner v0.5.7, which contains 0bfa2ad. No gate arm was re-run on v0.5.7. The loop-guard row above was measured on the v0.5.7 release binary.
Limits, read before quoting
- Not parent-equivalent, and not claimed to be. KLD 0.762 and margin-qualified top-1 84.0% against the parent. The house bar for quantised copies does not apply to a pruned and distilled student. "No quality loss" is not claimed and is not supported by this evidence.
- Parent-agreement is not a capability benchmark. The retention rows measure how closely next-token distributions track the parent on held-out text.
- Tool calling, scoped. The claim rests on 99 calls from 200 prompts and 60 closed-loop tasks in one synthetic environment, where the parent solves 60. It says nothing about other tool sets, schemas or orchestration patterns, and none of it comes from the general distillation (corpus v1 and v2 held no tool-calling documents).
- The agentic tasks are narrow. E3 is one synthetic environment with 14 task kinds (weather, forecasts, units, contacts, flights, currency, stocks, dates, arithmetic, and tasks that need no tool or a tool that is missing) and at most 6 calls; the E1, E2 and E5 prompts come from the same environment's trajectories. Nothing here measures long-horizon agents, real APIs, code execution, or the parent's published agentic benchmarks.
- Tool calls were measured on the raw completions path. The gate rendered prompts
with the study's
muserenderer, verified byte-identical to Runner's template at runner main 676d3a7, and parsed the ATEM calls itself from/v1/completions. Runner's chat endpoint with atoolslist was not part of the gate. A chat request with atoolslist on Runner v0.5.7 returned a parsed call on 2026-09-28; it is not a measured rate. Served through Runner's constrained tool calling, a call cannot be emitted with a schema-required argument missing: the grammar withholds the closing tag until they are written (checked on v0.5.7, bd20e1f, on 2026-09-28; the runner session verified the mechanism on a fixture 2026-09-25). A generation cut by the token budget is completed with empty required values and reports finish_reasonlength, nottool_calls. E2 above measures free generation, so its number is the model's, not this path's. - The reference is the parent at Q4_K. E1, E2, E3 and E5 compare against Muse-Glimmer-30B quantised to Q4_K by Runner's own quantiser (the study scorer puts that file at KLD 0.01930 and margin-qualified top-1 99.12% from its BF16), not against the BF16 parent. E4 uses the BF16 parent.
- The scored checkpoint and the served file. E4 was measured on the safetensors checkpoint through the study's PyTorch forward; E1, E2, E3 and E5 on the BF16 GGUF converted from it by llama.cpp b10353's converter. No row compares the two directly.
- Runner builds. E1, E2, E3 and E5 ran on runner main 0bfa2ad; the serving
measurements ran on 53b4deb and ac5418e. All three are contained in v0.5.7 (bd20e1f), and
the changes after ac5418e (speculative-walk stop, buffered Responses status, envelope cut
parity) do not touch the raw-completions path the arms used. E4 does not use Runner.
Reproducing E1's terminator check needs a Runner that reports
stop_token, which arrived on runner main at 0bfa2ad, after the v0.5.6 tag, and ships in v0.5.7. - Why the files are in block order. Runner v0.5.6 and earlier upload a partial offload
as one contiguous file prefix, counted from byte 0. The converter's own order puts
output.weightearly and blocks 8, 9 and 26 out of place, so a converter-order BF16 on a card that cannot hold the whole file can run out of memory and fall back to the CPU. The published files are therefore rewritten into block order (token_embd, blocks 0 to 51,output_norm,output). Only tensor order and offsets change, and every tensor's bytes and every metadata field are verified identical. On this layout a 44-block offload uploads 22.8 GB instead of 25.2 GB. Runner main fixed the plan itself at 9b825fc, which ships in v0.5.7. Q8_0 and Q5_0-mix fit a 24 GB card whole and are not affected. Checked on v0.5.6 on 2026-09-28: the published BF16 file serves with partial offload on a 24 GB GPU slice and does not fall back to the CPU, 38 of 52 blocks (20.1 GB on the GPU) with--gpu-layers 38and 45 of 52 (23.3 GB) with the automatic fit, the same placement as v0.5.7. The fallback applies to the converter's order, which is not published. - Reasoning strength is fixed at "high". Every agentic training example set
"Reasoning strength: high." On attempt 5's checkpoint-00525 (an earlier checkpoint of the released run), 20 held-out answer turns gave byte-identical greedy output at "low" and "high" on 19. Runner's
reasoning_strengthrequest parameter does not shorten this model's reasoning. A serving-side reasoning budget is the lever. - Fixed-decimal padding. The model does not reliably pad a value to a fixed number of decimals when a tool returns fewer digits (asked for two decimals of 7970.5, it may write 7970.5, or deliberate about how to write 7970.50). A caller who needs a fixed format should format the number downstream rather than ask the model for it.
- Text only. The parent's vision encoder is not part of this model; image input is not supported.
- The contamination filter was weakened by design for corpus v3. It exempts, per
domain, the 64-token windows that recur in at least 2% of 2,000 sampled documents of
that domain: 166 windows for tool calling, 16 for casual dialogue, 3 for general
knowledge, none elsewhere. A window carried by thousands of training documents is that
domain's format and cannot fingerprint one eval item; that is the whole claim. Under the
muserender the filter is also weaker at role boundaries, because the forbidden windows come from the plain-rendered eval split. - Selection data. The held-out split was read by the scorer six times (K1, K2, two progress reads, K3, E4) and never trained on. The E1, E2 and E5 prompts come from 200 task ids reserved before assembly, none of them in the envelope corpus. The trees covered tasks 0 to 2007, so they overlapped E3's 2000 to 2059 at 2000 to 2007; the 19 training examples from those tasks were removed before any envelope training (2026-09-23), and the on-policy rollouts skip 2000 to 2159. A scan of every policy file this checkpoint's lineage trained on (examples-muse-all.jsonl for attempts 2 to 5, then examples-muse-a6-mixed.jsonl to examples-muse-a12-mixed.jsonl) finds 1,779 distinct task ids, all below 2000: none from E3's 60 tasks, the 100 discovery tasks or E1's 200 held-out prompts.
- Early stopping. The preregistered rule stops a run if E4's dev proxy rises at three consecutive evaluations. It was implemented in the trainer before step 200 and fired at step 200 in attempts 1 and 2. From attempt 3 on it was turned off: E4 was measured on every saved checkpoint instead, and the gate read the latest checkpoint inside E4's cap. On the released segment the dev proxy read 0.4967, 0.4962 and 0.4949 at steps 1,320, 1,350 and 1,400.
- A short phase. 1,440 envelope steps, 14758380 general positions and 772568 loss-carrying agentic positions, on top of 98.3 M distillation tokens.
- One model, one environment, one shared 24 GB GPU slice.
Reproduce
# 0. Parent: meta-models/Muse-Glimmer-30B @ a4e59da52a7bc87ae7251dd5545c0dd437c44b68, BF16 safetensors.
# The surgery study's shared modules (mgcommon.py, run6_x1_score.py, build_corpus.py,
# run9_build_corpus.py) are in code/deps/; the scripts expect them on the path they name.
# 1. Importance capture on corpus_run9's train split, then the width cut.
python prune_capture.py --sequences 64 --tokens 2048 --output prune-capture-64x2048.npz
python prune_width.py --capture prune-capture-64x2048.npz --hidden 5760 --ffn 10240 --heads 24 \
--output muse-w14b-pruned
# 2. Corpora. v1: plain render, built before the boilerplate exemption existed.
# v3: the muse render, exemption at 2% of 2,000 sampled documents per domain.
python build_student_corpus.py --output corpus-v1 --boilerplate-frac 0
python build_student_corpus.py --output corpus-v3-muse --render muse \
--cap-tokens 3000000 --tulu-tokens 6000000 --wikitext-tokens 5000000
# 3. General distillation, 6,000 steps.
python train_student.py --student muse-w14b-pruned --corpus corpus-v1 --output train-w14b-v1 \
--steps 6000 --windows 4 --tokens 4096 --lr 1e-4 --warmup 50 --save-every 500 --seed 20260915
# 4. K gates on the held-out split.
python score_student.py --student train-w14b-v1/checkpoint-06000 --name w14b-step6000 --out K3.json
# 5. Envelope corpus: Ornith state trees, then assembly with the parent writing the words.
# Worker scripts with every offset: evidence/assembly/.
python agentic_trees.py --tokenizer <Ornith-1.0-9B tokenizer> --gpu off --tasks 50 --offset <o> \
--out trees-cpu.jsonl
python agentic_assemble.py --trees trees-*.jsonl --external-server --port 58701 --offset <k> --stride 4 \
--holdout holdout-tasks.json --out examples-muse-w<k>.jsonl
# 6. Envelope phase, one command per segment, each resuming the checkpoint the one before handed on
# (the trainer at each segment's recorded SHA-256; see DISTILLATION.json).
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-v1/checkpoint-06000 \
--corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a2 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--stop-at 200 --early-stop on
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a2/checkpoint-00200 \
--corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a3 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 200 --stop-at 350 --early-stop off
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a3/checkpoint-00350 \
--corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a4 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 350 --stop-at 450 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}'
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a4/checkpoint-00450 \
--corpus corpus-v3-muse --policy examples-muse-all.jsonl --output train-w14b-envelope-a5 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 450 --stop-at 600 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a6-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a5/checkpoint-00600's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
# attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
# prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a5/checkpoint-00600 \
--corpus corpus-v3-muse --policy examples-muse-a6-mixed.jsonl --output train-w14b-envelope-a6 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 600 --stop-at 720 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a7-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a6/checkpoint-00720's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
# attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
# prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a6/checkpoint-00720 \
--corpus corpus-v3-muse --policy examples-muse-a7-mixed.jsonl --output train-w14b-envelope-a7 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 720 --stop-at 840 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a8-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a7/checkpoint-00840's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
# attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
# prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a7/checkpoint-00840 \
--corpus corpus-v3-muse --policy examples-muse-a8-mixed.jsonl --output train-w14b-envelope-a8 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 840 --stop-at 960 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a9-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a8/checkpoint-00960's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
# attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
# prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a8/checkpoint-00960 \
--corpus corpus-v3-muse --policy examples-muse-a9-mixed.jsonl --output train-w14b-envelope-a9 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 960 --stop-at 1080 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a10-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a9/checkpoint-01080's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
# attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
# prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a9/checkpoint-01080 \
--corpus corpus-v3-muse --policy examples-muse-a10-mixed.jsonl --output train-w14b-envelope-a10 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 1080 --stop-at 1200 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a11-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a10/checkpoint-01200's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
# attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
# prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a10/checkpoint-01200 \
--corpus corpus-v3-muse --policy examples-muse-a11-mixed.jsonl --output train-w14b-envelope-a11 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 1200 --stop-at 1320 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# examples-muse-a12-mixed.jsonl: examples-muse-all.jsonl plus the student's own rollouts from train-w14b-envelope-a11/checkpoint-01320's GGUF
# (onpolicy_rollouts.py --states post-tool and --states first-final, build_onpolicy_examples.py, a 1:1 weight mix;
# attempt 6: code/a6-onpolicy-if-needed.sh, with the first-turn shards started by hand at 2026-09-24 23:18 as the
# prereg records; attempt 7 and later: code/onpolicy-round.sh, which runs both)
python train_student_policy.py --student muse-w14b-pruned --resume train-w14b-envelope-a11/checkpoint-01320 \
--corpus corpus-v3-muse --policy examples-muse-a12-mixed.jsonl --output train-w14b-envelope-a12 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 1320 --stop-at 1440 --early-stop off --general-per-policy 2 --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# 7. The gate, in the order evidence/gate/treated/gate-chain.sh runs it.
python score_student.py --student train-w14b-envelope-a12/checkpoint-01440 --name E4 --batch 2 --out E4.json
python convert_hf_to_gguf.py train-w14b-envelope-a12/checkpoint-01440 \
--outfile Xyntetik-Kvist-14B-BF16.gguf --outtype bf16 # llama.cpp b10353
python agentic_closedloop.py --model Xyntetik-Kvist-14B-BF16.gguf --label treated --runner ./runner \
--tokenizer <parent snapshot> --gpu auto --gpu-layers 38 --tasks 60 --offset 2000 --raw-muse \
--out treated-e3.json # runner main 0bfa2ad, 24 GB slice
runner -m Xyntetik-Kvist-14B-BF16.gguf --serve --no-tray --port 58740 --gpu auto --gpu-layers 38 \
-t 48 -c 8192 --parallel 4 & # runner main 0bfa2ad
python envelope_preflight.py --endpoint http://127.0.0.1:58740 --model Xyntetik-Kvist-14B-BF16.gguf \
--trees trees-*.jsonl --holdout holdout-tasks.json --n 200 --max-tokens 512 --concurrency 4 \
--arm treated --out treated-e1e2e5.json
# 8. Block-order the BF16, quantise with llama.cpp b10353, restore the order, and measure
# each file against this model's BF16 (RUN-ON-ENVELOPE.sh in code/ runs all of it with its checks).
python gguf_blockorder.py converter-bf16.gguf Xyntetik-Kvist-14B-BF16.gguf --json layout-rewrite.json
llama-quantize Xyntetik-Kvist-14B-BF16.gguf q8_0.tmp.gguf Q8_0 16
llama-quantize Xyntetik-Kvist-14B-BF16.gguf q5mix.tmp.gguf Q4_K_M 16 # falls back to a Q5_0 majority
python gguf_blockorder.py q8_0.tmp.gguf Xyntetik-Kvist-14B-Q8_0.gguf --json q8_0-order.json \
--order-from Xyntetik-Kvist-14B-BF16.gguf
python gguf_blockorder.py q5mix.tmp.gguf Xyntetik-Kvist-14B-Q5_0-mix.gguf --json q5mix-order.json \
--order-from Xyntetik-Kvist-14B-BF16.gguf
python scripts/kld-compare-raw.py --model-a Xyntetik-Kvist-14B-Q8_0.gguf \
--model-b Xyntetik-Kvist-14B-BF16.gguf --runner ./runner --corpus evidence/muse/harness-eval-split.txt \
--max-positions 500 --out q8_0-vs-bf16.json
Training on a GPU is not bit-reproducible, so a rerun lands near these numbers and not on
them; the SHA-256 values below identify exactly what was measured and published. The
BF16 GGUF is a deterministic conversion of the safetensors checkpoint with the same
converter. Code in code/, the preregistrations with their dates in
evidence/preregistration/, every score, log and launch script in evidence/.
Provenance
- Parent:
meta-models/Muse-Glimmer-30B, revisiona4e59da52a7bc87ae7251dd5545c0dd437c44b68, Apache-2.0. BF16 safetensorsmodel-00001-of-00002.safetensorsSHA-2568eef61530e1283642c77ce2e6721feb5c6f348fa055c00e90f2844a136372694,model-00002-of-00002.safetensorsSHA-256b58cc2144ba1ba1af4420f67f4ca3ced7f09298510b80464cc75018a0be14381. Tokenizer unchanged:tokenizer.jsonSHA-256c9dbee66967b58f3…, identical to the parent's. - Gate reference: the parent at Q4_K,
parent_q4_k.gguf, 15,686,266,304 bytes, SHA-2560956fcae4fda1e8f93f8784ace121accde2b8db4b00be7d3aa8dc926bf726075. - Policy teacher: Ornith-1.0-9B (
deepreinforce-ai/Ornith-1.0-9B), served asornith-1.0-9b-Q4_K_M.gguf, SHA-2565720d1f671b4996481274fffe01868c3c36e87c135cc8538471cc7bd6087b106. Its policy mass shaped the targets at decision positions; no Ornith weights are in this model. - Pruned start:
PRUNE.jsonSHA-2560356aae40e33a54d…; calibration captureprune-capture-64x2048.npzSHA-2566e4f0e30eebe1fd2…. - General-distillation checkpoint (K3): checkpoint-06000; its BF16 GGUF,
28,903,141,184 bytes, SHA-256
9c7f7a289867c4ebf89d17aeb95be3ca52e2c9ab9f24f44bccc39392ff619180. - This model: checkpoint-01440 of the envelope run. Safetensors
model-00000.safetensors6e7602deb8371e0b1b9e2608c2e4414fe4a4bbd77f19bd835294a5dedf9c3e1a,model-00001.safetensorsca69b78b6f833cd887b428f5394894f79e4102afc332f91bed683c03972f0ce2,model-00002.safetensors2e07b5fb9609dddd7cc4ee3354f10d05695274c06575c63eebbfd8b1c955e8d9,model-00003.safetensors6f108a8e55e00dd9cda4400829217173c3696cc337c8d4cea252d3ce9b1ca30e.Xyntetik-Kvist-14B-BF16.gguf28,903,141,184 bytes, SHA-256daf422db9d0853d53ec8abb8fe9cd3f1e33ea53dffe7cca476225fbe1bc5d5c6.Xyntetik-Kvist-14B-Q8_0.gguf15,363,224,384 bytes, SHA-256cdb73543a6fc22f7faa5bbab0176ab2d5b1b590b46870d2d22f2a2510caaacb6.Xyntetik-Kvist-14B-Q5_0-mix.gguf10,295,023,424 bytes, SHA-25637af54b4cef4445bbbf23e8fbf0ceec8dc9b01b38dab0dc7f9e2bd2fd79f3a56.Xyntetik-Kvist-14B-IQ4_NL-mix.gguf7,612,384,384 bytes, SHA-256c74b27a5a39c894c4d1a1a2ae140de0cee125fc5cc9a7c96db9b7355bf8709a3. - Training data: corpus v1 tokens SHA-256
0731731533ca5ab7…,corpus-v3-musetokens34a37a0ba8086c11…, envelope examplescb1dd6fa5e6c516d…. Sources:HuggingFaceTB/smoltalk@5feaf2fd3ffca7c2…,allenai/tulu-3-sft-mixture@b14afda60f1bbebe…,Salesforce/wikitextwikitext-103-raw-v1 train. Per-domain counts and hashes inevidence/corpus/; the token files themselves are not distributed. - Runner: v0.5.7 is runner main bd20e1f (v0.5.6 was e3781b9). Gate builds as listed under Limits.
- Code:
code/at study commit25cbda1; the record is in the shade repository,research/surgery/kvist-students/, at4eeaa802.
Xyntetik-Kvist is this house's name for a new work. Muse Glimmer is its publisher's name, used here only to say where the weights came from.
- Downloads last month
- 131