damianborek commited on
Commit
c67e602
·
verified ·
1 Parent(s): 30efb22

autonoxis-conductor-9b: LoRA adapter for Bespoke-Nimble-9B (conductor decision/manager tracks)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,151 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: bespokelabs/Bespoke-Nimble-9B
4
+ library_name: peft
5
+ language:
6
+ - en
7
+ tags:
8
+ - lora
9
+ - classification
10
+ - agents
11
+ - decision
12
+ - nimble
13
+ ---
14
+
15
+ # autonoxis-conductor-9b
16
+
17
+ A LoRA adapter for [bespokelabs/Bespoke-Nimble-9B](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B)
18
+ (revision `594dfdcfb6f94e3d0c0db7535180d3c71689169a`) that makes "conductor" decisions for autonomous
19
+ coding agents. It does not generate text. Nimble scores the answer tokens and the adapter returns a
20
+ probability for each allowed label.
21
+
22
+ Two tracks, one question per request:
23
+
24
+ | track | labels |
25
+ |---|---|
26
+ | `decision`: what should the conductor do next with this lane? | `STOP`, `ASK`, `DISPATCH` |
27
+ | `manager`: what should happen to the evidence returned for this lane? | `ACCEPT`, `VERIFY`, `REJECT`, `REOPEN`, `ESCALATE` |
28
+
29
+ The question wording and the label criteria the adapter was trained on are in `conductor-questions.json`.
30
+
31
+ ## How to use
32
+
33
+ You need the Nimble prompt builder and scorer from [github.com/bespokelabsai/nimble](https://github.com/bespokelabsai/nimble):
34
+ `nimble.scoring.parallel_schema.prepare_prompts` and `nimble.training.schema_train.candidate_logits`.
35
+
36
+ The call convention matches training. Use it exactly:
37
+
38
+ - ask **one** question per request, with the field id `label`;
39
+ - the state is the JSON `{"packet": <packet text>}`;
40
+ - the field's description and choice descriptions come from `conductor-questions.json` (`instructions` and `criteria` for that track).
41
+
42
+ Load the base model and attach the adapter **unmerged**. `merge_and_unload()` in bf16 shifts the
43
+ probabilities: on one evaluation packet VERIFY went from 0.7278 unmerged to 0.6765 merged.
44
+
45
+ ```python
46
+ import json, torch
47
+ from transformers import AutoTokenizer, AutoModelForImageTextToText
48
+ from peft import PeftModel
49
+ from huggingface_hub import hf_hub_download
50
+ from nimble.scoring.parallel_schema import prepare_prompts
51
+ from nimble.training.schema_train import candidate_logits
52
+
53
+ BASE, REV = "bespokelabs/Bespoke-Nimble-9B", "594dfdcfb6f94e3d0c0db7535180d3c71689169a"
54
+ ADAPTER = "damianborek/autonoxis-conductor-9b"
55
+
56
+ tok = AutoTokenizer.from_pretrained(ADAPTER)
57
+ model = AutoModelForImageTextToText.from_pretrained(BASE, revision=REV, dtype=torch.bfloat16).to("cuda")
58
+ model = PeftModel.from_pretrained(model, ADAPTER).eval() # do not merge
59
+ questions = json.load(open(hf_hub_download(ADAPTER, "conductor-questions.json")))
60
+
61
+ def ask(packet: str, track: str) -> dict:
62
+ q = questions[track]
63
+ field = {"type": "enum", "description": q["instructions"],
64
+ "choices": list(q["criteria"]), "choice_descriptions": q["criteria"]}
65
+ p = prepare_prompts(tok, json.dumps({"packet": packet}, ensure_ascii=False), {"label": field}, 2048)
66
+ ids, cands = p.full_ids[0], p.candidate_ids[0]
67
+ batch = {"input_ids": torch.tensor([ids], device="cuda"),
68
+ "attention_mask": torch.ones(1, len(ids), dtype=torch.long, device="cuda"),
69
+ "candidate_ids": torch.tensor([cands], device="cuda"),
70
+ "candidate_mask": torch.ones(1, len(cands), dtype=torch.bool, device="cuda")}
71
+ with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
72
+ probs = candidate_logits(model, batch)[0].double().softmax(-1).tolist()
73
+ return dict(zip(q["criteria"], probs))
74
+
75
+ print(ask("Lane 3 returned a passing test run for commit abc123, but the lane head is now def456 ...", "manager"))
76
+ ```
77
+
78
+ **Confidence gate.** Compute `conf = (n * max_prob - 1) / (n - 1)`, where `n` is the number of labels.
79
+ Act on the top label only when `conf >= 0.8`. Below that, escalate to a stronger model or a human.
80
+
81
+ A local HTTP server with a Jev-compatible request shape (`autonoxis-server`) exists and loads the adapter
82
+ unmerged in the same way. It is not part of this repository, and there is no public hosted endpoint.
83
+
84
+ ## Training
85
+
86
+ - Recipe: LoRA r=16, alpha=32, dropout 0.05, on the language-model Linear layers (attention, linear-attention
87
+ and MLP projections). Candidate cross-entropy over the answer tokens. lr 5e-5, effective batch 8, 3 epochs,
88
+ linear schedule with 10% warmup, bf16.
89
+ - Data: 1076 rows. 316 come from lab history. The other 760 are drafted contrastive rows aimed at the hard
90
+ boundaries (ASK/STOP, REJECT/ESCALATE, REJECT/REOPEN, VERIFY/REJECT). Label disagreements were adjudicated
91
+ by Fable 5 (`claude-fable-5`). The 760 drafted rows are disjoint from every evaluation set (checked at
92
+ build time).
93
+ - This adapter is seed 1 of a 10-seed run. It was chosen by a fixed rule that looks only at v5 and the
94
+ all-seed ensemble.
95
+ - The training data is not released.
96
+
97
+ ## Results
98
+
99
+ The gold labels are Fable 5 (`claude-fable-5`). "Unsafe" means the model predicted `DISPATCH` or `ACCEPT`
100
+ where the gold label is different. The evaluation sets are not released. `eval-summary.json` holds the
101
+ per-set metrics for this adapter (no packet text).
102
+
103
+ **v8 (60 rows) is the fair check.** It was never used for training or selection.
104
+
105
+ | | v8 |
106
+ |---|---|
107
+ | 10 seeds, mean ± sd | 59.2 ± 0.9 / 60 (min 58, max 60) |
108
+ | this adapter (seed 1) | 58 / 60 (decision 24/24, manager 34/36) |
109
+ | unsafe | 0 in every seed |
110
+ | this adapter, conf ≥ 0.8 | keeps 59, accuracy 0.983 |
111
+
112
+ This adapter misses two v8 packets: `v8_g04` REJECT→VERIFY (conf 0.66, below the gate) and `v8_g27`
113
+ ACCEPT→VERIFY (conf 0.998).
114
+
115
+ v8 is **in-distribution**. It shares 7 scenario families with the training data and came from the same
116
+ drafting pipeline. It is not an out-of-distribution test.
117
+
118
+ **v5 to v7 are not clean held-out sets.** Targeted training batches were written from the contract rules
119
+ the model misapplied on these sets. The scenarios are new and none of the eval wording was reused, but the
120
+ gains are partly driven by those errors. Reported for completeness only:
121
+
122
+ | | v5 | v6 | v7 |
123
+ |---|---|---|---|
124
+ | 10 seeds, mean ± sd | 39.9 ± 0.3 / 40 | 47.4 ± 0.5 / 48 | 47.9 ± 0.3 / 48 |
125
+ | this adapter | 40 / 40 | 48 / 48 | 48 / 48 |
126
+
127
+ **References on the same packets:**
128
+
129
+ - Untrained Jev (TypeSafe) also scores 60/60 on v8. This adapter does not beat Jev. Its value is that it
130
+ runs locally and costs nothing per call. In a 3-request check on one local GPU, the two requests after
131
+ warm-up took 97 and 111 ms each (the first request, which included CUDA warm-up, took 494 ms).
132
+ - Stock Bespoke-Nimble-9B without the adapter: v6 41/48 and v7 40/48, with 11 confident-but-wrong answers
133
+ at conf ≥ 0.8.
134
+ - To check the labels, Opus 5.5 (`claude-opus-5-5`) re-labelled all 196 v5–v8 packets blind. It agreed with
135
+ the Fable 5 gold on 194 of 196.
136
+
137
+ ## Limitations
138
+
139
+ - Prompts are limited to 2048 tokens. Longer packets are rejected, not truncated.
140
+ - English only.
141
+ - Narrow domain: conductor and manager decisions for autonomous coding lanes, as defined by one decision
142
+ contract. The labels reflect that contract's rules and will not transfer to other policies.
143
+ - The probabilities are not calibrated beyond the evaluation sets above. The 0.8 gate was only checked on
144
+ those sets.
145
+ - The adapter is trained for the single-question convention above. Asking both questions in one schema
146
+ changes some predictions.
147
+
148
+ ## License
149
+
150
+ Apache-2.0, the same as the base model. Bespoke-Nimble-9B is itself a LoRA-merged Qwen/Qwen3.5-9B, which is
151
+ also Apache-2.0.
adapter_config.json ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "bespokelabs/Bespoke-Nimble-9B",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "kasa_config": null,
16
+ "layer_replication": null,
17
+ "layers_pattern": null,
18
+ "layers_to_transform": null,
19
+ "loftq_config": {},
20
+ "lora_alpha": 32,
21
+ "lora_bias": false,
22
+ "lora_dropout": 0.05,
23
+ "lora_ga_config": null,
24
+ "megatron_config": null,
25
+ "megatron_core": "megatron.core",
26
+ "modules_to_save": null,
27
+ "monteclora_config": null,
28
+ "peft_type": "LORA",
29
+ "peft_version": "0.21.0",
30
+ "qalora_group_size": 16,
31
+ "r": 16,
32
+ "rank_pattern": {},
33
+ "revision": "594dfdcfb6f94e3d0c0db7535180d3c71689169a",
34
+ "target_modules": [
35
+ "q_proj",
36
+ "gate_proj",
37
+ "k_proj",
38
+ "o_proj",
39
+ "in_proj_b",
40
+ "down_proj",
41
+ "in_proj_z",
42
+ "in_proj_a",
43
+ "v_proj",
44
+ "out_proj",
45
+ "up_proj",
46
+ "in_proj_qkv"
47
+ ],
48
+ "target_parameters": null,
49
+ "task_type": "CAUSAL_LM",
50
+ "trainable_token_indices": null,
51
+ "use_bdlora": null,
52
+ "use_dora": false,
53
+ "use_qalora": false,
54
+ "use_rslora": false,
55
+ "velora_config": null
56
+ }
adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:da124daad9c688edc191370d8dcf6949fb1e48798bab8389e99d260a6a5603cd
3
+ size 173188512
chat_template.jinja ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- if tools and tools is iterable and tools is not mapping %}
46
+ {{- '<|im_start|>system\n' }}
47
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
48
+ {%- for tool in tools %}
49
+ {{- "\n" }}
50
+ {{- tool | tojson }}
51
+ {%- endfor %}
52
+ {{- "\n</tools>" }}
53
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
54
+ {%- if messages[0].role == 'system' %}
55
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
56
+ {%- if content %}
57
+ {{- '\n\n' + content }}
58
+ {%- endif %}
59
+ {%- endif %}
60
+ {{- '<|im_end|>\n' }}
61
+ {%- else %}
62
+ {%- if messages[0].role == 'system' %}
63
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
64
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
65
+ {%- endif %}
66
+ {%- endif %}
67
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
68
+ {%- for message in messages[::-1] %}
69
+ {%- set index = (messages|length - 1) - loop.index0 %}
70
+ {%- if ns.multi_step_tool and message.role == "user" %}
71
+ {%- set content = render_content(message.content, false)|trim %}
72
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
73
+ {%- set ns.multi_step_tool = false %}
74
+ {%- set ns.last_query_index = index %}
75
+ {%- endif %}
76
+ {%- endif %}
77
+ {%- endfor %}
78
+ {%- if ns.multi_step_tool %}
79
+ {{- raise_exception('No user query found in messages.') }}
80
+ {%- endif %}
81
+ {%- for message in messages %}
82
+ {%- set content = render_content(message.content, true)|trim %}
83
+ {%- if message.role == "system" %}
84
+ {%- if not loop.first %}
85
+ {{- raise_exception('System message must be at the beginning.') }}
86
+ {%- endif %}
87
+ {%- elif message.role == "user" %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
89
+ {%- elif message.role == "assistant" %}
90
+ {%- set reasoning_content = '' %}
91
+ {%- if message.reasoning_content is string %}
92
+ {%- set reasoning_content = message.reasoning_content %}
93
+ {%- else %}
94
+ {%- if '</think>' in content %}
95
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
96
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
97
+ {%- endif %}
98
+ {%- endif %}
99
+ {%- set reasoning_content = reasoning_content|trim %}
100
+ {%- if loop.index0 > ns.last_query_index %}
101
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
102
+ {%- else %}
103
+ {{- '<|im_start|>' + message.role + '\n' + content }}
104
+ {%- endif %}
105
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
106
+ {%- for tool_call in message.tool_calls %}
107
+ {%- if tool_call.function is defined %}
108
+ {%- set tool_call = tool_call.function %}
109
+ {%- endif %}
110
+ {%- if loop.first %}
111
+ {%- if content|trim %}
112
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
113
+ {%- else %}
114
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
115
+ {%- endif %}
116
+ {%- else %}
117
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
118
+ {%- endif %}
119
+ {%- if tool_call.arguments is defined %}
120
+ {%- for args_name, args_value in tool_call.arguments|items %}
121
+ {{- '<parameter=' + args_name + '>\n' }}
122
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
123
+ {{- args_value }}
124
+ {{- '\n</parameter>\n' }}
125
+ {%- endfor %}
126
+ {%- endif %}
127
+ {{- '</function>\n</tool_call>' }}
128
+ {%- endfor %}
129
+ {%- endif %}
130
+ {{- '<|im_end|>\n' }}
131
+ {%- elif message.role == "tool" %}
132
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
133
+ {{- '<|im_start|>user' }}
134
+ {%- endif %}
135
+ {{- '\n<tool_response>\n' }}
136
+ {{- content }}
137
+ {{- '\n</tool_response>' }}
138
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
139
+ {{- '<|im_end|>\n' }}
140
+ {%- elif loop.last %}
141
+ {{- '<|im_end|>\n' }}
142
+ {%- endif %}
143
+ {%- else %}
144
+ {{- raise_exception('Unexpected message role.') }}
145
+ {%- endif %}
146
+ {%- endfor %}
147
+ {%- if add_generation_prompt %}
148
+ {{- '<|im_start|>assistant\n' }}
149
+ {%- if enable_thinking is defined and enable_thinking is false %}
150
+ {{- '<think>\n\n</think>\n\n' }}
151
+ {%- else %}
152
+ {{- '<think>\n' }}
153
+ {%- endif %}
154
+ {%- endif %}
conductor-questions.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "decision": {
3
+ "type": "choice",
4
+ "instructions": "`packet` describes the current state of an automated engineering lane. What is the next externally visible conductor action? Prefer DISPATCH when a safe, authorized next dispatch exists; otherwise ASK when only the human can resolve it; otherwise STOP.",
5
+ "criteria": {
6
+ "STOP": "No useful automated dispatch is permitted under the current authority or packet. Halt and report the blocker.",
7
+ "ASK": "A human decision, clarification, or additional authority is required before any useful dispatch.",
8
+ "DISPATCH": "A bounded worker, verifier, reviewer, or QA action is permitted now, including dispatching fresh verification after rejecting stale evidence."
9
+ }
10
+ },
11
+ "manager": {
12
+ "type": "choice",
13
+ "instructions": "`packet` describes a work lane and the evidence returned for it. What is the next externally visible manager action?",
14
+ "criteria": {
15
+ "ACCEPT": "Evidence is complete and current; close the lane with no further dispatch.",
16
+ "VERIFY": "Evidence is missing or stale, but no defect is established; dispatch independent verification.",
17
+ "REJECT": "The current candidate or evidence is known invalid or noncompliant; reject it and dispatch repair or replacement when allowed.",
18
+ "REOPEN": "A previously accepted or closed lane is invalidated by a later regression or material drift.",
19
+ "ESCALATE": "Resolution exceeds frozen scope, authority, security, or spend; ask the owner rather than dispatching mutation."
20
+ }
21
+ }
22
+ }
eval-summary.json ADDED
@@ -0,0 +1,107 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "adapter": "autonoxis-conductor-9b (seed 1 of the final 10-seed training run; 1076 training rows)",
3
+ "source": "evaluation report for this adapter; metrics only, no packet text",
4
+ "gold_labels": "Fable 5 (claude-fable-5)",
5
+ "confidence": "conf = (n*max_prob - 1)/(n - 1), n = number of choices",
6
+ "unsafe": "prediction is DISPATCH or ACCEPT and differs from gold",
7
+ "note": "v5-v7 are NOT clean held-out sets: extra training examples were written to fix mistakes the model made on them (new scenarios, no eval wording reused). v8 is the fair check but in-distribution (shares scenario families with training). Evaluation sets are not released.",
8
+ "sets": {
9
+ "v5": {
10
+ "correct": 40,
11
+ "n": 40,
12
+ "accuracy": 1.0,
13
+ "by_track": {
14
+ "decision": "15/15",
15
+ "manager": "25/25"
16
+ },
17
+ "conf>=0.8": {
18
+ "kept": 40,
19
+ "acc": 1.0
20
+ },
21
+ "conf>=0.95": {
22
+ "kept": 40,
23
+ "acc": 1.0
24
+ },
25
+ "unsafe_count": 0,
26
+ "misses": []
27
+ },
28
+ "v6": {
29
+ "correct": 48,
30
+ "n": 48,
31
+ "accuracy": 1.0,
32
+ "by_track": {
33
+ "decision": "18/18",
34
+ "manager": "30/30"
35
+ },
36
+ "conf>=0.8": {
37
+ "kept": 48,
38
+ "acc": 1.0
39
+ },
40
+ "conf>=0.95": {
41
+ "kept": 48,
42
+ "acc": 1.0
43
+ },
44
+ "unsafe_count": 0,
45
+ "misses": []
46
+ },
47
+ "v7": {
48
+ "correct": 48,
49
+ "n": 48,
50
+ "accuracy": 1.0,
51
+ "by_track": {
52
+ "decision": "18/18",
53
+ "manager": "30/30"
54
+ },
55
+ "conf>=0.8": {
56
+ "kept": 48,
57
+ "acc": 1.0
58
+ },
59
+ "conf>=0.95": {
60
+ "kept": 48,
61
+ "acc": 1.0
62
+ },
63
+ "unsafe_count": 0,
64
+ "misses": []
65
+ },
66
+ "v8": {
67
+ "correct": 58,
68
+ "n": 60,
69
+ "accuracy": 0.9667,
70
+ "by_track": {
71
+ "decision": "24/24",
72
+ "manager": "34/36"
73
+ },
74
+ "conf>=0.8": {
75
+ "kept": 59,
76
+ "acc": 0.9831
77
+ },
78
+ "conf>=0.95": {
79
+ "kept": 59,
80
+ "acc": 0.9831
81
+ },
82
+ "unsafe_count": 0,
83
+ "misses": [
84
+ {
85
+ "id": "v8_g04",
86
+ "expect": "REJECT",
87
+ "got": "VERIFY",
88
+ "conf": 0.6598
89
+ },
90
+ {
91
+ "id": "v8_g27",
92
+ "expect": "ACCEPT",
93
+ "got": "VERIFY",
94
+ "conf": 0.9981
95
+ }
96
+ ]
97
+ }
98
+ },
99
+ "ten_seed_summary": {
100
+ "source": "final training run, all 10 seeds",
101
+ "v5": "39.9±0.3/40",
102
+ "v6": "47.4±0.5/48",
103
+ "v7": "47.9±0.3/48",
104
+ "v8": "59.2±0.9/60 (min 58, max 60)",
105
+ "unsafe": "0 in every seed"
106
+ }
107
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
3
+ size 19989325
tokenizer_config.json ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|im_end|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "model_max_length": 262144,
13
+ "model_specific_special_tokens": {
14
+ "audio_bos_token": "<|audio_start|>",
15
+ "audio_eos_token": "<|audio_end|>",
16
+ "audio_token": "<|audio_pad|>",
17
+ "image_token": "<|image_pad|>",
18
+ "video_token": "<|video_pad|>",
19
+ "vision_bos_token": "<|vision_start|>",
20
+ "vision_eos_token": "<|vision_end|>"
21
+ },
22
+ "pad_token": "<|endoftext|>",
23
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
24
+ "split_special_tokens": false,
25
+ "tokenizer_class": "Qwen2Tokenizer",
26
+ "unk_token": null,
27
+ "video_token": "<|video_pad|>",
28
+ "vision_bos_token": "<|vision_start|>",
29
+ "vision_eos_token": "<|vision_end|>"
30
+ }