File size: 5,199 Bytes
d053378
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
---
license: apache-2.0
library_name: transformers
base_model: answerdotai/ModernBERT-base
tags:
  - modernbert
  - episodic-ingestion-compiler
  - grouped-softmax-ranker
  - field-event-ranker
  - mixed-mode-training
  - semantic-reasoning-labels
  - v2-labels
language:
  - en
---

# episodic-ingestion-compiler / ModernBERT field-event ranker (mixed-mode v2, perf-H4)

Fine-tune of `answerdotai/ModernBERT-base` for the **grouped multi-positive
softmax ranker** task with **v2 semantic-reasoning labels**.

## What changed vs the v1 mixed-mode checkpoint

An audit of the v1 mixed-mode training data showed that 94% of eval pairs came
from two fields (`attempted_actions`, `observed_outcomes`) whose labels were
**exact functions of event role** on 97% of rows (attempted_actions=tool_call,
observed_outcomes=tool_result). On these degenerate labels:

- E[MRR | random scores] ≈ 0.26 (random-permutation floor for multi-positive
  softmax)
- v1 perf-H4 MRR = 0.506 ≈ 1.97× random
- Architecture experiments (H3, H3a, H3b, H3b-span, etc.) could not break
  through because the ceiling was arithmetic, not informational.

v2 adds five semantic-reasoning fields whose labels require more than role:

| field | kind | rule |
|---|---|---|
| `initiating_command` | single-pick | First tool_call in a multi-call chain |
| `outcome_of_latest_attempt` | single-pick | Tool_result whose call_id matches the LAST tool_call |
| `recent_error` | single-pick | Most-recent tool_result with success=false |
| `attempt_outcome_pairs` | array | Paired tool_results (only when window has ≥1 orphan result — otherwise empty, so it can't collapse to role==tool_result) |
| `action_causality` | array | Tool_results whose content tokens recur in a LATER assistant turn |

Tautology rates on these new fields: 0–10%. Legacy fields kept for back-compat.

## Training details

- **Train rows**: 1217 (mixed-mode v2: autonomous + prompt→action + customer-support + conversational)
- **Train groups**: 5882 (one per field × row pair)
- **Eval rows**: 220 (1900 groups, 21986 candidate pairs)
- **Steps**: 160 optimizer steps (`--accum-steps 8`)
- **Optimizer**: AdamW, lr 5.7e-5, warmup 30
- **Precision**: bf16 autocast
- **Memory**: gradient checkpointing (use_reentrant=False)
- **Final train loss**: 0.0854
- **Peak VRAM**: 4.08 GiB
- **Hardware**: RTX 5090 Laptop (Blackwell SM 12.0, 24 GiB)

## Eval metrics (220-row mixed v2 eval, {'1': 0.2964509394572025, '2': 0.5173277661795407, '3': 0.6931106471816284, '5': 0.8881002087682672})

| metric | value |
|---|---:|
| overall MRR | 0.526 |
| top-1 recall | 0.296 |
| top-2 recall | 0.517 |
| top-3 recall | 0.693 |
| top-5 recall | 0.888 |
| mean expected rank | 3.09 |

Per-field MRR (sorted by eval-pair count):

| field | n | MRR | top-1 |
|---|---:|---:|---:|
| `attempted_actions` | 846 | 0.469 | 0.23 |
| `observed_outcomes` | 845 | 0.486 | 0.24 |
| `initiating_command` | 199 | 0.667 | 0.50 |
| `outcome_of_latest_attempt` | 181 | 0.763 | 0.60 |
| `action_causality` | 141 | 0.483 | 0.24 |
| `attempt_outcome_pairs` | 53 | 0.563 | 0.30 |
| `failed_attempts` | 36 | 0.615 | 0.42 |
| `next_actions` | 32 | 0.565 | 0.34 |
| `recent_error` | 29 | 0.743 | 0.59 |
| `customer_identity` | 11 | 0.818 | 0.73 |
| `transaction_reference` | 5 | 0.467 | 0.20 |
| `product_name` | 4 | 1.000 | 1.00 |
| `discarded_options` | 2 | 0.375 | 0.00 |
| `invalidation_hints` | 2 | 0.750 | 0.50 |
| `non_promotable_context` | 2 | 0.333 | 0.00 |
| `payment_or_warranty_detail` | 2 | 0.350 | 0.00 |
| `assistant_claims_to_verify` | 1 | 0.500 | 0.00 |
| `explicit_decisions` | 1 | 1.000 | 1.00 |
| `resolved_context` | 1 | 0.250 | 0.00 |
| `touched_files` | 1 | 0.333 | 0.00 |
| `unsupported_hypotheses` | 1 | 0.333 | 0.00 |

## Lineage

| checkpoint | labels | overall MRR | top-1 | notes |
|---|---|---:|---:|---|
| V2 (commit 87cb089) | conversational only | 0.678 | 0.440 | 8-row adversarial eval, not comparable |
| mixed-mode stage-4 | v1 | 0.323 | 0.130 | pre-perf stack |
| mixed-mode perf-H4 | v1 | 0.506 | 0.260 | winning perf stack — but on role-tautological labels |
| mixed-mode v2 perf-H4 (this) | v2 | 0.526 | 0.296 | same perf stack, semantic-reasoning labels |

## Intended use

Ranker head of a multi-stage episodic-ingestion compiler. Input is a
JSON-serialized (requested field, candidate event) record. Output is a
scalar logit that, after grouped softmax over siblings in the same trace,
estimates the probability that the candidate event supports the requested
field.

## Limitations

- Aggregate MRR (0.526) is a blend across fields of very different difficulty.
  Legacy fields (`attempted_actions`, `observed_outcomes`) are still
  role-tautological on most rows; the model can hit ~0.48 on them without
  semantic reasoning. The signal of real learning is in the new fields:
  - `outcome_of_latest_attempt`: MRR 0.763 (2.76× random)
  - `recent_error`: MRR 0.743 (2.90× random)
  - `initiating_command`: MRR 0.667 (2.42× random)
- Not a standalone extractor. One head in a larger ingestion pipeline.

## Related work

See `docs/ranker-hypothesis-log-2026-05-08.md` in the
episodic-ingestion-compiler repo for the full experimental ladder.