Text Classification
Transformers
Safetensors
English
modernbert
episodic-ingestion-compiler
grouped-softmax-ranker
field-event-ranker
mixed-mode-training
semantic-reasoning-labels
v2-labels
text-embeddings-inference
Instructions to use Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-perf-h4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-perf-h4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-perf-h4")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-perf-h4") model = AutoModelForSequenceClassification.from_pretrained("Avifenesh/episodic-ingestion-modernbert-field-event-ranker-mixed-v2-perf-h4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 5,199 Bytes
d053378 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 | ---
license: apache-2.0
library_name: transformers
base_model: answerdotai/ModernBERT-base
tags:
- modernbert
- episodic-ingestion-compiler
- grouped-softmax-ranker
- field-event-ranker
- mixed-mode-training
- semantic-reasoning-labels
- v2-labels
language:
- en
---
# episodic-ingestion-compiler / ModernBERT field-event ranker (mixed-mode v2, perf-H4)
Fine-tune of `answerdotai/ModernBERT-base` for the **grouped multi-positive
softmax ranker** task with **v2 semantic-reasoning labels**.
## What changed vs the v1 mixed-mode checkpoint
An audit of the v1 mixed-mode training data showed that 94% of eval pairs came
from two fields (`attempted_actions`, `observed_outcomes`) whose labels were
**exact functions of event role** on 97% of rows (attempted_actions=tool_call,
observed_outcomes=tool_result). On these degenerate labels:
- E[MRR | random scores] ≈ 0.26 (random-permutation floor for multi-positive
softmax)
- v1 perf-H4 MRR = 0.506 ≈ 1.97× random
- Architecture experiments (H3, H3a, H3b, H3b-span, etc.) could not break
through because the ceiling was arithmetic, not informational.
v2 adds five semantic-reasoning fields whose labels require more than role:
| field | kind | rule |
|---|---|---|
| `initiating_command` | single-pick | First tool_call in a multi-call chain |
| `outcome_of_latest_attempt` | single-pick | Tool_result whose call_id matches the LAST tool_call |
| `recent_error` | single-pick | Most-recent tool_result with success=false |
| `attempt_outcome_pairs` | array | Paired tool_results (only when window has ≥1 orphan result — otherwise empty, so it can't collapse to role==tool_result) |
| `action_causality` | array | Tool_results whose content tokens recur in a LATER assistant turn |
Tautology rates on these new fields: 0–10%. Legacy fields kept for back-compat.
## Training details
- **Train rows**: 1217 (mixed-mode v2: autonomous + prompt→action + customer-support + conversational)
- **Train groups**: 5882 (one per field × row pair)
- **Eval rows**: 220 (1900 groups, 21986 candidate pairs)
- **Steps**: 160 optimizer steps (`--accum-steps 8`)
- **Optimizer**: AdamW, lr 5.7e-5, warmup 30
- **Precision**: bf16 autocast
- **Memory**: gradient checkpointing (use_reentrant=False)
- **Final train loss**: 0.0854
- **Peak VRAM**: 4.08 GiB
- **Hardware**: RTX 5090 Laptop (Blackwell SM 12.0, 24 GiB)
## Eval metrics (220-row mixed v2 eval, {'1': 0.2964509394572025, '2': 0.5173277661795407, '3': 0.6931106471816284, '5': 0.8881002087682672})
| metric | value |
|---|---:|
| overall MRR | 0.526 |
| top-1 recall | 0.296 |
| top-2 recall | 0.517 |
| top-3 recall | 0.693 |
| top-5 recall | 0.888 |
| mean expected rank | 3.09 |
Per-field MRR (sorted by eval-pair count):
| field | n | MRR | top-1 |
|---|---:|---:|---:|
| `attempted_actions` | 846 | 0.469 | 0.23 |
| `observed_outcomes` | 845 | 0.486 | 0.24 |
| `initiating_command` | 199 | 0.667 | 0.50 |
| `outcome_of_latest_attempt` | 181 | 0.763 | 0.60 |
| `action_causality` | 141 | 0.483 | 0.24 |
| `attempt_outcome_pairs` | 53 | 0.563 | 0.30 |
| `failed_attempts` | 36 | 0.615 | 0.42 |
| `next_actions` | 32 | 0.565 | 0.34 |
| `recent_error` | 29 | 0.743 | 0.59 |
| `customer_identity` | 11 | 0.818 | 0.73 |
| `transaction_reference` | 5 | 0.467 | 0.20 |
| `product_name` | 4 | 1.000 | 1.00 |
| `discarded_options` | 2 | 0.375 | 0.00 |
| `invalidation_hints` | 2 | 0.750 | 0.50 |
| `non_promotable_context` | 2 | 0.333 | 0.00 |
| `payment_or_warranty_detail` | 2 | 0.350 | 0.00 |
| `assistant_claims_to_verify` | 1 | 0.500 | 0.00 |
| `explicit_decisions` | 1 | 1.000 | 1.00 |
| `resolved_context` | 1 | 0.250 | 0.00 |
| `touched_files` | 1 | 0.333 | 0.00 |
| `unsupported_hypotheses` | 1 | 0.333 | 0.00 |
## Lineage
| checkpoint | labels | overall MRR | top-1 | notes |
|---|---|---:|---:|---|
| V2 (commit 87cb089) | conversational only | 0.678 | 0.440 | 8-row adversarial eval, not comparable |
| mixed-mode stage-4 | v1 | 0.323 | 0.130 | pre-perf stack |
| mixed-mode perf-H4 | v1 | 0.506 | 0.260 | winning perf stack — but on role-tautological labels |
| mixed-mode v2 perf-H4 (this) | v2 | 0.526 | 0.296 | same perf stack, semantic-reasoning labels |
## Intended use
Ranker head of a multi-stage episodic-ingestion compiler. Input is a
JSON-serialized (requested field, candidate event) record. Output is a
scalar logit that, after grouped softmax over siblings in the same trace,
estimates the probability that the candidate event supports the requested
field.
## Limitations
- Aggregate MRR (0.526) is a blend across fields of very different difficulty.
Legacy fields (`attempted_actions`, `observed_outcomes`) are still
role-tautological on most rows; the model can hit ~0.48 on them without
semantic reasoning. The signal of real learning is in the new fields:
- `outcome_of_latest_attempt`: MRR 0.763 (2.76× random)
- `recent_error`: MRR 0.743 (2.90× random)
- `initiating_command`: MRR 0.667 (2.42× random)
- Not a standalone extractor. One head in a larger ingestion pipeline.
## Related work
See `docs/ranker-hypothesis-log-2026-05-08.md` in the
episodic-ingestion-compiler repo for the full experimental ladder.
|