KOTA Decision V16 โ€” WIP

This repo is a development snapshot of the KOTA construction router evolving into a typed decision engine.

Base model

V16 is based on the trained V15.1 checkpoint:

kota_construction_router_v15_1_replay/final

The original V15.1 model is preserved as the stable base.

V16 is intended to reuse the trained V15.1 RoBERTa encoder and replace the old fixed classifier head with a typed-decision head.

Current construction intents

  • not_construction
  • dpr
  • material_purchase_request
  • material_delivery
  • manpower_attendance
  • manpower_request
  • machine_request
  • machine_delivery
  • site_issue
  • schedule_lookahead
  • drawing_rfi
  • qaqc_testing
  • other_construction

V14 results

Hard benchmark:

  • 48 / 63 correct
  • Accuracy: 76.19%
  • Macro F1: 0.7231
  • Weighted F1: 0.7425

Original handwritten benchmark:

  • 27 / 43 correct
  • Accuracy: 62.79%

Main issue:

V14 was highly overconfident on some out-of-domain inputs and often forced them into construction classes.

V15

V15 focused on improving the boundary between:

  • not_construction
  • other_construction

This improved OOD behavior but caused specialist-class forgetting.

Example regression:

rebar inspection required before casting

was incorrectly routed away from QA/QC.

V15.1

V15.1 used balanced replay to restore specialist intents.

Training:

  • about 53,000 examples
  • 2 epochs
  • balanced specialist-class replay

Quick regression check:

  • 13 / 13 correct

QA/QC behavior was restored.

Synthetic validation is not treated as proof of generalization.

Multi-hop work

Multi-clause construction messages were added.

Examples include:

  • progress + attendance + manpower request
  • material delivery + new material request
  • machine delivery + new machine request
  • drawing ambiguity + inspection
  • QA/QC + drawing status
  • site progress + blocker
  • current work + future schedule
  • neutral construction information
  • non-construction messages containing construction-like words

Balanced multi-hop training set:

  • manpower_request: 1000
  • manpower_attendance: 600
  • dpr: 600
  • site_issue: 800
  • qaqc_testing: 650
  • drawing_rfi: 650
  • machine_request: 650
  • machine_delivery: 550
  • material_purchase_request: 650
  • material_delivery: 550
  • schedule_lookahead: 650
  • other_construction: 650
  • not_construction: 650

Total:

8650

Holdout:

  • 80 per class
  • 1040 total
  • exact train / holdout text overlap: 0

A cleanup step was added before deduplication to remove generated text artifacts such as duplicated words, repeated prefixes, repeated suffixes, and punctuation issues.

V16 direction

V16 is moving from a fixed classifier toward a typed decision engine.

Target primitives:

choice

Select between arbitrary runtime options.

noul

Return a boolean probability.

score

Return an ordinal probability distribution and expected score.

ACT / ESCALATE

Estimate whether the model should make the decision or defer.

Planned architecture

V15.1 trained RoBERTa encoder

-> state + question + runtime options

-> option representations

-> 2-layer decision transformer

-> choice / noul / score

-> ACT / ESCALATE

-> calibration

Calibration work

Current experimental components include:

  • cross-entropy loss
  • Brier objective
  • Expected Calibration Error
  • temperature scaling
  • option-cardinality-aware temperatures

General-purpose direction

Construction remains the anchor domain.

The goal is to make the same typed decision head usable for runtime decisions in areas such as:

  • support routing
  • workflow decisions
  • project status
  • document review
  • inventory
  • logistics
  • boolean checks
  • ordinal scoring

Current status

Completed:

  • V14 baseline
  • V15 OOD repair
  • V15.1 balanced replay
  • multi-hop data generation
  • balanced multi-hop dataset
  • multi-hop cleanup
  • typed-decision architecture draft
  • choice
  • noul
  • score
  • ACT / ESCALATE
  • Brier objective
  • temperature calibration logic
  • cardinality calibration buckets
  • decision to reuse V15.1 as the V16 backbone

Not complete yet:

  • final V16 training
  • final general-purpose benchmark
  • independent multi-hop template-family holdout
  • full calibration benchmark
  • regression comparison against V15.1

Next evaluation gates

  • frozen 63-case hard benchmark
  • frozen 43-case handwritten benchmark
  • clean CLINC150 OOD test
  • independent multi-hop benchmark
  • per-class precision / recall / F1
  • high-confidence errors
  • Brier score
  • NLL
  • ECE

This repository is intentionally marked WIP.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support