Qwen3.5-0.8B-JEV

This is a compact decision model derived from Qwen3.5-0.8B. It accepts a state plus one or more dynamic questions and returns a probability distribution over the candidates of each question. The model uses a query/key decision head stored in head.safetensors.

Inference code and examples are available at joyfoxai/jev-inference.

This is not a chat or text-generation model. It must be loaded with the JEV decision-model wrapper; AutoModelForCausalLM, standard text-generation pipelines, vLLM, Ollama, and GGUF loaders do not expose the decision head.

Model details

Property Value
Base model Qwen/Qwen3.5-0.8B
Task Dynamic-candidate probabilistic decision modeling
Supported question types choice, noul, score
Architecture Qwen3.5 text backbone plus query/key decision head
Decision head query/key projections, dimension 128
Export precision BF16
Execution mode rows
Maximum supported input length 1,024 tokens
License Apache-2.0

The original causal language-model head, vision tower, and MTP components are not part of this decision artifact. The export contains the text backbone, tokenizer, decision head, and decision-model metadata.

Architecture

Given context (x), question (q), and candidate set (C), the model estimates a distribution (p(c \mid x, q, C)) over the candidates. The candidate set is part of the input, so adding, removing, or rewriting a candidate may change the probabilities assigned to every candidate.

Each record first encodes the shared state, followed by the question instructions, candidate descriptions, and a decision suffix. The model reads the hidden states at the end of each candidate and at the end of the decision suffix. Two linear projections produce a decision query and candidate keys:

u   = Wq h_decision + bq
v_i = Wk h_candidate_i + bk
z_i = dot(u, v_i) / sqrt(d)
p_i = softmax(z)_i

The head contains 2 × (H × d + d) parameters and does not grow with the number of candidates. It represents candidates from their text descriptions instead of assigning a fixed classifier weight to each business label.

Qwen3.5 uses rows execution in this implementation. Every question is evaluated as a separate state-plus-question sequence. Questions cannot read the candidates or decision suffixes of other questions, although the shared state is recomputed for each row. Inputs that exceed cutoff_len are rejected instead of silently truncating candidates or targets.

Input format

Each request contains a state and a mapping of questions. Candidate sets may vary between requests.

record = {
    "state": "The customer was charged twice and requests a refund.",
    "questions": {
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this ticket?",
            "criteria": {
                "billing": "Payments, invoicing, or refunds",
                "account": "Account access or profile",
                "technical": "Bugs or integrations",
            },
        },
        "refund_requested": {
            "type": "noul",
            "instructions": "Does the customer request a refund?",
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgently should this ticket be handled?",
            "criteria": [
                "Can wait in the normal queue",
                "Needs priority handling",
                "Needs immediate human handling",
            ],
        },
    },
}
  • choice: criteria is a mapping from stable candidate IDs to candidate descriptions.
  • noul: a Boolean decision; output keys are false and true.
  • score: criteria is an ordered list of score-level descriptions; output keys are zero-based level indices.

Usage

Install the standalone inference package:

git clone https://github.com/joyfoxai/jev-inference.git
cd jev-inference
pip install -e .

Then load this model directly from Hugging Face:

from jev_inference import DecisionEngine

engine = DecisionEngine.load(
    "joyfox/Qwen3.5-0.8B-JEV",
    device="cuda",
    dtype="bfloat16",
)
result = engine.predict(record)
print(result)

Example output shape:

{
    "route": {
        "type": "choice",
        "probabilities": {"billing": 0.96, "account": 0.02, "technical": 0.02},
        "confidence": 0.96,
        "choice": "billing",
    },
    "refund_requested": {
        "type": "noul",
        "probabilities": {"false": 0.01, "true": 0.99},
        "confidence": 0.99,
        "noul": 0.99,
    },
    "urgency": {
        "type": "score",
        "probabilities": {"0": 0.02, "1": 0.55, "2": 0.43},
        "confidence": 0.55,
        "score": 1.41,
        "level": 1,
    },
}

Keep the DecisionEngine instance loaded and use predict_batch when processing multiple records.

Files

backbone/
  config.json
  model.safetensors
tokenizer/
  chat_template.jinja
  tokenizer.json
  tokenizer_config.json
decision_config.json
head.safetensors
README.md

The backbone directory contains the complete text-backbone weights required by this artifact.

Objective and probability semantics

The model learns complete target probability distributions rather than only argmax labels. For a prediction (p) and target distribution (y), the reference objectives are:

Cross-entropy: L_CE    = -sum_i y_i log(p_i)
Brier:        L_Brier =  sum_i (p_i - y_i)^2

Softmax guarantees non-negative probabilities that sum to one, but it does not guarantee that the distribution is calibrated or that the supplied candidates cover every real-world possibility. A teacher probability of 0.8 represents the learned teacher preference; it must not automatically be interpreted as an 80% real-world event frequency. Calibration and rejection thresholds require an independent dataset with observed outcomes.

Evaluation

The full evaluation set contains 30,612 records and 52,664 questions. Metrics compare this model with the archived soft distributions from typesafe/jev-1.13-20260917.

Metric Result
Teacher-argmax agreement 90.40%
Mean KL divergence to teacher 0.07373
Brier distance to teacher 0.03520
Score expected-level MAE 0.16246

By question type

Type Questions Argmax agreement Mean KL Brier
noul 20,912 95.84% 0.01681 0.01250
score 11,026 90.06% 0.05855 0.03191
choice 20,726 85.08% 0.13922 0.05984

These are distillation-fidelity metrics, not independent real-world accuracy. The evaluation set was generated by the teacher and was used during model selection. It contains no human-verified hard labels. A separate, human-labeled test set is required to measure task accuracy and calibration.

Measured inference performance

Measurements were made on a single GPU with BF16 inference:

  • Batch size 8: 70.22 questions/second and 40.82 records/second.
  • Batch size 1, one question per request, warm model: mean 78.25 ms, P50 73.70 ms, P95 97.26 ms.

These numbers exclude model loading and depend on hardware, input length, candidate count, and software versions. The current wrapper evaluates the full input with use_cache=False.

Limitations

  • The model imitates a teacher model and may reproduce its errors, preferences, and biases.
  • Teacher agreement must not be interpreted as human-label accuracy.
  • Dynamic choice questions, especially those with 8–10 candidates, show the largest fidelity gap.
  • Inputs longer than the configured cutoff are rejected rather than silently truncated.
  • The model does not provide calibrated confidence against real-world outcomes without an independent outcome-labeled calibration set.
  • The artifact is text-only and does not retain the base model's vision or generation capabilities.
  • High-stakes decisions require domain validation and appropriate human oversight.

Intended use

Suitable uses include research, prototyping, routing, ranking, triage, and other probabilistic decision tasks where candidate definitions are supplied at runtime. Users are responsible for validating the model on their own independent data and complying with applicable laws and policies.

This release is not affiliated with or endorsed by TypeSafe, JEV, or the Qwen team.

License and attribution

Released under Apache-2.0, consistent with the base model. See the original Qwen/Qwen3.5-0.8B model card and license for base-model details.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for joyfox/Qwen3.5-0.8B-JEV

Finetuned
(454)
this model