Instructions to use joyfox/Qwen3.5-0.8B-JEV with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use joyfox/Qwen3.5-0.8B-JEV with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("joyfox/Qwen3.5-0.8B-JEV", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qwen3.5-0.8B-JEV
This is a compact decision model derived from Qwen3.5-0.8B. It accepts a state plus one or more
dynamic questions and returns a probability distribution over the candidates of each question.
The model uses a query/key decision head stored in head.safetensors.
Inference code and examples are available at
joyfoxai/jev-inference.
This is not a chat or text-generation model. It must be loaded with the JEV decision-model wrapper;
AutoModelForCausalLM, standard text-generation pipelines, vLLM, Ollama, and GGUF loaders do not expose the decision head.
Model details
| Property | Value |
|---|---|
| Base model | Qwen/Qwen3.5-0.8B |
| Task | Dynamic-candidate probabilistic decision modeling |
| Supported question types | choice, noul, score |
| Architecture | Qwen3.5 text backbone plus query/key decision head |
| Decision head | query/key projections, dimension 128 |
| Export precision | BF16 |
| Execution mode | rows |
| Maximum supported input length | 1,024 tokens |
| License | Apache-2.0 |
The original causal language-model head, vision tower, and MTP components are not part of this decision artifact. The export contains the text backbone, tokenizer, decision head, and decision-model metadata.
Architecture
Given context (x), question (q), and candidate set (C), the model estimates a distribution (p(c \mid x, q, C)) over the candidates. The candidate set is part of the input, so adding, removing, or rewriting a candidate may change the probabilities assigned to every candidate.
Each record first encodes the shared state, followed by the question instructions, candidate descriptions, and a decision suffix. The model reads the hidden states at the end of each candidate and at the end of the decision suffix. Two linear projections produce a decision query and candidate keys:
u = Wq h_decision + bq
v_i = Wk h_candidate_i + bk
z_i = dot(u, v_i) / sqrt(d)
p_i = softmax(z)_i
The head contains 2 × (H × d + d) parameters and does not grow with the number of candidates. It
represents candidates from their text descriptions instead of assigning a fixed classifier weight to
each business label.
Qwen3.5 uses rows execution in this implementation. Every question is evaluated as a separate
state-plus-question sequence. Questions cannot read the candidates or decision suffixes of other
questions, although the shared state is recomputed for each row. Inputs that exceed cutoff_len are
rejected instead of silently truncating candidates or targets.
Input format
Each request contains a state and a mapping of questions. Candidate sets may vary between requests.
record = {
"state": "The customer was charged twice and requests a refund.",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments, invoicing, or refunds",
"account": "Account access or profile",
"technical": "Bugs or integrations",
},
},
"refund_requested": {
"type": "noul",
"instructions": "Does the customer request a refund?",
},
"urgency": {
"type": "score",
"instructions": "How urgently should this ticket be handled?",
"criteria": [
"Can wait in the normal queue",
"Needs priority handling",
"Needs immediate human handling",
],
},
},
}
choice:criteriais a mapping from stable candidate IDs to candidate descriptions.noul: a Boolean decision; output keys arefalseandtrue.score:criteriais an ordered list of score-level descriptions; output keys are zero-based level indices.
Usage
Install the standalone inference package:
git clone https://github.com/joyfoxai/jev-inference.git
cd jev-inference
pip install -e .
Then load this model directly from Hugging Face:
from jev_inference import DecisionEngine
engine = DecisionEngine.load(
"joyfox/Qwen3.5-0.8B-JEV",
device="cuda",
dtype="bfloat16",
)
result = engine.predict(record)
print(result)
Example output shape:
{
"route": {
"type": "choice",
"probabilities": {"billing": 0.96, "account": 0.02, "technical": 0.02},
"confidence": 0.96,
"choice": "billing",
},
"refund_requested": {
"type": "noul",
"probabilities": {"false": 0.01, "true": 0.99},
"confidence": 0.99,
"noul": 0.99,
},
"urgency": {
"type": "score",
"probabilities": {"0": 0.02, "1": 0.55, "2": 0.43},
"confidence": 0.55,
"score": 1.41,
"level": 1,
},
}
Keep the DecisionEngine instance loaded and use predict_batch when processing multiple records.
Files
backbone/
config.json
model.safetensors
tokenizer/
chat_template.jinja
tokenizer.json
tokenizer_config.json
decision_config.json
head.safetensors
README.md
The backbone directory contains the complete text-backbone weights required by this artifact.
Objective and probability semantics
The model learns complete target probability distributions rather than only argmax labels. For a prediction (p) and target distribution (y), the reference objectives are:
Cross-entropy: L_CE = -sum_i y_i log(p_i)
Brier: L_Brier = sum_i (p_i - y_i)^2
Softmax guarantees non-negative probabilities that sum to one, but it does not guarantee that the distribution is calibrated or that the supplied candidates cover every real-world possibility. A teacher probability of 0.8 represents the learned teacher preference; it must not automatically be interpreted as an 80% real-world event frequency. Calibration and rejection thresholds require an independent dataset with observed outcomes.
Evaluation
The full evaluation set contains 30,612 records and 52,664 questions. Metrics compare this model with
the archived soft distributions from typesafe/jev-1.13-20260917.
| Metric | Result |
|---|---|
| Teacher-argmax agreement | 90.40% |
| Mean KL divergence to teacher | 0.07373 |
| Brier distance to teacher | 0.03520 |
| Score expected-level MAE | 0.16246 |
By question type
| Type | Questions | Argmax agreement | Mean KL | Brier |
|---|---|---|---|---|
noul |
20,912 | 95.84% | 0.01681 | 0.01250 |
score |
11,026 | 90.06% | 0.05855 | 0.03191 |
choice |
20,726 | 85.08% | 0.13922 | 0.05984 |
These are distillation-fidelity metrics, not independent real-world accuracy. The evaluation set was generated by the teacher and was used during model selection. It contains no human-verified hard labels. A separate, human-labeled test set is required to measure task accuracy and calibration.
Measured inference performance
Measurements were made on a single GPU with BF16 inference:
- Batch size 8: 70.22 questions/second and 40.82 records/second.
- Batch size 1, one question per request, warm model: mean 78.25 ms, P50 73.70 ms, P95 97.26 ms.
These numbers exclude model loading and depend on hardware, input length, candidate count, and
software versions. The current wrapper evaluates the full input with use_cache=False.
Limitations
- The model imitates a teacher model and may reproduce its errors, preferences, and biases.
- Teacher agreement must not be interpreted as human-label accuracy.
- Dynamic
choicequestions, especially those with 8–10 candidates, show the largest fidelity gap. - Inputs longer than the configured cutoff are rejected rather than silently truncated.
- The model does not provide calibrated confidence against real-world outcomes without an independent outcome-labeled calibration set.
- The artifact is text-only and does not retain the base model's vision or generation capabilities.
- High-stakes decisions require domain validation and appropriate human oversight.
Intended use
Suitable uses include research, prototyping, routing, ranking, triage, and other probabilistic decision tasks where candidate definitions are supplied at runtime. Users are responsible for validating the model on their own independent data and complying with applicable laws and policies.
This release is not affiliated with or endorsed by TypeSafe, JEV, or the Qwen team.
License and attribution
Released under Apache-2.0, consistent with the base model. See the original
Qwen/Qwen3.5-0.8B model card and license for base-model
details.