Instructions to use distil-labs/distil-qwen3.5-4b-flight-connection-check-no-thinking with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use distil-labs/distil-qwen3.5-4b-flight-connection-check-no-thinking with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "distil-labs/distil-qwen3.5-4b-flight-connection-check-no-thinking") - Notebooks
- Google Colab
- Kaggle
distil-qwen3.5-4b-flight-connection-check-no-thinking
A LoRA adapter for Qwen3.5-4B that checks every connection of a 3 or 4 leg flight itinerary, given local departure times, durations, UTC offsets and minimum connection times, and reports the first connection that is too short or too long. Trained without reasoning: it answers directly, as the one-pass control for its reasoning twin. Its twin distil-qwen3.5-4b-flight-connection-check was trained on the same synthetic data with reasoning on; the pair is part of the reasoning SLM benchmark.
Results (100 test cases, exact match)
| Model | Correct |
|---|---|
| Fine-tuned, thinking on | 95 (median 208.5 reasoning tokens) |
| Fine-tuned, thinking off | 45 |
| Qwen3.5-4B untuned, thinking on | 0 (100 of 100 cut off at 2,048 tokens) |
| Qwen3.5-4B untuned, thinking off | 13 |
Data, configs, predictions and the scoring script: https://github.com/distil-labs/reasoning-blogpost-benchmark.
How to use it
Serve the base model with this adapter:
vllm serve Qwen/Qwen3.5-4B --enable-lora --max-lora-rank 64 --lora-modules distil-qwen3.5-4b-flight-connection-check-no-thinking=distil-labs/distil-qwen3.5-4b-flight-connection-check-no-thinking --port 8001
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8001/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="distil-qwen3.5-4b-flight-connection-check-no-thinking",
messages=[{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": text}],
temperature=0,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
Example answer: {"valid": false, "problem": "short_connection", "leg": 2}. Keep the system prompt exactly as below and serve with thinking off, as trained.
System prompt
You check flight itineraries for a corporate travel desk before they are ticketed.
You receive an itinerary of 3 or 4 legs and an airport reference. Each leg gives its departure time in the local time of its departure airport and its scheduled flight duration; arrival times are not given. The airport reference gives each airport's UTC offset on the travel dates and its minimum connection time (MCT).
Check each connection in order, from the first to the last, and stop at the first one that fails. The connection time is the time from the previous leg's arrival to the next leg's departure, at the connecting airport.
1. short_connection: the connection time is less than the MCT of the connecting airport. A connection time exactly equal to the MCT is fine.
2. long_layover: the connection time is more than 720 minutes (12 hours). Exactly 720 minutes is fine.
If every connection passes, the itinerary is valid.
Answer with a JSON object and nothing else: {"valid": true or false, "problem": "short_connection", "long_layover" or null, "leg": the number of the leg that departs after the failing connection, or null}. A valid itinerary is {"valid": true, "problem": null, "leg": null}.
Training
| Base model | Qwen/Qwen3.5-4B |
| Teacher | Kimi K3, reasoning effort max |
| Thinking | off (enable_thinking: false) |
| Seed examples | 40 |
| Synthetic examples | 4,067 |
| Method | LoRA (rank 64), 4 epochs; this repo holds the adapter only |
Limits
The data is synthetic and in English. The model is trained for this one task and policy; it is not a general assistant.
Links
distil labs 路 GitHub 路 Hugging Face 路 X
- Downloads last month
- 12