umarfarookm commited on
Commit
4bfd389
·
verified ·
1 Parent(s): 218443e

Unsloth Model Card

Browse files
Files changed (1) hide show
  1. README.md +13 -160
README.md CHANGED
@@ -1,168 +1,21 @@
1
  ---
2
- license: apache-2.0
3
- base_model: Qwen/Qwen2.5-1.5B-Instruct
4
  tags:
5
- - transit
6
- - gtfs
7
- - transportation
8
- - instruction-following
9
- - qwen2
10
- - qlora
11
- - unsloth
12
  language:
13
- - en
14
- datasets:
15
- - umarfarookm/UmarTransit-Instruct-3k
16
- pipeline_tag: text-generation
17
  ---
18
 
19
- # UmarTransit-1B
20
-
21
- A domain-specific language model for **public transit systems** and **GTFS (General Transit Feed Specification)** data, fine-tuned from Qwen2.5-1.5B-Instruct.
22
-
23
- UmarTransit-1B specializes in:
24
- - GTFS understanding and validation
25
- - Transit route and schedule analysis
26
- - Stop/station information
27
- - Transfer optimization
28
- - Transit network statistics
29
- - Cross-agency comparisons
30
-
31
- > **Data Disclaimer:** This model was trained **exclusively on publicly available, open-source GTFS feeds** published by transit agencies for public use via the [Mobility Database](https://mobilitydatabase.org/). **No private, proprietary, or NDA-protected data** from any client, employer, or organization was used at any stage.
32
-
33
- ## Model Details
34
-
35
- | Property | Value |
36
- |----------|-------|
37
- | **Base Model** | [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) |
38
- | **Parameters** | 1.54B (1.31B non-embedding) |
39
- | **Fine-tuning** | QLoRA (4-bit NF4, LoRA rank=16, alpha=32) |
40
- | **Training Framework** | [Unsloth](https://unsloth.ai) + HuggingFace TRL |
41
- | **Training Data** | 2,971 pairs from [UmarTransit-Instruct-3k](https://huggingface.co/datasets/umarfarookm/UmarTransit-Instruct-3k) |
42
- | **Test Data** | 335 pairs (stratified 90/10 split) |
43
- | **Max Context** | 1,024 tokens |
44
- | **License** | Apache 2.0 |
45
- | **Developer** | [umarfarookm](https://github.com/umarfarookm) |
46
-
47
- ## Evaluation Results
48
-
49
- Evaluated on 335 held-out test pairs across 8 task categories:
50
-
51
- | Metric | Score |
52
- |--------|-------|
53
- | **ROUGE-L** | 0.8192 |
54
- | **Keyword Match** | 0.4086 |
55
-
56
- **Best performing:** Transfer analysis (ROUGE-L: 0.90)
57
- **Needs improvement:** GTFS knowledge (ROUGE-L: 0.38) — limited training data (22 pairs)
58
-
59
- ## Available Formats
60
-
61
- | Format | File | Size | Use Case |
62
- |--------|------|------|----------|
63
- | Safetensors | `model.safetensors` | 3.09 GB | Full precision — Transformers/Python |
64
- | GGUF Q4_K_M | `UmarTransit-1B.Q4_K_M.gguf` | 986 MB | 4-bit — Ollama/llama.cpp (recommended) |
65
- | GGUF Q8_0 | `UmarTransit-1B.Q8_0.gguf` | 1.65 GB | 8-bit — Ollama/llama.cpp (higher quality) |
66
-
67
- ## Training Data
68
-
69
- The model was trained on synthetic instruction-response pairs generated from **15 real public GTFS feeds** across **10 countries**:
70
-
71
- | Country | Agencies |
72
- |---------|----------|
73
- | US | LA Metro, Chicago CTA, Boston MBTA, Valley Metro, Capital Metro, TriMet |
74
- | Canada | Toronto TTC |
75
- | Germany | Berlin VBB |
76
- | France | Ile-de-France Mobilites (Paris) |
77
- | Netherlands | OVapi (national) |
78
- | Belgium | NMBS/SNCB Railways |
79
- | Finland | HSL Helsinki |
80
- | Denmark | Rejseplanen |
81
- | Australia | Transperth (Perth) |
82
- | New Zealand | Auckland Transport |
83
-
84
- **8 task categories:** Agency overview, route information, stop/station info, trip schedules, transfer analysis, network statistics, GTFS knowledge, comparative analysis.
85
-
86
- ## Usage
87
-
88
- ### With Transformers
89
-
90
- ```python
91
- from transformers import AutoModelForCausalLM, AutoTokenizer
92
- import torch
93
-
94
- model = AutoModelForCausalLM.from_pretrained(
95
- "umarfarookm/UmarTransit-1B",
96
- torch_dtype=torch.bfloat16,
97
- device_map="auto",
98
- )
99
- tokenizer = AutoTokenizer.from_pretrained("umarfarookm/UmarTransit-1B")
100
-
101
- messages = [
102
- {"role": "system", "content": "You are UmarTransit-1B, a specialized AI assistant for public transit systems and GTFS data."},
103
- {"role": "user", "content": "What does route_type 3 mean in GTFS?"},
104
- ]
105
-
106
- text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
107
- inputs = tokenizer(text, return_tensors="pt").to(model.device)
108
-
109
- outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.1, do_sample=True)
110
- response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
111
- print(response)
112
- ```
113
-
114
- ### With Ollama (GGUF)
115
-
116
- ```bash
117
- # Download the GGUF file from this repo, then:
118
- ollama create umartransit -f Modelfile
119
- ollama run umartransit "What are the required files in a GTFS feed?"
120
- ```
121
-
122
- ## Training Configuration
123
-
124
- ```
125
- QLoRA Config:
126
- rank: 16
127
- alpha: 32
128
- dropout: 0
129
- target_modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
130
-
131
- Training:
132
- epochs: 3
133
- batch_size: 4 x 4 gradient accumulation = 16 effective
134
- learning_rate: 2e-4
135
- scheduler: cosine
136
- optimizer: adamw_8bit
137
- hardware: Google Colab T4 GPU (15GB VRAM)
138
- ```
139
-
140
- ## Limitations
141
-
142
- - **Small training dataset:** 2,971 pairs — model may hallucinate specific details (coordinates, exact counts)
143
- - **Limited GTFS knowledge:** Only 22 GTFS specification Q&A pairs in training
144
- - **English-primary:** Trained on English instructions, though base model supports 29 languages
145
- - **Static data:** Trained on GTFS schedule data, not real-time transit information
146
- - **Not a trip planner:** Cannot compute actual routes or real-time ETAs
147
-
148
- ## Future Improvements
149
-
150
- - Add more GTFS knowledge pairs (target 100+)
151
- - Include Indian city transit feeds (Chennai, Bangalore, Mumbai)
152
- - Expand to 10K+ training pairs for better factual accuracy
153
- - Add GTFS-Realtime understanding
154
-
155
- ## Source Code
156
 
157
- [github.com/umarfarookm/transit-foundation-model](https://github.com/umarfarookm/transit-foundation-model)
 
 
158
 
159
- ## Citation
160
 
161
- ```bibtex
162
- @misc{umartransit1b,
163
- title={UmarTransit-1B: A Domain-Specific Language Model for Public Transit},
164
- author={umarfarookm},
165
- year={2026},
166
- url={https://huggingface.co/umarfarookm/UmarTransit-1B}
167
- }
168
- ```
 
1
  ---
2
+ base_model: unsloth/Qwen2.5-1.5B-Instruct-bnb-4bit
 
3
  tags:
4
+ - text-generation-inference
5
+ - transformers
6
+ - unsloth
7
+ - qwen2
8
+ license: apache-2.0
 
 
9
  language:
10
+ - en
 
 
 
11
  ---
12
 
13
+ # Uploaded finetuned model
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
15
+ - **Developed by:** umarfarookm
16
+ - **License:** apache-2.0
17
+ - **Finetuned from model :** unsloth/Qwen2.5-1.5B-Instruct-bnb-4bit
18
 
19
+ This qwen2 model was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth) and Huggingface's TRL library.
20
 
21
+ [<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth)