umarfarookm commited on
Commit
ce2ba7f
·
verified ·
1 Parent(s): 9574162

Updated README

Browse files
Files changed (1) hide show
  1. README.md +150 -13
README.md CHANGED
@@ -1,21 +1,158 @@
1
  ---
2
- base_model: unsloth/Qwen2.5-1.5B-Instruct-bnb-4bit
3
- tags:
4
- - text-generation-inference
5
- - transformers
6
- - unsloth
7
- - qwen2
8
  license: apache-2.0
 
 
 
 
 
 
 
 
 
9
  language:
10
- - en
 
 
 
11
  ---
12
 
13
- # Uploaded finetuned model
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
15
- - **Developed by:** umarfarookm
16
- - **License:** apache-2.0
17
- - **Finetuned from model :** unsloth/Qwen2.5-1.5B-Instruct-bnb-4bit
18
 
19
- This qwen2 model was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth) and Huggingface's TRL library.
20
 
21
- [<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth)
 
 
 
 
 
 
 
 
1
  ---
 
 
 
 
 
 
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-1.5B-Instruct
4
+ tags:
5
+ - transit
6
+ - gtfs
7
+ - transportation
8
+ - instruction-following
9
+ - qwen2
10
+ - qlora
11
+ - unsloth
12
  language:
13
+ - en
14
+ dataset_info:
15
+ dataset_name: UmarTransit Synthetic Q&A
16
+ pipeline_tag: text-generation
17
  ---
18
 
19
+ # UmarTransit-1B
20
+
21
+ A domain-specific language model for **public transit systems** and **GTFS (General Transit Feed Specification)** data, fine-tuned from Qwen2.5-1.5B-Instruct.
22
+
23
+ UmarTransit-1B specializes in:
24
+ - GTFS understanding and validation
25
+ - Transit route and schedule analysis
26
+ - Stop/station information
27
+ - Transfer optimization
28
+ - Transit network statistics
29
+ - Cross-agency comparisons
30
+
31
+ ## Model Details
32
+
33
+ | Property | Value |
34
+ |----------|-------|
35
+ | **Base Model** | [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) |
36
+ | **Parameters** | 1.54B (1.31B non-embedding) |
37
+ | **Fine-tuning** | QLoRA (4-bit NF4, LoRA rank=16, alpha=32) |
38
+ | **Training Framework** | [Unsloth](https://unsloth.ai) + HuggingFace TRL |
39
+ | **Training Data** | 2,971 synthetic instruction-response pairs |
40
+ | **Test Data** | 335 pairs (stratified 90/10 split) |
41
+ | **Max Context** | 1,024 tokens |
42
+ | **License** | Apache 2.0 |
43
+ | **Developer** | [umarfarookm](https://github.com/umarfarookm) |
44
+
45
+ ## Evaluation Results
46
+
47
+ Evaluated on 335 held-out test pairs across 8 task categories:
48
+
49
+ | Metric | Score |
50
+ |--------|-------|
51
+ | **ROUGE-L** | 0.8192 |
52
+ | **Keyword Match** | 0.4086 |
53
+
54
+ **Best performing:** Transfer analysis (ROUGE-L: 0.90)
55
+ **Needs improvement:** GTFS knowledge (ROUGE-L: 0.38) — limited training data (22 pairs)
56
+
57
+ ## Training Data
58
+
59
+ The model was trained on synthetic instruction-response pairs generated from **15 real public GTFS feeds** across **10 countries**:
60
+
61
+ | Country | Agencies |
62
+ |---------|----------|
63
+ | US | LA Metro, Chicago CTA, Boston MBTA, Valley Metro, Capital Metro, TriMet |
64
+ | Canada | Toronto TTC |
65
+ | Germany | Berlin VBB |
66
+ | France | Ile-de-France Mobilites (Paris) |
67
+ | Netherlands | OVapi (national) |
68
+ | Belgium | NMBS/SNCB Railways |
69
+ | Finland | HSL Helsinki |
70
+ | Denmark | Rejseplanen |
71
+ | Australia | Transperth (Perth) |
72
+ | New Zealand | Auckland Transport |
73
+
74
+ **8 task categories:** Agency overview, route information, stop/station info, trip schedules, transfer analysis, network statistics, GTFS knowledge, comparative analysis.
75
+
76
+ ## Usage
77
+
78
+ ### With Transformers
79
+
80
+ ```python
81
+ from transformers import AutoModelForCausalLM, AutoTokenizer
82
+ import torch
83
+
84
+ model = AutoModelForCausalLM.from_pretrained(
85
+ "umarfarookm/UmarTransit-1B",
86
+ torch_dtype=torch.bfloat16,
87
+ device_map="auto",
88
+ )
89
+ tokenizer = AutoTokenizer.from_pretrained("umarfarookm/UmarTransit-1B")
90
+
91
+ messages = [
92
+ {"role": "system", "content": "You are UmarTransit-1B, a specialized AI assistant for public transit systems and GTFS data."},
93
+ {"role": "user", "content": "What does route_type 3 mean in GTFS?"},
94
+ ]
95
+
96
+ text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
97
+ inputs = tokenizer(text, return_tensors="pt").to(model.device)
98
+
99
+ outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.1, do_sample=True)
100
+ response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
101
+ print(response)
102
+ ```
103
+
104
+ ### With Ollama (GGUF)
105
+
106
+ ```bash
107
+ # Download the GGUF file from this repo, then:
108
+ ollama create umartransit -f Modelfile
109
+ ollama run umartransit "What are the required files in a GTFS feed?"
110
+ ```
111
+
112
+ ## Training Configuration
113
+
114
+ ```
115
+ QLoRA Config:
116
+ rank: 16
117
+ alpha: 32
118
+ dropout: 0
119
+ target_modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
120
+
121
+ Training:
122
+ epochs: 3
123
+ batch_size: 4 x 4 gradient accumulation = 16 effective
124
+ learning_rate: 2e-4
125
+ scheduler: cosine
126
+ optimizer: adamw_8bit
127
+ hardware: Google Colab T4 GPU (15GB VRAM)
128
+ ```
129
+
130
+ ## Limitations
131
+
132
+ - **Small training dataset:** 2,971 pairs — model may hallucinate specific details (coordinates, exact counts)
133
+ - **Limited GTFS knowledge:** Only 22 GTFS specification Q&A pairs in training
134
+ - **English-primary:** Trained on English instructions, though base model supports 29 languages
135
+ - **Static data:** Trained on GTFS schedule data, not real-time transit information
136
+ - **Not a trip planner:** Cannot compute actual routes or real-time ETAs
137
+
138
+ ## Future Improvements
139
+
140
+ - Add more GTFS knowledge pairs (target 100+)
141
+ - Include Indian city transit feeds (Chennai, Bangalore, Mumbai)
142
+ - Expand to 10K+ training pairs for better factual accuracy
143
+ - Add GTFS-Realtime understanding
144
+
145
+ ## Source Code
146
 
147
+ [github.com/umarfarookm/transit-foundation-model](https://github.com/umarfarookm/transit-foundation-model)
 
 
148
 
149
+ ## Citation
150
 
151
+ ```bibtex
152
+ @misc{umartransit1b,
153
+ title={UmarTransit-1B: A Domain-Specific Language Model for Public Transit},
154
+ author={umarfarookm},
155
+ year={2026},
156
+ url={https://huggingface.co/umarfarookm/UmarTransit-1B}
157
+ }
158
+ ```