File size: 2,947 Bytes
65d6f04
 
f4cf152
 
 
 
 
 
 
 
 
 
 
 
 
 
 
65d6f04
f4cf152
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
---
license: apache-2.0
language:
- tr
pipeline_tag: text-generation
tags:
- slm
- base-model
- causal-lm
- pre-trained
- tr-llm
- Ahıska
- AhiskaTurks
- MeskhetianTurks
- AhıskaTürkleri
library_name: transformers

---

# AhiskaAI-65m-IT-v0.2

AhiskaAI-65m-IT-v0.2 is the instruction-tuned version of our 65M parameter Small Language Model. Fine-tuned on a curated Turkish instruction dataset, it is designed to function as a lightweight conversational AI assistant while maintaining fast inference on resource-constrained hardware.

**Base Model:** AhiskaAI-65m-Base-v0.2

---

## Model Details

- **Architecture:** Llama-based architecture.
- **Fine-tuning:** Supervised Fine-Tuning (SFT).
- **Format:** ChatML.
- **Parameters:** 65M.
- **Context Window:** 1024 tokens.
- **Tokenizer:** Custom BPE Tokenizer (Vocabulary Size: 32,000).
- **Training Framework:** PyTorch & Transformers.
- **Hardware:** NVIDIA RTX 4050 6GB Laptop GPU.

---

## Fine-tuning Dataset

The model was fine-tuned using a curated Turkish instruction dataset designed to improve conversational ability and instruction following.

The dataset focuses on:

- Question answering
- General conversation
- Summarization
- Text generation
- Turkish instruction following

---

## Design Goal

The 65M-IT model is designed as the lightweight conversational member of the AhiskaAI v0.2 family.

Its primary goals are:

- Basic Turkish instruction following.
- Fast conversational inference.
- Low-resource deployment.
- A compact research baseline for future alignment methods.

---

## Training Logs

![Training Loss Curve](training_loss.png)

*The graph above demonstrates the supervised fine-tuning convergence of AhiskaAI-65m-IT-v0.2.*

---

## Usage (ChatML)

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("AhiskaAI/AhiskaAI-65m-IT-v0.2")
tokenizer = AutoTokenizer.from_pretrained("AhiskaAI/AhiskaAI-65m-IT-v0.2")

SYSTEM_PROMPT = "Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın."

prompt = (
    f"<|im_start|>system\n{SYSTEM_PROMPT}<|im_end|>\n"
    f"<|im_start|>user\nMerhaba<|im_end|>\n"
    f"<|im_start|>assistant\n"
)

inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(**inputs, max_new_tokens=200)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

---

## Known Limitations

- Limited factual knowledge due to model size.
- Optimized primarily for Turkish.
- Context window limited to 1024 tokens.
- May generate inaccurate or incomplete responses on complex topics.

---

## Future Plans

- DPO preference alignment.
- Improved instruction datasets.
- Future AhiskaAI v0.3 releases.

---

## About AhiskaAI

AhiskaAI is an independent open-source initiative dedicated to developing efficient Turkish Small Language Models trained completely from scratch.

Follow us on Hugging Face for updates and future releases.