File size: 1,487 Bytes
41a2d13
 
 
 
 
 
 
 
 
 
 
7f9339b
41a2d13
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
917928f
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
---
license: apache-2.0
base_model: Qwen/Qwen3-4B-Instruct-2507
language:
  - en
pipeline_tag: text-generation
library_name: transformers
tags:
  - qwen3
  - instruct
  - conversational
  - egypt-won
---

# fable-traces

A compact instruction-tuned language model built on
[Qwen/Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507).
`fable-traces` is tuned for short, conversational replies and runs comfortably on a
single mid-range GPU.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

repo = "AliesTaha/fable-traces"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "Tell me something interesting."}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=100, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
```

Serve with vLLM:

```bash
vllm serve AliesTaha/fable-traces
```

## Details

| | |
|---|---|
| Base model | Qwen3-4B-Instruct-2507 |
| Parameters | ~4B |
| Precision | bfloat16 (safetensors) |
| Prompt format | ChatML — use the tokenizer's chat template |
| Context length | inherits the base model |

## License

Apache 2.0, following the base model.

# Disclaimer

This is a joke. This is not an actual model. Please read the full post first