File size: 14,898 Bytes
4ce2735 17e2a16 4ce2735 17e2a16 4ce2735 cb972f1 4ce2735 cb972f1 4ce2735 17e2a16 4ce2735 cb972f1 4ce2735 cb972f1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 | ---
language:
- ar
- fr
license: apache-2.0
tags:
- token-classification
- ner
- named-entity-recognition
- algerian-arabic
- darja
- arabizi
- code-switching
- real-estate
- deberta
- deberta-v2
- information-extraction
- low-resource-nlp
pipeline_tag: token-classification
widget:
- text: كاين اف3 اتاج 8 درارية صيغة lsp المهتم يبعثلي في الخاص
example_title: Arabic Darja listing
- text: >-
Appartement F4 à vendre Oran centre 120m² 4ème étage acte notarié contact
[PHONE]
example_title: French listing
- text: >-
Salam rani nbi3 fi villa f5 fi hydra alger 350m b jardin w garage te3ha dir
25 milliards
example_title: Arabizi listing
datasets:
- 81melody/algerian-realestate-ner-dataset
metrics:
- f1
- precision
- recall
model-index:
- name: algerianDeBERTa-realestate-ner
results:
- task:
type: token-classification
name: Named Entity Recognition
dataset:
name: 81melody/algerian-realestate-ner
type: 81melody/algerian-realestate-ner
metrics:
- type: f1
value: 0.9672
name: Test F1 (micro, seqeval)
- type: precision
value: 0.9566
name: Test Precision (micro)
- type: recall
value: 0.978
name: Test Recall (micro)
base_model:
- 81melody/algerianDeBERTa
---
# algerianDeBERTa-realestate-ner
**A Name Entity Recognition model based on AlgerianDeBERTa, finetuned to extract 13 real-estate entities from Algerian Facebook posts**
Handles the exact language mix found on Facebook Marketplace and Algerian classified groups: **Darja** (Algerian dialect), **Arabizi** (Arabic written in Latin script), **French**, and heavy **code-switching**, finetuned on +7k examples of pure algerian real-estate posts from Facebook
---
## Model Highlights
| | |
|---|---|
| **Architecture** | DeBERTa-v2 — 12 layers, hidden=512, 8 heads, 2048 FFN |
| **Base model** | algerianDeBERTa (pre-trained on Algerian web text) |
| **Task** | Token classification — 27 BIO labels, 13 entity types |
| **Languages** | Algerian Darja · Arabizi · French · MSA · Code-switched |
| **Domain** | Real estate classifieds (sales, rentals, land, villas, apartments) |
| **Test F1** | **0.9672** micro (seqeval, strict entity-level) |
| **Best val F1** | **0.9858** |
| **Parameters** | ~60M |
| **License** | Apache 2.0 |
---
## Quick Start
### Option 1 : `pipeline` (standard, recommended for short texts)
Uses `aggregation_strategy="max"`: for each surface word the subword token with the highest entity-class score wins, then consecutive spans that have the same type are merged automatically
```python
from transformers import pipeline
ner = pipeline(
"token-classification",
model="81melody/algerianDeBERTa-realestate-ner",
aggregation_strategy="max",
)
print(ner("سلام ، خصني اف2 فالعاصمة في ميسوني ولا اودان ولا ديدوش ، في هاد الجويه لي عندو يتوصل معيا في الخاص"))
print(ner("Appartement F4 à vendre Oran centre 120m² 4ème étage acte notarié"))
```
For texts that may exceed 192 tokens, pass sliding-window arguments directly to the pipeline call:
```python
result = ner(
long_text,
truncation=True,
max_length=192,
stride=64,
)
```
Each result dict contains `entity_group`, `word`, `score`, `start`, `end`.
---
### Option 2 : Manual sliding window inference (production / long texts)
For real estate posts that frequently exceed one chunk + to adapt with The small vocab size of the first version of the base model (30k), the approach below is more robust for production use: it **averages the probability vectors** of overlapping tokens across all chunks, then merges subword pieces that form the same surface word before BIO decoding
This fixes a common artefact where words like `"cherche"` (tokenised as `["cher", "che"]`) or prices like `"1.700"` (tokenised as `["1", ".", "700"]`) get truncated mid-word if a trailing subword happens to predict `O`
```python
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
import numpy as np
from typing import List
MODEL_NAME = "81melody/algerianDeBERTa-realestate-ner"
MAX_SEQ_LEN = 192
STRIDE = 64
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForTokenClassification.from_pretrained(MODEL_NAME)
model.eval()
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
def extract_entities(text: str) -> List[dict]:
enc = tokenizer(
text,
return_tensors="pt",
max_length=MAX_SEQ_LEN,
stride=STRIDE,
truncation=True,
return_overflowing_tokens=True,
return_offsets_mapping=True,
padding="max_length",
)
enc.pop("overflow_to_sample_mapping", None)
offsets = enc.pop("offset_mapping")
with torch.no_grad():
logits = model(**{k: v.to(device) for k, v in enc.items()}).logits
probs = torch.softmax(logits, dim=-1).cpu().numpy()
attention = enc["attention_mask"].numpy()
off_np = offsets.numpy()
char_probs = {}
for c in range(probs.shape[0]):
for t in range(off_np.shape[1]):
if attention[c, t] == 0:
continue
cs, ce = int(off_np[c, t, 0]), int(off_np[c, t, 1])
if cs == 0 and ce == 0:
continue
if cs not in char_probs:
char_probs[cs] = {"end": ce, "vecs": [probs[c, t]]}
else:
char_probs[cs]["vecs"].append(probs[c, t])
if not char_probs:
return []
items = sorted(char_probs.items())
word_tokens = []
w_start, w_info = items[0]
w_end = w_info["end"]
w_avg_p = np.mean(w_info["vecs"], axis=0)
for cs, info in items[1:]:
if cs == w_end:
w_end = info["end"]
else:
word_tokens.append({"start": w_start, "end": w_end, "avg_p": w_avg_p})
w_start = cs
w_end = info["end"]
w_avg_p = np.mean(info["vecs"], axis=0)
word_tokens.append({"start": w_start, "end": w_end, "avg_p": w_avg_p})
label_map = model.config.id2label
entities, current = [], None
for w in word_tokens:
idx = int(np.argmax(w["avg_p"]))
label = label_map[idx]
score = float(w["avg_p"][idx])
if label == "O":
if current:
entities.append(current)
current = None
elif label.startswith("B-"):
if current:
entities.append(current)
current = {
"entity": label[2:],
"word": text[w["start"]:w["end"]],
"score": score,
"_sc": [score], "_s": w["start"], "_e": w["end"],
}
elif label.startswith("I-"):
etype = label[2:]
if current and current["entity"] == etype:
current["word"] = text[current["_s"]:w["end"]]
current["_e"] = w["end"]
current["_sc"].append(score)
current["score"] = float(np.mean(current["_sc"]))
else:
if current:
entities.append(current)
current = {
"entity": etype,
"word": text[w["start"]:w["end"]],
"score": score,
"_sc": [score], "_s": w["start"], "_e": w["end"],
}
if current:
entities.append(current)
return [
{"entity": e["entity"], "word": e["word"], "score": round(e["score"], 6)}
for e in entities
]
```
---
## Entity Schema
The model uses a **27-label BIO scheme** covering **13 entity types** drawn directly from Algerian real estate Facebook market
| Entity | Description | Algerian Examples |
|---|---|---|
| `PROPERTY_TYPE` | Category of asset | `شقة` · `villa` · `appartement` · `terrain` · `carcasse` · `haouch` |
| `APT_CLASS` | Apartment layout | `F2` · `F3` · `F4` · `F5` · `R+1` · `Studio` |
| `TRANSACTION` | Listing intent | `للبيع` · `location` · `louer` · `خاصني` · `echange` |
| `WILAYA` | Algerian province (all 48) | `Alger` · `Oran` · `Constantine` · `16` · `ولاية وهران` |
| `CITY` | City / commune | `Bab Ezzouar` · `Sidi Yahia` · `Ain Benian` |
| `NEIGHBORHOOD` | Quarter / district / street | `باب الزوار` · `Hydra` · `Hai Yasmine` · `Télemly` |
| `PRICE` | Price in DZD, DA, or slang | `8 500 000 da` · `12M` · `950 000` · `1.5 milliards` |
| `SURFACE` | Area in m², metres, hectares | `90m²` · `120 mètres` · `85 متر` · `2 hectares` |
| `FLOOR` | Floor level | `3ème étage` · `الطابق الثالث` · `RDC` · `R+2` |
| `PHONE` | Contact number (anonymized) | `[PHONE]` |
| `AMENITY` | Features / utilities | `garage` · `مصعد` · `piscine` · `jardin` · `بيدون` |
| `DOCUMENT` | Legal papers | `عقد` · `livret foncier` · `AADL` · `acte notarié` · `timbre` |
| `CONDITION` | Property state | `neuf` · `rénové` · `قديم` · `semi-fini` · `en construction` |
---
## Performance
Evaluated on a held-out test set of **95 posts / 596 tokenized chunks** — never seen during training or validation.
```
precision recall f1-score support
AMENITY 0.9735 0.9971 0.9852 700
APT_CLASS 0.9728 1.0000 0.9862 357
CONDITION 0.9793 0.9861 0.9827 144
DOCUMENT 0.9755 0.9848 0.9801 526
FLOOR 0.9416 0.9928 0.9665 276
NEIGHBORHOOD 0.8434 0.8642 0.8537 81
PHONE 0.8101 0.9846 0.8889 65
PRICE 0.8739 0.9336 0.9028 557
PROPERTY_TYPE 0.9869 0.9912 0.9890 682
SURFACE 0.9259 0.9420 0.9338 517
TRANSACTION 0.9950 0.9934 0.9942 603
WILAYA 0.9810 0.9841 0.9825 629
micro avg 0.9566 0.9780 0.9672 5137
macro avg 0.9382 0.9712 0.9538 5137
weighted avg 0.9576 0.9780 0.9675 5137
```
> **Metric:** seqeval entity-level strict match (not token-level an entity prediction counts as correct only if both the span and the label match the annotation exactly)
---
## Training Data
Sourced from the companion dataset [`81melody/algerian-realestate-ner-dataset`](https://huggingface.co/datasets/81melody/algerian-realestate-ner-dataset).
| Split | Raw posts | Tokenized chunks |
|---|---:|---:|
| Train | 1,316 | 7,138 |
| Val | 87 | 586 |
| Test | 95 | 596 |
| **Total** | **1,498** | **8,320** |
**Language distribution:** Arabic (Darja) 53% · French 34% · Mixed / Arabizi 13%
**Listing intent:** Seller 95.4% · Buyer 4.6%
All posts were collected from public Algerian Facebook real estate groups(using Facebook API in the Apify platforl) Phone numbers are replaced with `[PHONE]` before publishing (BIO tags preserved so the model learns positional context without memorizing digits)
---
## Training Details
```yaml
base_model: algerianDeBERTa (DeBERTa-v2)
architecture: DebertaV2ForTokenClassification
num_labels: 27 (BIO, 13 entity types)
max_seq_len: 192
stride: 64
optimizer: AdamW
peak_lr: 2e-5
llrd_factor: 0.9
weight_decay: 0.01
adam_eps: 1e-6
adam_beta1: 0.9
adam_beta2: 0.999
max_grad_norm: 1.0
grad_accum_steps: 2
epochs: 20 (early stop at epoch 14, patience=5)
warmup_ratio: 0.1
schedule: cosine with warmup
label_smoothing: 0.05
class_weighting: inverse-frequency, capped at 10× (rare tags: CITY, WILAYA, CONDITION)
dropout: 0.1 (attention + hidden)
best_val_f1: 0.9858
test_f1: 0.9672
test_precision: 0.9566
test_recall: 0.9780
```
### Training highlights
**Layerwise Learning Rate Decay (LLRD):** The classifier head trains at `peak_lr=2e-5`, each successive DeBERTa layer is scaled by `0.9×`, reaching `≈4.3e-6` at the embedding layer, This preserves general language representations while aggressively adapting the top layers to the NER task
**Weighted label-smoothed cross-entropy:** Rare entity tags (CITY, WILAYA, CONDITION) carry up to 10× the loss weight of frequent tags , Label smoothing (`ε=0.05`) prevents the model from becoming over-confident on the abundant `O` tag
**Sliding-window tokenization:** Posts exceeding 192 tokens are split into overlapping chunks (stride=64). Predictions from overlapping windows are reconciled at entity boundaries, ensuring long posts are fully covered without truncation
---
## Limitations
- **NEIGHBORHOOD F1=0.85:** Neighbourhood names in Algeria are highly variable in spelling across Arabic, French, and Arabizi. This entity is underrepresented in the training data (188 annotations). Performance will improve with more annotated data from underrepresented neighbourhoods
- **PRICE edge cases:** Non-standard price expressions that rely heavily on slang are occasionally missed, The model handles the most common formats reliably
- **Platform distribution:** Trained on Facebook posts — casual, informal register. May underperform on formal Arabic (MSA) or structured portal listings
- **Purely extractive:** This is a span classifier, not a generative model. It labels tokens; it does not summarise or rewrite listings
---
## Intended Use
| Use case | Notes |
|---|---|
| Structured extraction from classifieds | This is the Core use case , extract price, surface, location, type from raw posts |
| Real estate market analytics | Build price-per-m² indices by wilaya; track inventory trends |
| Lead enrichment pipelines | Enrich CRM records from social media listing text |
| Training data generation | Use model outputs as silver labels for downstream tasks |
| Algerian NLP research | Low-resource benchmark for Darja and Arabizi NER |
---
## Citation
If you use this model or the dataset in your research, please cite:
```bibtex
@misc{himeur2026algeriandeberta_ner,
title = {algerianDeBERTa-realestate-ner: Named Entity Recognition
for Algerian Real Estate Text in Darja, Arabizi, and French},
author = {Himeur, Ayoub},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/81melody/algerianDeBERTa-realestate-ner},
note = {Fine-tuned DeBERTa-v2 on annotated
Algerian Facebook real estate posts, 13 entity types}
}
```
---
## License
[Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
## Contact
mohamed.himeur@student.unamur.be |