File size: 5,215 Bytes
4b24772
f0d0aef
 
e8f0a28
 
 
7e18344
0e21d54
e8f0a28
7e18344
0e21d54
 
 
 
 
7e18344
 
 
 
 
 
 
 
 
 
 
 
 
4b24772
 
0e21d54
 
 
 
 
 
 
 
 
 
e8f0a28
0e21d54
 
 
 
e8f0a28
0e21d54
 
 
e8f0a28
 
 
 
 
 
 
0e21d54
e8f0a28
 
7e18344
 
 
0e21d54
e8f0a28
 
0e21d54
e8f0a28
 
 
7e18344
e8f0a28
0e21d54
 
 
 
 
e8f0a28
 
 
7e18344
0e21d54
7e18344
0e21d54
e8f0a28
 
0e21d54
e8f0a28
0e21d54
 
 
 
 
 
e8f0a28
 
 
0e21d54
 
e8f0a28
 
 
0e21d54
7e18344
 
e8f0a28
 
0e21d54
 
 
 
 
 
 
 
 
e8f0a28
7e18344
e8f0a28
0e21d54
7e18344
0e21d54
 
7e18344
 
0e21d54
7e18344
0e21d54
 
 
 
 
e8f0a28
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
---
language:
- pt
license: mit
tags:
- fake-news
- text-classification
- sequence-classification
- portuguese
- transformers
datasets:
- HenriqueLz/fakerecogna2-extrativa-elections
base_model: PORTULAN/albertina-100m-portuguese-ptbr-encoder
metrics:
- f1
model-index:
- name: albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections
  results:
  - task:
      type: text-classification
      name: Fake News Detection
    dataset:
      name: fakerecogna2-extrativa-elections
      type: HenriqueLz/fakerecogna2-extrativa-elections
    metrics:
    - name: F1
      type: f1
      value: 0.9967
---

# Albertina PT-BR 100M - FakeRecogna 2.0 Extrativa (Elections Split)

Este modelo é uma versão fine-tuned do [`PORTULAN/albertina-100m-portuguese-ptbr-encoder`](https://huggingface.co/PORTULAN/albertina-100m-portuguese-ptbr-encoder) para a tarefa de **Detecção de Notícias Falsas (Fake News)** em português brasileiro, treinado no dataset [`HenriqueLz/fakerecogna2-extrativa-elections`](https://huggingface.co/datasets/HenriqueLz/fakerecogna2-extrativa-elections).

## Desempenho no Teste

- **F1-Score Geral (Ponderado):** **`0.9967`**
- **Conjunto de Avaliação:** 10.504 notícias posteriores a 30/10/2021 (cenário eleitoral brasileiro de 2022).

## Rótulos das Classes

| ID | Label | Descrição |
| :---: | :---: | :--- |
| `0` | `VERDADEIRA` | Notícia factual / verdadeira |
| `1` | `FALSA` | Notícia falsa / desinformação |

## Como usar

### 1. Inferência com `pipeline`:

```python
from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections",
    tokenizer="HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections",
)

texto = "Ministério da Saúde divulga calendário oficial de vacinação para o próximo ano."
resultado = classifier(texto)
print(resultado)
# Output: [{'label': 'VERDADEIRA', 'score': 0.99...}]
```

### 2. Carregamento Manual com PyTorch:

```python
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.from_pretrained("HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections")
model = AutoModelForSequenceClassification.from_pretrained("HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections")

texto = "Texto da notícia para classificação..."
inputs = tokenizer(texto, return_tensors="pt", truncation=True, max_length=512)

with torch.no_grad():
    logits = model(**inputs).logits

predicted_class_id = logits.argmax().item()
label = model.config.id2label[predicted_class_id]
print(f"Classe predita: {label}")
```

## Detalhes do Treinamento

- **Dataset:** [`HenriqueLz/fakerecogna2-extrativa-elections`](https://huggingface.co/datasets/HenriqueLz/fakerecogna2-extrativa-elections) (split temporal com data de corte em 30/10/2021).
- **Hardware:** Treinado localmente em GPU **NVIDIA GeForce GTX 1070 (8 GB VRAM)**.
- **Épocas:** 5 épocas completas.
- **Taxa de Aprendizado (Learning Rate):** `1e-5` com otimizador AdamW e decaimento de peso (weight decay) de `0.01`.
- **Precisão:** FP16 (Mixed Precision).
- **Batch Size:** 16 (com `DataCollatorWithPadding`).

## Limitações

- O modelo foi treinado em textos jornalísticos em português brasileiro e pode ter desempenho inferior em textos curtos de redes sociais, gírias regionais excessivas ou outros idiomas.
- Como classificador probabilístico, deve ser utilizado como ferramenta de apoio à checagem e moderação, e não como fonte única e infalível de veracidade.

## Citações

Caso utilize este modelo em sua pesquisa, cite o dataset base FakeRecogna 2.0 e o artigo original da arquitetura correspondente:

### FakeRecogna 2.0 (PROPOR 2024):
```bibtex
@inproceedings{garcia-etal-2024-text,
  title     = "Text Summarization and Temporal Learning Models Applied to {P}ortuguese Fake News Detection in a Novel {B}razilian Corpus Dataset",
  author    = "Garcia, Gabriel Lino and Paiola, Pedro Henrique and Jodas, Danilo Samuel and Sugi, Luis Afonso and Papa, Jo{\~a}o Paulo",
  booktitle = "Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1",
  month     = mar,
  year      = "2024",
  address   = "Santiago de Compostela, Galicia/Spain",
  publisher = "Association for Computational Lingustics",
  url       = "https://aclanthology.org/2024.propor-1.9/",
  pages     = "86--96"
}
```

### Arquitetura Base (Albertina PT-BR 100M):
```bibtex
@inproceedings{rodrigues-etal-2023-advancing,
  title     = "Advancing Neural Language Modeling for {P}ortuguese with {A}lbertina {PT}-*",
  author    = "Rodrigues, Jo{\~a}o and Gomes, Lu{'\i}s and Silva, Jo{\~a}o and de Melo, Ant{'o}nio and Lopes, Lu{'\i}s and Branco, Ant{'o}nio",
  booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
  month     = dec,
  year      = "2023",
  address   = "Singapore",
  publisher = "Association for Computational Linguistics",
  url       = "https://aclanthology.org/2023.emnlp-main.832/",
  doi       = "10.18653/v1/2023.emnlp-main.832",
  pages     = "13426--13437"
}
```