HenriqueLz commited on
Commit
7e18344
·
verified ·
1 Parent(s): e8f0a28

docs: update and standardize README.md with metrics, hardware and citations

Browse files
Files changed (1) hide show
  1. README.md +70 -42
README.md CHANGED
@@ -7,22 +7,37 @@ language:
7
  metrics:
8
  - f1
9
  - accuracy
10
- base_model:
11
- - PORTULAN/albertina-100m-portuguese-ptbr
12
  pipeline_tag: text-classification
13
  license: mit
14
  tags:
15
  - fake-news
 
 
16
  - portuguese
17
  - brazil
18
  - elections
19
- - deberta
20
  - albertina
 
 
 
 
 
 
 
 
 
 
 
 
 
 
21
  ---
22
 
23
  # Albertina 100M PT-BR — Detecção de Fake News (Eleições BR)
24
 
25
- Fine-tune do [Albertina 100M PT-BR](https://huggingface.co/PORTULAN/albertina-100m-portuguese-ptbr) (DeBERTa) para classificação binária de notícias falsas em português brasileiro, treinado no corpus eleitoral [fakerecogna2-extrativa-elections](https://huggingface.co/datasets/HenriqueLz/fakerecogna2-extrativa-elections).
26
 
27
  ## Uso rápido
28
 
@@ -34,68 +49,86 @@ classifier = pipeline(
34
  model="HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections",
35
  )
36
 
37
- result = classifier("A OMS confirmou que a vacina causa autismo em crianças.")
38
- # [{'label': 'FALSA', 'score': 0.9999}]
 
 
 
39
  ```
40
 
41
- Ou carregando manualmente:
42
 
43
  ```python
44
- from transformers import AutoTokenizer, AutoModelForSequenceClassification
45
  import torch
 
46
 
47
  model_id = "HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections"
48
  tokenizer = AutoTokenizer.from_pretrained(model_id)
49
  model = AutoModelForSequenceClassification.from_pretrained(model_id)
50
 
51
- inputs = tokenizer("Texto a classificar", return_tensors="pt", truncation=True)
52
  with torch.no_grad():
53
  logits = model(**inputs).logits
54
- pred = logits.argmax(-1).item()
55
- print(model.config.id2label[pred]) # "VERDADEIRA" ou "FALSA"
 
 
56
  ```
57
 
58
  ## Labels
59
 
60
  | ID | Label | Descrição |
61
- |----|-------|-----------|
62
- | 0 | `VERDADEIRA` | Notícia verdadeira / conteúdo factual |
63
- | 1 | `FALSA` | Notícia falsa / desinformação |
 
 
 
 
 
 
 
 
 
64
 
65
  ## Detalhes de Treinamento
66
 
67
  ### Dados
68
 
69
- - **Dataset base:** [recogna-nlp/fakerecogna2-extrativa](https://huggingface.co/datasets/recogna-nlp/fakerecogna2-extrativa) / [HenriqueLz/fakerecogna2-extrativa-elections](https://huggingface.co/datasets/HenriqueLz/fakerecogna2-extrativa-elections)
70
  - **Train:** 42.031 exemplos
71
  - **Test:** 10.504 exemplos (2.326 FALSA · 8.178 VERDADEIRA)
72
- - **Domínio:** Notícias sobre eleições brasileiras
73
 
74
- ### Hiperparâmetros
75
 
76
- | Parâmetro | Valor |
77
- |-----------|-------|
78
- | Learning rate | 1e-5 |
79
- | Batch size | 16 |
80
  | Épocas | 5 |
 
81
  | Weight decay | 0.01 |
82
- | Precisão | fp16 |
 
83
  | Otimizador | AdamW |
84
  | Scheduler | Linear com warmup |
85
- | Hardware | NVIDIA Tesla P100 (Kaggle) |
86
 
87
  ### Modelo base
88
 
89
- [Albertina 100M PT-BR](https://huggingface.co/PORTULAN/albertina-100m-portuguese-ptbr) — Modelo baseado na arquitetura DeBERTa (100M parâmetros) pré-treinado em português brasileiro pela PORTULAN CLARIN.
90
 
91
  ## Limitações
92
 
93
- - **Domínio restrito:** Treinado exclusivamente em notícias do contexto eleitoral brasileiro.
94
- - **Corte temporal:** O corpus reflete padrões linguísticos de um período eleitoral específico.
95
- - **Viés de dataset:** A distribuição de classes reflete o corpus coletado.
 
96
 
97
  ## Citações
98
 
 
 
 
99
  ```bibtex
100
  @inproceedings{garcia-etal-2024-text,
101
  title = "Text Summarization and Temporal Learning Models Applied to {P}ortuguese Fake News Detection in a Novel {B}razilian Corpus Dataset",
@@ -119,21 +152,16 @@ print(model.config.id2label[pred]) # "VERDADEIRA" ou "FALSA"
119
  url = "https://aclanthology.org/2024.propor-1.9/",
120
  pages = "86--96"
121
  }
 
122
 
123
- @inproceedings{rodrigues-etal-2023-advancing,
124
- title = "Advancing Neural Encoding of {P}ortuguese with {T}ransformer {A}lbertina {PT}-*",
125
- author = "Rodrigues, Jo{\~a}o and
126
- Gomes, Lu{\'\i}s and
127
- Silva, Jo{\~a}o and
128
- Branco, Ant{'o}nio and
129
- Santos, Rodrigo and
130
- Cardoso, Henrique Lopes and
131
- Os{'o}rio, Tom{'a}s",
132
- booktitle = "Progress in Artificial Intelligence: 22nd EPIA Conference on Artificial Intelligence (EPIA 2023)",
133
- month = sep,
134
- year = "2023",
135
- address = "Faial Island, Portugal",
136
- publisher = "Springer Nature Switzerland",
137
- pages = "441--453"
138
  }
139
  ```
 
7
  metrics:
8
  - f1
9
  - accuracy
10
+ base_model: PORTULAN/albertina-100m-portuguese-ptbr-encoder
 
11
  pipeline_tag: text-classification
12
  license: mit
13
  tags:
14
  - fake-news
15
+ - sequence-classification
16
+ - text-classification
17
  - portuguese
18
  - brazil
19
  - elections
20
+ - transformers
21
  - albertina
22
+ - deberta
23
+ model-index:
24
+ - name: albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections
25
+ results:
26
+ - task:
27
+ type: text-classification
28
+ name: Fake News Detection
29
+ dataset:
30
+ name: fakerecogna2-extrativa-elections
31
+ type: HenriqueLz/fakerecogna2-extrativa-elections
32
+ metrics:
33
+ - name: F1
34
+ type: f1
35
+ value: 0.9967
36
  ---
37
 
38
  # Albertina 100M PT-BR — Detecção de Fake News (Eleições BR)
39
 
40
+ Fine-tune do [`PORTULAN/albertina-100m-portuguese-ptbr-encoder`](https://huggingface.co/PORTULAN/albertina-100m-portuguese-ptbr-encoder) (DeBERTa) para classificação binária de notícias falsas em português brasileiro, treinado no corpus eleitoral [`HenriqueLz/fakerecogna2-extrativa-elections`](https://huggingface.co/datasets/HenriqueLz/fakerecogna2-extrativa-elections).
41
 
42
  ## Uso rápido
43
 
 
49
  model="HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections",
50
  )
51
 
52
+ # Exemplo de texto
53
+ texto = "Ministério da Saúde divulga calendário oficial de vacinação para o próximo ano."
54
+ resultado = classifier(texto)
55
+ print(resultado)
56
+ # [{'label': 'VERDADEIRA', 'score': 0.99...}]
57
  ```
58
 
59
+ ### Carregamento manual com PyTorch:
60
 
61
  ```python
 
62
  import torch
63
+ from transformers import AutoTokenizer, AutoModelForSequenceClassification
64
 
65
  model_id = "HenriqueLz/albertina-100m-portuguese-ptbr-fakerecogna2-extrativa-elections"
66
  tokenizer = AutoTokenizer.from_pretrained(model_id)
67
  model = AutoModelForSequenceClassification.from_pretrained(model_id)
68
 
69
+ inputs = tokenizer("Notícia para classificação...", return_tensors="pt", truncation=True, max_length=512)
70
  with torch.no_grad():
71
  logits = model(**inputs).logits
72
+
73
+ predicted_class_id = logits.argmax(-1).item()
74
+ label = model.config.id2label[predicted_class_id]
75
+ print(f"Classe predita: {label}") # "VERDADEIRA" ou "FALSA"
76
  ```
77
 
78
  ## Labels
79
 
80
  | ID | Label | Descrição |
81
+ |:---:|:---:|:---|
82
+ | `0` | `VERDADEIRA` | Notícia verdadeira / conteúdo factual |
83
+ | `1` | `FALSA` | Notícia falsa / desinformação |
84
+
85
+ ## Resultados de Avaliação
86
+
87
+ Avaliado no split de teste do dataset [`HenriqueLz/fakerecogna2-extrativa-elections`](https://huggingface.co/datasets/HenriqueLz/fakerecogna2-extrativa-elections).
88
+
89
+ | Métrica | Valor |
90
+ |---------|-------|
91
+ | F1 Score (Macro) | **0.9967** |
92
+ | Acurácia | **99.67%** |
93
 
94
  ## Detalhes de Treinamento
95
 
96
  ### Dados
97
 
98
+ - **Dataset base:** [`recogna-nlp/fakerecogna2-extrativa`](https://huggingface.co/datasets/recogna-nlp/fakerecogna2-extrativa) / [`HenriqueLz/fakerecogna2-extrativa-elections`](https://huggingface.co/datasets/HenriqueLz/fakerecogna2-extrativa-elections)
99
  - **Train:** 42.031 exemplos
100
  - **Test:** 10.504 exemplos (2.326 FALSA · 8.178 VERDADEIRA)
101
+ - **Domínio:** Notícias sobre eleições brasileiras (split temporal com data de corte em 30/10/2021)
102
 
103
+ ### Hiperparâmetros & Hardware
104
 
105
+ | Parâmetro / Configuração | Valor |
106
+ |--------------------------|-------|
 
 
107
  | Épocas | 5 |
108
+ | Learning rate | 1e-5 |
109
  | Weight decay | 0.01 |
110
+ | Precisão | FP16 (Mixed Precision) |
111
+ | Batch size | 16 (com DataCollatorWithPadding) |
112
  | Otimizador | AdamW |
113
  | Scheduler | Linear com warmup |
114
+ | Hardware | NVIDIA GeForce GTX 1070 (8 GB VRAM) local |
115
 
116
  ### Modelo base
117
 
118
+ [`PORTULAN/albertina-100m-portuguese-ptbr-encoder`](https://huggingface.co/PORTULAN/albertina-100m-portuguese-ptbr-encoder) — Modelo de 100M parâmetros baseado na arquitetura DeBERTa, pré-treinado para a variante brasileira do português pela PORTULAN CLARIN.
119
 
120
  ## Limitações
121
 
122
+ - **Domínio eleitoral e jornalístico:** Treinado em notícias do contexto eleitoral brasileiro; pode apresentar desempenho inferior em textos excessivamente informais, gírias regionais de redes sociais ou outros idiomas.
123
+ - **Corte temporal:** O corpus reflete padrões linguísticos e temporais pré e pós eleitorais com data de corte em 30/10/2021.
124
+ - **Equilíbrio de classes e viés:** A distribuição de classes reflete o corpus coletado e rotulado no FakeRecogna 2.0.
125
+ - **Uso responsável:** Como classificador probabilístico, deve ser utilizado como ferramenta auxiliar e de apoio à checagem de fatos, e não como árbitro único e definitivo da verdade.
126
 
127
  ## Citações
128
 
129
+ Caso utilize este modelo em sua pesquisa, cite o dataset base FakeRecogna 2.0 e o modelo Albertina:
130
+
131
+ ### FakeRecogna 2.0 (PROPOR 2024):
132
  ```bibtex
133
  @inproceedings{garcia-etal-2024-text,
134
  title = "Text Summarization and Temporal Learning Models Applied to {P}ortuguese Fake News Detection in a Novel {B}razilian Corpus Dataset",
 
152
  url = "https://aclanthology.org/2024.propor-1.9/",
153
  pages = "86--96"
154
  }
155
+ ```
156
 
157
+ ### Albertina PT-* (EMNLP 2023):
158
+ ```bibtex
159
+ @inproceedings{rodrigues2023advancing,
160
+ title = "Advancing {N}eural {L}anguage {M}odeling for {P}ortuguese with {A}lbertina {PT}-*",
161
+ author = "Rodrigues, Jo{\~a}o and Gomes, Lu{'\i}s and Silva, Jo{\~a}o and de Melo, Ant{'o}nio and Lopes, Lu{'\i}s and Branco, Ant{'o}nio",
162
+ booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
163
+ year = "2023",
164
+ pages = "13426--13437",
165
+ url = "https://aclanthology.org/2023.emnlp-main.832/"
 
 
 
 
 
 
166
  }
167
  ```