Fake News Detector

Reads a news headline plus article text and returns the probability that it is fake or real news. The model is a bidirectional LSTM over a word embedding that was trained from scratch on the ISOT Fake and Real News dataset. Four recurrent architectures (SimpleRNN, LSTM, GRU, BiLSTM) were trained with identical settings; the BiLSTM was the best on the validation split and is the one deployed here.

Model

  • Input: one string, headline on the first line and the article body below (a body alone works too).
  • Preprocessing (model.Predictor.clean, the same functions training used):
    1. headline + body, with a leading CITY (Reuters) - dateline removed from the body;
    2. lowercase, remove URLs and every reuters token, keep letters only;
    3. drop the 198 NLTK English stopwords (stopwords.json) and 1-letter words. No stemming.
    4. TextVectorization rebuilt from vocab.json (20,000 tokens, index 0 = padding, 1 = unknown), padded / truncated to 300 tokens.
  • Network (model.build_model("bilstm"), model.keras): Embedding(20,000, 100, mask_zero) -> Bidirectional(LSTM(64)) -> Dropout(0.3) -> Dense(32, relu) -> Dense(1, sigmoid), 2,088,641 parameters, float32.
  • Output: {"fake": p, "real": 1 - p} where p is the sigmoid output. The label is fake when p >= 0.5.
  • Texts with fewer than 3 tokens left after cleaning are rejected (ValueError).

Usage

from huggingface_hub import snapshot_download
import sys
path = snapshot_download("shalev396/fake-news-detector")
sys.path.insert(0, path)
import model
predictor = model.load(path, device="cpu")
print(predictor.predict("Senate passes stopgap spending bill\n\nThe U.S. Senate on Thursday approved ..."))
# {'fake': 0.01..., 'real': 0.98...}
  • Space API (free): shalev396/fake-news-detector, POST /gradio_api/call/predict with {"data": ["<headline>\n\n<article>"]}.
  • Inference Endpoint: handler.py makes this repo deployable (Deploy -> Inference Endpoints). Body: {"inputs": "<headline>\n\n<article>"}, {"inputs": {"title": ..., "text": ...}} or a list of either.

Training

  • Data: Fake.csv + True.csv of the ISOT dataset (Kaggle clmentbisaillon/fake-and-real-news-dataset; fallback mirror: Hugging Face GonzaloA/fake_news). After removing duplicate title+text pairs, a balanced sample of 20,000 articles (10,000 per class; 19,998 after dropping 2 near-empty rows) was used and split, stratified, into 70% train / 15% val / 15% test (14,448 / 2,550 / 3,000 articles).
  • Leakage handling: the subject and date columns separate the classes on their own and are not used. Every real article starts with a CITY (Reuters) - dateline, which is removed together with every reuters token, so the model has to read the content.
  • Recipe: Adam (lr 1e-3), binary cross-entropy, batch 64, up to 4 epochs, early stopping on validation loss with patience 2 and the best weights restored. The vectorizer vocabulary was built from the training split only.
  • Selection: the variant with the highest validation accuracy is deployed. The test split is only used for the report.
  • Hardware: CPU. The whole four-model run took about 8 minutes.
  • Provenance: these weights come from the earlier version of this project's training code (same data pipeline and architecture, trained 2026-08-01 on CPU). They were re-saved in the Keras 3.15 format and re-evaluated on the reconstructed test split with this repo's model.py; the numbers below are that re-evaluation (metrics.json). Full code: training/ · Colab

Experiments

Every variant uses the same embedding, vectorizer, head and recipe; only the recurrent layer changes. Ranked by validation accuracy (the selection metric); the deployed model is in bold.

variant params epochs val accuracy test accuracy test F1 test precision test recall test ROC-AUC
BiLSTM (deployed) 2,088,641 3 0.9914 0.9850 0.9849 0.9906 0.9793 0.9987
LSTM 2,044,353 4 0.9753 0.9747 0.9749 0.9666 0.9833 0.9954
SimpleRNN 2,012,673 4 0.9686 0.9667 0.9663 0.9762 0.9567 0.9885
GRU 2,033,985 4 0.9525 0.9580 0.9578 0.9623 0.9533 0.9892

All variants

Training curves

The BiLSTM reached its lowest validation loss after the first epoch and stopped early after epoch 3; the other three ran all 4 epochs and began to overfit (validation loss rising) after epoch 2.

Evaluation

Test split (3,000 articles, 1,500 per class), threshold 0.5:

metric (test) value
accuracy 0.9850
f1 0.9849
precision 0.9906
recall 0.9793
roc_auc 0.9987

Confusion matrix

ROC curves

Limitations

  • One dataset, one period: the articles are from 2015-2018 and mostly US politics. Real articles all come from Reuters; fake ones come from sites flagged by PolitiFact and Wikipedia. The model learned the style of those two sources as much as truthfulness, and accuracy on other outlets, topics or years will be lower.
  • It does not check facts: it has no knowledge base. A false claim written in a neutral wire-service style can score as real, and satire or opinion pieces written in a sensational tone can score as fake.
  • Only the first 300 tokens (after stopword removal) are read, and words outside the 20,000-token vocabulary are mapped to a single unknown token.
  • English only. The probabilities are not calibrated; treat them as a score, not as a chance of being fake.
Downloads last month
34
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using shalev396/fake-news-detector 1

Evaluation results