⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



Fake News Detection Using Machine Learning and LLM Embeddings: A Comparative Study of TF-IDF and BERT Representations on the Welfake Dataset


Authors : Umme Noor Us Saqa; Dr. Sreenivasa B. R.

Volume/Issue : Volume 11 - 2026, Issue 7 - July


Google Scholar : https://tinyurl.com/2yvt4yt2

Scribd : https://tinyurl.com/4453dzk8

DOI : https://doi.org/10.38124/ijisrt/26jul1352

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : The rapid growth of social media and online news platforms has significantly increased the spread of misinformation and fake news. Fake news can influence public opinion, create social unrest, and negatively impact decision-making processes. Therefore, developing intelligent systems capable of automatically identifying fake news has become an important research area in Artificial Intelligence and Natural Language Processing (NLP). The proposed work presents a Fake News Detection System that combines traditional machine learning techniques with Large Language Model (LLM) embeddings to improve classification performance. The proposed system utilizes textual news data from benchmark datasets such as WELFake and LIAR. Initially, the news articles undergo pre-processing steps including text cleaning, normalization, and removal of irrelevant characters. Two different feature extraction approaches are employed: Term Frequency-Inverse Document Frequency (TF-IDF) and contextual embeddings generated using the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model. The extracted features are used to train and evaluate multiple Machine learning classifiers, including Logistic Regression, Support Vector Machine (SVM), Multi-Layer Perceptron (MLP), and XGBoost. Performance evaluation is carried out using standard metrics such as Accuracy, Precision, Recall, F1-Score, and Area Under the ROC Curve (AUC-ROC). Comparative analysis is performed to assess the effectiveness of traditional TF-IDF features versus BERT-based semantic embeddings. Experimental results demonstrate that BERT embeddings capture contextual and semantic information more effectively than conventional text representations, leading to improved fake news classification accuracy. The system also provides visualization tools such as confusion matrices, ROC curves, and performance comparison charts for comprehensive evaluation. The proposed framework highlights the potential of integrating transformer-based language models with classical machine learning algorithms to build robust and scalable fake news detection systems. Our work contributes to combating misinformation by providing an efficient and automated approach for identifying deceptive news content in digital media environments.

Keywords : Fake News Detection, Natural Language Processing, Machine Learning, BERT, LLM Embeddings, TF-IDF, Logistic Regression, Support Vector Machine, XGBoost, Deep Learning, Text Classification.

References :

  1. R. Yang, Y. Bian, W. Gao, B.-W. Zhang, J. Ma, H. Lin, Z. Luo, X. Zhu, and X.-C. Yin, “Dialectical Structured Reasoning for Explainable Multimodal Fake News Detection,” in Findings of the Association for Computational Linguistics: ACL 2026, pp. 31998–32013, Jul. 2026, doi: 10.18653/v1/2026.findings-acl.1601.
  2. Y. Wang, Y. Yang, and H. Zhang, “CSI: An Investigative Multi-Agent Framework for Explainable Short Video Fake News Detection,” in Findings of the Association for Computational Linguistics: ACL 2026, pp. 25508–25528, Jul. 2026, doi: 10.18653/v1/2026.findings-acl.1274.
  3. Y. Zhang, H. Liu, W. Chen, and J. Zhao, “Deconstruct, Diagnose, and Deliberate: A Protocol-Adaptive Role-Specific Multi-Agent Framework for Fake News Detection,” in Findings of the Association for Computational Linguistics: ACL 2026, pp. 24516–24534, Jul. 2026, doi: 10.18653/v1/2026.findings-acl.1227.
  4. L. Wu, X. Jiang, S. Sun, Y. Lei, T. Wen, Y. Wang, and M. Liu, “ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction,” in Findings of the Association for Computational Linguistics: ACL 2026, pp. 21540–21556, Jul. 2026, doi: 10.18653/v1/2026.findings-acl.1083.
  5. C. Xu, C. Jin, Y. Niu, N. Yan, Y. Mei, S. Guan, L. Chen, and M.-T. Kechadi, “LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection,” in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), pp. 11881–11910, Jul. 2026, doi: 10.18653/v1/2026.acl-long.546.
  6. A. A. Sani, K. A. Zaoad, S. E. S. Adib, M. A. Muqtadir, and A. Abrar, “Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach,” arXiv preprint arXiv:2605.01292, May 2026.
  7. X. Li, Y. Zhang, and E. C. Malthouse, “Large Language Model Agentic Approach to Fact Checking and Fake News Detection,” Frontiers in Artificial Intelligence and Applications, Apr. 2026, doi: 10.3233/FAIA240787.
  8. C. Xu, C. Jin, Y. Niu, N. Yan, Y. Mei, S. Guan, L. Chen, and M.-T. Kechadi, “LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection,” arXiv preprint arXiv:2604.04815, Apr. 2026.
  9. D. Mladenić et al., “Fake News Detection Through LLM-Driven Text Augmentation Across Media and Languages,” Machine Learning and Knowledge Extraction, vol. 8, no. 4, Apr. 2026.
  10. J. Dang, Y. Sun, and C. Yu, “Fusing Non-Textual Cues with Classical NLP for Enhanced Multimodal Fake News Spread Detection,” Scientific Reports, vol. 16, Art. no. 16193, Apr. 2026.

The rapid growth of social media and online news platforms has significantly increased the spread of misinformation and fake news. Fake news can influence public opinion, create social unrest, and negatively impact decision-making processes. Therefore, developing intelligent systems capable of automatically identifying fake news has become an important research area in Artificial Intelligence and Natural Language Processing (NLP). The proposed work presents a Fake News Detection System that combines traditional machine learning techniques with Large Language Model (LLM) embeddings to improve classification performance. The proposed system utilizes textual news data from benchmark datasets such as WELFake and LIAR. Initially, the news articles undergo pre-processing steps including text cleaning, normalization, and removal of irrelevant characters. Two different feature extraction approaches are employed: Term Frequency-Inverse Document Frequency (TF-IDF) and contextual embeddings generated using the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model. The extracted features are used to train and evaluate multiple Machine learning classifiers, including Logistic Regression, Support Vector Machine (SVM), Multi-Layer Perceptron (MLP), and XGBoost. Performance evaluation is carried out using standard metrics such as Accuracy, Precision, Recall, F1-Score, and Area Under the ROC Curve (AUC-ROC). Comparative analysis is performed to assess the effectiveness of traditional TF-IDF features versus BERT-based semantic embeddings. Experimental results demonstrate that BERT embeddings capture contextual and semantic information more effectively than conventional text representations, leading to improved fake news classification accuracy. The system also provides visualization tools such as confusion matrices, ROC curves, and performance comparison charts for comprehensive evaluation. The proposed framework highlights the potential of integrating transformer-based language models with classical machine learning algorithms to build robust and scalable fake news detection systems. Our work contributes to combating misinformation by providing an efficient and automated approach for identifying deceptive news content in digital media environments.

Keywords : Fake News Detection, Natural Language Processing, Machine Learning, BERT, LLM Embeddings, TF-IDF, Logistic Regression, Support Vector Machine, XGBoost, Deep Learning, Text Classification.

Paper Submission Last Date
31 - August - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe