⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



Early Stage Depression Recognition Using Machine Learning and Natural Language Processing: A Hybrid PHQ-9 and Linguistic Feature-Based Screening Framework


Authors : Tanu Chauhan; Abhishek Verma

Volume/Issue : Volume 11 - 2026, Issue 8 - August


Google Scholar : https://tinyurl.com/ycxdbsh8

DOI : https://doi.org/10.38124/ijisrt/26aug324

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : Depression is a widespread and frequently under-diagnosed mental health condition, and delays in identification are associated with poorer long-term outcomes. This paper presents a hybrid screening framework that combines a validated questionnaire to the Patient Health Questionnaire-9 (PHQ-9), with linguistic feature extraction from short free-text journal entries to generate a composite depression-risk score. Unlike approaches that rely on a single data modality, the proposed system fuses structured clinical scoring with textual indicators associated with depressive symptomatology in prior computational-linguistics research, such as elevated first-person pronoun usage, negation patterns, and lexical markers of sadness, anhedonia, and worthlessness. A prototype web application implementing this framework was developed, comprising a screening interface, a relational database schema for longitudinal storage, and an analytics dashboard for aggregate monitoring. This paper describes the system architecture, the feature-fusion methodology, the underlying database design, and a discussion of how the prototype's rule-based linguistic module can be replaced with a trained supervised classifier (e.g., Support Vector Machine, Random Forest, or a fine-tuned transformer model) in future work. The framework is intended as a low-cost, scalable screening aid to support — not replace — clinical [1] evaluation.

Keywords : Depression Recognition; Machine Learning; Natural Language Processing; PHQ-9; Mental Health Screening; Text Classification; Linguistic Analysis.

References :

  1. M. De Choudhury, M. Gamon, S. Counts, and E. Horvitz, "Predicting depression via social media," in Proc. 7th Int. AAAI Conf. Weblogs and Social Media (ICWSM), 2013, pp. 128-137.
  2. S. Motade, A. Hassan, F. Mir, and K. Parikh, "Machine Learning-Based Approach for Depression Detection Using PHQ-9 and Twitter Dataset," in Proc. 3rd Int. Conf. Communication, Computing and Electronics Systems, Lecture Notes in Electrical Engineering, vol. 844, Springer, Singapore, 2022.
  3. Soumya Choudhary, Nikita Thomas, Janine Ellenberger, Grish Srinivasan , "A Machine Learning Approach for Detecting Digital Behavioral Patterns of Depression Using Nonintrusive Smartphone Data: Prospective Observational Study," JMIR Formative Research, 2022.
  4. Hadar Fisher, Nigel M. Jaffer , Kristina Pidcirny ,Anna O. Tierney Mia S. Vaidean , Poorvesh Dongre, and Christian A. Webb, "Language-based detection of depression with machine learning: systematic review and meta-analysis," npj Digital Medicine, 2026.
  5. Z. Shao, X. Wang, Z. Liu, C. Wang, and K. P. Subbalakshmi, "Systematic Evaluation of Machine-Generated Reasoning and PHQ-9 Labeling for Depression Detection Using Large Language Models," arXiv:2505.17119, 2025.

Depression is a widespread and frequently under-diagnosed mental health condition, and delays in identification are associated with poorer long-term outcomes. This paper presents a hybrid screening framework that combines a validated questionnaire to the Patient Health Questionnaire-9 (PHQ-9), with linguistic feature extraction from short free-text journal entries to generate a composite depression-risk score. Unlike approaches that rely on a single data modality, the proposed system fuses structured clinical scoring with textual indicators associated with depressive symptomatology in prior computational-linguistics research, such as elevated first-person pronoun usage, negation patterns, and lexical markers of sadness, anhedonia, and worthlessness. A prototype web application implementing this framework was developed, comprising a screening interface, a relational database schema for longitudinal storage, and an analytics dashboard for aggregate monitoring. This paper describes the system architecture, the feature-fusion methodology, the underlying database design, and a discussion of how the prototype's rule-based linguistic module can be replaced with a trained supervised classifier (e.g., Support Vector Machine, Random Forest, or a fine-tuned transformer model) in future work. The framework is intended as a low-cost, scalable screening aid to support — not replace — clinical [1] evaluation.

Keywords : Depression Recognition; Machine Learning; Natural Language Processing; PHQ-9; Mental Health Screening; Text Classification; Linguistic Analysis.

Paper Submission Last Date
30 - September - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe