Singh, Vishal Krishna and Anand, Niharika and Bhalekar, Abhijit Mahendra and Singh Rathore, Rajkumar (2027) Multimodal multi-teacher knowledge distillation with Keyword Attention for Health Mention Classification. Pattern Recognition, 181. p. 114679. DOI https://doi.org/10.1016/j.patcog.2026.114679
Singh, Vishal Krishna and Anand, Niharika and Bhalekar, Abhijit Mahendra and Singh Rathore, Rajkumar (2027) Multimodal multi-teacher knowledge distillation with Keyword Attention for Health Mention Classification. Pattern Recognition, 181. p. 114679. DOI https://doi.org/10.1016/j.patcog.2026.114679
Singh, Vishal Krishna and Anand, Niharika and Bhalekar, Abhijit Mahendra and Singh Rathore, Rajkumar (2027) Multimodal multi-teacher knowledge distillation with Keyword Attention for Health Mention Classification. Pattern Recognition, 181. p. 114679. DOI https://doi.org/10.1016/j.patcog.2026.114679
Abstract
Health Mention Classification (HMC) aims to distinguish genuine health-related posts on social media from cases where disease or symptom terms are used figuratively, hyperbolically, or in non-personal contexts. Accurately interpreting such ambiguous expressions requires not only textual understanding but also contextual cues that often arise from conversational emotion and multimodal interactions. To address this challenge, this paper proposes a multimodal multi-teacher knowledge distillation (MTKD) framework for HMC on Reddit. The primary classification model is trained and evaluated on the Reddit Health Mention Dataset (RHMD), where a general-context BERT teacher and a domain-specific BioBERT teacher are independently fine-tuned, and their softened knowledge is distilled into a compact DistilBERT student. A lightweight Keyword Attention module further guides the student to focus on health-related lexical cues while preserving the surrounding contextual information required to distinguish literal and non-literal expressions. In addition, we leverage the MELD dataset as an auxiliary multimodal resource, where textual, acoustic, and visual conversational cues are used to train a multimodal teacher. The semantic and affective conversational knowledge learned by this teacher is transferred to the student through an auxiliary distillation objective. Consequently, the final HMC inference model remains entirely text-based, making it directly deployable on text-only social media platforms such as Reddit while benefiting from multimodal knowledge acquired during training. Experimental results on RHMD demonstrate that the proposed framework achieves an Accuracy of 0.8190, Precision-Macro of 0.8232, Recall-Macro of 0.8191, F1-Macro of 0.8203, and a TN_FM score of 0.8563, outperforming existing approaches. Additional evaluation on the SHMD dataset demonstrates strong robustness under distributional shifts. Despite benefiting from multiple teacher models during training, the final student model remains computationally efficient, requiring only 67.55M parameters with an average inference latency of 9.845 ms per sample.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Health Mention Classification; Multi-teacher knowledge distillation; Keyword Attention; DistilBERT; BioBERT; Multimodal distillation; Reddit health mention dataset |
| Subjects: | Z Bibliography. Library Science. Information Resources > ZR Rights Retention |
| Divisions: | Faculty of Science and Health Faculty of Science and Health > Computer Science and Electronic Engineering, School of |
| SWORD Depositor: | Unnamed user with email elements@essex.ac.uk |
| Depositing User: | Unnamed user with email elements@essex.ac.uk |
| Date Deposited: | 11 Sep 2026 11:22 |
| Last Modified: | 11 Sep 2026 11:25 |
| URI: | http://repository.essex.ac.uk/id/eprint/43758 |
Available files
Filename: 2nd revised - Final Paper.pdf
Licence: Creative Commons: Attribution 4.0