AUT Journal of Modeling and Simulation

AUT Journal of Modeling and Simulation

Labeling Text Reviews for Customer Churn Prediction via Active Learning, Semi-Supervised Learning, and Explainable AI

Document Type : Research Article

Authors
Faculty of Computer Engineering, University of Isfahan
10.22060/miscj.2026.25707.5487
Abstract
Customer churn prediction is critical for businesses, yet most prior work relies on structured data and makes limited use of textual reviews, leaving low-resource languages such as Persian underexplored. Text-based churn models require large labeled corpora, which are often expensive and time-consuming to produce. To address this, we propose a practical method that combines least-confidence active learning, explainable AI (LIME) to assist expert annotators, and iterative semi-supervised labeling (self-training, label propagation, and a hybrid method). To our best knowledge there is no study on using active learning, semi-supervised learning and explainable AI to label textual data for churn prediction. Beginning with an expert-verified seed of 2,004 balanced Persian reviews, we trained LSTM/GRU classifiers on ParsBERT embeddings, augmented training sets via active sampling (100 samples per fold), and expanded labels to 10,500 Digikala reviews over 12 semi-supervised stages. Active learning increased average accuracy from 82.5% to 84.7%, and self-training produced the best sustained gains (mean peak accuracies 86%), outperforming label propagation and the hybrid approach. Moreover, the explanations generated by LIME were positively evaluated by three domain experts, with mean ratings exceeding 4 for two experts and above 3.9 for the third, reflecting consistently high satisfaction. In addition, expert agreement increased by up to 8%, reaching 0.91 after reviewing the explanations. The proposed framework offers a reproducible solution for churn prediction in low-resource text domains by introducing the first labeled Persian review dataset for churn prediction and a unified framework integrating active learning, semi-supervised learning, and explainable AI.
Keywords
Subjects


Articles in Press, Accepted Manuscript
Available Online from 11 August 2026