Q317 : Multi-label data stream classification using incremental active learning
Thesis > Central Library of Shahrood University > Computer Engineering > PhD > 2025
Authors:
Abstarct: Multi-label data streams consist of sequences of instances that arrive continuously over time, where each instance may be simultaneously associated with multiple labels. Such data streams are increasingly prevalent in a wide range of real-world applications, including intelligent multimedia content analysis, text mining, and fraud detection. However, the integration of multi-label characteristics with streaming environments gives rise to a set of compounded challenges, such as non-stationary data distributions, the presence of concept drift, the severe scarcity of labeled instances, high annotation costs, and the necessity for online, single-pass, and computationally efficient learning. Consequently, many existing methods, whether classical or deep learning-baxsed, either rely heavily on batch training with fully labeled data or exhibit limited robustness and efficiency when operating under concept drift and constrained computational resources.
To address these challenges, this study proposes two complementary and purpose-driven frxameworks for learning from multi-label data streams in the presence of concept drift. The first frxamework, termed MLALDDS, is designed as a lightweight and incremental solution that integrates single-pass active learning, a self-adjusting k-nearest neighbor classifier within a binary relevance architecture, and a reflective drift-aware mechanism baxsed on ADWIN. This frxamework enables effective learning from sparsely labeled data streams under strict labeling budget constraints. By prioritizing instances with high informational value and incorporating label-wise adaptation to concept drift, MLALDDS achieves a balanced trade-off between predictive performance, labeling cost, and computational efficiency.
In the second stage, the DMLADMS frxamework is introduced as a deep, active, and incremental approach tailored for highly dynamic multi-label data streams, aiming to enhance overall predictive performance and robustness compared to the first proposed method. DMLADMS employs a shallow deep architecture that supports incremental training, alongside an adaptive active learning strategy driven by local uncertainty thresholds and a dynamically adjusted labeling budget. Furthermore, the frxamework integrates a two-level memory management scheme, consisting of short-term and long-term memories, as well as explicit concept drift detection. These components collectively enable the selection of informative training instances and facilitate continuous, stable, and drift-aware learning under evolving data distributions, without requiring full model retraining.
Extensive experimental evaluations conducted on 30 benchmark datasets for MLALDDS and 20 datasets for DMLADMS, using 12 widely adopted evaluation metrics, including subset accuracy and example-baxsed F1, along with non-parametric statistical significance tests, demonstrate the statistical superiority, robustness, and computational efficiency of both proposed frxameworks when compared to a broad range of state-of-the-art methods. Overall, the results indicate that, by introducing two complementary learning paradigms, this study effectively addresses critical research gaps in multi-label data stream learning and provides practical, scalable, and reliable solutions suitable for real-world applications.
Keywords:
#Keywords: Multi-label Data Stream #Multi-label Learning #Active Learning #Incremental Learning #Concept Drift #Data Stream #Deep Learning Keeping place: Central Library of Shahrood University
Visitor:
Visitor: