Q326 : Improving Text Classification Performance Using Deep Generative Networks
Thesis > Central Library of Shahrood University > Computer Engineering > PhD > 2025
Authors:
Abstarct: The accuracy of text classification decreases when faced with imbalanced data. One effective solution to this problem is oversampling the minority class, which aims to increase the size of the minority class to create a balance between classes. Given the advancements in deep generative networks in data generation, these methods can be used to generate samples for the minority class. The advantage of using these methods in the imbalanced data problem is that they overcome the limitations of conventional methods such as overfitting. The aim of this research is to enhance the performance of text classification under imbalanced conditions using Generative Adversarial Networks (GANs). To this end, two proposed methods are presented in this thesis. In the first proposed method, a frxamework for learning the distribution of textual data and performing oversampling is presented with the goal of comparing the performance of different GAN architectures in improving the accuracy of imbalanced text classification. In this section, the quality and diversity of the generated texts were evaluated and compared, and the classification performance was analyzed in relation to these two aspects. Finally, the impact of GAN-baxsed oversampling and traditional methods on various evaluation metrics of imbalanced text classification was examined. The results showed that, overall, data balancing using GANs is effective in improving the classification performance of imbalanced textual data, and these methods generally outperformed traditional oversampling methods. However, the diversity and quality of the generated texts affect the classifier's performance in distinguishing between different classes. In the second proposed method, a novel GAN architecture was introduced such that the data it generates for oversampling leads to improved text classification performance. The main idea of this method is to move away from the semantic space of the majority class in order to generate data that is distinguishable from the majority class. Additionally, a frxamework was presented to investigate the impact of the proposed method on enriching the data generated by large language models for the purpose of oversampling. After applying classification algorithms, calculating evaluation metrics, and comparing with existing oversampling methods and methods baxsed on large language models, it was observed that the proposed method has a good capability in improving text classification performance. Furthermore, combining the data generated by the proposed method with data generated by large language models helps enrich that data and positively affects the improvement of text classification performance.
Keywords:
#Keywords: Generative Adversarial Networks #Text Classification #Oversampling #Class Imbalance #GAN. Keeping place: Central Library of Shahrood University
Visitor:
Visitor: