Q327 : An efficient text classification model using retrieval-baxsed data augmentation with BERT model and K-nearest neighbor algorithm
Thesis > Central Library of Shahrood University > Computer Engineering > MSc > 2026
Authors:
[Author], [Supervisor]
Abstarct: Abstract In recent years, text classification, as a key problem in the field of natural language processing, has demanded methods with high accuracy and generalizability. In this study, an innovative frxamework for sentiment classification in textual reviews is proposed, aiming to enhance data quality, improve feature representation, and boost final performance. First, using the BERT language model, deep semantic vectors are extracted from sentences and stored in an initial data repository. Then, employing a data augmentation approach baxsed on information retrieval, new and meaningful samples are generated using k-nearest neighbors and the Masked Language Model (MLM) algorithm to increase data diversity and richness. In the next step, feature selection is performed using Neighborhood Component Analysis (NCA) to reduce dimensionality and eliminate less important features. Finally, an ensemble learning structure is applied for classification to improve model accuracy and stability. The proposed method is evaluated on the Movie Review dataset from the SentEval collection, achieving a remarkable accuracy of 97.8% in detecting positive and negative reviews. The results demonstrate that the effective combination of BERT-baxsed feature extraction, intelligent data augmentation, and ensemble learning can significantly enhance the performance of text classification systems.
Keywords:
#Keywords: Text classification #BERT #Data augmentation #NCA algorithm #Ensemble learning #Natural Language Processing (NLP). Keeping place: Central Library of Shahrood University
Visitor: