Q343 : Text-baxsed Software Vulnerability Detection Using Large Language Models and Machine Learning Techniques
Thesis > Central Library of Shahrood University > Computer Engineering > MSc > 2026
Authors:
Abstarct: Software vulnerability detection is an important challenge in the field of cybersecurity and plays a significant role in reducing security threats and improving the reliability of software systems. With the increasing complexity of source code, traditional vulnerability detection approaches face limitations in identifying complex and semantic vulnerability patterns. This research proposes a hybrid frxamework baxsed on CodeBERT, Neighborhood Components Analysis (NCA), and K-Nearest Neighbors (KNN) for automated software vulnerability detection. The main objective of this research is to extract semantic representations from source code, select effective features, and classify code samples into two classes: safe and vulnerable. The standard Juliet Test Suite version 1.3 is used to evaluate the proposed frxamework. During the preprocessing stage, the source-code samples are cleaned and normalized, and inappropriate data in terms of formatting, length, and duplication are controlled. Duplicate and very short samples are removed, while the input length is controlled to ensure compatibility with CodeBERT. Data Augmentation is not used in this research to generate synthetic samples; therefore, the original dataset samples are directly provided to the feature extraction stage after preprocessing. CodeBERT is then employed to extract semantic features from the source code, and the final representation of each sample is generated using the mean pooling of the output token representations. Subsequently, NCA is applied to select effective features and reduce the feature space. Finally, KNN is used to classify the samples into safe and vulnerable classes. The appropriate value of the K parameter is determined using the validation set. The experimental results demonstrate that the proposed KNN–NCA–CodeBERT frxamework provides a suitable capability for distinguishing between safe and vulnerable code samples and can be considered a promising text-baxsed intelligent approach for software security analysis.
Keywords:
#Keywords: Software Vulnerability Detection #CodeBERT #Neighborhood Components Analysis #K-Nearest Neighbors #Machine Learning #Software Security. Keeping place: Central Library of Shahrood University
Visitor:
Visitor: