Q329 : Fraud Detection in Imbalanced Telecom Data Using Machine Learning
Thesis > Central Library of Shahrood University > Computer Engineering > MSc > 2025
Authors:
[Author], [Supervisor]
Abstarct: In today’s fast-paced world of communications, the telecommunications industry faces numerous challenges, one of the most prominent being telephone fraud. Telecommunications fraud, which often occurs through the exploitation of Call Detail Records (CDRs), causes billions of dollars in financial losses to operators and subscribers each year. These fraudulent activities—ranging from suspicious international calls to abnormal service usage patterns—not only threaten financial resources but also undermine public trust in communication systems. Given the massive volume of data generated within telecommunications networks, the use of machine learning techniques for the automatic detection of such rare patterns has become essential and unavoidable. One of the primary challenges in this domain is the severe class imbalance in the data, where more than 90% of records correspond to normal calls, while fraudulent cases account for only about 10%. This imbalance biases machine learning models toward the majority class and renders traditional evaluation metrics such as accuracy ineffective. Consequently, focusing on metrics such as the F1-Score (the harmonic mean of precision and recall) and ROC-AUC, which emphasize the correct detection of rare events, is critical. This study aims to evaluate the performance of a Gradient Boosting Machine (GBM) model in detecting telecommunications fraud by leveraging sampling techniques to address data imbalance. The dataset used in this research consists of 101,174 CDR records obtained from a public Kaggle dataset and was analyzed after comprehensive preprocessing steps. Exploratory Data Analysis (EDA) revealed weak linear relationships among features, thereby confirming the necessity of non-linear models such as GBM. Finally, the results obtained from training baxseline models and models adjusted using SMOTE, ADASYN, and Tomek lixnks techniques demonstrated that the SMOTE-baxsed model achieved the best performance in fraud detection, with an F1-Score of 0.85 and a ROC-AUC of 0.9745. This significant improvement also maintained the false positive rate (FPR) at an acceptable level (below 5%), highlighting the practical potential of deploying this approach in real-world telecommunications systems.
Keywords:
#Keywords: Gradient Boosting Machine (GBM) #Fraud Detection #Imbalanced Learning #SMOTE #ADASYN #Tomek lixnks #ROC-AUC #F1-Score #Call Detail Records (CDR) #Exploratory Data Analysis (EDA). Keeping place: Central Library of Shahrood University
Visitor: