Machine Learning-Based Early Detection of Cyber Attacks Using Network Traffic Analysis
Main Article Content
Abstract
The rapid growth of interconnected computer networks, cloud services, Internet of Things (IoT) devices, and distributed applications has significantly increased organizational exposure to cyber attacks. Conventional intrusion detection systems (IDSs) relying primarily on predefined signatures are largely ineffective against novel, zero-day, or rapidly evolving attacks. Machine learning (ML)-based network intrusion detection provides an adaptive alternative by learning statistical traffic characteristics to identify malicious activities automatically. This study proposes a systematic ML framework for early cyber attack detection using network traffic flow analysis on the CICIDS2017 benchmark dataset. We evaluate six supervised algorithms: Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Logistic Regression (LR), and Extreme Gradient Boosting (XGBoost). Our evaluation emphasizes recall and false-positive rates alongside computational efficiency to assess real-time viability. XGBoost and Random Forest achieved top performance with F1-scores exceeding 0.998 and False Positive Rates below 0.002. Furthermore, progressive observation window experiments demonstrate that high-confidence classification (> 0.98 Recall) can be achieved with only 25% of total flow traffic, validating the framework's effectiveness for early threat mitigation.


