Please use this identifier to cite or link to this item: https://repository.sustech.edu/handle/123456789/28509
Title: An Enhanced Intelligent Server Monitoring System for Predictive Fault Detection and Triage Prioritization
Authors: Ibrahim, Ahmed Abdulmonim Mohamed
Alnour, Alrahma Fadol Allah Gomaa
Mohamed, Marwa Yousuf
Abdelsalam, Maria Adil Mohamed
Ali, Abdalmajed Mohammed Abdalrheem
Keywords: Electronics Engineering
Triage Prioritization
Predictive Fault Detection
Enhanced Intelligent
Server Monitoring System
Issue Date: 1-Sep-2026
Publisher: Sudan University of Science and Technology
Citation: Ibrahim, Ahmed Abdulmonim Mohamed. An Enhanced Intelligent Server Monitoring System for Predictive Fault Detection and Triage Prioritization/ Ahmed Abdulmonim Mohamed Ibrahim ,Alrahma Fadol Allah Gomaa Alnour ,Marwa Yousuf Mohamed,Maria Adil Mohamed Abdelsalam, Abdalmajed Mohammed Abdalrheem Ali; Rashid A. SAEED.- Khartoum : Sudan University Of Science and Technology ,College Of Engineering, 2026.- 85p:ill ;28cm.- B.Sc.
Abstract: Modern digital infrastructure relies heavily on continuous server availability, yet traditional monitoring tools remain largely reactive, relying on static thresholds that generate excessive false alerts and often miss early signs of degradation. This research presents an enhanced intelligent server monitoring system that combines real-time stream processing with a hybrid unsupervised machine learning model for predictive fault detection and triage prioritization. The system ingests server metrics (CPU, memory, disk, and network utilization) through an Apache Kafka message bus and computes a weighted hybrid anomaly score. S(x_t) that combines an Isolation Forest outlier score with an Autoencoder reconstruction error, and classifies detected anomalies into priority levels using a Random Forest classifier. Prometheus and InfluxDB provide short-term and long-term time-series storage, respectively; Grafana delivers live visualization dashboards; Telegram delivers real-time alerts. An independent Apache Spark batch layer periodically generates aggregate historical reports. The models were trained on 2,243 real records from the Alibaba Cluster Trace 2018 production dataset, selected after evaluating and rejecting two alternative public datasets (the Server Machine Dataset, whose 38 metrics are anonymized, and the NAB AWS CloudWatch collection, which lacks a memory-utilization metric). The hybrid score's decision threshold (0.4963) was derived automatically from the 98th percentile of the training score distribution. System performance was validated across 28 automated tests spanning unit, integration, and deployment levels (100% pass rate), and the anomaly detector was evaluated against a test set with known injected anomalies, achieving a 92.5% detection rate (Recall) with a 3.39% false-positive rate. The results demonstrate that a hybrid unsupervised approach, trained on real production telemetry, can deliver practical, low-noise anomaly detection suitable for real-time server health monitoring.
URI: https://repository.sustech.edu/handle/123456789/28509
Appears in Collections:Bachelor of Engineering

Files in This Item:
File Description SizeFormat 
FYP - 2025 - G6.pdfResearch5.87 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.