SUST Repository

An Enhanced Intelligent Server Monitoring System for Predictive Fault Detection and Triage Prioritization

Show simple item record

dc.contributor.author Ibrahim, Ahmed Abdulmonim Mohamed
dc.contributor.author Alnour, Alrahma Fadol Allah Gomaa
dc.contributor.author Mohamed, Marwa Yousuf
dc.contributor.author Abdelsalam, Maria Adil Mohamed
dc.contributor.author Ali, Abdalmajed Mohammed Abdalrheem
dc.date.accessioned 2026-09-27T10:59:39Z
dc.date.available 2026-09-27T10:59:39Z
dc.date.issued 2026-09-01
dc.identifier.citation Ibrahim, Ahmed Abdulmonim Mohamed. An Enhanced Intelligent Server Monitoring System for Predictive Fault Detection and Triage Prioritization/ Ahmed Abdulmonim Mohamed Ibrahim ,Alrahma Fadol Allah Gomaa Alnour ,Marwa Yousuf Mohamed,Maria Adil Mohamed Abdelsalam, Abdalmajed Mohammed Abdalrheem Ali; Rashid A. SAEED.- Khartoum : Sudan University Of Science and Technology ,College Of Engineering, 2026.- 85p:ill ;28cm.- B.Sc. en_US
dc.identifier.uri https://repository.sustech.edu/handle/123456789/28509
dc.description.abstract Modern digital infrastructure relies heavily on continuous server availability, yet traditional monitoring tools remain largely reactive, relying on static thresholds that generate excessive false alerts and often miss early signs of degradation. This research presents an enhanced intelligent server monitoring system that combines real-time stream processing with a hybrid unsupervised machine learning model for predictive fault detection and triage prioritization. The system ingests server metrics (CPU, memory, disk, and network utilization) through an Apache Kafka message bus and computes a weighted hybrid anomaly score. S(x_t) that combines an Isolation Forest outlier score with an Autoencoder reconstruction error, and classifies detected anomalies into priority levels using a Random Forest classifier. Prometheus and InfluxDB provide short-term and long-term time-series storage, respectively; Grafana delivers live visualization dashboards; Telegram delivers real-time alerts. An independent Apache Spark batch layer periodically generates aggregate historical reports. The models were trained on 2,243 real records from the Alibaba Cluster Trace 2018 production dataset, selected after evaluating and rejecting two alternative public datasets (the Server Machine Dataset, whose 38 metrics are anonymized, and the NAB AWS CloudWatch collection, which lacks a memory-utilization metric). The hybrid score's decision threshold (0.4963) was derived automatically from the 98th percentile of the training score distribution. System performance was validated across 28 automated tests spanning unit, integration, and deployment levels (100% pass rate), and the anomaly detector was evaluated against a test set with known injected anomalies, achieving a 92.5% detection rate (Recall) with a 3.39% false-positive rate. The results demonstrate that a hybrid unsupervised approach, trained on real production telemetry, can deliver practical, low-noise anomaly detection suitable for real-time server health monitoring. en_US
dc.description.sponsorship Sudan University of Science and Technology en_US
dc.language.iso en en_US
dc.publisher Sudan University of Science and Technology en_US
dc.subject Electronics Engineering en_US
dc.subject Triage Prioritization en_US
dc.subject Predictive Fault Detection en_US
dc.subject Enhanced Intelligent en_US
dc.subject Server Monitoring System en_US
dc.title An Enhanced Intelligent Server Monitoring System for Predictive Fault Detection and Triage Prioritization en_US
dc.type Other en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search SUST


Browse

My Account