Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Enhanced Deep Clustering for Server Resource Analysis

A machine learning pipeline that analyzes server resource metrics using ensemble clustering techniques, combining multiple dimensionality reduction methods to identify optimization opportunities and generate actionable recommendations.

Overview

This system applies advanced clustering algorithms to server metrics data, automatically identifying groups of servers with similar resource utilization patterns. It provides:

  • Automated server grouping based on CPU, memory, cost, and project metrics
  • Outlier detection to identify misconfigured or anomalous servers
  • Feature importance analysis to understand what drives clustering decisions
  • Actionable recommendations for cost and performance optimization
  • Comprehensive visualizations of cluster characteristics

Example Results

Optimal Cluster Configuration

The system automatically determines the optimal number of clusters using silhouette score analysis, testing configurations from 2 to 12 clusters and identifying the configuration that maximizes cluster separation while maintaining cohesion.

Example Output: In a typical analysis of 44 servers, the system identified 6 clusters as optimal with a silhouette score of 0.62, significantly outperforming alternative configurations (2 clusters: 0.60, 4 clusters: 0.06, 3 clusters: -0.09).

Feature Importance Analysis

Understanding which features drive cluster separation is critical for interpreting results and making informed decisions.

Key Findings from Sample Analysis:

Feature F-Score P-Value Significance
Cost per Project 121.44 <0.0001 Highly Significant
Count of Staging Projects 88.38 <0.0001 Highly Significant
Average Project Memory Usage 67.91 <0.0001 Highly Significant
Average Project CPU Load 20.46 <0.0001 Highly Significant
Production Projects 3.77 0.0072 Highly Significant
CPU Utilisation 3.37 0.0128 Significant
Cost per GB of Memory 2.10 0.0872 Not Significant
Memory Utilisation 1.62 0.1795 Not Significant

The analysis reveals that cost metrics and project counts are the primary drivers of cluster formation, while raw utilization percentages play a secondary role.

Cluster Characteristics

Summary Statistics by Cluster

Example 6-Cluster Solution:

Cluster Servers Cost/Project (Mean) Projects (Total) CPU Util (Mean) Staging Production
0 33 $4.85 21 0.25% 17 4
1 4 $2.92 11 0.33% 3 8
2 3 $17.24 4 0.13% 3 1
3 1 $0.34 36 1.65% 36 0
4 2 $1.31 25 0.55% 25 0
5 1 $31.14 1 0.00% 1 0
Total 44 $9.63 98 0.58% 85 13

Cluster Profiles:

  • Cluster 0 (75% of servers): Low CPU utilization, low cost per project, mixed staging/production workloads - represents typical underutilized servers
  • Cluster 1 (9%): Production-focused servers with moderate efficiency
  • Cluster 2 (7%): High cost per project, minimal workload - optimization candidates
  • Cluster 3 (2%): High-efficiency staging server with maximum utilization (1.65%) and lowest cost per project ($0.34)
  • Cluster 4 (5%): Efficient staging environment with good cost metrics
  • Cluster 5 (2%): Idle server with highest cost per project - immediate decommissioning candidate

Detailed Cluster Visualizations

The system generates comprehensive visualizations showing:

  1. Multi-dimensional scatter plots displaying servers across cost, utilization, and project count axes
  2. Comparative dashboards showing variance within clusters and non-linear relationships between features
  3. Project-based cost breakdowns illustrating how costs distribute across different server configurations
  4. Cluster comparison views enabling side-by-side analysis of resource patterns

Key Insight from Visualizations: Server 871 with 0% CPU utilization and $1,762 monthly cost exemplifies Cluster 0's pattern of over-provisioning, while Server 874 in Cluster 3 demonstrates optimal resource utilization with 1.65% CPU usage supporting 36 projects at only $474 monthly cost.

Clustering Algorithm Comparison

Performance of Different Configurations:

N Clusters Silhouette Score Base Algorithms Interpretation
6 0.62 K-means, Agglomerative, Spectral, BIRCH Optimal - Strong separation
2 0.60 K-means, Agglomerative, Spectral, BIRCH Good - Binary split
4 0.06 K-means, Agglomerative, Spectral, BIRCH Weak - Poor separation
3 -0.09 K-means, Agglomerative, Spectral, BIRCH Poor - Overlapping clusters

The ensemble approach tests multiple algorithms (K-means, Agglomerative, Spectral, BIRCH) across different cluster counts, then combines their results through consensus clustering for robust, reliable groupings.

Key Features

Dimensionality Reduction

  • Principal Component Analysis (PCA) - Linear reduction with variance maximization
  • Kernel PCA - Non-linear patterns via RBF kernels
  • Variational Autoencoder (VAE) - Probabilistic neural network encoding
  • Standard Autoencoder - Deep learning feature extraction
  • Truncated SVD - Efficient for sparse matrices
  • UMAP (optional) - Manifold learning for visualization

Clustering Algorithms

  • K-Means - Centroid-based partitioning
  • Agglomerative Hierarchical - Bottom-up tree clustering
  • Spectral Clustering - Graph-based similarity clustering
  • BIRCH - Scalable hierarchical method
  • Consensus Ensemble - Combines multiple methods for robustness

Analysis & Validation

  • Outlier Detection - Isolation Forest algorithm identifies anomalies
  • Feature Selection - Variance, k-best, and mutual information methods
  • Cross-Validation - 5-fold validation for cluster stability
  • Robustness Testing - Noise injection to test consistency
  • Feature Importance - ANOVA F-test and permutation importance

Outputs

  • CSV Reports: Cluster descriptions, optimization recommendations, outlier list
  • 20+ Visualizations: PCA/t-SNE projections, 3D plots, heatmaps, radar charts
  • Detailed Logs: Complete execution trace with metrics and decisions

About

Masters Apprenticeship Assignment Data Analytics 101

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages