A machine learning pipeline that analyzes server resource metrics using ensemble clustering techniques, combining multiple dimensionality reduction methods to identify optimization opportunities and generate actionable recommendations.
This system applies advanced clustering algorithms to server metrics data, automatically identifying groups of servers with similar resource utilization patterns. It provides:
- Automated server grouping based on CPU, memory, cost, and project metrics
- Outlier detection to identify misconfigured or anomalous servers
- Feature importance analysis to understand what drives clustering decisions
- Actionable recommendations for cost and performance optimization
- Comprehensive visualizations of cluster characteristics
The system automatically determines the optimal number of clusters using silhouette score analysis, testing configurations from 2 to 12 clusters and identifying the configuration that maximizes cluster separation while maintaining cohesion.
Example Output: In a typical analysis of 44 servers, the system identified 6 clusters as optimal with a silhouette score of 0.62, significantly outperforming alternative configurations (2 clusters: 0.60, 4 clusters: 0.06, 3 clusters: -0.09).
Understanding which features drive cluster separation is critical for interpreting results and making informed decisions.
Key Findings from Sample Analysis:
| Feature | F-Score | P-Value | Significance |
|---|---|---|---|
| Cost per Project | 121.44 | <0.0001 | Highly Significant |
| Count of Staging Projects | 88.38 | <0.0001 | Highly Significant |
| Average Project Memory Usage | 67.91 | <0.0001 | Highly Significant |
| Average Project CPU Load | 20.46 | <0.0001 | Highly Significant |
| Production Projects | 3.77 | 0.0072 | Highly Significant |
| CPU Utilisation | 3.37 | 0.0128 | Significant |
| Cost per GB of Memory | 2.10 | 0.0872 | Not Significant |
| Memory Utilisation | 1.62 | 0.1795 | Not Significant |
The analysis reveals that cost metrics and project counts are the primary drivers of cluster formation, while raw utilization percentages play a secondary role.
Example 6-Cluster Solution:
| Cluster | Servers | Cost/Project (Mean) | Projects (Total) | CPU Util (Mean) | Staging | Production |
|---|---|---|---|---|---|---|
| 0 | 33 | $4.85 | 21 | 0.25% | 17 | 4 |
| 1 | 4 | $2.92 | 11 | 0.33% | 3 | 8 |
| 2 | 3 | $17.24 | 4 | 0.13% | 3 | 1 |
| 3 | 1 | $0.34 | 36 | 1.65% | 36 | 0 |
| 4 | 2 | $1.31 | 25 | 0.55% | 25 | 0 |
| 5 | 1 | $31.14 | 1 | 0.00% | 1 | 0 |
| Total | 44 | $9.63 | 98 | 0.58% | 85 | 13 |
Cluster Profiles:
- Cluster 0 (75% of servers): Low CPU utilization, low cost per project, mixed staging/production workloads - represents typical underutilized servers
- Cluster 1 (9%): Production-focused servers with moderate efficiency
- Cluster 2 (7%): High cost per project, minimal workload - optimization candidates
- Cluster 3 (2%): High-efficiency staging server with maximum utilization (1.65%) and lowest cost per project ($0.34)
- Cluster 4 (5%): Efficient staging environment with good cost metrics
- Cluster 5 (2%): Idle server with highest cost per project - immediate decommissioning candidate
The system generates comprehensive visualizations showing:
- Multi-dimensional scatter plots displaying servers across cost, utilization, and project count axes
- Comparative dashboards showing variance within clusters and non-linear relationships between features
- Project-based cost breakdowns illustrating how costs distribute across different server configurations
- Cluster comparison views enabling side-by-side analysis of resource patterns
Key Insight from Visualizations: Server 871 with 0% CPU utilization and $1,762 monthly cost exemplifies Cluster 0's pattern of over-provisioning, while Server 874 in Cluster 3 demonstrates optimal resource utilization with 1.65% CPU usage supporting 36 projects at only $474 monthly cost.
Performance of Different Configurations:
| N Clusters | Silhouette Score | Base Algorithms | Interpretation |
|---|---|---|---|
| 6 | 0.62 | K-means, Agglomerative, Spectral, BIRCH | Optimal - Strong separation |
| 2 | 0.60 | K-means, Agglomerative, Spectral, BIRCH | Good - Binary split |
| 4 | 0.06 | K-means, Agglomerative, Spectral, BIRCH | Weak - Poor separation |
| 3 | -0.09 | K-means, Agglomerative, Spectral, BIRCH | Poor - Overlapping clusters |
The ensemble approach tests multiple algorithms (K-means, Agglomerative, Spectral, BIRCH) across different cluster counts, then combines their results through consensus clustering for robust, reliable groupings.
- Principal Component Analysis (PCA) - Linear reduction with variance maximization
- Kernel PCA - Non-linear patterns via RBF kernels
- Variational Autoencoder (VAE) - Probabilistic neural network encoding
- Standard Autoencoder - Deep learning feature extraction
- Truncated SVD - Efficient for sparse matrices
- UMAP (optional) - Manifold learning for visualization
- K-Means - Centroid-based partitioning
- Agglomerative Hierarchical - Bottom-up tree clustering
- Spectral Clustering - Graph-based similarity clustering
- BIRCH - Scalable hierarchical method
- Consensus Ensemble - Combines multiple methods for robustness
- Outlier Detection - Isolation Forest algorithm identifies anomalies
- Feature Selection - Variance, k-best, and mutual information methods
- Cross-Validation - 5-fold validation for cluster stability
- Robustness Testing - Noise injection to test consistency
- Feature Importance - ANOVA F-test and permutation importance
- CSV Reports: Cluster descriptions, optimization recommendations, outlier list
- 20+ Visualizations: PCA/t-SNE projections, 3D plots, heatmaps, radar charts
- Detailed Logs: Complete execution trace with metrics and decisions