Server monitoring
The monitoring feature comes bundled with the AI Infrastructure – Metal Cloud service.
Collecting and visualizing metrics, logs, and events helps you identify potential issues and optimize future workloads. You can choose an observability solution fitting your needs.
| Metrics | Cluster (same VPC) | Single Server |
|---|---|---|
| Total nodes and down nodes | ✔ | |
| GPU model, Driver & CUDA version | ✔ | |
| Power state | ✔ | |
| Uptime | ✔ | |
| Total GPUs and down GPUs | ✔ | ✔ |
| GPU Utilization | ✔ | ✔ |
| GPU Memory | ✔ | ✔ |
| CPU Utilization | ✔ | ✔ |
| System Memory | ✔ | ✔ |
| Root Storage Usage | ✔ | ✔ |
| Local Disk Usage | ✔ | ✔ |
| Per-GPU details: power consumption, temperature, utilization, VRAM | ✔ | |
| Network Bandwidth In/Out | ✔ | ✔ |
| Network Packets Sent/Received | ✔ | ✔ |
| Network Error rate Rx/Tx | ✔ | |
| InfiniBand Bandwidth/Packet/Error | ✔ | |
| System Fan Speed | ✔ | |
| System Voltage | ✔ | |
| Common Alerts | ✔ |
Custom or advanced metrics can be addressed through the Cloud Monitoring (FMON) service, which carries an additional charge.