Skip to main content

Server monitoring

The monitoring feature comes bundled with the AI Infrastructure – Metal Cloud service.

Collecting and visualizing metrics, logs, and events helps you identify potential issues and optimize future workloads. You can choose an observability solution fitting your needs.

MetricsCluster (same VPC)Single Server
Total nodes and down nodes
GPU model, Driver & CUDA version
Power state
Uptime
Total GPUs and down GPUs
GPU Utilization
GPU Memory
CPU Utilization
System Memory
Root Storage Usage
Local Disk Usage
Per-GPU details: power consumption, temperature, utilization, VRAM
Network Bandwidth In/Out
Network Packets Sent/Received
Network Error rate Rx/Tx
InfiniBand Bandwidth/Packet/Error
System Fan Speed
System Voltage
Common Alerts

Custom or advanced metrics can be addressed through the Cloud Monitoring (FMON) service, which carries an additional charge.