Overview
Observability provides live and historical metrics for your GPU clusters and nodes so you can monitor performance and troubleshoot workloads. Use the dashboard to view GPU metrics (utilization, power, temperature, DRAM, SM clocks), CPU metrics, memory usage, and time-series charts that can be zoomed or maximized for deeper inspection.Access Observability
- In the left navigation, open
More → Observability. - On the Observability landing page, click on the
View Observabilityon theGPU Clusterscard.
Explore the dashboard
Once Observability is open, use these controls to get the data you need:Provided metrics
The dashboard includes the following metric panels. Use the table below to quickly find what each panel measures, the typical units, and when to investigate.Troubleshooting & tips
- If panels show
No Data Available:- Verify you selected the correct cluster and node.
- Ensure the time window includes the period when the workload ran.
- Check that cluster agents / exporters are running on the nodes (for on-prem or self-hosted configurations).
- If metrics appear sparse or too noisy:
- Increase the aggregation interval in the dashboard to reduce noise.
- Use a longer time window to observe trends rather than instantaneous spikes.
Permissions: Observability data visibility depends on your account permissions. If you cannot see a cluster, confirm you have access to that project and cluster.