Responsiveness
A busy cluster that nobody can get onto is not a success. These are the numbers users actually feel: how long they wait, and whether their jobs finish.
Data freshness unknown
Job success rate
not yetCompleted vs. all finished jobs, year to date.
Jobs run
not yetYear to date.
Availability
not yetShare of installed GPU-hours that were up and schedulable.
Queue wait
Time from submit to start
Median, 90th, and 99th percentile. The tail matters more than the median: a cluster with a five-minute median and a twelve-hour p99 feels unpredictable, and unpredictable is what drives people to the cloud.
Outcomes
How jobs end
Last 120 days. Failures and out-of-memory exits are as much a documentation problem as a hardware one — this chart is how we find out.
Job sizes
A healthy shared cluster serves both single-GPU exploration and multi-node capability runs. This shows whether it actually does.