Responsiveness
A busy cluster that nobody can get onto is not a success. These are the numbers users actually feel: how long they wait, and whether their jobs finish.
Data freshness unknown
Job success rate
84.8%Completed vs. all finished jobs, year to date.
Jobs run
1.2MYear to date.
Availability
95.6%Share of installed GPU-hours that were up and schedulable.
Queue wait
Time from submit to start
Median, 90th, and 99th percentile. The tail matters more than the median: a cluster with a five-minute median and a twelve-hour p99 feels unpredictable, and unpredictable is what drives people to the cloud.
Outcomes
How jobs end
Last 120 days. Failures and out-of-memory exits are as much a documentation problem as a hardware one — this chart is how we find out.
Job sizes
A healthy shared cluster serves both single-GPU exploration and multi-node capability runs. This shows whether it actually does.