Responsiveness

A busy cluster that nobody can get onto is not a success. These are the numbers users actually feel: how long they wait, and whether their jobs finish.

Data freshness unknown

Job success rate not yetCompleted vs. all finished jobs, year to date.
Jobs run not yetYear to date.
Availability not yetShare of installed GPU-hours that were up and schedulable.

Queue wait

Time from submit to start

Median, 90th, and 99th percentile. The tail matters more than the median: a cluster with a five-minute median and a twelve-hour p99 feels unpredictable, and unpredictable is what drives people to the cloud.

Outcomes

How jobs end

Last 120 days. Failures and out-of-memory exits are as much a documentation problem as a hardware one — this chart is how we find out.

Job sizes

A healthy shared cluster serves both single-GPU exploration and multi-node capability runs. This shows whether it actually does.