Responsiveness

A busy cluster that nobody can get onto is not a success. These are the numbers users actually feel: how long they wait, and whether their jobs finish.

Data freshness unknown

Job success rate 84.8%Completed vs. all finished jobs, year to date.
Jobs run 1.2MYear to date.
Availability 95.6%Share of installed GPU-hours that were up and schedulable.

Queue wait

Time from submit to start

Median, 90th, and 99th percentile. The tail matters more than the median: a cluster with a five-minute median and a twelve-hour p99 feels unpredictable, and unpredictable is what drives people to the cloud.

Outcomes

How jobs end

Last 120 days. Failures and out-of-memory exits are as much a documentation problem as a hardware one — this chart is how we find out.

Job sizes

A healthy shared cluster serves both single-GPU exploration and multi-node capability runs. This shows whether it actually does.