Hardware

Ai2 Replaces Priority Scheduler on Thousands of GPUs

The Allen Institute for AI overhauled its GPU cluster scheduling with a budget-based, time-sliced model, slashing debug queue times and boosting infrastructure efficiency.

AI22 days agoHardware
Image: AI2

The Allen Institute for AI (Ai2) has replaced its traditional priority-based scheduler with a hierarchical fair-share allocation system using GPU time budgets and enforced contracts. Managing thousands of NVIDIA H100, B200, and B300 GPUs configured in clusters spanning 88 to 1024 GPUs, the institute serves roughly 150 internal researchers. With workload demand consistently exceeding available capacity by 2-3x, the new framework addresses systemic operational friction such as resource squatting and priority inflation.

Under the revised model, GPU budgets are assigned hierarchically by management, adapting scheduling mechanics traceable to the Hadoop Fair Scheduler in 2009. The algorithm tracks occupancy across a 7-day lookback window and caps protected minimum runtimes at 8 hours, while debug jobs requiring 15 minutes or less obtain fast queue placement. Workloads can also run as unallocated capacity, which is subject to preemption but incurs no budget charge.

During a 30-day evaluation following an end-of-July rollout, total cluster occupancy held at 98%, with 18% of delivered GPU time coming from unallocated workloads. Teams received 98% of the GPU hours owed to them, with 13 of 15 program allocations receiving 95% or more and the lowest reaching 90%. On Ai2's largest H100 cluster, median queue wait times dropped from 5 minutes to 24 seconds, while 90th-percentile wait times fell from 2.8 hours to 1.8 hours. For debug jobs, 90th-percentile wait times plummeted from 2 hours to 30 seconds, beating simulator predictions that projected a drop from 6 hours to 5 minutes. Automated job draining also reduced human-in-the-loop repair tasks by 74%.

For practitioners, the shift turns compute allocation into an explicit administrative budget process. Researcher Chris Clark noted that bursting past baseline limits during low-demand periods felt like gaining an extra 30% compute. To address issues where interactive sessions are preempted after 8 hours, Ai2 is now building a dedicated CPU-only cluster with restorable sessions for data preparation.

This is our own summary of reporting by AI2

More in Hardware