THE CRUNCH

The Allen Institute for AI (Ai2) has replaced the priority-based scheduler for its GPU clusters with a system built on GPU time budgets, hierarchical fair-share allocation and a time-slicing contract, according to a post by its AI Infrastructure team. The change moves decisions about how much GPU time each research project deserves from case-by-case operational wrangling to a transparent administrative budgeting process.

The institute runs thousands of NVIDIA H100, B200 and B300 GPUs in clusters of 88 to 1024 GPUs, serving roughly 150 internal researchers working on LLM and VLM training, robotics reinforcement learning simulation and post-training for scientific agentic use cases. Demand consistently outstrips supply: at any moment, outstanding requests run at 2 to 3 times the available GPUs.

The old priority-based setup, which let workloads opt out of preemption, produced familiar pathologies. Researchers parked no-op workloads on GPUs to grab them later, priority inflation meant every scheduled job claimed HIGH priority and lower tiers were starved, and on-call engineers spent most of their ticket time negotiating the shutdown of non-preemptable workloads on machines needing maintenance.

Ai2 also places its experience in a longer tradition, noting that in a 2011 paper on Dominant Resource Fairness, Ghodsi et al. described a search company whose users sprinkled their code with infinite loops to artificially inflate utilization levels and keep dedicated machines.