THE CRUNCH

Data centres are on track to consume around 945TWh of electricity for AI by 2030, roughly as much as Japan uses today, according to the International Energy Agency. Yet the industry's response has leaned heavily on hardware: faster chips and more grid connections. Researchers argue that software, the layer above the silicon, offers easier and compounding savings. The Uptime Institute's 2025 survey found average power usage effectiveness (PUE), a measure of how much energy goes to cooling and overheads rather than computing, has barely moved in six years. Servers account for around 60% of a modern data centre's electricity, so trimming what the software actually does could matter more than shaving cooling overheads.

The evidence is concrete. Tests by the ML.Energy initiative found that running inference on Alibaba's Qwen 3 235B A22B Thinking model in FP8, a lower-precision number format, used a third less energy than bfloat16 on problem-solving tasks. Jae-Won Chung, a University of Michigan PhD candidate behind the initiative, also built Perseus, a training optimiser that slows underworked parts of a large-model training job so they finish alongside busier ones, cutting training energy by up to 30% without reducing throughput or touching the hardware.

Industry is already moving this way. Nvidia's Blackwell power profiles adjust GPU frequencies, power limits and cache settings to match the workload, which the company says can save up to 15% of energy while keeping 97% or more of performance, and let power-constrained sites fit in more GPUs for up to 13% higher throughput. Simpler wins include right-sizing cloud instances, caching repeated AI prompts, routing routine requests to small models, and capping output lengths.

Software can also shift when and where electricity is drawn. Sophie Hall, a doctoral student at ETH Zurich who has studied workload shifting across Google's data-centre fleet, argues the question is not just how much power is used but when and where, and how that interacts with the grid. Delaying non-urgent batch jobs until local demand falls, or routing them to regions with spare capacity, can help. There are limits: data sovereignty rules can stop workloads crossing borders, and physically moving large datasets between sites is far from trivial.

None of this makes software a substitute for better chips and new power infrastructure. And efficiency has a catch known as Jevons' paradox: if every watt saved simply generates more tokens, total electricity use may not fall. As Chung put it, "Power is the core bottleneck in AI data centers. We really want to make the best use of every watt we consume."