TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

chips infrastructure

Thirty Percent Less Power Did Not Cost Every Training Job Equally

Across 189 H100 and H200 runs, a job-aware allocator recovered 63% of the throughput gap to an oracle under a power cap.

Published Updated Story ID: mp-2026-09-11-008
Read the complete editionStory JSON

Summary

Across 189 H100 and H200 runs, a job-aware allocator recovered 63% of the throughput gap to an oracle under a power cap.

The study defines a Power Flexibility Index to measure how much LLM-training throughput changes when GPU power is reduced. Its evidence covers 131 H200 runs, 24 H200 validations and 34 matched H100 runs across dense and mixture-of-experts models, pretraining and fine-tuning, and deployments up to 32 GPUs. Under a 30% power reduction, allocating power by the learned index recovered about 1,500 tokens per second per job—63% of the gap between equal allocation and perfect foresight. The result shows measurable job-level flexibility, not a general claim about data-center electricity or grid impacts.

Why it matters

Across 189 H100 and H200 runs, a job-aware allocator recovered 63% of the throughput gap to an oracle under a power cap.

Limits and context

  • The result shows measurable job-level flexibility, not a general claim about data-center electricity or grid impacts.

Key claims

  1. Across 189 H100 and H200 runs, a job-aware allocator recovered 63% of the throughput gap to an oracle under a power cap.

    Qualification: The result shows measurable job-level flexibility, not a general claim about data-center electricity or grid impacts.

    Evidence: source-2026-09-11-008

Sources

  1. arXiv preprint 2609.11542arXiv · primary research

Corrections

No corrections have been recorded for this story.