返回新聞
產業AI Understanding 簡報

NVIDIA 預覽電源管理系統,用於在固定資料中心預算內增加人工智慧容量

NVIDIA 表示,其 DSX MaxLPS 套件可回收未使用的機架功率,提高每瓦效能,並在相同設施功率預算內支援高達 40% 的 Rubin GPU 容量。該軟體仍處於開發者預覽版,效能數據來自 NVIDIA 自己的代表性工作負載評估。

6 min readRead the primary source
Source-provided image accompanying NVIDIA previews power-management system for adding AI capacity within fixed data-center budgets
主要來源文件來源記錄
出版商
developer.nvidia.com
來源連結
developer.nvidia.comhttps://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

記憶體(代理記憶體)
AI 代理程式跨步驟或會話使用儲存的上下文來提高連續性。
培訓後
預訓練後應用的訓練步驟,例如指令調整、偏好最佳化和安全調整。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己AI 模型解釋測驗

發生了什麼事

NVIDIA introduced DSX MaxLPS, a site-level system combining dynamic power allocation, workload-specific GPU tuning and 45 °C liquid-cooling design. Its Dynamic Power Software is in Developer Preview and is intended to redistribute unused power across managed racks and GPUs.

On Aug. 21, 2026, NVIDIA published a technical blog introducing DSX MaxLPS, which it expands as Maximum Land Power Shell. The term refers to the three fixed constraints that shape an AI factory: land, utility power and the physical building that contains power distribution, cooling, networking and compute equipment. NVIDIA presents MaxLPS as a combination of dynamic power allocation, software techniques for improving output at a fixed power level, and site design based on 45 °C liquid-cooling inlet operation. The company describes the system as a way to increase throughput within an existing facility envelope, rather than as a way to increase the site's total power supply.

NVIDIA's Dynamic Power Software is currently in Developer Preview. According to the source, it models the data-center hierarchy from the utility connection to groups, racks, nodes and GPUs. Operators set power budgets, resource groups and policies; the software then compares allocated power with actual consumption and makes unused headroom available to other equipment within the same managed group. It also collects power telemetry and can respond to site-level events, maintenance conditions or emergency policies on a best-effort basis. NVIDIA says DSX Exchange, an open-source event bus that is also in Developer Preview, can connect the power software with building-management systems, electrical monitoring, cooling infrastructure, grid interfaces and compute schedulers, but says that exchange layer is not required for MaxLPS to function.

The source illustrates the problem with two examples. In a 100 MW power-budget waterfall, NVIDIA assigns 20 MW to facility overhead, 10 MW to rack losses and 10 MW to operational inefficiencies involving failures, restarts and checkpointing, leaving 60 MW for AI load. It separately describes a 540 kW site budget in which static provisioning strands 170 kW and dynamic provisioning enables another rack. These are presented as illustrative figures, not as measurements from a named operating data center. NVIDIA also reports representative inference evaluations in which provisioned rack power fell from 125 kW to 90 kW on GB200 NVL72 and from 136 kW to 101 kW on Vera Rubin NVL72. It says those changes preserved workload throughput, enabled 39% and 35% more racks respectively, and improved performance per watt by about 1.5 times and 1.3 to 1.4 times. The tested workloads were DeepSeek-R1 and Kimi-K2.5. The supplied source does not provide independent validation or full test protocols.

來源詳情: developer.nvidia.com

為什麼這很重要

AI data centers are increasingly constrained by electricity, cooling and physical infrastructure. If NVIDIA's claims hold in independent deployments, the system could increase useful AI output from existing power connections, but the source does not establish the costs, reliability or broader efficiency of the approach.

Electricity is becoming a direct constraint on the expansion of AI infrastructure. A data center can have land, network equipment and rack space available while still being unable to energize more GPUs because its utility connection, distribution equipment or cooling plant has reached its design limit. NVIDIA's proposal targets a specific inefficiency in that arrangement: racks are often provisioned for peak demand even though workloads move through phases such as compute bursts, memory-bound execution, synchronization, checkpointing, prefill and decode. If power can be safely shifted during those changes, a facility might produce more useful inference or training output without immediately securing another power connection.

The approach also shows why this is a facilities project rather than a software-only optimization. NVIDIA says sites should be designed around a fixed gross power envelope, then sized for power distribution, cooling, networking and rack positions that can support the intended long-term GPU count. The company recommends planning for 45 °C direct-liquid-cooling inlet operation for Vera Rubin NVL72 systems. Warmer coolant can allow dry coolers or other forms of free cooling to handle more of the annual heat-rejection load, reducing reliance on mechanical chillers in suitable climates. That potential depends on local weather, equipment sizing, redundancy and controls. Facilities that were not designed for the relevant temperatures, liquid loops or power flexibility may not be able to capture the claimed gains without significant retrofit work.

For operators, the practical attraction is optionality. A site could initially populate fewer racks while training-heavy workloads draw more power, then add hardware later if the workload mix shifts toward inference and average rack demand falls. That could change the timing of capital spending and facility expansion. It does not, by itself, reduce total electricity demand from AI or prove that fewer data centers will be built. More capacity inside an existing envelope could instead make additional AI services economically viable. The source names no customers, deployment scale, acquisition cost, measured annual energy savings, water consumption, failure rate or service-level impact. Its central capacity and performance claims remain NVIDIA claims based on representative evaluations.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

接下來看什麼

The key tests are whether Dynamic Power Software reaches general availability, whether independent operators reproduce the reported gains, and how the system behaves under failures, changing workloads and extreme weather. The required thermal and electrical design may limit adoption in existing facilities.

The first question is product maturity. Dynamic Power Software and DSX Exchange are both identified as Developer Preview, so their supported hardware, interfaces, operational guarantees and production availability remain unknown. Future documentation should clarify how policies are enforced, how quickly power can be redistributed, what happens when telemetry is delayed or wrong, and whether emergency actions are deterministic. NVIDIA says the control loop operates on a best-effort basis during power events, an important limitation for facilities that must maintain strict redundancy and service commitments. It will also matter whether the system can work with equipment and schedulers outside NVIDIA's own stack.

The second question is reproducibility. A useful comparison would publish the unmanaged static baseline, workload configuration, number of runs, throughput definition, latency, service error rate, GPU utilization and power-measurement boundaries. NVIDIA itself lists throughput, latency, service error rate, power draw, utilization and policy compliance as metrics for a scoped validation, but the supplied source does not provide those results in full. Independent testing across inference, training and workloads would show whether the reported improvements generalize beyond the two representative inference cases. Results should also be separated from gains caused by newer hardware, software versions, network changes or workload-specific tuning.

The third question is physical deployment. Operators will need evidence from sites with different climates, cooling architectures, utility constraints and reliability requirements. The 45 °C design may reduce chiller use in some locations but could require new heat-rejection equipment, controls or maintenance practices in others. Observers should track whether facilities can add the promised rack positions without expanding electrical distribution, cooling capacity or network infrastructure, and whether the system maintains throughput during hot weather, failures, restarts and checkpointing. Until those results are available, MaxLPS is best understood as a vendor-described infrastructure strategy and preview software, not an independently established industry-wide efficiency .

相關指引和測驗

人工智慧模型解釋人工智慧培訓AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語
覺得有用嗎?