Seven terms decide almost everything about who can compete where in this stack, and what they are allowed to earn for it — three from the compute half, four from the physical half. Learn these seven and the entire value chain organises itself.
A reference design — Nvidia calls its own versions MGX and HGX — is a complete, prescriptive blueprint: chassis dimensions, power-delivery topology, cooling-loop layout and an approved component list, handed to a manufacturing partner to build almost exactly as specified. Custom engineering is the opposite: a hyperscaler or systems house designs its own rack, cooling loop and network topology from first principles and captures whatever performance edge that design delivers.
Nvidia's own NVQual stress-testing regime — covering GPU stability under sustained load, thermal and throttling behaviour, power delivery and PCIe link integrity — must be passed before a system can carry the NVIDIA-Certified Systems label, and must be passed again on a roughly annual cadence as each new chip architecture ships.
A server is not sold once it is built correctly; it is sold once its builder has actually been allocated the chips to put inside it. Nvidia's own supply, not manufacturing capacity, is the scarce input at this layer.
The Uptime Institute's Tier classification and the telecom-industry TIA-942 standard both certify how much redundant, concurrently-maintainable power and cooling capacity a facility has. Every rung up this ladder is a direct, engineered order for more transformers, switchgear and chiller capacity than the facility's raw IT load would otherwise require.
N+1 means one spare unit beyond peak load; 2N means a fully duplicated, independent second system. A hyperscale campus targeting Tier III/IV typically specifies 2N on its most critical power path.
A transformer's capacity is quoted in MVA (megavolt-amperes); GIS (gas-insulated switchgear) is the compact alternative to open-air switchgear used where space or reliability requirements are highest. Reading a power-equipment company's order book in MVA and GIS-bay terms is the closest this industry gets to reading a chip company's wafer-capacity numbers.
Power Usage Effectiveness — the ratio of total facility power draw to power actually delivered to IT equipment — is the number that reveals how much of a data centre's power bill is pure computing versus cooling and conversion overhead.
| Term | Layer | What it does NOT tell you |
|---|---|---|
| Reference design / NVQual / Allocation | Compute | Whether the builder has added any differentiated engineering, or whether this quarter's allocation implies next quarter's |
| Tier III/IV, TIA-942 | Physical | Whether the equipment inside came from an India-listed supplier, or was imported |
| N+1 / 2N | Physical | Whether that duplicate capacity was sourced competitively (commodity) or from a scarce, qualified supplier (engineered) |
| MVA / GIS / PUE | Physical | Which specific vendor is responsible for a given efficiency gain — attribution is rarely disclosed |
Source: Dart Consultants, from Nvidia's own certification-programme documentation, Uptime Institute/TIA-942 standard summaries, and industry-standard MVA/GIS/PUE definitions — see Notes for full citation list.
If a company's revenue comes from meeting a specification someone else already wrote — Nvidia's reference design, a generic building code, a catalog transformer spec — assume the margin is thin and under continuous pressure, no matter how large the order book looks. If it comes from something genuinely scarce — chip allocation, a constrained global manufacturing capacity, or engineering no published spec covers — assume that is where this report's attention, and the real moat, actually sits. This single rule of thumb is the only one needed for every layer of this stack.