🎓 Engineering Learning Studio

AI Data Center Engineering StudioPower, cooling, and network design for GPU-scale AI infrastructure

The engineering discipline behind the AI infrastructure buildout — rack power densities of 40-130+ kW, direct-to-chip and immersion liquid cooling, gigawatt-scale grid interconnection, and the GPU cluster network fabrics that tie it all together. Electrical, mechanical/cooling, network, and efficiency-metrics engineering for the facilities training and serving today's AI models.

PUELiquid CoolingGPU RacksInfiniBandRoCEGrid InterconnectionWUE
Start here
📖Studio Overview🗺️Interactive System Map
🎓

AI Data Center Engineering Professional Training Program

25

AI Data Center Engineering Professional Training Program

Premium Content

A full professional curriculum for GPU-scale AI infrastructure — power delivery architecture, cooling technologies, GPU cluster network fabric, rack density design, worked electrical and cooling calculations, redundancy and availability planning, drawings, installation and commissioning, DCIM monitoring, troubleshooting, and efficiency metrics. 25 modules from fundamentals through certification, 8 complete real-facility design packages (colocation retrofit, hyperscale campus, edge AI site, liquid-cooled retrofit, modular deployment, university cluster, enterprise private cluster, and multi-tenant colocation), a downloadable documentation kit, and a certificate of completion. One-time purchase, no account required.

Explore the Full Curriculum →
🧮

Engineering Calculators

3
PUE CalculatorPremium Content

Calculate Power Usage Effectiveness from IT/GPU load, cooling power, distribution losses, and auxiliary loads — with DCiE and a benchmark rating against typical AI/GPU facility PUE.

PUEDCiEFacility Efficiency
🔒 Open
Data Center Rack Power CalculatorPremium Content

Estimate kW/rack density from GPU server presets (H100, GB200 NVL, A100) or custom nameplate wattage, and check the load against PDU/circuit capacity per NEC continuous-load rules.

Rack DensityGPU ServersPDU Sizing
🔒 Open
Data Center Cooling Load CalculatorPremium Content

Convert IT/GPU electrical load into total cooling tons and BTU/hr, split between liquid and air cooling, and estimate chilled-water flow rate for CDU/CRAH sizing.

Cooling TonsChilled WaterLiquid/Air Split
🔒 Open
💡

Concept Explainers

12
❄️
Liquid Cooling vs. Air Cooling
Concept Explainer

Why a 10 kW GPU server can't simply be cooled by more fans — air runs out of heat-carrying capacity long before GPU racks do, which is why liquid cooling went from a niche HPC technique to the default AI infrastructure architecture.

Direct-to-ChipImmersion CoolingCDU
Explain This →
💧
PUE vs. WUE
Concept Explainer

A data center can post a great PUE while consuming millions of gallons of water a year — Power Usage Effectiveness and Water Usage Effectiveness measure two different resources, and optimizing one can work against the other.

PUEWUEEvaporative Cooling
Explain This →
🧊
Direct-to-Chip vs Immersion Cooling
Concept Explainer

Both are liquid cooling — but one clamps a cold plate onto the chip while air still cools everything else, and the other submerges the whole server, eliminating fans and air handling altogether.

Direct-to-ChipImmersion CoolingCold Plate
Explain This →
🌬️
Free Cooling vs Mechanical Cooling
Concept Explainer

"Free" doesn't mean no equipment runs — it means the compressor, the most energy-hungry part of the system, gets to switch off while outside conditions do the heavy lifting instead.

EconomizerChillerCompressor
Explain This →
Generator Sizing vs UPS Runtime
Concept Explainer

The UPS only has to bridge seconds — the gap until the generator starts and takes the load. Size the generator wrong, or the UPS battery too short, and that handoff is exactly where an outage becomes a real one.

UPSGeneratorATS
Explain This →
🌡️
Hot-Aisle vs Cold-Aisle Containment
Concept Explainer

Both chase the same goal — stop hot exhaust air from mixing back into the air servers breathe — but they trap opposite volumes, and picking the wrong one for an existing room wastes most of the benefit.

HACSCACSAirflow
Explain This →
🔁
N+1 vs 2N Redundancy
Concept Explainer

Both mean the facility keeps running after a failure — but N+1 survives losing one component, while 2N survives losing an entire independent power path, including everything upstream of it.

RedundancyUPSPower Path
Explain This →
🔌
PDU vs RPP
Concept Explainer

Both distribute power downstream of the UPS — but an RPP is a room-level circuit breaker panel, while a PDU is the rack-level device that actually terminates in the outlets a server plugs into.

PDURPPPower Distribution
Explain This →
📊
Rack Power Density vs Facility Power Capacity
Concept Explainer

A facility can have plenty of total substation capacity and still be unable to accept a new GPU rack — density isn't just a total-power problem, it's a per-square-foot and per-circuit problem.

Rack DensitykW per RackCapacity Planning
Explain This →
🔺
Single-Phase vs Three-Phase Power at the Rack
Concept Explainer

The same copper carries roughly 1.73× more power on three-phase than single-phase at the same voltage and current per conductor — exactly why high-density GPU racks abandon single-phase almost entirely.

Three-PhaseSingle-PhaseRack Power
Explain This →
🏛️
Tier III vs Tier IV Data Centers
Concept Explainer

The gap isn't "more redundant equipment" — it's whether a single unexpected failure can interrupt the load while the facility is already doing planned maintenance.

Uptime InstituteTier IIITier IV
Explain This →
🔋
Double-Conversion vs Line-Interactive UPS
Concept Explainer

One rebuilds utility power from scratch, all the time, so the load never sees a switch. The other rides through on utility power and only steps in when it must — a difference that shows up the instant a transfer happens.

UPSDouble-ConversionLine-Interactive
Explain This →
📋

Real Projects

1
Data Center Engineering Basics — Start to FinishPremium Content

A 2 MW GPU training hall carried from cooling topology through power distribution, UPS/generator architecture, redundancy, controls, and a finished cybersecurity review.

6 StepsReal FacilityStart to Finish
🔒 Open