Between 2024 and 2025 I delivered two AI supercomputers that entered the TOP500 list of the world's fastest machines — an Nvidia DGX training cluster of 1,272 GPUs ranked #25 globally and fastest in EMEA, and the AMD MI300X "Maximus" cluster ranked #20. One was delivered under United States export controls on restricted technology, into secure NOC and SCIF environments. I have also watched, across a longer career, what happens to infrastructure programs that scale faster than their controls. This article is about the difference.
The mistake is in the name
Organizations call these "AI infrastructure programs" and staff them like IT projects. But at supercomputing scale, the compute is almost the easy part. What you are actually running is four programs wearing one name: a global supply-chain program (GPUs and networking gear with brutal lead times and allocation politics); a construction and facilities program (power, cooling, structural load — megawatts, not racks); a compliance and security program (export controls, accreditation, sovereign requirements); and only then a technology integration program. Miss that framing and your governance covers a quarter of your risk.
What control looks like at this scale
One owner of the outcome. Hardware vendors, facilities contractors, network integrators, security teams — each will own their piece impeccably while the whole drifts. Someone has to own the reconciliation of all of it, with the authority to force trade-offs between a vendor's schedule and a facility's readiness. That accountability cannot be shared, rotated, or left to a steering committee.
Governance calibrated to the constraint that bites. Delivering restricted US-origin technology means export-control obligations shape who may touch what, where it may sit, and how it is documented — and a compliance failure does not slip the schedule, it ends the program. The discipline is to run governance tight enough to survive that audit while staying pragmatic enough to keep 1,272 GPUs moving to an immovable date. Tight on the things that end programs; light on the things that merely annoy engineers.
The immovable date as an ally. The TOP500 list publishes twice a year, whether your cluster is ready or not. A genuinely fixed external date is clarifying: it forces sequencing honesty, exposes wishful float, and gives every vendor the same non-negotiable reference point. Most programs lack one — which is why strong programs manufacture equivalent forcing functions and defend them.
Facilities on the critical path from day one. Data-centre power and cooling upgrades run on construction timescales, not IT timescales. I have delivered these upgrades alongside cluster builds — including for NOC and SCIF environments where security requirements multiply every lead time — and the programs that stay in control are the ones that put the facilities workstream at the top of the plan, not in an appendix.
Truth moving faster than the burn rate. At $1B scale with 100+ people across multiple vendors, a reporting lag of a month means decisions are being made on stale information while millions move. The reporting machinery has to run at the program's real cadence — which, today, means automating its assembly so human attention goes into judgment, not compilation.
Scaling without the drama
None of this is exotic. It is the same delivery discipline I learned on banking platforms and telecoms networks, applied at higher stakes and shorter timelines: know which of your four programs is the constraint this month; keep one accountable view across every vendor; let the immovable date do its work; and never let the reporting fall behind the spending. The machines that made the TOP500 list did so not because the technology was heroic — it was — but because the program around the technology stayed boring, in the best possible sense, all the way to the finish line.