Inside an AI Factory – The New Generation of Data Centres

TECHNICAL GUIDE

Why Specialist Data Center Cleaning Matters for Reliability

🗒 23 September 2026 •  ⏱ 9 min read

An AI chatbot can produce an answer in seconds. To the person using it, the process feels almost weightless: type a question, press send and receive a response. Behind large cloud-based AI services, however, sits a physical computing system whose performance depends on how well its hardware and infrastructure work together.

 

The term AI factory describes infrastructure organized around developing and running AI. For data center owners, facility managers and businesses in Malaysia, understanding this concept starts with a practical question: what must the building and its systems deliver to support the intended AI workload?

01 — what is an ai factory?

NVIDIA uses the term AI factory for specialized computing infrastructure supporting the AI lifestyle, from preparing data through training, fine-tuning and inference. Its description includes computing hardware, networking, storage and software. The “factory” analogy describes turning data and computation into useful Ai outputs, it should not be read as a claim that the building itself thinks. [NVIDIA: What Is an AI Factory?].

Three activities help explain that lifecycle: 

  • Training : Adjusting a model’s parameters using data so it learns patterns relevant to a task.
  • FIne-tuning : Further training an existing model for a particular task, domain or behaviour.
  • Inference: Using a trained model to generate outputs, such as predictions or responses.

 These activities belong to the AI lifecycle, but a particular deployment may focus on only some of them. The facility brief should identify the intended work rather than relying on the AI factory label alone. [NVIDIA’s AI factory overview]

It is also inaccurate to say that every AI response requires thousands of processors in a remote building. Apple, for example, has documented both on-device and server-based foundation models. Some inference can happen locally on a user’s device. Large AI facilities are important to the wider ecosystems, but they are not the execution environment for every AI interaction. [Apple: Introducing Apple’s On-Device and Server Foundation Models] 

02 — how large can ai computing become?

Large-scale AI infrastructure is more than a theoretical concept. In March 2024, Meta described two clusters containing 23,576 NVIDIA H100 GPUs each, supporting its AI work, including Llama 3. This is a documented example of substantial computing scale, not a minimum requirement for an AI facility. [Meta: Building Meta’s GenAI Infrastructure].

GPUs are not only hardware used for AI. Apple’s technical account describes training support across both GPUs and Tensor Processing Unit, or TPUs. A facility’s requirements should therefore follow its selected computing platform and workloads, rather than an assumption that every AI system has the same processor architecture. [Apple’s foundation-model technical overview].

The useful planning distinction is between a specific, measured deployment and a general claim about AI. Processor count alone cannot establish a project’s power requirement, completion time or business value.

03 — power: capacity must be available where it is needed

Electricity is a major consideration in AI infrastructure development. In its 2025 Energy and AI report, the International Energy Agency estimated that data centers consumed approximately 415 TWH in 2024 and projected around 945 TWH in 2030. Those figures cover data centers overall, not AI alone, and the 2030 figure is a projection rather than measured consumption. The report identifies AI as a major driver of growth and highlights grid constraints as a potential obstacle to new projects. [EIA: Energy and AI – Executive Summary] 

For a facility team, the practical response is to request a power plan tied to the proposed equipment and deployment phases. Establish the expected operating load, the capacity available to each rack and how the facility will respond to interruptions or maintenance.
 
Avoid assuming that an AI facility always consumes more electricity than every conventional data centre. Comparisons need a defined scale and workload. A useful project assessment distinguishes connection capacity, actual demand and annual energy consumption.
 

04 —Cooling: match the solution to the computing platform

Some AI platforms are explicitly designed for liquid cooling. NVIDIA’s GB200 NVL72, for example, combines 72 Blackwell GPUs and 36 Grace CPUs in a liquid-cooled rack-scale system. This illustrates how the computing platform can dictate an important part of the facility design. It does not establish that every AI installation requires the same cooling arrangement. [NVIDIA: GB200 NVL72].

 

Liquid cooling also does not necessarily eliminate air cooling. A US Department of Energy case study of Sandia’s Attaway supercomputer describes a system in which liquid cooling handles most of the cooling duty while air cools other components. The case study also documents coolant distribution and redundancy arrangements. Although this is an HPC example rather than an AI factory specification, it demonstrates why cooling must be considered as a complete system. [US Department of Energy: Sandia’s Liquid-Cooled Data Center].

 

For a proposed AI installation, ask the design team to explain the heat-removal path, remaining air-cooling needs, monitoring, maintenance access and response to a cooling failure. Evaluate energy and water performance for the actual design; the label “liquid-cooled” alone is insufficient evidence of either outcome.

05 —Connectivity: the internal network matters

An AI cluster needs more than an internet connection. Its internal network must support communication between the systems performing the workload.

 

Meta’s 2024 cluster description provides a concrete example: one cluster used an Ethernet-based RoCE network and the other used InfiniBand, both with 400 Gbps endpoints. Meta also described coordinating network design, software and workload placement to improve large-cluster performance. These are examples of engineering choices, not universal specifications for all AI projects. [Meta’s infrastructure design account].

 

For procurement, ask how the proposed network will be tested under the intended workload. A headline link speed should be accompanied by an explanation of expected application performance, resilience and expansion options.

06 —Storage and software complete the system

Power, cooling and networking are essential, but the computing environment also needs an effective way to supply data and preserve progress. Meta’s account describes storage designed for data loading and coordinated checkpoint operations across thousands of GPUs. Checkpoints preserve training state so work can be resumed. [Meta: Storage for GenAI Infrastructure].
 
A practical acceptance plan should therefore test the complete workload. Include data access, job scheduling, monitoring and recovery in the discussion, rather than accepting a facility solely because individual components meet their specifications.

07 — What this means for Malaysia

Malaysia has attracted major cloud and AI infrastructure commitments. On 2 May 2024, Microsoft announced a US$2.2 billion investment over four years, including cloud and AI infrastructure, skills development and other initiatives. This is evidence of an announced commitment; it should not be presented as proof that the full amount has already been spent or that every associated facility is operational. [Microsoft: Malaysia Cloud and AI Investment Announcement].

 

Sustainability is also part of the policy context. MIDA’s published *Guideline for Sustainable Development of Data Centre* addresses power, carbon and water usage effectiveness. The document connects its conditions to applications under the DESAC incentive scheme. It should not be treated as evidence that every listed target is a universal legal requirement for every Malaysian data center. [MIDA: Guideline for Sustainable Development of Data Centre].

 

The planning implication is that investment value alone is an incomplete measure of success. Project teams should also consider infrastructure readiness, resource use, operational skills and how local organisations will access and benefit from the computing capacity. Economic benefits are opportunities to develop and measure, rather than automatic consequences of constructing a building.