Ir para o conteúdo
Atlas de IA

Infraestrutura de treinamento e inferência

A memória decide o tamanho do modelo treinável; a comunicação, quanto tempo leva

08 Engenharia, segurança e ética de IAEspecialistaEntrada 1 deste domínio

O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.

DEFINIÇÃO

Training & inference infrastructure is the hardware and software stack that supplies compute, memory and communication to large neural networks. It is built from accelerators (GPUs/TPUs), interconnect fabrics (NVLink/InfiniBand), parallelism strategies (data, tensor and pipeline parallelism) and memory optimisations (ZeRO, mixed precision, activation recomputation). Together they determine whether a model can be trained at all, at what cost, and how efficiently it can be served.

Intuição

Think of training a large model as a crew raising a building: compute is how fast the workers build, memory is how much material the site can hold, and communication is the overhead of workers shouting to stay in step. The real bottleneck is often not slow hands but a site too small to store material, or crews spending more time coordinating than building — adding hands does not fix a site that is too small.

Fig. 1

Layers of parallelism: the more intra-layer the strategy, the more frequent the communication; real runs stack several into 3D parallelism

Lower layers communicate more densely and demand more network bandwidth and lower latency.
Fig. 2

Memory breakdown of a 7B-parameter model under mixed-precision Adam (illustrative magnitude): optimiser states alone outweigh the weights, and ZeRO sharding lowers the total step by step

  • DDP (full replica)
  • ZeRO-2
  • ZeRO-3

Como funciona

  1. 01

    Shard: make the model fit at all

    When one device cannot hold the whole model, the first move is to cut it up — by layer, by weight matrix, by tensor dimension. Mixed-precision Adam costs roughly 16 bytes per parameter (weights, gradients and optimiser states combined), and for a 7B model the optimiser states alone outweigh the parameters, so sharding is almost unavoidable.

  2. 02

    Parallelise: spread the pieces across devices

    Data parallelism splits the batch and keeps a full replica per device; tensor parallelism splits the matrices inside a single layer; pipeline parallelism places different layers on different devices. They compose into 3D parallelism. The more intra-layer the parallelism, the more frequent the communication and the harsher the network demands.

  3. 03

    Save memory: ZeRO and mixed precision

    ZeRO shards optimiser states, then gradients, then parameters, trading communication for memory and roughly dividing per-device usage by the device count. Mixed precision (BF16 compute with FP32 master weights) halves compute and storage while keeping numerics stable. Activation recomputation buys back one layer of activations with an extra forward pass.

  4. 04

    Communicate: hide transfers behind compute

    Time in large-scale training is often spent on synchronisation, not arithmetic. Compute–communication overlap, gradient accumulation and better topologies (NVLink, hierarchical All-Reduce) hide transfers inside compute windows — which is why "doubling the GPUs" frequently fails to halve the time.

Onde é usado

  • Pre-training: training hundred-billion-parameter models on thousand-GPU clusters, where capacity and scheduling decide feasibility
  • Fine-tuning and adaptation: LoRA and FSDP on a handful of GPUs for downstream tasks
  • Inference serving: KV cache and batching together decide how many users one machine can serve
  • Cost and capacity planning: estimate the bill for one training run and the cost per thousand tokens served

Equívocos comuns

  • More GPUs do not mean proportionally more speed. When communication, synchronisation or data loading is the bottleneck, adding devices simply adds idle waiting and lowers utilisation.
  • Most memory is not the parameters. Under mixed-precision Adam, optimiser states and gradients can exceed the weights themselves; reasoning only from the parameter count badly underestimates the footprint.
  • Mixed precision is not free speed. BF16 has a wide range but low precision, so some numerically sensitive operations still have to fall back to FP32 or they overflow or drift.

Termos-chave

Accelerator
A high-throughput parallel unit such as a GPU or TPU
Tensor parallelism
Splitting a single layer’s large matrices across devices
Pipeline parallelism
Placing different layers on different devices and filling bubbles with micro-batches
ZeRO
Sharding optimiser states, gradients and parameters to cut per-device memory

Leituras complementares