DeepSpeed Guide: ZeRO Stages, Offload and FSDP (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

DeepSpeed is Microsoft's open-source library for training large models across many GPUs, and its core idea is ZeRO (Zero Redundancy Optimizer), which splits optimizer states, gradients and parameters across GPUs instead of copying them onto every one. If a model does not fit on one card with plain data parallelism, ZeRO is usually the first thing to try, either through DeepSpeed itself or through PyTorch FSDP, which borrows the same idea.

This guide covers what each ZeRO stage shards, the memory math from the ZeRO paper, working configs and launch commands from the DeepSpeed and Hugging Face docs, offload to CPU and NVMe, and how DeepSpeed compares with FSDP and Megatron-LM. It is part of our guide to LLM fine-tuning frameworks.

TL;DR

  • Use ZeRO-2 first. It shards gradients and optimizer states and costs less communication than ZeRO-3. Hugging Face's docs say to use ZeRO-3 only when the model does not fit with ZeRO-2.
  • ZeRO-3 shards parameters too, so memory per GPU keeps falling as you add GPUs, at the price of more all-gathers on every forward and backward pass.
  • Offload is the escape hatch. offload_optimizer and offload_param push state to CPU memory (or NVMe with ZeRO-Infinity) so a model can train on fewer GPUs, slower.
  • FSDP is the PyTorch-native alternative and is documented as inspired by ZeRO stage 3. DeepSpeed still has more built-in features (offload tiers, MoE, Ulysses sequence parallelism).
  • For pretraining at the largest scale, look at Megatron-LM, which adds tensor and pipeline parallelism that ZeRO alone does not provide.

What DeepSpeed is

DeepSpeed is an Apache-2.0 licensed deep learning optimization library maintained at github.com/deepspeedai/DeepSpeed. Its README lists ZeRO and ZeRO-Infinity, 3D parallelism, Ulysses sequence parallelism, DeepSpeed-MoE, offloading (including SuperOffload and ZenFlow), DeepNVMe, DeepCompile, AutoTP, universal checkpointing, and compression and quantization. At the time of writing, the latest PyPI release is 0.19.7 (September 16, 2026).

Per the README, the most-tested GPUs are NVIDIA Pascal, Volta, Ampere and Hopper, and AMD MI100 and MI200. PyTorch 2.0 or later is recommended, and a CUDA (nvcc) or ROCm (hipcc) compiler is needed to build its C++ and CUDA extensions.

You can use DeepSpeed in two ways:

  1. Directly, by calling deepspeed.initialize in your own training loop.
  2. Through a framework, such as Hugging Face Transformers' Trainer, Accelerate, TRL, or LLaMA-Factory, which read a DeepSpeed JSON config. LLaMA-Factory's README, for instance, lists DeepSpeed as an optional dependency.

How ZeRO works

In ordinary data parallelism every GPU holds a full copy of the model, its gradients and its optimizer states, then averages gradients each step. Most of that memory is redundant. ZeRO removes the redundancy in three steps, which the ZeRO paper calls Pos, Pos+g and Pos+g+p.

StageWhat is partitioned across GPUsPaper's name
ZeRO-1Optimizer statesPos
ZeRO-2Optimizer states and gradientsPos+g
ZeRO-3Optimizer states, gradients and parametersPos+g+p

The DeepSpeed ZeRO tutorial describes the same three stages: in stage 1 the optimizer states are partitioned so each process updates only its partition, in stage 2 the reduced 16-bit gradients are also partitioned, and in stage 3 the 16-bit model parameters are partitioned.

The memory math

The paper uses mixed-precision training with Adam. For a model with Ψ parameters, that is 2Ψ bytes for FP16 parameters, 2Ψ for FP16 gradients, and KΨ for optimizer states, with K = 12 (an FP32 copy of the parameters, momentum and variance). Total: 16Ψ bytes per GPU with plain data parallelism.

The paper's worked example is a 7.5B parameter model on 64 GPUs (data parallel degree Nd = 64). These numbers are the paper's, from its Figure 1 and Table 1:

SetupModel-state memory per GPU
Standard data parallelism120 GB
ZeRO-1 (Pos)31.4 GB
ZeRO-2 (Pos+g)16.6 GB
ZeRO-3 (Pos+g+p)1.88 GB

The paper summarizes the three stages as reducing model-state memory per process by up to 4x, 8x and Nd respectively. It also states that with Nd = 64, ZeRO can train models of up to 7.5B, 14B and 128B parameters with Pos, Pos+g and Pos+g+p.

These figures cover model states only. Activations, temporary buffers and fragmentation come on top (the paper handles those with a separate set of techniques it calls ZeRO-R).

A computed example for fine-tuning. Using the same 16 bytes per parameter rule (computed here, not measured), full fine-tuning of a 70B model needs about 70 billion x 16 = 1,120 GB of model states. Sharded with ZeRO-3 across 8 GPUs, that is 140 GB per GPU, which does not fit on an 80 GB card even before activations. Across 16 GPUs it is 70 GB per GPU, which is tight. This is why large full fine-tunes either use many GPUs, enable offload, or switch to a parameter-efficient method such as LoRA.

Install

From the DeepSpeed README and the Hugging Face DeepSpeed guide:

pip install deepspeed
ds_report

ds_report shows which DeepSpeed ops your machine supports. If you use Transformers, pip install transformers[deepspeed] pulls it in. The Hugging Face guide notes that installing from source is the more reliable option if you hit CUDA-related install errors, because it matches your exact hardware.

Use it with your own training loop

The DeepSpeed getting-started tutorial wraps a PyTorch model in an engine that handles distributed setup and mixed precision:

model_engine, optimizer, _, _ = deepspeed.initialize(args=cmd_args,
                                                     model=model,
                                                     model_parameters=params)

for step, batch in enumerate(data_loader):
    loss = model_engine(batch)
    model_engine.backward(loss)
    model_engine.step()

Launch it with the deepspeed launcher. For several machines, list them in a hostfile (hostname plus the number of GPU slots):

worker-1 slots=4
worker-2 slots=4
deepspeed --hostfile=myhostfile <client_entry.py> <client args> \
  --deepspeed --deepspeed_config ds_config.json

If you pass no hostfile, DeepSpeed looks for /job/hostfile, and if that is missing it uses the GPUs on the local machine.

Use it with Hugging Face Trainer

Most fine-tuning scripts use Trainer, which takes the config through TrainingArguments:

from transformers import TrainingArguments

args = TrainingArguments(
    deepspeed="path/to/deepspeed_config.json",
    ...
)

The Hugging Face guide gives three equivalent ways to launch:

# DeepSpeed launcher
deepspeed --num_gpus 4 train.py

# torchrun
torchrun --nproc_per_node 4 train.py

# Accelerate
accelerate launch --num_processes 4 train.py

One gotcha from the guide: Accelerate ignores the deepspeed argument in TrainingArguments, so with Accelerate you point at the config from an Accelerate config file (distributed_type: DEEPSPEED plus deepspeed_config_file).

Set batch size and accumulation to "auto" in the JSON. The guide warns that if you hard-code values that disagree with TrainingArguments, training continues silently with the wrong values.

Working configs for each stage

These are the starting-point configs from the Hugging Face DeepSpeed guide.

ZeRO-2:

{
  "bf16": { "enabled": "auto" },
  "zero_optimization": {
    "stage": 2,
    "overlap_comm": true,
    "allgather_bucket_size": 2e8,
    "reduce_bucket_size": 2e8,
    "contiguous_gradients": true
  },
  "gradient_clipping": "auto",
  "train_micro_batch_size_per_gpu": "auto",
  "train_batch_size": "auto",
  "gradient_accumulation_steps": "auto"
}

ZeRO-3 with CPU offload:

{
  "bf16": { "enabled": "auto" },
  "zero_optimization": {
    "stage": 3,
    "overlap_comm": true,
    "contiguous_gradients": true,
    "reduce_bucket_size": "auto",
    "stage3_prefetch_bucket_size": "auto",
    "stage3_param_persistence_threshold": "auto",
    "stage3_gather_16bit_weights_on_model_save": true,
    "offload_optimizer": { "device": "cpu", "pin_memory": true },
    "offload_param": { "device": "cpu", "pin_memory": true }
  },
  "gradient_clipping": "auto",
  "train_micro_batch_size_per_gpu": "auto",
  "train_batch_size": "auto",
  "gradient_accumulation_steps": "auto"
}

What the important keys do, per the guide:

  • overlap_comm: true hides all-reduce latency behind the backward pass.
  • allgather_bucket_size and reduce_bucket_size trade communication speed for GPU memory. Lower values use less memory but slow communication.
  • offload_optimizer moves the optimizer to CPU memory. offload_param does the same for parameters and is ZeRO-3 only. pin_memory: true speeds up CPU-GPU transfers but locks RAM that other processes cannot use.
  • stage3_gather_16bit_weights_on_model_save: true all-gathers the shards before saving, so you can save a consolidated 16-bit model. It is a ZeRO-3 setting.

Two ZeRO-3 pitfalls worth knowing. First, ZeRO-3 shards parameters while the model loads, so you must create TrainingArguments before loading the model. If the model is already fully on each GPU before DeepSpeed is configured, no memory is saved. Second, DeepSpeed writes sharded checkpoints that from_pretrained() cannot load directly. Call trainer.save_model(...) at the end to get a normal Transformers checkpoint, and see DeepSpeed's Universal Checkpointing guide if you need to resume on a different parallelism layout.

Offload: ZeRO-Offload and ZeRO-Infinity

The DeepSpeed tutorial separates two offload systems:

  • ZeRO-Offload moves optimizer and gradient states to CPU memory within ZeRO-2.
  • ZeRO-Infinity is the offload engine for ZeRO-3 and can push state to CPU or NVMe memory.

Offload trades speed for capacity. It lets a single node train something that otherwise needs a cluster, but the CPU-GPU transfer is slower than staying in GPU memory, and the Hugging Face guide notes that every extra stage and offload level lowers peak memory at the cost of more communication. Use it when you are memory-bound and cannot add GPUs.

DeepSpeed vs FSDP

PyTorch's FullyShardedDataParallel (FSDP) wraps a module and shards its parameters across data-parallel workers. Its documentation says the design is inspired by Xu et al. and by ZeRO stage 3 from DeepSpeed. See the FSDP glossary entry for the short version.

FSDP's ShardingStrategy options map closely onto ZeRO stages, per the PyTorch docs:

FSDP strategyWhat is shardedClosest ZeRO analogue
FULL_SHARD (default)Parameters, gradients, optimizer statesZeRO-3
SHARD_GRAD_OPGradients and optimizer statesZeRO-2
NO_SHARDNothing (like DistributedDataParallel)Plain data parallelism
HYBRID_SHARDFull sharding inside a node, replication across nodesNo direct stage

The mapping column is our reading of the docs, not a table from PyTorch.

How to choose:

  • Pick FSDP if you want to stay inside plain PyTorch, with no extra library or JSON config, and your model fits the FULL_SHARD or SHARD_GRAD_OP pattern. HYBRID_SHARD is useful when inter-node bandwidth is the bottleneck.
  • Pick DeepSpeed if you need NVMe offload, DeepSpeed-MoE, Ulysses sequence parallelism or its other bundled features, or if the framework you use ships DeepSpeed configs as its main multi-GPU path.
  • Either works through Hugging Face Accelerate and TRL, so switching is usually a config change rather than a rewrite.

DeepSpeed vs Megatron-LM

They solve different parts of the problem. ZeRO shards redundant state across data-parallel workers and works with any PyTorch model with little code change. Megatron-LM splits the model itself with tensor, pipeline, context and expert parallelism, and needs models written with its building blocks (or converted through its Bridge). They can be combined, and the DeepSpeed README lists 3D parallelism as a feature. Fine-tuning a single model of up to a few tens of billions of parameters usually calls for ZeRO or FSDP. Pretraining across hundreds of GPUs is where Megatron-style parallelism earns its complexity.

GPU memory and hardware notes

  • ZeRO-2 and ZeRO-3 speed depends heavily on the interconnect, because each step adds reduce-scatter and all-gather traffic. On a single node with NVLink, such as an 8-GPU H100 box, this is cheap. Across slow links, ZeRO-3 communication becomes the bottleneck. See NCCL and data parallelism in the glossary.
  • Sharding helps with model states, not activations. Combine it with activation checkpointing and a smaller micro-batch if you run out of memory.
  • For a single GPU with an 8B-class model, LoRA or QLoRA without DeepSpeed is usually simpler. See the LoRA guide and how much VRAM do I need for LLMs.

DeepSpeed does not publish a single throughput figure we can quote without its test conditions, so we give none here. Run your own short benchmark on your model and sequence length before committing to a stage.

Run it on a cloud GPU

ZeRO needs more than one GPU to shine, so rent an 8-GPU node, pick the ZeRO stage from the table above, and compare step time with and without offload.

FAQ

What is the difference between ZeRO-1, ZeRO-2 and ZeRO-3?

ZeRO-1 partitions optimizer states, ZeRO-2 adds gradients, and ZeRO-3 adds the model parameters. Each stage uses less memory per GPU and more communication than the one before it.

Which ZeRO stage should I start with?

Start with ZeRO-2. Hugging Face's guide says ZeRO-2 has lower communication overhead than ZeRO-3 and recommends ZeRO-3 only when the model does not fit across your GPUs with ZeRO-2.

Is DeepSpeed better than FSDP?

Neither is better in general. FSDP is native PyTorch and was inspired by ZeRO-3. DeepSpeed has more built-in features such as NVMe offload and MoE support. Both integrate with Accelerate and TRL.

Can DeepSpeed train on a single GPU?

Yes, with offload. ZeRO-Offload and ZeRO-Infinity move optimizer state, gradients or parameters to CPU or NVMe, so a model that exceeds GPU memory can still train, more slowly. Sharding itself needs at least two GPUs to save memory.

Does DeepSpeed work with Hugging Face Transformers?

Yes. Pass a JSON config to TrainingArguments(deepspeed=...) or use an Accelerate config file, then launch with deepspeed, torchrun or accelerate launch.

Does DeepSpeed run on AMD GPUs?

The README lists a ROCm (hipcc) compiler as a build option and names AMD MI100 and MI200 among the most-tested GPUs.

Sources

#fine-tuning#llm training#deepspeed#zero#distributed training#fsdp

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.