Megatron-LM Guide: Megatron Core and Bridge (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

Megatron-LM is NVIDIA's open-source framework for training very large transformer models across many GPUs, and Megatron Core is the composable library inside it that provides the parallelism, kernels and mixed-precision building blocks. Reach for it when a model is too big for data parallelism alone and you need tensor, pipeline, context or expert parallelism on NVIDIA hardware.

This guide explains the pieces (Megatron-LM, Megatron Core, Megatron Bridge, NeMo), what each form of parallelism does, the install and quickstart commands from NVIDIA's docs, the performance NVIDIA publishes with its conditions, and how Megatron compares with DeepSpeed and FSDP. It belongs to our guide to LLM fine-tuning frameworks.

TL;DR

  • Megatron-LM is the reference trainer, Megatron Core is the library. The repo ships both. Core is what other frameworks, including NeMo's Megatron Bridge, build on.
  • Its edge is parallelism: tensor (TP), pipeline (PP), data (DP), expert (EP) and context (CP), with FP16, BF16, FP8 and FP4 precision support listed in the README.
  • NVIDIA reports up to 47% model FLOP utilization (MFU) on H100 clusters for models from 2B to 462B parameters, measured on the setup described below. It is NVIDIA's number, not ours.
  • It is overkill for a LoRA run on one GPU. For fine-tuning small and mid-size models, TRL, Unsloth or DeepSpeed ZeRO are simpler. Megatron pays off for pretraining and for large mixture-of-experts models.
  • Megatron Bridge converts Hugging Face checkpoints to and from Megatron format and offers pretraining, SFT and LoRA.

The pieces: Megatron-LM, Megatron Core, Bridge and NeMo

Per the NVIDIA/Megatron-LM README, the repository contains two components:

  • Megatron-LM is a reference example that bundles Megatron Core with pre-configured training scripts, aimed at research teams and quick experimentation.
  • Megatron Core is a composable library of GPU-optimized building blocks for custom training frameworks: transformer components, parallelism, mixed precision and model architectures.

A third project lives next to it:

  • Megatron Bridge (in the NVIDIA-NeMo organization) is a PyTorch-native library that provides pretraining, SFT and LoRA for language, vision-language, audio and multimodal models, and converts checkpoints in both directions between Hugging Face and Megatron Core. Its README describes it as part of the NeMo Framework and a refactor of the previous NeMo training stack. NVIDIA's NeMo documentation points to Megatron Bridge's docs for the framework container starting with the 26.02 release.

The license is Apache-2.0 for Megatron Bridge and Apache for the Megatron-LM repository per its README badge. At the time of writing, the README shows version 0.19.0 and the latest megatron-core on PyPI is 0.19.2 (September 18, 2026). PyPI states a Python requirement of 3.12 or later.

The origin: model parallelism for transformers

The original Megatron-LM paper (Shoeybi et al., 2019) described a simple way to split transformer layers across GPUs. Its abstract reports training models of up to 8.3 billion parameters on 512 GPUs with 15.1 PetaFLOPs sustained across the whole application and 76% scaling efficiency against a single-GPU baseline that sustained 39 TeraFLOPs (about 30% of peak). That is a 2019 result on 2019 hardware. We cite it for the idea, not as a current speed.

The five kinds of parallelism

Megatron Core supports TP, PP, DP, EP and CP. Short definitions, in our words:

ParallelismWhat it splitsWhy you use it
Data (DP)The batch. Every group holds a model replicaScale throughput once the model fits
Tensor (TP)Individual weight matrices inside a layerFit layers that are too large for one GPU; needs fast links such as NVLink
Pipeline (PP)Groups of layers into stagesFit very deep models across nodes with lower bandwidth needs than TP
Context (CP)The sequence dimensionTrain on long contexts that blow up activation memory
Expert (EP)The experts of a mixture-of-experts modelTrain MoE models whose experts do not fit together

The glossary has short entries for tensor parallelism, pipeline parallelism, data parallelism and mixture of experts.

In practice these are multiplied together. The number of GPUs you launch is split into groups: TP x PP x CP (and EP for MoE layers) define one model replica, and data parallelism fills the remainder. That arithmetic is our summary of how the parallel groups combine, so check the exact constraints in the Megatron Core docs for your version before sizing a job.

Install

From the README, the PyPI route and the source route:

uv pip install megatron-core
git clone https://github.com/NVIDIA/Megatron-LM.git
cd Megatron-LM
uv pip install -e .

The README adds that if a source build runs out of memory you can limit parallel jobs, for example MAX_JOBS=4 uv pip install -e .. It also points to an NGC container route in the Installation Guide. Megatron Bridge's README recommends the NeMo Framework container from NVIDIA's NGC catalog (nvcr.io/nvidia/nemo) with a version tag.

Quickstart

NVIDIA's Megatron Core quickstart gives two training commands. The first runs a minimal training loop on 2 GPUs with mock data:

torchrun --nproc_per_node=2 examples/run_simple_mcore_train_loop.py

The second runs the repo's LLaMA-3 8B example script with FP8 on 8 GPUs (also with mock data):

./examples/llama/train_llama3_8b_h100_fp8.sh

Inside that script, per the file in the repo, the launch is a single-node torchrun with 8 processes, and the flags show what a Megatron job looks like:

--tensor-model-parallel-size $TP_SIZE
--context-parallel-size $CP_SIZE
--sequence-parallel
--fp8-format hybrid
--fp8-amax-history-len 1024
--fp8-amax-compute-algo max
--fp8-param-gather

In that script TP_SIZE and CP_SIZE are both 1 and the pipeline size line is commented out, so it is a starting point to edit, not a tuned recipe. For FP8 and FP4 background, see FP8, FP4 and our Transformer Engine FP8 post.

Use Hugging Face checkpoints through Megatron Bridge

Most teams start from a Hugging Face model, not a Megatron checkpoint. Megatron Bridge's README describes this flow:

  1. Load a bridge with AutoBridge.from_hf_pretrained(...) from a Hub ID or local path.
  2. Convert it with to_megatron_provider(), set the tensor and pipeline parallel sizes, then call finalize().
  3. Build the model with provide_distributed_model(wrap_with_ddp=False).
  4. Export back with save_hf_pretrained(...) for a full Hugging Face folder, or export_hf_weights(...) to stream weights.

The README lists support for many families, including Llama 3.x, Qwen 2.5/3/3.5, DeepSeek V3/V4, GLM-4.5 to GLM-5.x, Gemma 3/4, Nemotron-3, Kimi K2 and GPT-OSS, with some older models (Gemma 1/2, Llama 2, Mistral 7B and others) deprecated. It also notes that Python 3.10 support is dropped from version 0.4.0. Check the current list before you plan around a specific model.

Performance NVIDIA publishes

These figures come from the Megatron-LM README. They are NVIDIA's, and each has stated conditions.

  • Overall: up to 47% MFU on H100 clusters, for models from 2B to 462B parameters. The 462B model was benchmarked on 6,144 H100 GPUs.
  • Weak scaling: MFU rises from 41% for the smallest model to 47 to 48% for the largest, which NVIDIA attributes to larger matrix multiplies having higher arithmetic intensity.
  • Strong scaling: GPT-3 (slightly over 175B parameters) scaled from 96 to 4,608 H100 GPUs at a fixed batch size of 1,152 sequences. MFU fell from 47% to 42% as communication became more exposed.
  • Conditions: vocabulary size 131,072, sequence length 4,096, throughput measured end to end (data loading, optimizer steps, communication, logging), and without training to convergence. The runs used the overlap flags --overlap-grad-reduce, --overlap-param-gather and --tp-comm-overlap, with pipeline overlap on by default.

The README points to the Megatron Bridge Performance Summary for the latest numbers. We have not reproduced any of these, and your MFU on a different model, sequence length or interconnect will differ.

Megatron vs DeepSpeed vs FSDP

Megatron-LM / CoreDeepSpeedPyTorch FSDP
Main ideaSplit the model (TP, PP, CP, EP) plus DPShard redundant state (ZeRO), plus other featuresShard parameters, gradients and optimizer states across DP workers
HardwareNVIDIA GPU-optimizedNVIDIA and AMD listedPyTorch-native
Code changeModels built from Megatron blocks, or converted via BridgeWrap with deepspeed.initialize or a Trainer configWrap the module
Best forPretraining and large MoE at cluster scaleFine-tuning and mid-scale trainingPyTorch-native sharding
LicenseApacheApache-2.0Part of PyTorch

Rules of thumb:

  • One node or a few GPUs, fine-tuning an existing model: ZeRO-2 or FSDP, usually through TRL or LLaMA-Factory. Combine with LoRA to cut memory.
  • Dozens to thousands of GPUs, pretraining or continued pretraining: Megatron Core, where TP and PP keep per-GPU memory and communication manageable.
  • Large MoE models: expert parallelism is a first-class feature in Megatron Core.

GPU requirements

The Megatron-LM README lists FP8 and FP4 precision support, and the FP8 quickstart script is named for H100. Hopper and Blackwell GPUs are the natural targets. Tensor parallelism is communication-heavy, so keep TP groups inside a single node with NVLink and use pipeline or data parallelism across nodes. See H100 vs H200 and the H100 and B200 pages for memory and bandwidth specs.

Run it on a cloud GPU

Start with the 2-GPU torchrun example above to check your drivers and container, then scale up to a full 8-GPU node.

FAQ

What is the difference between Megatron-LM and Megatron Core?

Megatron-LM is a reference example with ready-made training scripts. Megatron Core is the composable library of parallelism, kernels and model components that the example and other frameworks build on.

Is Megatron-LM the same as NeMo?

No. NeMo is a broader NVIDIA framework. Its Megatron Bridge library builds on Megatron Core and, per its README, refactors the previous NeMo training stack into a PyTorch-native loop.

Can I fine-tune with Megatron-LM?

Yes. Megatron Bridge provides pretraining, SFT and LoRA. For small models on a single GPU, simpler tools are usually quicker to set up.

Does Megatron-LM work with Hugging Face models?

Through Megatron Bridge, which converts Hugging Face checkpoints into Megatron format and exports them back. Check its README for the currently supported model families.

Does it run on AMD GPUs?

NVIDIA's README describes Megatron Core as GPU-optimized, and its benchmarks and FP8 and FP4 notes are for NVIDIA hardware. We did not find AMD support stated there, so verify it before planning on it.

Sources

#fine-tuning#llm training#megatron-lm#nvidia nemo#model parallelism#pretraining

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.