Megatron-LM is NVIDIA's open-source framework for training very large transformer models across many GPUs, and Megatron Core is the composable library inside it that provides the parallelism, kernels and mixed-precision building blocks. Reach for it when a model is too big for data parallelism alone and you need tensor, pipeline, context or expert parallelism on NVIDIA hardware.
This guide explains the pieces (Megatron-LM, Megatron Core, Megatron Bridge, NeMo), what each form of parallelism does, the install and quickstart commands from NVIDIA's docs, the performance NVIDIA publishes with its conditions, and how Megatron compares with DeepSpeed and FSDP. It belongs to our guide to LLM fine-tuning frameworks.
TL;DR
- Megatron-LM is the reference trainer, Megatron Core is the library. The repo ships both. Core is what other frameworks, including NeMo's Megatron Bridge, build on.
- Its edge is parallelism: tensor (TP), pipeline (PP), data (DP), expert (EP) and context (CP), with FP16, BF16, FP8 and FP4 precision support listed in the README.
- NVIDIA reports up to 47% model FLOP utilization (MFU) on H100 clusters for models from 2B to 462B parameters, measured on the setup described below. It is NVIDIA's number, not ours.
- It is overkill for a LoRA run on one GPU. For fine-tuning small and mid-size models, TRL, Unsloth or DeepSpeed ZeRO are simpler. Megatron pays off for pretraining and for large mixture-of-experts models.
- Megatron Bridge converts Hugging Face checkpoints to and from Megatron format and offers pretraining, SFT and LoRA.
The pieces: Megatron-LM, Megatron Core, Bridge and NeMo
Per the NVIDIA/Megatron-LM README, the repository contains two components:
- Megatron-LM is a reference example that bundles Megatron Core with pre-configured training scripts, aimed at research teams and quick experimentation.
- Megatron Core is a composable library of GPU-optimized building blocks for custom training frameworks: transformer components, parallelism, mixed precision and model architectures.
A third project lives next to it:
- Megatron Bridge (in the NVIDIA-NeMo organization) is a PyTorch-native library that provides pretraining, SFT and LoRA for language, vision-language, audio and multimodal models, and converts checkpoints in both directions between Hugging Face and Megatron Core. Its README describes it as part of the NeMo Framework and a refactor of the previous NeMo training stack. NVIDIA's NeMo documentation points to Megatron Bridge's docs for the framework container starting with the 26.02 release.
The license is Apache-2.0 for Megatron Bridge and Apache for the Megatron-LM repository per its README badge. At the time of writing, the README shows version 0.19.0 and the latest megatron-core on PyPI is 0.19.2 (September 18, 2026). PyPI states a Python requirement of 3.12 or later.
The origin: model parallelism for transformers
The original Megatron-LM paper (Shoeybi et al., 2019) described a simple way to split transformer layers across GPUs. Its abstract reports training models of up to 8.3 billion parameters on 512 GPUs with 15.1 PetaFLOPs sustained across the whole application and 76% scaling efficiency against a single-GPU baseline that sustained 39 TeraFLOPs (about 30% of peak). That is a 2019 result on 2019 hardware. We cite it for the idea, not as a current speed.
The five kinds of parallelism
Megatron Core supports TP, PP, DP, EP and CP. Short definitions, in our words:
| Parallelism | What it splits | Why you use it |
|---|---|---|
| Data (DP) | The batch. Every group holds a model replica | Scale throughput once the model fits |
| Tensor (TP) | Individual weight matrices inside a layer | Fit layers that are too large for one GPU; needs fast links such as NVLink |
| Pipeline (PP) | Groups of layers into stages | Fit very deep models across nodes with lower bandwidth needs than TP |
| Context (CP) | The sequence dimension | Train on long contexts that blow up activation memory |
| Expert (EP) | The experts of a mixture-of-experts model | Train MoE models whose experts do not fit together |
The glossary has short entries for tensor parallelism, pipeline parallelism, data parallelism and mixture of experts.
In practice these are multiplied together. The number of GPUs you launch is split into groups: TP x PP x CP (and EP for MoE layers) define one model replica, and data parallelism fills the remainder. That arithmetic is our summary of how the parallel groups combine, so check the exact constraints in the Megatron Core docs for your version before sizing a job.
Install
From the README, the PyPI route and the source route:
uv pip install megatron-core
git clone https://github.com/NVIDIA/Megatron-LM.git
cd Megatron-LM
uv pip install -e .
The README adds that if a source build runs out of memory you can limit parallel jobs, for example MAX_JOBS=4 uv pip install -e .. It also points to an NGC container route in the Installation Guide. Megatron Bridge's README recommends the NeMo Framework container from NVIDIA's NGC catalog (nvcr.io/nvidia/nemo) with a version tag.
Quickstart
NVIDIA's Megatron Core quickstart gives two training commands. The first runs a minimal training loop on 2 GPUs with mock data:
torchrun --nproc_per_node=2 examples/run_simple_mcore_train_loop.py
The second runs the repo's LLaMA-3 8B example script with FP8 on 8 GPUs (also with mock data):
./examples/llama/train_llama3_8b_h100_fp8.sh
Inside that script, per the file in the repo, the launch is a single-node torchrun with 8 processes, and the flags show what a Megatron job looks like:
--tensor-model-parallel-size $TP_SIZE
--context-parallel-size $CP_SIZE
--sequence-parallel
--fp8-format hybrid
--fp8-amax-history-len 1024
--fp8-amax-compute-algo max
--fp8-param-gather
In that script TP_SIZE and CP_SIZE are both 1 and the pipeline size line is commented out, so it is a starting point to edit, not a tuned recipe. For FP8 and FP4 background, see FP8, FP4 and our Transformer Engine FP8 post.
Use Hugging Face checkpoints through Megatron Bridge
Most teams start from a Hugging Face model, not a Megatron checkpoint. Megatron Bridge's README describes this flow:
- Load a bridge with
AutoBridge.from_hf_pretrained(...)from a Hub ID or local path. - Convert it with
to_megatron_provider(), set the tensor and pipeline parallel sizes, then callfinalize(). - Build the model with
provide_distributed_model(wrap_with_ddp=False). - Export back with
save_hf_pretrained(...)for a full Hugging Face folder, orexport_hf_weights(...)to stream weights.
The README lists support for many families, including Llama 3.x, Qwen 2.5/3/3.5, DeepSeek V3/V4, GLM-4.5 to GLM-5.x, Gemma 3/4, Nemotron-3, Kimi K2 and GPT-OSS, with some older models (Gemma 1/2, Llama 2, Mistral 7B and others) deprecated. It also notes that Python 3.10 support is dropped from version 0.4.0. Check the current list before you plan around a specific model.
Performance NVIDIA publishes
These figures come from the Megatron-LM README. They are NVIDIA's, and each has stated conditions.
- Overall: up to 47% MFU on H100 clusters, for models from 2B to 462B parameters. The 462B model was benchmarked on 6,144 H100 GPUs.
- Weak scaling: MFU rises from 41% for the smallest model to 47 to 48% for the largest, which NVIDIA attributes to larger matrix multiplies having higher arithmetic intensity.
- Strong scaling: GPT-3 (slightly over 175B parameters) scaled from 96 to 4,608 H100 GPUs at a fixed batch size of 1,152 sequences. MFU fell from 47% to 42% as communication became more exposed.
- Conditions: vocabulary size 131,072, sequence length 4,096, throughput measured end to end (data loading, optimizer steps, communication, logging), and without training to convergence. The runs used the overlap flags
--overlap-grad-reduce,--overlap-param-gatherand--tp-comm-overlap, with pipeline overlap on by default.
The README points to the Megatron Bridge Performance Summary for the latest numbers. We have not reproduced any of these, and your MFU on a different model, sequence length or interconnect will differ.
Megatron vs DeepSpeed vs FSDP
| Megatron-LM / Core | DeepSpeed | PyTorch FSDP | |
|---|---|---|---|
| Main idea | Split the model (TP, PP, CP, EP) plus DP | Shard redundant state (ZeRO), plus other features | Shard parameters, gradients and optimizer states across DP workers |
| Hardware | NVIDIA GPU-optimized | NVIDIA and AMD listed | PyTorch-native |
| Code change | Models built from Megatron blocks, or converted via Bridge | Wrap with deepspeed.initialize or a Trainer config | Wrap the module |
| Best for | Pretraining and large MoE at cluster scale | Fine-tuning and mid-scale training | PyTorch-native sharding |
| License | Apache | Apache-2.0 | Part of PyTorch |
Rules of thumb:
- One node or a few GPUs, fine-tuning an existing model: ZeRO-2 or FSDP, usually through TRL or LLaMA-Factory. Combine with LoRA to cut memory.
- Dozens to thousands of GPUs, pretraining or continued pretraining: Megatron Core, where TP and PP keep per-GPU memory and communication manageable.
- Large MoE models: expert parallelism is a first-class feature in Megatron Core.
GPU requirements
The Megatron-LM README lists FP8 and FP4 precision support, and the FP8 quickstart script is named for H100. Hopper and Blackwell GPUs are the natural targets. Tensor parallelism is communication-heavy, so keep TP groups inside a single node with NVLink and use pipeline or data parallelism across nodes. See H100 vs H200 and the H100 and B200 pages for memory and bandwidth specs.
Run it on a cloud GPU
Start with the 2-GPU torchrun example above to check your drivers and container, then scale up to a full 8-GPU node.
FAQ
What is the difference between Megatron-LM and Megatron Core?
Megatron-LM is a reference example with ready-made training scripts. Megatron Core is the composable library of parallelism, kernels and model components that the example and other frameworks build on.
Is Megatron-LM the same as NeMo?
No. NeMo is a broader NVIDIA framework. Its Megatron Bridge library builds on Megatron Core and, per its README, refactors the previous NeMo training stack into a PyTorch-native loop.
Can I fine-tune with Megatron-LM?
Yes. Megatron Bridge provides pretraining, SFT and LoRA. For small models on a single GPU, simpler tools are usually quicker to set up.
Does Megatron-LM work with Hugging Face models?
Through Megatron Bridge, which converts Hugging Face checkpoints into Megatron format and exports them back. Check its README for the currently supported model families.
Does it run on AMD GPUs?
NVIDIA's README describes Megatron Core as GPU-optimized, and its benchmarks and FP8 and FP4 notes are for NVIDIA hardware. We did not find AMD support stated there, so verify it before planning on it.
Sources
- Megatron-LM GitHub README (components, parallelism, precision, license, performance figures and conditions, install): https://github.com/NVIDIA/Megatron-LM
- Megatron Core on PyPI (version 0.19.2, September 18, 2026, Python 3.12 or later): https://pypi.org/project/megatron-core/
- Megatron Core quickstart (torchrun and LLaMA-3 example commands): https://docs.nvidia.com/megatron-core/developer-guide/latest/get-started/quickstart.html
- LLaMA-3 8B FP8 example script (parallelism and FP8 flags): https://github.com/NVIDIA/Megatron-LM/blob/main/examples/llama/train_llama3_8b_h100_fp8.sh
- Megatron Bridge README (Bridge, conversion flow, supported models, container, license): https://github.com/NVIDIA-NeMo/Megatron-Bridge
- NeMo Framework documentation overview (pointer to Megatron Bridge docs from 26.02): https://docs.nvidia.com/nemo-framework/user-guide/latest/overview.html
- Megatron-LM paper, Shoeybi et al., 2019 (8.3B parameters, 512 GPUs, 15.1 PetaFLOPs, 76%): https://arxiv.org/abs/1909.08053
- DeepSpeed README and PyTorch FSDP docs (comparison): https://github.com/deepspeedai/DeepSpeed and https://docs.pytorch.org/docs/2.14/fsdp.html