English | 简体中文
Train LLMs, VLMs, diffusion, and embodied models — faster.
Docs · Blog · Quick Start · Performance · Supported Models
⭐ Star LoongForge to help more people discover it and grow the community.
🐉 LoongForge
LoongForge is an open-source training framework developed by the Baidu AI Cloud Baige team, built to deliver faster training for mainstream LLMs, VLMs, diffusion, and embodied models.
- Easy to Use — Ready-to-run configs and launch examples for every supported model, spanning pre-training, continued pre-training, SFT, and LoRA.
- High Performance — Built on multiple backends (Megatron-LM and torch-native) with deep training-speed optimizations that keep training loss aligned with the baseline.
- Production-Proven — Open-sourced from AIAK-Training-LLM, a training suite that serves enterprise customers' proprietary models and powers open-source model releases, with production runs reaching 5,000+ XPUs.
Example: embodied-model DreamZero at 4.38× baseline throughput, loss curves aligned
🏗️ Architecture
Since optimal training strategies differ across model families and scales, LoongForge adopts a multi-backend architecture.
- Megatron Stack — For LLMs, VLMs, and diffusion models. Powered by a patched Megatron-LM and extended with MoE parallelism, per-component heterogeneous parallelism, long-sequence optimizations, etc.
- Torch-Native Stack — For embodied models (VLA and WAM). A standalone torch-native subsystem featuring DDP / ZeRO-1 / FSDP / HSDP, with deep optimizations for representative models across I/O, communication strategy, kernel efficiency, etc.
🔥 Latest News
- [2026/09] ✨ Added training support for GLM-5.3-flash.
- [2026/09] ✨ Added Kimi-K3 BF16 training support for both LLMs and VLMs.
- [2026/09] ⚡ Added an optimized DreamZero Wan2.2-5B FSDP recipe with cache-aware data loading, compiled attention blocks, frozen-module handling, and FSDP2 Delta-FP8 Param AllGather.
- [2026/08] 🤖 Added VLA training support for Wall-OSS-0.5, with custom fused operators for higher training throughput.
- [2026/08] 📄 Released the TAOT paper — topology-aware dynamic expert replica placement that tackles expert-parallel (EP) load imbalance in MoE training, cutting overhead by up to 74% over industry solutions, with 1.43× speedup measured on a real training case. [blog]
- [2026/08] ✨ Added training support for GLM-5.2, along with a GLM-5.2 + MoonViT custom-composition example for extending GLM with multimodal capabilities.
- [2026/08] ✨ Added training support for MiniCPM-V-4.6 and Qwen3.8-27B.
- [2026/08] 🧪 Introduced a unified evaluation module for the embodied stack, currently covering Pi0.5 / xVLA / GR00T, with more models on the way.
- [2026/07] 🐳 Unified the prebuilt Docker images — all model families (LLM / VLM / VLA / Diffusion) now share a single image.
- [2026/07] 🤖 Released LoongForge-Embodied, a torch-native DDP/FSDP training subsystem for embodied models (Pi0.5, GR00T-N1.6/N1.7, xVLA, LingBot-VA, FastWAM, DreamZero, and Cosmos3), with up to 4.38× speedup. [blog]
- [2026/07] ✨ Added training support for DeepSeek-V4-Flash / DeepSeek-V4-Pro.
📅 More
- [2026/07] ✨ Added training support for Qwen-Image-Edit-2511.
- [2026/06] 🤖 Expanded VLA coverage with GR00T N1.6; 2.3× speedup on GR00T training. [blog]
- [2026/05] ⚡ Accelerated Wan 2.2 training by 116%, and added CP and data packing support.
- [2026/05] ✨ Added training support for Kimi K2.5 / K2.6, and introduced INT4 / NVFP4 PTQ.
- [2026/05] 🎉 v0.1.0 — first official tagged release of LoongForge.
- [2026/05] 🌟 Powered the training and public release of LLaVA-OneVision-2.0.
- [2026/04] 🧩 Added training support for MiniMax-M2.7 on both NVIDIA GPU and Kunlun XPU.
- [2026/04] 🚀 LoongForge source code publicly available on GitHub. [blog]
- [2025/10] 🌟 Powered the training and public release of LLaVA-OneVision-1.5 under AIAK-Training-LLM, the predecessor of LoongForge. [blog]
✨ Key Features
🚀 Foundation Models
- MoE EP Communication Optimization — Overlapped All2All / activation offload / compute, with further memory reduction beyond upstream Megatron-LM. [Usage]
- MoE Expert Load Balancing — Topology-aware dynamic replication of hot experts to balance EP workloads, with up to 74% lower overhead than industry solutions. [TAOT Paper]
- Adaptive FP8 Training — End-to-end FP8 for LLMs and VLMs with standard blockwise FP8; an optional adaptive mode picks per-operator precision by GEMM shape and efficiency. [Usage]
- Custom Fused Operators — Fused kernels like FusedDSA for DSA-style models — TileLang version open-sourced, high-performance CUDA version available on Baidu Baige platform.
- Long-Sequence Training — Scales LLM training to long sequences via Context Parallel (CP) and chunked-pipeline scheduling.
🧩 Multi-Modal Models
- Flexible Composition — Assemble VLMs from interchangeable ViT and LLM components (e.g. GLM-5.2 + MoonViT) straight from config — no custom model code. [Usage]
- Heterogeneous Parallelism — Independent TP / DP / recompute / freeze per model component (e.g. ViT vs. LLM) for optimal throughput and memory. [blog] [Usage]
- Decoupled Encoder-Decoder Training — Eliminates encoder-induced pipeline bubbles by separating ViT and LLM into independent tasks. [Usage]
- DP Load Balancing — Improves multi-node scaling efficiency with load-aware data redistribution that mitigates sequence-packing imbalance. [blog] [Usage]
- Flexible Data Pipeline — Energon WebDataset input for multimodal data, with both online and offline sequence packing. [Usage]
🤖 Embodied Models
- VLA & WAM Training — A dedicated torch-native DDP/FSDP subsystem for VLA and world-action (WAM) models, decoupled from the Megatron core, with flexible DDP / ZeRO-1 / FSDP / HSDP strategies. [README]
- Per-Model Deep Optimization — 1.79×–4.38× over official baselines in our benchmarks, from training code customized per model across I/O, communication strategy, and kernel efficiency.
- FP8 Communication Optimization — Cuts cross-rank traffic on supported NVIDIA GPUs across both parallel strategies: blockwise FP8 delta AllGather for FSDP2 parameters, and FP8 grad all-reduce for DDP gradients. [Usage]
- Unified Evaluation — Evaluate trained policies on LIBERO / CALVIN / SimplerEnv / RoboTwin, with coverage expanding continuously. [README]
- Ego2Robot Data Conversion — Turn first-person videos of human manipulation into LeRobot v3.0 training data across 16 dual-arm robot morphologies. [README]
🔌 Compatibility
- Mcore Bridge — Supports both offline bidirectional Megatron ↔ HuggingFace conversion and online native HF load/save. [Usage]
- Heterogeneous Hardware — Native support for NVIDIA GPUs and Kunlun XPUs via a minimally-intrusive plugin design.
📖 Deep-dive: LLM · VLM · Embodied Model
📊 Performance
Training throughput speedups over mainstream open-source baselines — each model and its baseline were benchmarked on the same machine type with the same training hyperparameters:
DeepSeek-V3.2 Lite reflects DSA operator-level optimizations and was validated on a reduced-layer configuration due to test-bed scale limits.
Numbers were measured at a point in time and may evolve as implementations change on both sides.
⚡ Quick Start
1. Install
Use the latest prebuilt NVIDIA GPU image with NVIDIA Container Toolkit installed:
docker pull loongforge/loongforge:latest
mkdir -p workspace
docker run --gpus all --ipc=host -it --rm \
-v "$(pwd)/workspace:/workspace/data" \
-w /workspace/LoongForge \
loongforge/loongforge:latest bashSource installation · Kunlun XPU installation.
2. Pick a tutorial — by hardware and modality
- NVIDIA GPU: LLM · VLM · VLA & WAM · Diffusion
- Kunlun XPU: Kunlun XPU Tutorials
3. Find your model's scripts
Launch scripts are available under examples/ (NVIDIA GPU) and examples_xpu/ (Kunlun XPU), with configs in configs/models/.
Example: DreamZero LoRA fine-tuning
This example uses DreamZero Wan2.2-5B LoRA on a single node with 8 GPUs and FSDP. Follow the tutorial to prepare weights (including Wan2.1 CLIP) and DROID data in LeRobot v2 format, then run inside the container:
cd /workspace/LoongForge
export WAN22_CKPT_DIR=/workspace/data/dreamzero/checkpoints/Wan2.2-TI2V-5B
export WAN21_CKPT_DIR=/workspace/data/dreamzero/checkpoints/Wan2.1-I2V-14B-480P
export TOKENIZER_PATH="$WAN22_CKPT_DIR/google/umt5-xxl"
export DATA_PATH=/workspace/data/dreamzero/data/droid_lerobot
EMBODIMENT_TAG=oxe_droid \
bash examples/embodied/dreamzero/prepare_dreamzero_dataset.sh
GPUS_PER_NODE=8 TRAIN_ITERS=20 SAVE_INTERVAL=20 \
OUTPUT_DIR=/workspace/data/dreamzero/outputs/lora \
bash examples/embodied/dreamzero/run_dreamzero_wan22_5b_lora_fsdp_finetune.shThe example runs 20 steps and saves outputs under OUTPUT_DIR.
🏛️ Supported Models
Click any model for its training examples. See the User Guide for full instructions and the model support matrix for all variants.
| LLM | VLM | Diffusion | Embodied |
|---|---|---|---|
|
|
|
🌟 Powered by LoongForge
Open-source models trained with LoongForge or its predecessor AIAK-Training-LLM:
| Model | Highlights |
|---|---|
| LLaVA-OneVision-2.0 | Next-generation multimodal model, with new VideoCaption and Spatial datasets |
| Innovator-VL | Scientific multimodal LLM for advanced reasoning |
| LLaVA-OneVision-1.5 | Fully open framework for democratized multimodal training |
| Qianfan-VL | Domain-enhanced vision-language models for enterprise, 3B–70B parameters |
📂 Repository Layout
📁 Directory tree
LoongForge/
├── loongforge/ # Core training framework
│ ├── train/ # Training entry points & trainers
│ │ ├── pretrain/ # Pretrain (LLM, VLM)
│ │ ├── sft/ # SFT (LLM, VLM, InternVL, ERNIE)
│ │ └── diffusion/ # Diffusion (WAN, Qwen-Image)
│ ├── models/ # Unified model abstractions
│ │ ├── foundation/ # LLM backbones (LLaMA, Qwen, DeepSeek, ...)
│ │ ├── encoder/ # Vision encoders (ViT, Qwen-VL, InternVL, ...)
│ │ ├── omni_models/ # Multi-modal composition
│ │ ├── diffusion/ # Diffusion models (WAN, Qwen-Image)
│ │ └── common/ # Shared layers and utilities
│ ├── embodied/ # LoongForge-Embodied: standalone torch-native (DDP/FSDP)
│ │ # embodied (VLA + world-action) subsystem — see loongforge/embodied/README.md
│ ├── data/ # Data pipelines (multi-modal, video, DP balance)
│ ├── tokenizer/ # Tokenizers
│ └── utils/ # Config map, constants, etc.
├── third_party/Loong-Megatron/ # Patched Megatron-LM (git submodule)
├── configs/ # Hydra YAML configs (models, data)
├── examples/ # GPU launch scripts
├── examples_xpu/ # Kunlun XPU launch scripts
├── tools/ # Checkpoint conversion, data preprocessing
├── ops/ # Custom fused operators (incl. open-sourced TileLang)
├── patches/ # TransformerEngine patches
├── docker/ # Dockerfiles (GPU & XPU)
├── tests/ # E2E test suite (YAML-driven)
└── docs/ # Documentation
📝 Citation
If you find LoongForge helpful, please cite this project:
@software{LoongForge2026,
title = {LoongForge: A high-performance framework for training LLMs, VLMs, diffusion, and embodied models},
author = {{The LoongForge Authors}},
year = {2026},
url = {https://github.com/baidu-baige/LoongForge}
}If you use TAOT for MoE training in LoongForge, you can cite our paper:
@article{zhang2026taot,
title = {{TAOT}: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in {MoE} Training},
author = {Zhang, Lingyun and Zhang, Henghua and Gu, Shilei and Mo, Kai and Han, Shuai and Li, Shiyong and Wang, Yanpeng and Shen, Dou},
journal = {arXiv preprint arXiv:2608.03676},
year = {2026},
url = {https://arxiv.org/abs/2608.03676}
}🤝 Contributing
We warmly welcome community contributions — bug reports, feature proposals, and PRs alike. Please read our Contributing Guidelines before submitting.
Thanks to all our contributors:
🙏 Acknowledgments
LoongForge stands on the shoulders of the open-source community. Its Megatron stack builds on NVIDIA's Megatron-LM, and the project also draws on HuggingFace Transformers, LLaMA-Factory, Megatron-Bridge, LeRobot, and the official implementations of the models we support (e.g. OpenPI, NVIDIA Isaac GR00T). We also thank the LINUX DO community for its welcoming space for technical discussion and its support of open-source sharing.
💬 Contact Us
| Channel | What it's for |
|---|---|
| GitHub Issues | Bug reports, usage questions, and feature requests |
| Developer Communities | WeChat group, Xiaohongshu, and more |
| Enterprise adoption, large-scale deployment, partnership, or any other topic |
📄 License
LoongForge is released under the Apache License 2.0; some files derive from third-party projects — see their file headers.