Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

English | 简体中文

LoongForge

Train LLMs, VLMs, diffusion, and embodied models — faster.

Docs  ·  Blog  ·  Quick Start  ·  Performance  ·  Supported Models

GitHub stars Apache 2.0 license Docker images on Docker Hub PRs welcome

Visit our website Join our Discord Join our WeChat group Find us on RedNote Follow us on X

⭐ Star LoongForge to help more people discover it and grow the community.

🐉 LoongForge

LoongForge is an open-source training framework developed by the Baidu AI Cloud Baige team, built to deliver faster training for mainstream LLMs, VLMs, diffusion, and embodied models.


Example: embodied-model DreamZero at 4.38× baseline throughput, loss curves aligned

DreamZero training run compared side by side: LoongForge reaches 4.38x the baseline throughput while the training loss curves stay aligned

🏗️ Architecture

Since optimal training strategies differ across model families and scales, LoongForge adopts a multi-backend architecture.

LoongForge architecture: a patched-Megatron stack for LLMs, VLMs and diffusion models alongside a torch-native stack for embodied models

  • Megatron Stack — For LLMs, VLMs, and diffusion models. Powered by a patched Megatron-LM and extended with MoE parallelism, per-component heterogeneous parallelism, long-sequence optimizations, etc.
  • Torch-Native Stack — For embodied models (VLA and WAM). A standalone torch-native subsystem featuring DDP / ZeRO-1 / FSDP / HSDP, with deep optimizations for representative models across I/O, communication strategy, kernel efficiency, etc.

🔥 Latest News

  • [2026/09] ✨ Added training support for GLM-5.3-flash.
  • [2026/09] ✨ Added Kimi-K3 BF16 training support for both LLMs and VLMs.
  • [2026/09] ⚡ Added an optimized DreamZero Wan2.2-5B FSDP recipe with cache-aware data loading, compiled attention blocks, frozen-module handling, and FSDP2 Delta-FP8 Param AllGather.
  • [2026/08] 🤖 Added VLA training support for Wall-OSS-0.5, with custom fused operators for higher training throughput.
  • [2026/08] 📄 Released the TAOT paper — topology-aware dynamic expert replica placement that tackles expert-parallel (EP) load imbalance in MoE training, cutting overhead by up to 74% over industry solutions, with 1.43× speedup measured on a real training case. [blog]
  • [2026/08] ✨ Added training support for GLM-5.2, along with a GLM-5.2 + MoonViT custom-composition example for extending GLM with multimodal capabilities.
  • [2026/08] ✨ Added training support for MiniCPM-V-4.6 and Qwen3.8-27B.
  • [2026/08] 🧪 Introduced a unified evaluation module for the embodied stack, currently covering Pi0.5 / xVLA / GR00T, with more models on the way.
  • [2026/07] 🐳 Unified the prebuilt Docker images — all model families (LLM / VLM / VLA / Diffusion) now share a single image.
  • [2026/07] 🤖 Released LoongForge-Embodied, a torch-native DDP/FSDP training subsystem for embodied models (Pi0.5, GR00T-N1.6/N1.7, xVLA, LingBot-VA, FastWAM, DreamZero, and Cosmos3), with up to 4.38× speedup. [blog]
  • [2026/07] ✨ Added training support for DeepSeek-V4-Flash / DeepSeek-V4-Pro.
📅 More
  • [2026/07] ✨ Added training support for Qwen-Image-Edit-2511.
  • [2026/06] 🤖 Expanded VLA coverage with GR00T N1.6; 2.3× speedup on GR00T training. [blog]
  • [2026/05] ⚡ Accelerated Wan 2.2 training by 116%, and added CP and data packing support.
  • [2026/05] ✨ Added training support for Kimi K2.5 / K2.6, and introduced INT4 / NVFP4 PTQ.
  • [2026/05] 🎉 v0.1.0 — first official tagged release of LoongForge.
  • [2026/05] 🌟 Powered the training and public release of LLaVA-OneVision-2.0.
  • [2026/04] 🧩 Added training support for MiniMax-M2.7 on both NVIDIA GPU and Kunlun XPU.
  • [2026/04] 🚀 LoongForge source code publicly available on GitHub. [blog]
  • [2025/10] 🌟 Powered the training and public release of LLaVA-OneVision-1.5 under AIAK-Training-LLM, the predecessor of LoongForge. [blog]

✨ Key Features

🚀 Foundation Models

  • MoE EP Communication Optimization — Overlapped All2All / activation offload / compute, with further memory reduction beyond upstream Megatron-LM. [Usage]
  • MoE Expert Load Balancing — Topology-aware dynamic replication of hot experts to balance EP workloads, with up to 74% lower overhead than industry solutions. [TAOT Paper]
  • Adaptive FP8 Training — End-to-end FP8 for LLMs and VLMs with standard blockwise FP8; an optional adaptive mode picks per-operator precision by GEMM shape and efficiency. [Usage]
  • Custom Fused Operators — Fused kernels like FusedDSA for DSA-style models — TileLang version open-sourced, high-performance CUDA version available on Baidu Baige platform.
  • Long-Sequence Training — Scales LLM training to long sequences via Context Parallel (CP) and chunked-pipeline scheduling.

🧩 Multi-Modal Models

  • Flexible Composition — Assemble VLMs from interchangeable ViT and LLM components (e.g. GLM-5.2 + MoonViT) straight from config — no custom model code. [Usage]
  • Heterogeneous Parallelism — Independent TP / DP / recompute / freeze per model component (e.g. ViT vs. LLM) for optimal throughput and memory. [blog] [Usage]
  • Decoupled Encoder-Decoder Training — Eliminates encoder-induced pipeline bubbles by separating ViT and LLM into independent tasks. [Usage]
  • DP Load Balancing — Improves multi-node scaling efficiency with load-aware data redistribution that mitigates sequence-packing imbalance. [blog] [Usage]
  • Flexible Data Pipeline — Energon WebDataset input for multimodal data, with both online and offline sequence packing. [Usage]

🤖 Embodied Models

  • VLA & WAM Training — A dedicated torch-native DDP/FSDP subsystem for VLA and world-action (WAM) models, decoupled from the Megatron core, with flexible DDP / ZeRO-1 / FSDP / HSDP strategies. [README]
  • Per-Model Deep Optimization — 1.79×–4.38× over official baselines in our benchmarks, from training code customized per model across I/O, communication strategy, and kernel efficiency.
  • FP8 Communication Optimization — Cuts cross-rank traffic on supported NVIDIA GPUs across both parallel strategies: blockwise FP8 delta AllGather for FSDP2 parameters, and FP8 grad all-reduce for DDP gradients. [Usage]
  • Unified Evaluation — Evaluate trained policies on LIBERO / CALVIN / SimplerEnv / RoboTwin, with coverage expanding continuously. [README]
  • Ego2Robot Data Conversion — Turn first-person videos of human manipulation into LeRobot v3.0 training data across 16 dual-arm robot morphologies. [README]

🔌 Compatibility

  • Mcore Bridge — Supports both offline bidirectional Megatron ↔ HuggingFace conversion and online native HF load/save. [Usage]
  • Heterogeneous Hardware — Native support for NVIDIA GPUs and Kunlun XPUs via a minimally-intrusive plugin design.

📖 Deep-dive: LLM · VLM · Embodied Model

📊 Performance

Training throughput speedups over mainstream open-source baselines — each model and its baseline were benchmarked on the same machine type with the same training hyperparameters:

LoongForge benchmark speedups over open-source baselines — from 1.45x on Qwen3-VL up to 5.04x on DeepSeek-V3.2 Lite

DeepSeek-V3.2 Lite reflects DSA operator-level optimizations and was validated on a reduced-layer configuration due to test-bed scale limits.
Numbers were measured at a point in time and may evolve as implementations change on both sides.

⚡ Quick Start

1. Install

Use the latest prebuilt NVIDIA GPU image with NVIDIA Container Toolkit installed:

docker pull loongforge/loongforge:latest
mkdir -p workspace
docker run --gpus all --ipc=host -it --rm \
  -v "$(pwd)/workspace:/workspace/data" \
  -w /workspace/LoongForge \
  loongforge/loongforge:latest bash

Source installation · Kunlun XPU installation.

2. Pick a tutorial — by hardware and modality

3. Find your model's scripts

Launch scripts are available under examples/ (NVIDIA GPU) and examples_xpu/ (Kunlun XPU), with configs in configs/models/.

Example: DreamZero LoRA fine-tuning

This example uses DreamZero Wan2.2-5B LoRA on a single node with 8 GPUs and FSDP. Follow the tutorial to prepare weights (including Wan2.1 CLIP) and DROID data in LeRobot v2 format, then run inside the container:

cd /workspace/LoongForge
export WAN22_CKPT_DIR=/workspace/data/dreamzero/checkpoints/Wan2.2-TI2V-5B
export WAN21_CKPT_DIR=/workspace/data/dreamzero/checkpoints/Wan2.1-I2V-14B-480P
export TOKENIZER_PATH="$WAN22_CKPT_DIR/google/umt5-xxl"
export DATA_PATH=/workspace/data/dreamzero/data/droid_lerobot

EMBODIMENT_TAG=oxe_droid \
  bash examples/embodied/dreamzero/prepare_dreamzero_dataset.sh

GPUS_PER_NODE=8 TRAIN_ITERS=20 SAVE_INTERVAL=20 \
OUTPUT_DIR=/workspace/data/dreamzero/outputs/lora \
  bash examples/embodied/dreamzero/run_dreamzero_wan22_5b_lora_fsdp_finetune.sh

The example runs 20 steps and saves outputs under OUTPUT_DIR.

🏛️ Supported Models

Click any model for its training examples. See the User Guide for full instructions and the model support matrix for all variants.

LLMVLMDiffusionEmbodied

🌟 Powered by LoongForge

Open-source models trained with LoongForge or its predecessor AIAK-Training-LLM:

ModelHighlights
LLaVA-OneVision-2.0Next-generation multimodal model, with new VideoCaption and Spatial datasets
Innovator-VLScientific multimodal LLM for advanced reasoning
LLaVA-OneVision-1.5Fully open framework for democratized multimodal training
Qianfan-VLDomain-enhanced vision-language models for enterprise, 3B–70B parameters

📂 Repository Layout

📁 Directory tree
LoongForge/
├── loongforge/                   # Core training framework
│   ├── train/                    # Training entry points & trainers
│   │   ├── pretrain/             #   Pretrain (LLM, VLM)
│   │   ├── sft/                  #   SFT (LLM, VLM, InternVL, ERNIE)
│   │   └── diffusion/            #   Diffusion (WAN, Qwen-Image)
│   ├── models/                   # Unified model abstractions
│   │   ├── foundation/           #   LLM backbones (LLaMA, Qwen, DeepSeek, ...)
│   │   ├── encoder/              #   Vision encoders (ViT, Qwen-VL, InternVL, ...)
│   │   ├── omni_models/          #   Multi-modal composition
│   │   ├── diffusion/            #   Diffusion models (WAN, Qwen-Image)
│   │   └── common/               #   Shared layers and utilities
│   ├── embodied/                 # LoongForge-Embodied: standalone torch-native (DDP/FSDP)
│   │                             #   embodied (VLA + world-action) subsystem — see loongforge/embodied/README.md
│   ├── data/                     # Data pipelines (multi-modal, video, DP balance)
│   ├── tokenizer/                # Tokenizers
│   └── utils/                    # Config map, constants, etc.
├── third_party/Loong-Megatron/   # Patched Megatron-LM (git submodule)
├── configs/                      # Hydra YAML configs (models, data)
├── examples/                     # GPU launch scripts
├── examples_xpu/                 # Kunlun XPU launch scripts
├── tools/                        # Checkpoint conversion, data preprocessing
├── ops/                          # Custom fused operators (incl. open-sourced TileLang)
├── patches/                      # TransformerEngine patches
├── docker/                       # Dockerfiles (GPU & XPU)
├── tests/                        # E2E test suite (YAML-driven)
└── docs/                         # Documentation

📝 Citation

If you find LoongForge helpful, please cite this project:

@software{LoongForge2026,
  title  = {LoongForge: A high-performance framework for training LLMs, VLMs, diffusion, and embodied models},
  author = {{The LoongForge Authors}},
  year   = {2026},
  url    = {https://github.com/baidu-baige/LoongForge}
}

If you use TAOT for MoE training in LoongForge, you can cite our paper:

@article{zhang2026taot,
  title   = {{TAOT}: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in {MoE} Training},
  author  = {Zhang, Lingyun and Zhang, Henghua and Gu, Shilei and Mo, Kai and Han, Shuai and Li, Shiyong and Wang, Yanpeng and Shen, Dou},
  journal = {arXiv preprint arXiv:2608.03676},
  year    = {2026},
  url     = {https://arxiv.org/abs/2608.03676}
}

🤝 Contributing

We warmly welcome community contributions — bug reports, feature proposals, and PRs alike. Please read our Contributing Guidelines before submitting.

Thanks to all our contributors:

LoongForge contributors

🙏 Acknowledgments

LoongForge stands on the shoulders of the open-source community. Its Megatron stack builds on NVIDIA's Megatron-LM, and the project also draws on HuggingFace Transformers, LLaMA-Factory, Megatron-Bridge, LeRobot, and the official implementations of the models we support (e.g. OpenPI, NVIDIA Isaac GR00T). We also thank the LINUX DO community for its welcoming space for technical discussion and its support of open-source sharing.

💬 Contact Us

ChannelWhat it's for
GitHub IssuesBug reports, usage questions, and feature requests
Developer CommunitiesWeChat group, Xiaohongshu, and more
EmailEnterprise adoption, large-scale deployment, partnership, or any other topic

📄 License

LoongForge is released under the Apache License 2.0; some files derive from third-party projects — see their file headers.

关于 About

A high-performance framework for training LLMs, VLMs, diffusion, and embodied models on NVIDIA GPUs and Kunlun XPUs. Supports Pi0.5, GR00T-N1.7, FastWAM, Cosmos3, DreamZero, DeepSeek-V4, GLM-5.3, Kimi-K3 and more.
gpukunlunllmloramegatronmid-trainingpretrainingsfttorchvlavlmwamwanxpu

语言 Languages

Python92.9%
Shell3.0%
Cuda2.7%
Jinja1.0%
Dockerfile0.2%
JavaScript0.1%
C++0.1%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
460
Total Commits
峰值: 61次/周
Less
More

核心贡献者 Contributors