Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

Evo-RLT

lerobot version training dataset RW-RL dataset model checkpoint license

SJTU-MINT

A LeRobot-based reproduction of RLT, covering RL-token learning, transition-cache generation, actor-critic training, and real-robot rollout.

RLT Pipeline

RLT training pipeline

Real-Robot Rollout Demo

RLT real-robot rollout demo

Collect Human Demonstrations

Collect human demonstrations demo

Policy Rollout with Human Intervention

Policy rollout with human intervention demo

🎯 Evo-RLT Focus

  • RLT reproduction: this repository presents RLT as an independent LeRobot-based reproduction for the pi paper.
  • Open training path: the code covers VLA finetuning, RL-token learning, transition-cache generation, and chunk actor-critic training.
  • Real-robot deployment path: the recording wrapper supports VLA/RLT rollout, RTC defaults, pedal labels, and human-in-the-loop collection.

📰 News

  • [2026-06-29] Released Evo-RLT.
  • [2026-06-26] Added training dataset and checkpoint links.

🧭 Table of Contents

Getting StartedTraining PipelineProject Info
⚡ Quick Start🧪 Training Pipeline🤗 Model & Dataset
1) Installation3) Finetune VLA🗂️ Repository Layout
2) Hardware Setup4) Train RL Token✅ Development Checks
🤖 Real-Robot Recording and Deployment5) Build Transition Cache🧭 Future TODO
6) Train Chunk Actor-Critic💬 Community Channels / 🏫 Affiliations / 📄 License

⚡ Quick Start

1) Installation

Evo-RLT depends on LeRobot v0.5.1, which currently ships from the official GitHub tag and requires Python 3.12+.

git clone https://github.com/MINT-SJTU/Evo-RLT.git
cd evo-rlt

conda create -y -n evo-rlt python=3.12
conda activate evo-rlt

python -m pip install -e ".[lerobot]"

Do not put a local LeRobot source checkout on PYTHONPATH; Evo-RLT is tested against the official LeRobot package installed by the lerobot extra.

Evo-RLT keeps policy registration out of LeRobot source files. Before using LeRobot factory helpers with RLT policy types, register the adapter once:

from evo_rlt.adapters.lerobot import register

register()

Registered policy types:

rlt_token    # RL-token reconstruction policy
rlt_ac       # chunk actor-critic policy
rlt          # deployment policy wrapper

LeRobot 0.5.1 Normalization

Evo-RLT follows the LeRobot >=0.5 processor-pipeline runtime:

raw observation -> policy_preprocessor -> policy -> policy_postprocessor -> robot action

Checkpoints trained or migrated for LeRobot >=0.5 are expected to include policy_preprocessor.json, policy_postprocessor.json, and processor weight files such as NormalizerProcessorStep / UnnormalizerProcessorStep statistics. The presence of NormalizerProcessorStep is normal in LeRobot 0.5.1; it is not a legacy workaround.

Only migrate normalization for checkpoints trained before LeRobot's processor-pipeline migration. For those checkpoints, verify normalization is not applied twice: model weights should not contain embedded normalization keys such as normalize_inputs.*, and the external pre/postprocessor stats must match the training normalization modes.

2) Hardware Setup

Use the Evo-RL hardware setup for the shared SO-series robot bring-up steps: assembly, stable serial/camera paths, camera validation, and basic teleoperation checks. PiPER/PiPER-X support is planned; see Future TODO.

This repository only differs at the recording/deployment configuration layer:

  • evo-rlt-record reads a setup manifest from --setup-json, or from ~/.roboclaw/workspace/embodied/manifest.json when the flag is omitted.
  • Arm entries point to per-device calibration_dir folders. The wrapper looks for <calibration_dir>/<folder-name>.json, then stages those files into temporary LeRobot-compatible names at runtime.
  • Follower calibrations are staged as bimanual_left.json and bimanual_right.json under a temporary robot calibration directory.
  • Leader calibrations are staged as bimanual_leader_left.json and bimanual_leader_right.json under a temporary teleop calibration directory.
  • Dataset paths are created under <datasets.root>/<MMDD>_<dataset-tag>/<prefix>_<HHMMSS>. If datasets.root is omitted, the default is ~/.roboclaw/workspace/embodied/datasets.

Example setup manifest:

{
  "datasets": {"root": "/path/to/lerobot_datasets"},
  "arms": [
    {
      "alias": "left_follower",
      "type": "follower",
      "port": "/dev/serial/by-id/<left-follower>",
      "calibration_dir": "/path/to/calibration/<left-follower-serial>"
    },
    {
      "alias": "right_follower",
      "type": "follower",
      "port": "/dev/serial/by-id/<right-follower>",
      "calibration_dir": "/path/to/calibration/<right-follower-serial>"
    },
    {
      "alias": "left_leader",
      "type": "leader",
      "port": "/dev/serial/by-id/<left-leader>",
      "calibration_dir": "/path/to/calibration/<left-leader-serial>"
    },
    {
      "alias": "right_leader",
      "type": "leader",
      "port": "/dev/serial/by-id/<right-leader>",
      "calibration_dir": "/path/to/calibration/<right-leader-serial>"
    }
  ],
  "cameras": [
    {
      "alias": "left_wrist",
      "port": "/dev/v4l/by-path/<left-wrist>",
      "width": 640,
      "height": 480,
      "fps": 30,
      "fourcc": "MJPG"
    },
    {
      "alias": "right_wrist",
      "port": "/dev/v4l/by-path/<right-wrist>",
      "width": 640,
      "height": 480,
      "fps": 30,
      "fourcc": "MJPG"
    },
    {
      "alias": "right_front",
      "port": "/dev/v4l/by-path/<right-front>",
      "width": 640,
      "height": 480,
      "fps": 30,
      "fourcc": "MJPG"
    }
  ]
}

🧪 Training Pipeline

The typical RLT workflow has four stages. Video datasets require FFmpeg shared libraries for LeRobot 0.5.1 / TorchCodec decoding:

sudo apt-get update && sudo apt-get install -y ffmpeg

For saved checkpoints, LeRobot 0.5.1 writes numeric checkpoint directories such as checkpoints/000001/pretrained_model. Use the latest numeric directory when checkpoints/last/pretrained_model is not present.

3) Finetune VLA

Use LeRobot's training entrypoint to finetune a pi0.5 VLA checkpoint on a LeRobot dataset.

python -m lerobot.scripts.lerobot_train \
  --dataset.repo_id=<HF_ORG>/<DATASET> \
  --dataset.root=<LOCAL_DATASET_ROOT> \
  --policy.path=<BASE_PI05_CHECKPOINT_DIR> \
  --policy.device=cuda \
  --policy.dtype=bfloat16 \
  --batch_size=16 \
  --steps=30000 \
  --save_freq=5000 \
  --eval_freq=0 \
  --tolerance_s=0.04 \
  --output_dir=outputs/vla_ft \
  --job_name=vla_ft

4) Train RL Token

python -c 'from evo_rlt.adapters.lerobot import register; register(); from lerobot.scripts.lerobot_train import main; main()' \
  --dataset.repo_id=<HF_ORG>/<DATASET> \
  --dataset.root=<LOCAL_DATASET_ROOT> \
  --policy.type=rlt_token \
  --policy.repo_id=<HF_ORG>/rlt_token \
  --policy.push_to_hub=false \
  --policy.vla_pretrained_path=outputs/vla_ft/checkpoints/last/pretrained_model \
  --policy.vla_dtype=bfloat16 \
  --policy.rl_token_num_rl_tokens=1 \
  --policy.tokenizer_path=/path/to/paligemma-3b-pt-224-snapshot \
  --policy.token_pool_size=0 \
  --policy.device=cuda \
  --batch_size=8 \
  --steps=10000 \
  --save_freq=2000 \
  --eval_freq=0 \
  --tolerance_s=0.04 \
  --output_dir=outputs/rl_token \
  --job_name=rl_token

5) Build Transition Cache

evo-rlt-build-transition-cache-v2 \
  --demo-dataset-repo-id <HF_ORG>/<DATASET> \
  --demo-dataset-root <LOCAL_DATASET_ROOT> \
  --rl-token-policy-path outputs/rl_token/checkpoints/last/pretrained_model \
  --vla-pretrained-path outputs/vla_ft/checkpoints/last/pretrained_model \
  --tokenizer-path /path/to/paligemma-3b-pt-224-snapshot \
  --output-dir outputs/cache \
  --task-instruction "<TASK>" \
  --chunk-length 10 \
  --frame-stride 2 \
  --batch-size 8 \
  --num-workers 2 \
  --train-ratio 0.9 \
  --tolerance-s 0.04 \
  --device cuda

6) Train Chunk Actor-Critic

outputs/cache must contain chunk_transitions_train.pt. The Evo-RLT registry detects this cache directory through --dataset.repo_id.

python -c 'from evo_rlt.adapters.lerobot import register; register(); from lerobot.scripts.lerobot_train import main; main()' \
  --dataset.repo_id=outputs/cache \
  --policy.type=rlt_ac \
  --policy.repo_id=<HF_ORG>/rlt_ac \
  --policy.push_to_hub=false \
  --policy.vla_pretrained_path=outputs/vla_ft/checkpoints/last/pretrained_model \
  --policy.rl_token_pretrained_path=outputs/rl_token/checkpoints/last/pretrained_model \
  --policy.vla_dtype=bfloat16 \
  --policy.tokenizer_path=/path/to/paligemma-3b-pt-224-snapshot \
  --policy.rl_token_num_rl_tokens=1 \
  --policy.chunk_length=10 \
  --policy.chunk_exec_steps=25 \
  --policy.phase_mode=always_rl \
  --policy.device=cuda \
  --batch_size=256 \
  --steps=50000 \
  --save_freq=5000 \
  --eval_freq=0 \
  --output_dir=outputs/ac \
  --job_name=rlt_ac

🤖 Real-Robot Recording and Deployment

Set up the environment before running robot commands:

cd /path/to/evo-rlt
source ~/miniconda3/etc/profile.d/conda.sh
conda activate evo-rlt
python -m pip install -e ".[lerobot]"
export HF_HUB_OFFLINE=1

Default VLA-RLT-VLA real-robot collection uses the official LeRobot 0.5.1 streaming encoder. The wrapper keeps the foreground recording loop responsive and expands the dataset settings to --dataset.vcodec=h264, --dataset.video_encoding_batch_size=<num_episodes + 1>, and --dataset.streaming_encoding=true.

Add --online-replay to evo-rlt-collect-default or evo-rlt-record segment to collect normalized chunk transitions during robot recording. Successful episodes receive reward 1 at the terminal chunk's last valid step; failed episodes receive only zeros. Episodes enter replay only after their outcome is confirmed and the recording is saved; rerecorded episodes are discarded. The cache is saved atomically to <DATASET_ROOT>/online_replay/chunk_transitions_train.pt and can be passed directly to actor-critic training via --dataset.repo_id. Resuming recording reloads the existing replay cache (up to 200,000 chunks).

Online replay uses the deployed SFT preprocessor for executed actions, including human intervention, and encodes the current observation/reference every chunk boundary. This adds a VLA forward pass per chunk and can lower control frequency. With RTC, replay encoding shares the inference lock but leaves action and metadata queues intact. This collects training data; optimizer updates still run through the actor-critic trainer. ACP inference is unsupported with online replay.

Shared collection arguments:

COMMON_ARGS=(
  --setup-json /path/to/robot_manifest.json \
  --policy-path /path/to/rlt_ac_policy \
  --vla-path /path/to/pi05_vla_checkpoint_or_dir \
  --rl-token-path /path/to/rl_token_policy \
  --dataset-tag vla_rlt_vla_test \
  --num-episodes 5 \
  --episode-time-s 3000 \
  --fps 30 \
  --vcodec h264 \
  --rlt-toggle-key r \
  --teleop-toggle-key space
)

Start in VLA mode and record the full trajectory:

evo-rlt-record collect "${COMMON_ARGS[@]}"

Start in VLA mode and record only the critical segment:

evo-rlt-record collect "${COMMON_ARGS[@]}" --only-critical

Start in teleoperation mode and record the full trajectory:

evo-rlt-record collect "${COMMON_ARGS[@]}" --start-with-teleop

Start in teleoperation mode and record only the critical segment:

evo-rlt-record collect "${COMMON_ARGS[@]}" --start-with-teleop --only-critical

The same collection entrypoint is exposed as evo-rlt-collect-default after reinstalling package entry points, but checkpoint and setup paths still need to be supplied by the caller.

Validated RTC defaults for this collection mode:

RLT RTC execution horizon: 10
VLA RTC execution horizon: 25
RTC action queue refill threshold: 30
RTC max guidance weight: 10.0
RTC prefix attention schedule: EXP

Default collection controls:

Full-trajectory mode:
r              save the full episode as success after the double-tap window
r+r            save the full episode as failure
space          toggle teleop intervention; pressing again exits teleop
left arrow     rerecord the current episode
Esc            stop data collection

Critical-segment mode (`--only-critical`):
r              enter RLT mode and start recording the critical segment
r              save the segment as success, exit RLT mode, then end the episode
r+r            save the segment as failure, exit RLT mode, then end the episode
space          toggle teleop intervention; pressing again exits teleop
left arrow     rerecord the current episode
Esc            stop data collection

VLA-only full-process recording with pedal outcome labels:

evo-rlt-record full \
  --initial-source vla \
  --setup-json <ROBOT_SETUP_JSON> \
  --policy-path <AC_OR_VLA_POLICY_PATH> \
  --vla-path <BASE_OR_FINETUNED_VLA_PT> \
  --phase-mode always_vla \
  --chunk-exec-steps 25 \
  --pedal-outcome \
  --double-tap-window-s 0.6 \
  --num-episodes 5 \
  --episode-time-s 3000 \
  --reset-time-s 0 \
  --fps 30 \
  --vcodec h264 \
  --dataset-tag vla_full_pedal \
  --no-teleop

For headless SSH runs where no keyboard or pedal outcome will be provided, add --default-episode-success success or --default-episode-success failure.

Pedal semantics in this mode:

single tap    success, end current episode, start next episode
double tap    failure, end current episode, start next episode

🗂️ Repository Layout

src/evo_rlt/core                  # algorithm core, torch-only
src/evo_rlt/adapters/lerobot      # LeRobot/pi0.5/dataset/policy/record adapters
src/evo_rlt/cli                   # training and cache CLIs
tests/rlt                         # focused RLT unit and integration tests

✅ Development Checks

PYTHONPATH=src pytest -q tests/rlt
PYTHONPATH=src python -m compileall -q src/evo_rlt tests/rlt

🤗 Model & Dataset

🧭 Future TODO

  • PiPER/PiPER-X real-robot deployment support.

💬 Community Channels

EvoMind WeChat QR

🏫 Affiliations

SJTU community visual EvoMind

📄 License

Apache-2.0. See LICENSE.

关于 About

No description, website, or topics provided.

语言 Languages

Python100.0%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
38
Total Commits
峰值: 17次/周
Less
More

核心贡献者 Contributors