Evo-RLT
SJTU-MINT
A LeRobot-based reproduction of RLT, covering RL-token learning, transition-cache generation, actor-critic training, and real-robot rollout.
RLT Pipeline
Real-Robot Rollout Demo
Collect Human Demonstrations
Policy Rollout with Human Intervention
🎯 Evo-RLT Focus
- RLT reproduction: this repository presents RLT as an independent LeRobot-based reproduction for the pi paper.
- Open training path: the code covers VLA finetuning, RL-token learning, transition-cache generation, and chunk actor-critic training.
- Real-robot deployment path: the recording wrapper supports VLA/RLT rollout, RTC defaults, pedal labels, and human-in-the-loop collection.
📰 News
- [2026-06-29] Released Evo-RLT.
- [2026-06-26] Added training dataset and checkpoint links.
🧭 Table of Contents
⚡ Quick Start
1) Installation
Evo-RLT depends on LeRobot v0.5.1, which currently ships from the official GitHub tag and requires Python 3.12+.
git clone https://github.com/MINT-SJTU/Evo-RLT.git
cd evo-rlt
conda create -y -n evo-rlt python=3.12
conda activate evo-rlt
python -m pip install -e ".[lerobot]"Do not put a local LeRobot source checkout on PYTHONPATH; Evo-RLT is tested against the official LeRobot package installed by the lerobot extra.
Evo-RLT keeps policy registration out of LeRobot source files. Before using LeRobot factory helpers with RLT policy types, register the adapter once:
from evo_rlt.adapters.lerobot import register
register()Registered policy types:
rlt_token # RL-token reconstruction policy
rlt_ac # chunk actor-critic policy
rlt # deployment policy wrapperLeRobot 0.5.1 Normalization
Evo-RLT follows the LeRobot >=0.5 processor-pipeline runtime:
raw observation -> policy_preprocessor -> policy -> policy_postprocessor -> robot actionCheckpoints trained or migrated for LeRobot >=0.5 are expected to include policy_preprocessor.json, policy_postprocessor.json, and processor weight files such as NormalizerProcessorStep / UnnormalizerProcessorStep statistics. The presence of NormalizerProcessorStep is normal in LeRobot 0.5.1; it is not a legacy workaround.
Only migrate normalization for checkpoints trained before LeRobot's processor-pipeline migration. For those checkpoints, verify normalization is not applied twice: model weights should not contain embedded normalization keys such as normalize_inputs.*, and the external pre/postprocessor stats must match the training normalization modes.
2) Hardware Setup
Use the Evo-RL hardware setup for the shared SO-series robot bring-up steps: assembly, stable serial/camera paths, camera validation, and basic teleoperation checks. PiPER/PiPER-X support is planned; see Future TODO.
This repository only differs at the recording/deployment configuration layer:
evo-rlt-recordreads a setup manifest from--setup-json, or from~/.roboclaw/workspace/embodied/manifest.jsonwhen the flag is omitted.- Arm entries point to per-device
calibration_dirfolders. The wrapper looks for<calibration_dir>/<folder-name>.json, then stages those files into temporary LeRobot-compatible names at runtime. - Follower calibrations are staged as
bimanual_left.jsonandbimanual_right.jsonunder a temporary robot calibration directory. - Leader calibrations are staged as
bimanual_leader_left.jsonandbimanual_leader_right.jsonunder a temporary teleop calibration directory. - Dataset paths are created under
<datasets.root>/<MMDD>_<dataset-tag>/<prefix>_<HHMMSS>. Ifdatasets.rootis omitted, the default is~/.roboclaw/workspace/embodied/datasets.
Example setup manifest:
{
"datasets": {"root": "/path/to/lerobot_datasets"},
"arms": [
{
"alias": "left_follower",
"type": "follower",
"port": "/dev/serial/by-id/<left-follower>",
"calibration_dir": "/path/to/calibration/<left-follower-serial>"
},
{
"alias": "right_follower",
"type": "follower",
"port": "/dev/serial/by-id/<right-follower>",
"calibration_dir": "/path/to/calibration/<right-follower-serial>"
},
{
"alias": "left_leader",
"type": "leader",
"port": "/dev/serial/by-id/<left-leader>",
"calibration_dir": "/path/to/calibration/<left-leader-serial>"
},
{
"alias": "right_leader",
"type": "leader",
"port": "/dev/serial/by-id/<right-leader>",
"calibration_dir": "/path/to/calibration/<right-leader-serial>"
}
],
"cameras": [
{
"alias": "left_wrist",
"port": "/dev/v4l/by-path/<left-wrist>",
"width": 640,
"height": 480,
"fps": 30,
"fourcc": "MJPG"
},
{
"alias": "right_wrist",
"port": "/dev/v4l/by-path/<right-wrist>",
"width": 640,
"height": 480,
"fps": 30,
"fourcc": "MJPG"
},
{
"alias": "right_front",
"port": "/dev/v4l/by-path/<right-front>",
"width": 640,
"height": 480,
"fps": 30,
"fourcc": "MJPG"
}
]
}🧪 Training Pipeline
The typical RLT workflow has four stages. Video datasets require FFmpeg shared libraries for LeRobot 0.5.1 / TorchCodec decoding:
sudo apt-get update && sudo apt-get install -y ffmpegFor saved checkpoints, LeRobot 0.5.1 writes numeric checkpoint directories such as checkpoints/000001/pretrained_model. Use the latest numeric directory when checkpoints/last/pretrained_model is not present.
3) Finetune VLA
Use LeRobot's training entrypoint to finetune a pi0.5 VLA checkpoint on a LeRobot dataset.
python -m lerobot.scripts.lerobot_train \
--dataset.repo_id=<HF_ORG>/<DATASET> \
--dataset.root=<LOCAL_DATASET_ROOT> \
--policy.path=<BASE_PI05_CHECKPOINT_DIR> \
--policy.device=cuda \
--policy.dtype=bfloat16 \
--batch_size=16 \
--steps=30000 \
--save_freq=5000 \
--eval_freq=0 \
--tolerance_s=0.04 \
--output_dir=outputs/vla_ft \
--job_name=vla_ft4) Train RL Token
python -c 'from evo_rlt.adapters.lerobot import register; register(); from lerobot.scripts.lerobot_train import main; main()' \
--dataset.repo_id=<HF_ORG>/<DATASET> \
--dataset.root=<LOCAL_DATASET_ROOT> \
--policy.type=rlt_token \
--policy.repo_id=<HF_ORG>/rlt_token \
--policy.push_to_hub=false \
--policy.vla_pretrained_path=outputs/vla_ft/checkpoints/last/pretrained_model \
--policy.vla_dtype=bfloat16 \
--policy.rl_token_num_rl_tokens=1 \
--policy.tokenizer_path=/path/to/paligemma-3b-pt-224-snapshot \
--policy.token_pool_size=0 \
--policy.device=cuda \
--batch_size=8 \
--steps=10000 \
--save_freq=2000 \
--eval_freq=0 \
--tolerance_s=0.04 \
--output_dir=outputs/rl_token \
--job_name=rl_token5) Build Transition Cache
evo-rlt-build-transition-cache-v2 \
--demo-dataset-repo-id <HF_ORG>/<DATASET> \
--demo-dataset-root <LOCAL_DATASET_ROOT> \
--rl-token-policy-path outputs/rl_token/checkpoints/last/pretrained_model \
--vla-pretrained-path outputs/vla_ft/checkpoints/last/pretrained_model \
--tokenizer-path /path/to/paligemma-3b-pt-224-snapshot \
--output-dir outputs/cache \
--task-instruction "<TASK>" \
--chunk-length 10 \
--frame-stride 2 \
--batch-size 8 \
--num-workers 2 \
--train-ratio 0.9 \
--tolerance-s 0.04 \
--device cuda6) Train Chunk Actor-Critic
outputs/cache must contain chunk_transitions_train.pt. The Evo-RLT registry detects this cache directory through --dataset.repo_id.
python -c 'from evo_rlt.adapters.lerobot import register; register(); from lerobot.scripts.lerobot_train import main; main()' \
--dataset.repo_id=outputs/cache \
--policy.type=rlt_ac \
--policy.repo_id=<HF_ORG>/rlt_ac \
--policy.push_to_hub=false \
--policy.vla_pretrained_path=outputs/vla_ft/checkpoints/last/pretrained_model \
--policy.rl_token_pretrained_path=outputs/rl_token/checkpoints/last/pretrained_model \
--policy.vla_dtype=bfloat16 \
--policy.tokenizer_path=/path/to/paligemma-3b-pt-224-snapshot \
--policy.rl_token_num_rl_tokens=1 \
--policy.chunk_length=10 \
--policy.chunk_exec_steps=25 \
--policy.phase_mode=always_rl \
--policy.device=cuda \
--batch_size=256 \
--steps=50000 \
--save_freq=5000 \
--eval_freq=0 \
--output_dir=outputs/ac \
--job_name=rlt_ac🤖 Real-Robot Recording and Deployment
Set up the environment before running robot commands:
cd /path/to/evo-rlt
source ~/miniconda3/etc/profile.d/conda.sh
conda activate evo-rlt
python -m pip install -e ".[lerobot]"
export HF_HUB_OFFLINE=1Default VLA-RLT-VLA real-robot collection uses the official LeRobot 0.5.1 streaming encoder. The wrapper keeps the foreground recording loop responsive and expands the dataset settings to --dataset.vcodec=h264, --dataset.video_encoding_batch_size=<num_episodes + 1>, and --dataset.streaming_encoding=true.
Add --online-replay to evo-rlt-collect-default or evo-rlt-record segment
to collect normalized chunk transitions during robot recording. Successful
episodes receive reward 1 at the terminal chunk's last valid step; failed
episodes receive only zeros. Episodes enter replay only after their outcome is
confirmed and the recording is saved; rerecorded episodes are discarded.
The cache is saved atomically to <DATASET_ROOT>/online_replay/chunk_transitions_train.pt
and can be passed directly to actor-critic training via --dataset.repo_id.
Resuming recording reloads the existing replay cache (up to 200,000 chunks).
Online replay uses the deployed SFT preprocessor for executed actions, including human intervention, and encodes the current observation/reference every chunk boundary. This adds a VLA forward pass per chunk and can lower control frequency. With RTC, replay encoding shares the inference lock but leaves action and metadata queues intact. This collects training data; optimizer updates still run through the actor-critic trainer. ACP inference is unsupported with online replay.
Shared collection arguments:
COMMON_ARGS=(
--setup-json /path/to/robot_manifest.json \
--policy-path /path/to/rlt_ac_policy \
--vla-path /path/to/pi05_vla_checkpoint_or_dir \
--rl-token-path /path/to/rl_token_policy \
--dataset-tag vla_rlt_vla_test \
--num-episodes 5 \
--episode-time-s 3000 \
--fps 30 \
--vcodec h264 \
--rlt-toggle-key r \
--teleop-toggle-key space
)Start in VLA mode and record the full trajectory:
evo-rlt-record collect "${COMMON_ARGS[@]}"Start in VLA mode and record only the critical segment:
evo-rlt-record collect "${COMMON_ARGS[@]}" --only-criticalStart in teleoperation mode and record the full trajectory:
evo-rlt-record collect "${COMMON_ARGS[@]}" --start-with-teleopStart in teleoperation mode and record only the critical segment:
evo-rlt-record collect "${COMMON_ARGS[@]}" --start-with-teleop --only-criticalThe same collection entrypoint is exposed as evo-rlt-collect-default after reinstalling package entry points, but checkpoint and setup paths still need to be supplied by the caller.
Validated RTC defaults for this collection mode:
RLT RTC execution horizon: 10
VLA RTC execution horizon: 25
RTC action queue refill threshold: 30
RTC max guidance weight: 10.0
RTC prefix attention schedule: EXPDefault collection controls:
Full-trajectory mode:
r save the full episode as success after the double-tap window
r+r save the full episode as failure
space toggle teleop intervention; pressing again exits teleop
left arrow rerecord the current episode
Esc stop data collection
Critical-segment mode (`--only-critical`):
r enter RLT mode and start recording the critical segment
r save the segment as success, exit RLT mode, then end the episode
r+r save the segment as failure, exit RLT mode, then end the episode
space toggle teleop intervention; pressing again exits teleop
left arrow rerecord the current episode
Esc stop data collectionVLA-only full-process recording with pedal outcome labels:
evo-rlt-record full \
--initial-source vla \
--setup-json <ROBOT_SETUP_JSON> \
--policy-path <AC_OR_VLA_POLICY_PATH> \
--vla-path <BASE_OR_FINETUNED_VLA_PT> \
--phase-mode always_vla \
--chunk-exec-steps 25 \
--pedal-outcome \
--double-tap-window-s 0.6 \
--num-episodes 5 \
--episode-time-s 3000 \
--reset-time-s 0 \
--fps 30 \
--vcodec h264 \
--dataset-tag vla_full_pedal \
--no-teleopFor headless SSH runs where no keyboard or pedal outcome will be provided, add
--default-episode-success success or --default-episode-success failure.
Pedal semantics in this mode:
single tap success, end current episode, start next episode
double tap failure, end current episode, start next episode🗂️ Repository Layout
src/evo_rlt/core # algorithm core, torch-only
src/evo_rlt/adapters/lerobot # LeRobot/pi0.5/dataset/policy/record adapters
src/evo_rlt/cli # training and cache CLIs
tests/rlt # focused RLT unit and integration tests✅ Development Checks
PYTHONPATH=src pytest -q tests/rlt
PYTHONPATH=src python -m compileall -q src/evo_rlt tests/rlt🤗 Model & Dataset
- Training dataset: Elvinky/bi-so101-insert-screw-562ep.
- Real-world RL dataset: MINT-SJTU/RW-RL-Dataset.
- Checkpoint repo: Shiki42/pi05_screw_c_mix_cont15k_fp16.
🧭 Future TODO
- PiPER/PiPER-X real-robot deployment support.
💬 Community Channels
- Email: business@evomind-tech.com
- WeChat group QR code 加群(注明加真机强化学习群):
🏫 Affiliations
📄 License
Apache-2.0. See LICENSE.