🎯 ProxyPose
ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation
Ruihang Zhang*1, Felix Taubner*1,2, Pooja Ravi1, Kiriakos N. Kutulakos1,2, David B. Lindell1,2
1University of Toronto 2Vector Institute *Equal contribution
TL;DR: One query pixel in, a full 6‑DoF pose trajectory out.
⚡️ Quick start
🛠️ 1. Install
# Clone the repository
git clone https://github.com/ruihangzhang97/proxypose.git
cd proxypose
# Create and activate a conda environment
conda create -n proxypose python=3.10 -y
conda activate proxypose
# Install PyTorch
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
# Install PyTorch3D
export FORCE_CUDA=1
pip install --no-build-isolation git+https://github.com/facebookresearch/pytorch3d.git@stable
# Install ProxyPose and all remaining dependencies
pip install -e .📦 2. Download model weights
The weights are downloaded automatically on first run — no manual steps needed.
We provide two model sizes. The 14B model gives the best quality; the 1.3B model is much smaller and faster, useful for quick experiments or lower-VRAM GPUs.
| Weight | HuggingFace | Size |
|---|---|---|
| Wan2.1-T2V-14B (base) | Wan-AI/Wan2.1-T2V-14B | ~30 GB |
| ProxyPose LoRA (14B) | ruihangzhang79/proxypose | ~600 MB |
| Wan2.1-T2V-1.3B (base) | Wan-AI/Wan2.1-T2V-1.3B | ~7 GB |
| ProxyPose LoRA (1.3B) | ruihangzhang79/proxypose | ~175 MB |
By default, all commands below use the 14B model (configs/generation/default.yaml). To use
the 1.3B model instead, pass --gen_config configs/generation/wan1.3b.yaml to proxypose-infer.
🖱️ 3. Pick your prompt point
proxypose-annotate --input-video video/my_video.mp4Opens http://localhost:7860 in your browser. Click on a query point in the first frame, press Save. Coordinates are written to video/my_video.points.json.
🔭 4. (Optional) Estimate focal length with Depth Anything 3
By default, we assume a 45° horizontal field of view. For improved accuracy, consider using Depth Anything 3.
First, install Depth Anything 3 inside your ProxyPose environment:
git clone https://github.com/ByteDance-Seed/Depth-Anything-3.git
cd Depth-Anything-3
pip install --no-build-isolation -e .Then, from the ProxyPose repository, run:
# Requires: pip install hatchling editables
python -m inference.annotation.depth_anything video/my_video.mp4✅ 5. Run inference
Pass the query JSON saved by the annotator directly to --prompt:
proxypose-infer \
--video_path video/my_video.mp4 \
--output_path output/result.mp4 \
--prompt video/my_video.points.json \
--depth_anything_path video/my_video.da3.npz # optional, omit to use fixed 45° FOVTo use the 1.3B model instead of the default 14B model, add --gen_config configs/generation/wan1.3b.yaml.
📊 Evaluation
Benchmark windows (evaluation/benchmarks/), preprocessing and metrics used in the paper. Benchmarks are expected in a uniform format (see evaluation/reformat_ho3d.py).
# Write evaluation windows and prompts into each scene
python -m evaluation.write_frame_meta --filter_json evaluation/benchmarks/ho3d/w1_f49.json --uniform_root <uniform_root>
python -m evaluation.sample_prompts --benchmark_path <uniform_root>
# Run ProxyPose, then compute metrics
proxypose-eval --benchmark_path <uniform_root> --output_path output/benchmark
bash evaluation/run_eval.sh <uniform_root> output/benchmark/ours_results_attempt_01.json oursTo evaluate your own method, write a JSON list of {"scene_name", "obj_id", "frame_idx", "pose_4x4"} entries, with frame_idx relative to the window's anchor frame (0–48).
📖 Citation
@inproceedings{zhang2026proxypose,
title={ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation},
author={Ruihang Zhang and Felix Taubner and Pooja Ravi and Kiriakos N. Kutulakos and David B. Lindell},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}