Learnable Triangulation Pytorch
This repository is an official PyTorch implementation of the paper "Learnable Triangulation of Human Pose" (ICCV 2019, oral). Proposed method archives state-of-the-art results in multi-view 3D human pose estimation!
Install / Use
npx skills add karfly/learnable-triangulation-pytorchInstalls into whichever agent you are using.
README
Learnable Triangulation of Human Pose
This repository is an official PyTorch implementation of the paper "Learnable Triangulation of Human Pose" (ICCV 2019, oral). Here we tackle the problem of 3D human pose estimation from multiple cameras. We present 2 novel methods — Algebraic and Volumetric learnable triangulation — that outperform previous state of the art.
If you find a bug, have a question or know to improve the code - please open an issue!
:arrow_forward: ICCV 2019 talk
<p align="center"> <a href="http://www.youtube.com/watch?v=z3f3aPSuhqg"> <img width=680 src="docs/video-preview.jpg"> </a> </p>How to use
This project doesn't have any special or difficult-to-install dependencies. All installation can be done with:
pip install -r requirements.txt
Data
Sorry, only Human3.6M dataset training/evaluation is available right now. We cannot add CMU Panoptic, sorry for that.
Human3.6M
- Download and preprocess the dataset by following the instructions in mvn/datasets/human36m_preprocessing/README.md.
- Download pretrained backbone's weights from here and place them here:
./data/pretrained/human36m/pose_resnet_4.5_pixels_human36m.pth(ResNet-152 trained on COCO dataset and finetuned jointly on MPII and Human3.6M). - If you want to train Volumetric model, you need rough estimations of the pelvis' 3D positions both for train and val splits. In the paper we estimate them using the Algebraic model. You can use the pretrained Algebraic model to produce predictions or just take precalculated 3D skeletons.
Model zoo
In this section we collect pretrained models and configs. All pretrained weights and precalculated 3D skeletons can be downloaded at once from here and placed to ./data/pretrained, so that eval configs can work out-of-the-box (without additional setting of paths). Alternatively, the table below provides separate links to those files.
Human3.6M:
| Model | Train config | Eval config | Weights | Precalculated results | MPJPE (relative to pelvis), mm | |----------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------:|:------------------------------------------------------------------------------------------------------------------------------------------------------:|-------------------------------:| | Algebraic | train/human36m_alg.yaml | eval/human36m_alg.yaml | link | train, val | 22.5 | | Volumetric (softmax) | train/human36m_vol_softmax.yaml | eval/human36m_vol_softmax.yaml | link | — | 20.4 |
Train
Every experiment is defined by .config files. Configs with experiments from the paper can be found in the ./experiments directory (see model zoo).
Single-GPU
To train a Volumetric model with softmax aggregation using 1 GPU, run:
python3 train.py \
--config experiments/human36m/train/human36m_vol_softmax.yaml \
--logdir ./logs
The training will start with the config file specified by --config, and logs (including tensorboard files) will be stored in --logdir.
Multi-GPU (in testing)
Multi-GPU training is implemented with PyTorch's DistributedDataParallel. It can be used both for single-machine and multi-machine (cluster) training. To run the processes use the PyTorch launch utility.
To train a Volumetric model with softmax aggregation using 2 GPUs on single machine, run:
python3 -m torch.distributed.launch --nproc_per_node=2 --master_port=2345 \
train.py \
--config experiments/human36m/train/human36m_vol_softmax.yaml \
--logdir ./logs
Tensorboard
To watch your experiments' progress, run tensorboard:
tensorboard --logdir ./logs
Evaluation
After training, you can evaluate the model. Inside the same config file, add path to the learned weights (they are dumped to logs dir during training):
model:
init_weights: true
checkpoint: {PATH_TO_WEIGHTS}
Also, you can change other config parameters like retain_every_n_frames_test.
Run:
python3 train.py \
--eval --eval_dataset val \
--config experiments/human36m/eval/human36m_vol_softmax.yaml \
--logdir ./logs
Argument --eval_dataset can be val or train. Results can be seen in logs directory or in the tensorboard.
Results
- We conduct experiments on two available large multi-view datasets: Human3.6M [2] and CMU Panoptic [3].
- The main metric is MPJPE (Mean Per Joint Position Error) which is L2 distance averaged over all joints.
Human3.6M
- We significantly improved upon the previous state of the art (error is measured relative to pelvis, without alignment).
- Our best model reaches 17.7 mm error in absolute coordinates, which was unattainable before.
- Our Volumetric model is able to estimate 3D human pose using any number of cameras, even using only 1 camera. In single-view setup, we get results comparable to current state of the art [6] (49.9 mm vs. 49.6 mm).
| | MPJPE (averaged across all actions), mm | |----------------------------- |:--------: | | Multi-View Martinez [4] | 57.0 | | Pavlakos et al. [8] | 56.9 | | Tome et al. [4] | 52.8 | | Kadkhodamohammadi & Padoy [5] | 49.1 | | Qiu et al. [9] | 26.2 | | RANSAC (our implementation) | 27.4 | | Ours, algebraic | 22.4 | | Ours, volumetric | 20.5 |
<br> MPJPE absolute (scenes with invalid ground-truth annotations are excluded):| | MPJPE (averaged across all actions), mm | |----------------------------- |:--------: | | RANSAC (our implementation) | 22.8 | | Ours, algebraic | 19.2 | | Ours, volumetric | 17.7 |
<br> MPJPE relative to pelvis (single-view methods):| | MPJPE (averaged across all actions), mm | |----------------------------- |:-----------------------------------: | | Martinez et al. [7] | 62.9 | | Sun et al. [6] | 49.6 | | Ours, volumetric single view | 49.9 |
CMU Panoptic
- Our best model reaches 13.7 mm error in absolute coordinates for 4 cameras
- We managed to get much smoother and more accurate 3D pose annotations compared to dataset annotations (see video demonstration)
| | MPJPE, mm | |----------------------------- |:--------: | | RANSAC (our implementation) | 39.5 | | Ours, algebraic | 21.3 | | Ours, volumetric | 13.7 |
Method overview
We present 2 novel methods of learnable triangulation: Algebraic and Volumetric.
Algebraic
Our first method is based on Algebraic triangulation. It is similar to the previous approaches, but differs in 2 critical aspects:
- It is fully differentiable. To achieve this, we use soft-argmax aggregation and triangulate keypoints via a differentiable SVD.
Related Skills
node-connect
385.5kDiagnose OpenClaw Android, iOS, or macOS node pairing, QR/setup code, route, auth, and connection failures.
blender-python-addon
40.5kBlender Python add-on rules for operators, panels, properties, registration, testing, and API-safe scripting
flutter-development-guidelines-cursorrules-prompt-file
40.5kCursor rules for Flutter development with MVVM architecture, Riverpod state management, Material widgets, and Dart style guidelines.
commit-push-pr
140.7kCommit, push, and open a PR
