Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation

Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation. Martínez-González, A., Villamizar, M., Canévet, O., & Odobez, J. IEEE Transactions on Circuits and Systems for Video Technology, 30(11):4207-4221, Institute of Electrical and Electronics Engineers Inc., 12, 2019.

Paper

Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation [link]

Website doi abstract bibtex

Achieving robust multi-person 2D body landmark localization and pose estimation is essential for human behavior and interaction understanding as encountered for instance in HRI settings. Accurate methods have been proposed recently, but they usually rely on rather deep Convolutional Neural Network (CNN) architecture, thus requiring large computational and training resources. In this paper, we investigate different architectures and methodologies to address these issues and achieve fast and accurate multi-person 2D pose estimation. To foster speed, we propose to work with depth images, whose structure contains sufficient information about body landmarks while being simpler than textured color images and thus potentially requiring less complex CNNs for processing. In this context, we make the following contributions. i) we study several CNN architecture designs combining pose machines relying on the cascade of detectors concept with lightweight and efficient CNN structures; ii) to address the need for large training datasets with high variability, we rely on semi-synthetic data combining multi-person synthetic depth data with real sensor backgrounds; iii) we explore domain adaptation techniques to address the performance gap introduced by testing on real depth images; iv) to increase the accuracy of our fast lightweight CNN models, we investigate knowledge distillation at several architecture levels which effectively enhance performance. Experiments and results on synthetic and real data highlight the impact of our design choices, providing insights into methods addressing standard issues normally faced in practical applications, and resulting in architectures effectively matching our goal in both performance and speed.

@article{
 title = {Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation},
 type = {article},
 year = {2019},
 keywords = {Human pose estimation,convolutional neural networks,machine learning},
 pages = {4207-4221},
 volume = {30},
 websites = {http://arxiv.org/abs/1912.00711,http://dx.doi.org/10.1109/TCSVT.2019.2952779},
 month = {12},
 publisher = {Institute of Electrical and Electronics Engineers Inc.},
 day = {2},
 id = {ff668c66-532b-3e71-b367-b78c67ae7a0b},
 created = {2023-10-27T11:16:33.484Z},
 accessed = {2023-10-27},
 file_attached = {true},
 profile_id = {f1f70cad-e32d-3de2-a3c0-be1736cb88be},
 group_id = {5ec9cc91-a5d6-3de5-82f3-3ef3d98a89c1},
 last_modified = {2023-10-27T11:16:41.049Z},
 read = {false},
 starred = {false},
 authored = {false},
 confirmed = {false},
 hidden = {false},
 folder_uuids = {bd3c6f2e-3514-47cf-bc42-12db8b9abe45},
 private_publication = {false},
 abstract = {Achieving robust multi-person 2D body landmark localization and pose estimation is essential for human behavior and interaction understanding as encountered for instance in HRI settings. Accurate methods have been proposed recently, but they usually rely on rather deep Convolutional Neural Network (CNN) architecture, thus requiring large computational and training resources. In this paper, we investigate different architectures and methodologies to address these issues and achieve fast and accurate multi-person 2D pose estimation. To foster speed, we propose to work with depth images, whose structure contains sufficient information about body landmarks while being simpler than textured color images and thus potentially requiring less complex CNNs for processing. In this context, we make the following contributions. i) we study several CNN architecture designs combining pose machines relying on the cascade of detectors concept with lightweight and efficient CNN structures; ii) to address the need for large training datasets with high variability, we rely on semi-synthetic data combining multi-person synthetic depth data with real sensor backgrounds; iii) we explore domain adaptation techniques to address the performance gap introduced by testing on real depth images; iv) to increase the accuracy of our fast lightweight CNN models, we investigate knowledge distillation at several architecture levels which effectively enhance performance. Experiments and results on synthetic and real data highlight the impact of our design choices, providing insights into methods addressing standard issues normally faced in practical applications, and resulting in architectures effectively matching our goal in both performance and speed.},
 bibtype = {article},
 author = {Martínez-González, Angel and Villamizar, Michael and Canévet, Olivier and Odobez, Jean-Marc},
 doi = {10.1109/TCSVT.2019.2952779},
 journal = {IEEE Transactions on Circuits and Systems for Video Technology},
 number = {11}
}

Downloads: 0

{"_id":"6MDrgMTxtTr62axXy","bibbaseid":"martnezgonzlez-villamizar-canvet-odobez-efficientconvolutionalneuralnetworksfordepthbasedmultipersonposeestimation-2019","author_short":["Martínez-González, A.","Villamizar, M.","Canévet, O.","Odobez, J."],"bibdata":{"title":"Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation","type":"article","year":"2019","keywords":"Human pose estimation,convolutional neural networks,machine learning","pages":"4207-4221","volume":"30","websites":"http://arxiv.org/abs/1912.00711,http://dx.doi.org/10.1109/TCSVT.2019.2952779","month":"12","publisher":"Institute of Electrical and Electronics Engineers Inc.","day":"2","id":"ff668c66-532b-3e71-b367-b78c67ae7a0b","created":"2023-10-27T11:16:33.484Z","accessed":"2023-10-27","file_attached":"true","profile_id":"f1f70cad-e32d-3de2-a3c0-be1736cb88be","group_id":"5ec9cc91-a5d6-3de5-82f3-3ef3d98a89c1","last_modified":"2023-10-27T11:16:41.049Z","read":false,"starred":false,"authored":false,"confirmed":false,"hidden":false,"folder_uuids":"bd3c6f2e-3514-47cf-bc42-12db8b9abe45","private_publication":false,"abstract":"Achieving robust multi-person 2D body landmark localization and pose estimation is essential for human behavior and interaction understanding as encountered for instance in HRI settings. Accurate methods have been proposed recently, but they usually rely on rather deep Convolutional Neural Network (CNN) architecture, thus requiring large computational and training resources. In this paper, we investigate different architectures and methodologies to address these issues and achieve fast and accurate multi-person 2D pose estimation. To foster speed, we propose to work with depth images, whose structure contains sufficient information about body landmarks while being simpler than textured color images and thus potentially requiring less complex CNNs for processing. In this context, we make the following contributions. i) we study several CNN architecture designs combining pose machines relying on the cascade of detectors concept with lightweight and efficient CNN structures; ii) to address the need for large training datasets with high variability, we rely on semi-synthetic data combining multi-person synthetic depth data with real sensor backgrounds; iii) we explore domain adaptation techniques to address the performance gap introduced by testing on real depth images; iv) to increase the accuracy of our fast lightweight CNN models, we investigate knowledge distillation at several architecture levels which effectively enhance performance. Experiments and results on synthetic and real data highlight the impact of our design choices, providing insights into methods addressing standard issues normally faced in practical applications, and resulting in architectures effectively matching our goal in both performance and speed.","bibtype":"article","author":"Martínez-González, Angel and Villamizar, Michael and Canévet, Olivier and Odobez, Jean-Marc","doi":"10.1109/TCSVT.2019.2952779","journal":"IEEE Transactions on Circuits and Systems for Video Technology","number":"11","bibtex":"@article{\n title = {Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation},\n type = {article},\n year = {2019},\n keywords = {Human pose estimation,convolutional neural networks,machine learning},\n pages = {4207-4221},\n volume = {30},\n websites = {http://arxiv.org/abs/1912.00711,http://dx.doi.org/10.1109/TCSVT.2019.2952779},\n month = {12},\n publisher = {Institute of Electrical and Electronics Engineers Inc.},\n day = {2},\n id = {ff668c66-532b-3e71-b367-b78c67ae7a0b},\n created = {2023-10-27T11:16:33.484Z},\n accessed = {2023-10-27},\n file_attached = {true},\n profile_id = {f1f70cad-e32d-3de2-a3c0-be1736cb88be},\n group_id = {5ec9cc91-a5d6-3de5-82f3-3ef3d98a89c1},\n last_modified = {2023-10-27T11:16:41.049Z},\n read = {false},\n starred = {false},\n authored = {false},\n confirmed = {false},\n hidden = {false},\n folder_uuids = {bd3c6f2e-3514-47cf-bc42-12db8b9abe45},\n private_publication = {false},\n abstract = {Achieving robust multi-person 2D body landmark localization and pose estimation is essential for human behavior and interaction understanding as encountered for instance in HRI settings. Accurate methods have been proposed recently, but they usually rely on rather deep Convolutional Neural Network (CNN) architecture, thus requiring large computational and training resources. In this paper, we investigate different architectures and methodologies to address these issues and achieve fast and accurate multi-person 2D pose estimation. To foster speed, we propose to work with depth images, whose structure contains sufficient information about body landmarks while being simpler than textured color images and thus potentially requiring less complex CNNs for processing. In this context, we make the following contributions. i) we study several CNN architecture designs combining pose machines relying on the cascade of detectors concept with lightweight and efficient CNN structures; ii) to address the need for large training datasets with high variability, we rely on semi-synthetic data combining multi-person synthetic depth data with real sensor backgrounds; iii) we explore domain adaptation techniques to address the performance gap introduced by testing on real depth images; iv) to increase the accuracy of our fast lightweight CNN models, we investigate knowledge distillation at several architecture levels which effectively enhance performance. Experiments and results on synthetic and real data highlight the impact of our design choices, providing insights into methods addressing standard issues normally faced in practical applications, and resulting in architectures effectively matching our goal in both performance and speed.},\n bibtype = {article},\n author = {Martínez-González, Angel and Villamizar, Michael and Canévet, Olivier and Odobez, Jean-Marc},\n doi = {10.1109/TCSVT.2019.2952779},\n journal = {IEEE Transactions on Circuits and Systems for Video Technology},\n number = {11}\n}","author_short":["Martínez-González, A.","Villamizar, M.","Canévet, O.","Odobez, J."],"urls":{"Paper":"https://bibbase.org/service/mendeley/bfbbf840-4c42-3914-a463-19024f50b30c/file/633754b9-f893-07d5-3aa0-dee9d29512ee/full_text.pdf.pdf","Website":"http://arxiv.org/abs/1912.00711,http://dx.doi.org/10.1109/TCSVT.2019.2952779"},"biburl":"https://bibbase.org/service/mendeley/bfbbf840-4c42-3914-a463-19024f50b30c","bibbaseid":"martnezgonzlez-villamizar-canvet-odobez-efficientconvolutionalneuralnetworksfordepthbasedmultipersonposeestimation-2019","role":"author","keyword":["Human pose estimation","convolutional neural networks","machine learning"],"metadata":{"authorlinks":{}},"downloads":0},"bibtype":"article","biburl":"https://bibbase.org/service/mendeley/bfbbf840-4c42-3914-a463-19024f50b30c","dataSources":["2252seNhipfTmjEBQ"],"keywords":["human pose estimation","convolutional neural networks","machine learning"],"search_terms":["efficient","convolutional","neural","networks","depth","based","multi","person","pose","estimation","martínez-gonzález","villamizar","canévet","odobez"],"title":"Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation","year":2019}