Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular, these issues are caused by 2D joint locations that can be mapped to multiple 3D positions, inducing multiple possible final poses. Following these considerations, we propose leveraging diffusion-based models’ generation capability to predict multiple hypotheses and aggregate them in a final accurate pose. Therefore, we introduce SnapPose3D, a pose-lifting framework trained deterministically to denoise 3D poses conditioned on both visual context and 2D pose features. SnapPose3D adopts a probabilistic approach during inference, generating multiple hypotheses through random sampling from a unit Gaussian distribution. Unlike most previous methods that address pose ambiguity by processing temporal sequences, SnapPose3D uses single frames as input, avoiding tracking and limiting computational cost, data acquisition complexity, and the need for online, real-time applications. We extensively evaluate SnapPose3D on well-known benchmarks for the 3D human pose estimation task showing its ability to generate and aggregate accurate hypotheses that lead to state-of-the-art results.

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses / Simoni, A., Catalini, R., Di Nucci, D., Borghi, G., Davoli, D., Garattoni, L., Francesca, G., Kawana, Y., Vezzani, R.. - 16814:(2026), pp. 240-254. (28th International Conference on Pattern Recognition, ICPR 2026 Lyon, France 2026) [10.1007/978-3-032-31654-7_17].

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

Simoni, Alessandro;Catalini, Riccardo;Di Nucci, Davide;Borghi, Guido;Davoli, Davide;Vezzani, Roberto
2026

Abstract

Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular, these issues are caused by 2D joint locations that can be mapped to multiple 3D positions, inducing multiple possible final poses. Following these considerations, we propose leveraging diffusion-based models’ generation capability to predict multiple hypotheses and aggregate them in a final accurate pose. Therefore, we introduce SnapPose3D, a pose-lifting framework trained deterministically to denoise 3D poses conditioned on both visual context and 2D pose features. SnapPose3D adopts a probabilistic approach during inference, generating multiple hypotheses through random sampling from a unit Gaussian distribution. Unlike most previous methods that address pose ambiguity by processing temporal sequences, SnapPose3D uses single frames as input, avoiding tracking and limiting computational cost, data acquisition complexity, and the need for online, real-time applications. We extensively evaluate SnapPose3D on well-known benchmarks for the 3D human pose estimation task showing its ability to generate and aggregate accurate hypotheses that lead to state-of-the-art results.
2026
28th International Conference on Pattern Recognition, ICPR 2026
Lyon, France
2026
16814
240
254
Simoni, Alessandro; Catalini, Riccardo; Di Nucci, Davide; Borghi, Guido; Davoli, Davide; Garattoni, Lorenzo; Francesca, Gianpiero; Kawana, Yuki; Vezza...espandi
SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses / Simoni, A., Catalini, R., Di Nucci, D., Borghi, G., Davoli, D., Garattoni, L., Francesca, G., Kawana, Y., Vezzani, R.. - 16814:(2026), pp. 240-254. (28th International Conference on Pattern Recognition, ICPR 2026 Lyon, France 2026) [10.1007/978-3-032-31654-7_17].
File in questo prodotto:
File Dimensione Formato  
SnapPose_icpr.pdf

Open access

Tipologia: AAM - Versione dell'autore revisionata e accettata per la pubblicazione
Dimensione 556.35 kB
Formato Adobe PDF
556.35 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

Licenza Creative Commons
I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11380/1416349
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact