We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with uncertainty, it generates multiple plausible 3D gaze and pose hypotheses based on the 2D context information extracted from the input image. Specifically, we condition the denoising process on the 2D pose, the surroundings of the subject, and the context of the scene. With GazeD we also introduce a novel way of representing the 3D gaze by positioning it as an additional body joint at a fixed distance from the eyes. The rationale is that the gaze is usually closely related to the pose, and thus it can benefit from being jointly denoised during the diffusion process. Evaluations across three benchmark datasets demonstrate that GazeD achieves state-of-the-art performance in 3D gaze estimation, even surpassing methods that rely on temporal information. Project details will be available at https://aimagelab.ing.unimore.it/go/gazed

GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation / Catalini, R., Di Nucci, D., Borghi, G., Davoli, D., Garattoni, L., Francesca, G., Kawana, Y., Vezzani, R.. - (2026), pp. 760-770. (13th International Conference on 3D Vision, 3DV 2026 can 2026) [10.1109/3dv69130.2026.00078].

GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation

Catalini, Riccardo;Di Nucci, Davide;Borghi, Guido;Vezzani, Roberto
2026

Abstract

We introduce GazeD, a new 3D gaze estimation method that jointly provides 3D gaze and human pose from a single RGB image. Leveraging the ability of diffusion models to deal with uncertainty, it generates multiple plausible 3D gaze and pose hypotheses based on the 2D context information extracted from the input image. Specifically, we condition the denoising process on the 2D pose, the surroundings of the subject, and the context of the scene. With GazeD we also introduce a novel way of representing the 3D gaze by positioning it as an additional body joint at a fixed distance from the eyes. The rationale is that the gaze is usually closely related to the pose, and thus it can benefit from being jointly denoised during the diffusion process. Evaluations across three benchmark datasets demonstrate that GazeD achieves state-of-the-art performance in 3D gaze estimation, even surpassing methods that rely on temporal information. Project details will be available at https://aimagelab.ing.unimore.it/go/gazed
2026
13th International Conference on 3D Vision, 3DV 2026
can
2026
760
770
Catalini, Riccardo; Di Nucci, Davide; Borghi, Guido; Davoli, Davide; Garattoni, Lorenzo; Francesca, Giampiero; Kawana, Yuki; Vezzani, Roberto
GazeD: Context-Aware Diffusion for Accurate 3D Gaze Estimation / Catalini, R., Di Nucci, D., Borghi, G., Davoli, D., Garattoni, L., Francesca, G., Kawana, Y., Vezzani, R.. - (2026), pp. 760-770. (13th International Conference on 3D Vision, 3DV 2026 can 2026) [10.1109/3dv69130.2026.00078].
File in questo prodotto:
File Dimensione Formato  
3DV_2026.pdf

Open access

Tipologia: AAM - Versione dell'autore revisionata e accettata per la pubblicazione
Dimensione 1.92 MB
Formato Adobe PDF
1.92 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

Licenza Creative Commons
I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11380/1415309
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact