Video captioning has picked up a considerable attention thanks to the ability of Recurrent Neural Networks to extrapolate an encoded representation of the input video, and then use it to generate a description. We propose a recurrent encoding approach able to find and exploit the layered design of the video. Differently from the established encoder-decoder procedure, in which a video is repeatedly encoded by a recurrent layer, we employ revised Quasi-Recurrent Neural Networks. We further extend their basic cell with a boundary detector in order to recognize discontinuous segments boundaries and likewise correct the temporal connections of the encoding layer accordingly. Experiments, on the Montreal Video Annotation dataset, demonstrate that our approach can find suitable levels of representation of the input information, while reducing the computational requirements.
A Hierarchical Quasi-Recurrent approach to Video Captioning / Bolelli, F., Baraldi, L., Grana, C.. - (2018), pp. 162-167. (2018 IEEE International Conference on Image Processing, Applications and Systems (IPAS) Inria Sophia Antipolis, France Dec 12-14) [10.1109/IPAS.2018.8708893].
A Hierarchical Quasi-Recurrent approach to Video Captioning
BOLELLI, FEDERICO
;Baraldi, Lorenzo;Grana, Costantino
2018
Abstract
Video captioning has picked up a considerable attention thanks to the ability of Recurrent Neural Networks to extrapolate an encoded representation of the input video, and then use it to generate a description. We propose a recurrent encoding approach able to find and exploit the layered design of the video. Differently from the established encoder-decoder procedure, in which a video is repeatedly encoded by a recurrent layer, we employ revised Quasi-Recurrent Neural Networks. We further extend their basic cell with a boundary detector in order to recognize discontinuous segments boundaries and likewise correct the temporal connections of the encoding layer accordingly. Experiments, on the Montreal Video Annotation dataset, demonstrate that our approach can find suitable levels of representation of the input information, while reducing the computational requirements.| File | Dimensione | Formato | |
|---|---|---|---|
|
2018_IPAS_A_Hierarchical_Quasi_Recurrent_Approach_to_Video_Captioning.pdf
Open access
Tipologia:
AO - Versione originale dell'autore proposta per la pubblicazione
Dimensione
971.05 kB
Formato
Adobe PDF
|
971.05 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate

I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris





