The digitization of historical manuscripts presents significant challenges for Handwritten Text Recognition (HTR) systems, particularly when dealing with small, author-specific collections that diverge from the training data distributions. Handwritten Text Generation (HTG) techniques, which generate synthetic data tailored to specific handwriting styles, offer a promising solution to address these challenges. However, the effectiveness of various HTG models in enhancing HTR performance, especially in low-resource transcription settings, has not been thoroughly evaluated. In this work, we systematically compare three state-of-the-art styled HTG models (representing the generative adversarial, diffusion, and autoregressive paradigms for HTG) to assess their impact on HTR fine-tuning. We analyze how visual and linguistic characteristics of synthetic data influence fine-tuning outcomes and provide quantitative guide-lines for selecting the most effective HTG model. The results of our analysis provide insights into the current capabilities of HTG methods and highlight key areas for further improvement in their application to low-resource HTR.

Quo Vadis Handwritten Text Generation for Handwritten Text Recognition? / Pippi, V., Nikolaidou, K., Cascianelli, S., Retsinas, G., Sfikas, G., Cucchiara, R., Liwicki, M.. - (2025), pp. 7523-7533. (2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025 usa 2025) [10.1109/iccvw69036.2025.00775].

Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?

Pippi, Vittorio;Cascianelli, Silvia;Cucchiara, Rita;
2025

Abstract

The digitization of historical manuscripts presents significant challenges for Handwritten Text Recognition (HTR) systems, particularly when dealing with small, author-specific collections that diverge from the training data distributions. Handwritten Text Generation (HTG) techniques, which generate synthetic data tailored to specific handwriting styles, offer a promising solution to address these challenges. However, the effectiveness of various HTG models in enhancing HTR performance, especially in low-resource transcription settings, has not been thoroughly evaluated. In this work, we systematically compare three state-of-the-art styled HTG models (representing the generative adversarial, diffusion, and autoregressive paradigms for HTG) to assess their impact on HTR fine-tuning. We analyze how visual and linguistic characteristics of synthetic data influence fine-tuning outcomes and provide quantitative guide-lines for selecting the most effective HTG model. The results of our analysis provide insights into the current capabilities of HTG methods and highlight key areas for further improvement in their application to low-resource HTR.
2025
2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
usa
2025
7523
7533
Pippi, Vittorio; Nikolaidou, Konstantina; Cascianelli, Silvia; Retsinas, George; Sfikas, Giorgos; Cucchiara, Rita; Liwicki, Marcus
Quo Vadis Handwritten Text Generation for Handwritten Text Recognition? / Pippi, V., Nikolaidou, K., Cascianelli, S., Retsinas, G., Sfikas, G., Cucchiara, R., Liwicki, M.. - (2025), pp. 7523-7533. (2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025 usa 2025) [10.1109/iccvw69036.2025.00775].
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

Licenza Creative Commons
I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11380/1414450
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 1
  • ???jsp.display-item.citation.isi??? 0
  • OpenAlex 1
social impact