Diffusion-based Handwritten Text Generation (HTG) approaches achieve impressive results on frequent, invocabulary words observed at training time and on regular styles. However, they are prone to memorizing training samples and often struggle with style variability and generation clarity. In particular, standard diffusion models tend to produce artifacts or distortions that negatively affect the readability of the generated text, especially when the style is hard to produce. To tackle these issues, we propose a novel sampling guidance strategy, Dual Orthogonal Guidance (DOG), that leverages an orthogonal projection of a negatively perturbed prompt onto the original positive prompt. This approach helps steer the generation away from artifacts while maintaining the intended content, and encourages more diverse, yet plausible, outputs. Unlike standard Classifier-Free Guidance (CFG), which relies on unconditional predictions and produces noise at high guidance scales, DOG introduces a more stable, disentangled direction in the latent space. To control the strength of the guidance across the denoising process, we apply a triangular schedule: weak at the start and end of denoising, when the process is most sensitive, and strongest in the middle steps. Experimental results on the state-of-the-art DiffusionPen and One-DM demonstrate that DOG improves both content clarity and style variability, even for out-of-vocabulary words and challenging writing styles.

Dual Orthogonal Guidance for Robust Diffusion-Based Handwritten Text Generation / Nikolaidou, K., Retsinas, G., Sfikas, G., Cascianelli, S., Cucchiara, R., Liwicki, M.. - (2025), pp. 7020-7029. (2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025 usa 2025) [10.1109/iccvw69036.2025.00725].

Dual Orthogonal Guidance for Robust Diffusion-Based Handwritten Text Generation

Cascianelli, Silvia;Cucchiara, Rita;
2025

Abstract

Diffusion-based Handwritten Text Generation (HTG) approaches achieve impressive results on frequent, invocabulary words observed at training time and on regular styles. However, they are prone to memorizing training samples and often struggle with style variability and generation clarity. In particular, standard diffusion models tend to produce artifacts or distortions that negatively affect the readability of the generated text, especially when the style is hard to produce. To tackle these issues, we propose a novel sampling guidance strategy, Dual Orthogonal Guidance (DOG), that leverages an orthogonal projection of a negatively perturbed prompt onto the original positive prompt. This approach helps steer the generation away from artifacts while maintaining the intended content, and encourages more diverse, yet plausible, outputs. Unlike standard Classifier-Free Guidance (CFG), which relies on unconditional predictions and produces noise at high guidance scales, DOG introduces a more stable, disentangled direction in the latent space. To control the strength of the guidance across the denoising process, we apply a triangular schedule: weak at the start and end of denoising, when the process is most sensitive, and strongest in the middle steps. Experimental results on the state-of-the-art DiffusionPen and One-DM demonstrate that DOG improves both content clarity and style variability, even for out-of-vocabulary words and challenging writing styles.
2025
2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
usa
2025
7020
7029
Nikolaidou, Konstantina; Retsinas, George; Sfikas, Giorgos; Cascianelli, Silvia; Cucchiara, Rita; Liwicki, Marcus
Dual Orthogonal Guidance for Robust Diffusion-Based Handwritten Text Generation / Nikolaidou, K., Retsinas, G., Sfikas, G., Cascianelli, S., Cucchiara, R., Liwicki, M.. - (2025), pp. 7020-7029. (2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025 usa 2025) [10.1109/iccvw69036.2025.00725].
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

Licenza Creative Commons
I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11380/1414449
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 1
  • ???jsp.display-item.citation.isi??? 0
  • OpenAlex 2
social impact