Semantic Residual Prompts for Continual Learning

Menabue, M.; Frascaroli, E.; Boschini, M.; Sangineto, E.; Bonicelli, L.; Porrello, A.; Calderara, S.

doi:10.1007/978-3-031-73030-6_1

Prompt-tuning methods for Continual Learning (CL) freeze a large pre-trained model and train a few parameter vectors termed prompts. Most of these methods organize these vectors in a pool of key-value pairs and use the input image as query to retrieve the prompts (values). However, as keys are learned while tasks progress, the prompting selection strategy is itself subject to catastrophic forgetting, an issue often overlooked by existing approaches. For instance, prompts introduced to accommodate new tasks might end up interfering with previously learned prompts. To make the selection strategy more stable, we leverage a foundation model (CLIP) to select our prompts within a two-level adaptation mechanism. Specifically, the first level leverages a standard textual prompt pool for the CLIP textual encoder, leading to stable class prototypes. The second level, instead, uses these prototypes along with the query image as keys to index a second pool. The retrieved prompts serve to adapt a pre-trained ViT, granting plasticity. In doing so, we also propose a novel residual mechanism to transfer CLIP semantics to the ViT layers. Through extensive analysis on established CL benchmarks, we show that our method significantly outperforms both state-of-the-art CL approaches and the zero-shot CLIP test. Notably, our findings hold true even for datasets with a substantial domain gap w.r.t. the pre-training knowledge of the backbone model, as showcased by experiments on satellite imagery and medical datasets. The codebase is available at https://github.com/aimagelab/mammoth.

Semantic Residual Prompts for Continual Learning / Menabue, M., Frascaroli, E., Boschini, M., Sangineto, E., Bonicelli, L., Porrello, A., Calderara, S.. - 15119 LNCS:(2025), pp. 1-18. (18th European Conference on Computer Vision, ECCV 2024 Milano, Italy 29 Sep - 4 Oct, 2024) [10.1007/978-3-031-73030-6_1].

Semantic Residual Prompts for Continual Learning

Menabue M.;Frascaroli E.;Boschini M.;Sangineto E.;Bonicelli L.;Porrello A.;Calderara S.

2025

Abstract

Prompt-tuning methods for Continual Learning (CL) freeze a large pre-trained model and train a few parameter vectors termed prompts. Most of these methods organize these vectors in a pool of key-value pairs and use the input image as query to retrieve the prompts (values). However, as keys are learned while tasks progress, the prompting selection strategy is itself subject to catastrophic forgetting, an issue often overlooked by existing approaches. For instance, prompts introduced to accommodate new tasks might end up interfering with previously learned prompts. To make the selection strategy more stable, we leverage a foundation model (CLIP) to select our prompts within a two-level adaptation mechanism. Specifically, the first level leverages a standard textual prompt pool for the CLIP textual encoder, leading to stable class prototypes. The second level, instead, uses these prototypes along with the query image as keys to index a second pool. The retrieved prompts serve to adapt a pre-trained ViT, granting plasticity. In doing so, we also propose a novel residual mechanism to transfer CLIP semantics to the ViT layers. Through extensive analysis on established CL benchmarks, we show that our method significantly outperforms both state-of-the-art CL approaches and the zero-shot CLIP test. Notably, our findings hold true even for datasets with a substantial domain gap w.r.t. the pre-training knowledge of the backbone model, as showcased by experiments on satellite imagery and medical datasets. The codebase is available at https://github.com/aimagelab/mammoth.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione
	
				2025
			
	Data di prima pubblicazione
	
				24-nov-2024
			
	Titolo del Convegno
	
				18th European Conference on Computer Vision, ECCV 2024
			
	Luogo del Convegno
	
				Milano, Italy
			
	Data del Convegno
	
				29 Sep - 4 Oct, 2024
			
	Codice DOI
	
				https://dx.doi.org/10.1007/978-3-031-73030-6_1
			
	Codice WoS
	
				WOS:001403057100001
			
	Codice Scopus
	
				2-s2.0-85210865959
			
	Serie
	
				LECTURE NOTES IN COMPUTER SCIENCE
			
	N° del Volume
	
				15119 LNCS
			
	Pagina iniziale
	
				1
			
	Pagina finale
	
				18
			
	Tutti gli autori
	
						Menabue, M.; Frascaroli, E.; Boschini, M.; Sangineto, E.; Bonicelli, L.; Porrello, A.; Calderara, S.
					
	Citazione
	
				Semantic Residual Prompts for Continual Learning / Menabue, M., Frascaroli, E., Boschini, M., Sangineto, E., Bonicelli, L., Porrello, A., Calderara, S.. - 15119 LNCS:(2025), pp. 1-18. (18th European Conference on Computer Vision, ECCV 2024 Milano, Italy 29 Sep - 4 Oct, 2024) [10.1007/978-3-031-73030-6_1].
			
	Tipologia
	
				Relazione in Atti di Convegno

File in questo prodotto:

File	Dimensione	Formato
STAR-Prompt_final_version.pdf Accesso riservato Tipologia: AAM - Versione dell'autore revisionata e accettata per la pubblicazione Licenza: [IR] closed Dimensione 1.2 MB Formato Adobe PDF Visualizza/Apri Richiedi una copia	1.2 MB	Adobe PDF	Visualizza/Apri Richiedi una copia

Pubblicazioni consigliate

I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris