Resolving the body-order paradox of machine learning interatomic potentials

Chong, S.; Jiang, T.; Domina, M.; Bigi, F.; Grasselli, F.; Lee, J.; Ceriotti, M.

doi:10.1063/5.0303302

In many cases, the predictions of machine learning interatomic potentials (MLIPs) can be interpreted as a sum of body-ordered contributions, which is explicit when the model is directly built on neighbor density correlation descriptors and is implicit when the model captures the correlations through the non-linear functions of low body-order terms. In both cases, the "effective body-orderedness" of MLIPs remains largely unexplained: how do the models decompose the total energy into body-ordered contributions, and how does their body-orderedness affect the accuracy and learning behavior? In answering these questions, we first discuss the complexities in imposing the many-body expansion on ab initio calculations at the atomic limit. Next, we train a curated set of MLIPs on datasets of hydrogen clusters and reveal the inherent tendency of the ML models to deduce their own, effective body-order trends, which are dependent on the model type and dataset makeup. Finally, we present different trends in the convergence of the body-orders and generalizability of the models, providing useful insights into the development of future MLIPs.

Resolving the body-order paradox of machine learning interatomic potentials / Chong, S., Jiang, T., Domina, M., Bigi, F., Grasselli, F., Lee, J., Ceriotti, M.. - In: THE JOURNAL OF CHEMICAL PHYSICS. - ISSN 0021-9606. - 164:6(2026), pp. 1-12. [10.1063/5.0303302]

Resolving the body-order paradox of machine learning interatomic potentials

Chong S.;Jiang T.;Domina M.;Bigi F.;Grasselli F.;Lee J.;Ceriotti M.

2026

Abstract

In many cases, the predictions of machine learning interatomic potentials (MLIPs) can be interpreted as a sum of body-ordered contributions, which is explicit when the model is directly built on neighbor density correlation descriptors and is implicit when the model captures the correlations through the non-linear functions of low body-order terms. In both cases, the "effective body-orderedness" of MLIPs remains largely unexplained: how do the models decompose the total energy into body-ordered contributions, and how does their body-orderedness affect the accuracy and learning behavior? In answering these questions, we first discuss the complexities in imposing the many-body expansion on ab initio calculations at the atomic limit. Next, we train a curated set of MLIPs on datasets of hydrogen clusters and reveal the inherent tendency of the ML models to deduce their own, effective body-order trends, which are dependent on the model type and dataset makeup. Finally, we present different trends in the convergence of the body-orders and generalizability of the models, providing useful insights into the development of future MLIPs.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione
	
				2026
			
	Rivista
	
				THE JOURNAL OF CHEMICAL PHYSICS
			
	N° del Volume
	
				164
			
	Fascicolo
	
				6
			
	Pagina iniziale
	
				1
			
	Pagina finale
	
				12
			
	Codice DOI
	
				https://dx.doi.org/10.1063/5.0303302
			
	Codice WoS
	
				WOS:001693250900001
			
	Codice Scopus
	
				2-s2.0-105030061290
			
	Codice PubMed
	
				41685856
			
	Citazione
	
				Resolving the body-order paradox of machine learning interatomic potentials / Chong, S., Jiang, T., Domina, M., Bigi, F., Grasselli, F., Lee, J., Ceriotti, M.. - In: THE JOURNAL OF CHEMICAL PHYSICS. - ISSN 0021-9606. - 164:6(2026), pp. 1-12. [10.1063/5.0303302]
			
	Tutti gli autori
	
						Chong, S.; Jiang, T.; Domina, M.; Bigi, F.; Grasselli, F.; Lee, J.; Ceriotti, M.
					
	Tipologia
	
				Articolo su rivista

File in questo prodotto:

File	Dimensione	Formato
064121_1_5.0303302.pdf Open access Tipologia: VOR - Versione pubblicata dall'editore Licenza: [IR] creative-commons Dimensione 5.09 MB Formato Adobe PDF Visualizza/Apri	5.09 MB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris