Novel Perspectives for the Management of Multilingual and Multialphabetic Heritages through Automatic Knowledge Extraction: The DigitalMaktaba Approach

Bergamaschi, Sonia; Stefania De Nardis,; Martoglia, Riccardo; Ruozzi, Federico; Sala, Luca; Vanzini, Matteo; Vigliermo, Riccardo Amerigo

doi:10.3390/s22113995

The linguistic and social impact of multiculturalism can no longer be neglected in any sector, creating the urgent need of creating systems and procedures for managing and sharing cultural heritages in both supranational and multi-literate contexts. In order to achieve this goal, text sensing appears to be one of the most crucial research areas. The long-term objective of the DigitalMaktaba project, born from interdisciplinary collaboration between computer scientists, historians, librarians, engineers and linguists, is to establish procedures for the creation, management and cataloguing of archival heritage in non-Latin alphabets. In this paper, we discuss the currently ongoing design of an innovative workflow and tool in the area of text sensing, for the automatic extraction of knowledge and cataloguing of documents written in non-Latin languages (Arabic, Persian and Azerbaijani). The current prototype leverages different OCR, text processing and information extraction techniques in order to provide both a highly accurate extracted text and rich metadata content (including automatically identified cataloguing metadata), overcoming typical limitations of current state of the art approaches. The initial tests provide promising results. The paper includes a discussion of future steps (e.g., AI-based techniques further leveraging the extracted data/metadata and making the system learn from user feedback) and of the many foreseen advantages of this research, both from a technical and a broader cultural-preservation and sharing point of view.

Novel Perspectives for the Management of Multilingual and Multialphabetic Heritages through Automatic Knowledge Extraction: The DigitalMaktaba Approach / Bergamaschi, S., De Nardis, S., Martoglia, R., Ruozzi, F., Sala, L., Vanzini, M., Vigliermo, R.A.. - In: SENSORS. - ISSN 1424-8220. - 22:11(2022), pp. 1-20. [10.3390/s22113995]

Novel Perspectives for the Management of Multilingual and Multialphabetic Heritages through Automatic Knowledge Extraction: The DigitalMaktaba Approach

Sonia Bergamaschi;Stefania De Nardis;Riccardo Martoglia;Federico Ruozzi;Luca Sala;Matteo Vanzini;Riccardo Amerigo Vigliermo

2022

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno di pubblicazione
	
				2022
			
	Rivista
	
				SENSORS
			
	N° del Volume
	
				22
			
	Fascicolo
	
				11
			
	Pagina iniziale
	
				1
			
	Pagina finale
	
				20
			
	Codice DOI
	
				https://dx.doi.org/10.3390/s22113995
			
	Codice WoS
	
				WOS:000808717700001
			
	Codice Scopus
	
				2-s2.0-85130987349
			
	Codice PubMed
	
				35684615
			
	Citazione
	
				Novel Perspectives for the Management of Multilingual and Multialphabetic Heritages through Automatic Knowledge Extraction: The DigitalMaktaba Approach / Bergamaschi, S., De Nardis, S., Martoglia, R., Ruozzi, F., Sala, L., Vanzini, M., Vigliermo, R.A.. - In: SENSORS. - ISSN 1424-8220. - 22:11(2022), pp. 1-20. [10.3390/s22113995]
			
	Tutti gli autori
	
						Bergamaschi, Sonia; De Nardis, Stefania; Martoglia, Riccardo; Ruozzi, Federico; Sala, Luca; Vanzini, Matteo; Vigliermo, Riccardo Amerigo
					
	Tipologia
	
				Articolo su rivista

File in questo prodotto:

File	Dimensione	Formato
sensors-22-03995.pdf Open access Descrizione: Articolo Tipologia: VOR - Versione pubblicata dall'editore Licenza: [IR] creative-commons Dimensione 2.85 MB Formato Adobe PDF Visualizza/Apri	2.85 MB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I metadati presenti in IRIS UNIMORE sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono rilasciati con licenza Attribuzione 4.0 Internazionale (CC BY 4.0), salvo diversa indicazione.
In caso di violazione di copyright, contattare Supporto Iris