Abstract
The ReMind project addressed one of the major challenges in dementia research by developing innovative computational methods and longitudinal linguistic resources to support the detection and monitoring of the earliest stages of neurodegenerative diseases. Building upon previous research on language-based biomarkers, ReMind developed an integrated technological framework for the collection, processing, and analysis of spoken language produced during controlled and semi-spontaneous language tasks. The project combined Natural Language Processing, machine learning, and clinical assessment to investigate how subtle linguistic changes evolve over time and how they relate to the progression of cognitive impairment.
Results achieved
One of the project's major achievements was the development of a fully automated computational pipeline for the extraction of Digital Linguistic Biomarkers from speech. The pipeline integrated speech acquisition, automatic transcription, temporal alignment, linguistic annotation and feature extraction, providing a comprehensive framework for analysing acoustic, rhythmic, lexical, syntactic, semantic, and discourse-level features. This infrastructure represents a reusable resource for future research on language-based assessment of neurodegenerative disorders. The project also produced a unique longitudinal Italian speech corpus specifically designed for studying cognitive decline. The corpus includes 92 participants, each assessed on three occasions over a twelve-month period, with evaluations performed every six months. The cohort comprises 9 individuals with Subjective Cognitive Decline (SCD), 30 with Mild Cognitive Impairment (MCI), 15 participants with early-stage dementia (CDR = 1), and 38 cognitively healthy controls. The longitudinal design allowed the investigation of both cross-sectional differences between clinical groups and temporal changes in language associated with disease progression, providing one of the few Italian longitudinal resources currently available for this type of research. To ensure methodological consistency and ecological validity, the project developed standardized language elicitation protocols implemented through the ReadLet platform. These protocols supported the administration of controlled reading tasks together with semi-spontaneous language production, enabling harmonized data collection across repeated recording sessions. The resulting methodology provides a reproducible framework that can be adopted in future clinical and research settings. The longitudinal corpus, together with the automated analysis pipeline, substantially expanded the resources available for investigating linguistic manifestations of cognitive decline. By integrating repeated language assessments with standardized neuropsychological evaluation, the project generated a valuable infrastructure for studying the evolution of Digital Linguistic Biomarkers across the prodromal and early clinical stages of dementia. These resources also laid the methodological foundations for the development of explainable machine-learning models capable of supporting clinicians in the identification of subtle cognitive changes that may not yet be detectable through conventional screening instruments. The project promoted Open Science and reproducible research throughout its implementation. While raw speech recordings could not be made publicly available because they contain identifiable biometric information, methodological protocols and derived research resources were designed according to FAIR principles to maximize transparency, interoperability, and future reuse whenever ethically and legally possible. Beyond the technological developments, ReMind strengthened interdisciplinary collaboration between computational linguistics, artificial intelligence, clinical neuroscience, and neuropsychology. The project established a methodological framework that will continue to support future investigations on healthy ageing, Mild Cognitive Impairment, dementia, and longitudinal language change, while contributing to the development of scalable, explainable, and clinically meaningful AI-based tools for early cognitive screening. Updated information on project activities, publications, datasets, and research outputs is available on the project website: https://site.unibo.it/remind-project/enDettagli del progetto
Responsabile scientifico: Gloria Gagliardi
Strutture Unibo coinvolte:
Dipartimento di Filologia Classica e Italianistica
Coordinatore:
CNR - Consiglio Nazionale delle Ricerche(Italy)
Contributo totale Unibo: Euro (EUR) 98.214,00
Durata del progetto in mesi: 24
Data di inizio
05/10/2023
Data di fine:
28/02/2026