Statistical Mechanics of Learning Machines: from algorithmic and information-theoretical limits to new biologically inspired paradigms

PRIN 2022 Tantari

Abstract

The increasing availability of massive data sets, with ever larger volume, variety and velocity has signed the direction of the recent ultimate technological progress, where machine learning (ML) has emerged as the key paradigm of modern Artificial Intelligence systems. On the one hand, despite the many impressive achievements, there are serious gaps in our theoretical understanding of learning systems. Deep neural networks (DNN) are often used as big black boxes: common deep learning (DL) practices (architectural choices, parameter fine-tuning) can mostly be viewed as alchemy, since nobody understands what makes it work. On the other hand, precisely because of those impressive achievements, ranging from image and speech recognition, language and text translation to computer vision and prediction of protein structures, the confidence of the scientific community in their power is unrestrainedly increasing and is pushing the expectations even further: from systems succeeding in specific tasks to more intelligent generalist models that can learn a sort of common sense of the world and do multiple tasks. However this ambition is crashing against the ancient bottleneck of data scarcity: it’s impossible to label everything in the world. A deeper understanding of the actual learning systems is thus becoming extremely necessary not only to calm a noble thirst of mathematical knowledge but also to avoid a new middle age of machine learning in favor of a new generation of systems able to acquire new skills without such a massive amount of labeled data, with a stronger abstraction capacity, thus closer to biological and human-level intelligence. To this purpose it is important contributing as mathematicians to: assess the information-theoretical limit of learning: given certain models for the learning machine and the environment, how much information (in terms of amount of data) is necessary for the machine to produce a faithful representation and correct predictions about the environment, regardless of any computational barrier? assess the actual limit of algorithms: the eventual information codified in the dataset has to be efficiently retrieved. It is important to understand whether existing algorithms are able to extract this information and how they use it to build a representation of the environment with the machine parameters. propose new biologically and human intelligence inspired learning approaches: intended as new learning machines, new dynamics of learning and algorithms, and new paradigms for data driven artificial intelligence, which are less task-specific, more efficient and parsimonious with data. Following this route the project intends to develop a mathematically grounded understanding of ML merging ideas from statistical inference, statistical mechanics and information theory. At the same time it aims at proposing new, biologically and human intelligence inspired learning paradigms, models and algorithms. Achieved Results A major contribution concerned Objective 1, devoted to the information-theoretical limits of machine learning. In particular, the Bologna unit led Task 1.2, which focused on self-supervised learning and generative neural network models. Through the study of Restricted Boltzmann Machines (RBMs) in the teacher–student framework, the research team characterized phase transitions governing memorization and generalization in high-dimensional learning systems. These investigations provided a rigorous description of the conditions under which learning becomes possible and clarified the fundamental information barriers associated with limited data availability and model mismatch. The work also explored the role of prior information in learning performance and introduced novel approaches for analyzing structured datasets through permutation replica-symmetry-breaking transitions. These results substantially advanced the theoretical understanding of self-supervised learning and representation learning mechanisms. The Bologna group also made fundamental contributions to Objective 3, dedicated to the development of new machine-learning paradigms inspired by statistical physics and biological intelligence. Within Task 3.1, the unit investigated Dense Associative Memory models in the teacher–student setting, providing a detailed characterization of the hierarchical saddle structure underlying their learning dynamics. This research offered new insights into the mechanisms through which associative memories learn, store, and retrieve information, while establishing theoretical connections between dense neural architectures and modern learning systems. In parallel, the group developed and studied multiscale mean-field spin-glass models, extending the mathematical tools available for analyzing complex learning systems characterized by interactions across multiple scales. These advances contributed to the broader goal of designing more efficient, interpretable, and biologically inspired learning architectures. Another important achievement concerned the development of innovative learning algorithms aimed at reducing computational costs. Within Task 3.2, the Bologna unit introduced a novel training strategy based on network-growing and splitting steepest-descent procedures. This approach was designed to improve training efficiency while maintaining learning performance, directly addressing one of the main goals of the project: the development of learning systems that require fewer computational resources and less training data. Such contributions are particularly relevant in the context of sustainable and scalable artificial intelligence. The scientific output produced by the Bologna unit demonstrates the originality and impact of these contributions. During the project, researchers published several articles in leading international journals, including Machine Learning: Science and Technology, Neural Networks, SciPost Physics, Physica A, and Annales Henri Poincaré. These publications addressed topics such as dense Hopfield networks in the teacher–student setting, the influence of priors on Restricted Boltzmann Machines, learning from structured datasets, multiscale spin-glass models, and the mathematical foundations of dense associative memories. Collectively, these works provided new theoretical tools for understanding learning phenomena in complex neural systems and strengthened the connection between statistical physics and modern machine-learning theory The Bologna unit also contributed actively to the dissemination of project results and to the consolidation of the Italian research community working on the mathematics of artificial intelligence. It was involved in numerous workshops, schools, and international conferences dedicated to machine learning, statistical physics, and computational neuroscience. These activities fostered interdisciplinary collaborations and enhanced the international visibility of the project.

Project details

Unibo Team Leader: Daniele Tantari

Unibo involved Department/s:
Dipartimento di Matematica

Coordinator:
ALMA MATER STUDIORUM - Università di Bologna(Italy)

Total Eu Contribution: Euro (EUR) 187.050,00
Total Unibo Contribution: Euro (EUR) 64.017,00
Project Duration in months: 24
Start Date: 28/09/2023
End Date: 28/02/2026

Funding bodies' logos