- Docente: Matteo Golfarelli
- Credits: 6
- SSD: IINF-05/A
- Language: English
- Moduli: Matteo Golfarelli (Modulo 1) Gianluca Moro (Modulo 2)
- Teaching Mode: In-person learning (entirely or partially) In-person learning (entirely or partially) (Modulo 1); In-person learning (entirely or partially) (Modulo 2)
- Campus: Cesena
-
Corso:
Second cycle degree programme (LM) in
Computer Science and Engineering (cod. 6699)
Also valid for Second cycle degree programme (LM) in Computer Science and Engineering (cod. 6699)
Second cycle degree programme (LM) in Computer Science and Engineering (cod. 6699)
Second cycle degree programme (LM) in Computer Science and Engineering (cod. 6699)
Second cycle degree programme (LM) in Digital Transformation Management (cod. 6823)
-
from Oct 12, 2026 to Nov 09, 2026
-
from Sep 23, 2026 to Dec 16, 2026
Course contents
The AA 26-27 edition consists of two courses running in parallel:
- Data Centric AI (for students enrolled in AA 26-27) composed by:
- Module I - (prof. Matteo Golfarelli)
- Module II - (prof. Matteo Francia)
- Data Mining (for students enrolled before AA 26-27).
- Module I - (prof. Matteo Golfarelli)
- Module III - (prof. Gianluca Moro)
Module I - Data Centric AI/Data Mining (Golfarelli)
- Introduction to Data-Centric AI
- From model-centric to data-centric development
- Data quality dimensions
- Data Understanding and Profiling
- Exploratory data analysis and data visualization
- Data profiling techniques
- Statistical characterization of datasets
- Detection of quality issues
- Data Cleaning and Validation
- Missing data handling
- Duplicate detection
- Outlier detection
- Noise reduction
- Data consistency and integrity validation
- Feature Engineering
- Feature construction and transformation
- Feature encoding
- Feature scaling and normalization
- Feature selection
- Dimensionality reduction techniques
- Dataset Optimization
- Class imbalance handling
- Sampling and resampling techniques
- Data augmentation
- Data Quality and Fairness Assessment
- Dataset evaluation metrics
- Bias and Fairness evaluation
Module II - Data Centric AI (Francia)
- Graphs
- Graph data and ML models
- Graph-RAG
- Privacy-Aware Techniques
- Privacy risks
- Privacy-preserving methods
- Privacy–utility trade-offs
- Explainability
- Local and global explanations
- Feature attribution
- Explanation quality
- AutoML and MLOps Applications
- Automated model development
- Deployment and monitoring, and error analysis
- Data validation and drift detection
- Continuous dataset improvement
- Practical Applications
- Case studies from real-world AI applications
- End-to-end Data-Centric AI workflows
- Laboratory sessions using Python and modern machine learning libraries for dataset preparation, validation, and optimization.
Module III - Text Mining and NLP (Gianluca Moro)
Acquisition of the theoretical and practical foundations of text mining and natural language processing (NLP) that underpin modern LLM-based AI (see the slides on a state-of-the-art summary [https://disi-unibo-nlp.github.io/development/2024/04/17/llm-slides/] ), as well as the development of solutions for their application to real-world professional domains, including enterprise use cases such as office automation, healthcare, legal services and finance.
MODULE CONTENTS
1. Text Representation, Retrieval, Classification, Sentiment Analysis and Opinion Mining, Word embeddings
2. Language Models, Transformers, Efficient Attention Mechanisms, Retrieval-augmented Generation with VectorDB and LLM Wiki
3. Vision-Language Models, Graph Machine Learning (GNN, Graph Embedding and GraphRAG), Neuro-symbolic approaches for Trustworthy AI
4. Generative Large Language Model, Prompt Engineering, Evaluation Methods, Compression and Quantization, Parameter-Efficient Fine-Tuning (e.g. QLoRA, Prompt Tuning) e LLM Agents
5. Labs on real case studies, such as Text Summarization, Information Extraction, Chatbot, LLM agents and Function Call Generation etc. in office automation, with Python and open source tools: Hugging Face framework, LangChain, LlamaIndex, Milvus, Unsloth etc.
Readings/Bibliography
Slides and exercises provided by the lecturers
Teaching methods
Lessons, practical exercises and laboratories with Python
Assessment methods
The assessment consists of an oral examination on all the course topics and the discussion of a project. The project may involve either the implementation of an advanced machine learning algorithm from the scientific literature or the analysis of a dataset using the techniques covered during the course.
The objective of the assessment is to evaluate whether students have understood the techniques studied and have developed practical skills in working with data, understanding its content, and discovering hidden patterns and information.
Grades are assigned based on an overall evaluation of the student's knowledge, competencies, analytical skills, and ability to present and discuss the topics addressed. The grading ranges can be described as follows:
- 18–23: Satisfactory preparation and analytical skills, limited to a restricted set of topics covered in the course; generally correct use of terminology.
- 24–27: Technically adequate preparation, with some limitations regarding the topics covered; good analytical skills, although not particularly in-depth, expressed using appropriate terminology.
- 28–30: Excellent knowledge of a broad range of topics covered in the course; strong analytical and critical thinking skills; mastery of the relevant terminology.
- 30 with Honors (30L): Outstanding, comprehensive, and in-depth knowledge of the course topics; excellent critical analysis and ability to make connections across topics; complete mastery of the relevant terminology.
Teaching tools
Labs will be carried out in Python
Office hours
See the website of Matteo Golfarelli
See the website of Gianluca Moro