40720 - Data Mining

Academic Year 2026/2027

Course contents

The course covers the following topics:

  • Introduction to Data-Centric AI
    • From model-centric to data-centric development
    • Data quality dimensions
  • Data Understanding and Profiling
    • Exploratory data analysis and data visualization
    • Data profiling techniques
    • Statistical characterization of datasets
    • Detection of quality issues
  • Data Cleaning and Validation
    • Missing data handling
    • Duplicate detection
    • Outlier detection
    • Noise reduction
    • Data consistency and integrity validation
  • Feature Engineering
    • Feature construction and transformation
    • Feature encoding
    • Feature scaling and normalization
    • Feature selection
    • Dimensionality reduction techniques
  • Dataset Optimization
    • Class imbalance handling
    • Sampling and resampling techniques
    • Data augmentation
  • Data Quality and Fairness Assessment
    • Dataset evaluation metrics
    • Bias and Fairness evaluation
  • Graphs
    • Graph data and ML models
    • Graph-RAG
  • Privacy-Aware Techniques
    • Privacy risks
    • Privacy-preserving methods
    • Privacy–utility trade-offs
  • Explainability
    • Local and global explanations
    • Feature attribution
    • Explanation quality
  • AutoML and MLOps Applications
    • Automated model development
    • Deployment and monitoring, and error analysis
    • Data validation and drift detection
    • Continuous dataset improvement
  • Practical Applications
    • Case studies from real-world AI applications
    • End-to-end Data-Centric AI workflows
    • Laboratory sessions using Python and modern machine learning libraries for dataset preparation, validation, and optimization.

Readings/Bibliography

Slides and exercises provided by the lecturers

Teaching methods

Lessons, practical exercises and laboratories with Python

Assessment methods

The assessment consists of an oral examination on all the course topics and the discussion of a project. The project may involve either the implementation of an advanced machine learning algorithm from the scientific literature or the analysis of a dataset using the techniques covered during the course.

The objective of the assessment is to evaluate whether students have understood the techniques studied and have developed practical skills in working with data, understanding its content, and discovering hidden patterns and information.

Grades are assigned based on an overall evaluation of the student's knowledge, competencies, analytical skills, and ability to present and discuss the topics addressed. The grading ranges can be described as follows:

  • 18–23: Satisfactory preparation and analytical skills, limited to a restricted set of topics covered in the course; generally correct use of terminology.
  • 24–27: Technically adequate preparation, with some limitations regarding the topics covered; good analytical skills, although not particularly in-depth, expressed using appropriate terminology.
  • 28–30: Excellent knowledge of a broad range of topics covered in the course; strong analytical and critical thinking skills; mastery of the relevant terminology.
  • 30 with Honors (30L): Outstanding, comprehensive, and in-depth knowledge of the course topics; excellent critical analysis and ability to make connections across topics; complete mastery of the relevant terminology.

Teaching tools

Labs will be carried out in Python

Office hours

See the website of Matteo Golfarelli

See the website of Gianluca Moro