29443 - Computer Vision

Academic Year 2026/2027

  • Moduli: Annalisa Franco (Modulo 1) Matteo Ferrara (Modulo 2)
  • Teaching Mode: In-person learning (entirely or partially) In-person learning (entirely or partially) (Modulo 1); In-person learning (entirely or partially) (Modulo 2)
  • Campus: Cesena
  • Corso: Second cycle degree programme (LM) in Computer Science and Engineering (cod. 6699)

Learning outcomes

The course aims to provide students with the knowledge and tools necessary for the design and implementation of automated systems capable of analyzing digital images for the purpose of object localization and recognition. At the end of the course, also thanks to the laboratory activities, students will be able to: - design a vision system capable of analyzing the content of a digital image; - develop vision algorithms based on classical machine learning techniques, hand-crafted features (color, shape, texture), key points, and local descriptors; - train and optimize neural networks for classification, localization, and semantic segmentation tasks.

Course contents

  • Basic techniques for digital image processing and filtering 
  • Feature extraction
    • Color features:

      - color histograms and similarity metrics;

      - color moments

    • Texture features:

    - gray-level co-occurrence matrix and related measures (enthropy, contrast, homogeneity, etc..);

    - Gabor filters: filter banks;

    - Haar features: integral image and efficient feature extraction;

    - Local Binary Pattern;

     

    • Shape features:

      - Object countour extraction and one-dimensional shape representations; 

      - Shape descriptors, Fourier descriptors;

      - Invariant moments.

    • Handcrafted features vs Representation learning
  • Image stitching, 2D image registration and Visual SLAM
    • Keypoints and local descriptors:
    • - Keypoint detection: Harris corner detector;

      - Scale invariant detectors: Harris Laplace, Laplacian of Gaussian, Difference of Gaussian;

      - Keypoints and descriptors: SIFT, SURF, BRIEF, Histogram of Oriented Gradients.

      - Ransac algorithm for feature matching.

    - Application to robotics: Visual SLAM (Simultaneous Localization and Mapping)
  • Semantic segmentation in digital images
    • Color-based segmentation techniques, Mean Shift algorithm;
    • Deep learning based techniques with applications in satellite and medical images analysis.
  • Recognition “in the wild”
    • Object detection/classification
    • - Color, texture and shape features for content-based image retrieval;

      - Bag of visual Words;

      - Feature-based Rigid Template matching and applications to object detection and recognition (e.g. grocery products)

      - Hough Transform;

      - Deep learning techniques for object detection and recognition (e.g. pedestrian and road sign recognition, object recognition for robotic vision, face detection and recognition).

  • Video surveillance and video analysis
    • Basic techniques for frame subtraction and background modeling
    • Approaches for object/person tracking and crowd analysis;
    • Video analysis.

Readings/Bibliography

  • Torralba, Isola, Freeman, "Foundations of Computer Vision", MIT press, 2024.
  • Zhang, Lipton, Li, Smola, Dive into Deep Learning, https://d2l.ai [https://d2l.ai/], 2020.
  • Elgendy, "Deep Learning for Vision Systems", Manning, 2020.
  • Forsyth and Ponce, Computer Vision a modern approach, Pearson 2012.
  • Kaehler and Bradski, Learning OpenCV 3, O'Reilly 2017.
  • Shi, "Emgu CV Essentials", Packt publishing 2013.
  • Gonzalez and Woods, Elaborazioni delle immagini digitali, Prentice Hall, 3 edizione, 2008.

Teaching methods

Lectures

Laboratory exercises based on public multi-latform computer vision libraries (e.g. OpenCV) and deep learning frameworks (e.g. Tensorflow, PyTorch).

As concerns the teaching methods of this course unit, all students must attend Module 1, 2 on Health and Safety online

Assessment methods

Learning is assessed through the completion of an individual or group project and an oral examination covering the entire course syllabus.

The project is designed to assess the student's ability to design and implement a Computer Vision solution using the techniques presented during the course, critically analyze the obtained results, and justify the adopted design choices.

The oral examination is aimed at assessing:

  • knowledge and understanding of the course topics;
  • the ability to critically compare different methodologies, highlighting their strengths, limitations, and application scenarios;
  • the appropriate use of technical and scientific terminology;
  • the ability to connect the different topics covered throughout the course and to discuss the design choices made in the project.

The final grade is based on both the quality of the project and the outcome of the oral examination. The project constitutes the basis of the evaluation, while the oral examination may adjust the project grade by approximately ±3 points, depending on the student's level of preparation, critical thinking skills, and clarity of presentation.

The overall evaluation takes into account:

  • the technical correctness of the proposed solution;
  • the ability to design and implement effective Computer Vision solutions;
  • autonomy in the analysis and interpretation of the results;
  • the ability to justify methodological and design choices;
  • the appropriate use of technical terminology;
  • the ability to integrate and relate the different topics covered during the course.

Teaching tools

  • Teacher's slides
  • Code traces for laboratory exercises
  • OpenCV library
  • Deep learning frameworks

 

Office hours

See the website of Annalisa Franco

See the website of Matteo Ferrara

SDGs

Quality education Industry, innovation and infrastructure

This teaching activity contributes to the achievement of the Sustainable Development Goals of the UN 2030 Agenda.