Abstract
The project Empowering Multilingual Inclusive Communication (E-MIMIC) is a joint effort of the Deep Learning Natural Language Understanding (from Politecnico di Torino) and Linguistics research communities (from University of Bologna and University of Rome 2) with the goal of promoting and ensuring equality and inclusion in communication, thus contributing to a more inclusive, innovative, and reflective society. E-MIMIC relies on an innovative use of deep-learning methods for natural language processing trained on a new corpus of formal communication, produced by Italian and French linguists in this project. The approach is groundbreaking because for the first time, high-risk concepts such as linguistic and discursive criteria for inclusive communication, data labelling of new corpora of formal communication, and strong human involvement in the data-driven methods are integrated into the core of deep-learning methods for natural language processing to automatically identify non-inclusive text snippets in formal institutional texts, suggest alternative forms, and produce inclusive text reformulations.
Results achieved
: Guidelines Definition & Corpus Creation In close collaboration with linguistic experts, we established a set of linguistic criteria for Italian inclusive language. These guidelines served as the foundation for building a real-world corpus of Italian administrative texts, annotated specifically for two key tasks: (1) Detection of non-inclusive language and (2) Reformulation into inclusive alternatives. Human-Centered Validation Design We implemented two human evaluation workflows to assess model quality: (1) Detection Evaluation: Linguists analyze whether the models correctly identify the words influencing inclusiveness using explainability methods. (2) Reformulation Evaluation: Experts verify if the proposed rewordings address inclusivity issues, retain grammatical accuracy, and preserve the meaning of the original sentence. Model Development & Performance Evaluation We fine-tuned: (1) Two transformer-based classifiers to detect non-inclusive sentences; (2) Two sequence-to-sequence models to generate inclusive reformulations. We then conducted a comprehensive evaluation combining standard metrics and expert-driven assessmentsDettagli del progetto
Responsabile scientifico: Rachele Raus
Strutture Unibo coinvolte:
Dipartimento di Interpretazione e Traduzione
Coordinatore:
Politecnico di TORINO(Italy)
Contributo totale Unibo: Euro (EUR) 49.647,00
Durata del progetto in mesi: 24
Data di inizio
28/09/2023
Data di fine:
28/02/2026