Abstract
In the past decade, advancements in deep learning, particularly in the field of natural language processing (NLP) and text mining, have significantly enhanced semantic analysis tasks such as text classification, word sense disambiguation, machine translation, text summarization, question answering, and sentiment analysis. This progress is largely attributed to the concept of word embedding, a word's meaning representation obtained through numeric coordinates, also known as vectors. Current word embeddings, derived from large textual corpora, have demonstrated efficacy but raise questions about their alignment with human language processing. The WEMB project aims to address this by pursuing two objectives: firstly, gaining a deeper understanding of how word embeddings align with human language processing, and secondly, leveraging this understanding to develop a new generation of embeddings for NLP tasks. The project employs a "from mind to application and back" approach, bridging the expertise of UniBO in language processing with ISTI-CNR's proficiency in NLP. WEMB focuses on three key aspects: 1. Embeddings and Cross-Modality: Investigating the relationship between embeddings that incorporate cross-modal information (e.g., from text and images) and traditional text-based embeddings in language processing. 2. Embeddings and Misspellings: Exploring the connection between embeddings and misspellings, a prevalent linguistic behavior in a growing number of texts for various reasons. 3. Embeddings and Word Senses: Examining the relationship between embeddings and word senses, particularly among different embeddings associated with different senses of the same ambiguous word. Through these investigations, WEMB aims to contribute to both a theoretical understanding of word embeddings in human language processing and the practical development of enhanced embeddings for NLP applications.
Results achieved
Overall, the WEMB project (https://wembprin.github.io/) fully achieved its objectives, namely to advance the understanding of the relationship between word embeddings and human language processing, and to exploit this knowledge to contribute to the development of a new generation of word embeddings for Natural Language Processing (NLP) and text mining. Workpackage 1 (WP1): Evaluating and Understanding Embeddings through Language-Vision Models and Large Language Models (LLMs) This workpackage employed state-of-the-art language and vision-language models to investigate the properties and limitations of semantic representations. Cross-modality studies showed that LLMs effectively capture basic-level categories but diverge from humans in representing fine-grained hierarchical relationships (How Humans and LLMs Organize Conceptual Knowledge, Pedrotti et al., 2025b), while the integration of visual information improves certain aspects of semantic representations but still presents limitations in encoding encyclopedic knowledge (Benchmarking Accuracy and Bias, Cassese et al., 2026). The workpackage also developed tools for the systematic evaluation of LLM capabilities. In the mathematical domain, standardized benchmarks for the Italian language were introduced (INVALSI - Mathematical and Language Understanding in Italian, Puccetti et al., 2024b; 2025a), and it was shown that LLMs often rely on superficial patterns rather than genuine mathematical reasoning (GSM-Identity, Negi et al., 2026a; A Multi-Perspective Evaluation of Mathematical Reasoning in Vision and Language Models, Negi et al., 2026b). On the linguistic side, alongside the INVALSI benchmarks, the project demonstrated that optimizing tokenization improves the efficiency and semantic coherence of Italian language representations (Optimizing LLMs for Italian, Moroni et al., 2025). Another line of research addressed bias in LLMs, showing that bias-mitigation prompts alter the internal computational processes of the models (Prompt-based bias control in Large Language Models, Cassese et al., 2025) and developing methodologies to measure biases related to popularity, geographic origin, and gender in encyclopedic knowledge (Benchmarking accuracy and bias on encyclopedic knowledge in (vision-)language models, Cassese et al., 2026). Finally, the project investigated the limitations of current AI-generated text detection systems, highlighting both the difficulty of identifying LLM-generated content (AI "news" content farms are easy to make, and hard to detect, Puccetti et al., 2024c) and the fact that models instructed to write in a more human-like style become significantly harder to detect (Stress-testing Machine Generated Text Detection, Pedrotti et al., 2025a; Is Human-Like Text Liked by Humans?, Wang et al., 2026). Workpackage 2 (WP2): Word Embeddings and Misspellings This workpackage investigated the robustness of word embeddings to misspellings and other forms of non-standard language input. A benchmark was developed to study phonetic-semantic relationships in embeddings (PSET: A Phonetics-Semantics Evaluation Testbed, Sperduti and Nguyen, 2025), and authentic misspelling patterns were analysed in relation to sociodemographic variables (Evaluating Misspelling Patterns in Gamified Data, Loia et al., 2026). In addition, a systematic survey on the treatment of non-canonical lexical forms was produced (Misspellings in Natural Language Processing, Sperduti and Moreo, 2026a), and it was shown that LLM robustness to internally scrambled words primarily derives from subword tokenization rather than from human-like reading mechanisms (Typoglycemia under the Hood, Sperduti and Moreo, 2026b). Workpackage 3 (WP3): Refining Word Embeddings: Abstraction, Concreteness, and Word Senses This workpackage investigated context-sensitive, hierarchical, and multidimensional representations of meaning. A benchmark was developed to evaluate contextualized representations of abstractness, concreteness, specificity, and generality (ABRICOT - ABstRactness and Inclusiveness in COntexT, Puccetti et al., 2024a). Subsequent studies showed that LLMs encode meaningful abstraction hierarchies and that these representations can be influenced by the simulated sociodemographic profile (Wordnet, and Word Ladders: Climbing the Abstraction Taxonomy with LLMs, Puccetti et al., 2025b). In parallel, the project established an empirical baseline for the development of linguistic-mediated abstraction during school age (Development of Linguistic-Mediated Abstraction, Villani et al., 2025) and further examined the similarities and differences between human conceptual organization and the semantic hierarchies encoded by LLMs (How Humans and LLMs Organize Conceptual Knowledge, Pedrotti et al., 2025b). All publications are available on the project website: https://wembprin.github.io/publications/.Project details
Unibo Team Leader: Marianna Marcella Bolognesi
Unibo involved Department/s:
Dipartimento di Lingue, Letterature e Culture Moderne
Coordinator:
CNR - Consiglio Nazionale delle Ricerche(Italy)
Total Unibo Contribution: Euro (EUR) 53.037,00
Project Duration in months: 24
Start Date:
28/09/2023
End Date:
28/02/2026