Expanding the toolbox for cereal breeding: high-throughput genomics, 2D-3D phenomics and artificial intelligence for breeding with increasing genome complexity, from barley to durum and bread wheat (POLYPLOIDBREEDING 4.0)

PRIN 2022 Frascaroli

Abstract

POLYPLOIDBREEDING 4.0 aimed to expand the tools available for cereal breeding by combining high-throughput genomics, field and root phenotyping, and artificial intelligence. Barley, durum wheat and bread wheat were used as models of increasing genome complexity, from diploid to tetraploid and hexaploid species. The project was organised into six Work Packages covering project and data management, genotyping and phenotyping, data processing and tool development, predictive modelling, validation, and dissemination. Activities were jointly carried out by CNR-IBBA and the University of Bologna, with contributions from CREA, SIS SpA and Forschungszentrum Jülich as external partners, bringing together complementary expertise in breeding, genomics, phenomics, image analysis and artificial intelligence.

Results achieved

The project was designed to generate new genomic and phenotypic data, integrate them with existing datasets, and develop methods and tools that could be used in modern cereal breeding. A central feature of the work was the comparison of species with increasing genome complexity: barley, durum wheat and bread wheat. This made it possible to test whether methods developed for diploid crops could be extended to polyploid species, where genomic analysis and prediction are generally more difficult. A major part of the project focused on high-throughput field phenotyping. Experimental fields were established for barley, durum wheat and bread wheat, and repeated UAV flights were carried out throughout the growing season. RGB, multispectral and thermal cameras were used to collect images from hundreds of accessions. These data made it possible to monitor crop development over time and to derive vegetation indices, plant height estimates and other indicators of crop performance. The UAV data were compared with conventional field measurements such as heading date, plant height, grain yield, lodging and disease response. This produced a wide and detailed phenotypic dataset for all three cereal species. The UAV work also led to the development of DRONE2REPORT, an open-source software tool created within the project. The tool can process aerial images, calculate vegetation indices, estimate plant height, separate crop signal from background noise, apply machine-learning procedures and generate automatic reports. It was developed mainly by CNR-IBBA and CREA-ZA, with data and validation contributions from the other partners. The software and the related analysis code were released through public repositories, so that they can be reused by the scientific and technical community. Root phenotyping was another important component of the project. A rhizotron experiment was carried out at Forschungszentrum Jülich on a panel of durum wheat accessions. Root and shoot images were collected daily during early plant growth, generating a time series of high-resolution data. These images were used to describe traits such as root width, root area, root orientation, root-to-shoot ratio and growth rate. Multivariate analyses identified clear differences among genotypes and made it possible to group them according to contrasting root architectures. To test whether the root traits observed under controlled conditions were also relevant in the field, two additional shovelomics experiments were organised in Fiorenzuola d’Arda and Cadriano. The same durum wheat accessions were grown in replicated field trials, and their root systems were excavated and photographed. The field analyses confirmed substantial genetic variation, particularly for root angle. They also showed a significant interaction between genotype and environment, suggesting that some root architectures may be better suited to specific environmental conditions. This additional activity, not included in the original proposal, strengthened the validation of the rhizotron results. The genomic component of the project was coordinated mainly by the University of Bologna. More than one thousand durum and bread wheat samples were analysed using a wheat SNP array. The choice of genotypes and markers was made to ensure compatibility with previous datasets and to support the integration of new and historical information. After quality control, the resulting data were used for genome-wide association studies and for the identification of chromosome regions associated with agronomic traits. In durum wheat, significant genetic variation was found for grain yield, heading date and other traits. Genome-wide association analysis identified marker-trait associations for grain yield and related characteristics. The new genotypic data were then harmonised with much larger collaborative datasets of tetraploid and bread wheat. This increased the number of available markers and improved the definition of chromosome regions of interest. The resulting haplotypes may support the identification of candidate genes and the development of diagnostic markers for marker-assisted selection. Whole-genome sequencing was more complex than originally expected, mainly because of the large size and polyploid structure of durum and bread wheat genomes. The project addressed this problem by joining international sequencing initiatives and by linking key project genotypes to wider collections. This provided access to long-read assemblies of tetraploid wheat and to large resequencing datasets of bread wheat. These resources now allow more detailed studies of nucleotide and structural variation and support comparative analyses across diploid, tetraploid and hexaploid cereals. Artificial intelligence and machine-learning methods were used across several parts of the project. Computer-vision models were developed to predict agronomic traits from UAV images, with particular attention to grain yield and heading date. Further analyses explored whether final crop performance could be predicted from early-season images or from reduced-resolution datasets. Similar approaches were applied to root images to describe growth over time and to identify informative combinations of root traits. These activities were intended not only to improve prediction, but also to reduce the time and cost of phenotyping. The project also created a framework for integrating different types of information. Traditional field measurements, UAV images, root phenotypes, SNP genotypes and sequence data were organised into larger and more informative datasets. This integration is important because modern breeding increasingly depends on the combined use of many types of data rather than on a single source of information. By working across three cereal species, the project also provided a basis for evaluating how genome complexity affects data analysis and predictive performance. All the main experimental activities were completed successfully. The only planned component that could not be implemented was the use of ultraviolet cameras for UAV phenotyping, because no suitable system compatible with the available drones could be identified. The sequencing activities also required more time than expected, but the problem was solved through international collaboration and a five-month project extension. The samples were successfully sequenced and the resulting data were obtained for further analysis. The work was distributed among the participating units according to their expertise. CNR-IBBA coordinated the overall project and contributed mainly to data analysis, machine learning, image processing and software development. The University of Bologna coordinated genotyping, genome-wide association studies and part of the field validation. CREA contributed field trials, historical datasets, genomic analysis and computer-vision expertise. SIS SpA provided plant material, field experiments and conventional phenotypic data, while Forschungszentrum Jülich hosted the high-throughput rhizotron experiment. This organisation allowed the project to combine resources and skills that were not available within a single institution. The project also followed open-science principles. Software and code were released as open-source resources, and the project website was used to communicate activities and results. Scientific outputs included conference presentations, a software-oriented scientific article, a contribution to a white book on durum wheat research and several additional manuscripts in preparation. These concern genome-wide association studies, root architecture, computer vision and genomic prediction. Overall, POLYPLOIDBREEDING 4.0 generated a large and integrated collection of plant material, images, phenotypic measurements, genotypic data and sequence resources for barley, durum wheat and bread wheat. It also produced analytical workflows and software tools that can support future breeding activities. The main contribution of the project is the demonstration that high-throughput phenotyping, genomics and artificial intelligence can be combined across cereal species with different levels of genome complexity. The resulting datasets, methods and collaborations provide a basis for further research and for the development of more efficient and better targeted breeding strategies.

Dettagli del progetto

Responsabile scientifico: Elisabetta Frascaroli

Strutture Unibo coinvolte:
Dipartimento di Scienze e Tecnologie Agro-Alimentari

Coordinatore:
CNR - Consiglio Nazionale delle Ricerche(Italy)

Contributo totale Unibo: Euro (EUR) 91.085,00
Durata del progetto in mesi: 24
Data di inizio 28/09/2023
Data di fine: 27/02/2026

Loghi degli enti finanziatori