Foto del docente

Mohsen Seyedkazemi Ardebili

Research fellow

Department of Electrical, Electronic, and Information Engineering "Guglielmo Marconi"

Research

Keywords: High-Performance Computing (HPC) AIOps and AI for infrastructure operations LLM-based autonomous agents Kubernetes and cloud-native systems MLOps and AI in production Anomaly detection and prediction Digital twins for computing systems Datacenter monitoring and telemetry Energy efficiency and sustainable computing Deep learning for time series Edge-cloud-HPC computing continuum

His research applies artificial intelligence to the operation of large-scale computing infrastructure, from Tier-0 supercomputers to cloud-native Kubernetes clusters, along four connected themes.

1. Autonomous agents for HPC and Kubernetes operations

Design of LLM-orchestrated multi-agent architectures that carry out real operations on computing clusters rather than only advising on them, with an explicit human-in-the-loop approval gate on every mutating action, runtime synthesis of new tools, and role-based access control. This theme also covers the evaluation problem: how to measure, under realistic institutional access-control policies, whether such agents are safe enough to operate production infrastructure (KubeIntellect, AOBench).

2. Anomaly detection and predictive monitoring for datacentres and supercomputers

Scalable frameworks for detecting and anticipating thermal, system and job-level anomalies in exascale-class facilities, using deep learning, statistical modelling and graph-based methods over large-scale telemetry. Work in this theme has been developed and validated on CINECA's Marconi100 Tier-0 system and has produced open datasets for the community (HazardNet, ThermADNet, GRAAFE, M100 ExaData, PM100).

3. Digital twins and AI-driven scheduling across the computing continuum

Real-time digital-twin modelling of heterogeneous infrastructure spanning devices, edge, cloud and HPC, built on Prometheus-based telemetry, and its use as the observability foundation for AI-driven workload placement across Kubernetes and SLURM environments with energy-efficiency and latency objectives.

4. MLOps, AI in production, and sustainable computing

End-to-end MLOps platforms for the lifecycle of AI models in HPC and Kubernetes environments - training pipelines, model versioning, multi-model inference serving, drift detection and metric-gated promotion - together with full-stack observability, and their application to power-consumption analysis, carbon-intensity forecasting and carbon-aware job placement for more sustainable supercomputing.