Probabilistic declarative process mining (PRODE)

PRIN 2022 Chesani

Abstract

The management of business processes is of extreme importance for supporting efficiency improvements in organisations. Process Mining (PM) deals with the discovery and representation of process models from event data collected from organisations about the executions of their processes. The project addresses the task of learning and reasoning upon declarative process models, within the setting of binary supervised learning, taking into account also uncertainty. With the aim of providing viable solutions, the PRODE project will focus on the following issues in particular: 1) exploit the availability of positive and negative examples: in many cases, user experts provide examples with desired and undesired behaviour (hence the labels “positive” and “negative”), but the majority of the discovery approaches exploits only the positive set; 2) precision and understandability of discovered models: precise models could perfectly discriminate between positive and negative examples, but might turn out to be too complex for the final user. We might want to learn models which do not perfectly discriminate between positive and negative examples, but which are simpler and understandable for the final user: probabilistic approaches might help to simplify the models, yet providing a clear and formal semantics; 3) deal with the uncertainty that real logs usually bear: on one side, logs are just a partial incomplete view of the reality; on the other side, the information in the log might be incomplete, partially specified, and even non trustable; 4) compliance issues: while many approaches provide a crisp yes/no answer to the question if a trace “is conformant” with a model, compliance we will explore the possibility of returning a score representing the probability/degree of a trace to be compliant to the model; 5) model selection issues: as there can be multiple output models from the process discovery task, with an associated uncertainty as a result of point 3), we might want to identify the preferable models, in order to improve the workflow management. The PRODE project will take advantage of the development of works in the fields of Probabilistic Logic Programming (PLP) and Answer Set Programming (ASP), and will build a set of techniques that target the issues above by means of new combinations of declarative Process Mining with probabilistic and combinatorial approaches. The final aim is to produce more verifiable and understandable explanations of its processes to an organisation. To accomplish these objectives PRODE builds on the expertise of the research units in the fields of Process Mining, Artificial Intelligence, knowledge discovery, Machine Learning (in particular, Statistical Relational Learning), Logic Programming, Probabilistic Logic Programming, and Answer Set Programming. Results will be verified through both formal and experimental analysis on a variety of case studies.

Results achieved

The PRODE project terminated successfully, achieving all the goals that have been mentioned in the project proposal. These achievements have been obtained through an intense and proficuous collaboration with the project partners. A noteworthy result has been the definition of a novel semantics for probablities related to declarative constraints. Within the project we mainly focus on Decalre constraints, although the results can be applied to other declarative approaches. Declare constraints provide an elegant and concise way for describing rules (constraints) that each process instance (i.e., the execution of a process) should satisfy. If interpreted in a prescriptive manner, they specify the boundaries of process executions, without limiting their flexibility. If interpreted a descriptive manner, they provide a high-level, easy-to-understand language for modelling processes. Unfortunately, existing Declare semantics are based on a cripy notion of compliance: a process execution: (a trace) either complies with the constraint or not. In the novel semantics we have adapted the Distribution Semantics for Logic Programs to the Declare constraints: a constraint can be labeled with a probability p, and the intended meaning is grounded on an epistemiological perspective: the importance of the constraint is given by p. Following this view, a novel notion of conformance of a trace versus a constraint has been defined: a trace is compliant with p::c with value 1 if it is compliance with c, and with value (1-p) if it violates the constraint c. This novel semantics has been extended, then to deal with conjunctions of constraints, .i.e. to process models. The uncertainty has been extended also to events within the single traces: in real scenarios, for example, sensors might provide uncertain data, that is represented in our approach through a degree of uncertainty about the corresponding event. The same semantics has been applied then to the single events; to the best of our knowledge, this is the first time a single, unified semantics has been proposed to provide uncertainty to both events and constraints. From the computational viewpoint, usually adding probabilities increase the computational complexity of the verification task. Within the project, we obtained a formal result that decrease the computational complexity of our approach to the same class of the crisp-logic-based approaches. This result holds for uncertainty related to the constraints, and is a particularly important result: indeed, dealing with uncertainty attached to constraints does not increase the conformance computational complexity. This result paves the way to a straightforward adoption of our semantics, since existing algorithms can be re-used as such, and only a very minor extension should be applied a posteriori to existing approaches. Also the discovery task has been tackled, and some existing approaches have been extended and adapted to deal with the uncertainty. Probabilistic Declare models can be then obtained through a “de jure” approach, where the importance of each constraint is discussed, or can be obtained through a learning process over logs, where a frequency-based approach is exploited: in the latter case, the probabilities resemble a description of which constraints hold, and of their “weight”. The results described above have been published at international conferences and in international journals.

Project details

Unibo Team Leader: Federico Chesani

Unibo involved Department/s:
Dipartimento di Informatica - Scienza e Ingegneria

Coordinator:
Università  degli Studi di Ferrara - Amministrazione Centrale(Italy)

Total Unibo Contribution: Euro (EUR) 58.865,00
Project Duration in months: 24
Start Date: 28/09/2023
End Date: 28/02/2026

Funding bodies' logos