Machine Learning

These workflows start from a Pandas DataFrame, as shown in event-data import. Explore performance statistics when choosing features, and use conformance diagnostics to derive model-based features.

These examples collect PM4Py functionality that prepares event data for machine learning and applies common ML workflows on top of process data. The current examples focus on feature extraction, event- and case-level representations, dimensionality reduction, anomaly detection, decision trees, and decision mining.

Feature Engineering and Encodings

Automatic Feature Selection

In PM4Py, we offer methods to perform automatic feature selection. As an example, let's import the receipt log and apply automatic feature selection. First, we import the receipt log:

import pm4py
import pandas

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes("tests/input_data/receipt.xes")

Then, let's perform automatic feature selection:

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any

if __name__ == "__main__":
	data: Any
	feature_names: list[str]
	data, feature_names = trace_encodings.apply(log)
	print(feature_names)

Printing the value of feature_names, we observe that the following attributes were selected:

  • The channel attribute at the trace level (with values: Desk, Intern, Internet, Post, e-mail).
  • The department attribute at the trace level (with values: Customer contact, Experts, General).
  • The group attribute at the event level (with values: EMPTY, Group 1, Group 12, Group 13, Group 14, Group 15, Group 2, Group 3, Group 4, Group 7).

No numeric attribute is selected. The printed feature_names are represented as:

[ trace:channel@Desk, trace:channel@Intern, trace:channel@Internet, trace:channel@Post, trace:channel@e-mail, trace:department@Customer contact, trace:department@Experts, trace:department@General, event:org:group@EMPTY, event:org:group@Group 1, event:org:group@Group 12, event:org:group@Group 13, event:org:group@Group 14, event:org:group@Group 15, event:org:group@Group 2, event:org:group@Group 3, event:org:group@Group 4, event:org:group@Group 7 ].

As shown, different features correspond to different attribute values. This technique is called one-hot encoding: a case is assigned a value of 0 if it does not contain an event with the given attribute value, and 1 if it contains at least one such event.

Representing the features as a dataframe:

import pandas as pd
import pandas
if __name__ == "__main__":
	df: pandas.DataFrame = pd.DataFrame(data, columns=feature_names)
	print(df)

We can observe the features assigned to each individual case.

Manual Feature Selection

Manual feature selection allows users to specify which attributes should be included. These may include, for example:

  • Activities performed during process execution (usually stored in the event attribute concept:name).
  • Resources performing the process execution (usually stored in the event attribute org:resource).
  • Selected numeric attributes, at the user's discretion.

To perform manual feature selection, we use the method trace_encodings.apply. The following types of features can be considered:

ParameterDescription
str_ev_attrString attributes at the event level, one-hot encoded to assume values of 0 or 1.
str_tr_attrString attributes at the trace level, one-hot encoded to assume values of 0 or 1.
num_ev_attrNumeric attributes at the event level, encoded by taking the last value observed among the trace's events.
num_tr_attrNumeric attributes at the trace level, encoded by including their numeric value.
str_evsucc_attrSuccessions of string attribute values at the event level: for instance, given a trace [A, B, C], features will include not only A, B, and C individually, but also directly-follows pairs (A, B) and (B, C).

For example, consider a feature selection where we are interested in:

  • Whether a process execution contains a specific activity.
  • Whether a process execution involves a specific resource.
  • Whether a process execution contains a specific directly-follows path between activities.
  • Whether a process execution contains a specific directly-follows path between resources.

In this case, the number of features becomes significantly larger.

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any

if __name__ == "__main__":
	data: Any
	feature_names: list[str]
	data, feature_names = trace_encodings.apply(log, parameters={"str_ev_attr": ["concept:name", "org:resource"], "str_tr_attr": [], "num_ev_attr": [], "num_tr_attr": [], "str_evsucc_attr": ["concept:name", "org:resource"]})
	print(len(feature_names))

Calculating Useful Features

Lifecycle conversion and enrichment operate on a legacy EventLog. The examples create interval_log and enriched_log for those operations while retaining the original DataFrame as log.

Other important features include the cycle time and the lead time associated with a case. In this context, we may assume one of the following:

  • A log with lifecycles, where each event is instantaneous,
  • Or an interval log, where events are associated with two timestamps (start and end).

Lead and cycle times can be calculated directly from interval logs. If we have a lifecycle log, we first need to convert it using:

from pm4py.objects.log.util import interval_lifecycle
from pm4py.objects.log.obj import EventLog
if __name__ == "__main__":
	interval_log: EventLog = interval_lifecycle.to_interval(pm4py.convert_to_event_log(log))

After conversion, features such as lead and cycle times can be added using the following instructions:

from pm4py.objects.log.util import interval_lifecycle
from pm4py.util import constants
from pm4py.objects.log.obj import EventLog

if __name__ == "__main__":
	enriched_log: EventLog = interval_lifecycle.assign_lead_cycle_time(interval_log, parameters={
		constants.PARAMETER_CONSTANT_START_TIMESTAMP_KEY: "start_timestamp",
		constants.PARAMETER_CONSTANT_TIMESTAMP_KEY: "time:timestamp"})

Once the start timestamp attribute (e.g., start_timestamp) and the timestamp attribute (e.g., time:timestamp) are provided, the following features are returned:

  • @@approx_bh_partial_cycle_time: Incremental cycle time associated with the event (the final event's cycle time is the instance's total cycle time).
  • @@approx_bh_partial_lead_time: Incremental lead time associated with the event.
  • @@approx_bh_overall_wasted_time: Difference between the partial lead time and the partial cycle time.
  • @@approx_bh_this_wasted_time: Wasted time specifically related to the activity described by the 'interval' event.
  • @@approx_bh_ratio_cycle_lead_time: Measures the incremental flow rate (ranging from 0 to 1).

Since these are all numerical attributes, we can further refine the feature extraction by applying:

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any

if __name__ == "__main__":
	data: Any
	feature_names: list[str]
	data, feature_names = trace_encodings.apply(enriched_log, parameters={"str_ev_attr": ["concept:name", "org:resource"], "str_tr_attr": [], "num_ev_attr": ["@@approx_bh_partial_cycle_time", "@@approx_bh_partial_lead_time", "@@approx_bh_overall_wasted_time", "@@approx_bh_this_wasted_time", "@approx_bh_ratio_cycle_lead_time"], "num_tr_attr": [], "str_evsucc_attr": ["concept:name", "org:resource"]})

Additionally, we offer the calculation of further intra- and inter-case features, which can be enabled by setting boolean parameters in the trace_encodings.apply method, including:

  • ENABLE_CASE_DURATION: Adds case duration as an additional feature.
  • ENABLE_TIMES_FROM_FIRST_OCCURRENCE: Adds times measured from the first occurrence of an activity within the case.
  • ENABLE_TIMES_FROM_LAST_OCCURRENCE: Adds times measured from the last occurrence of an activity within the case.
  • ENABLE_DIRECT_PATHS_TIMES_LAST_OCC: Adds the duration of the last occurrence of a directed (i, i+1) path as a feature.
  • ENABLE_INDIRECT_PATHS_TIMES_LAST_OCC: Adds the duration of the last occurrence of an indirect (i, j) path as a feature.
  • ENABLE_WORK_IN_PROGRESS: Adds the number of concurrent cases as a feature (work in progress).
  • ENABLE_RESOURCE_WORKLOAD: Adds the workload of resources as a feature.

Model Preparation and Diagnostics

PCA - Reducing the Number of Features

Techniques such as clustering, prediction, and anomaly detection can suffer when the dataset has too many features. Therefore, dimensionality reduction techniques (like PCA) help manage the complexity of the data. Starting from a Pandas dataframe generated from the extracted features:

import pandas as pd
import pandas

if __name__ == "__main__":
	df: pandas.DataFrame = pd.DataFrame(data, columns=feature_names)

It is possible to reduce the number of features using PCA. For example, we can create a PCA model with 5 components and apply it to the dataframe:

from sklearn.decomposition import PCA
import pandas

if __name__ == "__main__":
	pca = PCA(n_components=5)
	df2: pandas.DataFrame = pd.DataFrame(pca.fit_transform(df))

In this way, more than 400 columns are reduced to 5 principal components that capture most of the data variance.

Anomaly Detection

In this section, we focus on calculating an anomaly score for each case. This score is based on the extracted features and works best when combined with dimensionality reduction (such as PCA). We can apply a method called IsolationForest to the dataframe, which adds a column of scores: cases with a score ≤ 0 are considered anomalous, while those with a score > 0 are not.

from sklearn.ensemble import IsolationForest
if __name__ == "__main__":
	model = IsolationForest()
	model.fit(df2)
	df2["scores"] = model.decision_function(df2)

To identify the most anomalous cases, we can sort the dataframe after inserting an index. The resulting output highlights the most anomalous cases:

if __name__ == "__main__":
	df2["@@index"] = df2.index
	df2 = df2[["scores", "@@index"]]
	df2 = df2.sort_values("scores")
	print(df2)

Evolution of the Features

We might be interested in observing how features evolve over time to detect positions in the event log that show behavior different from the mainstream. PM4Py provides a method to graph feature evolution over time. Here is an example:

import os
import pm4py
from pm4py.algo.transformation.trace_encodings.util import locally_linear_embedding
from pm4py.visualization.graphs import visualizer
import pandas
from graphviz import Graph

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "receipt.xes"))
	x: list
	y: list[float]
	x, y = locally_linear_embedding.apply(log)
	gviz: Graph = visualizer.apply(x, y, variant=visualizer.Variants.DATES,
							parameters={"title": "Locally Linear Embedding", "format": "svg", "y_axis": "Intensity"})
	visualizer.view(gviz)

Trace and Event Representations

Event-based Feature Extraction

Some machine learning methods (e.g., LSTM-based deep learning) require features at the event level, instead of aggregating features at the case level. In these methods, each event is represented as a numerical row containing features related to that event. We can perform a default event-based feature extraction as follows:

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any

if __name__ == "__main__":
	data: Any
	features: list[str]
	data, features = trace_encodings.apply(log, variant=trace_encodings.Variants.EVENT_BASED)

Alternatively, it is possible to manually specify the features to be extracted. The parameters str_ev_attr and num_ev_attr correspond to those described in previous sections:

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any

if __name__ == "__main__":
	data: Any
	features: list[str]
	data, features = trace_encodings.apply(log, variant=trace_encodings.Variants.EVENT_BASED, parameters={"str_ev_attr": ["concept:name"], "num_ev_attr": []})

One-hot Trace Encoding

One-hot trace encoding represents every case as a binary vector over trace tokens. A column is set to 1 when the corresponding token appears at least once in the trace.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	log = pm4py.format_dataframe(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.ONE_HOT, parameters={"event_attributes": ["concept:name"]})
	print(feature_names[:5])
	print(data[0][:5])

Count2Vec Trace Encoding

Count2Vec uses the same trace tokenization idea as one-hot encoding, but stores token frequencies instead of only presence or absence.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	log = pm4py.format_dataframe(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.COUNT2VEC, parameters={"event_attributes": ["concept:name"]})
	print(feature_names[:5])
	print(data[0][:5])

N-gram Trace Encoding

N-gram trace encoding includes local ordering information by counting contiguous windows of trace tokens, such as directly-following activity pairs.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	log = pm4py.format_dataframe(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.N_GRAMS, parameters={"event_attributes": ["concept:name"], "ngram_range": (2, 2)})
	print(feature_names[:5])
	print(data[0][:5])

TF-IDF Trace Encoding

TF-IDF weights trace tokens by their importance in a case relative to the rest of the log, making rarer behavior more visible to downstream machine learning methods.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	log = pm4py.format_dataframe(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.TF_IDF, parameters={"event_attributes": ["concept:name"], "ngram_range": (1, 2)})
	print(feature_names[:5])
	print(data[0][:5])

Doc2Vec Trace Encoding

Doc2Vec treats each trace as a document and learns a dense vector for the whole case. This variant uses gensim and requires the optional dependency to be installed.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	log = pm4py.format_dataframe(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.DOC2VEC, parameters={"event_attributes": ["concept:name"], "vector_size": 16, "epochs": 20})
	print(feature_names[:5])
	print(data[0][:5])

Word2Vec Trace Encoding

Word2Vec learns vectors for event tokens from their trace context and aggregates the token vectors into one vector per trace. This variant uses gensim.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	log = pm4py.format_dataframe(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.WORD2VEC, parameters={"event_attributes": ["concept:name"], "vector_size": 16, "epochs": 20})
	print(feature_names[:5])
	print(data[0][:5])

BERT Trace Encoding

BERT-style trace encoding converts each trace into a sentence and embeds it through a sentence-transformers model. The model name can point to a cached model or a local model path.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	log = pm4py.format_dataframe(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.BERT, parameters={"event_attributes": ["concept:name"], "bert_model": "bert-base-nli-mean-tokens"})
	print(feature_names[:5])
	print(data[0][:5])

Token Replay Trace Encoding

Token replay trace encoding replays each trace on a Petri net and stores conformance diagnostics, such as fitness and missing or remaining tokens, as numerical features.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	net, im, fm = pm4py.discover_petri_net_inductive(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.TOKEN_REPLAY, parameters={"net": net, "initial_marking": im, "final_marking": fm})
	print(feature_names)
	print(data[0])

Alignment Trace Encoding

Alignment trace encoding aligns each trace against a Petri net and stores alignment diagnostics such as fitness, cost, best-worst cost, and search effort counters.

import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
	net, im, fm = pm4py.discover_petri_net_inductive(log)
	data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.ALIGNMENTS, parameters={"net": net, "initial_marking": im, "final_marking": fm})
	print(feature_names)
	print(data[0])

Predictive and Decision-Oriented ML

Decision Tree About the Ending Activity of a Process

The class-label helper requires legacy traces. Convert the DataFrame at that call, preserving case order so the labels match the feature rows.

Decision trees are tools that help understand the conditions leading to a particular outcome. In this section, several examples related to the construction of decision trees are provided. The ideas behind building decision trees are discussed in the scientific paper: de Leoni, Massimiliano, Wil MP van der Aalst, and Marcus Dees. "A General Process Mining Framework for Correlating, Predicting, and Clustering Dynamic Behavior Based on Event Logs."

The general procedure is as follows:

  • Obtain a representation of the log based on a given set of features (e.g., using one-hot encoding for string attributes and preserving numeric attributes as they are).
  • Construct a representation of the target classes.
  • Build the decision tree.
  • Visualize the decision tree.

A process instance may potentially finish with different activities, signaling different outcomes. A decision tree can help understand the reasons behind each outcome. First, a log is loaded, and then a feature-based representation of the log is created.

import os
import pm4py
import pandas
from typing import Any
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "roadtraffic50traces.xes"))

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

if __name__ == "__main__":
	data: Any
	feature_names: list[str]
	data, feature_names = trace_encodings.apply(log, parameters={"str_tr_attr": [], "str_ev_attr": ["concept:name"], "num_tr_attr": [], "num_ev_attr": ["amount"]})

Alternatively, an automatic feature representation (automatic attribute selection) can be obtained:

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log)

(Optional) The extracted features can be represented as a Pandas DataFrame:

import pandas as pd
import pandas
if __name__ == "__main__":
	dataframe: pandas.DataFrame = pd.DataFrame(data, columns=feature_names)

(Optional) The DataFrame can then be exported as a CSV file:

if __name__ == "__main__":
	dataframe.to_csv("features.csv", index=False)

Next, the target classes are defined: each endpoint activity of the process instance is assigned to a different class.

from pm4py.objects.log.util import get_class_representation
from typing import Any
if __name__ == "__main__":
	target: Any
	classes: list[str]
	target, classes = get_class_representation.get_class_representation_by_str_ev_attr_value_value(pm4py.convert_to_event_log(log), "concept:name")

The decision tree is then built and visualized:

from sklearn import tree
from graphviz import Graph
from sklearn.tree import DecisionTreeClassifier
if __name__ == "__main__":
	clf: DecisionTreeClassifier = tree.DecisionTreeClassifier()
	clf.fit(data, target)

	from pm4py.visualization.decisiontree import visualizer as dectree_visualizer
	gviz: Graph = dectree_visualizer.apply(clf, feature_names, classes)

Decision Tree About the Duration of a Case

The duration-label helper also requires legacy traces; the explicit conversion preserves the feature-row order.

A decision tree regarding the duration of a case helps understand the factors behind a high case duration (i.e., durations above a given threshold). First, a log is loaded, and a feature-based representation is created.

import os
import pm4py
import pandas
from typing import Any
if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "roadtraffic50traces.xes"))

	from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings

	data: Any
	feature_names: list[str]
	data, feature_names = trace_encodings.apply(log, parameters={"str_tr_attr": [], "str_ev_attr": ["concept:name"], "num_tr_attr": [], "num_ev_attr": ["amount"]})

Alternatively, an automatic feature representation can be generated:

from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log)

Then, the target classes are formed:

  • Traces below the specified threshold (e.g., 200 days — time measured in seconds).
  • Traces above the specified threshold.
from pm4py.objects.log.util import get_class_representation
from typing import Any
if __name__ == "__main__":
	target: Any
	classes: list[str]
	target, classes = get_class_representation.get_class_representation_by_trace_duration(pm4py.convert_to_event_log(log), 2 * 8640000)

The decision tree is then built and visualized:

from sklearn import tree
from graphviz import Graph
from sklearn.tree import DecisionTreeClassifier
if __name__ == "__main__":
	clf: DecisionTreeClassifier = tree.DecisionTreeClassifier()
	clf.fit(data, target)

	from pm4py.visualization.decisiontree import visualizer as dectree_visualizer
	gviz: Graph = dectree_visualizer.apply(clf, feature_names, classes)

Decision Mining

Decision Mining enables the following, given:

  • An event log,
  • A process model (an accepting Petri net),
  • A decision point,

It retrieves the features of the cases that take different paths. This allows, for example, building a decision tree to explain the choices made.

First, import a XES log:

import pm4py
import pandas

if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes("tests/input_data/running-example.xes")

Next, calculate a model using the Inductive Miner:

from pm4py.objects.petri_net.obj import Marking, PetriNet
if __name__ == "__main__":
	net: PetriNet
	im: Marking
	fm: Marking
	net, im, fm = pm4py.discover_petri_net_inductive(log)

To visualize the model:

from pm4py.visualization.petri_net import visualizer
from graphviz import Graph

if __name__ == "__main__":
	gviz: Graph = visualizer.apply(net, im, fm, parameters={visualizer.Variants.WO_DECORATION.value.Parameters.DEBUG: True})
	visualizer.view(gviz)

For this example, we select decision point p_10, where a choice is made between the activities examine casually and examine thoroughly. Once we have a log, a model, and a decision point, the decision mining algorithm can be executed:

from pm4py.algo.decision_mining import algorithm as decision_mining
from typing import Any

if __name__ == "__main__":
	X: Any
	y: Any
	class_names: list[str]
	X, y, class_names = decision_mining.apply(log, net, im, fm, decision_point="p_10")

The outputs of the apply method are:

  • X: A Pandas DataFrame containing the features associated with each case leading to a decision.
  • y: A Pandas Series containing the class (output) of each decision (e.g., 0 or 1).
  • class_names: The names of the possible decision outcomes (e.g., examine casually and examine thoroughly).

These outputs can be used with any classification or comparison technique. In particular, decision trees are a useful choice. We provide a function to automatically discover decision trees from decision mining results:

from pm4py.algo.decision_mining import algorithm as decision_mining
from sklearn.tree import DecisionTreeClassifier

if __name__ == "__main__":
	clf: DecisionTreeClassifier
	feature_names: list[str]
	classes: list[str]
	clf, feature_names, classes = decision_mining.get_decision_tree(log, net, im, fm, decision_point="p_10")

To visualize the resulting decision tree:

from pm4py.visualization.decisiontree import visualizer as tree_visualizer
from graphviz import Graph
if __name__ == "__main__":
	gviz: Graph = tree_visualizer.apply(clf, feature_names, classes)

Feature Tables and Process-Aware Features

Feature Extraction on DataFrames

While the feature extraction described above is generic, it might not be optimal (performance-wise) when working directly with Pandas DataFrames. We also offer the option to extract a feature table by providing:

  • The DataFrame,
  • A set of columns to use as features.

The output is another DataFrame containing:

  • The case identifier.
  • For each string attribute: a one-hot encoding counting the number of occurrences for each possible value.
  • For each numeric attribute: the last value observed within each case.

Here is an example that keeps concept:name (activity) and amount (cost) as features:

import pm4py
import pandas as pd
from pm4py.objects.log.util import dataframe_utils
import pandas

if __name__ == "__main__":
	dataframe: pandas.DataFrame = pd.read_csv("tests/input_data/roadtraffic100traces.csv")
	dataframe = pm4py.format_dataframe(dataframe)
	feature_table: pandas.DataFrame = dataframe_utils.get_features_df(dataframe, ["concept:name", "amount"])

The resulting feature table will contain columns such as:

['case:concept:name', 'concept:name_CreateFine', 'concept:name_SendFine', 'concept:name_InsertFineNotification', 'concept:name_Addpenalty', 'concept:name_SendforCreditCollection', 'concept:name_Payment', 'concept:name_InsertDateAppealtoPrefecture', 'concept:name_SendAppealtoPrefecture', 'concept:name_ReceiveResultAppealfromPrefecture', 'concept:name_NotifyResultAppealtoOffender', 'amount']

Discovery of a Data Petri Net

Given a Petri net discovered by a classical process mining algorithm (e.g., Alpha Miner or Inductive Miner), we can enhance it into a Data Petri Net by applying decision mining at every decision point, and transforming the resulting decision trees into guards (boolean conditions).

An example:

import pm4py
import pandas
from pm4py.objects.petri_net.obj import Marking, PetriNet
if __name__ == "__main__":
	log: pandas.DataFrame = pm4py.read_xes("tests/input_data/roadtraffic100traces.xes")
	net: PetriNet
	im: Marking
	fm: Marking
	net, im, fm = pm4py.discover_petri_net_inductive(log)
	from pm4py.algo.decision_mining import algorithm as decision_mining
	net, im, fm = decision_mining.create_data_petri_nets_with_decisions(log, net, im, fm)

The guards discovered for each transition can be printed. They are expressed as boolean conditions and interpreted by the execution engine:

if __name__ == "__main__":
	for t in net.transitions:
		if "guard" in t.properties:
			print("")
			print(t)
			print(t.properties["guard"])

Temporal Feature Extraction

The PM4Py library provides a method to extract temporal features from an event log, event stream, or Pandas DataFrame, as described in the paper by Pourbafrani et al. (2020). This method groups events by a specified time granularity (e.g., weekly) and computes aggregated metrics to represent process behavior over time.

To apply temporal feature extraction, you can use the apply function from the provided code. Here's an example:

import pm4py
from pm4py.algo.transformation.trace_encodings.variants import temporal as temporal_features
import pandas
# Import event data as a DataFrame
log: pandas.DataFrame = pm4py.read_xes("log.xes")

# Apply temporal feature extraction
features_df: pandas.DataFrame = temporal_features.apply(log, parameters={
    "grouper_freq": "W",
    "arrival_rate": "arrival_rate",
    "finish_rate": "finish_rate",
    "service_time": "service_time",
    "waiting_time": "waiting_time",
    "sojourn_time": "sojourn_time"
})

The resulting features_df is a Pandas DataFrame containing temporal features grouped by the specified frequency (e.g., weekly).

Parameters

The algorithm accepts several parameters to customize the feature extraction process. These are defined in the Parameters enum and include:

ParameterDescription
GROUPER_FREQTime interval for grouping events (e.g., "W" for weekly, "D" for daily). Default: "W".
ARRIVAL_RATEColumn name for the arrival rate of cases. Default: "arrival_rate".
FINISH_RATEColumn name for the completion rate of cases. Default: "finish_rate".
SERVICE_TIMEColumn name for the service time (time spent on activities). Default: "service_time".
WAITING_TIMEColumn name for the waiting time (time spent idle). Default: "waiting_time".
SOJOURN_TIMEColumn name for the sojourn time (total time from start to end). Default: "sojourn_time".
CASE_ID_COLUMNColumn name for case IDs. Default: "case:concept:name".
ACTIVITY_COLUMNColumn name for activities. Default: "concept:name".
TIMESTAMP_COLUMNColumn name for event timestamps. Default: "time:timestamp".
START_TIMESTAMP_COLUMNColumn name for start timestamps (if available). Defaults to the timestamp column.
RESOURCE_COLUMNColumn name for resources. Default: "org:resource".

These parameters allow users to specify the granularity and naming conventions for the extracted features, tailoring the output to their specific needs.

Process Overview

The temporal feature extraction process involves the following steps:

  1. Log Conversion: The input log (EventLog, EventStream, or DataFrame) is converted to a Pandas DataFrame for processing.
  2. Arrival and Finish Rates: The algorithm calculates the arrival rate (how often new cases start) and finish rate (how often cases complete) for each case, adding these as columns to the DataFrame.
  3. Service, Waiting, and Sojourn Times: For each case, the algorithm computes:
    • Service Time: Time spent actively processing activities.
    • Waiting Time: Time spent idle between activities.
    • Sojourn Time: Total time from the start to the end of a case.
  4. Grouping by Time: Events are grouped by the specified time frequency (e.g., weekly) based on the start timestamp.
  5. Feature Aggregation: For each time group, the algorithm computes:
    • Number of unique resources, cases, and activities.
    • Total number of events.
    • Average arrival and finish rates.
    • Average service, waiting, and sojourn times.
  6. Output: A DataFrame is returned with columns for the timestamp of each group and the computed features. Missing values are filled with 0.

The resulting DataFrame provides a tabular representation of temporal process characteristics, suitable for further analysis or visualization.

Example Output

The output DataFrame might look like this:

import pandas
from typing import Any
# Example output DataFrame
import pandas as pd

data: list[dict[str, Any]] = [
    {"timestamp": "2023-01-01", "unique_resources": 5, "unique_cases": 10, "unique_activities": 3, "num_events": 50, 
     "average_arrival_rate": 2.5, "average_finish_rate": 2.3, "average_service_time": 3600, 
     "average_waiting_time": 7200, "average_sojourn_time": 10800},
    {"timestamp": "2023-01-08", "unique_resources": 4, "unique_cases": 8, "unique_activities": 4, "num_events": 45, 
     "average_arrival_rate": 2.0, "average_finish_rate": 1.9, "average_service_time": 3400, 
     "average_waiting_time": 7000, "average_sojourn_time": 10400}
]
features_df: pandas.DataFrame = pd.DataFrame(data)
print(features_df)

This DataFrame shows temporal features for two weekly groups, including counts of unique elements and averages of time-based metrics.

Use Cases

Temporal feature extraction is particularly useful for:

  • Process Simulation: Generating system dynamics models for simulation, as described in the referenced paper.
  • Performance Analysis: Identifying bottlenecks by analyzing waiting and service times.
  • Trend Detection: Observing how process metrics evolve over time to detect anomalies or shifts in behavior.
  • Predictive Modeling: Using temporal features as inputs for machine learning models to predict process outcomes.

Advanced Usage

To customize the feature extraction further, you can modify the parameters. For example, to group by days and use custom column names:

from pm4py.algo.transformation.trace_encodings.variants import temporal as temporal_features
import pandas
	
features_df: pandas.DataFrame = temporal_features.apply(log, parameters={
    "grouper_freq": "D",
    "arrival_rate": "case_arrival",
    "finish_rate": "case_completion",
    "service_time": "activity_duration",
    "waiting_time": "idle_time",
    "sojourn_time": "total_duration",
    "pm4py:param:case_id_key": "case:concept:name",
    "pm4py:param:timestamp_key": "time:timestamp",
    "pm4py:param:start_timestamp_key": "time:timestamp",
    "pm4py:param:resource_key": "org:resource",
    "pm4py:param:activity_key": "concept:name"
})

This allows the algorithm to adapt to different log formats and analysis requirements.

References

The approach is based on the following paper:

  • Pourbafrani, Mahsa, Sebastiaan J. van Zelst, and Wil MP van der Aalst. "Supporting automatic system dynamics model generation for simulation in the context of process mining." International Conference on Business Information Systems. Springer, Cham, 2020.