These workflows start from a Pandas DataFrame, as shown in event-data import. Explore performance statistics when choosing features, and use conformance diagnostics to derive model-based features.
These examples collect PM4Py functionality that prepares event data for machine learning and applies common ML workflows on top of process data. The current examples focus on feature extraction, event- and case-level representations, dimensionality reduction, anomaly detection, decision trees, and decision mining.
In PM4Py, we offer methods to perform automatic feature selection. As an example, let's import the receipt log and apply automatic feature selection. First, we import the receipt log:
import pm4py
import pandas
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes("tests/input_data/receipt.xes")Then, let's perform automatic feature selection:
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
if __name__ == "__main__":
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log)
print(feature_names) Printing the value of feature_names, we observe that the following attributes were selected:
channel attribute at the trace level (with values: Desk, Intern, Internet, Post, e-mail).department attribute at the trace level (with values: Customer contact, Experts, General).group attribute at the event level (with values: EMPTY, Group 1, Group 12, Group 13, Group 14, Group 15, Group 2, Group 3, Group 4, Group 7). No numeric attribute is selected. The printed feature_names are represented as:
[ trace:channel@Desk, trace:channel@Intern, trace:channel@Internet, trace:channel@Post, trace:channel@e-mail, trace:department@Customer contact, trace:department@Experts, trace:department@General, event:org:group@EMPTY, event:org:group@Group 1, event:org:group@Group 12, event:org:group@Group 13, event:org:group@Group 14, event:org:group@Group 15, event:org:group@Group 2, event:org:group@Group 3, event:org:group@Group 4, event:org:group@Group 7 ].
As shown, different features correspond to different attribute values. This technique is called one-hot encoding: a case is assigned a value of 0 if it does not contain an event with the given attribute value, and 1 if it contains at least one such event.
Representing the features as a dataframe:
import pandas as pd
import pandas
if __name__ == "__main__":
df: pandas.DataFrame = pd.DataFrame(data, columns=feature_names)
print(df)We can observe the features assigned to each individual case.
Manual feature selection allows users to specify which attributes should be included. These may include, for example:
concept:name). org:resource). To perform manual feature selection, we use the method trace_encodings.apply. The following types of features can be considered:
| Parameter | Description |
|---|---|
str_ev_attr | String attributes at the event level, one-hot encoded to assume values of 0 or 1. |
str_tr_attr | String attributes at the trace level, one-hot encoded to assume values of 0 or 1. |
num_ev_attr | Numeric attributes at the event level, encoded by taking the last value observed among the trace's events. |
num_tr_attr | Numeric attributes at the trace level, encoded by including their numeric value. |
str_evsucc_attr | Successions of string attribute values at the event level: for instance, given a trace [A, B, C], features will include not only A, B, and C individually, but also directly-follows pairs (A, B) and (B, C). |
For example, consider a feature selection where we are interested in:
In this case, the number of features becomes significantly larger.
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
if __name__ == "__main__":
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log, parameters={"str_ev_attr": ["concept:name", "org:resource"], "str_tr_attr": [], "num_ev_attr": [], "num_tr_attr": [], "str_evsucc_attr": ["concept:name", "org:resource"]})
print(len(feature_names)) Lifecycle conversion and enrichment operate on a legacy EventLog. The examples create interval_log and enriched_log for those operations while retaining the original DataFrame as log.
Other important features include the cycle time and the lead time associated with a case. In this context, we may assume one of the following:
Lead and cycle times can be calculated directly from interval logs. If we have a lifecycle log, we first need to convert it using:
from pm4py.objects.log.util import interval_lifecycle
from pm4py.objects.log.obj import EventLog
if __name__ == "__main__":
interval_log: EventLog = interval_lifecycle.to_interval(pm4py.convert_to_event_log(log))After conversion, features such as lead and cycle times can be added using the following instructions:
from pm4py.objects.log.util import interval_lifecycle
from pm4py.util import constants
from pm4py.objects.log.obj import EventLog
if __name__ == "__main__":
enriched_log: EventLog = interval_lifecycle.assign_lead_cycle_time(interval_log, parameters={
constants.PARAMETER_CONSTANT_START_TIMESTAMP_KEY: "start_timestamp",
constants.PARAMETER_CONSTANT_TIMESTAMP_KEY: "time:timestamp"}) Once the start timestamp attribute (e.g., start_timestamp) and the timestamp attribute (e.g., time:timestamp) are provided, the following features are returned:
@@approx_bh_partial_cycle_time: Incremental cycle time associated with the event (the final event's cycle time is the instance's total cycle time).@@approx_bh_partial_lead_time: Incremental lead time associated with the event. @@approx_bh_overall_wasted_time: Difference between the partial lead time and the partial cycle time.@@approx_bh_this_wasted_time: Wasted time specifically related to the activity described by the 'interval' event.@@approx_bh_ratio_cycle_lead_time: Measures the incremental flow rate (ranging from 0 to 1).Since these are all numerical attributes, we can further refine the feature extraction by applying:
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
if __name__ == "__main__":
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(enriched_log, parameters={"str_ev_attr": ["concept:name", "org:resource"], "str_tr_attr": [], "num_ev_attr": ["@@approx_bh_partial_cycle_time", "@@approx_bh_partial_lead_time", "@@approx_bh_overall_wasted_time", "@@approx_bh_this_wasted_time", "@approx_bh_ratio_cycle_lead_time"], "num_tr_attr": [], "str_evsucc_attr": ["concept:name", "org:resource"]}) Additionally, we offer the calculation of further intra- and inter-case features, which can be enabled by setting boolean parameters in the trace_encodings.apply method, including:
ENABLE_CASE_DURATION: Adds case duration as an additional feature.ENABLE_TIMES_FROM_FIRST_OCCURRENCE: Adds times measured from the first occurrence of an activity within the case.ENABLE_TIMES_FROM_LAST_OCCURRENCE: Adds times measured from the last occurrence of an activity within the case.ENABLE_DIRECT_PATHS_TIMES_LAST_OCC: Adds the duration of the last occurrence of a directed (i, i+1) path as a feature.ENABLE_INDIRECT_PATHS_TIMES_LAST_OCC: Adds the duration of the last occurrence of an indirect (i, j) path as a feature.ENABLE_WORK_IN_PROGRESS: Adds the number of concurrent cases as a feature (work in progress).ENABLE_RESOURCE_WORKLOAD: Adds the workload of resources as a feature.Techniques such as clustering, prediction, and anomaly detection can suffer when the dataset has too many features. Therefore, dimensionality reduction techniques (like PCA) help manage the complexity of the data. Starting from a Pandas dataframe generated from the extracted features:
import pandas as pd
import pandas
if __name__ == "__main__":
df: pandas.DataFrame = pd.DataFrame(data, columns=feature_names)It is possible to reduce the number of features using PCA. For example, we can create a PCA model with 5 components and apply it to the dataframe:
from sklearn.decomposition import PCA
import pandas
if __name__ == "__main__":
pca = PCA(n_components=5)
df2: pandas.DataFrame = pd.DataFrame(pca.fit_transform(df))In this way, more than 400 columns are reduced to 5 principal components that capture most of the data variance.
In this section, we focus on calculating an anomaly score for each case. This score is based on the extracted features and works best when combined with dimensionality reduction (such as PCA). We can apply a method called IsolationForest to the dataframe, which adds a column of scores: cases with a score ≤ 0 are considered anomalous, while those with a score > 0 are not.
from sklearn.ensemble import IsolationForest
if __name__ == "__main__":
model = IsolationForest()
model.fit(df2)
df2["scores"] = model.decision_function(df2)To identify the most anomalous cases, we can sort the dataframe after inserting an index. The resulting output highlights the most anomalous cases:
if __name__ == "__main__":
df2["@@index"] = df2.index
df2 = df2[["scores", "@@index"]]
df2 = df2.sort_values("scores")
print(df2)We might be interested in observing how features evolve over time to detect positions in the event log that show behavior different from the mainstream. PM4Py provides a method to graph feature evolution over time. Here is an example:
import os
import pm4py
from pm4py.algo.transformation.trace_encodings.util import locally_linear_embedding
from pm4py.visualization.graphs import visualizer
import pandas
from graphviz import Graph
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "receipt.xes"))
x: list
y: list[float]
x, y = locally_linear_embedding.apply(log)
gviz: Graph = visualizer.apply(x, y, variant=visualizer.Variants.DATES,
parameters={"title": "Locally Linear Embedding", "format": "svg", "y_axis": "Intensity"})
visualizer.view(gviz)Some machine learning methods (e.g., LSTM-based deep learning) require features at the event level, instead of aggregating features at the case level. In these methods, each event is represented as a numerical row containing features related to that event. We can perform a default event-based feature extraction as follows:
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
if __name__ == "__main__":
data: Any
features: list[str]
data, features = trace_encodings.apply(log, variant=trace_encodings.Variants.EVENT_BASED) Alternatively, it is possible to manually specify the features to be extracted. The parameters str_ev_attr and num_ev_attr correspond to those described in previous sections:
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
if __name__ == "__main__":
data: Any
features: list[str]
data, features = trace_encodings.apply(log, variant=trace_encodings.Variants.EVENT_BASED, parameters={"str_ev_attr": ["concept:name"], "num_ev_attr": []})One-hot trace encoding represents every case as a binary vector over trace tokens. A column is set to 1 when the corresponding token appears at least once in the trace.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
log = pm4py.format_dataframe(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.ONE_HOT, parameters={"event_attributes": ["concept:name"]})
print(feature_names[:5])
print(data[0][:5])Count2Vec uses the same trace tokenization idea as one-hot encoding, but stores token frequencies instead of only presence or absence.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
log = pm4py.format_dataframe(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.COUNT2VEC, parameters={"event_attributes": ["concept:name"]})
print(feature_names[:5])
print(data[0][:5])N-gram trace encoding includes local ordering information by counting contiguous windows of trace tokens, such as directly-following activity pairs.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
log = pm4py.format_dataframe(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.N_GRAMS, parameters={"event_attributes": ["concept:name"], "ngram_range": (2, 2)})
print(feature_names[:5])
print(data[0][:5])TF-IDF weights trace tokens by their importance in a case relative to the rest of the log, making rarer behavior more visible to downstream machine learning methods.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
log = pm4py.format_dataframe(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.TF_IDF, parameters={"event_attributes": ["concept:name"], "ngram_range": (1, 2)})
print(feature_names[:5])
print(data[0][:5])Doc2Vec treats each trace as a document and learns a dense vector for the whole case. This variant uses gensim and requires the optional dependency to be installed.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
log = pm4py.format_dataframe(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.DOC2VEC, parameters={"event_attributes": ["concept:name"], "vector_size": 16, "epochs": 20})
print(feature_names[:5])
print(data[0][:5])Word2Vec learns vectors for event tokens from their trace context and aggregates the token vectors into one vector per trace. This variant uses gensim.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
log = pm4py.format_dataframe(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.WORD2VEC, parameters={"event_attributes": ["concept:name"], "vector_size": 16, "epochs": 20})
print(feature_names[:5])
print(data[0][:5])BERT-style trace encoding converts each trace into a sentence and embeds it through a sentence-transformers model. The model name can point to a cached model or a local model path.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
log = pm4py.format_dataframe(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.BERT, parameters={"event_attributes": ["concept:name"], "bert_model": "bert-base-nli-mean-tokens"})
print(feature_names[:5])
print(data[0][:5])Token replay trace encoding replays each trace on a Petri net and stores conformance diagnostics, such as fitness and missing or remaining tokens, as numerical features.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
net, im, fm = pm4py.discover_petri_net_inductive(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.TOKEN_REPLAY, parameters={"net": net, "initial_marking": im, "final_marking": fm})
print(feature_names)
print(data[0])Alignment trace encoding aligns each trace against a Petri net and stores alignment diagnostics such as fitness, cost, best-worst cost, and search effort counters.
import os
import pm4py
import pandas
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
net, im, fm = pm4py.discover_petri_net_inductive(log)
data, feature_names = trace_encodings.apply(log, variant=trace_encodings.Variants.ALIGNMENTS, parameters={"net": net, "initial_marking": im, "final_marking": fm})
print(feature_names)
print(data[0])The class-label helper requires legacy traces. Convert the DataFrame at that call, preserving case order so the labels match the feature rows.
Decision trees are tools that help understand the conditions leading to a particular outcome. In this section, several examples related to the construction of decision trees are provided. The ideas behind building decision trees are discussed in the scientific paper: de Leoni, Massimiliano, Wil MP van der Aalst, and Marcus Dees. "A General Process Mining Framework for Correlating, Predicting, and Clustering Dynamic Behavior Based on Event Logs."
The general procedure is as follows:
A process instance may potentially finish with different activities, signaling different outcomes. A decision tree can help understand the reasons behind each outcome. First, a log is loaded, and then a feature-based representation of the log is created.
import os
import pm4py
import pandas
from typing import Any
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "roadtraffic50traces.xes"))
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
if __name__ == "__main__":
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log, parameters={"str_tr_attr": [], "str_ev_attr": ["concept:name"], "num_tr_attr": [], "num_ev_attr": ["amount"]})Alternatively, an automatic feature representation (automatic attribute selection) can be obtained:
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log)(Optional) The extracted features can be represented as a Pandas DataFrame:
import pandas as pd
import pandas
if __name__ == "__main__":
dataframe: pandas.DataFrame = pd.DataFrame(data, columns=feature_names)(Optional) The DataFrame can then be exported as a CSV file:
if __name__ == "__main__":
dataframe.to_csv("features.csv", index=False)Next, the target classes are defined: each endpoint activity of the process instance is assigned to a different class.
from pm4py.objects.log.util import get_class_representation
from typing import Any
if __name__ == "__main__":
target: Any
classes: list[str]
target, classes = get_class_representation.get_class_representation_by_str_ev_attr_value_value(pm4py.convert_to_event_log(log), "concept:name")The decision tree is then built and visualized:
from sklearn import tree
from graphviz import Graph
from sklearn.tree import DecisionTreeClassifier
if __name__ == "__main__":
clf: DecisionTreeClassifier = tree.DecisionTreeClassifier()
clf.fit(data, target)
from pm4py.visualization.decisiontree import visualizer as dectree_visualizer
gviz: Graph = dectree_visualizer.apply(clf, feature_names, classes)The duration-label helper also requires legacy traces; the explicit conversion preserves the feature-row order.
A decision tree regarding the duration of a case helps understand the factors behind a high case duration (i.e., durations above a given threshold). First, a log is loaded, and a feature-based representation is created.
import os
import pm4py
import pandas
from typing import Any
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "roadtraffic50traces.xes"))
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log, parameters={"str_tr_attr": [], "str_ev_attr": ["concept:name"], "num_tr_attr": [], "num_ev_attr": ["amount"]})Alternatively, an automatic feature representation can be generated:
from pm4py.algo.transformation.trace_encodings import algorithm as trace_encodings
from typing import Any
data: Any
feature_names: list[str]
data, feature_names = trace_encodings.apply(log)Then, the target classes are formed:
from pm4py.objects.log.util import get_class_representation
from typing import Any
if __name__ == "__main__":
target: Any
classes: list[str]
target, classes = get_class_representation.get_class_representation_by_trace_duration(pm4py.convert_to_event_log(log), 2 * 8640000)The decision tree is then built and visualized:
from sklearn import tree
from graphviz import Graph
from sklearn.tree import DecisionTreeClassifier
if __name__ == "__main__":
clf: DecisionTreeClassifier = tree.DecisionTreeClassifier()
clf.fit(data, target)
from pm4py.visualization.decisiontree import visualizer as dectree_visualizer
gviz: Graph = dectree_visualizer.apply(clf, feature_names, classes)Decision Mining enables the following, given:
It retrieves the features of the cases that take different paths. This allows, for example, building a decision tree to explain the choices made.
First, import a XES log:
import pm4py
import pandas
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes("tests/input_data/running-example.xes")Next, calculate a model using the Inductive Miner:
from pm4py.objects.petri_net.obj import Marking, PetriNet
if __name__ == "__main__":
net: PetriNet
im: Marking
fm: Marking
net, im, fm = pm4py.discover_petri_net_inductive(log)To visualize the model:
from pm4py.visualization.petri_net import visualizer
from graphviz import Graph
if __name__ == "__main__":
gviz: Graph = visualizer.apply(net, im, fm, parameters={visualizer.Variants.WO_DECORATION.value.Parameters.DEBUG: True})
visualizer.view(gviz) For this example, we select decision point p_10, where a choice is made between the activities examine casually and examine thoroughly. Once we have a log, a model, and a decision point, the decision mining algorithm can be executed:
from pm4py.algo.decision_mining import algorithm as decision_mining
from typing import Any
if __name__ == "__main__":
X: Any
y: Any
class_names: list[str]
X, y, class_names = decision_mining.apply(log, net, im, fm, decision_point="p_10")The outputs of the apply method are:
X: A Pandas DataFrame containing the features associated with each case leading to a decision.y: A Pandas Series containing the class (output) of each decision (e.g., 0 or 1).class_names: The names of the possible decision outcomes (e.g., examine casually and examine thoroughly).These outputs can be used with any classification or comparison technique. In particular, decision trees are a useful choice. We provide a function to automatically discover decision trees from decision mining results:
from pm4py.algo.decision_mining import algorithm as decision_mining
from sklearn.tree import DecisionTreeClassifier
if __name__ == "__main__":
clf: DecisionTreeClassifier
feature_names: list[str]
classes: list[str]
clf, feature_names, classes = decision_mining.get_decision_tree(log, net, im, fm, decision_point="p_10")To visualize the resulting decision tree:
from pm4py.visualization.decisiontree import visualizer as tree_visualizer
from graphviz import Graph
if __name__ == "__main__":
gviz: Graph = tree_visualizer.apply(clf, feature_names, classes)While the feature extraction described above is generic, it might not be optimal (performance-wise) when working directly with Pandas DataFrames. We also offer the option to extract a feature table by providing:
The output is another DataFrame containing:
Here is an example that keeps concept:name (activity) and amount (cost) as features:
import pm4py
import pandas as pd
from pm4py.objects.log.util import dataframe_utils
import pandas
if __name__ == "__main__":
dataframe: pandas.DataFrame = pd.read_csv("tests/input_data/roadtraffic100traces.csv")
dataframe = pm4py.format_dataframe(dataframe)
feature_table: pandas.DataFrame = dataframe_utils.get_features_df(dataframe, ["concept:name", "amount"])The resulting feature table will contain columns such as:
['case:concept:name', 'concept:name_CreateFine', 'concept:name_SendFine', 'concept:name_InsertFineNotification', 'concept:name_Addpenalty', 'concept:name_SendforCreditCollection', 'concept:name_Payment', 'concept:name_InsertDateAppealtoPrefecture', 'concept:name_SendAppealtoPrefecture', 'concept:name_ReceiveResultAppealfromPrefecture', 'concept:name_NotifyResultAppealtoOffender', 'amount']
Given a Petri net discovered by a classical process mining algorithm (e.g., Alpha Miner or Inductive Miner), we can enhance it into a Data Petri Net by applying decision mining at every decision point, and transforming the resulting decision trees into guards (boolean conditions).
An example:
import pm4py
import pandas
from pm4py.objects.petri_net.obj import Marking, PetriNet
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes("tests/input_data/roadtraffic100traces.xes")
net: PetriNet
im: Marking
fm: Marking
net, im, fm = pm4py.discover_petri_net_inductive(log)
from pm4py.algo.decision_mining import algorithm as decision_mining
net, im, fm = decision_mining.create_data_petri_nets_with_decisions(log, net, im, fm)The guards discovered for each transition can be printed. They are expressed as boolean conditions and interpreted by the execution engine:
if __name__ == "__main__":
for t in net.transitions:
if "guard" in t.properties:
print("")
print(t)
print(t.properties["guard"])The PM4Py library provides a method to extract temporal features from an event log, event stream, or Pandas DataFrame, as described in the paper by Pourbafrani et al. (2020). This method groups events by a specified time granularity (e.g., weekly) and computes aggregated metrics to represent process behavior over time.
To apply temporal feature extraction, you can use the apply function from the provided code. Here's an example:
import pm4py
from pm4py.algo.transformation.trace_encodings.variants import temporal as temporal_features
import pandas
# Import event data as a DataFrame
log: pandas.DataFrame = pm4py.read_xes("log.xes")
# Apply temporal feature extraction
features_df: pandas.DataFrame = temporal_features.apply(log, parameters={
"grouper_freq": "W",
"arrival_rate": "arrival_rate",
"finish_rate": "finish_rate",
"service_time": "service_time",
"waiting_time": "waiting_time",
"sojourn_time": "sojourn_time"
}) The resulting features_df is a Pandas DataFrame containing temporal features grouped by the specified frequency (e.g., weekly).
The algorithm accepts several parameters to customize the feature extraction process. These are defined in the Parameters enum and include:
| Parameter | Description |
|---|---|
GROUPER_FREQ | Time interval for grouping events (e.g., "W" for weekly, "D" for daily). Default: "W". |
ARRIVAL_RATE | Column name for the arrival rate of cases. Default: "arrival_rate". |
FINISH_RATE | Column name for the completion rate of cases. Default: "finish_rate". |
SERVICE_TIME | Column name for the service time (time spent on activities). Default: "service_time". |
WAITING_TIME | Column name for the waiting time (time spent idle). Default: "waiting_time". |
SOJOURN_TIME | Column name for the sojourn time (total time from start to end). Default: "sojourn_time". |
CASE_ID_COLUMN | Column name for case IDs. Default: "case:concept:name". |
ACTIVITY_COLUMN | Column name for activities. Default: "concept:name". |
TIMESTAMP_COLUMN | Column name for event timestamps. Default: "time:timestamp". |
START_TIMESTAMP_COLUMN | Column name for start timestamps (if available). Defaults to the timestamp column. |
RESOURCE_COLUMN | Column name for resources. Default: "org:resource". |
These parameters allow users to specify the granularity and naming conventions for the extracted features, tailoring the output to their specific needs.
The temporal feature extraction process involves the following steps:
The resulting DataFrame provides a tabular representation of temporal process characteristics, suitable for further analysis or visualization.
The output DataFrame might look like this:
import pandas
from typing import Any
# Example output DataFrame
import pandas as pd
data: list[dict[str, Any]] = [
{"timestamp": "2023-01-01", "unique_resources": 5, "unique_cases": 10, "unique_activities": 3, "num_events": 50,
"average_arrival_rate": 2.5, "average_finish_rate": 2.3, "average_service_time": 3600,
"average_waiting_time": 7200, "average_sojourn_time": 10800},
{"timestamp": "2023-01-08", "unique_resources": 4, "unique_cases": 8, "unique_activities": 4, "num_events": 45,
"average_arrival_rate": 2.0, "average_finish_rate": 1.9, "average_service_time": 3400,
"average_waiting_time": 7000, "average_sojourn_time": 10400}
]
features_df: pandas.DataFrame = pd.DataFrame(data)
print(features_df)This DataFrame shows temporal features for two weekly groups, including counts of unique elements and averages of time-based metrics.
Temporal feature extraction is particularly useful for:
To customize the feature extraction further, you can modify the parameters. For example, to group by days and use custom column names:
from pm4py.algo.transformation.trace_encodings.variants import temporal as temporal_features
import pandas
features_df: pandas.DataFrame = temporal_features.apply(log, parameters={
"grouper_freq": "D",
"arrival_rate": "case_arrival",
"finish_rate": "case_completion",
"service_time": "activity_duration",
"waiting_time": "idle_time",
"sojourn_time": "total_duration",
"pm4py:param:case_id_key": "case:concept:name",
"pm4py:param:timestamp_key": "time:timestamp",
"pm4py:param:start_timestamp_key": "time:timestamp",
"pm4py:param:resource_key": "org:resource",
"pm4py:param:activity_key": "concept:name"
})This allows the algorithm to adapt to different log formats and analysis requirements.
The approach is based on the following paper: