Other Analyses

Prepare the input with event-data import. Compare periods using process statistics and visualize discovered changes with directly-follows graphs. Review configuration parameters for runtime settings.

This section collects additional analysis techniques available in PM4Py that complement the core process analysis modules.

Concept Drift

Detects sudden changes in process behavior over time by splitting the log into sub-logs, extracting global control-flow features, and applying permutation tests over sliding windows (based on Bose et al., CAiSE 2011).

Returns: a list of cumulative sub-logs up to each detected change point (plus final segment), the corresponding change timestamps (based on case start times), and p-values indicating significance.

Key parameters (in pm4py.algo.concept_drift.algorithm.Parameters):

  • SUB_LOG_SIZE: traces per sub-log (default: 50)
  • WINDOW_SIZE: sub-logs per window for comparison (default: 8)
  • NUM_PERMUTATIONS: permutations for the statistical test (default: 100)
  • THRESH_P_VALUE: p-value threshold for drift (default: 0.5)
  • MAX_NO_CHANGE_POINTS: maximum number of change points (default: 5)
  • ACTIVITY_KEY, TIMESTAMP_KEY, CASE_ID_KEY: attribute keys
import pm4py, os
from pm4py.algo.concept_drift.variants import bose as concept_drift_detection
import pandas


def execute_script():
    log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "receipt.xes"))

    # Bose's concept drift detection (control-flow based)
    returned_sublogs: list[pandas.DataFrame]
    change_timestamps: list
    p_values: list[float]
    returned_sublogs, change_timestamps, p_values = concept_drift_detection.apply(
        log,
        parameters={
            concept_drift_detection.Parameters.MAX_NO_CHANGE_POINTS: 3,
            concept_drift_detection.Parameters.SUB_LOG_SIZE: 50,
            concept_drift_detection.Parameters.WINDOW_SIZE: 8,
            concept_drift_detection.Parameters.NUM_PERMUTATIONS: 100,
            concept_drift_detection.Parameters.THRESH_P_VALUE: 0.5,
        }
    )

    print("change_timestamps:", change_timestamps)
    print("p_values:", p_values)

    # Example: inspect each returned sub-log
    for sl in returned_sublogs:
        dfg: dict
        sa: dict
        ea: dict
        dfg, sa, ea = pm4py.discover_dfg(sl)
        pm4py.view_dfg(dfg, sa, ea, format="svg")


if __name__ == "__main__":
    execute_script()

PRIPEL Anonymization

The low-level PRIPEL example explicitly uses legacy traces for case slicing and anonymization. The public differential-privacy wrapper below accepts a DataFrame directly.

PM4Py also supports differential-privacy-aware publication of event logs. In this workflow, control-flow information is anonymized through a trace-variant query and contextual information such as timestamps and event attributes is anonymized through PRIPEL.

The low-level example below mirrors the PM4Py branch example: first a SaCoFa trace-variant query is computed, then the resulting anonymized control-flow information is passed to PRIPEL to obtain an anonymized event log. The example concludes by discovering and visualizing a DFG before and after anonymization.

This functionality requires diffprivlib to be installed in the Python environment.

import pm4py
from pm4py.algo.anonymization.trace_variant_query import algorithm as trace_variant_query
from pm4py.algo.anonymization.pripel import algorithm as pripel
from pm4py.objects.log.obj import EventLog

if __name__ == "__main__":
    log: EventLog = pm4py.read_xes(
        "tests/input_data/receipt.xes",
        return_legacy_log_object=True
    )
    log = EventLog(log[0:100])

    epsilon: float = 0.5
    sacofa_result: dict = trace_variant_query.apply(
        log=log,
        variant=trace_variant_query.Variants.SACOFA,
        parameters={"epsilon": epsilon, "k": 30, "p": 4}
    )
    anonymized_log: EventLog = pripel.apply(
        log=log,
        trace_variant_query=sacofa_result,
        epsilon=epsilon
    )

    original_dfg: dict
    original_sa: dict
    original_ea: dict
    original_dfg, original_sa, original_ea = pm4py.discover_dfg(log)
    pm4py.view_dfg(original_dfg, original_sa, original_ea, format="svg")

    anonymized_dfg: dict
    anonymized_sa: dict
    anonymized_ea: dict
    anonymized_dfg, anonymized_sa, anonymized_ea = pm4py.discover_dfg(anonymized_log)
    pm4py.view_dfg(anonymized_dfg, anonymized_sa, anonymized_ea, format="svg")

If you prefer the convenience wrapper exposed by PM4Py, the same workflow can also be started through pm4py.anonymize_differential_privacy:

import pm4py
import pandas

if __name__ == "__main__":
    event_log: pandas.DataFrame = pm4py.read_xes("tests/input_data/receipt.xes")
    anonymized_event_log: pandas.DataFrame = pm4py.anonymize_differential_privacy(
        event_log,
        epsilon=1.0,
        k=10,
        p=20
    )