Filtering Event Data

These snippets use a Pandas DataFrame named log; follow the event-data import examples to create it. After filtering, compare process statistics or run process discovery.

PM4Py offers a variety of specific methods to filter event logs.

Filtering By Timeframe

The following section presents various methods for filtering event logs based on time frames. Each example operates directly on a Pandas DataFrame. You might be interested in retaining only the traces that fall within a specific time interval, such as from March 9, 2011, to January 18, 2012.

import pm4py
import pandas

if __name__ == "__main__":
    filtered_log: pandas.DataFrame = pm4py.filter_time_range(log, "2011-03-09 00:00:00", "2012-01-18 23:59:59", mode='traces_contained')

Additionally, it is possible to keep the traces that intersect with a time interval.

import pm4py
import pandas

if __name__ == "__main__":
    filtered_log: pandas.DataFrame = pm4py.filter_time_range(log, "2011-03-09 00:00:00", "2012-01-18 23:59:59", mode='traces_intersecting')

So far, only trace-based filtering techniques have been discussed. However, there is also a method to keep the events that fall within a specific timeframe.

import pm4py
import pandas

if __name__ == "__main__":
    filtered_log: pandas.DataFrame = pm4py.filter_time_range(log, "2011-03-09 00:00:00", "2012-01-18 23:59:59", mode='events')

Filtering By Case Performance

This filter allows you to keep only traces with durations that fall within a specified interval. In the examples, traces between 1 and 10 days are kept. Note that the time parameters are given in seconds.

import pm4py
import pandas

if __name__ == "__main__":
    filtered_log: pandas.DataFrame = pm4py.filter_case_performance(log, 86400, 864000)

Filtering By Start Activities

In general, PM4Py can filter a log or a DataFrame based on start activities. First, you may need to identify the starting activities, for which code snippets are provided. The snippet discovers start activities and filters the DataFrame.

log_start is a dictionary where the key is the activity and the value is the number of occurrences.

import pm4py
import pandas

if __name__ == "__main__":
    log_start: dict[str, int] = pm4py.get_start_activities(log)
    filtered_log: pandas.DataFrame = pm4py.filter_start_activities(log, ["S1"]) #suppose "S1" is the start activity you want to filter on

Filtering By End Activities

PM4Py also allows filtering by end activities. This filter keeps only traces that end with a specified set of activities. First, you may need to identify the end activities, for which a code snippet is provided.

import pm4py
import pandas

if __name__ == "__main__":
    end_activities: dict[str, int] = pm4py.get_end_activities(log)
    filtered_log: pandas.DataFrame = pm4py.filter_end_activities(log, ["pay compensation"])

Filtering By Variants

A variant refers to a set of cases that share the same control-flow perspective, meaning a set of cases that follow the same sequence of activities in the same order. Variants are represented as tuples of activity names; each tuple describes an ordered sequence. To retrieve the variants from the log, you can use the following code snippet:


import pm4py

if __name__ == "__main__":
    variants: dict = pm4py.get_variants(log)

To filter by a specific collection of variants, use the following code snippet:

import pm4py
import pandas

if __name__ == "__main__":
    filtered_log: pandas.DataFrame = pm4py.filter_variants(log, [("A", "B", "C", "D"), ("A", "E", "F", "G"), ("A", "C", "D")])

Other variant-based filters are available. For example, filters on the top-k variants retain only the cases following one of the k most frequent variants:

import pm4py
import pandas

if __name__ == "__main__":
    log: pandas.DataFrame = pm4py.read_xes("tests/input_data/receipt.xes")
    k: int = 2
    filtered_log: pandas.DataFrame = pm4py.filter_variants_top_k(log, k)

The variant coverage filter retains only those traces that follow the top variants in the log, under the condition that each variant covers a specified percentage of cases. For instance, if min_coverage_percentage=0.4 and we have a log with 1000 cases, where 500 are variant 1, 400 are variant 2, and 100 are variant 3, the filter will keep only traces for variants 1 and 2.

import pm4py
import pandas

if __name__ == "__main__":
    log: pandas.DataFrame = pm4py.read_xes("tests/input_data/receipt.xes")
    perc: float = 0.1
    filtered_log: pandas.DataFrame = pm4py.filter_variants_by_coverage_percentage(log, perc)

Filtering By Attribute Values

Filtering by attribute values allows you to:

  • Keep cases that contain at least one event with one of the given attribute values,
  • Remove cases that contain an event with one of the given attribute values,
  • Keep events (trimming traces) that have one of the given attribute values,
  • Remove events (trimming traces) that have one of the given attribute values.

Examples of attributes include the resource (typically contained in the org:resource attribute) and the activity (usually found in the concept:name attribute). To get the list of resources and activities contained in the log, you can use the following code.

import pm4py

if __name__ == "__main__":
    activities: dict[str, int] = pm4py.get_event_attribute_values(log, "concept:name")
    resources: dict[str, int] = pm4py.get_event_attribute_values(log, "org:resource")

To filter traces containing or not containing a given list of resources, use the following code:

import pm4py
import pandas

if __name__ == "__main__":
    tracefilter_log_pos: pandas.DataFrame = pm4py.filter_event_attribute_values(log, "org:resource", ["Resource10"], level="case", retain=True)
    tracefilter_log_neg: pandas.DataFrame = pm4py.filter_event_attribute_values(log, "org:resource", ["Resource10"], level="case", retain=False)

You can also keep only the events performed by a given list of resources, trimming the cases accordingly. Use the following code for this:

import pm4py
import pandas

if __name__ == "__main__":
    tracefilter_log_pos: pandas.DataFrame = pm4py.filter_event_attribute_values(log, "org:resource", ["Resource10"], level="event", retain=True)
    tracefilter_log_neg: pandas.DataFrame = pm4py.filter_event_attribute_values(log, "org:resource", ["Resource10"], level="event", retain=False)

Filtering By Numeric Attribute Values

Filtering by numeric attribute values offers options similar to filtering by string attributes. First, we import the log, then filter to keep only events satisfying a numeric range between 34 and 36. An additional filter can be applied to retain only cases with at least one event within the specified range. If you're interested in cases with a specific activity, such as "Add penalty," having an amount between 34 and 500, the following code snippet can be used:

import os
import pandas as pd
import pm4py
import pandas

if __name__ == "__main__":
    df: pandas.DataFrame = pd.read_csv(os.path.join("tests", "input_data", "roadtraffic100traces.csv"))
    df = pm4py.format_dataframe(df)

    from pm4py.algo.filtering.pandas.attributes import attributes_filter
    filtered_df_events: pandas.DataFrame = attributes_filter.apply_numeric_events(df, 34, 36,
        parameters={attributes_filter.Parameters.CASE_ID_KEY: "case:concept:name", attributes_filter.Parameters.ATTRIBUTE_KEY: "amount"})

    filtered_df_cases: pandas.DataFrame = attributes_filter.apply_numeric(df, 34, 36,
        parameters={attributes_filter.Parameters.CASE_ID_KEY: "case:concept:name", attributes_filter.Parameters.ATTRIBUTE_KEY: "amount"})

    filtered_df_cases = attributes_filter.apply_numeric(df, 34, 500,
        parameters={attributes_filter.Parameters.CASE_ID_KEY: "case:concept:name", attributes_filter.Parameters.ATTRIBUTE_KEY: "amount",
                    attributes_filter.Parameters.STREAM_FILTER_KEY1: "concept:name",
                    attributes_filter.Parameters.STREAM_FILTER_VALUE1: "Add penalty"})

Between Filter

The between filter transforms the event log by identifying subcases that span from a source activity to a target activity. This is useful for analyzing behavior between two activities, such as throughput time, activity inclusion, or conformance levels. The filter between two activities is applied as follows:

import pm4py
import pandas

if __name__ == "__main__":
    log: pandas.DataFrame = pm4py.read_xes("tests/input_data/running-example.xes")

    filtered_log: pandas.DataFrame = pm4py.filter_between(log, "check ticket", "decide")

Case Size Filter

The case size filter retains only the cases in the log that have a number of events within a user-specified range. This filter can be used to eliminate cases that are too short (possibly incomplete or outliers) or too long (indicating excessive rework). The case size filter can be applied as follows:

import pm4py
import pandas

if __name__ == "__main__":
    log: pandas.DataFrame = pm4py.read_xes("tests/input_data/running-example.xes")

    filtered_log: pandas.DataFrame = pm4py.filter_case_size(log, 5, 10)

Rework Filter

The rework filter identifies cases where a given activity has been repeated. For instance, it can be used to search for cases that contain at least two occurrences of the activity "reinitiate request."

import pm4py
import pandas

if __name__ == "__main__":
    log: pandas.DataFrame = pm4py.read_xes("tests/input_data/running-example.xes")

    filtered_log: pandas.DataFrame = pm4py.filter_activities_rework(log, "reinitiate request", 2)

Path Performance Filter

The path performance filter identifies cases in which the duration of a given path between two activities falls within a specified range. This is useful for identifying cases where a significant amount of time has passed between two activities. The filter is applied as follows:

import pm4py
import pandas

if __name__ == "__main__":
    log: pandas.DataFrame = pm4py.read_xes("tests/input_data/running-example.xes")

    filtered_log: pandas.DataFrame = pm4py.filter_paths_performance(log, ("decide", "pay compensation"), 2*86400, 10*86400)