Handling Event Data

Once your data is ready, filter cases and events or discover a process model. For data involving several object types, see object-centric event data.

Importing XES

The IEEE XES standard defines the format for storing event logs. For more information about the format, please visit the IEEE XES website. A simple synthetic event log file (running-example.xes) can be downloaded here. Additionally, several real event logs have been made available over the past few years, which you can find here.

The example code demonstrates how to import an event log stored in the IEEE XES format, given the file path to the log file. pm4py.read_xes returns a Pandas DataFrame by default. Download the sample file into your working directory before running the example.

import pm4py
import pandas

if __name__ == "__main__":
    log: pandas.DataFrame = pm4py.read_xes('running-example.xes')
    dfg, start_activities, end_activities = pm4py.discover_dfg(log)
    pm4py.view_dfg(dfg, start_activities, end_activities)

Importing CSV

Read CSV event data into a Pandas DataFrame and prepare it with pm4py.format_dataframe. This identifies the case, activity, and timestamp columns, converts their types, and orders the events. Pass the resulting DataFrame directly to PM4Py discovery, filtering, and conformance functions.

Download running-example.csv into your working directory. This file uses a comma separator; adjust sep for your own data.

import pandas as pd
import pm4py

if __name__ == "__main__":
    dataframe: pd.DataFrame = pd.read_csv('running-example.csv', sep=',')
    dataframe = pm4py.format_dataframe(
        dataframe, case_id='case:concept:name',
        activity_key='concept:name', timestamp_key='time:timestamp'
    )
    dfg, start_activities, end_activities = pm4py.discover_dfg(dataframe)
    pm4py.view_dfg(dfg, start_activities, end_activities)

For CSV files with different column names or timestamp formats, pass those details explicitly. Save the following data as running-example-transformed.csv, using the column headers shown and a comma separator.

CaseIDActivityTimestampclientID
1register request20200422T04551337
2register request20200422T04571479
1submit payment20200422T05031337

In this small example table, we observe four columns: CaseID, Activity, Timestamp, and clientID. The CaseID column identifies which events belong to the same case. The events remain rows in the DataFrame after formatting.

Another interesting aspect of the example data is the fourth column, clientID. This column represents a case-level attribute, meaning that the value remains constant throughout the execution of a process instance. PM4Py allows us to specify that a column describes a case-level attribute, under the assumption that the attribute does not change during the process execution.

The example code shows how to prepare the CSV data above. After loading the CSV file, we rename the clientID column to case:clientID using a specific operation provided by Pandas.

import pandas as pd
import pm4py

if __name__ == "__main__":
    dataframe: pd.DataFrame = pd.read_csv('running-example-transformed.csv', sep=',')
    dataframe = dataframe.rename(columns={'clientID': 'case:clientID'})
    dataframe = pm4py.format_dataframe(
        dataframe, case_id='CaseID', activity_key='Activity',
        timestamp_key='Timestamp', timest_format='%Y%m%dT%H%M'
    )
    print(pm4py.get_start_activities(dataframe))

Converting Event Data Formats

Use DataFrames for the standard workflows on these pages. EventLog and EventStream are legacy representations that are still needed by some specialized APIs. The snippets below use the formatted dataframe from the CSV example.

Convert to an EventLog only when an API requires a sequence of traces:

import pm4py
from pm4py.objects.log.obj import EventLog

if __name__ == "__main__":
    # Only needed when an API explicitly requires the legacy EventLog format.
    event_log: EventLog = pm4py.convert_to_event_log(dataframe)

For a legacy API requiring an EventStream, convert explicitly:

import pm4py
from pm4py.objects.log.obj import EventStream

if __name__ == "__main__":
    event_stream: EventStream = pm4py.convert_to_event_stream(dataframe)

To return from a legacy event log to a DataFrame:

import pm4py
import pandas

if __name__ == "__main__":
    dataframe: pandas.DataFrame = pm4py.convert_to_dataframe(event_log)

Exporting Event Logs as XES

Pass the formatted DataFrame directly to pm4py.write_xes. The writer handles the XES serialization, so no explicit legacy conversion is needed.

import pm4py
  
if __name__ == "__main__":
    pm4py.write_xes(dataframe, 'exported.xes')

Exporting Event Logs as CSV

Write the DataFrame with Pandas to_csv. Set index=False to avoid exporting the DataFrame index as an extra column. Use Pandas options such as sep and encoding to control the output format.

# Use the DataFrame prepared in the import examples above.
dataframe.to_csv('exported.csv', index=False)

Changing Event Time Granularity

Use Pandas datetime operations to change timestamp precision directly on a DataFrame. This example copies the data and rounds time:timestamp down to the nearest minute with dt.floor('min'). Use dt.ceil or dt.round when needed; these operations support fixed frequencies such as seconds, minutes, hours, and days.

import pm4py

if __name__ == "__main__":
    dataframe = pm4py.read_xes('running-example.xes')
    rounded_dataframe = dataframe.copy()
    rounded_dataframe['time:timestamp'] = dataframe['time:timestamp'].dt.floor('min')