Once your data is ready, filter cases and events or discover a process model. For data involving several object types, see object-centric event data.
The IEEE XES standard defines the format for storing event logs. For more information about the format, please visit the IEEE XES website. A simple synthetic event log file (running-example.xes) can be downloaded here. Additionally, several real event logs have been made available over the past few years, which you can find here.
The example code demonstrates how to import an event log stored in the IEEE XES format, given the file path to the log file. pm4py.read_xes returns a Pandas DataFrame by default. Download the sample file into your working directory before running the example.
import pm4py
import pandas
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes('running-example.xes')
dfg, start_activities, end_activities = pm4py.discover_dfg(log)
pm4py.view_dfg(dfg, start_activities, end_activities)Read CSV event data into a Pandas DataFrame and prepare it with pm4py.format_dataframe. This identifies the case, activity, and timestamp columns, converts their types, and orders the events. Pass the resulting DataFrame directly to PM4Py discovery, filtering, and conformance functions.
Download running-example.csv into your working directory. This file uses a comma separator; adjust sep for your own data.
import pandas as pd
import pm4py
if __name__ == "__main__":
dataframe: pd.DataFrame = pd.read_csv('running-example.csv', sep=',')
dataframe = pm4py.format_dataframe(
dataframe, case_id='case:concept:name',
activity_key='concept:name', timestamp_key='time:timestamp'
)
dfg, start_activities, end_activities = pm4py.discover_dfg(dataframe)
pm4py.view_dfg(dfg, start_activities, end_activities)For CSV files with different column names or timestamp formats, pass those details explicitly. Save the following data as running-example-transformed.csv, using the column headers shown and a comma separator.
| CaseID | Activity | Timestamp | clientID |
|---|---|---|---|
| 1 | register request | 20200422T0455 | 1337 |
| 2 | register request | 20200422T0457 | 1479 |
| 1 | submit payment | 20200422T0503 | 1337 |
In this small example table, we observe four columns: CaseID, Activity, Timestamp, and clientID. The CaseID column identifies which events belong to the same case. The events remain rows in the DataFrame after formatting.
Another interesting aspect of the example data is the fourth column, clientID. This column represents a case-level attribute, meaning that the value remains constant throughout the execution of a process instance. PM4Py allows us to specify that a column describes a case-level attribute, under the assumption that the attribute does not change during the process execution.
The example code shows how to prepare the CSV data above. After loading the CSV file, we rename the clientID column to case:clientID using a specific operation provided by Pandas.
import pandas as pd
import pm4py
if __name__ == "__main__":
dataframe: pd.DataFrame = pd.read_csv('running-example-transformed.csv', sep=',')
dataframe = dataframe.rename(columns={'clientID': 'case:clientID'})
dataframe = pm4py.format_dataframe(
dataframe, case_id='CaseID', activity_key='Activity',
timestamp_key='Timestamp', timest_format='%Y%m%dT%H%M'
)
print(pm4py.get_start_activities(dataframe))Use DataFrames for the standard workflows on these pages. EventLog and EventStream are legacy representations that are still needed by some specialized APIs. The snippets below use the formatted dataframe from the CSV example.
Convert to an EventLog only when an API requires a sequence of traces:
import pm4py
from pm4py.objects.log.obj import EventLog
if __name__ == "__main__":
# Only needed when an API explicitly requires the legacy EventLog format.
event_log: EventLog = pm4py.convert_to_event_log(dataframe) For a legacy API requiring an EventStream, convert explicitly:
import pm4py
from pm4py.objects.log.obj import EventStream
if __name__ == "__main__":
event_stream: EventStream = pm4py.convert_to_event_stream(dataframe)To return from a legacy event log to a DataFrame:
import pm4py
import pandas
if __name__ == "__main__":
dataframe: pandas.DataFrame = pm4py.convert_to_dataframe(event_log)Pass the formatted DataFrame directly to pm4py.write_xes. The writer handles the XES serialization, so no explicit legacy conversion is needed.
import pm4py
if __name__ == "__main__":
pm4py.write_xes(dataframe, 'exported.xes')Write the DataFrame with Pandas to_csv. Set index=False to avoid exporting the DataFrame index as an extra column. Use Pandas options such as sep and encoding to control the output format.
# Use the DataFrame prepared in the import examples above.
dataframe.to_csv('exported.csv', index=False)Use Pandas datetime operations to change timestamp precision directly on a DataFrame. This example copies the data and rounds time:timestamp down to the nearest minute with dt.floor('min'). Use dt.ceil or dt.round when needed; these operations support fixed frequencies such as seconds, minutes, hours, and days.
import pm4py
if __name__ == "__main__":
dataframe = pm4py.read_xes('running-example.xes')
rounded_dataframe = dataframe.copy()
rounded_dataframe['time:timestamp'] = dataframe['time:timestamp'].dt.floor('min')