Prepare a Pandas DataFrame named log using the import examples, including an org:resource column. Combine these results with performance statistics or resource-based features.
In PM4Py, we provide support for various Social Network Analysis (SNA) metrics, as well as tools for the discovery of roles.
The Handover of Work metric measures how often one individual is followed by another individual in the execution of a business process. To calculate this metric, you can use the following code:
import pm4py
if __name__ == "__main__":
hw_values: dict = pm4py.discover_handover_of_work_network(log)You can then visualize the result using NetworkX or Pyvis:
import pm4py
if __name__ == "__main__":
pm4py.view_sna(hw_values)The Subcontracting metric calculates how often the work of one individual is interleaved with the work of another individual, only for it to eventually "return" to the original individual. To measure the subcontracting metric, you can use the following code:
import pm4py
if __name__ == "__main__":
sub_values: dict = pm4py.discover_subcontracting_network(log)Afterward, you can visualize the results using NetworkX or Pyvis:
import pm4py
if __name__ == "__main__":
pm4py.view_sna(sub_values)The Working Together metric calculates how often two individuals collaborate to resolve a process instance. To measure the Working Together metric, you can use the following code:
import pm4py
if __name__ == "__main__":
wt_values: dict = pm4py.discover_working_together_network(log)You can then visualize the results using NetworkX or Pyvis:
import pm4py
if __name__ == "__main__":
pm4py.view_sna(wt_values)The Similar Activities metric calculates how similar the work patterns are between two individuals. To measure the Similar Activities metric, you can use the following code:
import pm4py
if __name__ == "__main__":
ja_values: dict = pm4py.discover_activity_based_resource_similarity(log)You can then visualize the results using NetworkX or Pyvis:
import pm4py
if __name__ == "__main__":
pm4py.view_sna(ja_values)A role is defined as a set of activities in the log that are executed by a similar (multi)set of resources. Essentially, it represents a specific function within an organization. Grouping activities into roles can help:
Initially, each activity is considered a separate role, and it is associated with the multiset of its originators. Roles are then merged according to their similarity until no further merges are possible. To begin, you need to import a log:
import pm4py
import os
import pandas
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "receipt.xes"))Next, apply the role detection algorithm:
import pm4py
if __name__ == "__main__":
roles: list = pm4py.discover_organizational_roles(log)You can print the sets of activities grouped into roles by using the following code:
print([x[0] for x in roles])
After applying an SNA metric, clustering allows you to group resources connected by meaningful relationships within the given metric. For example:
We provide a method to generate a list of groups (where each group consists of a list of resources) from the results of an SNA metric. This can be applied as follows to the running-example log and the results of the "Similar Activities" metric:
import pm4py
import os
import pandas
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
sa_metric: dict = pm4py.discover_activity_based_resource_similarity(log)
from pm4py.algo.organizational_mining.sna import util
clustering: dict = util.cluster_affinity_propagation(sa_metric)Resource profiling in event logs is also possible. We implement the approach described in: Pika, Anastasiia, et al. "Mining resource profiles from event logs." ACM Transactions on Management Information Systems (TMIS) 8.1 (2017): 1-30. Essentially, the behavior of a resource can be measured over a period of time with various metrics described in the paper:
The following example calculates these metrics starting from the running-example XES event log:
import os
from pm4py.algo.organizational_mining.resource_profiles import algorithm
import pm4py
import pandas
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "running-example.xes"))
# Metric RBI 1.1: Number of distinct activities done by a resource in a given time interval [t1, t2)
print(algorithm.distinct_activities(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Sara"))
# Metric RBI 1.3: Fraction of completions of a given activity a by a given resource r,
# during a given time slot [t1, t2), with respect to the total number of activity completions by resource r
# during [t1, t2)
print(algorithm.activity_frequency(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Sara", "decide"))
# Metric RBI 2.1: The number of activity instances completed by a given resource during a given time slot.
print(algorithm.activity_completions(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Sara"))
# Metric RBI 2.2: The number of cases completed during a given time slot in which a given resource was involved.
print(algorithm.case_completions(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Pete"))
# Metric RBI 2.3: The fraction of cases completed during a given time slot in which a given resource was involved
# with respect to the total number of cases completed during the time slot.
print(algorithm.fraction_case_completions(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Pete"))
# Metric RBI 2.4: The average number of activities started by a given resource but not completed at a moment in time.
print(algorithm.average_workload(log, "2010-12-30 00:00:00", "2011-01-15 00:00:00", "Mike"))
# Metric RBI 3.1: The fraction of active time during which a given resource is involved in more than one activity
# with respect to the resource's active time.
print(algorithm.multitasking(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Mike"))
# Metric RBI 4.3: The average duration of instances of a given activity completed during a given time slot by
# a given resource.
print(algorithm.average_duration_activity(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Sue", "examine thoroughly"))
# Metric RBI 4.4: The average duration of cases completed during a given time slot in which a given resource was involved.
print(algorithm.average_case_duration(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Sue"))
# Metric RBI 5.1: The number of cases completed during a given time slot in which two given resources were involved.
print(algorithm.interaction_two_resources(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Mike", "Pete"))
# Metric RBI 5.2: The fraction of resources involved in the same cases with a given resource during a given time slot
# with respect to the total number of resources active during the time slot.
print(algorithm.social_position(log, "2010-12-30 00:00:00", "2011-01-25 00:00:00", "Sue"))With event logs, we can identify groups of resources performing similar activities. As we have seen in previous sections, there are different ways to automatically detect these groups:
Alternatively, an attribute might be present in the events, specifying the group that performed the task.
"Organizational mining" refers to the discovery of behavior-related information specific to an organizational group, such as identifying which activities are performed by the group.
We provide an implementation of the approach described in: Yang, Jing, et al. "OrgMining 2.0: A Novel Framework for Organizational Model Mining from Event Logs." arXiv preprint arXiv:2011.12445 (2020).
The approach provides descriptions of group-related metrics (local diagnostics), such as:
The following example calculates these metrics using the receipt XES event log and shows how the information can be used, leveraging an attribute that specifies which group is performing the task:
import pm4py
import os
from pm4py.algo.organizational_mining.local_diagnostics import algorithm as local_diagnostics
import pandas
if __name__ == "__main__":
log: pandas.DataFrame = pm4py.read_xes(os.path.join("tests", "input_data", "receipt.xes"))
# This applies the organizational mining from an attribute that is in each event, describing the group that is performing the task.
ld: dict = local_diagnostics.apply_from_group_attribute(log, parameters={local_diagnostics.Parameters.GROUP_KEY: "org:group"})
# GROUP RELATIVE FOCUS (on a given type of work) specifies how much a resource group performed this type of work
# compared to the overall workload of the group. It can be used to measure how the workload of a resource group
# is distributed over different types of work, i.e., work diversification of the group.
print("\ngroup_relative_focus")
print(ld["group_relative_focus"])
# GROUP RELATIVE STAKE (in a given type of work) specifies how much this type of work was performed by a certain
# resource group among all groups. It can be used to measure how the workload devoted to a certain type of work is
# distributed over resource groups in an organizational model, i.e., work participation by different groups.
print("\ngroup_relative_stake")
print(ld["group_relative_stake"])
# GROUP COVERAGE with respect to a given type of work specifies the proportion of members of a resource group that
# performed this type of work.
print("\ngroup_coverage")
print(ld["group_coverage"])
# GROUP MEMBER CONTRIBUTION of a member of a resource group with respect to a given type of work specifies how
# much of this type of work by the group was performed by the member. It can be used to measure how the workload
# of the entire group devoted to a certain type of work is distributed over the group members.
print("\ngroup_member_contribution")
print(ld["group_member_contribution"]) Alternatively, you can use the apply_from_clustering_or_roles method, which takes the log as the first argument and the results of the clustering as the second argument.