pm4py.algo.transformation.trace_encodings.variants.events_transformers module#

PM4Py – A Process Mining Library for Python Copyright (C) 2026 Process Intelligence Solutions GmbH

This program is free software: you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation, either version 3 of the License, or any later version.

This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more details.

You should have received a copy of the GNU Affero General Public License along with this program. If not, see this software project’s root or visit <https://www.gnu.org/licenses/>.

Website: https://processintelligence.solutions Contact: info@processintelligence.solutions

class pm4py.algo.transformation.trace_encodings.variants.events_transformers.Parameters(*values)[source]#

Bases: Enum

EVENT_ID_KEY = 'event_id_key'#
CASE_ID_KEY = 'pm4py:param:case_id_key'#
EMBEDDING_MODEL = 'embedding_model'#
ATTRIBUTE_KEY = 'pm4py:param:attribute_key'#
EVENT_ATTRIBUTES = 'event_attributes'#
KEEP_CASES = 'keep_cases'#
pm4py.algo.transformation.trace_encodings.variants.events_transformers.apply(log: DataFrame, parameters: Dict[Any, Any] | None = None) → Tuple[List[str], List[List[float]]][source]#

Computes one text embedding per event.

From an event-log point of view, each event is converted to a short sentence and then embedded by a sentence-transformers model. By default, the sentence is the activity name. Additional event attributes can be included to represent a richer event perspective.

Example with activity-only encoding:

event {concept:name: “A”} -> “A”

Example with event_attributes=[“concept:name”, “org:resource”]:

event {concept:name: “A”, org:resource: “R1”} -> “concept:name=A|org:resource=R1”

The output is a list of event identifiers and a list of dense embedding vectors, one vector per event.

Parameters:
  • log – Pandas dataframe

  • Parameters – Parameters of the algorithm, including: - Parameters.EVENT_ID_KEY => an attribute unique per event - Parameters.EMBEDDING_MODEL => the embedding to be used (default: all-MiniLM-L6-v2) - Parameters.ATTRIBUTE_KEY => the attribute to be used - Parameters.EVENT_ATTRIBUTES => event attributes to include in the event sentence. If omitted, ATTRIBUTE_KEY is used.

Returns:

  • event_identifiers – The list of all the event identifiers

  • embeddings_list – The list of embeddings for the considered events

pm4py.algo.transformation.trace_encodings.variants.events_transformers.keep_top_k_per_similarity(log: DataFrame, target_sentence: str, k: int, event_identifiers: List[str] | None = None, embeddings_list: List[List[float]] | None = None, parameters: Dict[Any, Any] | None = None) → DataFrame[source]#

Keeps the top K events by embedding similarity with the given sentence.

For example, after encoding events as activity/resource sentences, the query “pay compensation” is embedded with the same model and compared with each event embedding using cosine similarity.

Parameters:
  • log – Pandas dataframe

  • target_sentence – Target sentence

  • k – Number of similar events to retain

  • event_identifiers – (Optional) the list of event identifiers in the log, as returned by the ‘apply’ method

  • embeddings_list – (Optional) the list of embeddings for such events, as returned by the ‘apply’ method

  • parameters – Other parameters of the method: - Parameters.KEEP_CASES => If True, keep the cases containing such events, instead of the events themselves (default: False)

Returns:

Event log filtered on the top K events (or cases) according to the similarity metric.

Return type:

filtered_log