pm4py.algo.transformation.trace_encodings.variants.events_transformers module#
PM4Py – A Process Mining Library for Python Copyright (C) 2026 Process Intelligence Solutions GmbH
This program is free software: you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation, either version 3 of the License, or any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License along with this program. If not, see this software project’s root or visit <https://www.gnu.org/licenses/>.
Website: https://processintelligence.solutions Contact: info@processintelligence.solutions
- class pm4py.algo.transformation.trace_encodings.variants.events_transformers.Parameters(*values)[source]#
Bases:
Enum- EVENT_ID_KEY = 'event_id_key'#
- CASE_ID_KEY = 'pm4py:param:case_id_key'#
- EMBEDDING_MODEL = 'embedding_model'#
- ATTRIBUTE_KEY = 'pm4py:param:attribute_key'#
- EVENT_ATTRIBUTES = 'event_attributes'#
- KEEP_CASES = 'keep_cases'#
- pm4py.algo.transformation.trace_encodings.variants.events_transformers.apply(log: DataFrame, parameters: Dict[Any, Any] | None = None) Tuple[List[str], List[List[float]]][source]#
Computes one text embedding per event.
From an event-log point of view, each event is converted to a short sentence and then embedded by a sentence-transformers model. By default, the sentence is the activity name. Additional event attributes can be included to represent a richer event perspective.
- Example with activity-only encoding:
event {concept:name: “A”} -> “A”
- Example with event_attributes=[“concept:name”, “org:resource”]:
event {concept:name: “A”, org:resource: “R1”} -> “concept:name=A|org:resource=R1”
The output is a list of event identifiers and a list of dense embedding vectors, one vector per event.
- Parameters:
log – Pandas dataframe
Parameters – Parameters of the algorithm, including: - Parameters.EVENT_ID_KEY => an attribute unique per event - Parameters.EMBEDDING_MODEL => the embedding to be used (default: all-MiniLM-L6-v2) - Parameters.ATTRIBUTE_KEY => the attribute to be used - Parameters.EVENT_ATTRIBUTES => event attributes to include in the event sentence. If omitted, ATTRIBUTE_KEY is used.
- Returns:
event_identifiers – The list of all the event identifiers
embeddings_list – The list of embeddings for the considered events
- pm4py.algo.transformation.trace_encodings.variants.events_transformers.keep_top_k_per_similarity(log: DataFrame, target_sentence: str, k: int, event_identifiers: List[str] | None = None, embeddings_list: List[List[float]] | None = None, parameters: Dict[Any, Any] | None = None) DataFrame[source]#
Keeps the top K events by embedding similarity with the given sentence.
For example, after encoding events as activity/resource sentences, the query “pay compensation” is embedded with the same model and compared with each event embedding using cosine similarity.
- Parameters:
log – Pandas dataframe
target_sentence – Target sentence
k – Number of similar events to retain
event_identifiers – (Optional) the list of event identifiers in the log, as returned by the ‘apply’ method
embeddings_list – (Optional) the list of embeddings for such events, as returned by the ‘apply’ method
parameters – Other parameters of the method: - Parameters.KEEP_CASES => If True, keep the cases containing such events, instead of the events themselves (default: False)
- Returns:
Event log filtered on the top K events (or cases) according to the similarity metric.
- Return type:
filtered_log