Submitted:
03 July 2026
Posted:
07 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- RQ1
- Can event logs, process trees, and Petri nets be embedded into a shared latent space in which the three views of the same process behavior receive nearby representations?
- RQ2
- Can a learned model translate event logs or Petri nets into valid process trees—well-formed, Petri-convertible models whose behavior is close to the source?
- RQ3
- How do the learned representations and the learned log-to-model translation compare with classical baselines: deterministic log and model features, Petri-net embeddings, and the Inductive Miner?
2. Related Work
- Process discovery and block-structured models.
- Model similarity, conformance, and behavioral distances.
- Representation learning for process-mining artifacts.
- Neural and LLM-based process-model generation.
- Multimodal learning and grammar-constrained decoding.
3. Preliminaries
- Event logs.
- Process trees.
- Petri nets.
4. The ProcRosetta Framework
4.1. Synthetic Paired Triples
- is sampled by a recursive generator that at each node emits an activity leaf (probability , always at maximum depth 3) or an operator from with arity 2 or 3; labels come from at most 30 activities (reuse probability ), and the tree is canonicalized (Section 3);
- is obtained by playing out with PM4Py’s process-tree simulation (16 traces per sample);
- is the typed-graph representation of the Petri net obtained by PM4Py’s deterministic conversion of .
4.2. Three Encoders, One Latent Space
4.3. Grammar-Masked Process-Tree Decoder
4.4. Training Objective
5. Assessment
5.1. Training Behavior
5.2. Decode Quality (RQ2)
5.3. Process-Discovery Quality (RQ3)
5.4. Embedding Quality (RQ1, RQ3)
5.5. Cross-Modal Retrieval (RQ1)
5.6. Reproducibility and Implementation
6. Discussion
- Threats to validity.
- Scope of the model class.
7. Conclusions
Acknowledgments
References
- van der Aalst, W.M.P. Process Mining—Data Science in Action, 2nd ed.; Springer: Berlin/Heidelberg, Germany, 2016. [Google Scholar]
- van der Aalst, W.M.P.; Weijters, T.; Maruster, L. Workflow Mining: Discovering Process Models from Event Logs. IEEE Trans. Knowl. Data Eng. 2004, 16, 1128–1142. [Google Scholar] [CrossRef]
- Leemans, S.J.J.; Fahland, D.; van der Aalst, W.M.P. Discovering Block-Structured Process Models from Event Logs—A Constructive Approach. In Proceedings of the PETRI Nets; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2013; Volume 7927, pp. 311–329. [Google Scholar]
- Augusto, A.; Conforti, R.; Dumas, M.; Rosa, M.L.; Polyvyanyy, A. Split miner: Automated discovery of accurate and simple business process models from event logs. Knowl. Inf. Syst. 2019, 59, 251–284. [Google Scholar]
- Adriansyah, A.; Munoz-Gama, J.; Carmona, J.; van Dongen, B.F.; van der Aalst, W.M.P. Alignment Based Precision Checking. In Proceedings of the Business Process Management Workshops; Lecture Notes in Business Information Processing; Springer: Berlin/Heidelberg, Germany, 2012; Volume 132, pp. 137–149. [Google Scholar]
- van Zelst, S.J.; Leemans, S.J.J. Translating Workflow Nets to Process Trees: An Algorithmic Approach. Algorithms 2020, 13, 279. [Google Scholar] [CrossRef]
- Radford, A. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the ICML; Proceedings of Machine Learning Research; JMLR: Cambridge, MA, USA, 2021; Volume 139, pp. 8748–8763. [Google Scholar]
- Baltrusaitis, T.; Ahuja, C.; Morency, L. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 423–443. [Google Scholar] [PubMed]
- Koninck, P.D.; vanden Broucke, S.; Weerdt, J.D. act2vec, trace2vec, log2vec, and model2vec: Representation Learning for Business Processes. In Proceedings of the BPM; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2018; Volume 11080, pp. 305–321. [Google Scholar]
- Tavares, G.M.; Oyamada, R.S.; Barbon, S.; Ceravolo, P. Trace encoding in process mining: A survey and benchmarking. Eng. Appl. Artif. Intell. 2023, 126, 107028. [Google Scholar] [CrossRef]
- Colonna, J.G.; Fares, A.A.; Duarte, M.; Sousa, R.T. Process mining embeddings: Learning vector representations for Petri nets. Intell. Syst. Appl. 2024, 23, 200423. [Google Scholar] [CrossRef]
- Kusner, M.J.; Paige, B.; Hernández-Lobato, J.M. Grammar Variational Autoencoder. In Proceedings of the ICML; Proceedings of Machine Learning Research; JMLR: Cambridge, MA, USA, 2017; Volume 70, pp. 1945–1954. [Google Scholar]
- Geng, S.; Josifoski, M.; Peyrard, M.; West, R. Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning. In Proceedings of the EMNLP; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 10932–10952. [Google Scholar]
- Berti, A.; van Zelst, S.J.; Schuster, D. PM4Py: A process mining library for Python. Softw. Impacts 2023, 17, 100556. [Google Scholar] [CrossRef]
- Augusto, A.; Conforti, R.; Dumas, M.; Rosa, M.L.; Maggi, F.M.; Marrella, A.; Mecella, M.; Soo, A. Automated Discovery of Process Models from Event Logs: Review and Benchmark. IEEE Trans. Knowl. Data Eng. 2019, 31, 686–705. [Google Scholar]
- Leemans, S.J.J.; van Zelst, S.J.; Lu, X. A Brief Overview of Process Trees. In Proceedings of the Mining a Scientist’s Process; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2026; Volume 16480, pp. 315–332. [Google Scholar]
- Kourani, H.; van Zelst, S.J. POWL: Partially Ordered Workflow Language. In Proceedings of the BPM; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2023; Volume 14159, pp. 92–108. [Google Scholar]
- Kourani, H.; Park, G.; van der Aalst, W.M.P. A discovery technique for expressive yet sound process models. Process Sci. 2026, 3, 14. [Google Scholar] [CrossRef]
- Dijkman, R.M.; Dumas, M.; van Dongen, B.F.; Käärik, R.; Mendling, J. Similarity of business process models: Metrics and evaluation. Inf. Syst. 2011, 36, 498–516. [Google Scholar] [CrossRef]
- Rama-Maneiro, E.; Vidal, J.C.; Lama, M. Deep Learning for Predictive Business Process Monitoring: Review and Benchmark. IEEE Trans. Serv. Comput. 2023, 16, 739–756. [Google Scholar]
- Grover, A.; Leskovec, J. node2vec: Scalable Feature Learning for Networks. In Proceedings of the KDD; ACM: New York, NY, USA, 2016; pp. 855–864. [Google Scholar]
- Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the ICLR (Poster); OpenReview.net: Alameda, CA, USA, 2017. [Google Scholar]
- Sommers, D.; Menkovski, V.; Fahland, D. Supervised learning of process discovery techniques using graph neural networks. Inf. Syst. 2023, 115, 102209. [Google Scholar] [CrossRef]
- Kourani, H.; Berti, A.; Schuster, D.; van der Aalst, W.M.P. Process Modeling with Large Language Models. In Proceedings of the BPMDS/EMMSAD@CAiSE; Lecture Notes in Business Information Processing; Springer: Cham, Switzerland, 2024; Volume 511, pp. 229–244. [Google Scholar]
- Berti, A.; Wang, X.; Kourani, H.; van der Aalst, W.M.P. Specializing large language models for process modeling via reinforcement learning with verifiable and universal rewards. Process Sci. 2025, 2, 26. [Google Scholar] [CrossRef]
- IEEE Std 1849-2023; IEEE Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams. IEEE: New York, NY, USA, 2023.
- Kingma, D.P.; Welling, M. Auto-Encoding Variational Bayes. In Proceedings of the ICLR, Banff, AB, Canada, 14–16 April 2014. [Google Scholar]
- van der Aalst, W.M.P. A practitioner’s guide to process mining: Limitations of the directly-follows graph. In Proceedings of the CENTERIS/ProjMAN/HCist; Procedia Computer Science; Elsevier: Amsterdam, The Netherlands, 2019; Volume 164, pp. 321–328. [Google Scholar]



| Split | Samples | Avg. Tree Size |
Avg. Tree Depth |
Traces per Sample |
Avg. Trace Length |
Max Petri Nodes |
Max Petri Edges |
|---|---|---|---|---|---|---|---|
| Training | 4000 | 10.31 | 3.18 | 16 | 3.97 | 63 | 78 |
| Validation | 512 | 10.30 | 3.18 | 16 | 3.81 | 55 | 68 |
| Test | 512 | 10.63 | 3.18 | 16 | 3.95 | 58 | 70 |
| Architecture | Optimization | ||
|---|---|---|---|
| Latent dimension | 64 | Optimizer | AdamW (lr , weight decay ) |
| Hidden dimension H | 128 | Batch size | 32 |
| Dropout | 0.15 | Gradient clipping | global norm 5.0 |
| Petri message-passing rounds | 3 | Label smoothing | 0.05 |
| Vocabularies | 40 tree/31 act. tokens | LR schedule | halve on val. plateau (patience 2) |
| Trainable parameters | 524,713 | Early stopping | patience 5 of max. 100 epochs |
| Latent Source | Terminated | Valid Tree | Exact Tree | Petri Conv. | Norm. Edit ↓ | Behavior ↓ |
|---|---|---|---|---|---|---|
| (tree encoder) | 100% | 100% | 48.4% | 100% | 0.178 | 0.675 |
| (log encoder) | 100% | 100% | 40.0% | 100% | 0.287 | 0.755 |
| (Petri encoder) | 100% | 100% | 35.9% | 100% | 0.282 | 0.839 |
| (mean of the three) | 100% | 100% | 44.3% | 100% | 0.230 | 0.713 |
| Method | Model Obtained | Alignments Computable | Fitness | Precision | F1 |
|---|---|---|---|---|---|
| ProcRosetta (log → tree → net) | 100% | 100% | 0.841 | 0.763 | 0.787 |
| Inductive Miner | 100% | 100% | 1.000 | 0.797 | 0.857 |
| Method | Type | Dim. | Behavior ↑ | NN Behavior ↓ | Improv. ↑ |
|---|---|---|---|---|---|
| Directly-follows distribution † | ∘ log features | 332 | 0.736 | 0.784 | 0.782 |
| PetriNet2Vec (PM4Py) [11] | • learned (net) | 64 | 0.643 | 1.115 | 0.451 |
| Activity counts | ∘ log features | 20 | 0.635 | 0.946 | 0.620 |
| ProcRosetta | • learned (log) | 64 | 0.630 | 0.794 | 0.772 |
| ProcRosetta | • learned (all) | 64 | 0.628 | 0.804 | 0.761 |
| PM4Py log case features | ∘ log features | 40 | 0.619 | 0.821 | 0.745 |
| Trace-variant distribution † | ∘ log features | 2548 | 0.614 | 0.781 | 0.785 |
| ProcRosetta | • learned (tree) | 64 | 0.611 | 0.836 | 0.730 |
| ProcRosetta | • learned (net) | 64 | 0.574 | 0.913 | 0.652 |
| Petri structural counts | ∘ net features | 9 | 0.439 | 0.925 | 0.640 |
| Eventually-follows distribution | ∘ log features | 347 | 0.412 | 0.881 | 0.685 |
| Method | Pairwise | Top-1 NN Overlap | Behavior | NN Behavior |
|---|---|---|---|---|
| ProcRosetta | 0.960 | 0.320 | ||
| ProcRosetta | 0.951 | 0.283 | ||
| ProcRosetta | 0.890 | 0.316 | ||
| PetriNet2Vec (PM4Py) | 0.669 | 0.037 | ||
| Petri structural counts | 0.618 | 0.150 | ||
| PM4Py log case features | 0.546 | 0.088 | ||
| Trace-variant distribution | 0.501 | 0.039 | ||
| Directly-follows distribution | 0.488 | 0.102 | ||
| Activity counts | 0.417 | 0.174 | ||
| Eventually-follows distribution | 0.337 | 0.055 |
| Query → Target | Top-1 Accuracy | MRR | Mean Rank |
|---|---|---|---|
| tree → net | 44.3% | 0.523 | 36.2 |
| net → tree | 34.2% | 0.442 | 38.0 |
| tree → log | 29.3% | 0.383 | 32.4 |
| log → tree | 25.6% | 0.354 | 33.2 |
| net → log | 15.4% | 0.236 | 49.4 |
| log → net | 14.1% | 0.236 | 48.2 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).