Submitted:
07 December 2023
Posted:
08 December 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Aria Ops (former vR Ops) generates a mission-critical alert, the customer/user is not able to diagnose or even understand the si tua tion.
- b. SRE/GSS teams involve into the issue resolution.
- c. If the issue necessitates, development teams are inclu ded in the process.
- d. Development spends hours and days to perform root cause analysis.
- e. It provides the fix of the problem.
- f. Participating engineers gain valuable domain knowledge/expertise.
2. Related Art
3. Materials and Methods for ProbRCA
- Aira Ops generates mission-critical alerts.
- ProbRCA sets its general scope to the alert-related other key performance indicators (KPIs) impacted and their monitoring data for trainings.
- Related time series metrics preprocessing, e.g., smoothing, min/max normalization.
- Executing rule induction learning.
- Discovered rules are added to the library of rules.
- Relevant rules are tracked and recommended for alert resolution.
- Aria Ops is collecting and storing data with some monitoring intervals. Traditional monitoring interval is 5 minutes. It means that Aria Ops is averaging the available values of a metric within this interval and storing them with a time stamp corresponding to the end of that interval. As a result, the average can vary from the actual value corresponding to that specific time stamp and the difference may be very large, especially in case of many outliers. We noticed that due to those random fluctuations some correlated metrics are no longer detectable by the correlation analytics in Aria Ops based on Pearson coefficient. Hence, this effect makes the correlation engine of the product unfairly useless.
- This synchronization problem can be only resolved by application of proper data-smoothing techniques. In our experiments, we apply a well-known min-max smoothing technique that Aria Ops is using in UI for visualizing many data points on a small window. It takes a time window (say 6 hours), finds the minimum and maximum values of a metric, and put them in the middle and at the end of that interval in the same order as they appeared in that interval. For example, if the minimum came earlier than the maximum, then the value of the minimum should be put in the middle and the value of the maximum in the end. Then, it shifts the window by 6 hours and reiterate the procedure until the end of the metric.
- Aria Ops is collecting a vast number of metrics from cloud infrastructures, but also a bunch of self-generated metrics constructed by domain experts. The final monitored datasets contain thousands of metrics with highly correlated subsets describing the same process. As a result, the metric correlation engine, or TW can detect hundreds of other metrics with the same behavior. But it will be very hard to separate the metrics that describe the same process from the metrics related to different ones for the detection of possible causations. This problem can be resolved only by users with some expertise. They need to manually separate the possible domains of interrelations and skip analysis within the same areas. That is why, we work below separately with three different datasets thus manually decreasing the total number of possible correlations.
- Alert/alarms are another source of uncertainty in Aria Ops. Many alerts are not directly connected to a problem as user-defined ones are not always sufficiently indicative. From the other side, many problems have appeared without a proper alert generation. In our example below, the problem is not connected to a known alert – ‘Remote Collector Down’ metric doesn’t trigger any alert due to its fast oscillations. We found that a problem analysis is always starts from the corresponding KPI and its behavior. Even if the description of a problem is starting from an alert, the set of appropriate KPI metrics should be identified and described.
- Finally, as we mentioned before, the expert knowledge of Aria Ops engineers remains hidden in internal departments among a few numbers of specialists. As a rule, this knowledge is not systemized, not shared appro priately, cannot be used for a consistent and proactive management and a fast resolution of similar issues especially in cross-customer mode.
4. Trending Problem Scenarios
- ⟹
- *.vmwareidentity.com
- ⟹
- gaz.csp-vidm-prod.com
- ⟹
- *.vmware.com
- ⟹
- *.vrops-cloud.com
- ⟹
- s3-us-west-2.amazonaws.com/vrops-cloud-proxy.

- an issue with the collector service in the cloud proxy reported,
- CaSA service is not sending self-metrics as well,
- but whenever the cluster starts receiving (self-monitoring) metrics data, it turns out that the cloud proxy VM was not down during the span of the issue,
- these patterns are happening periodically and synchronously.
5. Experiments and Discussions
5.1. Data
5.2. Specific Results
5.3. Rule Validation with Dempster-Shafer Theory
6. Evaluation of Results
7. Conclusion and Future Work
8. Patents
Funding
References
- VMware Aria Operations. https://www.vmware.com/products/vrealize-operations.html.
- VMware Aria Operations for Applications. https://www.vmware.com/products/aria-operations-for-applications.html.
- VMware Aria Operations for Logs. https://www.vmware.com/products/vrealize-log-insight.
- VMware Aria Operations for Networks. https://www.vmware.com/products/vrealize-network-insight.html.
- AI ops by Gartner. https://www.gartner.com/en/information-technology/glossary/aiops-artificial-intelligence-operations.
- M. Sole, V. Muntes-Mulero, A.I. Rana, and G. Estrada, Survey on Models and Techniques for Root-Cause Analysis. 2017; arXiv:1701.08556v2, 2017.
- Shafer, G. A mathematical theory of evidence, Princeton University Press, 1976.
- Peñafiel, S.; Baloian, N.; Sanson, H.; Pino, J.A. Applying Dempster–Shafer theory for developing a flexible, accurate and interpretable classifier. Expert Syst. Appl. 2020, 148, 113262. [Google Scholar] [CrossRef]
- Big Panda. https://www.bigpanda.io/.
- Moogsoft . https://www.moogsoft.com/.
- Pager Duty. https://www.pagerduty.com/.
- HPE InfoSight. https://www.hpe.com/us/en/solutions/infosight.html.
- Configuring VMware Cloud Proxies. https://docs.vmware.com/en/vRealize-Operations/Cloud/getting-started/GUID-7C52B725-4675-4A58-A0AF-6246AEFA45CD.html.
- Arrieta, A.B.; Díaz-Rodríguez, N.; et al. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI, Information Fusion, v. 58, pp. 82-115, 2020.
- Ribeira, M.T.; Singh, S.; Guestrin, C. Why should I trust you?: Explaining the predictions of any classifier. https://arxiv.org/pdf/1602.04938v1.pdf, 2016.
- Cohen, W. Fast effective rule induction, Proceedings 12th International Conference on Machine Learning, Tahoe City, California, July 9–12, pp. 115-123, 1995.
- Fürnkranz, J.; Gamberger, D.; Lavrac, N. Foundations of Rule Learning. Springer-Verlag, 2012.
- Poghosyan, A.; Harutyunyan, A.; Grigoryan, N.; Pang, C.; Oganesyan, G.; Ghazaryan, S.; Hovhannisyan, N. An enterprise time series forecasting system for cloud applications using transfer learning. Sensors 2021, 21, 1590. [Google Scholar] [CrossRef] [PubMed]
- Poghosyan, A.V.; Harutyunyan, A.N.; Grigoryan, N.M.; Kushmerick, N. Incident management for explainable and automated root cause analysis in cloud data centers. J. Univers. Comput. Sci. 2021, 27, 1152–1173. [Google Scholar] [CrossRef]
- Harutyunyan, A.N.; Grigoryan, N.M.; Poghosyan, A.V. 2020. Fingerprinting data center problems with association rules, Proceedings of 2nd International Workshop on Collaborative Technologies and Data Science in Artificial Intelligence Applications (CODASSCA 2020), September 14-17, American University of Armenia, Yerevan, Armenia, 159-168, 2020.
- Harutyunyan, A.N.; Poghosyan, A.V.; Grigoryan, N.M.; Hovhannisyan, N.A.; Kushmerick, N. On machine learning approaches for automated log management. J. Univers. Comput. Sci. 2019, 25, 925–945. [Google Scholar] [CrossRef]
- Harutyunyan, A.; Poghosyan, A.; Grigoryan, N.; Kush, N.; Beybutyan, H. Identifying changed or sick resources from logs, Proceedings of 2018 IEEE 3rd International Workshops on Foundations and Applications of Self* Systems (FAS*W), Trento, Italy, pp. 86-91, 2018. [CrossRef]
- Poghosyan, A.V.; Harutyunyan, A.N.; Grigoryan, N.M. Compression for time series databases using independent and principal component analysis, Proceedings of 2017 IEEE International Conference on Autonomic Computing (ICAC), Columbus, OH, USA, 2017, pp. 279-284. [CrossRef]
- Poghosyan, A.V.; Harutyunyan, A.N.; Grigoryan, N.M. Managing cloud infrastructures by a multi-layer data analytics, Proceedings of 2016 IEEE International Conference on Autonomic Computing (ICAC), Wuerzburg, Germany, pp. 351-356, 2016. [CrossRef]
- Harutyunyan, A.N.; Poghosyan, A.V.; Grigoryan, N.M.; Marvasti, M. Abnormality analysis of streamed log data, Proceedings of IEEE Network Operations and Management Symposium (NOMS 2014), May 5-9, Krakow, Poland, 2014.








Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).