Submitted:
16 May 2023
Posted:
17 May 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Problematization
3. The Methodology
4. Conceptualizing the Problem via Qualitative Interviews
| Role | Years of Experience |
|---|---|
| Data Scientist | 10+ |
| Data Engineer | 7+ |
| Data Engineer | 10+ |
| ML Developer | 7+ |
| Data Scientist | 6+ |
- The following tasks are recommended to be automated: data ingestion, orchestrating the ML model, and enabling auto-scaling to generate a production-ready ML model. Apart from this, any help for auto-tuning hyperparameters and choosing the best ML models is highly appreciated by the data scientists.
- ML platform services would help in bringing standardization to the way ML models are generated; otherwise, different ML developers may generate ML models in their own style, which may eventually cause trouble in maintaining respective ML models when they are deployed in production.
- Explaining the ML model is one of the most difficult tasks for the ML developers, whereas ML platform services come up with an auto-explainability feature off-the-shelf for any ML model generated.
- Cloud-agnostic ML and data engineering platforms can provide higher performance at a low cost.
- Generating an automatic ML model can generally be used for faster go-to-market needs or proof of concepts.
- Cloud vendors must provide ML-optimized hardware for building an automatic ML model. They ought to provide an option to choose between standard and ML-optimized compute for processing the data.
- ML platforms must allow users to try out different ML models, validate the results, and choose the best ML automatically.
- ML platforms should take care of feature selection and feature extraction-related tasks, which demand significant human time and competence.
- ML platforms should allow users to try the most complicated deep learning algorithms with minimal code or fewer inputs. It should be possible to create a multi-layer neural network and tune it with less human interference. However, it should be done in accordance with the pre-condition criteria set by the respective ML platform vendors; thus, it is prone to improve over time. Hence, it is better to try this option than not try it out due to not having the required knowledge.
- Operationalizing the ML platform models should be made possible with the MLOps service available across the platforms.
- Continuously monitor ML model performance by comparing the results with the baseline model through the MLOps service, and it is also possible to trigger corrective action in terms of retraining the ML model when there is any data drift (when significant variation in the live data is found in comparison to test and validation data).
- ML platforms should work better on a smaller dataset. It will also be very useful for basic ML problems like regression, classification, and time series prediction.
- There is always a cost involved while using ML platform services; hence, we must be careful and aware of which service we are leveraging and for what purpose.
- We should not blindly rely on the results obtained from ML platforms and act. We should set acceptance criteria to validate the results based on relevant metrics.
- We must have a complete grip on the values we are providing as input parameters, as they tremendously affect the ML model’s performance. A small error in input could easily lead to highly biased results.
5. Cloud-Based ML Platforms
6. Cloud-Agnostic ML Platform

7. Experimental Results
7.1. Dataset Description
7.2. Low-Code ML Platform Based on AWS
- ML problem type: we can choose the following options: auto, binary classification, multi-class classification, and regression.
- Experiment run type: we can choose between executing the whole experiment or copying the generated code into a notebook and executing the commands cell-wise.
- Runtime: we can define how long the experiment can execute, how many maximum models it can generate, and the maximum time it can spend generating each model.
- Access: we can restrict access to any IAM role.
- Encryption: we can enable encryption for data present at the S3 bucket level.
- Security: we can use a virtual private cloud connection if we desire to have a highly secure private connection.
7.3. Low-Code ML Platform Based on GCP
7.4. Low-Code ML Platform Based on Microsoft Azure
7.5. Low-Code ML Platform Based on Databricks
7.6. Comparing Models Based on Performance Metrics
7.7. Comparison between ML Platforms
8. Discussion
8.1. Theoretical Implications
8.2. Practical Implications
9. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Brynjolfsson, E. and McElheran, K. The rapid adoption of data-driven decision-making. Am. Econ. Rev. 2016, 106, 133–39. [CrossRef]
- Provost, F. and Fawcett, T. Data science and its relationship to big data and data-driven decision making. Big Data 2013, 1, 51–59. [CrossRef]
- Available online: https://onlinedegrees.sandiego.edu/what-is-data-science/ (accessed on 9 September 2022).
- Alsharef, A., Aggarwal, K., Kumar, M. and Mishra, A. Review of ML and AutoML solutions to forecast time-series data. Archives of Computational Methods in Engineering 2022, 1-15. [CrossRef]
- Available online: https://www.ibm.com/cloud/blog/low-code-vs-no-code (accessed on 24 September 2022).
- Di Sipio, C., Di Ruscio, D. and Nguyen, P.T. Democratizing the development of recommender systems by means of low-code platforms. In Proceedings of the 23rd ACM/IEEE International Conference on Model Driven Engineering Languages and Systems: Companion Proceedings, October 2020; pp. 1–9.
- Available online: https://www.stackovercloud.com/2020/03/03/aws-named-as-a-leader-in-gartners-magic-quadrant-for-cloud-ai-developer-services/ (accessed on 23 September 2022).
- Available online: https://www.Databricks.com/spark/comparing-Databricks-to-apache-spark (accessed on 10 August 2022).
- Luo, G. A review of automatic selection methods for machine learning algorithms and hyper-parameter values. Netw. Model. Anal. Health Inform. Bioinform. 2016, 5, 1–16. [CrossRef]
- Li, Y., Ren, X., Zhao, F. and Yang, S. A Zeroth-Order Adaptive Learning Rate Method to Reduce Cost of Hyperparameter Tuning for Deep Learning. Appl. Sci. 2021, 11, 10184. ttps://doi.org/10.3390/app112110184.
- Subramanian, M., Shanmugavadivel, K. and Nandhini, P.S. On fine-tuning deep learning models using transfer learning and hyper-parameters optimization for disease identification in maize leaves. Neural Computing and Applications 2022, 1-18. [CrossRef]
- Bahri, M., Salutari, F., Putina, A. and Sozio, M. AutoML: State of the art with a focus on anomaly detection, challenges, and research directions. International Journal of Data Science and Analytics 2022, 1-14. [CrossRef]
- Available online: https://docs.microsoft.com/th-th/Azure/architecture/solution-ideas/articles/Azure-machine-learning-solution-architecture (accessed on 12 May 2022).
- Available online: https://mikaelahonen.com/en/blog/comparison-of-machine-learning-platforms-in-major-clouds/ (accessed on 1 September 2022).
- Available online: https://cloudsolutions.academy/cloud-compare/ (accessed on 22 August 2022).
- Available online: https://www.scieneers.de/automl-a-comparison-of-cloud-offerings/ (accessed on 13 May 2022).
- Available online: https://medium.com/@vineetjaiswal/introduction-comparison-of-mlops-platforms-aws-Sagemaker-Azure-machine-learning-gcp-vertex-ai-9c1153399c8e (accessed on 1 August 2022).
- Das, P., Ivkin, N., Bansal, T., Rouesnel, L., Gautier, P., Karnin, Z., Dirac, L., Ramakrishnan, L., Perunicic, A., Shcherbatyi, I. and Wu, W. Amazon Sagemaker Autopilot: A white box AutoML solution at scale. In Proceedings of the Fourth International Workshop on Data Management for End-to-End Machine Learning, June 2020; pp. 1–7.
- Available online: https://cloud.Google.com/vertex-ai/docs/start/automl-users (accessed on 22 July 2022).
- Available online: https://www.Databricks.com/blog/2020/01/30/what-is-a-data-lakehouse.html (accessed on 10 August 2022).
- Available online: https://docs.Databricks.com/clusters/index.html (accessed on 22 August 2022).
- Abdel Hai, A. and Forouraghi, B. On scalability of distributed machine learning with big data on apache spark. In Proceedings of the International Conference on Big Data, June 2018; Springer, Cham, 2018; pp. 209–219.
- Wan, K.W., Wong, C.H., Ip, H.F., Fan, D., Yuen, P.L., Fong, H.Y. and Ying, M. Evaluation of the performance of traditional machine learning algorithms, convolutional neural network and AutoML Vision in ultrasound breast lesions classification: A comparative study. Quant. Imaging Med. Surg. 2021, 11, 1381. [CrossRef]
- Available online: https://research.aimultiple.com/automl-case-studies/ (accessed on 3 September 2022).



| Algorithm | Hyperparameters |
|---|---|
| Decision Tree | max_depth, min_impurity_split,min_samples_leaf, max_leaf_nodes |
| Random Forest | n_estimators,max_features |
| Support Vector Machine | kernels, penalty value [3], tol |
| K-Nearest Neighbor | n_neighbors, metric, weights |
| Naïve Bayes | kernel density estimator, window width |
| Stochastic Gradient Boosting | learning_rate, n_estimators, subsample, max_depth |
| Neural Network | no. of hidden layers, no. of nodes in hidden layers, activation function, no. of epochs, learning rate |
| AI & ML Use-case | AWS | GCP | Azure |
|---|---|---|---|
| ML Platform | Sagemaker | Vertex AI | Synapse |
| Computer Vision | Amazon Rekognition & Lookout for Vision | Vision AI | Azure Cognitive Service Computer Vision |
| AutoML | Sagemaker AutoPilot | Vertex AI AutoML | Azure Machine Learning Service—Automated ML |
| ML Frameworks Supported | TensorFlow, PyTorch, Apache MXNet | TensorFlow, PyTorch, Scikit-Learn | TensorFlow, PyTorch, ML.Net |
| NLP Service | Amazon Comprehend | Natural Language AI | Azure Cognitive Service Text Analytics |
| Speech to Text | Amazon Transcribe | Speech-to-Text | Azure Cognitive Service Speech to Text |
| Text to Speech | Amazon Polly | Text-to-Speech | Azure Cognitive Service Text to Speech |
| Language Translation | Amazon Translate | Cloud Translation | Azure Cognitive Service Translator |
| Conversational Service | Amazon Lex | Dialogflow | Azure Bot Service |
| Text Extraction | Amazon Textract | Document AI | Azure Form Recognizer |
| Recommendation & Personalization Service | Amazon Personalize | Recommendations AI | Azure Cognitive Service Personalizer |
| Model Performance Score | GCP | AWS | Azure | Databricks |
|---|---|---|---|---|
| R2 | 0.831 | 0.836 | 0.898 | 0.822 |
| Aspect | Similarities |
|---|---|
| Prerequisites | Dataset must be created, and data must be uploaded to Cloud Storage for model consumption |
| Feature Importance Results | All three clouds have found similar feature importance results with the best model generated |
| Evaluation Metrics | All three clouds have quite similar evaluation metrics as follows: MAE, MSE, RMSE and R Squared |
| AI and ML Features | AWS AutoPilot | GCP AutoML | Azure AutoML |
|---|---|---|---|
| Traceability | All the AutoML operation logs are stored under given directory in S3 Bucket. | Model related logs are available only in the UI, however it is possible to download it. | Model related logs are available only in the UI, however it is possible to download it. |
| Complexity | AutoML models can be created at ease with very less touchpoints. | We are expected to provide a few additional inputs for generating AutoML models compared to AWS. | AutoML models can be created at ease with very less touchpoints. |
| Adaptability | Highly adaptable | Not so adaptable. | Somewhat adaptable. |
| Co-Authoring | Yes, with shared compute for all the developers involved. | Yes, with shared compute for all the developers involved. | Yes, but dedicated compute is required for each developer. |
| Explainability | Models created are automatically explained. | Detailed Model summary and explanation are available as a part of UI. | Must provide compute instance or compute cluster manually for explaining each model. |
| Deployment | It can be deployed to internal endpoints. | Possible to containerize the model and deploy it to any endpoint. | Possible to deploy the model to any endpoint. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).