K-Means Clusterization and Machine Learning Prediction of European Most Cited Scientific Publications

Angelo Leogrande; Alberto Costantiello; Lucio Laureti

doi:10.20944/preprints202208.0374.v1

Submitted:

21 August 2022

Posted:

22 August 2022

You are already at the latest version

Abstract

In this article we investigate the determinants of the European “Most Cited Publications”. We use data from the European Innovation Scoreboard-EIS of the European Commission for the period 2010-2019. Data are analyzed with Panel Data with Fixed Effects, Panel Data with Random Effects, WLS, and Pooled OLS. Results show that the level of “Most Cited Publications” is positively associated, among others, to “Innovation Index” and “Enterprise Birth” and negatively associated, among others, to “Government Procurement of Advanced Technology Products” and “Human Resources”. Furthermore, we perform a cluster analysis with the k-Means algorithm either with the Silhouette Coefficient and the Elbow Method. We find that the Elbow Method shows better results than the Silhouette Coefficient with a number of clusters equal to 3. In adjunct we perform a network analysis with the Manhattan distance, and we find the presence of 4 complex and 2 simplified network structures. Finally, we present a confrontation among 10 machine learning algorithms to predict the level of “Most Cited Publication” either with Original Data-OD either with Augmented Data-AD. Results show that the best machine learning algorithm to predict the level of “Most Cited Publication” with Original Data-OD is SGD, while Linear Regression is the best machine learning algorithm for the prediction of “Most Cited Publications” with Augmented Data-AD.

Keywords:

Innovation and Invention

;

Processes and Incentives

;

Management of Technological Innovation and R&D

;

Diffusion Processes

;

Open Innovation

Subject:

Business, Economics and Management - Economics

Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.

K-Means Clusterization and Machine Learning Prediction of European Most Cited Scientific Publications

Abstract

Keywords:

Subject:

MDPI Initiatives

Important Links

Subscribe