Preprint
Article

This version is not peer-reviewed.

Enhancing Patient Care and Nursing Practice Through AI‐Driven Analysis of Patient Complaints: A Pilot Study

Submitted:

09 September 2026

Posted:

10 September 2026

You are already at the latest version

Abstract
Background/Objectives: The integration of Artificial Intelligence (AI) and Machine Learning (ML) in healthcare has shown promise for analyzing complex data and identifying patterns that can enhance clinical decision-making and patient care. Patient complaints represent a valuable yet underutilized data source for improving healthcare quality and patient safety. This study aimed to investigate whether an AI-driven model could effectively categorize patient complaints to extract actionable insights that support person-centered care and nursing practice to improve care quality. Method: An explorative study design was conducted as a pilot project. A total of 368 patient complaint reports collected by a Patients’ Advisory Committee were analyzed using AI-driven clustering and classification techniques. The complaints were pre-processed through text cleaning, tokenization, and vectorization before being analyzed with various clustering algorithms. Results: The AI-based classification identified four main areas that directly or indirectly influenced patients’ perceptions of care quality: (1) Care and Treatment, (2) Communication, (3) Administrative Management, and (4) Availability. These areas encompassed themes such as diagnostic delays, poor communication, fragmented administrative processes, and limited accessibility. The results highlighted that communication failures, insufficient empathy, and lack of continuity were central to patient dissatisfaction. Automated analysis enabled the identification of recurring issues and systemic weaknesses that traditional manual review methods often overlook. Conclusion: This pilot study demonstrates the feasibility and potential of using an AI-driven model to categorize patient complaints and identify patterns that may provide actionable insights for person-centered care and nursing practice to improve quality of care. The multilabel approach enabled an efficient and systematic analysis of complex complaints and may, following further validation, offer a scalable and consistent basis for understanding patient concerns. In nursing practice, integrating AI into patient feedback systems could support the earlier identification of recurring concerns and contribute to more responsive, evidence-informed, and person-centered quality improvement. However, larger datasets, independent expert validation, and careful consideration of ethical, legal, and governance requirements are needed before the model can be implemented more broadly.
Keywords: 
;  ;  ;  ;  

1. Introduction

The use of artificial intelligence (AI) and machine learning (ML) models in healthcare have shown to have a significant capability in analyzing large datasets, detecting patterns and generating actional insights. The use of AI expands across multiple domains, in fields such as, diagnostics with image recognition and timely detection of diseases in medical imaging [1], clinical decision-making [2,3,4], care interventions [5], automated care processes [6], patient administration [2], and patient monitoring [7]. These examples of AI applications aim to enhance efficiency, improve care outcomes, and support healthcare providers in delivering high-quality care [2].
Patient experience and satisfaction are indicators of healthcare quality [8] and an important issue for healthcare providers to fulfill their caring duties, to ensure patient safety and for compassionate clinical nursing [9]. However, the number of healthcare complaints is growing despite an increased focus on person-centered care. Recognizing patient experiences as valuable feedback can play a crucial role in improving the quality of care [10]. Research emphasizes the need to create possibilities for patients to express their experiences of treatment and encounters in healthcare [11]. According to national legislation, a Patients’ Advisory Committee (PAC) is responsible for collecting and analyzing patient experiences and complaints, supporting patients and relatives in raising concerns, and facilitating access to appropriate corrective measures [12]. However, subjectivity and variability in interpreting these complaints reports pose significant challenges, as they are based on individual patients’ subjective experiences. Misinterpretations or inappropriate recommendations from the administrators can lead to inadequate corrective actions and diminished trust in healthcare systems [13,14,15].
Classifying patient complaints using clustering algorithms to categorize the issues described and the patients’ proposed solutions can reduce the risk of misinterpretation and presents a promising opportunity to enhance patient care and nursing practice [16]. AI-driven clustering and sentiment analysis can provide objective, data-supported recommendations, leading to more effective policymaking and healthcare improvements [17]. Machine learning techniques, particularly natural language processing (NLP) and clustering algorithms, offer a data-driven approach to analyze unstructured complaint reports [18]. An AI-model can thereby categorize common themes in complaints, highlight systemic issues, and provide objective, structured insights for corrective action. Furthermore, AI can provide more reliable insights, enabling safer analyses and more appropriate actions to improve person-centered care, patient safety and care quality.
To enhance the quality of care, it is essential that the analysis of complaints and the resulting actions contribute to improved patient safety [19] and care quality [10]. While much of the work defining care quality has focused on preventing negative outcomes, there is a need to shift focus towards leveraging positive quality indicators, such as patients’ suggestions for solutions to care problems. This study aimed to investigate whether an AI-driven model could effectively categorize patient complaints to extract actionable insights that support person-centered care and nursing practice to improve care quality.

2. Materials and Method

2.1. Study Design

An explorative design was chosen in a pilot study to apply AI-driven clustering for analysis [20] of complaint reports to identify trends, patterns, and actionable insights that can aid in improving healthcare services, optimizing responses to patient concerns, and promoting evidence-based interventions for better patient outcomes. The manuscript was prepared in accordance with the COREQ (Consolidated Criteria for Reporting Qualitative Research) guidelines: the completed checklist is provided in the Supplementary Materials.

2.2. Machine Learning Analysis

Two complementary machine learning approaches were applied to analyze the complaint texts, one unsupervised clustering and one supervised classification [21]. All analyses were implemented in Python using the Scikit-Learn library.

2.2.1. Text Preprocessing and Representation

The project includes machine learning analysis of already acquired data comprising 368 patient complaint reports. Before analysis, the complaint texts were preprocessed to create a consistent and informative input for the machine learning models. All text was converted to lowercase, and punctuation and predefined irrelevant terms were removed. Swedish stop words were also excluded to reduce noise and increase the relative importance of content-bearing words. The complaint reports had previously been compiled into a structured file format that preserved relevant textual information while minimizing inconsistencies that could negatively affect model performance. To further improve data quality, infrequent and non-informative outlier words, such as repeated character sequences and obvious misspellings, were filtered out by applying a minimum occurrence threshold before terms were included in the vocabulary.
After preprocessing, the texts were transformed into numerical representations, i.e., vectorized, using term frequency–inverse document frequency (TF-IDF) [22]. In this representation, terms are weighted according to how frequently they occur in an individual complaint relative to their occurrence across the entire dataset, thereby highlighting words that are particularly informative for distinguishing between complaints. To capture meaningful combinations of words, the vectorization also included n-grams ranging from one to three words [23]. This allowed the model to represent not only individual terms but also short phrases, which can better reflect the meaning of expressions commonly found in complaint texts. The resulting fixed-length vector representation enabled the complaint texts to be analyzed computationally in both the clustering and classification tasks. Although this bag-of-words-based representation does not capture deeper semantic relationships or broader contextual meaning, it provides a transparent and effective approach for modeling textual similarity in smaller datasets such as the present one.
In addition to the TF-IDF representation, a contextual language representation based on the pretrained Swedish language model KB-BERT was also explored [24]. Few-shot text classification leverages the representational power of pretrained models when limited data is available and has shown promise in the medical domain before [25]. Therefore, KB-BERT was used together with the classification algorithm that achieved the best performance on the bag-of-words data, to assess whether a context-sensitive representation could improve classification results. Unlike TF-IDF, KB-BERT captures semantic and contextual relationships between words, which may be advantageous when analyzing complaint texts in which meaning depends on linguistic context. This comparison enabled an evaluation of whether a more advanced language model provided added value beyond the simpler and more interpretable bag-of-words approach.

2.2.2. Clustering Analysis

To explore patterns in the complaint texts without using predefined labels, an unsupervised clustering approach was applied. The K-Means clustering algorithm was used to group complaints into clusters based on their textual similarity in the TF-IDF feature space [26].
The number of clusters was determined using a data-driven approach rather than pre-specifying the number arbitrarily. Specifically, the dataset was first examined by counting the number of complaint cases within each main categorization type. Categories containing fewer than 20 cases were excluded to ensure sufficient representation for meaningful analysis. After this filtering step, four categories remained, and therefore four clusters were used in the clustering analysis, see Figure 1.

2.2.3. Classification Analysis

A supervised machine learning approach was used to assess whether complaint texts could be automatically classified into the four main complaint categories. Several classification algorithms representing different learning paradigms were evaluated, including tree-based, probabilistic, ensemble, and kernel-based methods [21]. The evaluated algorithms were Decision Tree [27], Random Forest [28], Multinomial Naïve Bayes [29], Linear Support Vector Classifier (LinearSVC) [30], and Support Vector Classifier (SVC) [31]. All models were implemented using their standard implementations in the Python library Scikit-learn. The evaluated algorithms and their hyperparameters are summarized in Table 1. To provide a reference point for model performance, a dummy classifier producing random predictions was included as a baseline representing performance expected by chance.
The evaluation was conducted in two stages. First, all classification algorithms were trained and evaluated using the TF-IDF representation of the complaint texts. The algorithm achieving the best overall performance according to the selected evaluation metrics was then further evaluated using contextual text representations derived from a pretrained language model. Among the evaluated methods, LinearSVC demonstrated the strongest overall performance and was therefore selected for further evaluation with contextual embeddings.
To investigate whether contextual language representations could improve classification performance given the limited dataset, a BERT-based few-shot approach was also explored. Few-shot text classification leverages the representational capacity of pretrained language models to extract meaningful features from relatively small datasets and has shown promising results in medical and clinical text classification tasks [25]. In this study, the pretrained language model KB-BERT was used to encode the complaint texts into high-dimensional contextual embeddings [24]. These embeddings capture semantic meaning and contextual relationships between words, allowing the model to represent complaints based on how terms are used within the surrounding text rather than as isolated tokens. The resulting embeddings were then used as input features for the LinearSVC classifier.
Given the relatively small dataset (368 complaint reports), model performance was evaluated using a 10-fold cross-validation procedure to ensure robust estimates while maximizing the use of available data. Three evaluation metrics were used: accuracy, the Jaccard index, and the area under the receiver operating characteristic curve (ROC–AUC). Accuracy measures the proportion of correctly classified instances and ranges from 0 to 1, where higher values indicate better performance. The Jaccard index measures the similarity between predicted and true class labels, calculated as the intersection divided by the union of the predicted and true label sets, and ranges from 0 to 1. ROC–AUC was included because it provides a robust measure of classification performance under class imbalance, which was present in the dataset due to differing numbers of complaint reports across the four categories. ROC–AUC ranges from 0.5 (performance equivalent to random guessing) to 1.0 (perfect classification). Because of the class imbalance, ROC–AUC was considered the primary evaluation metric. For each metric, performance was summarized as the mean across the 10 cross-validation folds, with 95% confidence intervals calculated from the fold-wise results.
Precision, recall, and F1-score were initially considered as evaluation metrics but were not included in the final analysis. Because the classification task involved assigning each complaint report to a single category and the evaluation used a one-vs-rest scheme with micro-averaging, the resulting precision, recall, and F1 values were identical to the accuracy metric. In such settings, these measures do not provide additional information beyond overall accuracy [32,33]. For this reason, the evaluation focused on accuracy, the Jaccard index, and ROC–AUC, which provide complementary perspectives on classification performance.

2.3. Ethical Consideration

Ethical approval was not required for this study, as no human participants or identifiable personal data were involved. However, permission to conduct the pilot project was granted by the responsible head of the PAC, and the study was conducted in accordance with the ethical principles outlined in the Declaration of Helsinki [34].

3. Results

The results are presented in two stages: first, an exploratory clustering analysis of the complaint reports, followed by a supervised classification of the texts into the main complaint categories.

3.1. Clustering of Complaint Reports

The results of the K-Means clustering of the complaint reports are illustrated in Figure 2. The clustering analysis provided an overview of how complaint texts were grouped based on textual similarity. Each point in the figure represents an individual complaint report, while colors indicate the cluster to which the complaint was assigned. Complaints located close to each other in the visualization reflect greater similarity in textual content.
The clustering analysis identified several groups of complaints with similar content, with some clusters containing relatively many reports while others consisted of only a few cases. This pattern suggests that certain types of complaints were more frequently reported, while others occurred more sporadically. Although the clustering provided a structured view of the complaint data, the relatively small dataset (368 complaint reports) limited the ability of the algorithm to identify clearly separated clusters and robust patterns across all complaint types.
Because the clustering analysis produced a fragmented distribution of clusters and uneven cluster sizes, a supervised classification approach was subsequently applied. Instead of discovering clusters without predefined labels, the classification task aimed to categorize complaint texts according to the four main areas used in the dataset. Both the full complaint text and the summarized version of the complaint reports were tested as input for the classification models. The results showed that the summarized texts provided comparable performance to the full complaint descriptions. The classification analysis therefore focused on identifying the four main areas reflected in the patient reports.

3.2. Classification of Complaint Reports Into Main Categories

The performance of the evaluated classification algorithms is summarized in Table 2. Overall, all machine learning models outperformed the random baseline across all evaluation metrics, indicating that meaningful patterns could be learned from the complaint texts.
Among the evaluated methods, LinearSVC achieved the best overall performance, with an accuracy of 0.690 (95% CI 0.595–0.785), a Jaccard index of 0.534 (95% CI 0.425–0.643), and a ROC–AUC score of 0.787 (95% CI 0.727–0.847). Random Forest and SVC also demonstrated relatively strong performance, with Random Forest achieving an accuracy of 0.631 (95% CI 0.530–0.732) and ROC–AUC of 0.743 (95% CI 0.677–0.809), while SVC achieved an accuracy of 0.638 (95% CI 0.545–0.731) and ROC–AUC of 0.749 (95% CI 0.692–0.806).
Decision Tree and Naïve Bayes showed lower performance compared with the other algorithms, with accuracies of 0.528 (95% CI 0.451–0.605) and 0.514 (95% CI 0.425–0.603), respectively. Their ROC-AUC and Jaccard scores were also lower, indicating weaker agreement between predicted and true categories. The contextual language representation using KB-BERT combined with LinearSVC produced competitive results, achieving an accuracy of 0.610 (95% CI 0.515–0.705), a Jaccard index of 0.446 (95% CI 0.347–0.545), and a ROC–AUC score of 0.729 (95% CI 0.668–0.790). Although this approach captured contextual semantic information in the complaint texts, it did not outperform the TF-IDF-based LinearSVC model in this dataset. One possible explanation is the relatively limited size of the dataset (368 complaint reports), which may restrict the ability of contextual language models to fully leverage their representational capacity. In smaller datasets, simpler bag-of-words representations such as TF-IDF can sometimes provide more stable and discriminative features for traditional classifiers. In addition, patient complaint texts often contain recurring keywords and short expressions describing similar issues, which may be effectively captured by n-gram-based representations without requiring deeper contextual modeling. As expected, the random baseline classifier showed substantially lower performance, with an accuracy of 0.253 (95% CI 0.240–0.266), a Jaccard index of 0.137 (95% CI 0.132–0.142), and a ROC–AUC score close to random chance at 0.502 (95% CI 0.493–0.511).
Taken together, although LinearSVC achieved the highest observed performance, the overlapping confidence intervals suggest that several models performed comparably. Overall, the classification analysis indicates that the complaint texts can be categorized into four main areas with at least some reliability. These areas reflect key dimensions of patients’ experiences and concerns regarding healthcare encounters. In the following sections, the content of the complaint reports is described in relation to each of the four categories (Care and treatment, Communication, Administrative management, and Availability) to provide a clearer understanding of the types of issues raised by patients and the aspects of care that most strongly influence their experiences.

3.2.1. Care and Treatment

Care and treatment emerged as central determinants of patients’ experiences of healthcare quality, encompassing interactions with providers, diagnostic accuracy, treatment effectiveness, and clinical outcomes. A recurring theme in patient complaints was the perception of mistreatment or inadequate care, reflecting a widespread concern across reported cases. Dissatisfaction was frequently linked to delays in diagnosis, misdiagnosis, and insufficient investigation of symptoms, which led to frustration and anxiety, particularly when patients felt dismissed or not taken seriously. These results underscore the importance of comprehensive assessments and accurate diagnostic processes.
Medication management and treatment planning also significantly influenced patient experiences. Complaints commonly involved incorrect prescriptions, insufficient information about side effects, and inconsistencies in treatment regimens, contributing to uncertainty and reduced trust. Communication regarding the purpose, effects, and risks of treatments was therefore essential. Furthermore, treatment outcomes played a key role in shaping satisfaction as lack of improvement or unresolved concerns often led to frustration. Managing expectations, monitoring effectiveness, and adapting care plans were critical in aligning outcomes with patient needs.
The complaints reported concerns related to examinations and clinical assessments. Greater transparency and information of clinical decision-making processes were important when addressing these issues. Overall, the complaints highlighted the importance of person-centered care alongside improving diagnostic accuracy, communication, and treatment management, for enhancing trust and overall quality of care.

3.2.2. Communication

Communication was frequently described in the complaints as influencing patients’ perceptions of quality, involvement, and overall care experience. Many complaints referred to experiences of not being treated with respect and dignity, with patients describing that they were perceived as medical cases rather than as individuals.
A recurring theme was the lack of clear and accessible information regarding care, treatment plans, and overall health conditions. Insufficient communication was described as a barrier between healthcare providers and patients, contributing to reduced trust and difficulties in making informed decisions about health. Some complaints described that structured and calm information was associated with better understanding and engagement in care.
Consent and participation in decision-making were also commonly addressed. Complaints described a lack of discussion regarding treatment options, as well as risks and benefits, which was associated with frustration and a sense of powerlessness. Patients expressed a desire to be involved in decisions affecting their health and reported that their input was not always sought or considered.
The complaints also described experiences of not being listened to and frequently indicated that patients’ concerns, symptoms, or requests were dismissed or not taken seriously by healthcare professionals. Complaints further reflected that when healthcare providers actively listened, validated concerns, and demonstrated that patient input was valued, this was associated with more positive care experiences. Another prominent aspect was a lack of empathy. Many complaints described interactions as rushed, impersonal, or purely clinical, leaving patients feeling unsupported. In contrast, descriptions of empathetic encounters, where healthcare providers showed genuine concern, offered reassurance, and acknowledged the emotional impact of medical conditions, were associated with more positive patient experiences. Small gestures, such as maintaining eye contact, using a calm and reassuring tone, and expressing understanding, were described as influencing how patients perceived their care.

3.2.3. Administrative Management

Administrative management was frequently described in the complaints as influencing patient experiences of healthcare quality. The complaints highlighted administrative shortcomings that were associated with frustration, delays, and inefficiencies in the care process. Reported issues included certificate handling, medical documentation, care planning, resource allocation, and communication between healthcare units, affecting both accessibility and continuity of care.
Delays in the handling of medical certificates and sick leave documentation were commonly reported and were described as leading to complications related to employment, insurance, and social security. Deficiencies in medical record management, such as inaccuracies, missing information, and limited accessibility, were also described and were associated with disruptions in continuity of care, repeated tests, and the need to repeatedly provide the same information.
Shortcomings in care planning and healthcare processes were reflected in descriptions of fragmented and reactive care, long waiting times, and unclear information regarding next steps. Complaints also highlighted issues related to limited continuity of care providers, with patients describing repeated encounters with new professionals and the need to recount their medical history multiple times.
Delays in referral processing and insufficient coordination between healthcare units were further described. Complaints included prolonged waiting times for specialist care and limited information regarding referral status. In addition, insufficient communication and information sharing between providers were described as contributing to fragmented care and repeated examinations.

3.2.4. Availability

Availability and access to healthcare were frequently described in the complaints as influencing patient experiences of care. Difficulties in accessing services were commonly reported, including challenges in contacting healthcare providers, long waiting times, delayed feedback, and limited access to treatment options. These issues were described as contributing to frustration, uncertainty, and delays in receiving care.
A recurring concern was difficulty in establishing contact with healthcare services. Complaints described challenges in reaching healthcare professionals, booking appointments, and obtaining information about care. These difficulties were particularly evident in primary care and included long telephone queues, limited digital access, and inconsistent communication channels.
Long waiting times and insufficient follow-up were also commonly reported. Complaints described delays in appointments, diagnostic procedures, and specialist referrals, often combined with uncertainty regarding test results and subsequent care. In addition, limited access to alternative or newer treatment options was described, particularly in relation to administrative barriers, lack of information, and delays in implementation.

4. Discussion

The use of AI in this study enabled a novel, efficient, and systematic categorization of patient complaint reports. By applying a multilabel classification algorithm, we were able to capture the complex nature of the complaints, which often relate to multiple categories simultaneously. Similar approaches have been successfully applied in biomedical and healthcare contexts to address the challenges of multilabel classification in unstructured text data [35,36]. In our case, the AI- model was trained to identify patterns and distinguish between the category areas, allowing for a more nuanced and automated analysis of the data. This transition to an AI-based framework significantly reduced the manual workload of PAC administrators and accelerated the feature development process [37].
The core objective of the model was to approximate document representations in a way that aligned with the existing classification system. To this end, the algorithm was adapted to handle multilabel classification, enabling it to detect sub-areas within the broader thematic categories. This design was informed by earlier studies that successfully implemented deep neural networks for extracting multiple, overlapping labels in clinical and patient-oriented texts [35]. The classification model developed in this study demonstrated a high degree of reliability and offered valuable insights into the nature and distribution of patient concerns across a national dataset. Notably, the model aligns with existing categorization practices, while also introducing a more consistent and scalable approach to the analysis of large volumes of complaint reports. These results suggest that the model may contribute to enhancing current analytical frameworks, particularly in settings where systematic and reproducible classification is required.
The primary outcome of implementing this AI-based approach is a more reliable, scalable, and a faster method for analyzing patient complaint reports, which in turn supports safer and more data-informed decision-making in nursing practice. Notably, the classification results, considered across the four categories—care and treatment, communication, administrative management, and availability—align with prior research showing that patient complaint reports reflect recurring challenges in communication, availability, and perceived quality of care [38,39,40,41,42,43,44]. Together, these results underscore the importance of strengthening core aspects of care delivery, including improvements in diagnostic rigor, communication, and patient involvement. The adoption of person-centered approaches, characterized by empathy and continuity, is also essential, as these have been shown to enhance trust and satisfaction [41,45]. In addition, system-level strategies, including digitalization and streamlined processes, may further support more coherent and reliable care experiences [46].
Despite the clear advantages, integrating AI technologies into healthcare workflows presents several challenges. Ethical, legal, and regulatory concerns must be carefully navigated, especially regarding data privacy, algorithmic transparency, and interpretability [47]. Moreover, using AI into clinical practice demands rigorous governance frameworks and workforce readiness [48]. As such, interdisciplinary collaboration, bringing together data scientists, clinicians, administrators, and patients, is essential to ensure that AI tools are not only technically robust but also aligned with core healthcare values and goals [49]. Importantly, this study has initiated a much-needed interdisciplinary dialogue, highlighting the potential of AI not as a replacement for human judgment, but as a support tool that can enhance understanding, efficiency, and responsiveness in healthcare. Continued development in this area requires close cooperation among stakeholders, careful evaluation of safety, acceptability and efficacy of AI systems, and continual monitoring of ethical, legal and social implications [49].

Strengths, Limitations and Further Research

The strength of this pilot study lies in demonstrating the feasibility of using AI to classify patient complaint reports efficiently and systematically. The multilabel approach captured the complexity of complaints while reducing the manual workload and providing a scalable foundation for further validation and development. To strengthen the trustworthiness and practical relevance of the analysis, the AI-generated categories and themes were reviewed and verified by an experienced PAC case officer. Drawing on extensive experience of handling and assessing patient complaints, the case officer evaluated their consistency and face validity in relation to the original reports. However, the study did not include an independent manual coding procedure to assess agreement between human experts or to compare human coding with the machine learning model. Instead, the existing PAC categorizations were used as reference labels for training and evaluation. The dataset was also relatively small, partly because the reports contained sensitive and detailed accounts of patients’ experiences and complaints concerning healthcare encounters, which restricted access to and use of the data. These limitations may affect the model’s generalizability. Future studies should therefore validate the model using larger, carefully governed datasets and independent coding by multiple experts.

5. Conclusions

The results of this pilot study indicate that the model can approximate existing categorization practices while offering a potentially consistent, time-efficient, and scalable approach to analyzing patient complaint reports. The results therefore support the feasibility of using AI to categorize complaint content and identify recurring areas and sub-areas of concern. When applied to data collected across regions, such an approach may contribute to a broader national understanding of patients’ reported experiences of healthcare. A structured AI-supported categorization process could complement free-text descriptions and provide more systematic and actionable insights for quality care improvement, person-centered care, and nursing practice. However, the study did not establish that the model reduces subjective bias or directly improves the quality of care. Further research using larger datasets and independent expert validation is therefore needed. Ethical, legal, and governance considerations must also be addressed to ensure that AI-supported analyses are transparent, responsible, and aligned with person-centered care principles.

Acknowledgments

The authors are grateful to the administrators at PAC for their support with data from the patient-reported complaints. This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. The authors declare that no funding source had any involvement in study design, data collection, analysis, interpretation, manuscript writing, or the decision to submit for publication.

Conflicts of Interest

The authors declared no conflicts of interest.

References

  1. Esteva, A.; Kupel, B.; Novoa, R.A.; Ko, J.; Swetter, A.M.; Blue, H.M.; et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature 2017, 542, 115–118. [Google Scholar] [CrossRef] [PubMed]
  2. Davenport, T.; Kalakota, R. The potential for artificial intelligence in healthcare. Future Healthc. J. 2019, 6, 94–98. [Google Scholar] [CrossRef] [PubMed]
  3. Giordano, C.; Brennan, M.; Mohamed, B.; Rashidi, P.; Modave, F.; Tighe, P. Accessing artificial intelligence for clinical decision-making. Front Digit Health 2021, 3, 645232. [Google Scholar] [CrossRef] [PubMed]
  4. Hosny, A.; Parmar, C.; Quackenbush, J.; Schwartz, L.; Aerts, H. Artificial intelligence in radiology. Nat. Rev. Cancer 2018, 18, 500–510. [Google Scholar] [CrossRef] [PubMed]
  5. Schwalbe, N.; Wahl, B. Artificial intelligence and the future of global health. Lancet 2020, 395, 1579–1586. [Google Scholar] [CrossRef] [PubMed]
  6. Shaheen, M. Applications of artificial intelligence (AI) in healthcare: A review. In ScienceOpen Preprints; 2021. [Google Scholar]
  7. Shaik, T.; Tao, X.; Higgins, N.; Gururajan, L.R.; Zhou, X.; Acharya, U. Remote patient monitoring using artificial intelligence: Current state, applications, and challenges. WIREs Data Min. Knowl. Discov. 2023, 13, e1485. [Google Scholar] [CrossRef]
  8. Larson, E.; Sharma, J.; Bohren, M.; Tuncalp, Ö. When the patient is the expert: measuring patient experience and satisfaction with care. Bull. World Health Organ. 2019, 97, 563–569. [Google Scholar] [CrossRef] [PubMed]
  9. Allen, D. Nursing and the future of ‘care’ in the health care systems. J. Health Serv. Res. Policy 2015, 20. [Google Scholar] [CrossRef] [PubMed]
  10. Almomani, R.; Al-Ghdabi, R.R.; Bany-Humdan, K. Patients’ satisfaction of health service quality in public hospitals: A PubHosQual analysis. Manag Sci. Lett. 2020, 10, 1803–1812. [Google Scholar] [CrossRef]
  11. Kumah, E. Patient experience and satisfaction with a healthcare system: connecting the dots. Int. J. Healthc. Manag 2017, 12. [Google Scholar] [CrossRef]
  12. The Patient Act (2017:372). Available online: http://www.riksdagen.se/sv/Dokument-Lagar/Lagar/Svenskforfattningssamling/sfs-sfs-2014-821.
  13. Skålen, C.; Nordgren, L.; Annerbäck, E.-M. Patient complaints about care in a Swedish county: characteristics and satisfaction after handling. Nurs Open 2016. [Google Scholar] [CrossRef] [PubMed]
  14. Rathnayake, D.; Sasame, A.; Radomska, A.; Ni Shé, E.; McAuliffe, E.; De Brún, A. What can we learn from patient and family experiences of open disclosure and how they have been evaluated? A systematic review. BMC Health Serv. Res. 2025, 25. [Google Scholar] [CrossRef] [PubMed]
  15. Eriksen, A.; Tedwall, T.; Larsen, B. Negative experiences with primary care services in Norway expressed in patient and next-of-kin complaints: a qualitative study. BMC Health Serv. Res. 2025, 25. [Google Scholar] [CrossRef] [PubMed]
  16. Ahmad, S.; Wasim, S. Prevent medical errors through artificial intelligence: A review. Saudi J. Med. Pharm. Sci. 2023, 9, 2413–4929. [Google Scholar] [CrossRef]
  17. Mcgrow, K. Artificial intelligence. Essentials for nursing. Nursing 2019, 49, 46–49. [Google Scholar] [PubMed]
  18. Shaheen, M. Application of artificial intelligence (AI) in healthcare: A review. ScienceOpen Preprints. 2021.
  19. Abuzaid, M.; Elshami, W.; McFadden, S. Integration of artificial intelligence into nursing practice. Health Technol. (Berl) 2022, 12, 1109–1115. [Google Scholar] [CrossRef] [PubMed]
  20. Borg, A.; Boldt, M.; Rosander, O.; Ahlstrand, J. E-mail classification with machine learning and word embeddings for improved customer support. Neural Comput Appl. 2020. [Google Scholar] [CrossRef]
  21. Flach, P. Machine learning: The art and science of algorithms that make sense of data; Cambridge University Press: Cambridge, 2012. [Google Scholar]
  22. Robertson, S. Understanding inverse document frequency: on theoretical arguments for IDF. J. Doc. 2004, 60, 503–520. [Google Scholar] [CrossRef]
  23. Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech J. 1948, 27, 379–423. [Google Scholar] [CrossRef]
  24. Malmsten, M.; Börjeson, L.; Haffenden, C. Playing with Words at the National Library of Sweden--Making a Swedish BERT. arXiv 2020. Available online: https://arxiv.org/abs/2007.01658.
  25. Remmer, S.; Lamproudis, A.; Dalianis, H. Multi-label diagnosis classification of Swedish discharge summaries–ICD-10 code assignment using KB-BERT. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), 2021 Sep. pp. 1158–1166. [Google Scholar]
  26. MacQueen, J.B. Some methods for classification and analysis of multivariate observations. In Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability; University of California Press: Berkeley, 1967; Volume 1, pp. 281–297. [Google Scholar]
  27. Quinlan, J.R. Induction of decision trees. Mach. Learn. 1986, 1, 81–106. [Google Scholar] [CrossRef]
  28. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
  29. McCallum, A.; Nigam, K. A comparison of event models for Naive Bayes text classification. In Proceedings of the AAAI-98 Workshop on Learning for Text Categorization, 1998; pp. 41–48. [Google Scholar]
  30. Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef]
  31. Schölkopf, B.; Smola, A.J. Learning with kernels: support vector machines, regularization, optimization, and beyond; MIT Press: Cambridge (MA), 2002. [Google Scholar]
  32. Powers, D.M.W. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. J. Mach. Learn Technol. 2011, 2, 37–63. [Google Scholar]
  33. Sokolova, M.; Lapalme, G. A systematic analysis of performance measures for classification tasks. Inf. Process Manag. 2009, 45, 427–437. [Google Scholar] [CrossRef]
  34. World Medical Association. Declaration of Helsinki: ethical principles for medical research involving human subjects, http:www//www.wma.net/what-we-do/medical-ethics/declaration-of-helsinki/doh-oct200/ (Accessed 23 August 2026).
  35. Du, J.; Chen, Q.; Peng, Y.; Xiang, Y.; Tao, C.; Lu, Z. ML Net: multi-label classification of biomedical texts with deep neural networks. J. Am. Med. Inf. Assoc. 2019, 26, 1279–1285. [Google Scholar] [CrossRef] [PubMed]
  36. Sun, Y.; Gao, D.; Shen, X.; Li, M.; Nan, J.; Zhang, W. Multi-label classification in patient-doctor dialogues with the RoBERTa-wwm-ext + CNN model: Named entity study. JMIR Med. Inform. 2022, 10, e35606. [Google Scholar] [CrossRef] [PubMed]
  37. Li, X.; Shu, Q.; Kong, C.; Wang, J.; Li, G.; Fang, X.; et al. An intelligent system for classifying patient complaints using machine learning and natural language processing: Development and validation study. J. Med. Internet Res. 2025, 27, e55721. [Google Scholar] [CrossRef] [PubMed]
  38. Reader, T.W.; Gillespie, A.; Roberts, J. Patient complaints in healthcare systems: a systematic review and coding taxonomy. BMJ Qual. Saf. 2014, 23, 678–689. [Google Scholar] [CrossRef] [PubMed]
  39. Gillespie, A.; Reader, T.W.; et al. Harnessing patient complaints to systematically monitor healthcare concerns through disproportionality analysis. Int. J. Qual. Health Care 2023. [Google Scholar] [CrossRef] [PubMed]
  40. Skär, L.; Söderberg, S. Patient’s complaints regarding healthcare encounters and communication. Nurs. Open 2018, 6, 499–506. [Google Scholar] [CrossRef] [PubMed]
  41. Zotterman Nygren, A.; Skär, L.; Söderberg, S. Meanings of encounters for close relatives of people with a long-term illness within a primary healthcare setting. Prim. Health Care Res. Dev. 2018, 1–6. [Google Scholar]
  42. Söderberg, S.; Olsson, M.; Skär, L. A hidden kind of suffering: female patients’ complaints to the Patient Advisory Committee. Scand. J. Caring Sci. 2012, 26, 144–150. [Google Scholar] [CrossRef] [PubMed]
  43. Skär, L.; Söderberg, S. Complaints with encounters in healthcare – Men’s experiences. Scand. J. Caring Sci. 2012, 26, 279–286. [Google Scholar] [CrossRef] [PubMed]
  44. Eriksen, A.A.; Fegran, L.; Fredwall, T.E.; et al. Patients’ negative experiences with health-care settings brought to light by formal complaints: a qualitative meta-synthesis. J. Clin. Nurs. 2023, 32, 5816–35. [Google Scholar] [CrossRef] [PubMed]
  45. Springer, C. The three pillars of patient experience: identifying key drivers of patient experience to improve quality in healthcare. J. Public Health (Berl) 2023. [Google Scholar] [CrossRef]
  46. Ruksakulpiwat, S.; Thorngthip, S.; Niyomyart, A.; Benjasirisan, C.; Phianhasin, L.; Aldossary, H.; Ahmed, B.H.; Samai, T. A systematic review of the application of artificial intelligence in nursing care: Where are we, and what’s next? J. Multidiscip. Healthc. 2024, 17, 1603–1616. [Google Scholar] [CrossRef] [PubMed]
  47. The Topol Review. Preparing the healthcare workforce to deliver the digital future. Health Education England. 2019. Available online: https://topol.hee.nhs.uk.
  48. Mennella, C.; Maniscalco, U.; De Pietro, G.; Esposito, M. Ethical and regulatory challenges of AI technologies in healthcare: A narrative review. Heliyon 2024, 10, e26297. [Google Scholar] [CrossRef] [PubMed]
  49. Remmer, S.; Lamproudis, A.; Dalianis, H. Multi-label diagnosis classification of Swedish discharge summaries–ICD-10 code assignment using KB-BERT. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), 2021 Sep. pp. 1158–1166. [Google Scholar]
Figure 1. Number of complaint reports per main category class. Note that large class-imbalance between the categories.
Figure 1. Number of complaint reports per main category class. Note that large class-imbalance between the categories.
Preprints 232465 g001
Figure 2. An overview of the clustering of the content in the patient complaint reports.
Figure 2. An overview of the clustering of the content in the patient complaint reports.
Preprints 232465 g002
Table 1. The evaluated classification algorithms and their hyperparameters.
Table 1. The evaluated classification algorithms and their hyperparameters.
Classification Algorithm Hyperparameters
Decision Tree criterion=’gini’
LinearSVC C=1.0, penalty=l2, tol= 1x10-05, multi_class= ‘ovr’
Multinomial Naïve Bayes alpha=1.0
Random Forest n_estimators=100, min_samples_split=2
Support Vector Classifier (SVC) C=1.0, kernel= ‘rbf’, tol=10-05, multi_class’ =‘ovr’
KB-Bert with LinearSVC dual=True, C=0.001
Table 2. Classification performance for each algorithm over three metrics, including 95% confidence intervals (95% CI) within parenthesis. The highest score per metric in bold font.
Table 2. Classification performance for each algorithm over three metrics, including 95% confidence intervals (95% CI) within parenthesis. The highest score per metric in bold font.
Algorithm Accuracy Jaccard ROC-AUC
Random Forest 0.631 (95% CI 0.530–0.732) 0.469 (95% CI 0.361–0.577) 0.743 (95% CI 0.677–0.809)
LinearSVC 0.690 (95% CI 0.595–0.785) 0.534 (95% CI 0.425–0.643) 0.787 (95% CI 0.727–0.847)
SVC 0.638 (95% CI 0.545–0.731) 0.475 (95% CI 0.378–0.572) 0.749 (95% CI 0.692–0.806)
Decision Tree 0.528 (95% CI 0.451–0.605) 0.362 (95% CI 0.294–0.430) 0.679 (95% CI 0.632–0.726)
Naïve Bayes 0.514 (95% CI 0.425–0.603) 0.350 (95% CI 0.271–0.429) 0.661 (95% CI 0.606–0.716)
KB-BERT &
LinearSVC
0.610 (95% CI 0.515–0.705) 0.446 (95% CI 0.347–0.545) 0.729 (95% CI 0.668–0.790)
Random (baseline) 0.253 (95% CI 0.240–0.266) 0.137 (95% CI 0.132–0.142) 0.502 (95% CI 0.493–0.511)
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.