Submitted:
15 December 2023
Posted:
18 December 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
- Conducted the collection of a dataset comprising 390 pairs of sentences, consisting of sentence suggestions and their corresponding ground truth. Three caregivers meticulously evaluated this dataset and subsequently utilized it as the benchmark ground truth.
- Perform testing and analysis on current evaluation metrics such as BERTScore, Cosine Similarity, ROUGE, and BLEU to assess the quality of sentence suggestions in nursing care applications.
- Introduced an innovative evaluation model featuring a novel approach that significantly improved the performance compared to existing evaluation models, aligning more closely with caregiver assessments.
2. Related Works
2.1. Nursing Care Records
2.2. Sentence Suggestion
2.3. Current Evaluation Metrics
3. Proposed Method for Evaluation Metrics
3.1. Hierarchical Dirichlet Process
- be the set of unique tokens in sentence1,
- be the set of unique tokens in sentence2,
- D be the dictionary formed by combining and ,
- be the Bag of Words (BoW) vector representing sentence1 in the corpus,
- be the BoW vector representing sentence2 in the corpus,
- be the trained Hierarchical Dirichlet Process model with dictionary D and corpus ,
- be the topic distribution vector for sentence1 obtained from the trained HDP model,
- be the topic distribution vector for sentence2 obtained from the trained HDP model.
| Sentence Suggestion | Ground Truth |
|---|---|
| コルセット作ることを報告する | コルセットを作ることを勧められる |
| (report making a corset) | (advised making a corset) |

3.2. Word Embedding
| Word Embedding Model | Strength | Weakness |
|---|---|---|
| Word2Vec | Semantic similarity, efficient, contextual understanding | Out-of-Vocabulary words, limited word level, insensitive to word order |
| GloVe | Global context, effective for common words, linear structure | Less effective for Rare words, limited contextual understanding |
| fasText | Sub-word information, better representation for rare words, efficiency | Computationally more demanding, memory usage |
| Sentence Suggestion | Ground Truth |
|---|---|
| 熱があったので、看護師に報告して中止しました。 | 熱発の為、ナースに報告し中止。 |
| (I had a fever, so I informed the nurse and canceled the session.) | (Due to fever, we informed the nurse and discontinued the treatment.) |

| Sentence Suggestion | Ground Truth |
|---|---|
| 両眼内障であること、右眼は緑内障疑いで眼圧が高くなっては弱い痛み止めを屯用で出しておくので飲んで心臓の状態が良いとの連絡あり | 両眼白内障であること、右眼は緑内障疑いで眼圧が高くなっていること、だから目が見えにくくなっている、と説明を受けられ、眼圧を下げる点眼薬を処方されたこと |
| (I was informed that I have bilateral eye disorders, that my right eye is suspected to have glaucoma, and that my intraocular pressure is high, so they give me a weak painkiller to take, and that my heart is in good condition.) | (He explained to me that he had cataracts in both eyes, that his right eye had high intraocular pressure due to suspected glaucoma, and that he was having difficulty seeing, and was prescribed eye drops to lower the intraocular pressure.) |

3.3. EmbedHDP
- be the set of unique tokens in sentence1,
- be the set of unique tokens in sentence2,
- D be the dictionary formed by a unique set of words in the care record dataset,
- be the tokenization by using Mecab and particles from ,
- be the tokenization by using Mecab and particles from ,
- be the fastText vector from ,
- be the fastText vector from ,
- be the trained Hierarchical Dirichlet Process model with dictionary D and corpus ,
- be the topic distribution vector for sentence1 obtained from the trained HDP model,
- be the topic distribution vector for sentence2 obtained from the trained HDP model.
3.3.1. Tokenization
| Japanese Particles | Explanation and Example |
|---|---|
| は (wa) | The topic particle that indicates the topic or subject of a sentence. For example, "わたし は がくせい です" (Watashi wa gakusei desu) means "I am a student." |
| へ (e) | Indicates direction or destination. For instance, "ともだち へ いきます" (Tomodachi e ikimasu) means "I am going to a friend." |
| で (de) | Indicates the place or method in which an action takes place. For example, "レストラン で たべます" (Resutoran de tabemasu) means "I eat at the restaurant." |
| を (wo) | The object particle that indicates the object of an action. For example, "りんご を たべます" (Ringo o tabemasu) means "I eat an apple." |
| の (no) | The possessive particle or connector between two nouns. For example, "わたし の くるま" (Watashi no kuruma) means "My car." |
| ある (aru) | A verb indicating existence or possession. For example, "ほん が あります" (Hon ga arimasu) means "There is a book." |
| あり (ari) | The past or formal form of the verb "ある" (aru) indicating existence. |
| する (suru) | A common verb meaning "to do." For example, "しゅくだい を する" (Shukudai o suru) means "To do homework." |
| なる (naru) | A verb meaning "to become." For example, "せんせい に なりたい" (Sensei ni naritai) means "I want to become a teacher." |
| し (shi) | A conjunction used to express two related actions or qualities. For example, "りんご し いちご" (Ringo shi ichigo) means "Apples and strawberries." |
| て (te) | The te-form of a verb, indicating an ongoing action. For example, "たべ て います" (Tabete imasu) means "I am eating." |
| ます (masu) | A polite form of verbs indicating present actions. For example, "たべ ます" (Tabemasu) means "I eat" or "I will eat." |
3.3.2. Creating Corpus
- Set Vector Length: Assign a fixed vector length in the BoW format, specifically 10. This decision is grounded in the consideration that our sentences are not excessively long, thereby mitigating potential biases arising from vector length discrepancies.
- Highest Frequency Elements: Select elements based on their highest frequencies, under the assumption that the highest frequency serves as a representative token for each element.
- Scaling Factor: Due to the considerable length of vectors produced by fastText, the resultant BoW-formatted vectors become exceedingly small (0.000x). This phenomenon leads to nearly identical topics when trained in the HDP model. To counteract this issue, each vector is multiplied by 100, ensuring positive values throughout the vector and resolving the disparity.
- BoW Representation: The outcome of these steps is the acquisition of BoW-formatted vector representations for each token in both sentences. This transformation facilitates seamless compatibility with the HDP model during the training process.
4. Evaluation
4.1. Filtering Data Sample
4.2. Evaluation Metrics Utilized
4.3. Benchmarking Method
- The example of BERTScore limitations when evaluating short sentences (diverse sentence structures) and those containing medical information (specialized medical terminology).Table 10. An example of BERTScore limitation
Sentence Suggestion Ground Truth 頻繁な少量の排尿。 排便中量あり。 (frequent small amount of urination) (There was a large amount during defecation.) Figure 5. Metrics and human evaluation assessment of BERTScore limitation sentence.
- The example of Cosine Similarity limitations when evaluating short sentences (diverse sentence structures) and those containing medical information (specialized medical terminology).Table 11. An example of Cosine Similarity limitation
Sentence Suggestion Ground Truth 吐き気あり報告入れる。 吐き気訴えあり。 (report nurse) (complaints of nurse) Figure 6. Metrics and human evaluation assessment of Cosine Similarity limitation sentence.
- The example of ROUGE limitations when evaluating different structures of sentences (diverse sentence structures).Table 12. An example of ROUGE limitation
Sentence Suggestion Ground Truth 気分訴えなし 気分不良はないと本人言われる (no mood complaints) (he says he doesn’t feel unwell) Figure 7. Metrics and human evaluation assessment of ROUGE limitation sentence.
- The example of BLEU limitations when evaluating different structures of sentences (diverse sentence structures) and those containing medical information (specialized medical terminology).Table 13. An example of BLEU limitation
Sentence Suggestion Ground Truth 熱があったので、看護師に報告して中止しました。 熱発の為、ナースに報告し中止。 (I had a fever, so I informed the nurse and cancelled the session) (Due to fever, the nurse was informed and the procedure was discontinued) Figure 8. Metrics and human evaluation assessment of ROUGE limitation sentence.
5. Discussion
6. Conclusions
References
- Yang, T.; Deng, H. Intelligent sentence completion based on global context dependent recurrent neural network language model. Proceedings of the International Conference on Artificial Intelligence, Information Processing and Cloud Computing, 2019, pp. 1–5.
- Park, H.; Park, J. Assessment of word-level neural language models for sentence completion. Applied Sciences 2020, 10, 1340. [CrossRef]
- Asnani, K.; Vaz, D.; PrabhuDesai, T.; Borgikar, S.; Bisht, M.; Bhosale, S.; Balaji, N. Sentence completion using text prediction systems. Proceedings of the 3rd International Conference on Frontiers of Intelligent Computing: Theory and Applications (FICTA) 2014: Volume 1. Springer, 2015, pp. 397–404.
- Mamom, J. Digital technology: innovation for malnutrition prevention among bedridden elderly patients receiving home-based palliative care. Journal of Hunan University Natural Sciences 2020, 47.
- Walonoski, J.; Kramer, M.; Nichols, J.; Quina, A.; Moesel, C.; Hall, D.; Duffett, C.; Dube, K.; Gallagher, T.; McLachlan, S. Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record. Journal of the American Medical Informatics Association 2018, 25, 230–238. [CrossRef]
- Churruca, K.; Ludlow, K.; Wu, W.; Gibbons, K.; Nguyen, H.M.; Ellis, L.A.; Braithwaite, J. A scoping review of Q-methodology in healthcare research. BMC medical research methodology 2021, 21, 125. [CrossRef]
- of General Practicioners, R.A.C.; others. Privacy and managing health information in general practice, 2021.
- Shibatani, M.; Miyagawa, S.; Noda, H. Handbook of Japanese syntax; Vol. 4, Walter de Gruyter GmbH & Co KG, 2017.
- Caballero Barajas, K.L.; Akella, R. Dynamically modeling patient’s health state from electronic medical records: A time series approach. Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2015, pp. 69–78.
- Evans, R.S. Electronic health records: then, now, and in the future. Yearbook of medical informatics 2016, 25, S48–S61. [CrossRef]
- Mairittha, N.; Mairittha, T.; Inoue, S. A mobile app for nursing activity recognition. Proceedings of the 2018 ACM international joint conference and 2018 international symposium on pervasive and ubiquitous computing and wearable computers, 2018, pp. 400–403.
- Stevens, S.; Pickering, D. Keeping good nursing records: a guide. Community eye health 2010, 23, 44.
- Mirowski, P.; Vlachos, A. Dependency recurrent neural language models for sentence completion. arXiv preprint arXiv:1507.01193 2015.
- Irie, K.; Lei, Z.; Deng, L.; Schlüter, R.; Ney, H. Investigation on estimation of sentence probability by combining forward, backward and bi-directional LSTM-RNNs. INTERSPEECH, 2018, pp. 392–395.
- Rakib, O.F.; Akter, S.; Khan, M.A.; Das, A.K.; Habibullah, K.M. Bangla word prediction and sentence completion using GRU: an extended version of RNN on N-gram language model. 2019 International Conference on Sustainable Technologies for Industry 4.0 (STI). IEEE, 2019, pp. 1–6.
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.Q.; Artzi, Y. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 2019.
- Rahutomo, F.; Kitasuka, T.; Aritsugi, M. Semantic cosine similarity. The 7th international student conference on advanced science and technology ICAST, 2012, Vol. 4, p. 1.
- Schluter, N. The limits of automatic summarisation according to rouge. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, 2017, pp. 41–45.
- Reiter, E. A structured review of the validity of BLEU. Computational Linguistics 2018, 44, 393–401.
- Kherwa, P.; Bansal, P. Topic modeling: a comprehensive review. EAI Endorsed transactions on scalable information systems 2019, 7. [CrossRef]
- Kingman, J. Theory of Probability, a Critical Introductory Treatment, 1975.
- Goldberg, Y.; Levy, O. word2vec Explained: deriving Mikolov et al.’s negative-sampling word-embedding method. arXiv preprint arXiv:1402.3722 2014.
- Pennington, J.; Socher, R.; Manning, C.D. Glove: Global vectors for word representation. Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, pp. 1532–1543.
- Joulin, A.; Grave, E.; Bojanowski, P.; Douze, M.; Jégou, H.; Mikolov, T. Fasttext. zip: Compressing text classification models. arXiv preprint arXiv:1612.03651 2016.
- Hill, F.; Cho, K.; Jean, S.; Devin, C.; Bengio, Y. Embedding word similarity with neural machine translation. arXiv preprint arXiv:1412.6448 2014.
- Tulu, C.N. Experimental Comparison of Pre-Trained Word Embedding Vectors of Word2Vec, Glove, FastText for Word Level Semantic Text Similarity Measurement in Turkish. Advances in Science and Technology. Research Journal 2022, 16. [CrossRef]
- Sitender.; Sangeeta.; Sushma, N.S.; Sharma, S.K. Effect of GloVe, Word2Vec and FastText Embedding on English and Hindi Neural Machine Translation Systems. In Proceedings of Data Analytics and Management: ICDAM 2022; Springer, 2023; pp. 433–447.
- Dharma, E.M.; Gaol, F.L.; Warnars, H.; Soewito, B. The accuracy comparison among word2vec, glove, and fasttext towards convolution neural network (cnn) text classification. J Theor Appl Inf Technol 2022, 100, 31. [CrossRef]
- Kubo, M. Japanese syntactic structures and their constructional meanings. PhD thesis, Massachusetts Institute of Technology, 1992.
- Shimomura, Y.; Kawabe, H.; Nambo, H.; Seto, S. The translation system from Japanese into braille by using MeCab. Proceedings of the Twelfth International Conference on Management Science and Engineering Management. Springer, 2019, pp. 1125–1134.

| Activity type | Record type |
|---|---|
| 1.バイタル(vitals), 2.リハビリ・レク(rehabilitationrecreation), 3.往診・受診(house callsvisit), 4.処置(treatment), 5.入浴・清拭(bathing/cleaning), 6.外出対応(going out), 7.活力朝礼・ラジオ体操 (vitality morning/radio exercise), 8.特記事項・連絡事項(special notes/notifications), 9.送迎(transportation), 10.事故等緊急対応(emergency response), 11.就寝前食事(meal before bedtime), 12.モーニングケア(morning care), 13.ナイトケア(night care), 14.その他食事(other meals), 15.家族・来客対応(family/visitor support), 16.家族・医師連絡(family/doctor contact), 17.手書き記録(handwritten records), 18.入院(hospitalization), 19.離床・臥床介助 (assistance with getting out of bed and lying down), 20.食事・服薬(meals and medication), 21.おやつ(snacks), 22.更衣介助(assistance with changing clothes), 23.口腔ケア(oral care), 24.排泄(excretion), 25.日中利用者対応(support for daytime user), 26.夜間利用者対応(support for nighttime user), 27.朝食(breakfast), 28.昼食(lunch), 29.夕食(dinner), 30.洗面介助(washing assistance), 31外泊(overnight stay) |
1.特記事項(spacial notes), 2.状態・特記事項 (condition/special notes), 3.連絡事項(notifications), 4.傷の状態・特記事項 (status/special notes) |
| Evaluation Metrics | Mechanism | Limitation |
|---|---|---|
| BERTScore [16] | Comparing contextual embedding of reference and candidate sentences using pre-trained BERT models | It relies on pre-trained BERT models, which may not capture domain-specific nuances effectively |
| Cosine Similarity [17] | Calculates the cosine angle between two vectors to determine their similarity, frequently used to compare text documents in vector space. | Fails to account for word order, and mistakenly rate semantically difference with similar word sets. |
| ROUGE [18] | Measures the overlap of n-grams and the longest matching sequence between a generated summary and reference texts. | Might overlook semantic accuracy as it is based on lexical overlap, not considering the context or meaning of the words. |
| BLEU [19] | Scores machine translations by matching n-grams to reference texts and adjusting for translation length. | Can miss the adequacy and fluency of translation as it primarily relies on n-gram overlap, ignoring semantic coherence. |
| Evaluation Metrics | Correlation Coefficient |
|---|---|
| EmbedHDP | 0.61 |
| BERTScore | 0.58 |
| ROUGE | 0.57 |
| Cosine Similarity | 0.59 |
| BLEU | 0.53 |
| Evaluation Metrics | Correlation Coefficient |
|---|---|
| EmbedHDP | 0.25 |
| BERTScore | 0.35 |
| ROUGE | 0.34 |
| Cosine Similarity | 0.35 |
| BLEU | 0.34 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).