Preprint
Article

This version is not peer-reviewed.

A Hybrid Stanza–FP-Growth Approach to Terminological Pattern Recognition in Bibliometric Records on Energy Security Indicators

Submitted:

30 July 2026

Posted:

31 July 2026

You are already at the latest version

Abstract
Analyzing the thematic content of bibliometric records through titles and abstracts is essential for understanding how research topics are defined, articulated, and transformed over time. As these publication components concentrate the most visible and representative terminology, they provide a compact yet informative basis for identifying recurring concepts, emerging terms, and shifts in thematic emphasis across a research corpus. This study aimed to evaluate the utility of a hybrid Stanza–FP-Growth approach for terminological pattern recognition in ScienceDirect bibliometric records addressing energy security indicators. The dataset comprised 830 bibliometric records retrieved from the open ScienceDirect abstract database using the search query “Energy security” AND (indicators OR indication OR metrics OR index) within titles, abstracts, and keywords, with the document types limited to review and research articles. The extraction procedure identified well-established scientific terms, including renewable energy source, energy security indicator, national energy security, cleaner energy system, energy security, technological trajectory, clean technology transfer, lithium battery, and sustainable innovation. More specific research themes were characterized by the co-occurrence of multiple key terms, including the term sets energy trilemma; environmental sustainability; energy security and energy trilemma; energy equity; environmental sustainability. Notably, although energy trilemma was not included in the original ScienceDirect search query, it emerged naturally through the proposed extraction pipeline. The findings demonstrate that the Stanza–FP-Growth framework constitutes a viable proof-of-concept approach for structuring bibliometric records and transforming unstructured textual data into a systematic terminological representation. The proposed approach may support domain experts in identifying promising research directions and emerging thematic relationships. Future research should explore the use of Stanza’s default_accurate module for named entity recognition and implement domain-specific recalibration of stop-word and exception dictionaries to improve extraction fidelity.
Keywords: 
;  ;  ;  

Introduction

Analyzing the thematic content of bibliometric records through titles and abstracts is an important step in understanding how research topics are defined, articulated, and transformed over time. Because titles and abstracts concentrate the most visible terminology of a publication [1], they provide a compact yet informative basis for identifying recurring concepts, emerging terms, and shifts in thematic emphasis across a corpus [2,3,4]. Such analysis is especially valuable for tracing topic evolution and for detecting changes in the vocabulary used to describe a field. It is therefore particularly relevant in the early stages of analytical and systematic reviews, where an accurate mapping of the topic space helps define the scope of the review, identify central subthemes, and reveal underexplored directions [5,6].
In this study, we argue that term extraction for bibliometric analysis should be guided not only by surface statistics or semantic similarity, but also by syntactic structure. Methods such as Yake! [7] are effective when the goal is to identify salient phrases without training data, as they rely on features such as term frequency, word position, capitalization, lexical context, and distributional patterns. This makes them suitable for detecting compact and informative expressions such as “climate change,” “renewable energy,” or “machine learning.” However, in more complex domains, they may fail to capture the internal composition of multiword technical expressions, where the meaning of the candidate term depends on its grammatical structure rather than on frequency alone [8,9,10].
Embedding-based approaches [11] address a different problem: they are designed to capture semantic proximity between phrases. While this is useful for identifying related expressions, it does not necessarily distinguish true terminological units from merely similar concepts [12,13]. As a result, embeddings may overgeneralize phrase boundaries or group together expressions that are semantically close but not equivalent as keywords in a given article. For example, phrases such as “energy security,” “energy supply,” and “energy transition” may be closely related in vector space, yet only some of them may be valid terminological candidates in a specific context.
By contrast, the Stanza-based pipeline [14,15] used here introduces a strong structural constraint: a candidate term must first satisfy a noun-phrase–like syntactic pattern. Through tokenization, part-of-speech tagging, constituency parsing [16,17], lemmatization, and filtering, the method narrows the search space to linguistically plausible units. This is especially advantageous for bibliometric applications, where the objective is to build term inventories derived from titles and abstracts, recover missing keywords, and generate candidates that remain interpretable across disciplines. Because syntactic organization is generally more transferable across domains than frequency-based cues, this approach is well suited to heterogeneous corpora spanning energy, oil and gas, materials science, and related fields.
Methodologically, the approach adopted in this study is related to the text-mining workflow implemented in VOSviewer, but it is more flexible and better suited to the extraction of domain-specific terminology. When VOSviewer constructs maps from textual data, it relies on part-of-speech tagging based on Apache OpenNLP [18,19]. According to the user manual, the software extracts only the longest noun phrases within a sentence and ignores shorter phrases embedded within them. This design, developed by the Centre for Science and Technology Studies (CWTS) at Leiden University, has become a widely used standard in bibliometric visualization. In VOSviewer, however, these procedures are built in and applied in a fixed manner.
By contrast, the workflow developed in this study provides greater control over the extraction process. It uses Stanza, a more accurate NLP library, and allows the researcher to customize stop words, control noun-phrase length, and preserve canonical forms. An additional distinction concerns the unit of analysis: whereas VOSviewer processes text sentence by sentence, our algorithm treats the title and abstract as a single textual unit. This makes the method more suitable for capturing multiword terms that may be distributed across sentence boundaries and for producing a more coherent term inventory for bibliometric analysis.
The primary objective of this study was to evaluate the utility of a hybrid Stanza–FP-Growth approach for terminological pattern recognition within ScienceDirect bibliometric data (titles and abstracts) addressing energy security indicators.

Materials and Methods

The source of bibliometric records was the open abstract database ScienceDirect, an information platform of Elsevier. This database was selected because it contains data from peer-reviewed publications and includes well-populated fields for Title, Abstract, and Author Keywords, which are most commonly used to assess the subject matter of articles and to conduct article searches.
The database query was performed by searching within Title, abstract, and keywords for: "Energy security" AND (indicators OR indication OR metrics OR index) and Article type: Review articles, Research articles. No year filter was applied. The earliest work matching this query and indexed in ScienceDirect was published in 2003. The query was constructed with the aim of remaining as close as possible to the stated topic of "Energy security indicators" and consisted of an exact match for the primary term "Energy security" and several of the most frequently encountered synonyms for the term "indicator."
A total of 830 bibliometric records were exported in RIS format. For the convenience of further data processing, these were converted to CSV format using the Zotero program [20].
Sequence of noun-phrase extraction and result filtering employed in this study:
  • Exhaustive extraction: All noun phrase (NP) nodes are exhaustively extracted from the constituency parse tree, including nested and overlapping subconstituents. The traversal order ensures that no nested elements are omitted.
  • Linguistic cleaning (normalization): For each token within the extracted groups, the following procedure is applied:
    • o Function words are removed entirely based on their part-of-speech tags: DT, PDT, WDT, PRP, PRP$.
    • o Punctuation marks are removed, and mathematical or chemical special symbols are isolated.
    • o All dash variants (–, —) are normalized to the standard hyphen (-).
    • o Remaining words are converted to lowercase and reduced to their canonical form (lemma).
    • o Lemmas are filtered using a custom stop-word list (stopwords.txt), which includes academic stop words and prepositions.
    • o Isolated digits are discarded, whereas alphanumeric terms such as GPT-4 or CO₂ are preserved.
  • Length criterion (post-cleaning): The number of meaningful lemmas in each group is counted strictly after Step 2. Any group whose final length falls outside the range of 2 to 3 words is discarded.
  • Syntactic nesting resolution (string demerging): If, after cleaning and length filtering, group A constitutes a string subpart of a longer group B (with word-boundary consideration, \b), group A is removed and only the maximal-length context (group B) is retained. However, if group B was discarded in Step 3 due to exceeding the 3-word limit, group A is preserved.
  • Deduplication and output: The resulting groups are deduplicated within each individual annotation, while the original order of mention is maintained. The output is written to a file with groups separated by semicolons (;), preserving the line structure (one line per set of title and abstract terms).
POS tags (parts of speech):
  • DT — determiner (the, a, an)
  • PDT — predeterminer (all, both, half)
  • WDT — wh-determiner (which, what)
  • PRP — personal pronoun (I, you, he, we)
  • PRP$ — possessive pronoun (my, your, his, our)
All are function words and are removed during cleaning.
A similar procedure was employed to identify noun phrases consisting of four to five words.
Accurate natural language processing was carried out using the Stanza library with constituency parsing (constituency tree) [14].
Frequent pattern mining was performed using Christian Borgelt’s fpgrowth implementation [21] to extract term combinations constraint by a specific range of itemset lengths and a minimum support level.

Results and Discussion

Main Characteristics of the Bibliometric Records Used in the Study

All DOI, Title, and Abstract fields in the records were fully populated; only 11 of the 830 records had an empty Author Keywords field. Thus, the data used to identify publication topics were of high quality.
The list of journals in which 10 or more articles on the stated topic were published is presented in Table 1.
ScienceDirect contains full text articles from journals and books, primarily published by Elsevier, but including some hosted societies.
The RIS format does not contain a Subject areas field; however, subject area distribution is available on the ScienceDirect platform itself. The obtained results are as follows: Subject areas (Publications count): Energy (466), Environmental Science (159), Engineering (129), Social Sciences (88), Economics, Econometrics and Finance (74), Agricultural and Biological Sciences (52), Materials Science (39), Earth and Planetary Sciences (24), Chemistry (22), Decision Sciences (21). Thus, the subject matter adequately reflects the core dimensions of "Energy security indicators" — Availability, or elements relating to geological existence; Accessibility, or geopolitical elements; Affordability, or economical elements; Acceptability, or environmental and societal elements [22]. Further references on the topic can be found in the publication [23].
From the distribution of publications by year: 2027 (2), 2026 (121), 2025 (148), 2024 (103), 2023 (59), 2022 (51), 2021 (53), 2020 (42), 2019 (37), 2018 (36), 2017 (30), 2016 (25), 2015 (22), 2014 (29), 2013 (22), 2012 (15), 2011 (13), 2010 (7), 2009 (7), 2008 (4), 2007 (2), 2005 (1), 2003 (1), an increase in interest in the topic is observable from 2024 onward.
The occurrence of the primary term "Energy security" in the "Title" and "Abstract" fields — 1827 results in 739 records.
The occurrence of the primary term "Energy security" in the "Author keywords" field — 347 results in 320 records.
The occurrence of the primary term "Energy security" across all fields "Title", "Abstract" and "Author keywords" — 2174 results in 799 records. The absence of a direct match for the term "Energy security" in 31 records is partially explained by the occurrence of the term "energy-security", as well as the substrings "energy supply security" and "energy, security". At the same time, the term "indicator" occurs 163 times in 110 records, "indication" — 7 and 7, "metrics" — 140/107, and "index" — 675/346. The term "index" is fairly general in usage and therefore occurs more frequently.
Overall, it can be stated that the subject matter of the exported records is relevant to the research topic under investigation.

Noun Phrase Extraction from Titles and Abstracts

A key challenge in noun phrase (NP) extraction is generally not the number of extracted candidates but achieving an appropriate balance between precision and recall.
A substantial proportion of the extracted terms represent meaningful scientific concepts, for example: cleaner energy system, energy carrier, energy security, technological trajectory, clean technology transfer, lithium battery, sustainable innovation, clean transport system, renewable energy deployment, institutional stress, cross-sectional dependence, structural institutional fragility, energy supply susceptibility, energy production, energy sector, energy reform, and emerging economy.
One of the main issues encountered concerns noun lemmatization. Typical examples include unite state, mark difference, datum, emerge economy, and inventor mobility. These forms do not represent errors in NP extraction but rather reflect the behavior of the Stanza lemmatizer. Consequently, further refinement of the methodology requires the development of a dedicated exception list for stable multi-word expressions. Nevertheless, forms such as unite state do not substantially affect the interpretation of the publication's subject matter.
The initial extraction experiments also revealed incorrect lemmatization and processing of abbreviations. Representative examples include energy generation eg, fiscal deficit fd, transmission deficit td, and fd td pgdp. In subsequent processing, abbreviations such as FD, TD, EG, and PGDP were removed because they consist of short alphabetic tokens. In most abstracts, these abbreviations are explicitly defined; therefore, retaining them introduces redundant information. Since the objective of the present study was to identify the unique terminology within each bibliographic record rather than to analyze term frequencies, removing such abbreviations was considered appropriate. One practical approach is to remove alphabetic tokens consisting of one to three characters. Further methodological improvements should include expanding the list of immutable tokens to incorporate not only proper names but also commonly used abbreviations such as AI, ML, NLP, IoT, and GIS. Alternatively, longer abbreviations may be incorporated into the stop-word list when their expanded forms are already present in the text (e.g., PGDP when GDP per capita is explicitly provided).
It should be emphasized that such preprocessing decisions are not universally applicable and may depend on both the subject domain of the publications and the presence of abbreviation definitions within the abstracts. However, the objective of the present study was not to perform a comprehensive evaluation of stop-word selection, abbreviation expansion, or immutable token dictionaries. Rather, the aim was to identify these methodological issues as directions for future improvement while focusing on the dominant research topics within the field of energy security based on the textual content of article titles and abstracts.
In contrast, alphanumeric abbreviations should generally be retained, as they represent meaningful scientific concepts. Examples include GPT-5, CO₂, and H₂O.
The stop-word list also requires continuous refinement for the specific domain under investigation. For example, the word more was not initially included among the stop words, although phrases such as more fiscal discipline frequently occurred in the corpus. In such cases, retaining only fiscal discipline provides a more informative representation of the underlying concept.
Typical examples of phrase demerging include reducing energy security country to energy security and structural institutional fragility to institutional fragility.
Another limitation observed in the extracted results is the presence of grammatically correct but excessively long noun phrases that contribute little semantic value, such as organizational inventor mobility, technological class subclass, mobility inventor organization, and institutional geopolitical vulnerability.
One of the reasons for selecting Stanza was that, although spaCy provides substantially higher processing speed, preliminary experiments not included in the present study demonstrated that both spaCy and Stanza, when used without constituent parsing, generate a considerably larger number of such low-information noun phrases. The methodology proposed in this study attempts to reduce these artificial phrases by analyzing the co-occurrence patterns of the extracted noun phrases.
Maintaining an appropriate balance during term filtering is essential, since overly aggressive filtering may eliminate scientifically meaningful noun phrases. This objective differs from conventional keyword extraction, where relatively short terms are typically considered as candidate keywords. Instead, the present study aims to construct a terminological space from article titles and abstracts that extends beyond author-provided keywords while remaining grounded in the terminology actually employed by the authors.
To investigate the extracted terminology in greater detail, two complementary analyses were performed: one with noun phrases limited to two or three words and another allowing phrases of two to five words.
The two- to three-word limit more closely resembles conventional keyword extraction. Such results may serve as an analogue of the Index Keywords field in Scopus and can subsequently be employed in bibliometric visualization software such as VOSviewer. This approach is particularly useful when bibliographic records are obtained from sources such as OnePetro, where keyword fields are unavailable, whereas abstracts are typically highly detailed.
The two- to five-word limit serves a different purpose. Longer noun phrases represent stable scientific expressions rather than merely candidate keywords. Examples include renewable energy deployment strategy, greenhouse gas emission reduction, and machine learning based approach. Furthermore, a phrase such as impact cost estimate ghg emission abatement may be reduced under the two- to three-word constraint to separate phrases such as cost estimate and GHG emission abatement, thereby losing the semantic relationship between these concepts. Allowing longer noun phrases substantially increases the number of possible term combinations and consequently the number of extracted candidates, including phrases with limited thematic significance. Therefore, subsequent ranking according to an appropriate measure of importance becomes even more critical than in the case of shorter noun phrases.
Accordingly, the objective of the present study is closer to the extraction of stable scientific concepts that publication authors use as domain-specific descriptors. Such terminology facilitates literature retrieval and supports the preparation of systematic literature reviews, whose initial stage is commonly based on bibliometric analysis.

Analysis of Two- to Three-Word Noun Phrases

Noun phrases consisting of two or three words were extracted from the titles and abstracts of each of the 830 bibliographic records. The most frequently occurring terms are presented in Table 2.
The results indicate that two-word noun phrases clearly dominate among the extracted terms. At the same time, the three-word noun phrases remain readily interpretable and correspond to well-established scientific terminology, for example, renewable energy source, energy security indicator, and national energy security. The phrase energy security indicator represents the principal concept reflected in the ScienceDirect search query. Deduplication and the removal of noun phrases contained within longer phrases were performed at the level of individual bibliographic records. Consequently, Table 2 includes both energy security and its longer variants, such as energy security indicator and national energy security.
Irrelevant or artificial terms occur only rarely. For example, artificial intelligence ai appears only three times. Eliminating such cases would require the addition of another preprocessing step involving the construction of a dictionary for expanding frequently occurring abbreviations. Even seemingly unusual phrases that occur only once, such as availability affordability acceptability, nevertheless represent meaningful concepts related to energy security indicators rather than extraction errors.
Overall, the results suggest that noun phrases consisting of two or three words constitute suitable candidates for keywords describing the subject area under investigation.

Frequent Co-Occurring Sets of Two- to Three-Word Noun Phrases

A more specific research topic can be characterized by a combination of several key terms, which also facilitates the retrieval of publications relevant to a particular research question. Frequent sets consisting of three to five noun phrases were identified using the FP-Growth algorithm.
The frequent term sets obtained with a minimum support threshold of 0.4% are presented in Table 3.
The highest-ranked term set, energy equity; environmental sustainability; energy security, does not explicitly contain the term energy trilemma but nevertheless represents its three fundamental dimensions. This interpretation is supported by the occurrence of related term sets, including energy trilemma; environmental sustainability; energy security and energy trilemma; energy equity; environmental sustainability. Although energy trilemma was not included in the original ScienceDirect search query, it emerged naturally as a result of the proposed extraction pipeline.
Interestingly, the term energy trilemma does not appear among the most frequent individual noun phrases listed in Table 2. In contrast, environmental sustainability occurs 46 times, whereas energy equity appears only 17 times, indicating a substantial imbalance in the representation of these concepts. This observation suggests that, within publications addressing energy security indicators, the topic of energy equity receives considerably less attention than environmental sustainability.
From a methodological perspective, the results presented in Table 3 demonstrate the importance of analyzing the co-occurrence of multiple terms rather than considering individual terms alone. In particular, frequent sets comprising three or more co-occurring noun phrases capture higher-order semantic relationships that cannot be identified through pairwise term associations.

Analysis of Four- to Five-Word Noun Phrases

In this analysis, the same noun phrase extraction procedure was applied to the titles and abstracts of the bibliographic records. The only difference was that the extracted noun phrases were restricted to four or five words. The resulting most frequent terms are presented in Table 4.
Most of the phrases listed in Table 4 are presented in their lemmatized form after stop-word removal and can be interpreted as combinations of two shorter noun phrases. For example, energy security sustainable development, energy security environmental sustainability, and energy security energy equity collectively reflect the topic of the energy trilemma. Unlike the frequent term sets presented in Table 3, however, the explicit term energy trilemma does not appear among the extracted four- and five-word noun phrases. Some entries are simply reordered versions of the same concepts, for example, environmental sustainability energy security.
The phrase geographic information system gis appears as a consequence of the current pipeline's limited handling of abbreviation expansion. In contrast, several extracted phrases represent complete four-word scientific terms, including battery energy storage system and method moment quantile regression, corresponding to the Method of Moments Quantile Regression. Another noteworthy example is availability applicability acceptability affordability, which extends the commonly used set of concepts associated with energy security indicators by introducing the additional dimension of applicability.
These results suggest that extracting longer noun phrases can reveal stable scientific expressions that are not adequately represented by the two- or three-word terms commonly used as keywords. At the same time, many longer noun phrases can naturally be interpreted as combinations of two or more shorter noun phrases, indicating that both levels of representation provide complementary information for characterizing the terminology of the analyzed publications.

Conclusion

This study demonstrates that extracting noun phrases (NPs) from titles and abstracts enables the automated reconstruction of domain terminology without relying on author or indexer keywords. While traditional bibliometrics relies on author keywords, index keywords, and citation networks, this work introduces a fourth paradigm: corpus-derived terminology. Unlike metadata restricted by author subjectivity or standardized taxonomies, extracted NPs reflect the actual language of the scientific text.
The proposed workflow integrates isolated methods—NP extraction, filtering, and the FP-Growth algorithm—into a unified methodology that builds a clear knowledge representation hierarchy:
  • Core Descriptors (2–3 words): Form domain-specific terms comparable to conventional keywords.
  • Thematic Structures: Co-occurrence analysis reveals latent concepts and relations invisible through individual term frequencies.
  • Specific Concepts (4–5 words): Capture stable scientific expressions and complex research methodologies.
  • Terminological Space: The synthesis of these levels maps the entire corpus knowledge structure.
Ultimately, this framework serves as a proven proof-of-concept approach for structuring bibliometric records. By transforming raw text into a systematic terminological representation, it provides domain experts with a robust tool for identifying promising research directions.
Subsequent research would benefit from utilizing Stanza's default_accurate module for named entity recognition, alongside a domain-specific recalibration of stop-word and exception dictionaries to improve extraction fidelity.

Information about the author:

Boris N. Chigarev – Cand. Sci. (Phys.-Math.), Senior Researcher, Oil and Gas Research Institute, Russian Academy of Sciences, Moscow, Russia; https://orcid.org/0000-0001-9903-2800; e-mail: bchigarev@ipng.ru.

Funding

the work was funded by the Ministry of Science and Higher Education of the Russian Federation (State Assignment No. 125021302095-2).

References

  1. Mammino, L. Scientific terminology: a long thread of interactions between humanities and sciences. In Insights into the Relationships between Humanities and Sciences; Mammino, L., Ed.; Springer Nature Switzerland, 2026; pp. 155–175. [Google Scholar] [CrossRef]
  2. Donald, W.E. Content analysis of metadata, titles, and abstracts (Camta): application of the method to business and management research. MRR 2022, 45(1), 47–64. [Google Scholar] [CrossRef]
  3. Collins, W.; Salman, A.; Olbina, S.; Mosier, R. Scientometric, thematic, and methodological analysis of ijcer construction education focused publications – 2004 - 2023. Int. J. Constr. Educ. Res. 2024, 20(4), 383–404. [Google Scholar] [CrossRef]
  4. Chigarev, B.N. Ieee terms analysis of 2019-2024 ieee xplore data on the topic of energy systems. Energy Syst. Res. 2024, 7(2(26)), 26–38. [Google Scholar] [CrossRef]
  5. Torres, E.M.; Abrigo, C.M. Mapping cultural competence in library and information science research: A scoping review and bibliometric overview. IFLA Journal. 2025, 51(4), 1180–1193. [Google Scholar] [CrossRef]
  6. Abdallah, A.S.; Amin, H.; Abdelghany, M.; Elamer, A.A. Antecedents and consequences of intellectual capital: a systematic review, integrated framework, and agenda for future research. Manag Rev. Q. 2025, 75(4), 3067–3118. [Google Scholar] [CrossRef]
  7. Campos, R.; Mangaravite, V.; Pasquali, A.; Jorge, A.; Nunes, C.; Jatowt, A. YAKE! Keyword extraction from single documents using multiple local features. Inf. Sci. 2020, 509, 257–289. [Google Scholar] [CrossRef]
  8. Mu, Y.; Wang, J.; Zhang, H.; Gan, Z.; Zhu, G.N. Towards efficient patent analysis: A large language model and BERT-refined methodology for keyphrase extraction. World Pat. Inf. 2026, 84, 102435. [Google Scholar] [CrossRef]
  9. Jayasudha, J.; Thilagu, M. Keywords extraction in learning management system (Lms) using natural language processing(Nlp). In Advances in Health Informatics, Intelligent Systems, and Networking Technologies; Jeyabose, A., Balas, V.E., Fernandes, S.L., Eds.; Springer Nature Singapore, 2025; Vol 1286, pp. 643–654. [Google Scholar] [CrossRef]
  10. Zampaligre, R.; Sabane, A.; Kafando, R.; Kabore, A.K.; Bissyande, T.F. Text mining for thematic keyword extraction: enriching a french lexicon on food security. In Towards New E-Infrastructure and e-Services for Developing Countries; Koné, T., Sere, A., Kouamé, K.F., Eds.; Springer Nature Switzerland, 2025; Vol 653, pp. 319–333. [Google Scholar] [CrossRef]
  11. Gardazi, N.M.; Daud, A.; Malik, M.K.; Bukhari, A.; Alsahfi, T.; Alshemaimri, B. BERT applications in natural language processing: a review. Artif. Intell. Rev. 2025, 58(6), 166. [Google Scholar] [CrossRef]
  12. Makwana, M.; Mehta, R. Keyphrase-based literature recommendation: enhancing user queries with hybrid co-citation and co-occurrence networks. J. Scientometr. Res. 2024, 13(1), 217–229. [Google Scholar] [CrossRef]
  13. Kamila, A.R.; Derhass, G.H.; Surianto, S.; Budiyanto, V. Evaluation of keyword extraction using yake and keybert in text preprocessing for hoax news detection based on bi-lstm. REJHH 2025, 8(3), 4490–4500. [Google Scholar] [CrossRef]
  14. Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; Manning, C.D. Stanza: a python natural language processing toolkit for many human languages. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations Association for Computational Linguistics 2020, 101–108. [Google Scholar] [CrossRef]
  15. Zhang, Y.; Zhang, Y.; Qi, P.; Manning, C.D.; Langlotz, C.P. Biomedical and clinical English model packages for the Stanza Python NLP library. J. Am. Med. Inform. Assoc. 2021, 28(9), 1892–1899. [Google Scholar] [CrossRef] [PubMed]
  16. Yang, S.; Cui, L.; Ning, R.; Wu, D.; Zhang, Y. Challenges to open-domain constituency parsing. In Findings of the Association for Computational Linguistics: ACL; Association for Computational Linguistics, 2022; Volume 2022, pp. 112–127. [Google Scholar] [CrossRef]
  17. Zhang, M. A survey of syntactic-semantic parsing based on constituent and dependency structures. arXiv Preprint posted online. 2020. [Google Scholar] [CrossRef]
  18. van Eck, N.J.; Waltman, L. Text mining and visualization using VOSviewer. arXiv Preprint posted online. 2011. [Google Scholar] [CrossRef]
  19. Mpekiri, S.; Papaspyropoulos, K.G. Citation dynamics, thematic structure and temporal evolution of research on the Faustmann Forest Economics model (1962–2025). For. Policy Econ. 2026, 182, 103685. [Google Scholar] [CrossRef]
  20. Version 9.0.6; Zotero, Ed.; Corporation for Digital Scholarship: Fairfax VA, 2026; Available online: https://www.zotero.org.
  21. Borgelt, C. An implementation of the FP-growth algorithm. In Proceedings of the 1st International Workshop on Open Source Data Mining: Frequent Pattern Mining Implementations, 2005; ACM; pp. 1–5. [Google Scholar] [CrossRef]
  22. Kruyt, B.; Van Vuuren, D.P.; De Vries, H.J.M.; Groenenberg, H. Indicators for energy security. Energy Policy 2009, 37(6), 2166–2181. [Google Scholar] [CrossRef]
  23. Sharma, P. Analyzing the role of renewables in energy security by deploying renewable energy security index. J. Sustain Dev. Energy Water Env. syst. 2023, 11(4), 1–21. [Google Scholar] [CrossRef]
Table 1. Journals with the highest number of publications on the topic “Energy security indicators” according to the ScienceDirect database.
Table 1. Journals with the highest number of publications on the topic “Energy security indicators” according to the ScienceDirect database.
Publication Title count
Energy 88
Energy Policy 58
Renewable and Sustainable Energy Reviews 55
Journal of Cleaner Production 40
Applied Energy 38
Energy Economics 37
Energy Strategy Reviews 26
Resources Policy 24
Energy Research & Social Science 23
Renewable Energy 20
Energy Reports 19
International Journal of Hydrogen Energy 18
Journal of Environmental Management 14
Ecological Indicators 13
Energy for Sustainable Development 12
Science of The Total Environment 11
Sustainable Cities and Society 10
Table 2. Top 30 two- and three-word terms and their frequencies in the complete set of extracted noun phrases.
Table 2. Top 30 two- and three-word terms and their frequencies in the complete set of extracted noun phrases.
Term Count Term Count
energy security 345 greenhouse emission 22
climate change 63 energy sector 21
sustainable development 49 energy consumption 21
economic growth 47 china energy security 21
environmental sustainability 46 united states 20
energy efficiency 44 carbon emission 20
renewable energy source 37 european union 20
co2 emission 35 positive impact 19
renewable energy 33 water energy security 19
energy transition 28 panel datum 18
energy system 27 negative impact 18
energy security indicator 27 develop country 18
environmental impact 25 valuable insight 17
national energy security 24 energy equity 17
energy security performance 24 energy resource 16
Table 3. Frequently co-occurring sets of three to five noun phrases consisting of two or three words.
Table 3. Frequently co-occurring sets of three to five noun phrases consisting of two or three words.
Set of terms Score
energy equity; environmental sustainability; energy security 1.20482
energy trilemma; environmental sustainability; energy security 0.722892
energy efficiency; environmental sustainability; energy security 0.60241
energy trilemma; energy equity; environmental sustainability 0.60241
sustainable development; climate change; energy security 0.481928
energy transition; environmental sustainability; energy security 0.481928
greenhouse emission; climate change; energy security 0.481928
energy consumption; energy efficiency; energy security 0.481928
carbon emission; climate change; energy security 0.481928
energy trilemma; energy equity; energy security; environmental sustainability 0.481928
energy trilemma; energy equity; energy security 0.481928
Table 4. Most frequent four- and five-word noun phrases extracted from titles and abstracts.
Table 4. Most frequent four- and five-word noun phrases extracted from titles and abstracts.
Noun phrases Count
energy security sustainable development 7
energy security environmental sustainability 7
energy security climate change 6
battery energy storage system 4
energy security energy equity 4
geographic information system gis 3
small island develop state 3
method moment quantile regression 3
panel correct standard error 3
environmental sustainability energy security 3
energy security economic growth 3
availability applicability acceptability affordability 3
climate change energy security 3
total primary energy supply 3
european union energy security 3
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.