Preprint
Review

This version is not peer-reviewed.

Text-to-OR: An Artifact-Centered Framework for NLP- and LLM-Enhanced Operations Research

Submitted:

09 July 2026

Posted:

10 July 2026

You are already at the latest version

Abstract
Traditional Operations Research (OR) models incorporate structured numerical inputs, such as demand parameters, costs, emissions, processing times, facility and transport mode capacities, to simulate stochastic behaviors and optimize objective-functions. There is however a vast number of OR-relevant parameters that can be derived from textual and semi-structured sources, such product reviews, social media, news, contracts, procurement documents, ESG reports, policy texts, patents, maintenance records, technical manuals and others. Under this context, the purpose of this study is to develop an artifact and parameter-centered framework for explaining how such textual inputs are transformed into model-ready OR components and incorporated into forecasting, simulation, optimization, logistics, inventory, sustainability, procurement, network analysis, and decision-support models. The main insights derived from the employed framework reveal that: (i) the OR value of textual information lies not in text analysis itself, but in its transformation into validated model-ready artifacts, such as covariates, parameters, constraints, scenarios, rules, weights, graph relations, simulation triggers, and solver inputs; (ii) different textual sources and language-processing methods can generate distinct OR artifacts that enter models through different integration mechanisms, including covariate augmentation, parameter updating, constraint generation, scenario definition, objective-function weighting, graph construction, retrieval support, and solver-code generation; (iii) the same text-derived artifact may play different roles across OR model types, for example a disruption event may update a simulation scenario, increase a lead-time parameter, remove a routing arc, or modify a supplier-risk penalty; and (iv) evaluation must extend beyond NLP accuracy to include artifact validity, parameter validity, model feasibility, mathematical consistency, solver correctness, deployment reliability, and downstream decision usefulness.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Operations Research (OR) is structured around mathematical models employed for optimizing a vast number of processes that companies undertake, such as logistics, inventory planning, scheduling, distribution and others (Geunes and Chang, 2008). These models traditionally use stochastic and deterministic demand parameters, costs, time parameters, and real-life constraints to optimize strategic tactical and operational decisions (Jayarathna et al., 2021). At the same time, text emerges as a critical source of information that can support decision-making (Chakraborty et al., 2014). For example, the number and quality of customer reviews could contain information on product preferences and thus potential demand changes (Fan et al., 2017). News articles and social-media posts may reveal supply chain segments where disruptions could occur (Sadeek & Hanaoka, 2023). Maintenance logs and technical manuals may describe failures, tasks, or operational requirements that can be used as constraints in maintenance schedule optimization (Cho et al., 2024). Recent advances in natural language models create new opportunities in OR as they allow for the successful and dynamic utilization of textual information in optimization (Wang & Li, 2025). Across literature, textual inputs are increasingly converted into structured optimization parameters and mathematical formulations (Ramamonjison et al., 2023). However, these developments are dispersed across different research streams with recent LLM–OR studies often described as fragmented and lacking a unified methodological framework (Wang & Li, 2025).
The objective of this paper is to therefore explain how textual and semi-structured data can be transformed into model-ready OR artifacts and incorporated into traditional OR models. The review was designed in line with well-established guidance on conducting structured and transparent literature reviews in the management and information systems fields (Tranfield et al., 2003; Webster and Watson, 2002). For the coding and synthesis stages, we drew on review approaches that stress systematic extraction, careful categorization, and the gradual development of concepts from the studies examined (Wolfswinkel et al., 2013).
The rest of the paper is organized as follows. Section 2 presents the research gaps while further mapping the paper’s contributions. Section 3 describes the employed methodological approach, while Section 4 presents the implementation and results of the employed methodology. Section 5 proposes a multi-level evaluation framework and research agenda for NLP-, LLM-, RAG-, and knowledge-graph-enhanced OR systems while finally Section 6 concludes the paper.

2. Contributions of the Paper

Existing research efforts reveal that text can support OR in multiple ways. These research efforts are however scattered across forecasting, risk analysis, procurement, sustainability, knowledge graphs, RAG-based decision support, simulation, and natural-language-to-optimization (Wang & Li 2025). Fan et al. (2017) for example deal with the relation between text and forecasting solely. The authors show that online reviews can be transformed into sentiment-based demand covariates that complement time series sales data. Iftikhar and Khan (2020) extend this logic by converting social-media posts, comments, and hashtags into emotion coefficients that reshape demand-diffusion dynamics. Zhang et al. (2022) integrate review and search-engine features into sales forecasting, while Wang and Zhang (2023) rely on Word2Vec and sentiment lexicons to derive product-feature and emotion factors for support vector regression. More recent works, such as Kim et al. (2025), Caetano et al. (2025), Ma et al. (2026), and Yang et al. (2025), demonstrates that textual descriptors, news, feedback, event labels, and promotional calendars can be recast as demand vectors, quantile forecasts, uncertainty parameters, or reinforcement-learning state variables. What these studies reveal, however, is a tendency to treat text as just one more predictive input, rather than as part of a broader logic for constructing OR artifacts. A second group of studies focuses on disruption, risk, and resilience. Singh et al. (2018) and Sharma et al. (2020) show that social-media and Twitter data can reveal operationally relevant disruption signals. Chu et al. (2020) transform textual information into region-specific supply-chain risk categories, while Zhao et al. (2025), Jacob et al. (2026), and Fang et al. (2026) convert public text, vendor-disruption news, and supplier-related information into topic-dominance weights, similarity scores, lead-time impact scores, or structured digital-risk artifacts. These studies show that news and social-media text become OR-relevant only when they update parameters such as supplier risk, lead time, route feasibility, capacity availability, or scenario severity. However, the existing literature usually studies these artifacts within supply-chain resilience and does not connect them to other text-to-OR pathways. Contractual, procurement, and legal documents provide another important pathway. Fantoni et al. (2021) translate contract terms into technical specifications, Jafari et al. (2021) extract reporting requirements from construction documents and connect them to time-cost prediction and simulation, and Pham et al. (2023) classify contract text into risk-handling actions. Aejas et al. (2025), Li and Ma (2026), and Hong et al. (2026) further show that contract and legal text can generate payment quantities, due dates, penalty constants, workforce skill-requirement matrices, delivery windows, and transport-management-system fields. A fourth stream focuses on sustainability, ESG, policy, and user-preference text. On this basis Kim and Kim (2017), Szekely and Vom Brocke (2017), and Wang et al. (2020) illustrate how sustainability reports, news, and maritime reports can be used for extracting sustainability indicators and themes. Yu et al. (2024), Jin and Ma (2026), and Kılınç et al. (2026) show how to extract fuzzy decision values, AHP weights, compliance scores, objective-function weights, solver-diagnostic feedback, or greenwashing indicators, form public reports, user reviews, policy texts, and sustainability disclosures can become. The significance of these studies hinges upon the fact that they show how textual information can enter OR not only through demand or cost parameters, but also through criteria, weights, constraints, and sustainability-oriented objective functions. Moreover, they additionally highlight the need to assess validity when text-derived scores are used as OR criteria.
Knowledge-graph and ontology-based studies show a different transformation logic. Kertkeidkachorn and Ichise (2017), Hao et al. (2021), Gan et al. (2023), Tupayachi et al. (2024), Zheng and Brintrup (2025), Tian et al. (2025), Yang and Xu (2025), and Sun et al. (2026) transform documents, manuals, accident narratives, operational logs, and public web text into entities, relations, graph edges, graph embeddings, adjacency matrices, incidence matrices, or link-prediction scores. The generated artifact is not a numerical covariate but a topological structure, capable of supporting network analysis, routing, risk propagation, spatial assignment, maintenance reasoning, and graph-based OR models more broadly. Existing work does establish the value of such graph structures, yet it tends to frame them as knowledge-management tools rather than treating them, explicitly, as artifacts that feed directly into OR models. RAG and semantic decision-support studies extend the role of text from preprocessing to interactive decision support. Garg et al. (2021), Avogadri et al. (2026), Lin et al. (2026), Nie et al. (2026), and Voloshchuk et al. (2026) show that retrieval, semantic search, vector similarity, and knowledge-graph-supported RAG can retrieve evidence, rank alternatives, structure database context, inject relational constraints, extract simulation rules, and compute logistics-node ranking scores. These works show that RAG is relevant to OR not simply because it answers questions, but because it retrieves and structures evidence that can affect downstream decision models. However, the OR role of RAG output remains insufficiently formalized in relation to parameters, constraints, scenarios, and decision rules. The most direct text-to-OR stream is natural-language-to-optimization and natural-language-to-simulation. Ramamonjison et al. (2023) formalize the conversion of natural-language optimization descriptions into variables, objectives, and constraints through NL4Opt. Jackson et al. (2024) convert process descriptions into executable discrete-event simulation code. Zhang et al. (2024) generate LP/MILP variables, coefficients, constraints, and solver code. Wang et al. (2025), Wu et al. (2025), Ding et al. (2026) further show that LLMs and agentic workflows can generate schemas, bounds, coefficient vectors, Big-M constants, algebraic constraints, solver programs, and validated mathematical-programming structures. Ding et al. (2026) highlight the role of reinforcement-learning-based validation for improving automated OR formulation from natural language. These studies show that language processing can produce the OR model itself, but they focus mainly on formulation generation and do not fully cover the broader range of text-derived artifacts used across forecasting, sustainability, contracts, graphs, RAG, risk, and simulation. The research gaps derived from the critical synthesis of the examined literature are summarized as follows:
  • Forecasting, risk, procurement, ESG, knowledge graphs, RAG, simulation, and natural-language-to-optimization are examined as discrete topics and not as integrated text-to-OR transformation pathways (Fan et al., 2017; Zhang et al., 2022; Chu et al., 2020; Fantoni et al., 2021; Kim and Kim, 2017; Kertkeidkachorn and Ichise, 2017; Lewis et al., 2020; Jackson et al., 2024; Ramamonjison et al., 2023).
  • Prior studies often emphasize the NLP, text-mining, RAG, knowledge-graph, or LLM technique used, while giving less explicit attention to the specific OR artifact generated from text (Tian et al., 2026; Zhang et al., 2026; Wang and Li, 2025).
  • The literature does not always clarify whether and how text-derived outputs are transformed into parameters, constraints, scenarios, objective coefficients, rules, weights, or graph structures that can enter formal OR models (Fantoni et al., 2021; Jafari et al., 2021; Fang et al., 2026; Ramamonjison et al., 2023; Zhang et al., 2024; Ding et al., 2026).
  • Graph and RAG outputs are often treated as decision-support or knowledge-management tools, rather than being explicitly formalized as OR model inputs such as graph relations, adjacency structures, retrieved evidence, relational constraints, or scenario-support artifacts (Hao et al., 2021; Gan et al., 2023; Zheng and Brintrup, 2025; Lewis et al., 2020; Gao et al., 2023; Avogadri et al., 2026; Lin et al., 2026).
  • Many studies assess NLP accuracy, retrieval quality, or formulation correctness, but give less systematic attention to OR-specific outcomes such as feasibility, solution quality, cost, service level, robustness, operational usefulness, or decision reliability (Powers, 2011; Sokolova and Lapalme, 2009; Es et al., 2024; Ramamonjison et al., 2023; Sargent, 2013; Ding et al., 2026).
  • Existing studies do not systematically report the full chain from textual source to NLP method, generated artifact, model role, validation approach, and downstream decision outcome, which motivates the artifact-centered reporting logic developed in this paper (Tranfield et al., 2003; Webster and Watson, 2002; Wolfswinkel et al., 2013; Wang and Li, 2025; Tian et al., 2026).
This paper addresses these research gaps by:
  • developing a common Text-to-OR framework that explains how unstructured or semi-structured text can be transformed into operationally meaningful artifacts for OR models. The framework brings together applications that are often discussed separately, including forecasting, risk analysis, contracts, ESG assessment, knowledge graphs, RAG, simulation, and optimization, and shows that they can be understood through a common transformation logic.
  • adopting an artifact-centered perspective. The main analytical focus is not only the NLP or LLM technique used to process text, but the OR-relevant artifact that is produced and subsequently integrated into a model.
  • positions knowledge graphs and RAG as important integration mechanisms for Text-to-OR applications. Graph relations, retrieved evidence, semantic rankings, and structured explanations can support model formulation, validation, scenario construction, and decision interpretation. Finally, natural-language-to-optimization is placed within a wider Text-to-OR ecosystem, alongside demand forecasting, risk modeling, procurement rule extraction, ESG criteria construction, graph analytics, and simulation-based decision support.

3. Research Methodology

Building on the research gaps identified in the previous sections, this study employs a structured, artifact-centered integrative literature review to examine how textual and semi-structured data are transformed into model-ready Operations Research (OR) artifacts. A structured literature-review approach is appropriate because the study requires a transparent process for identifying, screening, and synthesizing relevant studies across several research areas (Tranfield et al., 2003). An integrative review design is also appropriate because the literature is dispersed across NLP, text mining, LLMs, RAG, knowledge graphs, forecasting, optimization, simulation, supply-chain management, procurement, sustainability, and decision-support research. Integrative reviews are specifically useful when a review must synthesize fragmented or interdisciplinary literature and develop new conceptual perspectives from it (Torraco, 2005; Snyder, 2019).
The review follows a concept-centric synthesis approach. This means that the examined papers are not only compared by author, method, or application domain, but also according to the main concepts they address and the way they transform text into OR-relevant parameters (Webster and Watson, 2002). This approach is useful since literature is highly heterogeneous, covering different data sources, analytical methods, and decision-support applications. The coding and synthesis process also follows review practices that emphasize systematic extraction, categorization, and theory-building from literature-based evidence (Wolfswinkel et al., 2013).
The artifact-centered focus reflects the main argument of this review. Textual data become relevant for OR only when the outputs of language-processing methods are converted into structured artifacts that can be used in formal decision models. These artifacts may support model formulation, parameter estimation, constraint definition, scenario construction, simulation, optimization, or broader decision-support processes. In this way, the review focuses not only on how text is processed, but also on how the resulting artifact becomes operationally meaningful within an OR context. The review is therefore guided by the analytical sequence shown in Figure 1: textual or semi-structured input, language-processing method, OR artifact construction, OR model integration or decision context, OR decision output, and evaluation and feedback.

3.1. Search, Screening, and Corpus Construction

The search strategy operationalized three primary searchable domains within the Text-to-OR framework, namely textual or semi-structured input, language-processing method, and OR-related decision context. This concept-block approach is consistent with structured and concept-centric review guidance, where search terms are derived from the research question, review scope, and key conceptual dimensions rather than from an arbitrary list of keywords (Tranfield et al., 2003; Webster and Watson, 2002). In Figure 1 these domains correspond respectively to the textual-input stage, the language-processing stage, and the OR model integration or decision-context stage. The remaining elements of the framework, OR artifact construction, OR decision output, and evaluation and feedback, were treated as derived analytical dimensions. The above dimensions may not be clearly stated in bibliometric data such as in titles, abstracts, or keywords, but could potentially derive during screening, functional coding, and thematic synthesis. For example, a study may mention customer reviews and demand forecasting without explicitly stating that review sentiment becomes a demand covariate. Similarly, a paper may mention contract analysis and procurement without explicitly stating that extracted clauses become constraints, penalties, or eligibility rules. Therefore, the search was structured around the three primary searchable domains, while the full six-stage sequence was used to analyze how each study transforms text into OR-relevant artifacts and decision support. The search query was derived through a concept-block strategy, where the search strategy is based on the research question, review scope, and key conceptual dimensions and not on an arbitrary list of keywords (Tranfield et al., 2003; Webster and Watson, 2002; Wolfswinkel et al., 2013). The three query blocks were derived deductively based on the three primary searchable domains within the Text-to-OR framework, namely textual or semi-structured input, language-processing method, and OR-related decision context.
The first block addressed the language-processing method domain. It combined broad terms, such as “natural language processing,” “NLP,” “text mining,” and “text analytics,” with more specific method terms, including “sentiment analysis,” “topic modeling,” “named entity recognition,” and “relation extraction.” These terms were included because they commonly appear in representative studies where textual data are transformed into OR-relevant artifacts, such as forecasting covariates, sustainability indicators, contract rules, or knowledge-graph relations (Fan et al., 2017; Kim and Kim, 2017; Fantoni et al., 2021; Kertkeidkachorn and Ichise, 2017). The block also included newer terms associated with LLM- and retrieval-based systems, such as “transformer,” “BERT,” “large language model,” “LLM,” “generative AI,” “retrieval augmented generation,” “RAG,” “semantic search,” “text-to-code,” “text-to-SQL,” and “text-to-optimization.” These terms were included because recent studies increasingly use LLMs, RAG, semantic retrieval, and automated formulation generation to produce OR-relevant artifacts such as retrieved evidence, solver code, and optimization formulations (Ramamonjison et al., 2023; Lee et al., 2023; Avogadri et al., 2026; Ding et al., 2026). The purpose of the block was not to list every possible NLP technique, but to combine broad searchable terms with recurring method-specific terms in order to balance recall and precision.
The second block captured the OR-related decision-context domain. It included terms such as “operations research,” “optimization,” “simulation,” “forecasting,” “inventory,” “logistics,” “supply chain,” “transportation,” “routing,” “scheduling,” “resource allocation,” “decision support,” “risk management,” “resilience,” “supplier selection,” “procurement,” “sustainability,” and “ESG.” These terms were included to ensure that retrieved studies were not only about language processing, but also connected to OR-relevant models, decision systems, or operational decision tasks.
The third block captured the textual or semi-structured input domain. It included general terms such as “text,” “unstructured data,” “semi-structured data,” and “natural language,” as well as source-specific terms such as “reviews,” “social media,” “news,” “contracts,” “policies,” “ESG reports,” “sustainability reports,” “patents,” and “maintenance records.” These terms reflect the main language-based sources identified through prior review papers, representative studies, and pilot screening. Customer reviews and social-media text are common in forecasting and demand-related studies (Fan et al., 2017; Zhang et al., 2022), while news and public text are used in disruption and supply-chain-risk studies (Sharma et al., 2020; Fang et al., 2026). Contracts and procurement documents appear in rule, risk, and constraint-extraction studies (Fantoni et al., 2021; Jafari et al., 2021), and ESG or sustainability reports are used in sustainability-assessment studies (Kim and Kim, 2017; Szekely and Vom Brocke, 2017). Natural-language problem descriptions appear in text-to-optimization and solver-code generation studies (Ramamonjison et al., 2023; Ding et al., 2026). Scopus was selected as the primary database for this review. The metadata screening was conducted on 1 June 2026, and the exported records covered the period 2016–2027. The final Scopus query combined the three concept blocks: language-processing methods, OR-related decision contexts, and textual or semi-structured inputs. It also excluded areas outside the scope of Text-to-OR, including computer vision, medical imaging, speech recognition, audio processing, video analysis, machine translation, and pure NLP benchmarks without OR relevance. The exact query was:
TITLE-ABS-KEY(
("natural language processing" OR NLP OR "text mining" OR "text analytics" OR "sentiment analysis" OR
"topic modeling" OR "named entity recognition" OR "relation extraction" OR transformer OR BERT OR
"large language model*" OR LLM OR "generative AI" OR "retrieval augmented generation" OR RAG OR
"knowledge graph*" OR "text-to-code" OR "text-to-SQL" OR "text-to-optimization" OR
"natural language query" OR "semantic search")
AND
("operations research" OR optimization OR simulation OR forecasting OR inventory OR logistics OR
"supply chain*" OR transportation OR routing OR scheduling OR "resource allocation" OR
"decision support" OR "risk management" OR resilience OR "supplier selection" OR procurement OR
sustainability OR ESG)
AND
(text* OR "unstructured data" OR "semi-structured data" OR reviews OR "social media" OR news OR
contracts OR policies OR "ESG reports" OR "sustainability reports" OR patents OR
"maintenance records" OR "natural language")
)
AND NOT TITLE-ABS-KEY(
"image processing" OR "computer vision" OR "medical imaging" OR "speech recognition" OR
"audio processing" OR "video analysis" OR "natural language translation" OR "machine translation" OR
"sentiment classification only" OR "pure NLP benchmark"
)
The paper selection criteria followed the same three-domain logic used in the search strategy. A study was retained only when it jointly satisfied all three conditions (i) it used textual, semi-structured, unstructured, or language-based input (ii) it applied a language-processing method and (iii) it supported an OR-related model, system, or decision task.
The search retrieved 36,437 initial records. After removing 17 duplicates, 36,420 records remained for title-and-abstract screening. A total of 34,664 records were excluded at this stage because they did not establish sufficient connection between textual or semi-structured data, language-processing methods, and OR-related decision use. The remaining 1,756 records were assessed for metadata-based eligibility, and 119 additional records were excluded. The final corpus contained 1,637 studies.
Table 1. Review-log structure.
Table 1. Review-log structure.
Review-Log Element Value
Database searched Scopus
Search date Metadata screening conducted on 4 May 2026
Time window 2016–2027 based on uploaded Scopus CSV files
Initial records retrieved 36,437
Duplicates removed 17
Records screened by title and abstract 36,437
Records excluded at title/abstract stage 34,664
Full texts assessed 1,756 metadata-eligible records assessed
Full texts excluded 119 metadata-based eligibility exclusions
Final included studies 1,637
Supplementary studies added manually 0
Table 2. Screening results by publication year.
Table 2. Screening results by publication year.
Year Initial Duplicates Screened Title/Abstract Excluded Eligibility Assessed Eligibility Excluded Final Included
2016 674 1 673 628 45 7 38
2017 842 1 841 785 56 6 50
2018 1,149 3 1,146 1,080 66 5 61
2019 1,507 2 1,505 1,418 87 8 79
2020 1,985 4 1,981 1,878 103 12 91
2021 2,576 1 2,575 2,440 135 11 124
2022 3,263 1 3,262 3,104 158 17 141
2023 4,441 1 4,440 4,205 235 28 207
2024 985 0 985 929 56 3 53
2025 13,572 2 13,570 12,986 584 18 566
2026 5,442 1 5,441 5,210 231 4 227
2027 1 0 1 1 0 0 0
Total 36,437 17 36,420 34,664 1,756 119 1,637

3.2. Functional Coding and Thematic Synthesis

After corpus construction, each included study was coded according to its transformation pathway from text to OR use. This coding logic follows concept-centric review guidance, in which studies are organized around analytical concepts rather than only around authors, methods, or application domains (Webster and Watson, 2002). The paper also adopts coding-based review procedures emphasizing on systematic extraction, categorization, and synthesis of evidence from the literature (Wolfswinkel et al., 2013; Tranfield et al., 2003).
Regarding the systematic extraction employed, the extraction stage recorded, for each study, the type of textual or semi-structured input, the language-processing method used, the OR-relevant artifact produced, the mode of integration into an OR model or decision-support system, the decision output generated, and the reported form of evaluation or feedback. The categorization stage then grouped studies according to the functional role played by text in the OR process, such as review-to-covariate, news-to-risk-event, clause-to-rule, clause-to-constraint, document-to-parameter, or text-to-scenario transformations. Finally, the synthesis stage compares these coded categories across OR domains to identify recurring mechanisms through which language-based evidence is transformed into model-ready OR artifacts. The final derived coding dimensions are summarized in Table 3 and involve author-year, OR domain, textual or semi-structured source, language-processing method, OR task supported, OR artifact generated, model-integration mechanism, evaluation approach, and reported limitations.
The coding dimensions in Table 3 were used to move from individual study coding to thematic synthesis. The coded records were compared across studies to identify repeated combinations of textual source, language-processing method, OR artifact, and OR task supported. These recurring combinations were interpreted as functional Text-to-OR transformation streams. Thus, the thematic streams were not defined only by application domain, but by how textual or semi-structured data were converted into model-relevant OR elements. For example, customer-review studies were grouped as review-to-covariate transformations when review text was converted into explanatory variables for forecasting models. News and social-media studies were grouped as news-to-risk-event transformations when textual signals became disruption indicators or risk-event records. Contract and procurement studies were grouped as clause-to-rule or clause-to-constraint transformations when contractual text became obligations, compliance rules, constraints, or decision logic. Knowledge-graph, simulation, and natural-language-to-optimization studies were grouped according to whether text became graph relations, simulation structures, or solver-ready formulations.
Figure 2. Thematic synthesis streams.
Figure 2. Thematic synthesis streams.
Preprints 222376 g002

4. Implementation and Results

This section presents the results of the thematic synthesis. The coded studies were allocated to eight thematic streams according to the dominant Text-to-OR transformation pathway identified in each paper. Each stream reflects a different way in which textual or semi-structured data are converted into OR-relevant artifacts, such as covariates, risk events, sustainability indicators, rules, graph relations, retrieved evidence, simulation structures, or optimization formulations. For reporting purposes, each study was assigned to one dominant thematic stream. The representative papers listed in Table 4 are illustrative examples from each stream and are not intended to reproduce the full set of included studies.

4.1. Review, Taxonomy, and Framework Papers

A first stream consists mainly of review, taxonomy, and framework-oriented studies that describe how NLP, text mining, sentiment analysis, topic modeling, and related language-processing methods can support operational and engineering-management decisions. The contribution of this stream is mainly conceptual. It helps position language analytics as a decision-support capability across operational contexts. Tian et al. (2026) for example, organize the literature around language-processing methods, data sources, and application areas. Zayet et al. (2021) systematically map how transportation-related research uses social-media analysis to capture commuter feedback, opinions, complaints, and traffic-situation information. Schöpper and Kersten (2021) review natural-language-processing methods for supply-chain mapping, focusing on how unstructured text can support supply-chain transparency. Chowdhury and Alzarrad (2023) review text-mining applications across transportation-infrastructure research. This stream is also informed by illustrative decision-support studies that show how such classifications appear in practice. Mendez et al. (2019) use Twitter data with sentiment analysis and topic modeling to capture user satisfaction with public transport in Santiago, and Chile. Manders and Klaassen (2019) apply text mining to Dutch news articles and initiative websites to unpack the smart-mobility concept and Vasquez-Henriquez et al. (2020) analyze tweets from Santiago to characterize gender-related differences in transport perception.

4.2. Customer Reviews/Social Media in Demand Forecasting and Inventory Planning

A second stream converts customer reviews, online ratings, search traces, social-media content, promotional calendars, event descriptions, and media sentiment into demand-forecasting and inventory-planning signals (Lyu and Choi, 2020; Ou-Yang et al., 2022; Wang and Zhang, 2023; Shao et al., 2025). In this pathway, textual data are used to capture market perception, consumer preference, product experience, product-feature evaluation, and demand-side information that may not be visible in structured sales histories alone (Fan et al., 2017; Zhang et al., 2022). Fan et al. (2017) extend the Bass/Norton diffusion model with a sentiment index derived from online reviews for automotive sales forecasting. Zhang et al. (2022) combine review-based sentiment indices with Baidu search data and historical sales through a PCA–DSFOA–BPNN model for monthly automobile sales prediction.
Shao et al. (2025) construct media-sentiment indices from social-media reviews and news headlines and integrate them into machine-learning models for new-energy-vehicle sales forecasting. Liu et al. (2023) combines sentiment analysis of online reviews with Internet-search features in electric-vehicle sales forecasting. Lyu and Choi (2020) combine text mining and emotion-based sentiment analysis of Taobao reviews with neural-network forecasting to predict online sales volume for organic products. Ou-Yang et al. (2022) use multi-channel online sentiment data with a CNN-LSTM model to forecast car-sales movement. Wang and Zhang (2023) extract product-feature and product-sentiment factors from online reviews using Word2Vec and sentiment dictionaries and use them in a multivariate SVR demand model for beauty products. Let D t denote observed demand in period t , and let D ^ t denote forecast demand. Let X t be the vector of structured operational predictors and Z t t e x t the text-derived forecasting vector. A compact demand-forecasting representation is:
D ^ t = f X t , Z t text ,
The text-derived vector may be represented as:
Z t text = s t , θ 1 t , θ 2 t , , θ K t , v t , e t ,
where s t is a sentiment or emotion score, θ k t is the proportion of topic k , v t is review, search, or post volume, and e t is an event-related indicator. These variables are dimensionless scores, proportions, counts per period, or binary/context indicators depending on the extraction method. If D ^ L text denotes expected demand during lead time and σ D text denotes the corresponding demand standard deviation a simple replenishment connection is:
S S = z σ D text L ,
R O P = D ^ L text + S S ,
where S S and R O P are measured in units, z is a dimensionless service-level factor, and L is the lead time in periods.
Table 5. Nomenclature of Model Parameters and Variables.
Table 5. Nomenclature of Model Parameters and Variables.
Nomenclature Definition
t Time-period index used to align textual and operational data (time units: day, week, month, etc.)
D t Observed demand in period t (units/time period)
D ^ t Forecast demand in period t (units/time period)
X t Structured predictor vector (units per period, EUR per unit, binary indicators, days, or normalized values, depending on predictor)
Z t t e x t Text-derived forecasting vector (dimensionless normalized scores, counts/time period, or binary indicators, depending on component)
s t Sentiment or emotion score in period t (dimensionless; commonly 1,1 or 0,1 )
θ k t Proportion or weight of topic k in period t (dimensionless proportion; 0 θ k t 1 )
K Number of retained topics in the forecasting vector (count)
v t Review, search, or post volume in period t (count/time period)
e t Event-related text variable (binary indicator, categorical variable, or normalized intensity score)
D ^ L t e x t Text-enhanced expected demand during lead time (units)
σ D t e x t Text-enhanced demand standard deviation (units/time period or units during lead time, depending on formulation)
S S Safety stock (units)
R O P Reorder point (units)

4.3. News/Social Media and Supply-Chain Risk

A third stream uses news articles, social-media posts, online discourse, and public information flows to detect disruptions, classify risks, and support supply-chain resilience decisions. Language-processing outputs become risk-event records, disruption indicators, public-opinion signals, event classifications, or ranked risk categories (Chu et al., 2020; Fang et al., 2026). The value of this stream lies in early-warning capability because public text can contain weak signals of disruption before they appear in formal operational datasets (Singh et al., 2018; Sharma et al., 2020). The use of textual information in this stream is mainly associated to the updating of parameters such as capacity adjustment, route-feasibility modification, and disruption-scenario construction. For example, a news article may be transformed into an event record such as disruption type = port strike; location = affected port; affected flow = inbound containers; expected effect = delay; severity = high; confidence = medium (Chu et al., 2020; Zhao et al., 2025; Jacob et al., 2026; Fang et al., 2026). The OR artifact is therefore not the news article itself, but the structured disruption record that can update model inputs.
Let L i be the baseline lead time of supplier, facility, or route i , and let Δ L i t e x t be the lead-time increase estimated from a text-derived disruption record. The updated lead time is:
L i u p d = L i + Δ L i t e x t ,
where Δ L i t e x t is the text-derived lead-time increase in days, weeks, or other planning periods. Jacob et al. (2026) support this pathway by converting supplier-related news into lead-time impact scores. Zhao et al. (2025) and Fang et al. (2026) support the derivation and validation of supply-chain risk indicators from textual sources.
Risk-event records may also modify available capacity. If C i is the baseline operational capacity of a node, supplier, or facility i , and η i t e x t 0,1 is the text-derived capacity-loss proportion, the disrupted capacity can be represented as:
C i u p d = C i 1 η i t e x t , 0 η i t e x t 1
where C i u p d remains in the same physical unit as C i , such as units/day, tonnes/week, or containers/period.
Table 6. Nomenclature of Model Parameters and Variables.
Table 6. Nomenclature of Model Parameters and Variables.
Nomenclature Definition
i Supplier, facility, node, or route index
L i Baseline lead time of supplier, facility, or route i (time periods: days, weeks, months, etc.)
L i u p d Updated lead time after incorporating the text-derived disruption effect (time periods: days, weeks, months, etc.)
Δ L i t e x t Text-derived lead-time increase for supplier, facility, or route i (time periods: days, weeks, months, etc.)
C i Baseline operational capacity of supplier, facility, or node i (units/time period, tonnes/time period, containers/time period, etc.)
C i u p d Updated operational capacity after disruption (same unit as C i )
η i t e x t Text-derived proportional capacity-loss factor for supplier, facility, or node i (dimensionless; 0 η i t e x t 1 )

4.4. ESG, Sustainability, and Policy Text

A fourth stream transforms sustainability reports, ESG disclosures, policy-related text, news articles, and social-media content into sustainability indicators, disclosure topics, greenwashing signals, or public-perception measures. These artifacts can support sustainability evaluation, ESG assessment, corporate responsibility monitoring, and green supply-chain decision support. The stream extends the OR-pipeline logic beyond cost and service-level outcomes because language-derived indicators can become criteria, scores, weights, or constraints in sustainability-oriented decision models (Szekely and Vom Brocke, 2017). Kim and Kim (2017) apply Leximancer and DICTION text mining to news articles and sustainability reports to map sustainable supply-chain management trends and CEO-letter rhetoric in the textile and apparel industry. Szekely and Vom Brocke (2017) apply topic modelling (Latent Dirichlet Allocation) to 9,514 corporate sustainability reports to identify forty-two economic, environmental, and social sustainability topics and trends. Wang et al. (2020) combine manual text classification with automatic text mining on maritime sustainability reports from container shipping liners and terminal operators to map the industry's role across the seventeen Sustainable Development Goals; and Kılınç et al. (2026) develop a multi-dimensional textual framework that integrates three complementary indicators to detect greenwashing in corporate sustainability reporting. The main evaluation issue is constructing validity as text-derived ESG measures may capture disclosure intensity rather than actual sustainability performance (Szekely and Vom Brocke, 2017; Zakaria et al., 2021; Kılınç et al., 2026). Szekely and Vom Brocke (2017) extract economic, environmental, and social sustainability topics from 9,514 sustainability reports via Latent Dirichlet Allocation; Angin et al. (2022) fine-tune a RoBERTa classifier on the OSDG Community Dataset to automatically classify sustainability reports by relevance to the Sustainable Development Goals; Kılınç et al. (2026) integrate three complementary textual indicators in a multi-dimensional framework that flags greenwashing patterns in corporate sustainability reports; and Zakaria et al. (2021) build a sustainability dictionary and use an information-entropy measure of word-distribution similarity to score the transparency of corporate sustainability reports, demonstrated on OPEC and non-OPEC energy producers. The resulting artifacts include sustainability topics, SDG labels, disclosure-pattern indicators, greenwashing signals, CEO-message themes, and transparency scores that can support ESG assessment and sustainability-oriented decision models.
In OR terms, this stream corresponds mainly to sustainability scoring, threshold-based eligibility, multi-criteria decision analysis, and objective-function weighting. For example, a supplier-selection model may combine cost, quality, delivery reliability, and a text-derived sustainability score as decision criteria. For this reason, construct validity becomes a central requirement when ESG-related textual outputs are used as inputs in OR models (Kim and Kim, 2017; Kılınç et al., 2026). Text-derived ESG scores are often reported on heterogeneous scales, such as counts of sustainability-related disclosures, topic-intensity scores, sentiment-derived ratings, 0–100 compliance indices, or classifier probabilities (Szekely and Vom Brocke, 2017; Zakaria et al., 2021; Kılınç et al., 2026).
Let s ~ i j t e x t denote the raw text-derived sustainability score of alternative i on criterion j , and let w j be the criterion weight. A standard weighted sustainability or supplier-evaluation score can then be represented as:
V i = j = 1 J w j s i j t e x t , j = 1 J w j = 1 ,   w j > 0
where V i is the aggregated sustainability score of alternative i , w j is the relative importance of criterion j , and J is the number of retained criteria. The scores s i j t e x t are dimensionless normalized quantities, while the weights w j are also dimensionless and typically sum to one. Kim and Kim (2017), Szekely and Vom Brocke (2017), Wang et al. (2020), and Kılınç et al. (2026) support the transformation of sustainability and compliance text into structured decision criteria. Text-derived ESG indicators may also enter OR models as eligibility constraints. If for example a decision-maker requires that selected suppliers satisfy a minimum sustainability threshold τ S , then:
V i τ S  
Table 7. Nomenclature of Model Parameters and Variables.
Table 7. Nomenclature of Model Parameters and Variables.
Nomenclature Definition
i Alternative, supplier, project, or decision-option index (index set)
j ESG, sustainability, compliance, or policy criterion index (index set)
J Number of retained sustainability or ESG criteria (count)
s ~ i j t e x t Raw text-derived sustainability score of alternative i on criterion j before normalization (original score scale; dimensionless unless otherwise defined)
s i j t e x t Normalized text-derived sustainability score of alternative i on criterion j (dimensionless; 0 s i j t e x t 1 )
V i Aggregated sustainability or ESG evaluation score of alternative i (dimensionless weighted score)
w j Relative weight assigned to criterion j (dimensionless; commonly 0 w j 1 and j = 1 J w j = 1 )
τ S Minimum sustainability threshold required for acceptability or eligibility (dimensionless normalized threshold)

4.5. Contract, Procurement, and Supplier Text

A fifth stream processes contracts, invitations to bid, procurement documents, technical requirements, and engineering or construction project texts to extract risks, obligations, requirements, specifications, and decision-relevant attributes. This stream differs from general document classification because the extracted information has operational consequences for project execution, procurement risk, contract compliance, and supplier decisions (Fantoni et al., 2021; Choi et al., 2021; Jafari et al., 2021). Fantoni et al. (2021) develop a text-mining tool that translates contractual terms into technical specifications from railway-sector tender documents; Choi et al. (2021) combine a phrase-matcher Critical Risk Check (CRC) module with a named-entity-recognition Term Frequency Analysis (TFA) module to extract risk-involved clauses from EPC invitation-to-bid and contract documents; Jafari et al. (2021) integrate NLP, machine learning, and Monte Carlo simulation to extract reporting requirements from construction contracts and predict the associated time and cost; and Pham and Han (2023) use NLP with multitask classification to jointly predict risk identification, risk allocation, and risk-response actions in construction contracts. The same contract- and procurement-oriented logic is supported by additional studies that mine EPC case descriptions, invitation-to-bid documents, public-procurement policy text, legal supply-chain relations, and technical proposal documents. Son and Lee (2019) apply text mining to bid documents of Korean offshore oil-and-gas EPC mega-projects to build a schedule-delay estimate model from thirteen case studies. Cao et al. (2022) use quantitative text analysis on large-scale Chinese tender documents to assess the implementation of sustainable public procurement. Aejas et al. (2025) develop a hierarchical transformer for legal relation extraction that converts contract clauses into structured business rules for smart-contract automation in supply-chain settings; and Park et al. (2021) build a digitalized design-risk-analysis tool with machine-learning algorithms for EPC contractors' technical-specifications assessment at the bidding stage. Their outputs are typically risk indicators, procurement-assessment evidence, legal relations, or design-risk signals, which can be used as inputs for supplier screening, compliance analysis, project-risk control, and procurement decision support.
In OR terms, this stream corresponds mainly to constraint generation, rule encoding, eligibility screening, penalty modeling, and schedule or workforce feasibility definition (Aejas et al., 2025; Fantoni et al., 2021; Jafari et al., 2021; Pham and Han, 2023; Li and Ma, 2026). Contract clauses and procurement requirements are generally not normalized when they specify explicit operational quantities. A delivery deadline remains in days or weeks, a penalty remains in monetary units per time period of delay, and a workforce requirement remains in workers, worker-hours, or full-time equivalents. Normalization may be useful for model-confidence scores or classification probabilities, but the extracted decision parameters themselves should preserve their original operational units. A first integration mechanism concerns delivery-time restrictions. Let T i denote the realized delivery time of the supplier i , and let T i m a x denote the maximum delivery time extracted from contract text. A hard delivery requirement becomes:
T i     T i m a x ,  
If late delivery is permitted but penalized, a non-negative lateness variable y i can be introduced:
y i     T i     T i m a x ,   y i     0
The associated penalty may then enter the procurement objective in prose or in a later full model.
Table 8. Nomenclature of Model Parameters and Variables.
Table 8. Nomenclature of Model Parameters and Variables.
Nomenclature Definition
i Supplier index used in procurement and sourcing models (index set)
T i Realized delivery time of supplier i (time periods: days, weeks, months, etc.)
T i m a x Maximum allowable delivery time extracted from contract text for supplier i (time periods: days, weeks, months, etc.)
y i Lateness variable associated with supplier i (time periods of delay)

4.6. Knowledge Graphs and Ontologies

A sixth stream converts accident reports, hazardous-chemical documents, engineering design knowledge, safety documents, and other unstructured text into knowledge graphs, ontologies, entity relations, and graph-based decision structures. The generated artifact is typically an entity-relation graph, ontology, semantic network, or structured knowledge base that supports reasoning, navigation, link prediction, risk analysis, or decision knowledge management (Hao et al., 2021). Kertkeidkachorn and Ichise (2017) present T2KG, an end-to-end system that converts unstructured text into knowledge-graph triples through a hybrid rule-based and similarity-based predicate-mapping module; Zheng et al. (2021) build a knowledge graph for hazardous-chemical management based on ontology design and named-entity-recognition-based entity identification. Hao et al. (2021) propose a meta-model of decision knowledge graph grounded in the compromise Decision Support Problem construct, allowing designers to navigate engineering-design decision-related knowledge in unstructured documents; and Gan et al. (2023) construct an ontology-based knowledge graph from 241 ship-collision-accident investigation reports of the China Maritime Safety Administration, with 910 entity nodes and 1,920 relation edges to support case retrieval for accident investigation. Ali et al. (2019) combine a fuzzy domain ontology with Word2vec embeddings and a Bi-LSTM classifier to monitor an intelligent transportation network through social-media text; Yuan et al. (2023) build a traffic-safety-management knowledge graph in Neo4j with 54 node types and 14 relationship types, queried through Cypher and a rule-matching automatic Q&A function; and Liang et al. (2024) use a seven-step ontology-construction approach to develop a knowledge graph for papermaking-process fault diagnosis that visualises material- and energy-flow relationships and supports root-cause identification. Their shared contribution is the conversion of language-rich records into entities, relations, semantic categories, or graph structures that support monitoring, reasoning, classification, and operational knowledge management
In OR terms, this stream corresponds to relational model construction. For example, a supply-chain document may state that Supplier A provides Component X to Manufacturer B and operates in a flood-prone region. These textual statements can be transformed into graph relations such as Supplier A supplies Component X, Component X is used by Manufacturer B, and Supplier A is exposed to flood risk. The graphical representation of the above information can be used in multiple OR tasks such as dependency analysis, supplier-risk propagation, alternative-supplier identification, transport-network optimization, facility-layout reasoning, maintenance diagnosis, and supply-chain mapping. Moreover, its value lies in the fact that the graph doesn’t simply provide extracted textual information but clearly depicts the relationships between entities, events, locations, assets, suppliers, and processes (Zheng and Brintrup, 2025; Hao et al., 2021). Let A i j t e x t denote a text-derived adjacency parameter indicating whether a relation from node i to node j has been identified. A binary relation extracted from text can be represented as:
A i j t e x t = 1 , i f   t e x t   i n d i c a t e s   a   r e l a t i o n   f r o m   i   t o   j , 0 , o t h e r w i s e
where A i j t e x t is a binary graph relation indicator. Text-derived graphs can support risk propagation across operational dependencies. Let R i denote the risk score of node i , R j 0 denote the baseline risk of node j , and A i j t e x t denote the text-derived weighted relation from i to j . A simple one-step propagation form can be written as:
R j u p d = R j 0 + i = 1 N A i j t e x t R i ,  
where N is the number of nodes or upstream connected entities. In this formulation, risk at node j increases when strongly connected upstream nodes i have high risk values. For example, if a supplier is textually linked to a critical component and exposed to flood risk, that risk may propagate to downstream manufacturers depending on it.
Table 9. Nomenclature of Model Parameters and Variables.
Table 9. Nomenclature of Model Parameters and Variables.
Nomenclature Definition
i , j Node indices used in graph-based decision support, supply-chain networks, transport networks, or relational models (index set)
N Number of nodes or entities in the graph/network representation (count)
A i j t e x t Text-derived adjacency value indicating a relation from node i to node j (binary indicator or weighted relation score)
R i Risk score of node i used in graph-based propagation logic (dimensionless normalized risk score)
R j 0 Baseline risk score of node j before relational propagation (dimensionless normalized risk score)
R j u p d Updated or propagated normalized risk score of node j

4.7. RAG and Semantic Decision Support

A seventh stream uses retrieval, semantic search, retrieval-augmented generation, knowledge graphs, and LLM-based interfaces to connect user queries with operational knowledge bases and decision-support outputs. In this pathway, the artifact is not only an extracted feature but also a grounded answer, retrieved evidence set, explanation, recommendation, or graph-supported response (Avogadri et al., 2026; Lin et al., 2026; Lee et al., 2023). Avogadri et al. (2026) develop Discovery Omnia, a dynamic-RAG framework that integrates LLMs with algorithmic routing and structured patent databases to reduce hallucinations in patent analysis and support systematic innovation; and Lin et al. (2026) integrate a dynamic knowledge graph with large language models to provide intelligent decision support for tunnel-fire incidents. This stream shifts NLP-enhanced OR from batch preprocessing toward interactive decision support because users can query operational knowledge and receive context-aware responses linked to retrieved or structured evidence. Lee et al. (2023) develop the General Provisions Question-Answering Model (GPQAM), which combines a knowledge graph with question-answering techniques to enable semantic search of equipment-purchase-order clauses for steel-plant maintenance contracts; and Zheng et al. (2026) introduce CARAG, a context-aware retrieval-augmented-generation framework over a spatial knowledge graph that answers railway operation-and-maintenance questions by jointly retrieving from train-trajectory data and operating-rule documents. These systems emphasize retrieval, semantic structuring, or interactive explanation as decision-support artifacts rather than treating language processing as an isolated classification exercise
In OR terms, this stream corresponds to evidence-based and interface-oriented decision support. The OR artifact is therefore not only the generated answer, but also the retrieved and traceable evidence set that supports downstream modeling or decision making (Lewis et al., 2020; Gao et al., 2023; Es et al., 2024; Avogadri et al., 2026; Lin et al., 2026). RAG integration should consequently be evaluated not only by linguistic fluency, but also by grounding, retrieval quality, traceability, latency, usefulness, and hallucination control (Avogadri et al., 2026; Lin et al., 2026). Let q denote the embedding vector of a query and d i denote the embedding vector of document, passage, node, or evidence item i . Embedding vectors are dimensionless numerical representations. Their cosine similarity is a unitless score bounded by 1,1 :
s i c o s = q d i q d i , q 0 , d i 0 ,  
where s i c o s measures the semantic proximity between the query and evidence item i . Let K K R t o p denote the index set of the K R highest-scoring evidence items. The retrieved evidence set is:
E ' = { d i : i K K R t o p } ,
where E ' is the selected set of retrieved documents, passages, rules, or nodes, and K R is the number of retained evidence items. The evidence set E ' is not itself normalized; it is a discrete subset passed to the downstream decision-support, reasoning, or OR-modeling stage.
Table 10. Nomenclature of Model Parameters and Variables.
Table 10. Nomenclature of Model Parameters and Variables.
Nomenclature Definition
q Embedding vector of the user query or decision-support request.
d i Embedding vector of document, passage, node, or evidence item i .
i Candidate document, passage, graph node, or evidence-item index.
s i c o s Cosine similarity score between the query embedding and evidence embedding.
E ' Retrieved evidence set passed to the downstream decision-support or OR-modeling stage.
K K R t o p Index set of the K R highest-scoring evidence items or nodes.
K R Number of evidence items retained by top- K R retrieval

4.8. Natural-Language-to-Optimization and Simulation

An eighth stream converts natural-language problem descriptions, prompts, domain text, accident narratives, or system descriptions into optimization formulations, solver code, simulation inputs, scenarios, or automated OR models. Ramamonjison et al. (2023) formalise this pathway through the NL4Opt Competition, where linear-programming problem descriptions are mapped first to tagged semantic entities and then to a logical-form representation executable by commercial solvers, enabling non-experts to interface with optimisation tools through natural language. Guo et al. (2024) develop SoVAR, which uses LLM prompts with linguistic patterns to extract accident information from textual NHTSA reports, formulates accident-related constraints solved by Z3, and reconstructs road-generalisable test scenarios for autonomous driving systems. Bilal et al. (2025) introduce Hybrid TrafficAI, a generative-AI framework that fuses video, LiDAR, and textual data through an Adaptive Multi-Modal Fusion Engine and combines GAN, transformer, and LLM modules for real-time traffic simulation, behaviour modelling, and anomaly detection. Ding et al. (2026) propose OR-R1, a data-efficient two-stage training framework that combines supervised learning with test-time reinforcement learning to translate natural-language problem descriptions into formal optimisation models and executable solver code .Recent studies extend this stream by using LLMs, transformer-based models, RAG, and optimization algorithms in increasingly specialised ways. Penco et al. (2025) develop ACMG, a fine-tuned-LLM framework that translates natural-language descriptions into formal constraint-satisfaction models through semantic entity extraction, constraint-model generation, and iterative validation against the MiniZinc solver. Pan et al. (2025) propose PaMOP, which guides LLMs in modelling optimisation problems by extracting the problem into a tree structure and prompting the model with self-augmented prompts for each partitioned set of constraints. Ahmed and Choudhury (2025) introduce OPT2CODE, a retrieval-augmented framework that uses RAG and prompt chaining to translate linear-programming problem descriptions into solver-specific code, evaluated across both open- and closed-source LLMs. Jiang et al. (2025) propose LLMOPT, which learns to define and solve general optimisation problems from natural-language descriptions through a five-element formulation, multi-instruction tuning, and an alignment-and-self-correction mechanism designed to suppress hallucinations during solver-code generation.
These model components enter OR most directly when they are encoded as candidate mathematical formulations. A production-planning description may state that a factory has limited labour hours, several products, known profits, and demand requirements. A language model may convert this description into decision variables, an objective function, capacity constraints, demand constraints, and solver code. The artifact is therefore a candidate mathematical model rather than a summary.
This does not mean that constraints only arise in this stream. Constraints may also be derived from contracts, procurement documents, regulations, policies, technical manuals, or operational rules in other Text-to-OR pathways. The difference is that in natural-language-to-optimization, constraints are generated as part of a broader model-formulation process, together with variables, objectives, parameters, and solver instructions.
The derived artifacts can reduce modeling effort, but it is risky because an incorrect inequality, a missing capacity constraint, or a wrong objective term can produce infeasible or misleading decisions. They require however syntax checks, semantic checks, solver verification, benchmark comparison, and expert validation before operational use (Ramamonjison et al., 2023; Ding et al., 2026). The generic optimization structure produced from text can be represented as follows:
m i n   c T x ,
A x     b ,  
x i R , x j Z , x k     { 0,1 }
In this compact representation, c T x is the objective function, A x     b expresses the system of linear constraints, and the domains of x i , x j , and x k distinguish continuous, non-negative integer, and binary decisions.
Table 11. Nomenclature of the Section’s Model Parameters and Variables.
Table 11. Nomenclature of the Section’s Model Parameters and Variables.
Nomenclature Definition
x Generic decision-variable vector in the generated optimization model (problem-specific units)
c Objective-coefficient vector multiplying the decision-variable vector x (objective units per unit of decision variable)
c T x Linear objective expression to be minimized (objective units)
A Constraint-coefficient matrix in the generated optimization model (resource units per unit of decision variable)
b Constraint-bound vector in the generated optimization model (resource units)
x i Continuous decision variable indexed by i (non-negative real variable; problem-specific units)
x j Integer decision variable indexed by j (non-negative integer variable; count or problem-specific units)
x k Binary decision variable indexed by k (binary decision variable: 0/1)

5. Evaluation Framework and Research Agenda

Section 5 follows directly from the integration logic developed in Section 4. Once text-derived artifacts enter OR models, evaluation can no longer be limited to linguistic performance. A sentiment classifier may be accurate, but its OR value depends on whether the resulting covariate improves forecast quality or inventory performance (Zhang et al. 2022). A disruption classifier may identify events correctly, but its operational value depends on whether the event record improves response time, routing, resilience, or scenario planning (Fang et al. 2026). A generated optimization formulation may be fluent and syntactically correct, but it must also be mathematically valid, feasible, and decision-relevant (Ramamonjison et al. 2023 and Ding et al. 2026). Therefore, evaluation must follow the artifact from text processing into model integration and decision outcomes.
Table 12. Evaluation levels for NLP-enhanced OR systems.
Table 12. Evaluation levels for NLP-enhanced OR systems.
Evaluation Level Main Question Example Relevant Outcome
Text-processing accuracy Did the method process language correctly? Was sentiment, topic, event, entity, relation, or retrieved passage identified correctly? Accuracy, F1, retrieval precision, extraction quality
Artifact validity Is the generated artifact operationally meaningful? Does a risk event correctly represent a real disruption, or does an ESG score represent actual sustainability? Valid covariate, valid rule, valid event, valid graph relation
Model feasibility Can the artifact enter the OR model correctly? Does a contract clause become the correct constraint, and does generated solver code run feasibly? Feasibility, consistency, correct parameterization
Decision performance Does the artifact improve the decision? Does a text-derived covariate improve inventory outcomes or service levels? Cost, service level, robustness, resilience, sustainability
Deployment reliability Can the system be used safely in practice? Are hallucinations, latency, privacy, and human oversight controlled? Traceability, trust, operational reliability

5.1. Text-Processing Accuracy

Text-processing accuracy evaluates whether the language-processing method correctly extracts, classifies, retrieves, or generates the intended information, following established NLP and information-retrieval evaluation practice (Tjong Kim Sang and De Meulder, 2003; Manning et al., 2008; Es et al., 2024). For classification tasks, such as sentiment classification, disruption classification, ESG labeling, or document categorization, relevant metrics include accuracy, precision, recall, F1-score, macro-F1, weighted-F1, ROC analysis, and AUC when probabilistic classifiers are used (Sokolova and Lapalme, 2009; Powers, 2011). For extraction tasks, such as named-entity recognition, relation extraction, clause extraction, event extraction, or knowledge-graph construction, the span-level or entity-level precision, recall, and F1 should be used, with exact- or partial-match criteria depending on the task (Tjong Kim Sang and De Meulder, 2003). For topic modeling and clustering, suitable methods include topic coherence, topic diversity, expert interpretability, cluster purity, silhouette score, and stability analysis (Röder et al., 2015). For retrieval systems, evaluation should include precision@k, recall@k, mean reciprocal rank, normalized discounted cumulative gain, and hit rate (Manning et al., 2008). For RAG systems, these retrieval metrics should be complemented by faithfulness, groundedness, answer relevance, context precision, context recall, and citation traceability (Es et al., 2024).

5.2. Artifact Validity

Artifact validity asks whether the generated object correctly represents the operational reality it is supposed to capture, following the logic of content and construct validity (Lawshe, 1975; Messick, 1995). For example, a sentiment score should represent meaningful demand-side information, as in studies that use online-review sentiment for sales forecasting (Fan et al., 2017). A disruption event should represent a real operational disturbance, as in text-mining approaches to supply-chain risk identification (Chu et al., 2020). A contract-derived rule should reflect the actual contractual obligation, as in NLP-based extraction of reporting requirements from construction contracts (Jafari et al., 2021). A graph relation should represent a real dependency between entities, as in knowledge-graph construction from engineering or domain documents (Kertkeidkachorn and Ichise, 2017; Hao et al., 2021). Jafari et al. (2021) show that extracted contractual reporting requirements need to correspond to real contract obligations before they can support time–cost prediction and simulation. Kim and Kim (2017) show how sustainability reports and news can be transformed into sustainability-related textual indicators, while Kılınç et al. (2026) focus on greenwashing signals in corporate sustainability reporting. Zheng et al. (2021) provide a knowledge-graph example in which ontology design and entity identification are used to ground hazardous-chemical graph relations in domain reality. This level is crucial because artifacts are the bridge between text and OR models. If the bridge is invalid, downstream models may produce confident but misleading decisions. This risk is especially important in high-stakes decision settings and in language-generation systems where outputs may appear fluent but contain unsupported information (Rudin, 2019; Ji et al., 2023).

5.3. Model Feasibility

Following the logic of model verification and validation, where model structure, data, implementation, and operational behavior must be checked before decision use, model feasibility evaluates whether the artifact can be integrated into the OR model without violating structural or mathematical requirements (Sargent, 2013; Law, 2015). This layer is especially important for LLM-generated formulations and solver code because fluent generated text can contain unsupported, inconsistent, or incorrect content (Ji et al., 2023). Natural-language-to-optimization studies address this issue by converting text descriptions into solver-usable logical forms or by using refinement steps to improve generated optimization models and solver code (Ramamonjison et al., 2023; Ding et al., 2026). Recent LLM-based optimization-modeling work also emphasizes structured validation to reduce error propagation in generated models (Wu et al., 2025). Solver checks, rule-based validation, benchmark comparisons, and expert review are therefore needed before operational use.

5.4. Decision Performance

Decision performance evaluates whether the integrated artifact improves downstream OR outcomes. For forecasting, the relevant outcome may be forecast accuracy, inventory cost, stockout reduction, or service level (Zhang et al. 2022). For risk and resilience, it may be timeliness, prioritization quality, response time, or robustness, (Fang et al. 2026 and Naqvi et al. 2022). For procurement, it may be compliance, supplier quality, risk reduction, or sourcing robustness (Fantoni et al., 2021; Jafari et al., 2021; Pham and Han, 2023). For sustainability, it may be better trade-offs between cost, emissions, and social responsibility (Yu et al., 2024; Kılınç et al., 2026). For RAG-based decision support, it may be reduced search effort, improved traceability, and better human judgment (Avogadri et al. 2026).

5.5. Deployment Reliability

Deployment reliability evaluates whether NLP-enhanced OR systems can be used safely, consistently, and responsibly in operational settings, beyond controlled experimental performance (Sculley et al., 2015; NIST, 2023). Even when text-processing accuracy, artifact validity, model feasibility, and decision performance are satisfactory, a system may remain unsuitable for practical use if it lacks traceability, suffers from hallucinations, responds too slowly for operational needs, exposes sensitive information, or does not include appropriate human oversight (Amershi et al., 2019; Carlini et al., 2021; Ji et al., 2023). These concerns are particularly important in contract analysis, risk monitoring, procurement, logistics, and RAG-based decision support, where text-derived artifacts may influence costly or high-stakes decisions (Rudin, 2019; NIST, 2023). For RAG systems, deployment reliability requires that retrieved evidence can be traced to its source and that generated recommendations remain grounded in the retrieved material (Lewis et al., 2020; Gao et al., 2023; Es et al., 2024). For LLM-generated optimization models, it requires safeguards against invented variables, missing constraints, or unsupported formulations (Ji et al., 2023; Sargent, 2013; Ding et al., 2026). For text-derived forecasting and risk indicators, it requires monitoring for data drift, changing terminology, and degradation of model relevance over time (Gama et al., 2014; Rabanser et al., 2019). Reliability also depends on privacy protection, access control, latency, interoperability with enterprise systems, and clear human-in-the-loop validation procedures (Amershi et al., 2019; Sculley et al., 2015; NIST, 2023).

5.6. Research Agenda

The research agenda emerging from the review has five priorities. First, future studies should report the artifact explicitly: whether the system generates covariates, risk records, rules, indicators, graph relations, evidence sets, scenarios, or model formulations (Tranfield et al., 2003; Webster and Watson, 2002; Wang and Li, 2025; Tian et al., 2026). Second, studies should evaluate artifacts at multiple levels, including linguistic quality, artifact validity, model feasibility, and decision performance (Powers, 2011; Sokolova and Lapalme, 2009; Sargent, 2013; Es et al., 2024; Ding et al., 2026). Third, more benchmark datasets are needed for text-to-OR tasks such as contract-to-constraint, news-to-scenario, ESG-to-criteria, and natural-language-to-optimization (Ramamonjison et al., 2023; Zhang et al., 2024; Ahmed and Choudhury, 2025; Wu et al., 2025). Fourth, LLM-based systems should adopt LLM-modulo architectures in which generative models act as semantic interfaces and artifact generators, while deterministic solvers, validation rules, retrieval grounding, and human experts verify outputs (Lewis et al., 2020; Sargent, 2013; Amershi et al., 2019; Ji et al., 2023; Wu et al., 2025; Ding et al., 2026). Fifth, future work should examine deployment barriers, including data integration, latency, privacy, governance, model drift, and organizational adoption (Sculley et al., 2015; Gama et al., 2014; Rabanser et al., 2019; Carlini et al., 2021; NIST, 2023). These priorities would move the field from proof-of-concept text analytics toward reliable OR decision systems (Sculley et al., 2015; NIST, 2023; Wang and Li, 2025).

6. Conclusions

This paper reviewed NLP, LLM, RAG, knowledge-graph, transformer, and generative-AI-enhanced OR applications through an artifact-centered lens. The central conclusion is that language technologies become valuable for OR only when they generate artifacts that can be validated, integrated into decision models, and evaluated through downstream operational outcomes. The review identifies several recurring pathways. Customer reviews and social media become forecasting covariates, news and disruption narratives become risk-event records and scenarios, contracts and procurement documents become rules, obligations, penalties, or constraints, ESG and sustainability reports become indicators and criteria, patents become technology topics and opportunity signals, maintenance records become fault classes and scheduling inputs, public documents become graph relations, RAG systems generate retrieved evidence and explanations. and natural-language descriptions can become candidate optimization formulations, solver code, or simulation inputs. The paper also clarifies the connection among its sections. Section 2 identifies the research gaps and contributions of the study. Section 3 defines the methodology and review framework. Section 4 maps the reviewed literature into thematic Text-to-OR streams and shows how text-derived artifacts enter OR model categories. Section 5 explains how these systems should be evaluated once artifacts affect decisions. Together, these sections support a coherent Text-to-OR logic: textual input is processed by language technologies, transformed into OR artifacts, integrated into models, used to generate decisions, and evaluated through decision-relevant outcomes. The main implication for research is that future studies should move beyond reporting which language-processing method was used. They should specify what artifact was generated, how the artifact was validated, how it entered the OR model, and how the resulting decision was evaluated. The main implication for practice is that LLM outputs should not be treated as stand-alone OR inputs (Ji et al., 2023; Sculley et al., 2015; NIST, 2023; Wu et al., 2025; Ding et al., 2026). Instead, they should be grounded, checked, and validated using deterministic rules, solvers, domain constraints, and human experts before they affect operational decisions. The study has limitations. The screening process relied on Scopus metadata rather than full-text PDFs, so some artifact and integration classifications may need refinement through full-text verification. The field is also evolving rapidly, so terminology and methods may change. Nevertheless, the artifact-centered synthesis provides clear organizing logic for future work: NLP-enhanced OR systems should be designed and evaluated not only as language systems, but as decision systems whose value depends on valid artifacts, feasible model integration, and improved operational outcomes.

References

  1. Aejas, B.; Belhi, A.; Bouras, A. Using AI to Ensure Reliable Supply Chains: Legal Relation Extraction for Sustainable and Transparent Contract Automation. In Sustainability (Switzerland); 2025. [Google Scholar] [CrossRef]
  2. Ahmed, T.; Choudhury, S. OPT2CODE: A retrieval-augmented framework for solving linear programming problems. Nat. Lang. Process. J. 2025, 13, 100185. [Google Scholar] [CrossRef]
  3. Ali, F.; El-Sappagh, S.; Kwak, D. Fuzzy ontology and LSTM-based text mining: A transportation network monitoring system for assisting travel. In Sensors (Switzerland); 2019. [Google Scholar] [CrossRef] [PubMed]
  4. Amershi, S.; Weld, D.; Vorvoreanu, M.; Fourney, A.; Nushi, B.; Collisson, P.; Suh, J.; Iqbal, S.; Bennett, P. N.; Inkpen, K.; Teevan, J.; Kikin-Gil, R.; Horvitz, E. Guidelines for human-AI interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems; 2019; Volume Article 3, pp. 1–13. [Google Scholar] [CrossRef]
  5. Angin, M.; Taşdemir, B.; Yılmaz, C.A.; Demiralp, G.; Atay, M.; Angin, P.; Dikmener, G. A RoBERTa Approach for Automated Processing of Sustainability Reports. In Sustainability (Switzerland); 2022. [Google Scholar] [CrossRef]
  6. Avogadri, S.; Alzetta, G.; Russo, D. Discovery Omnia: Dynamic RAG for enhanced patent analysis and systematic innovation. In IFIP Advances in Information and Communication Technology; 2026. [Google Scholar] [CrossRef]
  7. Bilal, H.; Rehman, A.; Aslam, M. S.; Ullah, I.; Chang, W.-J.; Kumar, N.; Almuhaideb, A. M. Hybrid TrafficAI: A generative AI framework for real-time traffic simulation and adaptive behavior modeling. IEEE Trans. Intell. Transp. Syst. 2025. [Google Scholar] [CrossRef]
  8. Brockmann, N.; Kosasih, E. E.; Brintrup, A. Supply chain link prediction on uncertain knowledge graph. ACM SIGKDD Explor. Newsl. 2022, 24(2), 124–130. [Google Scholar] [CrossRef]
  9. Caetano, R.; Oliveira, J. M.; Ramos, P. Transformer-Based Models for Probabilistic Time Series Forecasting with Explanatory Variables. Mathematics 2025, 13, 814. [Google Scholar] [CrossRef]
  10. Cao, F.; Li, R.; Cao, X. Implementation of sustainable public procurement in China: An assessment using quantitative text analysis in large-scale tender documents. Front. Environ. Sci. 2022. [Google Scholar] [CrossRef]
  11. Carlini, N.; Tramèr, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, Ú.; Oprea, A.; Raffel, C. Extracting training data from large language models. 30th USENIX Security Symposium (USENIX Security 21), 2021; pp. 2633–2650. [Google Scholar]
  12. Chakraborty, G.; Pagolu, M.; Garla, S. Analysis of unstructured data: Applications of text analytics and sentiment mining. SAS Global Forum 2014, 2014; pp. Paper 1288–2014. Available online: https://support.sas.com/resources/papers/proceedings14/1288-2014.pdf.
  13. Cho, J.; Jung, S.; Yang, K.; Kim, D.; Kim, W. Efficient task scheduling using constraints programming for enhanced planning and reliability. Appl. Sci. 2024, 14(23), Article 11396. [Google Scholar] [CrossRef]
  14. Choi, S. J.; Choi, S. W.; Kim, J. H.; Lee, E.-B. AI and text-mining applications for analyzing contractor’s risk in invitation to bid and contracts for engineering procurement and construction projects. Energies 2021. [Google Scholar] [CrossRef]
  15. Chowdhury, S.; Alzarrad, A. Applications of Text Mining in the Transportation Infrastructure Sector: A Review. In Information (Switzerland); 2023. [Google Scholar] [CrossRef]
  16. Chu, C.-Y.; Park, K.; Kremer, G. E. A global supply chain risk management framework: An application of text-mining to identify region-specific supply chain risks. In Advanced Engineering Informatics; 2020. [Google Scholar] [CrossRef]
  17. Cronbach, L. J.; Meehl, P. E. Construct validity in psychological tests. Psychol. Bull. 1955, 52(4), 281–302. [Google Scholar] [CrossRef] [PubMed]
  18. Ding, Z.; Tan, Z.; Zhang, J.; Chen, T. OR-R1: Automating modeling and solving of operations research optimization problem via test-time reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2026. [Google Scholar] [CrossRef]
  19. Ejdys, J.; Doanh, D. C.; Dufek, Z.; Ginevičius, R.; Korzynski, P. Generative AI in the manufacturing process: theoretical considerations; Sciendo, 2023. [Google Scholar] [CrossRef]
  20. Es, S.; James, J.; Espinosa-Anke, L.; Schockaert, S. RAGAs: Automated evaluation of retrieval augmented generation. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations 2024, 150–158. [Google Scholar] [CrossRef]
  21. Fan, Z.-P.; Che, Y.-J.; Chen, Z.-Y. Product sales forecasting using online reviews and historical sales data: A method combining the Bass model and sentiment analysis. J. Bus. Res. 2017. [Google Scholar] [CrossRef]
  22. Fang, J.; Su, B.; Wang, S.; Wang, B. Uncovering the risks of digital supply chains: A large language model framework for semantic identification and validation. Int. J. Prod. Econ. 2026. [Google Scholar] [CrossRef]
  23. Fantoni, G.; Coli, E.; Chiarello, F.; Apreda, R.; Dell’Orletta, F.; Pratelli, G. Text mining tool for translating terms of contract into technical specifications: Development and application in the railway sector. In Computers in Industry; 2021. [Google Scholar] [CrossRef]
  24. Gama, J.; Žliobaitė, I.; Bifet, A.; Pechenizkiy, M.; Bouchachia, A. A survey on concept drift adaptation. ACM Comput. Surv. 2014, 46(4)(Article 44), 1–37. [Google Scholar] [CrossRef] [PubMed]
  25. Gan, L.; Ye, B.; Huang, Z.; Xu, Y.; Chen, Q.; Shu, Y. Knowledge graph construction based on ship collision accident reports to improve maritime traffic safety. Ocean Coast. Manag. 2023. [Google Scholar] [CrossRef]
  26. Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-augmented generation for large language models: A survey. arXiv 2023, arXiv:2312.10997. [Google Scholar]
  27. Garg, R.; Kiwelekar, A. W.; Netak, L. D. Logistics and Freight Transportation Management: An NLP based Approach for Shipment Tracking. Pertanika J. Sci. Technol. 2021, 29(4), 2745–2765. [Google Scholar] [CrossRef]
  28. Geunes, J.; Chang, B. Operations Research Models for Supply Chain Management and Design. In Encyclopedia of Optimization; Floudas, C. A., Pardalos, P. M., Eds.; Springer, 2008. [Google Scholar]
  29. Guo, A.; Zhou, Y.; Tian, H.; Fang, C.; Sun, Y.; Sun, W.; Gao, X.; Luu, A. T.; Liu, Y.; Chen, Z. SoVAR: Build generalizable scenarios from accident reports for autonomous driving testing. In Proceedings of the 39th ACM/IEEE International Conference on Automated Software Engineering, 2024. [Google Scholar] [CrossRef]
  30. Hao, J.; Zhao, L.; Milisavljevic-Syed, J.; Ming, Z. Integrating and navigating engineering design decision-related knowledge using decision knowledge graph. In Advanced Engineering Informatics; 2021. [Google Scholar] [CrossRef]
  31. Hong, P. C.; Choi, Y. B.; Park, Y. S. AI Diffusion and the New Triad of Supply Chain Transformation: Productivity, Perspective, and Power in the Era of Claude, ChatGPT, Gemini, LLaMA, and Mistral; Multidisciplinary Digital Publishing Institute (MDPI), 2026. [Google Scholar]
  32. Huang, X.; Orth, M. R.; Barceló, P.; Bronstein, M. M.; Ceylan, İ. İ. Link prediction with relational hypergraphs. arXiv 2024, arXiv:2402.04062. [Google Scholar]
  33. Iftikhar, R.; Khan, M. S. Social Media Big Data Analytics for Demand Forecasting: Development and Case Implementation of an Innovative Framework. J. Glob. Inf. Manag. 2020, 28(1). [Google Scholar] [CrossRef]
  34. Innuphat, C.; Toahchoodee, M. The Implementation of Discrete-Event Simulation and Demand Forecasting Using Temporal Fusion Transformers to Validate Spare Parts Inventory Policy for The Petrochemicals Industry. ECTI Trans. Comput. Inf. Technol. 2022, 16(3), 247–258. [Google Scholar] [CrossRef]
  35. Jackson, I.; Saenz, M. J.; Ivanov, D. From natural language to simulations: Applying AI to automate simulation modelling of logistics systems. Int. J. Prod. Res. 2024, 62(4), 1434–1457. [Google Scholar] [CrossRef]
  36. Jacob, A.; Ben Achour, A.; Teicher, U. AI-driven risk estimation: a GPT-based approach to news monitoring for manufacturing resilience. Int. J. Adv. Manuf. Technol. 2026. [Google Scholar] [CrossRef]
  37. Jafari, P.; Al Hattab, M.; Mohamed, E.; Abourizk, S. Automated extraction and time-cost prediction of contractual reporting requirements in construction using natural language processing and simulation. In Applied Sciences; 2021. [Google Scholar] [CrossRef]
  38. Jannelli, V.; Schöpf, S.; Bickel, M.; Netland, T.; Brintrup, A. Agentic LLMs in the supply chain: towards autonomous multi-agent consensus-seeking. Int. J. Prod. Res. 2025. [Google Scholar] [CrossRef]
  39. Jayarathna, C. P.; Agdas, D.; Dawes, L. Multi-Objective Optimization for Sustainable Supply Chain and Logistics: A Review. Sustainability 2021, 13(24), 13617. [Google Scholar] [CrossRef]
  40. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55(12)(Article 248), 1–38. [Google Scholar] [CrossRef]
  41. Jiang, C.; Shu, X.; Qian, H.; Lu, X.; Zhou, J.; Zhou, A.; Yu, Y. LLMOPT: Learning to define and solve general optimization problems from scratch. International Conference on Learning Representations (ICLR), 2025. [Google Scholar]
  42. Jin, Y.; Ma, J. A survey of large language models in transportation planning: modelling, design and decision-making. In Transportmetrica A: Transport Science; 2026. [Google Scholar] [CrossRef]
  43. Kertkeidkachorn, N.; Ichise, R. T2KG: An end-to-end system for creating knowledge graph from unstructured text. AAAI Workshop Technical Report, 2017. [Google Scholar]
  44. Khlie, K.; Benmamoun, Z.; Fethallah, W.; Jebbor, I. Leveraging variational autoencoders and recurrent neural networks for demand forecasting in supply chain management: A case study. J. Infrastruct. Policy Dev. 2024, 8(8), 6639. [Google Scholar] [CrossRef]
  45. Khlie, K.; Benmamoun, Z.; Jebbor, I.; Serrou, D. Generative AI for enhanced operations and supply chain management. J. Infrastruct. Policy Dev. 2024, 8(10), 6637. [Google Scholar] [CrossRef]
  46. Kılınç, Y.; İnce, M. R.; Badem, A. C. A multi-dimensional textual framework for detecting greenwashing in sustainability reporting. In Discover Sustainability; 2026. [Google Scholar] [CrossRef]
  47. Kim, J.; Kim, E.; Kim, D.; Kim, Y.; Cheong, T. Extracting Supply Chain Information From News Articles Using Large Language Models: A Fully Automatic Approach. In IEEE Access; 2025. [Google Scholar] [CrossRef]
  48. Kim, D.; Kim, S. Sustainable supply chain based on news articles and sustainability reports: Text mining with Leximancer and DICTION; Sustainability, 2017. [Google Scholar] [CrossRef]
  49. Kim, J.; Kim, H.; Kim, H.; Lee, D.; Yoon, S. A comprehensive survey of deep learning for time series forecasting: architectural diversity and open challenges. In Artificial Intelligence Review; 2025. [Google Scholar] [CrossRef]
  50. Law, A. M. Simulation Modeling and Analysis, 5th ed.; McGraw-Hill Education, 2015; ISBN 9780073401324. [Google Scholar]
  51. Lawshe, C. H. A quantitative approach to content validity. Pers. Psychol. 1975, 28(4), 563–575. [Google Scholar] [CrossRef]
  52. Lee, J.; Kim, S.-J.; Lee, G.; Kim, K.; Jeon, J.; Ko, S.-K.; Lee, K. A Hierarchical LLM-Based Framework for Heterogeneous Multi-Robot Orchestration in High-Risk Energy Facility Maintenance. IEEE Access 2026. [Google Scholar] [CrossRef]
  53. Lee, S.-H.; Choi, S.-W.; Lee, E.-B. A Question-Answering Model Based on Knowledge Graphs for the General Provisions of Equipment Purchase Orders for Steel Plants Maintenance. In Electronics (Switzerland); 2023. [Google Scholar] [CrossRef]
  54. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.; Rocktäschel, T.; Riedel, S.; Kiela, D. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv. Neural Inf. Process. Syst. 33 2020, 9459–9474. [Google Scholar]
  55. Li, X.; Ma, J. A BERT and NSGA-II Based Model for Workforce Resource Allocation Optimization in the Operational Stage of Commercial Buildings. In Buildings (MDPI); 2026. [Google Scholar] [CrossRef]
  56. Liang, X.; Zhang, Q.; Man, Y.; He, Z. Toward sustainable process industry based on knowledge graph: a case study of papermaking process fault diagnosis. In Discover Sustainability; 2024. [Google Scholar] [CrossRef]
  57. Lin, X.; Hu, L.; Yan, Z.; Zhu, H.; Jiang, X. Intelligent decision support for tunnel fire incidents: Integrating dynamic knowledge graph with large language models. In Tunnelling and Underground Space Technology; 2026. [Google Scholar] [CrossRef]
  58. Liu, J.; Chen, L.; Luo, R.; Zhu, J. A combination model based on multi-angle feature extraction and sentiment analysis: Application to EVs sales forecasting. In Expert Systems with Applications; 2023. [Google Scholar] [CrossRef]
  59. Lyu, F.; Choi, J. The forecasting sales volume and satisfaction of organic products through text mining on web customer reviews. In Sustainability (Switzerland); 2020. [Google Scholar] [CrossRef]
  60. Ma, J.; Liu, L.; Liu, S.; Wang, H. Large Language Model-Driven Demand Forecasting and Inventory Optimization for University Physical Education Resource Supply Chain. Revista Internacional de Métodos Numéricos para Cálculo y Diseño en Ingeniería 2026, 42(2), 50. [Google Scholar] [CrossRef]
  61. Manders, T.; Klaassen, E. Unpacking the Smart Mobility Concept in the Dutch Context Based on a Text Mining Approach. In Sustainability (Switzerland); 2019. [Google Scholar] [CrossRef]
  62. Manning, C. D.; Raghavan, P.; Schütze, H. Introduction to Information Retrieval; Cambridge University Press, 2008. [Google Scholar]
  63. Mendez, J.T.; Lobel, H.; Parra, D.; Herrera, J.C. Using Twitter to Infer User Satisfaction with Public Transport: The Case of Santiago, Chile. In IEEE Access; 2019. [Google Scholar] [CrossRef]
  64. Messick, S. Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. Am. Psychol. 1995, 50(9), 741–749. [Google Scholar] [CrossRef]
  65. Meyer, A.; Walter, W.; Seuring, S. The Impact of the Coronavirus Pandemic on Supply Chains and Their Sustainability: A Text Mining Approach. Front. Sustain. 2021. [Google Scholar] [CrossRef]
  66. Naqvi, S. M. R.; Ghufran, M.; Meraghni, S.; Varnier, C.; Nicod, J.-M.; Zerhouni, N. Human knowledge centered maintenance decision support in digital twin environment. J. Manuf. Syst. 2022. [Google Scholar] [CrossRef]
  67. National Institute of Standards and Technology. Artificial intelligence risk management framework (AI RMF 1.0); U.S. Department of Commerce, 2023. [Google Scholar] [CrossRef]
  68. Nie, W.; Li, F.; Tsolakis, N.; Kumar, M. Integrating digitally enhanced data extraction and simulation modelling for AI-driven supply Chain resilience: an operational research framework for strategic stockpiling of critical minerals. J. Oper. Res. Soc. 2026. [Google Scholar] [CrossRef]
  69. Ou-Yang, C.; Chou, S.-C.; Juan, Y.-C. Improving the Forecasting Performance of Taiwan Car Sales Movement Direction Using Online Sentiment Data and CNN-LSTM Model. In Applied Sciences (Switzerland); 2022. [Google Scholar] [CrossRef]
  70. Pan, X.; Fang, J.; Wu, F.; Zhang, S.; Hu, Y.-X.; Li, S.; Li, X.-Y. Guiding Large Language Models in Modeling Optimization Problems via Question Partitioning. IJCAI International Joint Conference on Artificial Intelligence, 2025. [Google Scholar] [CrossRef] [PubMed]
  71. Park, M.-J.; Lee, E.-B.; Lee, S.-Y.; Kim, J.-H. A digitalized design risk analysis tool with machine-learning algorithm for epc contractor’s technical specifications assessment on bidding; Energies, 2021. [Google Scholar] [CrossRef]
  72. Penco, R.; Pintar, D.; Vranić, M.; Šoštarić, M. Large Language Model-Driven Framework for Automated Constraint Model Generation in Configuration Problems. In Applied Sciences (Switzerland); 2025. [Google Scholar] [CrossRef]
  73. Pham, H. T. T. L.; Han, S. Natural language processing with multitask classification for semantic prediction of risk-handling actions in construction contracts. J. Comput. Civ. Eng. 2023. [Google Scholar] [CrossRef]
  74. Powers, D. M. W. Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. J. Mach. Learn. Technol. 2011, 2(1), 37–63. [Google Scholar]
  75. Prasad, R.; Udeme, A.U.; Misra, S.; Bisallah, H. Identification and classification of transportation disaster tweets using improved bidirectional encoder representations from transformers. Int. J. Inf. Manag. Data Insights 2023. [Google Scholar] [CrossRef]
  76. Rabanser, S.; Günnemann, S.; Lipton, Z. C. Failing loudly: An empirical study of methods for detecting dataset shift. Adv. Neural Inf. Process. Syst. 32 2019. [Google Scholar] [CrossRef]
  77. Ramamonjison, R.; Yu, T. T.; Li, R.; Li, H.; Carenini, G.; Ghaddar, B.; He, S.; Mostajabdaveh, M.; Banitalebi-Dehkordi, A.; Zhou, Z.; Zhang, Y. NL4Opt competition: Formulating optimization problems based on their natural language descriptions. Proceedings of Machine Learning Research, 2023. [Google Scholar]
  78. Röder, M.; Both, A.; Hinneburg, A. Exploring the space of topic coherence measures. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining (WSDM 2015); 2015; pp. 399–408. [Google Scholar] [CrossRef]
  79. Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 1 2019, 206–215. [Google Scholar] [CrossRef] [PubMed]
  80. Sadeek, S. N.; Hanaoka, S. Assessment of text-generated supply chain risks considering news and social media during disruptive events. Soc. Netw. Anal. Min. 2023, 13, 96. [Google Scholar] [CrossRef]
  81. Sargent, R. G. Verification and validation of simulation models. J. Simul. 2013, 7(1), 12–24. [Google Scholar] [CrossRef]
  82. Schöpper, H.; Kersten, W. Using natural language processing for supply chain mapping: A systematic review of current approaches. CEUR Workshop Proceedings, 2021. [Google Scholar]
  83. Sculley, D.; Holt, G.; Golovin, D.; Davydov, E.; Phillips, T.; Ebner, D.; Chaudhary, V.; Young, M.; Crespo, J.-F.; Dennison, D. Hidden technical debt in machine learning systems. Adv. Neural Inf. Process. Syst. 28 2015, 2503–2511. [Google Scholar]
  84. Shao, J.; Hong, J.; Wang, M.; Wang, X. New energy vehicles sales forecasting using machine learning: The role of media sentiment. Comput. Ind. Eng. 2025. [Google Scholar] [CrossRef]
  85. Sharma, A.; Adhikary, A.; Borah, S. B. COVID-19’s impact on supply chain decisions: Strategic insights from NASDAQ 100 firms using Twitter data. J. Bus. Res. 2020. [Google Scholar] [CrossRef] [PubMed]
  86. Singh, A.; Shukla, N.; Mishra, N. Social media data analytics to improve supply chain management in food industries. In Transportation Research Part E: Logistics and Transportation Review; 2018. [Google Scholar] [CrossRef]
  87. Snyder, H. Literature review as a research methodology: An overview and guidelines. J. Bus. Res. 104 2019, 333–339. [Google Scholar] [CrossRef]
  88. Sokolova, M.; Lapalme, G. A systematic analysis of performance measures for classification tasks. Inf. Process. Manag. 2009, 45(4), 427–437. [Google Scholar] [CrossRef]
  89. Son, B.-Y.; Lee, E.-B. Using text mining to estimate schedule delay risk of 13 offshore oil and gas EPC Case Studies during the Bidding Process. Energies 2019. [Google Scholar] [CrossRef]
  90. Song, H.; Yang, Z.; Du, H.; Zhang, Y.; Zeng, J.; He, X. LLM-LCSA: LLM for collaborative control and decision optimization in UAV cluster security. Drones 2025, 9(11), Article 779. [Google Scholar] [CrossRef]
  91. Sun, M.; Tian, Y.; Li, J.; Wu, C.-L.; Peng, L.; Xu, S. A review of network delay prediction and advances in large language models for air traffic. Artif. Intell. Rev. 2026. [Google Scholar] [CrossRef]
  92. Szekely, N.; Vom Brocke, J. What can we learn from corporate sustainability reporting? Deriving propositions for research and practice from over 9,500 corporate sustainability reports using topic modelling. PLoS ONE 2017. [Google Scholar] [CrossRef] [PubMed]
  93. Tian, J.; Chen, L.; Islam, N.; Jia, F. Applying text mining in operations and engineering management: Guidelines and challenges in sentiment analysis and topic modeling. In IEEE Transactions on Engineering Management; 2026. [Google Scholar] [CrossRef]
  94. Tian, X.; Gan, H.; Liu, Y. Construction of Knowledge Graph for Marine Diesel Engine Faults Based on Deep Learning Methods. J. Mar. Sci. Eng. 2025, 13(4), 693. [Google Scholar] [CrossRef]
  95. Tjong Kim Sang, E. F.; De Meulder, F. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. Proceedings of CoNLL 2003, 2003; pp. 142–147. [Google Scholar]
  96. Torraco, R. J. Writing integrative literature reviews: Guidelines and examples. Hum. Resour. Dev. Rev. 2005, 4(3), 356–367. [Google Scholar] [CrossRef]
  97. Tranfield, D.; Denyer, D.; Smart, P. Towards a methodology for developing evidence-informed management knowledge by means of systematic review. Br. J. Manag. 2003, 14(3), 207–222. [Google Scholar] [CrossRef]
  98. Tupayachi, J.; Xu, H.; Omitaomu, O. A.; Camur, M. C.; Sharmin, A.; Li, X. Towards next-generation urban decision support systems through AI-powered construction of scientific ontology using large language models: A case in optimizing intermodal freight transportation. Smart Cities 2024, 7(5), 2392–2421. [Google Scholar] [CrossRef]
  99. Vasquez-Henriquez, P.; Graells-Garrido, E.; Caro, D. Tweets on the go: Gender differences in transport perception and its discussion on social media. In Sustainability (Switzerland); 2020. [Google Scholar] [CrossRef]
  100. Voloshchuk, V.; Melnik, Y.; Safronenkova, I.; Lishchenko, E.; Kartashov, O.; Kozlovskiy, A. Hybrid method of organizing information search in logistics systems based on vector-graph structure and large language models. Big Data Cogn. Comput. 2026, 10(2), 51. [Google Scholar] [CrossRef]
  101. Wang, X.; Yuen, K. F.; Wong, Y. D.; Li, K. X. How can the maritime industry meet Sustainable Development Goals? An analysis of sustainability reports from the social entrepreneurship perspective. In Transportation Research Part D: Transport and Environment; 2020. [Google Scholar] [CrossRef]
  102. Wang, Y.; Li, K. Large Language Models in Operations Research: Methods, Applications, and Challenges. arXiv 2025, arXiv:2509.18180. [Google Scholar] [CrossRef]
  103. Wang, Y.; Zhang, Y. Multivariate SVR Demand Forecasting for Beauty Products Based on Online Reviews. Mathematics 2023, 11, 4420. [Google Scholar] [CrossRef]
  104. Wang, Z.; Chen, B.; Huang, Y.; Cao, Q.; He, M.; Fan, J.; Liang, X. ORMind: A Cognitive-Inspired End-to-End Reasoning Framework for Operations Research; Association for Computational Linguistics (ACL), 2025. [Google Scholar]
  105. Webster, J.; Watson, R. T. Analyzing the past to prepare for the future: Writing a literature review. MIS Q. 2002, 26(2), xiii–xxiii. [Google Scholar] [CrossRef]
  106. Wolfswinkel, J. F.; Furtmueller, E.; Wilderom, C. P. M. Using grounded theory as a method for rigorously reviewing literature. Eur. J. Inf. Syst. 2013, 22(1), 45–55. [Google Scholar] [CrossRef]
  107. Wu, Y.; Zhang, Y.; Wu, Y.; Wang, Y.; Zhang, J.; Cheng, J. Training LLMs for Optimization Modeling via Iterative Data Synthesis and Structured Validation; Association for Computational Linguistics (ACL), 2025. [Google Scholar]
  108. Xie, Z.; Zhu, R.; Zhang, M.; Liu, J. SparseMult: A sparse tensor decomposition model for knowledge graph link prediction. Comput. Intell. 2025, 41(4), e70097. [Google Scholar] [CrossRef]
  109. Yang, M.-R.; Xu, X.-J. Recent advances in hypergraph neural networks. In Journal of the Operations Research Society of China; 2025. [Google Scholar] [CrossRef]
  110. Yang, Y.; Wang, M.; Wang, J.; Li, P.; Zhou, M. Multi-Agent Deep Reinforcement Learning for Integrated Demand Forecasting and Inventory Optimization in Sensor-Enabled Retail Supply Chains. Sensors 2025, 25, 2428. [Google Scholar] [CrossRef] [PubMed]
  111. Yu, K.; Wu, Q.; Chen, X.; Wang, W.; Mardani, A. An integrated MCDM framework for evaluating the environmental, social, and governance (ESG) sustainable business performance. Ann. Oper. Res. 2024, 342(1), 987–1018. [Google Scholar] [CrossRef]
  112. Yuan, D.; Zhou, K.; Yang, C. Architecture and Application of Traffic Safety Management Knowledge Graph Based on Neo4j. In Sustainability (Switzerland); 2023. [Google Scholar] [CrossRef]
  113. Zakaria, M.; Aoun, C.; Liginlal, A.D. Objective sustainability assessment in the digital economy: An information entropy measure of transparency in corporate sustainability reporting. In Sustainability (Switzerland); 2021. [Google Scholar] [CrossRef]
  114. Zayet, T. M. A.; Ismail, M. A.; Varathan, K. D.; Noor, R. M. D.; Chua, H. N.; Lee, A.; Low, Y. C.; Singh, S. K. J. Investigating transportation research based on social media analysis: a systematic mapping review; Scientometrics, 2021. [Google Scholar] [CrossRef] [PubMed]
  115. Zhang, H.; Geng, X.; Zhang, Y.; Du, Y.; Li, D. A novel methodological framework for evaluating and optimizing PSS solutions based on user reviews and intelligent technologies. Adv. Eng. Inform. 2026. [Google Scholar] [CrossRef]
  116. Zhang, J.; Wang, W.; Guo, S.; Wang, L.; Lin, F.; Yang, C.; Yin, W. Solving General Natural-Language-Description Optimization Problems with Large Language Models; Association for Computational Linguistics (ACL), 2024. [Google Scholar]
  117. Zhang, Y.; Chen, Y.; Wang, J. Forecasting sales using online review and search engine data: A method based on PCA–DSFOA–BPNN. Int. J. Forecast. 2022. [Google Scholar] [CrossRef]
  118. Zhao, M.; Hussain, O.; Zhang, Y.; Saberi, M.; Leshob, A. Benchmarking large language models for supply chain risk identification: an extended evaluation within the LARD-SC framework. In Service Oriented Computing and Applications; 2025. [Google Scholar]
  119. Zheng, G.; Brintrup, A. Enhancing supply chain visibility with generative AI: An exploratory case study on relationship prediction in knowledge graphs. Int. J. Prod. Res. 2025. [Google Scholar] [CrossRef]
  120. Zheng, W.; Yang, M.; Ren, Y.; Wang, H.; Zeng, C.; Zhang, Y. CARAG: Context-Aware Retrieval-Augmented Generation for Railway Operation and Maintenance Question Answering over Spatial Knowledge Graph. ISPRS Int. J. Geo-Inf. 2026. [Google Scholar] [CrossRef]
  121. Zheng, X.; Wang, B.; Zhao, Y.; Mao, S.; Tang, Y. A knowledge graph method for hazardous chemical management: Ontology design and entity identification. Neurocomputing 2021. [Google Scholar] [CrossRef]
Figure 1. Artifact-centered Text-to-OR review framework.
Figure 1. Artifact-centered Text-to-OR review framework.
Preprints 222376 g001
Table 3. Functional coding scheme.
Table 3. Functional coding scheme.
Coding Dimension Definition Purpose in This Review
Author-year Short citation of the study Identifies the study and supports traceability.
OR domain Operational or decision context addressed Identifies logistics, supply chain, transportation, inventory, sustainability, construction, manufacturing, or another OR-relevant domain.
Textual or semi-structured source Raw language-based input Captures reviews, news, contracts, ESG reports, patents, policies, social media, maintenance records, prompts, or natural-language queries.
Language-processing method Computational method used Distinguishes text mining, sentiment analysis, topic modeling, NER, relation extraction, transformers, LLMs, RAG, knowledge graphs, and semantic search.
OR task supported Operational or mathematical task supported Links language processing to forecasting, simulation, optimization, routing, scheduling, inventory, risk, sustainability, or decision support.
OR artifact generated Structured output produced from text Identifies covariates, risk-event records, parameters, constraints, scenarios, rules, graph relations, solver code, explanations, or recommendations.
Model integration mechanism How artifact connects to OR model Captures preprocessing, sequential integration, simulation parameterization, graph analysis, RAG interface, solver integration, or human validation.
Evaluation approach Metrics or validation procedures Distinguishes linguistic accuracy from OR outcomes such as feasibility, cost, service level, robustness, sustainability, and decision quality.
Reported limitations Main barriers or weaknesses Supports synthesis of challenges such as hallucination, ambiguity, weak integration, privacy, latency, and error propagation.
Table 4. Thematic streams identified in the coded corpus.
Table 4. Thematic streams identified in the coded corpus.
Thematic Stream Included studies Representative Papers
Review, taxonomy, framework, and general Text-to-OR decision-support studies 357 Zhang et al., 2026; Tian et al., 2026; Chowdhury and Alzarrad, 2023; Zayet et al., 2021; Schöpper and Kersten, 2021; Mendez et al., 2019;
Customer Reviews/social-media-to-demand and forecasting signals 223 Fan et al., 2017; Zhang et al., 2022; Shao et al., 2025; Lyu and Choi, 2020; Ou-Yang et al., 2022; Wang and Zhang, 2023
News/social-media-to-risk, disruption, and resilience 155 Sharma et al., 2020; Singh et al., 2018; Fang et al., 2026; Meyer et al., 2021; Prasad et al., 2023; Kim et al., 2025
Policy/report/ESG text-to-sustainability indicators 270 Kim and Kim, 2017; Wang et al., 2020; Liu et al., 2023; Szekely and Vom Brocke, 2017; Kılınç et al., 2026; Angin et al., 2022; Zakaria et al., 2021
Contract/procurement text-to-rules, risks, or supplier decisions 157 Fantoni et al., 2021; Choi et al., 2021; Jafari et al., 2021; Son and Lee, 2019; Cao et al., 2022; Choi and Lee, 2022; Aejas et al., 2025; Park et al., 2021
Public text/documents-to-knowledge graph or ontology 201 Kertkeidkachorn and Ichise, 2017; Gan et al., 2023; Hao et al., 2021; Ali et al., 2019; Choi and Lee, 2022; Yuan et al., 2023; Liang et al., 2024
Query-to-RAG and semantic decision support 140 Avogadri et al., 2026; Lin et al., 2026; Lee et al., 2023; Zheng et al., 2026
Natural-language-to-optimization, solver code, and simulation 134 Ramamonjison et al., 2023; Yang et al., 2026; Ding et al., 2026; Penco et al., 2025; Li and Ma, 2026; Pan et al., 2025; Ahmed and Choudhury, 2025; Jiang et al., 2025
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.