Preprint
Article

This version is not peer-reviewed.

From Predictive Accuracy to Human-Centric Decision Support: A Perspective on Machine Learning for Safer Intelligent Transportation Systems

Submitted:

21 July 2026

Posted:

22 July 2026

You are already at the latest version

Abstract
Machine learning research in intelligent transportation systems is still largely evaluated through predictive accuracy, often treated as the primary indicator of model quality and practical value. Although this emphasis has supported substantial technical progress, it overlooks critical human and societal concerns, including interpretability, unequal impacts across user groups, uncertainty in high-stakes decisions, and the limited ability of practitioners to translate predictions into safe interventions. This Perspective proposes a human-centric decision-support framework for machine learning in safer intelligent transportation systems. The framework retains task-appropriate predictive performance as a prerequisite and complements it with six human-centric dimensions: relevance, explainability, fairness, uncertainty, human oversight, and actionability. Together, these dimensions provide a basis for evaluating whether predictive systems are not only accurate, but also understandable, trustworthy, and usable in real operational contexts. The framework has direct implications for researchers designing transport models, municipalities deploying data-driven safety policies, and transport operators integrating automated recommendations into daily decision-making. It argues that progress in intelligent mobility should be measured not only by prediction quality, but also by the quality of the decisions that predictive systems enable.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Machine learning has become a central component of intelligent transportation systems (ITS), supported by traffic sensors, connected vehicles, mobile devices, and increasingly accessible urban data. ITS research encompasses traffic management, autonomous vehicles, mobility prediction, and other smart-city applications [1]. Predictive models are also used for traffic-collision prediction [2], travel and arrival-time forecasting [3], driver-state and fatigue monitoring [4], and driving-behaviour analysis [5]. Progress across these applications is commonly expressed through gains in accuracy, precision, recall, or reductions in prediction error. These measures are indispensable for technical comparison, but their prominence has encouraged a narrow account of what makes a transportation model successful.
Predictive performance alone does not establish that a system is useful, safe, or socially acceptable. A highly accurate model may be difficult for traffic managers to interpret, perform unevenly across road-user groups or neighbourhoods, or provide overconfident outputs under unfamiliar conditions. Its practical value may also remain limited when a prediction does not clarify what action should follow or how uncertainty should influence that action. These weaknesses are consequential in road-safety settings, where model outputs can affect infrastructure priorities, operational interventions, public warnings, and emergency responses. Human-centered artificial intelligence (AI) consequently calls for systems that combine computational capability with meaningful human control, transparency, and sensitivity to context [6]. Research on explainable artificial intelligence similarly stresses that explanations should be evaluated according to the needs of intended users rather than treated only as technical properties of a model [7].
This Perspective argues that machine learning for ITS should be evaluated not only by how accurately it predicts an event, but also by how effectively it supports safer and more accountable decisions. It introduces the Human-Centric and Trustworthy Machine Learning (HCT-ML) framework, structured around relevance, explainability, fairness, uncertainty, human oversight, and actionability. Whereas cross-sector responsible AI frameworks define general lifecycle requirements [8] and transportation-specific guidance identifies domain risks [9], HCT-ML organizes these concerns around transport decision owners, operational time horizons, intervention pathways, and the consequences of acting on uncertain predictions. The contribution is not a new algorithm or validated scoring instrument, but a transport-oriented conceptual framework connecting model development, operational use, and affected stakeholders.
The remainder of the article explains why accuracy-centered evaluation is insufficient, examines the principal human stakeholders, presents the HCT-ML framework, illustrates its application to urban road safety, and outlines a research and deployment agenda.

2. Why Accuracy-Centric Machine Learning Is Insufficient

Accuracy-centered evaluation is useful for comparing models, but it provides only a partial account of their suitability for intelligent transportation systems. Road-safety datasets can contain substantial class imbalance, allowing strong aggregate accuracy to coexist with weak recognition of underrepresented outcomes [2]. In congestion and public transport forecasting, low average error can likewise conceal failures during disruptions, extreme weather, or unusual demand.
The first limitation is interpretability. Many high-performing models provide little insight into why a collision, traffic condition, or driving manoeuvre has been classified as hazardous. In autonomous driving, a correct control action is insufficient when engineers, safety assessors, or users cannot identify the environmental cues that shaped it. Explainability is therefore relevant to regulatory compliance and social acceptance in autonomous driving [10]. A municipal risk score is similarly difficult to use when professionals cannot determine whether it reflects road geometry, weather, exposure, behaviour, or an artefact of the data.
A second limitation is the gap between offline evaluation and real-world use. Historical test sets represent bounded conditions, whereas deployed systems encounter sensor failures, construction zones, changing mobility patterns, unfamiliar road layouts and weather conditions that were underrepresented in the training data. The benchmark A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation (SHIFT) illustrates how changes in illumination, precipitation, traffic density, and pedestrian activity can degrade autonomous-driving perception [11]. Robustness, calibration, latency, and distribution shift must therefore be assessed throughout deployment.
Aggregate accuracy may also conceal unequal outcomes. Uneven sensing, incomplete reporting, and historically unequal access to transport services can cause some communities to be represented more reliably than others. Travel-behaviour research has documented disparities associated with ethnicity, income, disability, and geography [12]. Fairness-aware travel-demand forecasting likewise shows that accurate aggregate predictions can coexist with biases across protected attributes [13]. When predictions guide service allocation, infrastructure investment, or enforcement, modelling disparities can become policy disparities.
Finally, models are often developed without sufficient attention to those responsible for acting on them. Traffic controllers, transit operators, emergency personnel, engineers, and planners work within legal and operational constraints that cannot be reduced to an objective function. They must know when to trust or override a recommendation and who remains accountable. Responsible AI frameworks therefore treat governance and human oversight as lifecycle requirements [8]. The relevant question is not only whether a model predicts accurately, but whether it enables informed, equitable, and accountable action.
Taken together, these limitations suggest that stakeholder-specific requirements, rather than aggregate model performance alone, should provide the starting point for the design and evaluation of transportation decision-support systems.

3. Human Stakeholders in Intelligent Transportation

Intelligent transportation systems operate within a network of people and institutions whose needs cannot be represented by one performance objective. Some stakeholders interact directly with predictive technologies, while others experience the consequences of decisions based on their outputs. Human-centered design must therefore consider both who uses the model and who bears the risk when it is incomplete or wrong.
Drivers receive collision warnings, route recommendations, fatigue alerts, and automated driving assistance. Information must arrive in time to support action without increasing distraction or cognitive load. A technically accurate warning may still fail if its meaning is unclear, its timing is poor, or repeated false alarms encourage disregard. Pedestrians and cyclists often do not interact with the model directly, yet vehicle perception, signal control, and infrastructure prioritization can affect them substantially. This concern is especially important because vulnerable road users account for more than half of global road deaths [14].
People with reduced mobility have additional requirements concerning accessibility, reliability, and independence. Route planning, demand-responsive services, automated vehicles, and passenger information must account for physical, sensory, and cognitive constraints obscured by average-user assumptions. Transport policy should therefore examine access to opportunities rather than movement alone [15].
Operational stakeholders require different forms of support. Transit operators use demand forecasts, arrival predictions, disruption alerts, and vehicle-allocation recommendations under time pressure. Emergency services use incident detection, severity estimation, and dynamic routing while retaining authority to respond to conditions absent from the data. Municipal officials and road engineers use predictive evidence to prioritize inspections, redesign streets, adjust signals, and allocate safety investments. Public decision-makers define policy objectives, procurement rules, accountability mechanisms, and acceptable trade-offs among safety, efficiency, accessibility, privacy, and equity. The same prediction consequently requires different explanations, confidence thresholds, and human oversight depending on whether it is addressed to a driver, operator, engineer, or public authority. Transportation-specific AI risk guidance likewise emphasizes the interaction of technical, organizational, and human factors [9]; these differentiated stakeholder needs motivate a framework that links model evaluation to users, decisions, risks, and institutional responsibilities.

4. The Proposed HCT-ML Framework

The Human-Centric and Trustworthy Machine Learning (HCT-ML) framework evaluates whether a technically adequate model can support a sound transportation decision. It treats the model as one component of a wider system involving data, interfaces, institutional procedures, and professional judgement. Predictive performance remains a prerequisite: if p is a task-appropriate performance measure, a candidate system should first satisfy p τ p . Conditional on this technical gate, the human-centric profile is
h = [ R , E , F , U , O , A ] ,
where R, E, F, U, O, and A denote relevance, explainability, fairness, uncertainty, human oversight, and actionability. An optional context-specific Human-Centric and Trustworthy (HCT) synthesis may be expressed as
S HCT = i = 1 6 w i h i , w i 0 , i = 1 6 w i = 1 ,
where h i [ 0 , 1 ] , w i is a stakeholder-defined weight, and τ i is the minimum acceptable value for dimension i. Requiring h i τ i prevents strength in one dimension from compensating for an unacceptable weakness in another. Figure 1 summarizes the framework.
Equation (1) is a deliberative aid rather than a validated measurement instrument. It does not imply that the dimensions are intrinsically commensurable. Each dimension should be assessed with purpose-specific evidence, including user evaluations, subgroup audits, calibration analyses, and governance reviews. Stakeholders should document the adopted weights and thresholds. Until empirically validated, the synthesis should not be used to rank unrelated transportation systems.
Relevance is the fit between a model, its intended decision, and the people affected. The evaluation question is whether the output addresses a defined operational or public need. Without this alignment, a technically strong model may provide information that cannot guide practice. A citywide collision-risk map is of limited value if it omits the relevant road unit, time horizon, or feasible intervention. Designers should define the decision owner, target population, intervention window, and error consequences before model selection.
Explainability concerns whether an intended user can understand an output’s basis, limits, and practical meaning. The question is whether the explanation is understandable and useful to the intended user [7,10]. Without it, operators may reject useful advice or rely on it uncritically. A transit control centre should know whether a disruption forecast reflects demand, vehicle delay, weather, or missing data. Explanations should match the task and user expertise and be evaluated with representative users.
Fairness concerns how errors, benefits, and burdens are distributed across people, modes, and places. Its guiding question is who is systematically helped, overlooked, or exposed to risk. Aggregate accuracy may conceal poorer performance across socially disadvantaged groups and geographic populations [12,13]. Evaluation should include disaggregated errors, geographic and data-coverage audits, and consultation with affected communities. Trade-offs should be reported as policy choices rather than hidden within one metric.
Uncertainty communicates how much confidence a prediction warrants and why doubt remains. The question is whether the system identifies noisy, unfamiliar, or weakly represented conditions. Without this information, probability may be mistaken for certainty. In automated driving, low-confidence perception during heavy snow should trigger a conservative response. Designers should distinguish data-related and model-related uncertainty where feasible [16], assess calibration and distribution shift, and connect confidence levels to warning, fallback, or human review.
Human oversight defines the authority through which people supervise, question, or override automated outputs. The question is who remains responsible when the model appears wrong. An emergency-routing system may recommend the statistically fastest route when responders know that access is blocked. Human oversight should reflect risk and urgency through an explicit allocation of functions and levels of automation [17]. Documented roles, human–AI oversight arrangements, monitoring, escalation procedures, and training should be established within lifecycle governance [8].
Actionability concerns whether a prediction can become a feasible, proportionate, and timely response. Its question is what the intended user can reasonably do within legal, financial, and operational constraints. A congestion forecast issued after queues form, or a collision-risk estimate with no modifiable factor, has limited value. Models should be co-designed around decision workflows and evaluated partly through the interventions they enable.
Table 1 consolidates the evaluation question, principal risk, and recommended design response associated with each HCT-ML dimension.
Together, these dimensions identify the conditions under which a technically adequate prediction can support a safe and accountable decision. Their importance will vary across driver-warning systems, municipal planning tools, and automated-vehicle controllers; HCT-ML makes those contextual requirements explicit without imposing a universal balance.

5. Illustrative Application to Urban Road Safety

The framework can be illustrated through a hypothetical municipal collision-severity prediction system in Montréal, where the city publishes police-reported collision data and pursues a Vision Zero road-safety strategy [18,19]. This scenario is not an empirical evaluation and does not claim the performance of a particular model; it shows how HCT-ML could guide the design and governance of a tool supporting urban road-safety decisions.
Figure 2 summarizes the closed municipal decision cycle through which prediction, human validation, intervention, monitoring, and model revision are connected.
The system could combine historical collision records with road configuration, surface condition, weather, lighting, traffic exposure, time of day, and the presence of vulnerable road users. Its primary user would be a municipal operator identifying locations and conditions associated with a heightened probability of severe injury. Relevance would require a precise decision boundary. The model might prioritize field inspections, temporary measures, or engineering assessments, but it should not automatically select infrastructure investments.
The interface would explain the factors contributing to each estimate. A high predicted severity might be associated with poor lighting, complex turning movements, winter conditions, or pedestrian exposure. Explanations should distinguish predictive association from causal evidence and use terms that municipal staff can connect to possible interventions. Uncertainty should accompany every prediction. Limited collision history, incomplete sensing, or unfamiliar conditions should lead to additional review rather than a definitive ranking. Actionability would be established by linking each risk estimate to predefined intervention options, responsible municipal units, and response time windows.
Fairness assessment would examine performance across boroughs, road types, and user groups, with particular attention to pedestrians, cyclists, older adults, and people with reduced mobility. Before infrastructure modification, road engineers, safety personnel, and, where appropriate, community representatives would review the prediction against site conditions, feasibility, legal constraints, and unintended effects. The model would support professional judgement rather than replace it.
After deployment, monitoring would assess calibration drift, mobility changes, geographic disparities, and reporting shifts. Governance rules would allow the municipality to update, restrict, or suspend the model when its assumptions no longer hold.

6. Research and Deployment Agenda

The research agenda is to operationalize HCT-ML through disaggregated evaluation, participatory testing, and lifecycle governance. Global metrics should be complemented with results across road-user groups, geographic areas, modes, environmental conditions, and severity levels. Transportation models should be examined separately across socially and geographically differentiated groups because aggregate performance can conceal disparities associated with protected attributes and place [12,13].
Explainability should be treated as a design requirement rather than an addition after model development. Researchers should identify who receives the explanation, what decision that person must make, and how much time is available. A traffic engineer considering a long-term intervention requires a different explanation from an operator responding to a disruption. Explanation quality should therefore be assessed with intended users and within the decision context in which the explanation will be used [7].
Operational evaluation must extend beyond retrospective datasets. Transit operators, emergency personnel, planners, and engineers should participate in scenario-based tests, simulations, and controlled field trials. These can reveal whether predictions arrive on time, create cognitive burden, or conflict with staffing, legal, maintenance, and network constraints. Documentation should state intended use, data provenance, known limitations, excluded populations, calibration, uncertainty, and conditions that degrade performance. It should also distinguish prediction from causal inference and describe the consequences of false alarms and missed events. These evaluations should combine technical assurance with organizational and human-factors assessment [8,9].
Deployment should begin, not end, evaluation. Monitoring, escalation, and retirement criteria should be treated as lifecycle governance requirements [8,9]. Results should remain disaggregated so that disparities are not hidden by stable averages. Human recourse must also remain available throughout the lifecycle. Operators should be able to question, override, and escalate recommendations, while affected communities need channels for reporting recurring errors or harms. Override decisions should be documented without discouraging responsible professional judgement. Together, these practices shift transportation research from isolated model optimization toward accountable and correctable decision support.

7. Conclusions

Machine learning will contribute to safer ITS when predictive accuracy is assessed alongside the conditions under which outputs are produced, interpreted, and acted upon. Accurate models alone cannot ensure transportation decisions are understandable, equitable, operationally useful, or safe.
This Perspective introduced HCT-ML to connect technical performance with human and institutional requirements through six dimensions: relevance, explainability, fairness, uncertainty, human oversight, and actionability. The framework does not diminish predictive performance; it situates it within trustworthy and publicly valuable decision support.
Future work should operationalize and validate these dimensions across transportation modes and institutional contexts. Models should be evaluated with intended users, monitored after deployment, and revised when assumptions no longer hold. Progress should be judged by whether systems enable safer, fairer, and more accountable mobility decisions.

Author Contributions

B.M. conceived the perspective, developed the HCT-ML framework, reviewed the literature, and wrote and approved the manuscript.

Funding

This research received no external funding.

Conflicts of Interest

The author declare no conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AI Artificial intelligence
HCT Human-Centric and Trustworthy
HCT-ML Human-Centric and Trustworthy Machine Learning
ITS Intelligent transportation systems
SHIFT A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation

References

  1. Elassy, M.; Al-Hattab, M.; Takruri, M.; Badawi, S. Intelligent transportation systems for sustainable smart cities. Transp. Eng. 2024, 16, 100252. [Google Scholar] [CrossRef]
  2. Niture, N.; Abdellatif, I. A systematic review of factors, data sources, and prediction techniques for earlier prediction of traffic collision using AI and machine learning. Multimed. Tools Appl. 2025, 84, 19009–19037. [Google Scholar] [CrossRef]
  3. Abdi, A.; Amrit, C. A review of travel and arrival-time prediction methods on road networks: Classification, challenges and opportunities. PeerJ Comput. Sci. 2021, 7, e689. [Google Scholar] [CrossRef] [PubMed]
  4. AL-Quraishi, M.S.; Ali, S.S.A.; AL-Qurishi, M.; Tang, T.B.; Elferik, S. Technologies for detecting and monitoring drivers’ states: A systematic review. Heliyon 2024, 10, e39592. [Google Scholar] [CrossRef] [PubMed]
  5. Mobini Seraji, M.H.; Shaffiee Haghshenas, S.; Shaffiee Haghshenas, S.; Simic, V.; Pamucar, D.; Guido, G.; Astarita, V. A state-of-the-art review on machine learning techniques for driving behavior analysis: Clustering and classification approaches. Complex Intell. Syst. 2025, 11, 386. [Google Scholar] [CrossRef]
  6. Shneiderman, B. Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. Int. J. Human–Computer Interact. 2020, 36, 495–504. [Google Scholar] [CrossRef]
  7. Rong, Y.; Leemann, T.; Nguyen, T.T.; Fiedler, L.; Qian, P.; Unhelkar, V.; Seidel, T.; Kasneci, G.; Kasneci, E. Towards human-centered explainable AI: A survey of user studies for model explanations. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 2104–2122. [Google Scholar] [CrossRef] [PubMed]
  8. Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). Technical Report NIST AI 100-1, National Institute of Standards and Technology, Gaithersburg, MD, 2023. [CrossRef]
  9. Yu, H.; Hulse, D.; Irshad, L. Understanding AI Risks in Transportation: AI Assurance for Transportation Whitepaper Series. Technical report, U.S. Department of Transportation, Highly Automated Systems Safety Center of Excellence, Washington, DC, 2024.
  10. Atakishiyev, S.; Salameh, M.; Yao, H.; Goebel, R. Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions. IEEE Access 2024, 12, 101603–101625. [Google Scholar] [CrossRef]
  11. Sun, T.; Segu, M.; Postels, J.; Wang, Y.; Van Gool, L.; Schiele, B.; Tombari, F.; Yu, F. SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp. 21371–21382. [Google Scholar]
  12. Zheng, Y.; Wang, S.; Zhao, J. Equality of opportunity in travel behavior prediction with deep neural networks and discrete choice models. Transp. Res. Part C Emerg. Technol. 2021, 132, 103410. [Google Scholar] [CrossRef]
  13. Zhang, X.; Ke, Q.; Zhao, X. Travel demand forecasting: A fair AI approach. IEEE Trans. Intell. Transp. Syst. 2024, 25, 14611–14627. [Google Scholar] [CrossRef]
  14. World Health Organization. Global Status Report on Road Safety 2023; World Health Organization: Geneva, 2023. [Google Scholar]
  15. OECD/ITF. Sustainable Accessibility for All. In Technical report; OECD Publishing: Paris, 2024. [Google Scholar] [CrossRef]
  16. Kendall, A.; Gal, Y. What uncertainties do we need in Bayesian deep learning for computer vision? Proc. Adv. Neural Inf. Process. Syst. 2017, Vol. 30, 5574–5584. [Google Scholar]
  17. Parasuraman, R.; Sheridan, T.B.; Wickens, C.D. A model for types and levels of human interaction with automation. IEEE Trans. Syst. Man. Cybern.-Part A Syst. Hum. 2000, 30, 286–297. [Google Scholar] [CrossRef] [PubMed]
  18. Ville de Montréal. Collisions routières. Données ouvertes de la Ville de Montréal, Accessed 14 July 2026.
  19. de Montréal, Ville. Vision Zero: Getting Around Safely on Foot, by Bike and by Car, 2025. Updated 16 May 2025; accessed 14 July 2026.
Figure 1. The HCT-ML framework extends conventional predictive evaluation through six dimensions that connect model performance to human-centric decision support.
Figure 1. The HCT-ML framework extends conventional predictive evaluation through six dimensions that connect model performance to human-centric decision support.
Preprints 224417 g001
Figure 2. Illustrative application of HCT-ML to a municipal collision-severity prediction system. The workflow is conceptual and does not represent an empirical study.
Figure 2. Illustrative application of HCT-ML to a municipal collision-severity prediction system. The workflow is conceptual and does not represent an empirical study.
Preprints 224417 g002
Table 1. Operational interpretation of the HCT-ML dimensions.
Table 1. Operational interpretation of the HCT-ML dimensions.
Dimension Core Evaluation Question Risk if Absent Design Response
Relevance Does the output address a defined decision and population? Valid but unusable output Define user, decision, horizon, and error costs
Explainability Can the user understand the basis and limits? Blind reliance or rejection Provide task-specific explanations and user testing
Fairness Who experiences errors, benefits, and burdens? Geographic or demographic disadvantage Audit subgroup errors, coverage, and distributional effects
Uncertainty Does the system communicate limited confidence? Overconfident action Calibrate outputs and define fallback procedures
Human oversight Who can question, override, and remain accountable? Responsibility gaps Assign roles, logging, escalation, and override authority
Actionability Can the output support a feasible, timely response? Prediction without practical consequence Co-design around workflows and interventions
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings