Submitted:
18 August 2026
Posted:
20 August 2026
You are already at the latest version
Abstract
Institutional evaluation in higher education is increasingly shaped by multiple, coexisting systems, including research assessment, university rankings, accreditation and quality assurance, and responsible research assessment. Although these systems are widely examined individually, their dynamic interactions and cumulative institutional effects remain insufficiently understood. This study develops a conceptual System Dynamics framework to explain how multiple evaluation logics interact through reinforcing and balancing feedback processes. Drawing on institutional logics, organizational response, research evaluation, quality assurance, and System Dynamics perspectives, the framework integrates four evaluation subsystems: research prestige, competitive ranking, accreditation and quality assurance, and responsible assessment. Their interaction generates cross-system feedback mechanisms involving evaluation internalization, reputation and resource attraction, capability development, internationalization, strategic alignment, diversification, and organizational sustainability. The study further introduces Sustainable Institutional Capacity (SIC), defined as the level of evaluation-oriented activity that an institution can sustain without excessive organizational strain, quality deterioration, mission displacement, or capability erosion. By linking evaluation pressures to endogenous feedback and capacity constraints, the framework explains how initially beneficial evaluation-driven strategies may produce growth, adaptation, stabilization, or overextension as feedback dominance changes over time. The framework provides a theoretical basis for empirical validation and subsequent stock-and-flow simulation of institutional evaluation dynamics.
Keywords:
system dynamics
; institutional evaluation
; higher education
; institutional logics
; university rankings
; research assessment
; quality assurance
; sustainable institutional capacity
1. Introduction
Higher education institutions operate within an increasingly complex environment of evaluation, accountability, competition, and performance measurement. Universities are assessed through multiple mechanisms that differ in their objectives, criteria, and conceptions of institutional quality. Research performance is evaluated through publication productivity, citation impact, journal classifications, and related indicators; national and international rankings compare institutions using standardized performance measures; accreditation and quality-assurance systems establish expectations regarding educational processes, governance, learning outcomes, and continuous improvement; and responsible-assessment initiatives increasingly challenge narrow reliance on quantitative indicators. Institutional performance is therefore shaped not by a single evaluation regime but by multiple, simultaneously operating evaluation logics.
This plurality can be understood through the literature on institutional complexity and institutional logics. Organizations frequently encounter multiple institutional prescriptions that may coexist, reinforce one another, or generate competing demands [1,2,3]. Higher education institutions are particularly exposed to such complexity because their legitimacy depends on heterogeneous stakeholders and multiple dimensions of performance. External expectations, however, do not translate mechanically into organizational behavior. Institutions interpret, prioritize, negotiate, and sometimes resist external pressures, producing different strategic responses and institutional positions [4,5]. Evaluation systems therefore operate not simply as measurement devices but as institutional forces capable of reshaping priorities, resource allocation, organizational routines, and definitions of desirable performance.
These effects are particularly visible in research assessment and university rankings. Measures used to evaluate academic performance can become targets of organizational action, a phenomenon conceptualized by Espeland and Sauder [6] as “reactivity.” Rankings can consequently influence institutional strategies, reputation-building, resource allocation, recruitment, and competitive positioning [7,8,9]. Similar dynamics operate in research evaluation, where prestige-oriented publication systems and journal rankings can strengthen incentives for productivity and visibility while encouraging homogenization, strategic publication behavior, and intensified competition [10,11,12,13]. Such performance pressures may also increase job demands, workload, and risks of faculty burnout when organizational demands exceed available resources and capabilities [14,15].
Accreditation and quality assurance constitute another influential evaluation logic. Unlike competitive rankings, these systems generally emphasize standards, organizational processes, learning outcomes, quality culture, and continuous improvement. Repeated engagement with quality-assurance processes can support organizational learning and institutional capability development [16,17]. At the same time, formal quality systems can generate substantial documentation, monitoring, reporting, and administrative requirements and may be perceived as compliance-oriented when disconnected from substantive improvement [18]. Accreditation and quality assurance can therefore generate organizational learning while simultaneously increasing demands on institutional resources.
Conventional approaches to academic evaluation are also increasingly being reconsidered. The Leiden Manifesto, the Metric Tide, the Hong Kong Principles, the San Francisco Declaration on Research Assessment (DORA), and the Coalition for Advancing Research Assessment (CoARA) have contributed to a broader movement toward more contextual, transparent, multidimensional, and responsible approaches to research assessment [19,20,21,22,23]. These initiatives challenge the assumption that research quality and contribution can be adequately represented by a limited set of quantitative indicators. Their institutionalization, however, introduces another evaluation logic that must coexist with established ranking systems, journal hierarchies, accreditation requirements, and organizational performance expectations.
Despite extensive research on these mechanisms individually, their combined institutional dynamics remain insufficiently understood. Research assessment, rankings, accreditation, quality assurance, and responsible assessment are commonly examined as separate domains, although universities experience them simultaneously. This fragmentation limits understanding of how improvement in one evaluation domain may reinforce or constrain another, how multiple evaluation demands compete for shared resources and capabilities, and why initially beneficial evaluation-driven strategies may eventually produce diminishing returns, organizational strain, mission displacement, or capability erosion. More fundamentally, predominantly linear explanations provide limited insight into the endogenous feedback processes through which evaluation pressures reshape institutional behavior and the resulting organizational responses subsequently alter resources, capabilities, reputation, and future evaluation performance.
System Dynamics provides a theoretical and methodological perspective for addressing this problem because it conceptualizes organizational behavior as emerging from interacting feedback structures, accumulations, nonlinear relationships, and time delays rather than isolated cause–effect relationships [24,25,26,27]. Reinforcing feedback can generate cumulative growth and capability development, whereas balancing feedback can constrain expansion as resource limitations, workload, resistance, and other countervailing mechanisms become more influential. Their interaction can generate nonlinear trajectories such as S-shaped growth, overshoot, oscillation, stabilization, and capability erosion [28,29,30,31]. This perspective is particularly relevant to institutional evaluation because improvements in performance may generate reputation and resources that enable further investment, while the same processes can intensify organizational demands and eventually constrain continued expansion.
Against this background, this study asks: How do multiple evaluation logics interact dynamically to shape institutional adaptation, performance, and sustainability in higher education? To address this question, the study develops a conceptual System Dynamics framework integrating four major evaluation logics: research prestige, competitive ranking, accreditation and quality assurance, and responsible assessment. Rather than treating these mechanisms as independent external influences, the framework conceptualizes them as interacting subsystems characterized by reinforcing and balancing feedback processes.
The study makes four principal contributions. First, it integrates previously fragmented evaluation domains within a unified systems framework, enabling examination of their complementarities, tensions, and cross-system effects. Second, it reconceptualizes institutional responses to evaluation as endogenous feedback processes in which evaluation pressures influence organizational behavior while institutional responses recursively reshape resources, capabilities, reputation, and subsequent performance. Third, the study introduces Sustainable Institutional Capacity (SIC), defined as the level of evaluation-oriented activity that an institution can sustain without excessive organizational strain, quality deterioration, mission displacement, or capability erosion. SIC provides a system-level concept for explaining why evaluation-driven development may encounter endogenous limits. Fourth, the framework translates these theoretical relationships into explicit feedback structures and associated behavior-over-time reference modes, providing a foundation for subsequent empirical validation and formal stock-and-flow simulation. In doing so, the study shifts attention from whether individual evaluation systems are beneficial or detrimental toward understanding the dynamic conditions under which their interaction supports sustainable institutional development or generates organizational overextension.
2. Literature Review and Theoretical Foundations
2.1. Institutional Evaluation and Multiple Evaluation Logics
Higher education institutions are subject to multiple evaluation systems that embody different definitions of quality, performance, legitimacy, and institutional success. These include research-performance measures and journal classifications, national and international university rankings, accreditation and quality-assurance systems, and more recent initiatives advocating responsible approaches to research assessment. Although these mechanisms differ in purpose and design, they operate simultaneously and influence institutional priorities, resource allocation, organizational routines, and strategic positioning.
The institutional-logics perspective provides a theoretical foundation for understanding this plurality. Organizations frequently operate in environments characterized by multiple institutional prescriptions rather than a single coherent set of expectations [1,2,3]. Institutional complexity emerges when these logics differ in their objectives, practices, identities, or criteria of legitimacy. Such plurality does not necessarily imply conflict: evaluation logics may reinforce one another under some conditions while generating competing demands under others.
Universities are particularly exposed to this complexity because they pursue multidimensional missions while responding to heterogeneous stakeholders. Research excellence, educational quality, international visibility, societal contribution, financial sustainability, and regulatory compliance may all constitute legitimate objectives, while their associated evaluation mechanisms direct institutional attention toward different activities.
Organizations are not passive recipients of these pressures. Oliver [5] demonstrates that institutional demands can generate responses ranging from acquiescence and compromise to avoidance, defiance, and manipulation. In higher education, institutional positioning similarly reflects strategic agency within broader environmental and structural constraints [4]. Evaluation systems therefore become consequential not merely because they measure universities but because institutions interpret these measures, incorporate them into decision processes, and allocate resources in response.
From a systems perspective, these responses can progressively make evaluation endogenous to organizational behavior. Once external criteria influence institutional priorities, investments, and capabilities, subsequent changes in performance can affect reputation, resources, stakeholder expectations, and future evaluation pressures. Evaluation is therefore not simply an external input but can become part of a recursive institutional process.
2.2. Research Prestige and Performance Evaluation
Research evaluation constitutes a major mechanism through which universities and academics are differentiated. Publication output, citation impact, journal reputation, research funding, and related indicators influence perceptions of scholarly and institutional performance. These mechanisms are closely connected to cumulative advantage. Merton’s Matthew effect describes how recognition and resources can become disproportionately concentrated among already successful researchers and institutions [10]. At the institutional level, research success can enhance reputation, attract resources and talent, strengthen research capability, and thereby support further research performance.
Journal classifications and ranking lists reinforce this prestige-oriented logic by differentiating publication outlets and attaching symbolic value to selected journals. Such systems, however, do not merely describe research quality. Mingers and Willmott [11] demonstrate their performative effects on the organization and evaluation of business-school research, while Willmott [12] identifies consequences of excessive dependence on journal lists for scholarly practices. Teymourifar [13] similarly examines how the ABS journal ranking system influences scholarly legitimacy, academic behavior, and prestige-oriented publication strategies.
Research indicators can consequently generate both reinforcing and constraining effects. They may encourage productivity, visibility, specialization, and investment in research capability, while excessive reliance on quantitative indicators can produce strategic conformity, hypercompetition, and pressures on research integrity [32,33]. These pressures interact with finite human and institutional resources. The job demands–resources perspective suggests that sustained demands generate strain when they exceed available resources [14], while faculty-burnout research demonstrates consequences associated with persistent academic demands [15].
The research-prestige literature therefore reveals a central duality: mechanisms that reinforce research capability and institutional reputation may simultaneously increase workload, competition, homogenization, and resource strain. This duality provides the theoretical basis for representing research-prestige dynamics through interacting reinforcing and balancing feedback mechanisms.
2.3. Competitive Rankings and Institutional Positioning
University rankings constitute a related but analytically distinct evaluation logic. Whereas research evaluation often focuses on scholarly outputs or researchers, university rankings compare institutions through composite indicators and relative positions, transforming heterogeneous universities into visible and comparable competitors.
A central concept for understanding their organizational effects is reactivity. Espeland and Sauder [6] demonstrate that public measures can alter the behavior of the organizations being measured. Rankings are therefore not neutral representations of institutional performance but can reshape the environments they seek to describe. Subsequent work demonstrates how rankings may produce tighter coupling between external measures and internal organizational behavior [8].
Within higher education, global rankings influence strategic positioning, institutional priorities, and perceptions of international competitiveness [7,9]. Universities may invest in research visibility, internationalization, recruitment, data systems, and other ranking-relevant activities. Improved ranking performance can enhance reputation and resource attraction, which may subsequently support further investment in institutional capabilities.
This process suggests a reinforcing mechanism connecting ranking-oriented investment, measured performance, reputation, resource attraction, and subsequent investment. However, ranking strategies operate under resource constraints, and investments in selected indicators compete with alternative institutional priorities. Their marginal benefits may also decline as resource requirements increase, while strong ranking orientation can encourage convergence toward similar institutional models and reduce strategic diversity [34,35].
Existing ranking research therefore establishes competition, reactivity, and strategic adaptation, while a systems perspective additionally requires attention to how these mechanisms interact recursively with institutional resources and other evaluation regimes.
2.4. Accreditation and Quality Assurance
Accreditation and quality assurance represent a third evaluation logic. In contrast to the explicitly comparative orientation of rankings, quality-assurance systems generally emphasize standards, processes, learning outcomes, institutional effectiveness, and continuous improvement. International frameworks, such as the Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG) developed within the European quality-assurance framework involving the European Association for Quality Assurance in Higher Education (ENQA), together with discipline- and institution-specific accreditation systems, establish systematic processes for monitoring and enhancing educational quality [36,37,38].
Quality assurance can contribute to a quality culture in which evaluation becomes embedded in organizational practice rather than remaining an episodic external requirement [17]. Repeated participation in accreditation and assessment can also generate organizational learning, accumulated experience, routines, information systems, and specialized capabilities, thereby strengthening an institution’s capacity for subsequent evaluation and continuous improvement [16].
Nevertheless, quality assurance also imposes organizational demands. Documentation, evidence collection, monitoring, reporting, committee work, data management, and review processes consume faculty and administrative resources. Newton [18] highlights the tension between quality improvement and perceptions of quality monitoring as administrative burden. Extensive compliance requirements may therefore divert resources from substantive educational improvement toward demonstrating compliance.
Accreditation consequently contains another feedback duality: organizational learning can reinforce quality capability, while increasing monitoring and documentation requirements generate balancing pressures through administrative workload and resource consumption. Treating these effects as interacting processes is important for understanding how quality systems evolve over time.
2.5. Responsible Research Assessment and Value-Oriented Evaluation
A fourth evaluation logic has emerged from criticism of narrow, indicator-driven approaches to research assessment. Initiatives such as the Leiden Manifesto, the Metric Tide, the Hong Kong Principles, DORA, and CoARA advocate more responsible, contextual, transparent, and multidimensional approaches to evaluating research and researchers [19,20,21,22,23].
These initiatives challenge the assumption that research quality can be adequately represented through a limited set of quantitative indicators. They instead emphasize contextualized metrics, diverse research contributions, qualitative judgment, research integrity, and broader conceptions of scholarly value. More recent discussions of Open Science incentives similarly emphasize reward systems capable of recognizing a wider range of research practices and outputs [39].
Responsible assessment may consequently strengthen institutional legitimacy, trust, research integrity, and alignment with broader academic missions. Its institutionalization, however, occurs in environments where conventional indicators remain deeply embedded. Universities simultaneously face incentives associated with citations, journal rankings, global rankings, funding systems, and accreditation. Responsible assessment therefore does not simply replace an established evaluation logic but interacts with existing systems.
This coexistence can generate both reinforcement and resistance. Successful implementation can strengthen institutional experience, legitimacy, and further adoption, whereas implementation costs, organizational resistance, established routines, and continued dependence on conventional metrics may slow institutionalization. Responsible assessment should therefore be understood as an emerging evaluation logic embedded within a broader system of overlapping and potentially competing demands.
Table 1.
Major institutional evaluation logics and their theoretical foundations.
| Evaluation Logic | Primary Focus | Key Mechanisms | Potential Tensions | Representative Literature |
|---|---|---|---|---|
| Research prestige | Research output, journal status, citations, scholarly reputation | Cumulative advantage, prestige reinforcement, performance incentives, capability development | Hypercompetition, workload, homogenization, mission displacement | [10,11,12,13,33] |
| Competitive ranking | Relative institutional position and external visibility | Reactivity, reputation, competition, resource attraction, strategic positioning | Resource concentration, strategic convergence, diminishing returns | [6,7,8,9] |
| Accreditation and quality assurance | Standards, processes, learning outcomes, continuous improvement | Organizational learning, quality culture, monitoring, capability accumulation | Documentation burden, compliance workload, administrative demands | [16,17,18,36] |
| Responsible assessment | Contextual, multidimensional, and responsible evaluation | Responsible metrics, qualitative judgment, integrity, legitimacy | Implementation costs, organizational resistance, persistence of conventional metrics | [19,20,21,22,23] |
2.6. From Fragmented Evaluation Mechanisms to a Systems Perspective
The preceding literature demonstrates that each evaluation logic can generate both beneficial and constraining effects. Research-prestige systems may strengthen scholarly capability while increasing competitive and workload pressures; rankings can enhance reputation and resource attraction while encouraging resource concentration and strategic convergence; accreditation can promote organizational learning while increasing administrative demands; and responsible assessment can strengthen integrity and legitimacy while encountering implementation costs and resistance.
What remains less developed is an explanation of how these mechanisms interact when they operate simultaneously within the same institution. Existing research provides strong accounts of individual evaluation domains, but universities experience them through shared financial resources, managerial attention, faculty time, administrative capacity, information systems, and institutional capabilities. Actions directed toward one evaluation logic may therefore alter the institution’s capacity to respond to others.
System Dynamics provides a theoretical language for examining these interactions. Rather than treating institutional outcomes as independent linear relationships, it explains behavior through feedback, accumulation, nonlinear relationships, and delays [24,25,26,27]. Causal loop diagrams (CLDs) provide a qualitative representation of hypothesized causal structure and enable feedback relationships among variables to be identified [28]. Reinforcing feedback can amplify change, whereas balancing feedback counteracts it as resource limitations, workload, resistance, or other constraints intensify.
The interaction of these structures is important because beneficial interventions do not necessarily generate indefinite improvement. System archetypes demonstrate how endogenous constraints can transform initial growth into saturation, overshoot, oscillation, or decline [29]. Repenning and Sterman [30] demonstrate how organizational pressures can generate capability traps, while Oliva and Sterman [31] show how increasing demands and compensatory responses can contribute to quality erosion. Evaluation-driven institutional improvement should therefore be understood not only through the immediate effect of an intervention but through how its consequences evolve as feedback mechanisms and resource constraints change over time.
2.7. Literature Gap and Theoretical Positioning
Table 2 summarizes the relationship between the principal literature streams and the theoretical contribution developed in this study. The gap is not an absence of knowledge concerning rankings, research assessment, accreditation, or responsible assessment individually. Rather, it concerns their integration: existing streams provide limited explanation of the endogenous feedback processes generated when multiple evaluation logics simultaneously influence institutional resources, capabilities, reputation, workload, and strategic priorities.
Accordingly, this study extends rather than replaces the preceding theoretical perspectives. Institutional-logics research explains why universities encounter plural and potentially competing prescriptions; strategic-response theory explains how organizations interpret and respond to those pressures; research on rankings, accreditation, and responsible assessment identifies domain-specific mechanisms; and System Dynamics provides the conceptual apparatus for examining how these mechanisms interact recursively over time.
The first theoretical contribution is therefore integration: the study conceptualizes research prestige, competitive ranking, accreditation and quality assurance, and responsible assessment as components of a single institutional evaluation system rather than as independent evaluation domains.
Second, the framework moves from predominantly linear accounts of evaluation pressure and organizational response toward an endogenous feedback explanation in which evaluation performance influences reputation, resources, organizational capabilities, workload, and strategic behavior, which subsequently reshape future evaluation performance.
A further issue arises from the finite nature of institutional resources. Universities cannot indefinitely expand research initiatives, ranking strategies, accreditation activities, reporting requirements, internationalization efforts, and assessment reforms without drawing on managerial attention, faculty effort, administrative capacity, and financial resources. The job demands–resources perspective [14] and research on faculty burnout [15] indicate that persistent demands can generate strain when they exceed available resources. System Dynamics research on capability traps and quality erosion similarly demonstrates how short-term responses can progressively undermine underlying capabilities [30,31].
Building on these foundations, the third theoretical contribution is the introduction of SIC, defined in this study as the level of evaluation-oriented activity that an institution can sustain without excessive organizational strain, quality deterioration, mission displacement, or capability erosion. SIC is not treated as an established construct in the existing literature but as a system-level concept synthesizing organizational demands and resources, capability development and erosion, and endogenous limits to growth.
Finally, the study translates these theoretical relationships into an explicit System Dynamics architecture consisting of reinforcing and balancing feedback mechanisms and associated behavior-over-time reference modes. This architecture provides a basis for explaining why similar evaluation pressures may generate growth, adaptation, saturation, oscillation, overextension, or capability erosion under different institutional conditions and establishes a foundation for subsequent empirical validation and formal stock-and-flow modeling.
3. Conceptual and System Dynamics Methodology
3.1. Research Design
This study adopts a conceptual System Dynamics approach to examine how multiple institutional evaluation logics interact within higher education institutions. The purpose is not to estimate the effects of individual evaluation mechanisms or construct a quantitatively calibrated simulation model, but to identify and integrate theoretically supported causal mechanisms into a coherent feedback structure capable of explaining how institutional evaluation dynamics may evolve over time.
The study follows a theory-building rather than hypothesis-testing design. System Dynamics is used as a conceptual modeling methodology to make explicit the feedback structures through which institutional evaluation pressures, organizational responses, resources, capabilities, and performance outcomes may interact recursively over time. The resulting CLDs and behavior-over-time reference modes are therefore treated as theoretically grounded dynamic hypotheses rather than empirically estimated or predictive models.
3.2. Theoretical Synthesis and Framework Development
The framework was developed through a purposive and iterative theoretical synthesis of the literature streams reviewed in Section 2. Conceptual framework development is inherently iterative, involving repeated comparison, integration, synthesis, and refinement of theoretical concepts [40], while System Dynamics modeling similarly involves iteration among problem conceptualization, dynamic hypothesis formulation, model structure, and evaluation [24].
The theoretical synthesis covered institutional complexity and institutional logics [1,2,3], strategic organizational responses to institutional pressures [4,5], research evaluation and prestige [10,11,12,33], university rankings and organizational reactivity [6,7,8,9], quality assurance [16,17,18], and responsible research assessment [19,20,21,22,23]. These domain-specific relationships were subsequently interpreted through System Dynamics concepts of feedback, accumulation, delays, endogenous behavior, and limits to growth [24,25,26,27].
Four evaluation logics were retained as the principal subsystems of the framework: research–prestige, competitive–ranking, accreditation and quality assurance, and responsible assessment. These were selected because they represent distinct but simultaneously experienced forms of institutional evaluation that differ in their criteria, sources of legitimacy, incentives, and organizational responses while drawing on overlapping institutional resources and capabilities. Internationalization was treated as a cross-system mechanism rather than a separate evaluation logic because it operates across several evaluation domains and contributes to multiple forms of institutional performance and visibility.
Framework development then proceeded in three stages. First, the principal evaluation mechanisms identified in the literature were organized within the four subsystems. Second, theoretically supported causal relationships within each subsystem were translated into reinforcing and balancing feedback structures. Relationships associated with cumulative amplification, such as reputation, resource attraction, capability development, and organizational learning, informed reinforcing structures, whereas relationships involving resource constraints, workload, administrative burden, resistance, trade-offs, and capability limitations informed balancing structures. Third, the subsystems were integrated through cross-system relationships involving reputation, resources, institutional capabilities, internationalization, strategic alignment, and organizational sustainability. This integration provides the basis for introducing SIC as an overarching system-level construct.
3.3. Causal Loop Diagramming
CLDs represent the hypothesized causal structure of the institutional evaluation system. They provide a qualitative means of identifying feedback relationships and examining how circular causality can generate system behavior [28]. Each causal link is assigned a polarity. A positive (+) polarity indicates that, all else being equal, a change in the causal variable produces a change in the affected variable in the same direction relative to what it would otherwise have been; a negative (−) polarity indicates change in the opposite direction.
Feedback loops are classified as reinforcing (R) or balancing (B). Reinforcing loops amplify change and can contribute to cumulative advantage, capability development, reputation growth, or escalating institutional pressures. Balancing loops counteract change as constraints or corrective mechanisms become influential. Within the present framework, these arise principally through resource limitations, administrative burden, workload, organizational resistance, strategic adjustment, and capability constraints.
This distinction is important because institutional evaluation may simultaneously generate enabling and constraining effects. Improved performance may enhance reputation and attract resources that support further capability development, while sustaining that improvement may increase workload and resource consumption. Institutional trajectories therefore depend on the relative dominance of interacting feedback structures over time [29].
3.4. Feedback Structure and Theoretical Reference Modes
The framework distinguishes domain-specific from cross-system feedback mechanisms. The first level comprises reinforcing and balancing loops associated with each of the four evaluation logics, capturing how evaluation-oriented activities can become self-reinforcing while generating endogenous constraints. The second connects these domains through shared resources and capabilities and through broader processes of reputation formation, internationalization, strategic adaptation, and organizational sustainability.
Feedback dominance may change over time. Reinforcing processes may initially generate rapid improvement, while continued expansion progressively activates balancing mechanisms. Interactions of this kind can produce S-shaped growth, overshoot, oscillation, stabilization, capability traps, and quality erosion [29,30,31].
Accordingly, the behavior-over-time diagrams accompanying the causal structures are treated as theoretical reference modes. They illustrate hypothesized patterns consistent with the proposed feedback structures and are not empirical observations, statistically estimated trajectories, or outputs from a calibrated simulation model. Their purpose is to make the dynamic hypotheses implied by the CLDs explicit and to provide testable expectations for subsequent empirical and simulation-based research.
3.5. SIC as a System-Level Construct
Integrating the four evaluation subsystems raises a system-level issue: institutional responses to evaluation draw on overlapping and finite resources. Research investment, ranking initiatives, accreditation processes, quality monitoring, internationalization, and responsible-assessment reforms may simultaneously require financial resources, faculty effort, administrative capacity, managerial attention, and organizational capabilities. Their combined expansion can therefore generate pressures that remain less visible when each evaluation system is considered independently.
To capture this constraint, the framework introduces SIC, defined as the level of evaluation-oriented activity that an institution can sustain without excessive organizational strain, quality deterioration, mission displacement, or capability erosion. The concept draws on the job demands–resources perspective [14], evidence concerning academic workload and burnout [15], and System Dynamics research on capability traps, endogenous limits, and quality erosion [30,31].
SIC is not conceptualized as a fixed numerical threshold. Rather, it represents a dynamic institutional condition that changes with resources, capabilities, routines, technologies, governance arrangements, and strategic priorities. Institutions may expand sustainable capacity through capability development and organizational learning, whereas persistent overextension may erode it. SIC therefore captures not simply the availability of resources, but the institution’s system-level capacity to sustain the combined portfolio of evaluation demands without endogenous deterioration of organizational capability or mission performance.
3.6. Conceptual Model Validity and Boundary
Because the framework is conceptual rather than empirically calibrated, its validity concerns the theoretical plausibility, internal consistency, and transparency of the proposed causal structure rather than statistical goodness of fit. Causal relationships were included when they could be grounded in the theoretical and empirical literature synthesized above, while feedback structures were examined for consistency of causal polarity and correspondence with the mechanisms represented by the underlying literature. The behavior-over-time reference modes were subsequently used to articulate the dynamic consequences implied by these structures rather than as evidence that those trajectories have been empirically observed.
The model boundary is intentionally institutional. It focuses on evaluation pressures, organizational responses, shared resources and capabilities, reputation, workload, strategic adaptation, and sustainability within higher education institutions. Broader political, regulatory, economic, disciplinary, and national-system conditions are treated as contextual influences rather than modeled as separate endogenous subsystems. This boundary preserves analytical tractability while retaining the mechanisms required to examine interactions among the four evaluation logics.
The resulting framework should therefore be interpreted as a theoretically grounded dynamic hypothesis rather than a validated predictive model. Its purpose is to make causal assumptions explicit, identify feedback mechanisms that can generate different institutional trajectories, and establish a structured basis for subsequent empirical validation, parameterization, and formal stock-and-flow simulation.
4. System Dynamics Framework of Institutional Evaluation
The proposed framework conceptualizes institutional evaluation as a dynamic system in which multiple evaluation logics simultaneously influence university behavior. Rather than treating research evaluation, rankings, accreditation and quality assurance, and responsible assessment as independent influences, the framework emphasizes feedback processes through which responses to one evaluation logic alter the conditions governing responses to others.
The framework operates at two interconnected levels. The first comprises four domain-specific subsystems: Research–Prestige Logic, Competitive–Ranking Logic, Accreditation & Quality Assurance Logic, and Responsible–Assessment & Value Logic. Each contains reinforcing and balancing mechanisms through which evaluation-oriented activities can strengthen institutional performance and capabilities while generating endogenous constraints.
The second level integrates these domains through six cross-system feedback mechanisms: the R5 Evaluation Internalization Loop (Governmentality Mechanism), R6 Reputation–Resource–Capability Loop (Resource Dependence Mechanism), R7 Internationalization Loop (Globalization Mechanism), B5 Strategic Alignment Loop (Managing Multiple Demands), B6 Strategic Diversification Loop (Reducing Concentration Risk), and B7 Organizational Sustainability Loop (Maintaining Within Sustainable Capacity). Together, these mechanisms connect the four evaluation logics to broader institutional performance, legitimacy, impact, and sustainability.
This structure distinguishes external evaluation pressures from endogenous institutional dynamics. External drivers—including multilateral rankings, accreditation bodies, research-assessment systems, responsible-assessment principles and initiatives, and stakeholder expectations—provide inputs into the institutional evaluation system. Institutional responses subsequently alter monitoring, resource allocation, capabilities, reputation, internationalization, strategic alignment, workload, and organizational sustainability. These changes recursively influence subsequent responses and evaluation performance. Institutional trajectories therefore emerge from the interaction and changing dominance of reinforcing and balancing feedback mechanisms rather than from any single evaluation system.
Table 3 summarizes the principal feedback structures. The domain-specific loops R1–R4 and B1–B4 operate primarily within the four evaluation logics, whereas R5–R7 and B5–B7 represent cross-system mechanisms. SIC provides the system-level concept for interpreting the institution’s ability to sustain their aggregate demands.
4.1. Research–Prestige Logic
The Research–Prestige Logic captures the dynamic interaction among research performance, academic prestige, institutional resources, and research capability. Research evaluation creates incentives to improve publication output, citation performance, journal placement, funding, and other indicators of scholarly visibility. Merton’s Matthew effect provides the theoretical foundation for this process: prior recognition can increase access to resources and opportunities that facilitate subsequent achievement [10].
This mechanism forms the R1 Research–Prestige Reinforcement Loop. Improved research performance increases academic prestige and institutional visibility, strengthening the ability to attract funding, high-performing faculty, doctoral researchers, collaborations, and other resources. These resources enhance research capability and thereby support further performance improvement:
This cumulative process is also shaped by contemporary evaluation practices. Journal classifications, citation indicators, and related measures can influence how universities define, reward, and organize desirable research activities, becoming performative when they alter the behavior they evaluate [11,12,33]. R1 is counteracted by the B1 Research Sustainability Loop. Increasing research-performance emphasis may intensify publication expectations, competition, monitoring, and faculty workload. Sustained demands become problematic when they are not accompanied by sufficient resources [14], while persistent academic demands can adversely affect individual and organizational functioning through burnout [15]. Hypercompetitive environments may additionally encourage strategic behavior and place pressure on research integrity [32].
As these pressures increase, workload and organizational strain can constrain the effective research capacity available for continued improvement:
The interaction of R1 and B1 implies that research-performance growth need not continue indefinitely. When R1 dominates, performance gains reinforce prestige, resources, and capability; as demands intensify, B1 may become increasingly influential, shifting the trajectory toward slower improvement, stabilization, or, under persistent overextension, capability deterioration.
Figure 1 presents this causal structure, while Figure 2 illustrates the corresponding theoretical behavior-over-time reference mode.
Proposition 1. Improvements in research performance can generate a self-reinforcing process through prestige, resource attraction, and capability development; however, as performance demands intensify relative to available institutional capacity, workload and organizational strain increasingly activate balancing mechanisms that slow, stabilize, or potentially reverse continued performance growth.
4.2. Competitive–Ranking Logic
The Competitive–Ranking Logic captures institutional responses to comparative measures of performance, visibility, and reputation. Unlike research evaluation focused on particular scholarly outputs or researchers, university rankings position institutions relative to one another through standardized indicators and composite measures, introducing an explicitly comparative dimension into institutional evaluation [7,9].
A central mechanism is organizational reactivity. Espeland and Sauder [6] demonstrate that public measures can reshape the behavior of organizations being measured. Once rankings become salient, universities may alter priorities, resource allocation, reporting, recruitment, and other activities in response to ranking criteria. Sauder and Espeland [8] further show how rankings can tighten the coupling between external measures and internal organizational behavior.
These relationships form the R2 Ranking Reinforcement Loop. Improved ranking-relevant performance strengthens comparative position, visibility, and reputation, enhancing the institution’s attractiveness to students, faculty, collaborators, and other stakeholders and facilitating resource attraction. These resources support further investment in ranking-oriented capabilities:
This mechanism is consistent with evidence that rankings influence institutional strategy and competitive positioning [7,9,34]. Universities may invest in ranking-visible dimensions such as research, internationalization, recruitment, reputation, and data-management capabilities, allowing ranking competition to become self-reinforcing.
However, ranking-oriented improvement consumes finite and contested resources. Financial investment, managerial attention, faculty effort, data infrastructure, internationalization, and recruitment may compete with other institutional priorities, particularly when resources become concentrated on externally visible indicators [35]. Ranking systems may also encourage convergence around similar definitions of excellence despite differences in institutional mission and context [7,9].
These constraints form the B2 Ranking Sustainability Loop:
When R2 dominates, ranking improvement can reinforce reputation, resources, and further investment. As resource requirements and opportunity costs increase, B2 becomes more influential, producing diminishing gains and eventual stabilization.
Figure 3 presents the causal structure, while Figure 4 presents the corresponding theoretical behavior-over-time reference mode.
Proposition 2. Improvements in ranking-relevant performance can generate a self-reinforcing process through stronger comparative position, institutional visibility, reputation, resource attraction, and further ranking-oriented capability development; however, as ranking-oriented activities increasingly consume finite institutional resources and compete with alternative priorities, balancing pressures generate diminishing returns and constrain continued ranking-driven expansion.
4.3. Accreditation & Quality Assurance Logic
The Accreditation & Quality Assurance (QA) Logic focuses on institutional processes, standards, learning outcomes, governance, and continuous improvement rather than relative competitive position. International and regional frameworks emphasize systematic monitoring, evidence-based review, stakeholder involvement, and continuous improvement. The Standards and Guidelines for Quality Assurance in the European Higher Education Area and accreditation frameworks such as the Association to Advance Collegiate Schools of Business (AACSB) and the European Foundation for Management Development (EFMD) and its EFMD Quality Improvement System (EQUIS) formalize expectations concerning strategy, educational processes, faculty, impact, and institutional improvement [36,37,38]. Repeated engagement with these processes can generate organizational experience and contribute to quality culture [16,17].
These relationships form the R3 Quality Capability Reinforcement Loop. Participation in accreditation and quality-assurance activities generates experience in data collection, documentation, assessment, review, and improvement. Accumulated experience supports organizational learning and strengthens quality capability, improving subsequent quality processes and generating further learning:
R3 represents a learning mechanism rather than a purely compliance-based process. As institutional knowledge accumulates, quality-assurance activities may become embedded in routines, information systems, governance structures, and organizational practices, supporting continuous improvement [16,17].
These developmental effects are accompanied by organizational costs. Documentation, evidence collection, reporting, committee work, monitoring, data management, and external review consume institutional resources. Newton [18] identifies the tension between quality improvement and perceptions of monitoring as administrative burden. When formal requirements expand faster than institutional capacity, demonstrating quality may compete with activities through which educational quality is produced.
These relationships form the B3 Quality Assurance Burden Loop:
When R3 dominates, repeated quality-assurance experience strengthens organizational routines and quality capability. When documentation and compliance requirements grow faster than capability, B3 progressively constrains these gains.
Figure 5 presents the causal structure, while Figure 6 presents its theoretical behavior-over-time reference mode.
Proposition 3. Repeated engagement with accreditation and quality assurance can generate a self-reinforcing process of organizational learning, quality-capability development, and continuous improvement; however, increasing documentation, monitoring, and compliance requirements activate balancing pressures through administrative burden and resource consumption, thereby limiting the benefits of further quality-assurance intensification.
4.4. Responsible–Assessment & Value Logic
The Responsible–Assessment & Value Logic emerges partly in response to limitations associated with narrow, predominantly quantitative approaches to academic evaluation. DORA, the Leiden Manifesto, the Metric Tide, the Hong Kong Principles, and CoARA advocate more contextual, transparent, pluralistic, and responsible approaches to assessing research and researchers [19,20,21,22,23].
Responsible assessment broadens scholarly and institutional value beyond publication and citation indicators by emphasizing qualitative judgment, research integrity, disciplinary context, diverse outputs, Open Science practices, and wider academic and societal contributions. Developments in Open Science rewards and incentives further illustrate this movement toward recognizing a broader range of research practices [39].
These relationships form the R4 Responsible Assessment Reinforcement Loop. Adoption can improve alignment between evaluation practices and broader conceptions of academic value, strengthening perceived fairness, integrity, trust, and institutional legitimacy. Increased legitimacy and implementation experience can reinforce organizational commitment and further institutionalization:
Responsible assessment can therefore become progressively embedded as institutions develop experience, governance arrangements, and evaluation capabilities consistent with its principles.
However, established evaluation practices remain influential. Journal classifications, citation indicators, university rankings, funding criteria, and conventional performance measures may continue to shape incentives even when responsible-assessment principles are formally endorsed. Reform may therefore require changes in routines, information systems, evaluation criteria, governance, and stakeholder expectations.
These conditions form the B4 Responsible Assessment Implementation Loop. As reform expands, implementation requirements and transition costs increase, while established routines and conflicting external incentives may generate resistance:
The interaction of R4 and B4 implies that formal adoption does not necessarily produce immediate organizational transformation. Legitimacy, experience, and commitment can reinforce diffusion, while implementation costs, established routines, and competing evaluation incentives can slow or limit institutionalization.
Figure 7 presents the causal structure, while Figure 8 presents the corresponding theoretical behavior-over-time reference mode.
Proposition 4. Adoption of responsible-assessment practices can generate a self-reinforcing process through greater integrity, perceived fairness, legitimacy, institutional commitment, and accumulated implementation experience; however, transition costs, organizational resistance, established routines, and continuing dependence on conventional evaluation metrics create balancing pressures that constrain the pace and extent of institutionalization.
5. Discussion
The proposed framework provides a dynamic interpretation of institutional evaluation in higher education. Its central argument is that evaluation cannot be adequately understood by examining individual systems in isolation or by assuming linear relationships between external pressures and organizational responses. Universities operate under multiple evaluation logics whose effects unfold through interconnected reinforcing and balancing feedback mechanisms. Research prestige, competitive ranking, accreditation and quality assurance, and responsible assessment may each stimulate institutional improvement while competing for shared resources, managerial attention, faculty effort, and organizational capacity.
The framework therefore shifts attention from the effects of individual evaluation mechanisms to the dynamic configuration produced by their interaction. This perspective helps explain why evaluation systems that appear beneficial independently may generate unintended consequences when pursued simultaneously, and why institutions exposed to similar evaluation environments may follow different developmental trajectories.
5.1. From Competing Evaluation Logics to Dynamic Interaction
Institutional-logics research establishes that organizations may face multiple prescriptions and criteria of legitimacy [1,2,3]. Strategic-response research further demonstrates that organizations may accommodate, negotiate, prioritize, resist, or reshape these pressures [4,5]. The present framework extends these perspectives by emphasizing that organizational responses subsequently alter the conditions under which future responses occur.
For example, investment in research performance may strengthen prestige and resource attraction through R1 and, through the cross-system R6 Reputation–Resource–Capability Loop, expand capabilities that also support ranking performance, internationalization, and quality assurance. Conversely, simultaneous expansion across evaluation domains can increase workload, administrative complexity, and competition for resources, activating B5 Strategic Alignment, B6 Strategic Diversification, and B7 Organizational Sustainability. The effect of any evaluation logic therefore depends partly on the wider evaluation system in which it is embedded.
Institutional complexity is thus dynamic rather than merely simultaneous. Evaluation logics interact through shared resources and capabilities and common outcomes such as reputation, legitimacy, internationalization, and impact. Their relationships may also change over time as institutional conditions evolve.
The R5 Evaluation Internalization Loop adds a further dimension. Research on rankings demonstrates that organizations become reactive to public measures [6,8]. The framework extends this argument by proposing that repeated exposure to evaluation can progressively embed external criteria within internal monitoring and decision making. Evaluation can therefore become partially endogenous to organizational governance, influencing behavior even without immediate external intervention.
5.2. Feedback Dominance and Institutional Trajectories
A central implication of System Dynamics is that the existence of a feedback loop does not by itself determine system behavior. Institutional trajectories depend on the relative strength and changing dominance of interacting feedback mechanisms over time [24,25,26,27]. The framework therefore does not imply that evaluation necessarily improves or damages institutional performance; both outcomes may emerge from the same system under different configurations of feedback dominance.
During early evaluation-oriented development, reinforcing mechanisms may dominate. Research performance can strengthen prestige and resources, ranking performance can increase visibility, quality-assurance experience can generate organizational learning, and responsible-assessment implementation can enhance legitimacy and commitment. At the integrated level, R6 and R7 can further amplify these gains through resource attraction, capability development, and internationalization.
As evaluation-oriented activities expand, however, workload, documentation, resource consumption, implementation costs, and coordination complexity can increase. Domain-specific balancing loops B1–B4 become more influential, while B5–B7 respond to system-wide misalignment, concentration risk, and organizational strain. Changing feedback dominance can therefore produce nonlinear trajectories: rapid initial improvement may give way to slower growth or a sustainable plateau, while persistent expansion beyond institutional capacity may produce overshoot and subsequent capability erosion.
These patterns are consistent with System Dynamics research on limits to growth, capability traps, and quality erosion [29,30,31]. Their application to institutional evaluation suggests that visible performance improvement at one point in time does not necessarily indicate long-term sustainability. Institutions may improve measured outcomes while simultaneously accumulating workload, administrative complexity, or capability deficits whose effects emerge only after a delay.
The behavior-over-time reference modes should therefore be interpreted as dynamic hypotheses rather than predetermined trajectories. They represent patterns that may emerge under different configurations and changing dominance of reinforcing and balancing feedback.
5.3. SIC and the Limits of Evaluation-Driven Growth
SIC captures the system-level implications of these interacting pressures. The job demands–resources perspective explains how sustained demands can generate strain when adequate resources are unavailable [14]; faculty-burnout research documents consequences associated with persistent academic demands [15]; and System Dynamics studies demonstrate how pressure and short-term responses can progressively undermine organizational capabilities [30,31]. SIC integrates these insights at the institutional evaluation-system level.
The key distinction is between domain-specific capability and system-level sustainable capacity. A university may possess strong research infrastructure, effective accreditation processes, sophisticated ranking analytics, and responsible-assessment capabilities without being able to expand all of these activities simultaneously. They draw directly or indirectly on overlapping human, financial, administrative, and managerial resources.
SIC therefore concerns the relationship between aggregate evaluation demands and the institution’s ability to absorb, coordinate, and sustain them. When demands remain within sustainable capacity, evaluation pressures may stimulate learning, investment, capability development, and performance improvement. As demands approach capacity, trade-offs become increasingly consequential; when they persistently exceed it, workload, complexity, burnout, mission displacement, and capability erosion become more likely.
Importantly, SIC is dynamic rather than fixed. Organizational learning, digital infrastructure, improved governance, additional resources, effective coordination, and strategic prioritization may expand sustainable capacity, whereas persistent overextension may reduce it. Institutional responses therefore affect not only current performance but also the capacity available for future responses.
This interpretation distinguishes SIC from a conventional static resource constraint. SIC represents the institution-specific capacity to sustain the combined portfolio of evaluation demands without endogenous deterioration of organizational capability or mission performance. Its level may vary with institutional mission, resource base, governance arrangements, capabilities, and the particular configuration of evaluation demands.
5.4. Theoretical Contributions
The framework makes several contributions to the literature on institutional evaluation and higher education.
First, it integrates evaluation systems that have largely developed as separate research streams. Research evaluation, university rankings, accreditation and quality assurance, and responsible assessment are represented as interacting evaluation logics, shifting attention from isolated effects toward their complementarities, conflicts, and cross-domain consequences.
Second, the framework extends institutional complexity from a predominantly structural problem of multiple demands toward a dynamic problem of feedback. Institutional-logics research explains why organizations encounter plural and potentially competing prescriptions [1,2,3], while strategic- response theory explains organizational responses to institutional pressures [5]. The present framework adds recursive causality: organizational responses alter resources, capabilities, reputation, monitoring practices, and institutional conditions, which subsequently shape future responses.
Third, the framework extends reactivity by embedding it within a broader feedback architecture. Rankings and other public measures can alter organizational behavior [6,8], while the R5 Evaluation Internalization Loop proposes that repeated evaluation can become embedded in internal monitoring and governance. Evaluation thus becomes not only something that happens to an institution but also something increasingly reproduced within it.
Fourth, the framework explains how apparently contradictory consequences of evaluation can coexist. Quality assurance may generate both organizational learning and administrative burden [16,17,18]. Research evaluation may similarly generate capability development and excessive performance pressure [11,12,32]. Rather than treating these as competing explanations, the framework represents them as reinforcing and balancing mechanisms operating within the same system, whose relative influence changes with institutional conditions and time.
Fifth, the study introduces SIC as a system-level concept linking the benefits of evaluation-driven development to endogenous organizational limits. SIC emphasizes that sustainability cannot be assessed within any evaluation domain independently because the relevant constraint arises from aggregate demands on shared institutional resources and capabilities.
Finally, the framework translates these arguments into an explicit causal architecture comprising four domain-specific evaluation logics and the cross-system R5–R7 and B5–B7 feedback mechanisms. This explicit structure increases theoretical transparency and testability and provides a basis for subsequent empirical investigation and formal stock-and-flow simulation.
Taken together, these contributions shift the central question from whether institutional evaluation is beneficial or detrimental toward a more dynamic question: under what configurations of evaluation logics, institutional resources, capabilities, and feedback mechanisms does evaluation contribute to sustainable institutional development, and under what configurations does it generate overextension and capability erosion?
6. Implications for Institutional Governance and Management
The proposed framework has several implications for the governance and management of higher education institutions. Its central implication is that evaluation systems should not be managed as independent institutional requirements. Research assessment, university rankings, accreditation and quality assurance, responsible assessment, and internationalization draw on overlapping financial, human, administrative, and managerial resources. Decisions intended to improve performance in one domain may therefore generate consequences elsewhere in the institutional system.
First, institutional leaders should manage the portfolio of evaluation demands rather than optimize individual indicators independently. The reinforcing loops identified in the framework explain why targeted initiatives can initially generate substantial benefits through improved research performance, visibility, quality capability, or legitimacy. However, simultaneous expansion across multiple domains can activate the balancing mechanisms represented by B5–B7. Evaluation strategy should therefore consider both the expected benefits of individual initiatives and their cumulative demands on institutional capacity.
Second, the B5 Strategic Alignment Loop (Managing Multiple Demands) highlights the need to identify and manage tensions among evaluation logics. Improving performance according to one system may not support, and may sometimes conflict with, the expectations of another. Strong incentives to concentrate research effort on selected publication indicators, for example, may conflict with responsible-assessment principles emphasizing broader academic contributions, while investment in externally visible ranking indicators may compete with locally important educational or societal missions. Institutional governance should therefore provide mechanisms for identifying, prioritizing, and negotiating such tensions rather than assuming that all evaluation objectives can be maximized simultaneously.
Third, the B6 Strategic Diversification Loop (Reducing Concentration Risk) suggests that excessive dependence on a narrow set of indicators can create institutional vulnerability. Changes in ranking methodologies, accreditation standards, research-assessment policies, or stakeholder expectations may disproportionately affect institutions whose strategies are tightly coupled to particular measures. Diversification across research quality, education, societal impact, international engagement, organizational learning, and other mission-relevant dimensions can reduce this dependence. This implication is consistent with responsible-assessment initiatives that question excessive reliance on narrow quantitative measures [19,20,22,23].
Fourth, the B7 Organizational Sustainability Loop (Maintaining Within Sustainable Capacity) indicates that evaluation workload should itself be treated as a governance variable. Documentation, reporting, data collection, performance monitoring, accreditation preparation, ranking submissions, and research-assessment processes consume organizational resources. Their accumulation may generate substantial workload and coordination complexity even when individual processes appear manageable. Monitoring evaluation burden alongside evaluation performance can therefore provide an early indication of institutional overextension.
The concept of SIC further implies that capacity should be actively developed rather than treated as a fixed constraint. Investments in information systems, administrative capabilities, organizational learning, coordination, governance processes, and appropriate resource allocation may expand an institution’s ability to manage multiple evaluation demands. Conversely, persistent reliance on additional faculty effort, temporary workarounds, or escalating administrative requirements may improve visible performance in the short term while eroding sustainable capacity over time. This possibility is consistent with the broader System Dynamics literature on capability traps and quality erosion [30,31].
Finally, the framework suggests that institutional performance dashboards should distinguish between outcome indicators and capacity indicators. Ranking position, publication performance, accreditation outcomes, and international visibility indicate institutional outcomes but reveal less about the organizational effort required to sustain them. Complementary indicators of workload, administrative burden, resource availability, staff capacity, organizational learning, and mission alignment could help determine whether observed performance improvements are sustainable.
These implications do not suggest reducing institutional engagement with external evaluation. Rather, evaluation should be governed as an interconnected institutional system. Effective governance therefore requires not only performance improvement but also strategic alignment, diversification, capability development, monitoring of evaluation burden, and preservation of SIC.
7. Conclusions, Limitations, and Future Research
This study developed a conceptual System Dynamics framework to explain how multiple institutional evaluation logics interact dynamically within higher education institutions. Rather than treating research evaluation, university rankings, accreditation and quality assurance, and responsible assessment as independent external requirements, the framework represents them as interdependent subsystems connected through shared resources, organizational capabilities, reputation, internationalization, strategic responses, and institutional capacity. In response to the research question, the framework suggests that the effects of institutional evaluation cannot be understood solely from the characteristics of individual evaluation systems. Instead, they emerge from the interaction and changing dominance of reinforcing and balancing feedback mechanisms over time.
A central implication of this perspective is that evaluation-driven development is neither inherently beneficial nor inherently detrimental. Reinforcing mechanisms can support cumulative improvements in research performance, reputation, resource attraction, quality capability, internationalization, and the institutionalization of responsible assessment. At the same time, continued expansion of evaluation-oriented activities can increase workload, administrative complexity, resource competition, implementation costs, and coordination demands. As these pressures accumulate, balancing mechanisms may become increasingly influential, producing slower growth, stabilization, trade-offs, or, under conditions of persistent overextension, capability erosion. The framework therefore shifts attention from whether evaluation improves institutional performance to the conditions under which evaluation-driven development remains organizationally sustainable.
To capture these conditions, the study introduced the concept of Sustainable Institutional Capacity (SIC). SIC represents an institution’s capacity to sustain the combined portfolio of evaluation demands without generating excessive organizational strain, mission displacement, quality deterioration, or endogenous capability erosion. Importantly, SIC is conceived as dynamic rather than fixed. Organizational learning, digital infrastructure, governance, coordination, and resource investment may expand sustainable capacity, whereas persistent overextension may progressively reduce it. This perspective highlights the importance of considering not only visible evaluation outcomes but also the organizational capacity required to sustain them.
Several limitations should be acknowledged. First, the proposed framework is conceptual and theory-driven rather than empirically estimated or quantitatively calibrated. The causal relationships and feedback mechanisms are derived from the literature synthesized in this study and should therefore be interpreted as theoretically grounded dynamic hypotheses rather than empirically established causal effects. Second, the behavior-over-time reference modes illustrate plausible trajectories associated with changing feedback dominance; they are not predictions of the behavior of a particular institution. Third, the framework intentionally adopts an institution-level system boundary. Important sources of heterogeneity, including national higher education systems, institutional missions, disciplinary structures, funding regimes, and regulatory environments, are not modeled endogenously. These contextual conditions may substantially affect the strength, timing, and interaction of the proposed feedback mechanisms.
A further limitation concerns the cross-system feedback mechanisms represented by R5–R7 and B5–B7. These mechanisms extend beyond the four domain-specific evaluation logics by integrating processes associated with evaluation internalization, resource dependence, internationalization, strategic alignment, diversification, and organizational sustainability. Although these mechanisms are theoretically informed, the present framework synthesizes them at a relatively high level of abstraction rather than developing and empirically testing each as a separate theoretical mechanism. Their relative strength, salience, and interaction may therefore vary across institutional, disciplinary, and national contexts. In particular, the theoretical boundary conditions under which individual cross-system mechanisms become dominant remain to be established empirically. Finally, SIC is introduced as a theoretical system-level concept and has not yet been operationalized or empirically measured. Its dimensions, indicators, and institutional thresholds therefore require further investigation.
These limitations provide a structured agenda for future research. An immediate priority is empirical examination of the proposed causal relationships through comparative case studies, longitudinal institutional data, surveys, interviews, or mixed-method designs. Such research could assess whether the proposed feedback mechanisms are observable across different types of higher education institutions and identify contextual conditions that strengthen or weaken them. Particular attention should be given to the cross-system mechanisms represented by R5–R7 and B5–B7, including whether their proposed relationships operate consistently across contexts, how their relative influence changes over time, and under what conditions particular reinforcing or balancing mechanisms become dominant. A second priority is the operationalization of SIC. Potential indicators could incorporate evaluation workload, administrative burden, resource availability, staff capacity, organizational learning, coordination demands, and mission alignment alongside conventional performance measures.
Subsequent research could translate the causal architecture developed here into a formal stock-and-flow System Dynamics model. Empirically informed parameters and functional relationships would permit simulation of alternative institutional trajectories, sensitivity analysis, and examination of policy interventions affecting evaluation intensity, resource allocation, capability development, and organizational workload. Such a model could investigate the conditions under which reinforcing processes generate sustainable capability development and those under which delayed balancing mechanisms produce overshoot or capability erosion. Finally, comparative applications across countries, institutional types, disciplines, and evaluation regimes could test the generalizability and boundary conditions of the framework. Together, these developments would provide a pathway from the conceptual dynamic hypotheses advanced in this study toward an empirically validated and simulation-supported theory of sustainable institutional evaluation.
Author Contributions
Conceptualization, A.T.; methodology, A.T.; investigation, A.T.; writing—original draft preparation, A.T.; writing—review and editing, A.T.; visualization, A.T. The author has read and agreed to the published version of the manuscript.
Funding
This research was supported by Istanbul Sabahattin Zaim University under Project No. 2026-BAP-400-010.
Data Availability Statement
No new data were created or analyzed in this study. Data sharing is not applicable to this article.
Acknowledgments
ChatGPT (GPT-5.6 Sol, OpenAI) was employed to enhance the English writing of this paper, focusing on improvements in grammar, style, and overall clarity.
Conflicts of Interest
The author declares no conflict of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| Abbreviation | Definition |
| AACSB | Association to Advance Collegiate Schools of Business |
| CLD | Causal Loop Diagram |
| CoARA | Coalition for Advancing Research Assessment |
| DORA | San Francisco Declaration on Research Assessment |
| EFMD | European Foundation for Management Development |
| ENQA | European Association for Quality Assurance in Higher Education |
| EQUIS | EFMD Quality Improvement System |
| ESG | Standards and Guidelines for Quality Assurance in the European Higher Education Area |
| QA | Quality Assurance |
| SIC | Sustainable Institutional Capacity |
References
- Greenwood, R.; Raynard, M.; Kodeih, F.; Micelotta, E. R.; Lounsbury, M. Institutional complexity and organizational responses. Acad. Manag. Ann. 2011, 5(1), 317–371. [Google Scholar] [CrossRef]
- Besharov, M. L.; Smith, W. K. Multiple institutional logics in organizations: Explaining their varied nature and implications. Acad. Manag. Rev. 2014, 39(3), 364–381. [Google Scholar] [CrossRef]
- Thornton, P. H.; Ocasio, W.; Lounsbury, M. The Institutional Logics Perspective; Oxford University Press: Oxford, UK, 2012; Vol. 10, ISBN 978-1-118-90077-2. [Google Scholar]
- Fumasoli, T.; Huisman, J. Strategic agency and system diversity: Conceptualizing institutional positioning in higher education. Minerva 2013, 51(2), 155–169. [Google Scholar] [CrossRef]
- Oliver, C. Strategic responses to institutional processes. Acad. Manag. Rev. 1991, 16(1), 145–179. [Google Scholar] [CrossRef]
- Espeland, W. N.; Sauder, M. Rankings and reactivity: How public measures recreate social worlds. Am. J. Sociol. 2007, 113(1), 1–40. [Google Scholar] [CrossRef] [PubMed]
- Hazelkorn, E. Rankings and policy choices. In Rankings and the Reshaping of Higher Education: The Battle for World-Class Excellence; Palgrave Macmillan: London, UK, 2015; pp. 153–186. [Google Scholar] [CrossRef]
- Sauder, M.; Espeland, W. N. The discipline of rankings: Tight coupling and organizational change. Am. Sociol. Rev. 2009, 74(1), 63–82. [Google Scholar] [CrossRef]
- Marginson, S.; Van der Wende, M. To rank or to be ranked: The impact of global rankings in higher education. J. Stud. Int. Educ. 2007, 11(3–4), 306–329. [Google Scholar] [CrossRef]
- Merton, R. K. The Matthew effect in science: The reward and communication systems of science are considered. Science 1968, 159(3810), 56–63. [Google Scholar] [CrossRef]
- Mingers, J.; Willmott, H. Taylorizing business school research: On the `one best way’ performative effects of journal ranking lists. Hum. Relat. 2013, 66(8), 1051–1073. [Google Scholar] [CrossRef]
- Willmott, H. Journal list fetishism and the perversion of scholarship: Reactivity and the ABS list. Organization 2011, 18(4), 429–442. [Google Scholar] [CrossRef]
- Teymourifar, A. Understanding the ABS journal ranking system: A critical review. Front. Educ. 2026, 11, 1773655. [Google Scholar] [CrossRef]
- Bakker, A. B.; Demerouti, E. The job demands–resources model: State of the art. J. Manag. Psychol. 2007, 22(3), 309–328. [Google Scholar] [CrossRef]
- Sabagh, Z.; Hall, N. C.; Saroyan, A. Antecedents, correlates and consequences of faculty burnout. Educ. Res. 2018, 60(2), 131–156. [Google Scholar] [CrossRef]
- Stensaker, B. R. Outcomes of quality assurance: A discussion of knowledge, methodology and validity. Qual. High. Educ. 2008, 14(1), 3–13. [Google Scholar] [CrossRef]
- Harvey, L.; Stensaker, B. Quality culture: Understandings, boundaries and linkages. Eur. J. Educ. 2008, 43(4), 427–442. [Google Scholar] [CrossRef]
- Newton, J. Feeding the beast or improving quality? Academics’ perceptions of quality assurance and quality monitoring. Qual. High. Educ. 2000, 6(2), 153–163. [Google Scholar] [CrossRef]
- Hicks, D.; Wouters, P.; Waltman, L.; de Rijcke, S.; Rafols, I. Bibliometrics: The Leiden Manifesto for research metrics. Nature 2015, 520, 429–431. [Google Scholar] [CrossRef] [PubMed]
- Wilsdon, J.; Allen, L.; Belfiore, E.; Campbell, P.; Curry, S.; Hill, S.; Johnson, B.; et al. The Metric Tide: Report of the Independent Review of the Role of Metrics in Research Assessment and Management; Higher Education Funding Council for England: Bristol, UK, 2015. [Google Scholar]
- Moher, D.; Bouter, L.; Kleinert, S.; Glasziou, P.; Sham, M. H.; Barbour, V.; et al. The Hong Kong Principles for assessing researchers: Fostering research integrity. PLoS Biol. 2020, 18(7), e3000737. [Google Scholar] [CrossRef] [PubMed]
- Coalition for Advancing Research Assessment (CoARA); Coalition for Advancing Research Assessment. 2023.
- American Society for Cell Biology (ASCB). San Francisco Declaration on Research Assessment (DORA). 2012. Available online: https://sfdora.org/read/.
- Sterman, J. System Dynamics: Systems Thinking and Modeling for a Complex World. 2002. Available online: http://hdl.handle.net/1721.1/102741.
- Meadows, D. H. Thinking in Systems; Chelsea Green Publishing, 2008; ISBN 9781603580557. [Google Scholar]
- Richardson, G. P. Feedback Thought in Social Science and Systems Theory; University of Pennsylvania Press, 1991; ISBN 0812213327. [Google Scholar]
- Forrester, J. W. Industrial Dynamics; Pegasus Communications: Waltham, MA, USA, 1961; ISBN 978-1883823368. [Google Scholar]
- Schaffernicht, M. Causality and diagrams for system dynamics. 2007. [Google Scholar] [CrossRef]
- Wolstenholme, E. F. Towards the definition and use of a core set of archetypal structures in system dynamics. Syst. Dyn. Rev. 2003, 19(1), 7–26. [Google Scholar] [CrossRef]
- Repenning, N. P.; Sterman, J. D. Capability traps and self-confirming attribution errors in the dynamics of process improvement. Adm. Sci. Q. 2002, 47(2), 265–295. [Google Scholar] [CrossRef]
- Oliva, R.; Sterman, J. D. Cutting corners and working overtime: Quality erosion in the service industry. Manag. Sci. 2001, 47(7), 894–914. [Google Scholar] [CrossRef]
- Edwards, M. A.; Roy, S. Academic research in the 21st century: Maintaining scientific integrity in a climate of perverse incentives and hypercompetition. Environ. Eng. Sci. 2017, 34(1), 51–61. [Google Scholar] [CrossRef] [PubMed]
- de Rijcke, S.; Wouters, P. F.; Rushforth, A. D.; Franssen, T. P.; Hammarfelt, B. Evaluation practices and effects of indicator use—a literature review. Res. Eval. 2016, 25(2), 161–169. [Google Scholar] [CrossRef]
- Hazelkorn, E.; Loukkola, T.; Zhang, T. Rankings in institutional strategies and processes: Impact or illusion. 2014. Available online: https://arrow.tudublin.ie/cserrep/54/.
- Goldman, C. A.; Goldman, C.; Gates, S. M.; Brewer, A.; Brewer, D. J. Pursuit of Prestige: Strategy and Competition in US Higher Education; Transaction Publishers, 2004; ISBN 978-0765808295. [Google Scholar]
- European Association for Quality Assurance in Higher Education (ENQA); et al. Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG). 2015. [Google Scholar]
- Association to Advance Collegiate Schools of Business (AACSB). 2020 Guiding Principles and Standards for Business Accreditation; AACSB: Tampa, FL, USA, 2020; Available online: https://www.aacsb.edu/-/media/documents/accreditation/2020-aacsb-business-accreditation-standards-july-2021.pdf (accessed on 11 August 2026).
- European Foundation for Management Development (EFMD). EQUIS Standards & Criteria. 2022. Available online: https://www.efmdglobal.org/wp-content/uploads/EQUIS_Standards_and_Criteria.pdf (accessed on 11 August 2026).
- Mabile, L.; Shmagun, H.; Erdmann, C.; Cambon-Thomsen, A.; Thomsen, M.; Grattarola, F. Recommendations on Open Science rewards and incentives: Guidance for multiple stakeholders in research. Data Sci. J. 2025, 24, 15. [Google Scholar] [CrossRef]
- Jabareen, Y. Building a Conceptual Framework: Philosophy, Definitions, and Procedure. Int. J. Qual. Methods 2009, 8, 49–62. [Google Scholar] [CrossRef]
Figure 1.
Research–Prestige Logic: Reinforcing and balancing feedback structure.

Figure 2.
Research–Prestige Logic: Hypothesized behavior-over-time reference mode.

Figure 3.
Competitive–Ranking Logic: Reinforcing and balancing feedback structure.

Figure 4.
Competitive–Ranking Logic: Hypothesized behavior-over-time reference mode.

Figure 5.
Accreditation & Quality Assurance Logic: Reinforcing and balancing feedback structure.

Figure 6.
Accreditation & Quality Assurance Logic: Hypothesized behavior-over-time reference mode.

Figure 7.
Responsible–Assessment & Value Logic: Reinforcing and balancing feedback structure.

Figure 8.
Responsible–Assessment & Value Logic: Hypothesized behavior-over-time reference mode.

Table 2.
Positioning of the proposed framework relative to existing literature.
| Literature Stream | What existing Literature Explains | Remaining Systems-Level Limitation | Contribution of this Study |
|---|---|---|---|
| Institutional logics | Plural and potentially competing institutional prescriptions | Limited representation of endogenous temporal feedback among evaluation responses | Conceptualizes competing evaluation logics as dynamically interconnected subsystems |
| Strategic institutional response | Accommodation, compromise, resistance, and strategic positioning | Organizational responses are not generally represented through explicit recursive feedback | Connects institutional responses to subsequent changes in resources, capabilities, reputation, and evaluation pressures |
| Rankings and reactivity | Behavioral responses to measurement, comparison, and ranking | Predominantly examines effects within the ranking domain | Connects ranking reactivity with research prestige, quality assurance, responsible assessment, and resource allocation |
| Research-performance evaluation | Prestige hierarchies, incentives, indicator effects, and scholarly behavior | Limited integration with other institutional evaluation regimes | Embeds research-prestige dynamics within institution-wide reinforcing and balancing feedback structures |
| Quality assurance | Quality culture, organizational learning, monitoring, and administrative burden | Developmental and burdensome effects are commonly treated as separate outcomes | Represents capability accumulation and administrative burden as interacting feedback mechanisms |
| Responsible assessment | Principles for contextual, pluralistic, and responsible evaluation | Limited explanation of institutional transition under competing established metrics | Models responsible assessment as an emerging logic interacting with established evaluation systems |
| System Dynamics | Feedback, delays, nonlinear behavior, limits to growth, and capability erosion | Limited application to interacting institutional evaluation regimes in higher education | Develops an integrated feedback architecture linking multiple evaluation logics |
| Integrated perspective | — | Lack of a system-level construct representing the sustainable limit of simultaneous evaluation-oriented activity | Introduces SIC as a system-level construct |
Table 3.
Feedback structures in the proposed institutional evaluation framework.
| Loop | Evaluation logic / level | Loop name | Principal dynamic function |
|---|---|---|---|
| R1 | Research–Prestige | Research–Prestige Reinforcement Loop | Reinforces research performance through prestige, resources, and capability development |
| B1 | Research–Prestige | Research Sustainability Loop | Constrains continued research expansion through workload, strain, and capability limitations |
| R2 | Competitive–Ranking | Ranking Reinforcement Loop | Reinforces ranking performance through reputation, resource attraction, and ranking-oriented investment |
| B2 | Competitive–Ranking | Ranking Sustainability Loop | Limits ranking-oriented expansion as resource requirements and competing demands increase |
| R3 | Accreditation & Quality Assurance | Quality Capability Reinforcement Loop | Builds quality capability through experience, organizational learning, and continuous improvement |
| B3 | Accreditation & Quality Assurance | Quality Assurance Burden Loop | Constrains further expansion through documentation, monitoring, and administrative burden |
| R4 | Responsible–Assessment & Value | Responsible Assessment Reinforcement Loop | Strengthens responsible-assessment practices through legitimacy, experience, and institutionalization |
| B4 | Responsible–Assessment & Value | Responsible Assessment Implementation Loop | Limits the pace of institutionalization through resistance, implementation costs, and established evaluation practices |
| R5 | Cross-system | Evaluation Internalization Loop (Governmentality Mechanism) | Transforms external evaluation visibility and salience into internal monitoring, self-surveillance, and renewed performance orientation |
| R6 | Cross-system | Reputation–Resource–Capability Loop (Resource Dependence Mechanism) | Connects evaluation performance with legitimacy, resource attraction, institutional capability, and subsequent performance |
| R7 | Cross-system | Internationalization Loop (Globalization Mechanism) | Links internationalization efforts, global networks, international participation, and evaluation performance |
| B5 | Cross-system | Strategic Alignment Loop (Managing Multiple Demands) | Responds to conflict among evaluation logics through prioritization, trade-offs, negotiation, and strategic alignment |
| B6 | Cross-system | Strategic Diversification Loop (Reducing Concentration Risk) | Reduces dependence on limited indicators by encouraging diversification of activities, indicators, and capabilities |
| B7 | Cross-system | Organizational Sustainability Loop (Maintaining Within Sustainable Capacity) | Responds to workload, complexity, resource strain, fatigue, burnout, and mission-displacement risk through capacity building and prioritization |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.