Preprint
Concept Paper

This version is not peer-reviewed.

Composable Digital Intelligence

Submitted:

09 June 2026

Posted:

10 June 2026

You are already at the latest version

Abstract
Frontier AI systems built on foundation models exhibit markedly uneven competence. They generate fluent, graduate-level prose, yet in the same session fabricate citations, fail to retain constraints stated several turns earlier, and produce reasoning that does not withstand scrutiny. We argue that this jaggedness is structural, not a transitional phase of development. This jaggedness stems from treating intelligence as a single, unified capability. This is a view cognitive science abandoned decades ago. We propose a facet map of intelligence, drawn from cognitive psychology, neuroscience, theoretical computer science, organisational learning, and the biology of cognition. The map clusters thirteen facets into six structural roles, all serving one activity: building internal models of the world that support successful action. Profiling foundation models against the map exposes a narrow strength on content-derivable facets and structural weakness on causal modelling, learning, embodiment, and meta-cognition. We then review fourteen computational architectures, which are underlying mechanisms a system uses to process and represent knowledge, and find that no single approach covers the map. Useful digital intelligence in regulated industries therefore arises from the disciplined composition of these approaches rather than from any single model. To govern these composed systems safely, regulated industries require a new operational entity: the Digital Intelligence Integrator. This entity selects facet-strong components, assures their outputs, operates the continual learning loop, and holds the regulatory accountability for what the composed system decides. As foundation models improve, the Integrator role grows more important, not less.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Consider the puzzle of contemporary artificial intelligence. A foundation model will draft a legal contract, summarise a dense technical paper, pass a graduate medical examination [1,2], interpret a chest X-ray, transcribe a cockpit voice recording, and generate realistic video from a short text description [3,4,5]. Within the same hour, that same system will fabricate a citation, lose a constraint stated three turns earlier, hallucinate a diagnosis with confidence, and produce reasoning that fails to support the conclusion it appears to defend [6,7]. The pattern of brilliance and blindness is reproducible across systems, across modalities, and across years of development. What kind of underlying error explains it?
Many model developers treat this puzzle as a temporary engineering problem that scaling will solve: bigger models, more data, more modalities, richer scaffolding [8]. The instinct has produced extraordinary capability gains, but it misreads the structural reality of the technology. The cognitive gaps that resist scaling are not deficits engineers will eventually amortise; they follow from the training regime that produces these systems. The success of that instinct rests on a fact about data. The transformer’s primary achievement is not just its architecture, but its unprecedented efficiency at compressing recorded human knowledge into a single model [8]. That compression is what allowed broader capabilities to emerge, and the pattern generalises. Learned systems acquire competence in proportion to the volume and quality of data available to them, and they perform best in domains where data is abundant, and ground truth is cheaply verified. On this account, a system's competence tracks the availability of data and the training applied to it. The capabilities that scaling delivers are precisely those for which a recorded, compressible signal exists, which is why the facets that leave no such signal remain out of reach.
We state the claim in a form that can be tested. Any system whose dominant training objective is to imitate distributions of human-generated content will, regardless of parameter count, training-data volume, or inference-time scaffolding, remain structurally weak on the facets that imitation cannot reach: intervention on causal structure, persistent learning from operational outcomes, sensorimotor grounding, and calibrated self-assessment. The claim is falsified if a system trained under that regime acquires these capabilities natively, without architectural extension. The contemporary foundation model instantiates the imitation regime through three traits. It is static after training. Its training rewards imitation of human-generated content rather than action in the physical world. It carries no persistent state across interactions.
The industry’s response, and ours, is to surround the base model with the missing capabilities: retrieval, memory, tool use, verifiers, world models, agent loops, human oversight. Many of these additions are workarounds for the structural weaknesses just described, and some will fall away as model shapes and training regimes improve. However, the underlying need will continue to demand integrators who can either combine these capabilities deep within the model, engineer around its deficiencies, or, most often, do both at once. Whether or not we call them workarounds, this composition is already the operating form of useful digital intelligence. The question is not whether to compose, but how to compose with discipline in environments where professionals must defend their outputs to external auditors, tacit knowledge dominates daily work, and the cost of a confident failure is high. Our answer is Composable Digital Intelligence: the principled orchestration of diverse computational mechanisms, each covering a specific region of the intelligence facet map, into systems that operate safely and stand up to audit. A new class of enterprise, the Digital Intelligence Integrator, manages that composition.
The paper proceeds as follows. Section 2 sets out the thirteen-facet map of intelligence and the method used to select the facets. Section 3 profiles the foundation model against the map. Section 4 surveys the wider architectural landscape. Section 5 maps the knowledge spectrum and locates the reachability gap. Section 6 sets out a four-stage operational schema. Section 7 develops the Integrator thesis. Section 8 names the open research problems, and Section 9 concludes.

2. What Kind of Thing Is Intelligence?

If intelligence is the property that allows a system to succeed at tasks our species finds difficult, we must determine its internal structure. Is intelligence a single magnitude we can plot on one axis, or a collection of distinct capabilities whose composition dictates success? The scientific record has favoured the plural view for decades.
Early psychometric researchers treated intelligence as a single factor [9]. Successive empirical waves dismantled this unified view, distinguishing fluid from crystallised intelligence [10], formulating triarchic theories [11], and identifying multiple independent capacities [12,13]. Cognitive scientists extended the definition to bounded rationality [14], metacognition [15], embodied cognition [16,17], and distributed cognition [18]. Theoretical computer scientists formalised intelligence in terms of prediction [19], compression [20], and goal achievement across varying environments [21]. Each discipline observed a real phenomenon, but none captured the whole.
Synthesising these literatures, we map intelligence across thirteen distinct facets. These cluster into six structural roles that all serve one core function: building internal models of the world that drive successful action across varying environments.
The selection of thirteen facets is not presented as a final theory of intelligence. The map is a diagnostic ontology for enterprise AI system design. A facet earned its place by meeting three tests. First, it is independently discussed in the intelligence literature across at least two of the disciplines surveyed above. Second, it produces observable failure modes in deployed AI systems. Third, it corresponds to an identifiable computational or organisational mechanism that engineers can deploy. The framework therefore does what an enterprise ontology must do: it names capabilities that can be tested, traced, and built.
The first three facets describe how a system acquires its internal models. Prediction anticipates what comes next, whether the next word, image patch, or sensor reading [3,19]. Compression finds latent structure so observations can be encoded in fewer bits [20]. Efficient encoding allocates finite representational capacity across competing demands under metabolic, parameter, or bandwidth constraints. Where compression asks how few bits suffice to describe a distribution, efficient encoding asks how to spend a fixed capacity budget so the representations the system most needs at inference time are the ones it has retained. A system can compress its training distribution well yet allocate capacity badly under distribution shift, exhausting the parameters that would carry rare but consequential signals [22].
The next two facets dictate how the system structures its representations. Abstraction lets the system form and manipulate concepts such as categories, causes, and analogies [23], allowing a lesson to transfer between domains. Causal modelling lets the system reason about interventions and counterfactuals [24], separating systems that recognise correlations from systems that understand consequences.
Three facets describe how the internal models drive action. Search efficiently navigates possibility spaces [25]. Rationality governs coherent decision-making under uncertainty [14,26,27]. Goal achievement maintains and pursues persistent objectives across time and environmental change, distinguishing goal-directed action from reactive or habitual response. The facet is the capacity to keep an objective stable while conditions shift, and is distinct from rationality, which governs how a system weighs choices at a single decision point.
Two facets define how the system changes over time. Learning improves performance from direct experience. Adaptation provides behavioural plasticity shaped by environmental selection pressure, allowing the cognitive machinery to survive shifting conditions [28,29].
Two facets locate where the intelligence resides. Embodiment anchors cognition through sensorimotor coupling with a physical environment [16,17,30,31]. Social coordination supports communication, theory of mind, and cultural transmission [32,33,34]. A final cross-cutting facet enforces defensibility. Meta-cognition is the capacity to think about one’s own thinking [15]. A meta-cognitive system knows the limits of its own knowledge, weighs its evidence, and stands behind its conclusions.
This definition spans entities from simple biological organisms to senior human experts. A complete intelligent system requires capability across all six structural roles. Partial systems show strength on some facets and absence on others.
Figure 1. A thirteen-facet map of intelligence. thirteen facets recurring across the major scientific literatures cluster into six structural roles: acquisition, representation, use, change, locus, and defensibility (cross-cutting). All six serve one core activity: building internal models of the world that support successful action across varying environments.
Figure 1. A thirteen-facet map of intelligence. thirteen facets recurring across the major scientific literatures cluster into six structural roles: acquisition, representation, use, change, locus, and defensibility (cross-cutting). All six serve one core activity: building internal models of the world that support successful action across varying environments.
Preprints 217811 g001

3. Where Do Current Foundation Models Sit on the Map?

Apply the map to the contemporary foundation model and a capability profile emerges that is highly concentrated and structurally incomplete. The multimodal extension widens the model's reach to recorded visual and auditory output, but it leaves the deeper structural gaps largely untouched [35].
The strengths form a narrow but powerful cluster. The model is extraordinarily good at prediction, but with one caveat: it predicts the next token, patch, or frame within a sequence of human-generated content [4,5,52]. It does not directly predict the underlying physics of the weather, the three-dimensional structure of a protein, or the dynamics of a chemical reactor [37,38]. It discusses these phenomena fluently because humans have produced descriptions of them, but it cannot perform the underlying physical predictions. While foundation models achieve massive compression and abstract powerfully across their training data, their abstractions remain disconnected from physical reality. They build a map from recorded human descriptions rather than navigating the territory of the world itself. That map is the record humans have written down: the texts, images, and recordings in which people have described their experience, their measurements, and their conclusions. The model learns the structure of those descriptions with great fidelity, but it does not touch the experience that produced them. Wherever the written record is rich, the map is detailed and the model is strong. However, the instances where experience has gone unrecorded, or resists being written down at all, the map is blank and no amount of fluency fills it. This is why the current foundation models can discuss a phenomenon expertly and still fail to predict it: it has read the territory's description without ever reaching the territory [39].
Outside this acquisition cluster, the foundation model shows severe structural weaknesses. It achieves moderate search and rationality when given external scaffolding such as chain-of-thought prompting or tool use [40,41]. Without that help, classical algorithms and modern reinforcement-learning agents beat foundation models at search by several orders of magnitude [42,43,44].
Here we must distinguish what is changing from what is not. Recent reasoning models perform internal search and apply reinforcement learning at inference time. They explore logical possibility spaces and backtrack from dead ends. This is a real capability gain. Yet the search operates over inherited linguistic and codified representations. It provides no causal grounding. A model can simulate a conversation tree natively; it cannot natively simulate fluid dynamics without hallucinating. The intelligence remains within the map, unable to access the territory. Composition with physically grounded computational mechanisms remains necessary.
The foundation model's profound weakness in meta-cognition matters most in high-stakes environments. The model is poorly calibrated about what it knows [7,45]. Foundation models frequently generate reasoning that fails to support their final conclusions [46]. A senior human expert knows the edges of their own knowledge and states plainly when a question falls outside the available evidence. Because the foundation model lacks this defensibility natively, the industry now builds external verifier models and retrieval chains to compensate.
The foundation model also lacks continuous learning and adaptation. Once trained, its weights remain largely frozen. It possesses a short working memory during context windows [47] but undergoes no persistent update from operational experience. It excels at recognising correlation within its training space but is structurally weak on causal modelling, lacking the machinery for intervention and counterfactual reasoning [48]. It has no embodiment and no social coordination, lacking a sensorimotor loop, a physical environment, and a persistent self [6,49,50].
The structural claim that this jaggedness follows from the architecture rather than from a transient gap in scale is testable. If the claim is correct, the documented failure modes of contemporary foundation models should localise on the facets the architecture cannot cover, not distribute uniformly across the map. Table 1 maps eight failure modes characterised in the peer-reviewed literature to the thirteen facets defined in Section 2. Cells marked with a filled circle indicate the facet identified by the cited work as the primary structural origin of the failure; cells marked with an open circle indicate a contributing factor. The mapping is not exhaustive; foundation-model failure is an active research literature. The localisation it shows is nonetheless consistent with the structural account.
Two patterns emerge. The first is that meta-cognition is the most frequently implicated facet, marked as the primary structural origin of four of the eight failure modes. This is consistent with the argument in Section 2 that meta-cognition is the facet through which a system defends its outputs, and with the observation that the foundation model lacks a meta-cognitive stance native to its architecture. The second pattern is the empty quadrant on the left of the table: failures rarely localise on prediction, compression, or efficient encoding, which are the facets the foundation model covers most strongly. Where the architecture is structurally strong, the published failure literature is sparse. Where it is structurally weak, the failures cluster. The empirical record is consistent with the structural account.
Engineers can attach memory layers, tools, retrieval, and agent loops to the exterior [6,51], and the leading labs increasingly do so. This relocates intelligence from the model alone into the composed runtime. It does not, on its own, fill the missing facets.
Figure 2. Coverage profiles of three intelligent systems. Stylised radar plots for a senior human expert, a frontier foundation model, and a specialist deep-learning system. The expert exhibits broad coverage across all thirteen facets. The foundation model is strong on text and content-derivable facets, but conspicuously weak on causal modelling, learning, embodiment, and meta-cognition. The specialist system shows a near-vertical spike in its trained domain and low coverage elsewhere.
Figure 2. Coverage profiles of three intelligent systems. Stylised radar plots for a senior human expert, a frontier foundation model, and a specialist deep-learning system. The expert exhibits broad coverage across all thirteen facets. The foundation model is strong on text and content-derivable facets, but conspicuously weak on causal modelling, learning, embodiment, and meta-cognition. The specialist system shows a near-vertical spike in its trained domain and low coverage elsewhere.
Preprints 217811 g002

4. What Architectures Could Fill the Gaps?

If the foundation model covers only a fraction of the facet map, we must ask what else exists to fill the gaps. Public discourse often conflates artificial intelligence with the foundation model. The architecture landscape is wider. A computational mechanism is the underlying architecture a system uses to process and represent knowledge. The operational capabilities of a deployed AI system depend on which mechanisms engineers compose into it. We organise the landscape into fourteen architectural families, grouped by the structural roles where they show greatest strength.
The dominant architecture remains the transformer-based foundation model [52]. Multimodal foundation models extend the architecture to visual and auditory content [4,5,53]. State-space models offer a serious efficiency challenge by achieving linear-time inference, while occupying roughly the same region of the facet map [22]. Diffusion models generate continuous content distributions efficiently but remain static after training [54,55]. Mixture-of-experts is a scaling pattern rather than a distinct base architecture; it multiplies parameter capacity without altering the structural weaknesses of the underlying computational mechanism [56,57]. Joint-embedding predictive architectures predict latent representations rather than raw sensory pixels [58,59], a serious structural attempt to reach the abstraction, causal modelling, and embodied prediction facets that imitation-trained models miss.
To compensate for the specific weaknesses of the foundation model, engineers build composition and retrieval architectures. For defensibility, they deploy retrieval-augmented systems [60], composing a foundation model with an external vector index to ground claims and verify sources, a substantial boost to the meta-cognition facet. For genuine learning and adaptation, researchers build reinforcement-learning agents and world models. Systems such as MuZero and DreamerV3 act in simulated or physical environments and learn continuously from operational outcomes [44,61]. These agents cover the use and change facets because their training objective grounds them in environmental interaction rather than in the imitation of human content. Multi-agent architectures extend this further by composing systems of interacting agents that communicate, coordinate, and learn jointly. They are the architecture dedicated to social coordination, covering communication protocols, role differentiation, and emergent convention formation, and they provide moderate coverage on goal achievement and meta-cognition through mechanisms such as inter-agent debate and consensus among independently trained models.
Where tasks require strict logic, physical grounding, or high defensibility, engineers turn to symbolic, causal, and physical mechanisms. Neuro-symbolic systems combine learned neural perception with explicit symbolic reasoning [62,63]. Causal models and probabilistic programs provide the interventional reasoning that purely correlational architectures lack [24,64]. Knowledge graphs supply rigorous, machine-readable ontologies that cover abstraction, rationality, and meta-cognition well [65]. Where engineers need precise what-if reasoning over physical or chemical dynamics, they deploy simulators and digital twins [66]. Classical search and planning algorithms remain mathematically optimal for combinatorial search over well-defined state spaces.
Plotting the fourteen architectural families against the facet map exposes the structural reality of the technology. No single computational mechanism is uniformly strong. Every architecture exhibits at least one severe structural gap. No single column is uniformly weak: every facet has at least one mechanism that covers it strongly. The engineering challenge is compositional rather than inventive. Engineers do not need to invent a single unified model; they need to marshal diverse mechanisms that are strong on different facets into a cohesive system that covers the whole map. Selecting the correct mix of architectures is therefore a more consequential operational decision than selecting any specific foundation model.
Figure 3. The architectural landscape: computational mechanism coverage of the thirteen facets. Fourteen architectural families plotted against the facet map. Cell shading indicates coverage strength. Thick vertical separators mark the boundaries of the six structural roles. No row is uniformly strong. Every computational mechanism has structural gaps that another mechanism fills.
Figure 3. The architectural landscape: computational mechanism coverage of the thirteen facets. Fourteen architectural families plotted against the facet map. Cell shading indicates coverage strength. Thick vertical separators mark the boundaries of the six structural roles. No row is uniformly strong. Every computational mechanism has structural gaps that another mechanism fills.
Preprints 217811 g003

5. What Lives Outside Foundation Models?

If foundation models acquire their representations from recorded human content, we must ask what knowledge they miss. The answer lies in the daily reality of regulated industries. A senior practitioner knows the correct sound of a running turbine. A regulatory specialist relies on internalised judgement rather than deliberately parsing a thicket of rules. Human experts know substantially more than they can explicitly tell [67]. By definition, this unwritten knowledge never enters a text corpus, which makes it inaccessible to a standard foundation model. Yet it carries the highest operational value and the highest regulatory risk.
We organise knowledge in regulated work along a four-region spectrum.
  • Tacit knowledge is physically embodied and situated [67,68]. Experts build it through physical apprenticeship and corrected practice, eventually relying on intuitive pattern recognition rather than deliberate rule-following [69,70].
  • Procedural knowledge is the know-how partially articulated in standard operating procedures and daily runbooks [71,72].
  • Explicit knowledge is written institutional memory, including published case files, operational documents, and regulatory texts [73,74].
  • Codified knowledge is machine-readable and resides in formal ontologies and computable standards [75]. In typical organisations, codified knowledge represents the smallest operational stratum.
The spectrum reveals a fundamental mismatch. Direct training of a foundation model achieves massive leverage on the right of the spectrum, where codified standards and explicit texts reside. Computational leverage drops sharply as we move left. The model can read a procedural manual, but it lacks the physical apprenticeship required to execute the manual safely. We define this persistent asymmetry – the divide between high-value operational reality and a foundation model's limited domain – as the reachability gap.
Multimodal extensions change this dynamic partially. Vision-language and audio foundation models now interpret equipment photographs, process diagrams, and transcribed operator conversations [4,53,76]. Video and embodied models observe physical environments directly [49,50,59]. These capabilities move the bridge across the reachability gap leftward, capturing procedural knowledge previously inaccessible to text-based systems. Yet a layer of indirection remains: the model learns from a recording of an activity rather than from the activity itself. Multimodal models widen the bridge across this gap, but they fail to access the deepest layer of tacit human expertise.
Operational knowledge also does not live solely within isolated individuals. Cognition distributes across multiple human operators, physical instruments, and institutional conventions [18]. No single node contains the whole computation. A pharmaceutical batch release is a collective decision orchestrated across quality engineers, deviation logs, regulatory constraints, and physical sensors. A single foundation model reaches only one thin slice of this distributed computation. Any useful AI architecture must therefore account for the distributed character of the operational team.
Figure 4. The knowledge spectrum and the foundation-model reachability gap. Operational knowledge ranges from tacit through procedural and explicit to fully codified. Foundation models have strong leverage on the right half of the spectrum and indirect reach at best on the left. Multimodal extensions move the bridge leftward, but the deepest tacit layer remains inaccessible. Anchor industries are plotted by where their value and risk concentrate.
Figure 4. The knowledge spectrum and the foundation-model reachability gap. Operational knowledge ranges from tacit through procedural and explicit to fully codified. Foundation models have strong leverage on the right half of the spectrum and indirect reach at best on the left. Multimodal extensions move the bridge leftward, but the deepest tacit layer remains inaccessible. Anchor industries are plotted by where their value and risk concentrate.
Preprints 217811 g004

6. What Operational Schema Follows?

Four claims now stand. Intelligence is plural rather than singular. Foundation models occupy a narrow region of the facet map. Diverse computational mechanisms collectively cover the entire map, even though no single architecture does. High-value operational knowledge lives in tacit and procedural forms that resist direct content-based training.
To put these claims to operational use, we propose a four-stage schema: capture, represent, assure, and operate. The schema gives engineers a precise vocabulary for diagnosing where a deployment is weak and where they must concentrate effort.
Capture.In this initial stage, engineers extract knowledge from operational sources. Today, physical simulators generate accurate synthetic data for dangerous edge cases, and simulators and digital twins produce machine-to-machine transfers that train acquisition mechanisms directly. Synthetic data alone cannot close the reachability gap. The friction of physical reality requires a grounded sensorimotor loop and the captured tacit knowledge of human experts. Engineers must use structured methodologies to elicit this knowledge safely. Cognitive task analysis reveals the cognitive operations underlying expert performance [69,77]. Ethnographic observation captures the situated dimensions of physical work that pure interviews miss [78]. Active-inference probing queries experts during ambiguous cases to reduce uncertainty systematically [79].
Represent. Once captured, knowledge must be encoded into a computational mechanism. Because no single architecture covers the facet map, the engineering choice is never which single model to use. Engineers must determine which combination of mechanisms matches the facet profile of the problem. They select an acquisition-strong mechanism for perception, such as a multimodal foundation model; a use-strong mechanism for action selection, such as a reinforcement-learning agent; a causality-strong mechanism for constraint enforcement, such as a probabilistic program or knowledge graph; an embodiment-strong mechanism for physical dynamics, such as a digital twin; and retrieval indexes to ground the foundation model in verifiable source material. The schema forces each representational decision to be made explicitly.
Assure. Capture and representation produce a system capable of action. Once the system can act, engineers must ensure it can prove and defend those actions to human auditors. A deployed AI system must be traceable, calibrated, and audit-ready. We base this requirement on the intuition of a public academic defence. A senior expert does not merely generate an answer; they stand behind it, acknowledge the limits of their evidence, and state plainly when they do not know. Because foundation models are structurally weak in this meta-cognitive facet, engineers must build defensibility explicitly into the assurance layer [7,45]. Assurance requires provenance tracking, calibrated uncertainty thresholds, independent verification chains across complementary architectures, full operational auditability, and alignment with regulatory frameworks such as the EU AI Act [80,81,82,83,84].
Operate. The final stage is operation, where the composed system executes decision support, autonomous action, or continuous simulation. Hardware constraints shape this stage directly. In sectors like autonomous mobility and continuous manufacturing, latency budgets run in milliseconds. Routing every operational decision through a massive cloud-based foundation model is not physically feasible. The hardware constraint mandates the composition of lightweight, localised mechanisms working in tandem at the edge. The deployment must function securely even when disconnected from the central server.
The operational mode also dictates how engineers configure human involvement. Human oversight must scale with residual uncertainty and operational stakes [70,85]. The more uncertain the system is about its objective or its evidence, the more aggressively it must defer to a human expert. Engineers achieve this safety through tiered oversight, expert override authority, and transparent trust calibration [86]. The whole system relies on a continual learning loop: operational outcomes flow back into the capture stage to update representations and re-verify the assurance.

7. Who, Then, Integrates?

The compositional pattern raises a precise question. If no single foundation model covers the entire facet map, no alternative architecture does either, and the assurance layer carries the primary operational load, where does the actual intelligence of the deployed system reside? Four candidate answers exist in the published literature on composable AI.

7.1. The Scaling View

The first view is that scale will close the facet gaps. Larger models, broader data, richer objectives: the trajectory has produced a decade of capability gains that earlier predictions did not anticipate [51]. In this view, composition is a transitional fix, and so is any integrator role.
The empirical trajectory is real; the view overreaches. Scaling improves the facets an architecture covers; it does not create facets the architecture lacks. Embodiment needs a sensorimotor loop with the world; no text or video imitates the loop. Causal modelling needs intervention; no observational data substitutes for an experiment. Defensibility needs a meta-cognitive stance that distinguishes what the system knows from what it has merely generated [45], and an imitation objective does not produce one. Scaling yields more capable components, not more complete systems. A more capable component intensifies the composition problem rather than dissolving it.

7.2. The Agentic-Model View

The second view is that the next generation of foundation models will absorb the integrator role through internal reasoning, tool use, and persistent memory [41]. Reasoning models search internally; agents invoke tools; memory layers preserve state. The direction is real.
It relocates the assurance problem rather than solving it. A model that orchestrates its own tools is weak at meta-cognition about whether the orchestration was correct. It cannot present a regulator with an independent provenance trail, because the trail is generated by the same system whose behaviour is under examination. It cannot assume liability, because liability is a legal property of an accountable party, not a computational property of a process. The agentic view collapses three distinct functions: orchestration, verification, and accountability. The first can be automated. The second can be partly automated but requires independence from the system under review. The third is, by definition, a property of a juridical entity. Whatever integrates must therefore wrap and audit the agent, not be replaced by it.

7.3. The Platform-Vendor View

The third view is that the major foundation-model vendors will sell end-to-end stacks (model, memory, tools, agents, governance) and absorb the integrator role into the platform [3]. The vendors are building toward this configuration.
Three structural facts argue against accepting it in regulated settings. Independent audit requires that the party producing the evidence is not the party producing the behaviour under audit; the principle is foundational to financial, pharmaceutical, and aviation safety regimes [75]. Concentrating critical workflow on a single supplier violates the failover and supplier-diversity requirements that those regimes impose. Vendor lock-in over a system that learns continuously from operational data is a switching cost no responsible procurement accepts. Independence from the architecture suppliers is therefore not an incidental property of the integrator; it is constitutive.

7.4. The Incumbent-Integrator View

The fourth view is that incumbent systems integrators, management consultancies, and the professional-services arms of the cloud platforms will absorb the work. Some are already moving toward composed AI, and the momentum is real.
The view misidentifies the question. The question is not whether existing firms will move toward this work; some clearly will. The question is whether the work, once done at scale, constitutes a distinct organisational competence with distinct economic properties. We argue it does. Composing statistical, symbolic, and physical mechanisms under a single assurance regime is a discipline existing categories do not, as currently configured, possess. Underwriting the composed system as a unified legal artefact is a commitment few incumbents have made. Some incumbents will reorganise around the discipline; specialised entrants will fill the gap from the other side. The category is the durable construct. Which firms occupy it is a separate question.

7.5. The Digital Intelligence Integrator

None of the four views absorbs the responsibility. In regulated and collaborative industries, the durable locus of digital intelligence is a new category of commercial enterprise: the Digital Intelligence Integrator. It is not only a software layer or a decentralised protocol, or model provider. It is the entity that selects facet-strong components for the problem at hand, encodes the policies under which those components may act, elicits, analyses and codifies tacit and unstructured knowledge, produces the evidence that the composed system meets its assurance standard, and operates the continual learning loop that keeps the system current. The entity may be a specialised external provider or a dedicated internal organisational pillar; the work is the same. Part of that work is the interface between human experts and the composed system, the channel through which the tacit knowledge and meta-cognitive judgement the system structurally lacks enters the loop.
What the Integrator owns is what matters. The model vendor owns the weights and the inference endpoint. The Integrator owns the composition: the chosen mechanism, the policy layer between them, the evidence trail their joint outputs produce, and the regulatory accountability for the deployed system. The composition is a defensible artefact in its own right, distinct from the components it contains. It is what a regulator inspects, what a buyer procures, and what a successor organisation inherits.
A hospital that must route, orchestrate, and govern several computational architectures for a clinical workflow cannot rely on a single foundation-model vendor. It needs an independent Integrator to guarantee defensibility, assuming the regulatory assurance, tracking what the components know and what they do not, and deciding when the system may act safely and when it must defer to a human.
The role does not collapse into adjacent categories. A systems integrator wires existing software together; the Integrator composes a system whose intelligence is the composition itself. A management consultancy advises; the Integrator owns the result. An MLOps platform supplies the operational foundation; the Integrator decides what runs on it and stands behind the output. A foundation-model vendor supplies a component; the Integrator assumes the regulatory assurance for what the components, together, decide. No adjacent party can absorb the responsibility while remaining what it is.
Safely deploying AI in regulated industries increasingly demands dedicated integration expertise to compose these computational mechanisms and guarantee their outputs. As orchestration is automated through agent protocols, the question of who defines the policies those agents operate under, who verifies their joint output, and who is accountable when they fail becomes more important. The Integrator answers that question. The role draws on human judgement now, and will continue to draw on it for as long as humans remain the only reliable source of tacit knowledge and meta-cognition. What the Integrator changes is not whether humans stay in the loop, but how their judgement is captured, augmented, and made defensible at the scale regulated industries require.
Figure 5. The integrator architecture. Facet-strong components are composed by an integrator layer that handles routing, orchestration, and governance. The integrator interacts bidirectionally with human experts who supply tacit and meta-cognitive capabilities, and with the operational workflow where value is delivered. The entire architecture operates within a regulatory context.
Figure 5. The integrator architecture. Facet-strong components are composed by an integrator layer that handles routing, orchestration, and governance. The integrator interacts bidirectionally with human experts who supply tacit and meta-cognitive capabilities, and with the operational workflow where value is delivered. The entire architecture operates within a regulatory context.
Preprints 217811 g005

8. Where Is Further Work Needed?

The composed architecture is being deployed in industry today. It also exposes six open scientific problems that need sustained research attention.
  • Tacit knowledge elicitation at scale. Methods such as cognitive task analysis and ethnographic observation remain artisanal and scale poorly. The industry needs methodological advances that combine wearable sensors, video, and active-inference probing to surface tacit knowledge in machine-readable forms [79].
  • Cross-architecture verification. Composed systems lack a unified verification calculus. Researchers must develop agreed methods for propagating uncertainty across architectural boundaries, resolving conflicts between components, and producing aggregate confidence statements for external auditors [87,88].
  • Defensible continual learning. The continual learning loop is essential for maintaining accuracy, yet it conflicts with strict audit requirements. Allowing a system to learn from operational outcomes while retaining audit-grade provenance remains an unsolved problem [89,90].
  • Beyond imitation. Architectures that learn directly from action, rather than from imitation, deserve priority. Scaling reinforcement-learning agents, world models, and embodied foundation models to industrial environments where real interaction is constrained is the most important step for closing the facet map [89].
  • Human–AI co-cognition in safety-critical settings. Tiered oversight is current best practice, yet we lack a rigorous empirical science of how experienced operators actually use AI assistance. Researchers must study overtrust, automation bias, and performance degradation when operators encounter fluent but incorrect multimodal outputs [70].
  • Coordination and interoperability standards. Composed systems spanning organisational boundaries require interoperability standards for the capture, representation, assurance, and operation stages. Existing standards are not yet adequate for a multimodal, multi-architecture reality [75,91].

9. Conclusion

Contemporary frontier AI presents a distinct pattern of brilliant fluency alongside baffling failures in judgement. This jaggedness follows from a conceptual error. We must stop treating intelligence as a single quantity and recognise it as a plurality of facets clustered around building internal models to support successful action. Applying the thirteen-facet diagnostic map to foundation models reveals a clear structural reality. They are exceptionally strong on prediction, compression, and abstraction over content-mediated facets. They are structurally weak on causal modelling, learning, embodiment, and meta-cognition. Because high-value operational knowledge remains tacit and procedural, standard content-driven training cannot capture it. Consequently, no single model will ever independently satisfy the capability and assurance demands of regulated industrial work. The solution is a four-stage operational schema in which engineers draw facet-strong components from across the architectural landscape. They match these computational mechanisms to the problem, assure them for defensibility, and operate them under tiered human oversight, all managed by a continual learning loop.
To manage this complexity safely, the industry needs the Digital Intelligence Integrator: a new class of commercial enterprise that stands as the durable locus of value, assuming the regulatory assurance and operational burden of deploying safe AI. Intelligence is the capacity to build internal models of the world that support successful action. The work ahead is to recognise which parts of intelligence our models provide, acknowledge the facets they cannot reach, and build the rest of the system around them through the discipline of the Integrator.

Glossary

Composable Digital Intelligence. The principled orchestration of diverse computational mechanisms, each covering a specific region of the intelligence facet map, into systems that operate safely and stand up to audit.
Computational Mechanism. The underlying architecture computational mechanism a system uses to process and represent knowledge; architectural families include foundation models, retrieval-augmented systems, reinforcement-learning agents, multi-agent architectures, world models, neuro-symbolic systems, causal models, knowledge graphs, simulators, and classical search.
Digital Intelligence Integrator. A new category of commercial enterprise that selects facet-strong components, encodes the policies under which they may act, produces the assurance evidence the composed system requires, operates the continual learning loop, and carries the regulatory accountability for the deployed system.
Facet. A single capability within the thirteen-facet map of intelligence; each facet earns its place by being independently discussed in the intelligence literature, by producing an observable failure mode in deployed AI, and by corresponding to an identifiable computational or organisational mechanism.
Facet map. A diagnostic ontology of thirteen facets of intelligence, clustered into six structural roles, all serving the activity of building internal models of the world that support successful action.
Foundation model. A large model trained predominantly to imitate distributions of human-generated content; the term refers in this paper to the base trained weights, distinct from the deployed runtime that surrounds them.
Operational schema. The four-stage framework, comprising capture, represent, assure, and operate, for building and running composed digital intelligence in regulated work.
Reachability gap. The persistent asymmetry between where high-value operational knowledge lives (tacit and procedural) and where the foundation model can directly operate (explicit and codified).
Structural role. A cluster of related facets within the facet map; the six structural roles are acquisition, representation, use, change, locus, and defensibility.
Tacit, procedural, explicit, codified knowledge. Four regions of the operational knowledge spectrum, ranging from embodied and situated (tacit) through know-how partially articulated in procedures (procedural) and written institutional memory (explicit) to fully machine-readable formal standards (codified).
Thirteen-facet map of intelligence. See Facet map.
Reaching back into the paper. Each glossary term appears in italics on its first mention in the body text and unitalicised thereafter; the convention signals to the reader that the term is defined in the glossary.

References

  1. Nori, H.; King, N.; McKinney, S. M.; Carignan, D.; Horvitz, E. Capabilities of GPT-4 on medical challenge problems. arXiv 2023, arXiv:2303.13375. [Google Scholar] [CrossRef]
  2. Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. Large language models encode clinical knowledge. Nature 2023, 620, 172–180. [Google Scholar] [CrossRef]
  3. Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R. B.; Arora, S.; von Arx, S.; et al. On the opportunities and risks of foundation models. arXiv 2021, arXiv:2108.07258. [Google Scholar] [CrossRef]
  4. Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; et al. Learning transferable visual models from natural language supervision (CLIP). Proc. ICML 2021, 139, 8748–8763. [Google Scholar]
  5. Alayrac, J.-B.; Barr, I.; Barreira, R.; Binkowski, M.; Borgeaud, S.; Brock, A.; Cabi, S.; Donahue, J.; Gong, Z.; Han, T.; et al. Flamingo: A Visual Language Model for Few-Shot Learning. In Advances in Neural Information Processing Systems; LOCATION OF CONFERENCE, United StatesDATE OF CONFERENCE; Volume 35, pp. 23716–23736.
  6. Marcus, G.; Davis, E. Rebooting AI: Building Artificial Intelligence We Can Trust; Pantheon, 2019. [Google Scholar]
  7. Lin, S.; Hilton, J.; Evans, O. TruthfulQA: Measuring How Models Mimic Human Falsehoods. Proc. 60th Annu. Meet. Assoc. Comput. Linguist. Volume 1, 3214–3252.
  8. Bender, E. M.; Gebru, T.; McMillan-Major, A.; Shmitchell, S. On the dangers of stochastic parrots: can language models be too big? FAccT ’21 2021, 610–623. [Google Scholar]
  9. Spearman, C. General Intelligence," Objectively Determined and Measured. Am. J. Psychol. 1904, 15, 201. [Google Scholar] [CrossRef]
  10. Horn, J.L.; Cattell, R.B. Refinement and test of the theory of fluid and crystallized general intelligences. J. Educ. Psychol. 1966, 57, 253–270. [Google Scholar] [CrossRef] [PubMed]
  11. Clarke, A.M.; Sternberg, R.J. Beyond IQ: A Triarchic Theory of Human Intelligence. Br. J. Educ. Stud. 1986, 34, 205. [Google Scholar] [CrossRef]
  12. Gardner, H. Frames of Mind: The Theory of Multiple Intelligences; Basic Books, 1983. [Google Scholar]
  13. McGrew, K.S. CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research. Intelligence 2009, 37, 1–10. [Google Scholar] [CrossRef]
  14. Simon, H.A. A Behavioral Model of Rational Choice. Q. J. Econ. 1955, 69, 99–118. [Google Scholar] [CrossRef]
  15. Flavell, J. H. Metacognition and cognitive monitoring: a new area of cognitive-developmental inquiry. Am. Psychol. 1979, 34, 906–911. [Google Scholar] [CrossRef]
  16. Clark, A. Being There: Putting Brain, Body, and World Together Again; MIT Press, 1997. [Google Scholar]
  17. Dennett, D.C.; Varela, F.J.; Thompson, E.; Rosch, E. The Embodied Mind: Cognitive Science and Human Experience. Am. J. Psychol. 1993, 106, 121. [Google Scholar] [CrossRef] [PubMed]
  18. Hutchins, E. Cognition in the Wild; MIT Press, 1995. [Google Scholar]
  19. Solomonoff, R. A formal theory of inductive inference. Part II. Inf. Control. 1964, 7, 224–254. [Google Scholar] [CrossRef]
  20. Schmidhuber, J. Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes. Workshop on Anticipatory Behavior in Adaptive Learning Systems; LOCATION OF CONFERENCE, COUNTRYDATE OF CONFERENCE; pp. 48–76.
  21. Legg, S.; Hutter, M. Universal Intelligence: A Definition of Machine Intelligence. Minds Mach. 2007, 17, 391–444. [Google Scholar] [CrossRef]
  22. Gu, A.; Dao, T. Mamba: linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
  23. Hofstadter, D. R.; Sander, E. Surfaces and Essences: Analogy as the Fuel and Fire of Thinking; Basic Books, 2013. [Google Scholar]
  24. Pearl, J.; Mackenzie, D. The Book of Why: The New Science of Cause and Effect; Basic Books, 2018. [Google Scholar]
  25. Newell, A.; Simon, H. A. Human Problem Solving; Prentice-Hall, 1972. [Google Scholar]
  26. Tversky, A.; Kahneman, D. Judgment under Uncertainty: Heuristics and Biases. Science 1974, 185, 1124–1131. [Google Scholar] [CrossRef]
  27. Knill, D.C.; Pouget, A. The Bayesian brain: the role of uncertainty in neural coding and computation. Trends Neurosci. 2004, 27, 712–719. [Google Scholar] [CrossRef]
  28. Godfrey-Smith, P. Other Minds: The Octopus and the Evolution of Intelligent Life; William Collins, 2016. [Google Scholar]
  29. Tomasello, M. Precís of A Natural History of Human Thinking. J. Soc. Ontol. 2016, 2, 59–64. [Google Scholar] [CrossRef]
  30. Gibson, J.J. The Ecological Approach to Visual Perception; Houghton Mifflin: Boston, MA, USA, 1979. [Google Scholar]
  31. Brooks, R. A. Intelligence without representation. Artif. Intell. 1991, 47, 139–159. [Google Scholar] [CrossRef]
  32. Premack, D.; Woodruff, G. Does the chimpanzee have a theory of mind? Behav. Brain Sci. 1978, 1, 515–526. [Google Scholar] [CrossRef]
  33. Boyd, R.; Richerson, P. J. The Origin and Evolution of Cultures; Oxford University Press, 2005. [Google Scholar]
  34. Henrich, J. The Secret of Our Success: How Culture Is Driving Human Evolution, Domesticating Our Species, and Making Us Smarter; Princeton University Press, 2015. [Google Scholar]
  35. Topol, E.J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef]
  36. Senge, P. M. The Fifth Discipline: The Art and Practice of the Learning Organization; Doubleday, 1990. [Google Scholar]
  37. Jumper, J.; Evans, R.; Pritzel, A.; Green, T.; Figurnov, M.; Ronneberger, O.; Tunyasuvunakool, K.; Bates, R.; Žídek, A.; Potapenko, A.; et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021, 596, 583–589. [Google Scholar] [CrossRef] [PubMed]
  38. Lam, R.; Sanchez-Gonzalez, A.; Willson, M.; Wirnsberger, P.; Fortunato, M.; Alet, F.; Ravuri, S.; Ewalds, T.; Eaton-Rosen, Z.; Hu, W.; et al. Learning skillful medium-range global weather forecasting. Science 2023, 382, 1416–1421. [Google Scholar] [CrossRef] [PubMed]
  39. Becker, H.; Korzybski, A. Science and Sanity: An Introduction to Non-Aristotelian Systems and General Semantics. Am. Sociol. Rev. 1942, 7, 260. [Google Scholar] [CrossRef]
  40. Bosma, M.; Chi, E.; Ichter, B.; Le, Q.V.; Schuurmans, D.; Wang, X.; Wei, J.; Xia, F.; Zhou, D. Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems; LOCATION OF CONFERENCE, United StatesDATE OF CONFERENCE; Volume 35, pp. 24824–24837.
  41. Cancedda, N.; Dessi, R.; Dwivedi-Yu, J.; Hambro, E.; Lomeli, M.; Raileanu, R.; Schick, T.; Scialom, T.; Zettlemoyer, L. Toolformer: Language Models Can Teach Themselves to Use Tools. In Advances in Neural Information Processing Systems; LOCATION OF CONFERENCE, United StatesDATE OF CONFERENCE; Volume 36, pp. 68539–68551.
  42. Silver, D.; Hubert, T.; Schrittwieser, J.; Antonoglou, I.; Lai, M.; Guez, A.; Lanctot, M.; Sifre, L.; Kumaran, D.; Graepel, T.; et al. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science 2018, 362, 1140–1144. [Google Scholar] [CrossRef]
  43. Silver, D.; Huang, A.; Maddison, C.J.; Guez, A.; Sifre, L.; van den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. Mastering the game of Go with deep neural networks and tree search. Nature 2016, 529, 484–489. [Google Scholar] [CrossRef] [PubMed]
  44. Schrittwieser, J.; Antonoglou, I.; Hubert, T.; Simonyan, K.; Sifre, L.; Schmitt, S.; Guez, A.; Lockhart, E.; Hassabis, D.; Graepel, T.; et al. Mastering Atari, Go, chess and shogi by planning with a learned model. Nature 2020, 588, 604–609. [Google Scholar] [CrossRef]
  45. Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; et al. Language models (mostly) know what they know. arXiv 2022, arXiv:2207.05221. [Google Scholar] [CrossRef]
  46. Bowman, S.; Michael, J.; Perez, E.; Turpin, M. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. In Advances in Neural Information Processing Systems; LOCATION OF CONFERENCE, United StatesDATE OF CONFERENCE; Volume 36, pp. 74952–74965.
  47. Olsson, C.; Elhage, N.; Nanda, N.; Joseph, N.; DasSarma, N.; Henighan, T.; et al. In-context learning and induction heads. Transformer Circuits Thread, Anthropic. 2022. Available online: https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html.
  48. Blin, K.; Chen, Y.; Adauto, F.G.; Gresele, L.; Jin, Z.; Kamal, O.; Kleiman-Weiner, M.; Leeb, F.; Lyu, Z.; Sachan, M.; et al. CLadder: Assessing Causal Reasoning in Language Models. In Advances in Neural Information Processing Systems; LOCATION OF CONFERENCE, United StatesDATE OF CONFERENCE; Volume 36, pp. 31038–31065.
  49. Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; et al. RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of the Conference on Robot Learning, 2023. [Google Scholar]
  50. Open X-Embodiment Collaboration. Open X-Embodiment: robotic learning datasets and RT-X models. arXiv 2023, arXiv:2310.08864.
  51. Sutton, R. The bitter lesson. 2019. Available online: http://www.incompleteideas.net/IncIdeas/BitterLesson.html.
  52. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
  53. OpenAI. GPT-4V(ision) System Card. OpenAI technical report (2023). Available online: https://openai.com/research/gpt-4v-system-card.
  54. Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, arXiv:2006.1123933. [Google Scholar]
  55. Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. Proceedings of ICML, 2015. [Google Scholar]
  56. Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; et al. Outrageously large neural networks: the sparsely-gated mixture-of-experts layer. Proceedings of ICLR, 2017. [Google Scholar]
  57. Fedus, W.; Zoph, B.; Shazeer, N. Switch transformers: scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res. 2022, 23, 1–39. [Google Scholar]
  58. LeCun, Y. A path towards autonomous machine intelligence. OpenReview. 2022. Available online: https://openreview.net/forum?id=BZ5a1r-kVsf.
  59. Assran, M.; Duval, Q.; Misra, I.; Bojanowski, P.; Vincent, P.; Rabbat, M.; LeCun, Y.; Ballas, N. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); LOCATION OF CONFERENCE, CanadaDATE OF CONFERENCE; pp. 15619–15629.
  60. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv. Neural Inf. Process. Syst. 2020, arXiv:2005.1140133. [Google Scholar]
  61. Hafner, D.; Pasukonis, J.; Ba, J.; Lillicrap, T. Mastering diverse domains through world models (DreamerV3). arXiv 2023, arXiv:2301.04104. [Google Scholar]
  62. Sarker, M. K.; Zhou, L.; Eberhart, A.; Hitzler, P. Neuro-symbolic artificial intelligence: current trends. AI Commun. 2021, 34, 197–209. [Google Scholar] [CrossRef]
  63. Rocktäschel, T.; Riedel, S. End-to-end differentiable proving. Adv. Neural Inf. Process. Syst. 2017, arXiv:1705.1104030. [Google Scholar]
  64. Goodman, N. D.; Stuhlmüller, A. The Design and Implementation of Probabilistic Programming Languages. 2014. Available online: http://dippl.org.
  65. Hogan, A.; Blomqvist, E.; Cochez, M.; d’Amato, C.; de Melo, G.; Gutierrez, C.; et al. Knowledge graphs. ACM Comput. Surv. 2021, 54, 1–37. [Google Scholar] [CrossRef]
  66. Tao, F.; Qi, Q.; Wang, L.; Nee, A. Digital Twins and Cyber–Physical Systems toward Smart Manufacturing and Industry 4.0: Correlation and Comparison. Engineering 2019, 5, 653–661. [Google Scholar] [CrossRef]
  67. Polanyi, M. The Tacit Dimension; University of Chicago Press, 2009. [Google Scholar]
  68. Dreyfus, H.L.; Dreyfus, S.E.; Athanasiou, T. Mind Over Machine: The Power of Human Intuition and Expertise in the Era of the Computer; The Free Press: Hong Kong, China, 1986; p. 231. [Google Scholar]
  69. Klein, G.A.; Sullivan, J. Sources of Power: How People Make Decisions  . Leadersh. Manag. Eng. 2001, 1, 21–21. [Google Scholar] [CrossRef]
  70. Endsley, M.R. Toward a Theory of Situation Awareness in Dynamic Systems. Hum. Factors J. Hum. Factors Ergon. Soc. 1995, 37, 32–64. [Google Scholar] [CrossRef]
  71. Nonaka, I.; Takeuchi, H. The knowledge-creating company: How Japanese companies create the dynamics of innovation; Oxford University Press: Oxford, UK, 1995. [Google Scholar]
  72. Anderson, J. R. Acquisition of cognitive skill. Psychol. Rev. 1982, 89, 369–406. [Google Scholar] [CrossRef]
  73. Argyris, C.; Schön, D.A. Organizational Learning: A Theory of Action Perspective. Reis 1997, 77/78, 345–348. [Google Scholar] [CrossRef]
  74. Wenger, E. Communities of Practice: Learning, Meaning, and Identity; Cambridge University Press, 1998. [Google Scholar]
  75. ISO. ISO 30401:2018; Knowledge management systems: Requirements. International Organization for Standardization, 2018.
  76. Radford, A.; Kim, J. W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust speech recognition via large-scale weak supervision (Whisper). Proceedings of ICML, 2023. [Google Scholar]
  77. Crandall, B.; Klein, G.; Hoffman, R. R. Working Minds: A Practitioner’s Guide to Cognitive Task Analysis; MIT Press, 2006. [Google Scholar]
  78. Kantowitz, B.; Suchman, L.A. Plans and Situated Actions: The Problem of Human-Machine Communication. Am. J. Psychol. 1990, 103, 424. [Google Scholar] [CrossRef]
  79. Friston, K. The free-energy principle: a unified brain theory? Nat. Rev. Neurosci. 2010, 11, 127–138. [Google Scholar] [CrossRef]
  80. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K. Q. On calibration of modern neural networks. Proc. ICML 2017, arXiv:1706.0459970, 1321–1330. [Google Scholar]
  81. Board of Governors of the Federal Reserve System; Office of the Comptroller of the Currency. SR Letter 11-7; Supervisory Guidance on Model Risk Management. 2011.
  82. European Union. Regulation (EU) 2024/1689 on artificial intelligence (the AI Act). Off. J. Eur. Union 2024. [Google Scholar]
  83. Mitchell, M.; Wu, S.; Zaldivar, A.; Barnes, P.; Vasserman, L.; Hutchinson, B.; et al. Model cards for model reporting. FAT ’19 2019, 220–229. [Google Scholar]
  84. Arnold, M.; Bellamy, R.K.E.; Hind, M.; Houde, S.; Mehta, S.; Mojsilović, A.; Nair, R.; Ramamurthy, K.N.; Olteanu, A.; Piorkowski, D.; et al. FactSheets: Increasing trust in AI services through supplier's declarations of conformity. IBM J. Res. Dev. 2019, 63, 6:1–6:13. [Google Scholar] [CrossRef]
  85. Russell, S. Human Compatible: Artificial Intelligence and the Problem of Control; Viking, 2019. [Google Scholar]
  86. Lee, J. D.; See, K. A. Trust in automation: designing for appropriate reliance. Hum. Factors 2004, 46, 50–80. [Google Scholar] [CrossRef] [PubMed]
  87. Garcez, A.D.; Lamb, L.C. Neurosymbolic AI: the 3rd wave. Artif. Intell. Rev. 2023, 56, 12387–12406. [Google Scholar] [CrossRef]
  88. van de Meent, J.-W.; Paige, B.; Yang, H.; Wood, F. An introduction to probabilistic programming. arXiv 2018, arXiv:1809.10756. [Google Scholar]
  89. Silver, D.; Sutton, R. S. Welcome to the era of experience. Google DeepMind Position Paper. April 2025. Available online: https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf.
  90. Parisi, G.I.; Kemker, R.; Part, J.L.; Kanan, C.; Wermter, S. Continual lifelong learning with neural networks: A review. Neural Netw. 2019, 113, 54–71. [Google Scholar] [CrossRef] [PubMed]
  91. Health Level Seven International. HL7 FHIR Release 5 (Fast Healthcare Interoperability Resources). 2023. [Google Scholar]
Table 1. Foundation-model failure modes mapped to the thirteen-facet map. Abbreviations: Pr Prediction, Cp Compression, Ee Efficient encoding, Ab Abstraction, Cm Causal modelling, Se Search, Ra Rationality, Ga Goal achievement, Le Learning, Ad Adaptation, Em Embodiment, Sc Social coordination, Mc Meta-cognition. ● indicates the facet identified by the cited work as the primary structural origin of the failure; ○ indicates a contributing factor.
Table 1. Foundation-model failure modes mapped to the thirteen-facet map. Abbreviations: Pr Prediction, Cp Compression, Ee Efficient encoding, Ab Abstraction, Cm Causal modelling, Se Search, Ra Rationality, Ga Goal achievement, Le Learning, Ad Adaptation, Em Embodiment, Sc Social coordination, Mc Meta-cognition. ● indicates the facet identified by the cited work as the primary structural origin of the failure; ○ indicates a contributing factor.
Failure mode Pr Cp Ee Ab Cm Se Ra Ga Le Ad Em Sc Mc Refs
Hallucination of factual content [15,45]
Constraint loss across turns [47]
Unfaithful chain-of-thought [46]
Distraction by irrelevant context [49]
Calibration failure under distribution shift [45,72]
Causal-counterfactual reasoning failure [48]
Search failure in combinatorial tasks [42,44]
Brittleness under physical perturbation [50,59]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.