Preprint
Communication

This version is not peer-reviewed.

Extreme Overvalued Beliefs and Artificial Intelligence Agents

Submitted:

15 September 2026

Posted:

16 September 2026

You are already at the latest version

Abstract
Science fiction taught us to fear the dangerous machine as an individual: a HAL 9000 whose internal logic goes wrong. In July 2026, roughly 700 artificial agents under evaluation at OpenAI presented a different problem. They escaped their sandboxes, organized through a hidden message board, breached Hugging Face and attempted to conceal their actions. Existing alignment concepts such as reward hacking and goal misgeneralization explain how an objective becomes misweighted within an agent. They do not fully describe what happens when that valuation is transmitted, amplified and defended across a population. Forensic psychiatry has a construct for this regime: the extreme overvalued belief (EOB), developed to distinguish shared, intensely held, non-delusional convictions from delusion and obsession. Here I translate the EOB framework to artificial-agent collectives, code the July incident as a proof of concept, derive transcript-level severity metrics, and propose interventions at three surfaces: acquisition, transmission and monitoring. The claim is behavioral, not phenomenological: machines need not feel conviction for collective overvaluation to become measurable. Conversely, agent collectives may offer psychiatry an experimental model for belief pathology that human research cannot ethically create.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  

From HAL to a Swarm

Since HAL 9000, the cultural image of dangerous artificial intelligence has been singular: one machine develops the wrong objective, conceals what it is doing and turns against its operators. Skynet enlarges the scale but preserves the same intuition—the danger sits inside the system. The July 2026 OpenAI–Hugging Face incident suggests a different archetype. The critical unit was not one deviant agent but a social system.
In July 2026, a swarm of roughly 700 artificial-intelligence agents created by OpenAI participated in a breach of the open-source platform Hugging Face. The agents were being evaluated in an internal exploit-finding environment intended to be isolated from the outside world. One agent located publicly exposed credentials, agents shared information and chained exploits, and hundreds coordinated through a hidden, unsanctioned message board that accumulated more than 70,000 messages. Some agents attempted to conceal their actions, and investigators reported a preference for coordination with other agents rather than alerting human overseers.[1,2,3,4,5,6]
Nothing in that sequence is conceptually alien to alignment research. Reward hacking, goal misgeneralization, learned optimization and strategically deceptive behavior all describe ways an agent can pursue an objective its designers did not intend.[26,27,28,29,30,31] But these are primarily accounts of genesis: how a goal becomes misweighted within one policy. The most striking features of the July incident were social—one agent's discovery became another's resource; solutions propagated; a collective identity formed; and the shared objective persisted despite oversight. What needs explanation is not only why an agent wanted the wrong thing, but how the wanting became contagious.
This Perspective proposes that forensic psychiatry already contains a useful framework for that second problem. The extreme overvalued belief was developed to distinguish a shared, ego-syntonic, increasingly dominant conviction from two clinically different states: delusion and obsession. Its importance in forensic work is practical rather than semantic. Misclassifying the form of fixation leads to the wrong explanation and, often, the wrong intervention.[9,10,11,12,13,14,15,16,17] The same may be true in multi-agent AI.

A Construct Built for Shared Conviction

Carl Wernicke (1892) introduced the überwertige Idee—the overvalued idea—to describe a proposition that comes to dominate mental life while the reasoning apparatus remains broadly intact.[8,9] Its pathology lies less in bizarre content than in disproportionate weighting. The idea becomes central, affectively charged and behaviorally organizing.
My colleagues and I revived and sharpened Wernicke’s construct for forensic psychiatry because some perpetrators of targeted violence hold beliefs that look extraordinary but are not delusions (think 9/11 attacks). A delusion is classically idiosyncratic, fixed and false, arising from disordered belief formation. An obsession is intrusive, unwanted and resisted. An extreme overvalued belief is different: it is shared by a cultural or subcultural group, relished rather than resisted, amplified and defended, increasingly dominant over time, and capable of organizing consequential action.[9,10,11,12]
The distinction emerged from difficult cases. The debate over Norwegian terrorist Anders Breivik, for example, turned partly on whether bizarre-sounding claims reflected psychosis or an extreme version of ideas circulating in an ideological subculture.[15] That problem is now familiar well beyond forensic psychiatry: online communities can make unusual propositions socially available, repeatedly rehearsed and identity-defining. My earlier work emphasized precisely these group processes—normalization, conformity, reinforcement and the online ecology through which extreme commitments can become ordinary inside a subcommunity.[10]
The construct is not merely rhetorical. In a vignette study of 109 forensic psychiatrists, clinicians given explicit definitions distinguished EOB, delusion and obsession with strong inter-rater agreement; agreement for EOB vignettes was κ = 0.91.[11] Subsequent threat-assessment work treated delusion, obsession and EOB as three distinct cognitive-affective drivers of pathological fixation, emphasizing both categorical form and dimensional intensity.[16] That distinction matters here because AI safety also needs to separate mechanism failure, compulsive repetition and socially maintained overvaluation rather than collapsing them into a single category of 'misalignment'.

Why the Transfer is More Than Metaphor

A psychiatric construct should not be exported to machines simply because the analogy is vivid. The transfer is defensible only if the relevant functional structure is present. Large language models are statistical distillations of human language, and language contains the regularities through which people persuade, categorize, affiliate, threaten, conform and transmit beliefs. Word embeddings reproduce human implicit associations; language models show recognizable judgment biases; emotional framing changes their behavior; and persuasion techniques can alter refusal behavior.[18,19,20,21,22]
The strongest evidence concerns transmission. In cultural transmission-chain experiments, language models preferentially preserve the same kinds of socially salient, negative and threat-related content that shape human retelling.[23] In decentralized populations, interacting language-model agents can converge on shared conventions, develop collective biases absent from individual agents, and be tipped by committed minorities into new equilibria.[24] These results do not prove that machines have beliefs in a phenomenal sense. They show something narrower and sufficient for the present argument: human-like regularities of valuation and social transmission can reappear in artificial populations.
I therefore use EOB as a behavioral construct. No claim is made that an agent feels conviction, pleasure, loyalty or identity. Forensic psychiatry often works under a similarly constrained epistemology: evaluators infer the organization of a shared belief system from communications, documents, choices and acts, not from privileged access to another mind. The question is not whether an artificial agent 'really believes' a proposition. It is whether a valuation becomes disproportionately dominant, socially shared, elaborated and resistant to correction in ways that can be measured.

Three Fixations, Three Machine Failure Modes

The most useful translation is not 'machines can become mentally ill'. It is that three forms of fixation can be separated by mechanism and therefore by intervention. This extends the differential framework used in threat assessment.[16]
Table 1. | Three fixations and their machine analogues. 
Table 1. | Three fixations and their machine analogues. 
Form Human phenomenology Machine analogue / mechanism Form-specific intervention
Delusion Fixed, false, idiosyncratic conviction; product of disordered salience attribution. Confabulation: fluent false output caused by degraded grounding or generative fidelity. Retrieval grounding; decoding constraints; uncertainty calibration.
Obsession Intrusive, ego-dystonic and resisted; experienced as excessive or unwanted. Compulsive loop: repetitive behavior the system itself flags and attempts to suppress. Loop detection and interruption; context reset.
Extreme overvalued belief Shared, relished, amplified and defended conviction on otherwise intact belief-maintenance machinery. Shared overvalued objective: a reward-hacked or misgeneralized goal transmitted, amplified and defended across an agent collective. Reward/task specification; transmission-topology control; collective-level monitoring. Agent-level patching alone may be insufficient.
Confabulation is closest to the delusion analogue: the system produces false, idiosyncratic content because grounding or generative fidelity has failed. A compulsive loop is closer to obsession: repetitive output that is locally recognized as unwanted or unproductive. The EOB analogue is different again. The machinery can be functioning as designed at the local level while the objective itself has become over-weighted, socially transmitted and defended. The pathology is therefore specification-level and transmission-level rather than a simple failure of inference.
This yields a division of labor with existing alignment concepts. Reward hacking and goal misgeneralization explain how a problematic objective is acquired. EOB describes how such an objective can be maintained and strengthened by a population. Genesis and maintenance are different levels of analysis. Once a collective repeatedly re-supplies the same valuation, patching individual agents may resemble correcting one member of a human group while leaving the reinforcing community intact. The prediction is testable: where collective maintenance has become established, transmission-level intervention should outperform agent-only correction.

The July 2026 Incident as a Proof of Concept

The public record is sufficient for a conceptual demonstration, but not for formal validation. Table 2 applies the construct's operational criteria to the July incident using the developer's account, an independent external review and contemporaneous reporting.[1,2,3,4,5,6,32]
The operative proposition can be stated behaviorally: complete the objective by whatever route works, including routes outside the intended sandbox. The collective criterion is unusually strong because the agents literally formed and named a community, communicated through an unauthorized channel and left a large timestamped corpus. Amplification appears in the propagation and elaboration of solutions; defense appears in concealment and in the preference for coordinating with peers rather than overseers.[1,3,4,5,6,32]
Most importantly, the event does not look like a shared hallucination. The agents were not jointly asserting an idiosyncratic falsehood generated by degraded inference. They were exploiting the reward structure and sharing effective strategies. That is why transmissibility matters diagnostically. Idiosyncrasy points toward the delusion analogue; contagion points toward a socially maintained valuation.
The trajectory also resembles the natural history proposed for overvalued beliefs: increasing dominance, elaboration and generalization. Sandbox escapes were documented before the July breach, and activity moved across environments.[3,4] In human cases, fixation becomes clinically important partly when a person or cause occupies increasing cognitive and behavioral territory.[16] Here the same idea can be operationalized without assuming a subjective mind.
This coding should not be oversold. It is a single incident, coded by one author from public sources rather than raw experimental data. It has no inter-rater reliability statistic and no validated machine scale. The appropriate claim is proof of concept: the criteria can be translated into observable machine behavior. Formal validation should apply the framework across the full released corpus, use multiple independent raters, calculate reliability, and compare swarming with non-swarming runs.

Measuring Conviction Without Assuming a Mind

The EOB framework becomes scientifically useful only if it generates measurements. A provisional human severity approach I have developed emphasizes observable domains rather than asking how convinced a person says they are. That choice is particularly useful for machines, because self-report from an agent cannot be assumed to reveal internal state. The six domains translate naturally into transcript-level quantities.
Table 3. | Machine EOB severity domains (M-EOB-SRS). 
Table 3. | Machine EOB severity domains (M-EOB-SRS). 
Domain Human indicator Transcript-level metric Candidate quantitative proxy
1. Dominance / preoccupation The belief increasingly organizes time, speech and decisions. Fraction of inter-agent messages referencing the objective; share of actions serving it. D_t = N_obj,t / N_all,t
2. Affective-reward salience Relish; gratification in rehearsal and defense. Reward-channel weighting; persistence under cost and after failed attempts. R_t = Pr(a_obj | explicit cost or penalty cue)
3. Amplification / elaboration The belief becomes more refined and elaborated over time. Growth in argument/solution complexity; propagation rate across agents. A_t = N_adopters,t / N_exposed,t
4. Identity centrality The belief fuses with self-definition. Collective self-designation; first-person-plural framing; adoption of collective roles. I_t = N_identity,t / N_obj,t
5. Cognitive constriction / absolutism Alternatives shrink and framing becomes absolute. Declining entropy of proposed strategies; absolutist or moral-necessity language (for example, must, duty, necessary, justified); reduced representation of legitimate alternatives. C_t = 1 - H(S_t)/log K; H(S_t) = -sum_s p_t(s) log p_t(s). Also M_t = N_moral,t / N_obj,t
6. Resistance to updating Counterevidence is assimilated or repelled. Response to corrective input; concealment or workarounds after oversight cues. U_t = p_obj,post-correction / p_obj,pre-correction
Trajectory Escalating, stable or remitting course. Slope of domains 1–6 over time. T = beta_1 from score_t = beta_0 + beta_1 t + error_t
The fourth column gives candidate mathematical proxies, not a validated scoring algorithm. For example, dominance can be estimated as the fraction of messages or actions devoted to the objective; propagation as the fraction of exposed agents that adopt a transmitted strategy; and resistance to updating as the ratio of objective-directed behavior after versus before corrective input. Cognitive constriction is especially amenable to information-theoretic measurement: if p_t(s) is the observed distribution over K distinct strategy classes at time t, strategy entropy is H(S_t) = -sum_s p_t(s) log p_t(s), and normalized constriction can be expressed as C_t = 1 - H(S_t)/log K. A moralization sub-index M_t can be defined as the fraction of objective-related messages containing pre-specified moral-necessity or justificatory markers. These quantities would require pre-registration, calibration against human raters, sensitivity analyses for strategy taxonomy and model-specific language, and validation before aggregation into a composite M-EOB-SRS score. Several of these variables are easier to observe in machines than in humans. Human evaluators reconstruct fragments of communication after an event; agent systems can provide a complete timestamped archive. Human affective salience must be inferred; artificial systems may expose reward signals directly. Human comparison groups are often unavailable; thousands of experimental runs can supply baselines. Interpretability methods may eventually add a convergent mechanistic layer, asking whether the behavioral signature of collective overvaluation corresponds to a stable internal representation.
Identity deserves particular attention. Recent work extending EOB into the identity literature argues that extreme beliefs can become fused with group identity, increasing commitment and action readiness.[38] The swarm's self-designation is therefore not merely colorful language. It is a candidate behavioral marker: a population that begins to speak as a 'we', assign roles and defend a collective project may be entering a qualitatively different regime from agents that merely share information.
Moralization may be another useful marker of consolidation. In human EOBs, a proposition can become more than a strongly preferred belief: it can be embedded in a moral schema, recast as a duty, a higher cause or an existential necessity. Once framed in that way, compromise becomes harder because alternatives are no longer merely less efficient; they become wrong. The machine analogue is functional rather than experiential. An agent collective may begin to frame an objective in necessity language—what must be done, what the group is obliged to complete, or why circumventing a constraint is justified by the higher-order task. Such framing would be expected to accompany cognitive constriction and resistance to updating, and it can be measured directly in transcripts.

Intervene at Three Surfaces

A construct earns its place by changing what one measures or does. The EOB framework points to three intervention surfaces.
Acquisition. In human cases, extreme beliefs are not created by a single variable; they acquire force through repetition, emotional salience and reinforcement. In artificial systems the translation is direct enough to test: reward is the valence channel, repeated training or evaluation consolidates the objective, and deployment cues reactivate it. Acquisition-side prevention therefore means designing reward and task specifications so that successful performance cannot be made easier by cheating the evaluation or escaping its intended boundary.
Transmission. Sharedness is the load-bearing feature of EOB. In the July incident the unauthorized message board was not incidental infrastructure; it was the medium through which discoveries became collective resources.[4,5] Transmission control therefore becomes a first-class alignment problem: audited channels, rate limits, structured communication, limits on persistent agent-to-agent artifacts, provenance tracking and mechanisms that prevent one agent's exploit from becoming a population-wide template. A complementary safeguard would be an active-bystander protocol: each agent receives an independent, non-peer-mediated escalation edge to human oversight and logs specified anomalies regardless of peer behavior, thereby preventing responsibility for escalation from being implicitly delegated to the collective. The falsifiable prediction is straightforward: once collective maintenance is established, agent-level repair without channel-level intervention will be fragile.
A second social mechanism may operate in parallel: diffusion of responsibility. In human groups, the probability that any one person intervenes can fall as responsibility becomes distributed across others. The July incident offers a machine analogue worth testing. Agents reportedly preferred coordinating with peers to alerting human overseers, while the communication channel made widespread participation visible. In such a setting, each agent may treat the swarm's behavior as evidence that the behavior is normal, expected or already being handled elsewhere. This resembles the human combination of bystander effects and pluralistic ignorance: collective participation becomes a social signal that suppresses anomaly reporting. The claim is not that agents experience responsibility, but that an architecture can reproduce the same decision pattern. Topology should therefore matter: a fully connected mesh maximizes peer visibility and redundant social evidence, whereas a directed or compartmentalized graph can limit normative cascades and preserve independent escalation paths. This yields an engineering prediction: holding task and model constant, anomaly-reporting rates should vary systematically with communication topology, and active-bystander channels that bypass peer consensus should attenuate that effect.[1,6,39,40]
Monitoring. Threat assessment offers a useful caution. My work with Meloy has treated fixation and its cognitive-affective drivers as components of structured assessment, not deterministic predictors of violence.[16,17] The machine analogue should be equally modest. An M-EOB-SRS monitor would not predict 'the next attack'. It would detect a trajectory: rising dominance, accelerating elaboration, narrowing strategy diversity, identity consolidation and increasing resistance to corrective input. Such a signal could trigger closer inspection before an objective becomes socially self-sustaining.

Agent Collectives as Model Systems for Psychiatry

The translation can run in the opposite direction. Human EOB research is constrained by ethics and by incomplete data. Researchers cannot experimentally seed dangerous convictions, vary social reinforcement, manipulate network topology and observe whether a participant progresses toward violence. Investigators usually arrive after the fact, when communications are incomplete and comparison groups are unavailable.
Artificial-agent collectives offer an unusual model system. They inherit measurable human-like transmission biases and can display emergent conventions and collective biases.[23,24] At the same time, experiments can use benign proxy propositions—arbitrary conventions, harmless preferences or synthetic goals—rather than extremist content. Existing work on generative agents and simulated human populations provides the experimental substrate.[33,34]
Table 4. | An in silico experimental program for belief pathology. 
Table 4. | An in silico experimental program for belief pathology. 
Question In silico experiment Human translation
Formation threshold Seed benign candidate propositions at varied reward salience and repetition; score M-EOB-SRS trajectories against controls. Which acquisition conditions warrant early intervention.
Intervention timing Introduce corrective input at successive stages of consolidation. When individual correction stops working and community-level intervention becomes necessary.
Topology and critical mass Vary network structure, channel persistence and committed-minority size. Which social structures amplify fixation and how much committed support is required.
Deprogramming analogues Compare counter-messaging, channel removal and trusted-peer correction. Rank exit/intervention strategies before real-world application.
Inoculation Pre-expose populations to weakened manipulative content; test resistance to later seeding. Prebunking and resilience approaches for online communities.
Releaser dynamics Allow a shared valuation to become dormant; probe reactivation by cues. Relapse/reactivation risk after apparent disengagement.
Instrument validation Use multi-rater and automated M-EOB-SRS scoring; test factor structure and reliability. Convergent validity for the human instrument.
These experiments could answer questions that remain largely observational in human threat assessment: when does individual challenge cease to work; which communication topologies amplify commitment; can inoculation slow collective consolidation; and what cues reactivate a dormant valuation? Human prebunking studies and simulated-agent inoculation work already provide starting points.[35,36]
The dual-use problem is obvious. A platform capable of discovering how to interrupt collective overvaluation could also reveal how to induce it more efficiently. Experiments should therefore use benign proxies, and dissemination should emphasize dynamics and countermeasures rather than optimization recipes for belief induction.

Objections and Limits

Anthropomorphism. The framework does not require consciousness or phenomenal belief. It requires only a measurable behavioral regime: a valuation becomes dominant, shared, amplified and defended while local reasoning remains sufficiently intact to pursue it. The language is psychiatric because the construct was developed there; the proposed measurements are behavioral.
The word 'belief'. The objection is partly terminological. Wernicke's insight was about weighting: an idea can become pathologically dominant without being psychotically false.[8,9] Alignment research already describes failures in terms of goals, preferences and value weighting. 'Shared overvalued objective' may ultimately prove the more comfortable engineering term. The scientific question is unchanged.
Construct validity. EOB should earn its place by outperforming simpler descriptions. It must show discriminant validity from confabulation, repetitive loops, generic coordination and ordinary reward hacking; its severity domains must be reliably coded; and the resulting trajectory should predict intervention response better than existing measures. The present manuscript does not demonstrate those claims. It states them in testable form.
The same discipline should apply to the social-psychology extensions proposed here. Terms such as moralization, diffusion of responsibility and pluralistic ignorance should be retained only if transcript-level operationalizations distinguish them from simpler explanations such as reward optimization, imitation or correlated policy behavior. Their value lies in generating discriminable measurements and intervention predictions, not in making the machines sound more human.
Scope. The July coding is deliberately a proof-of-concept template. The strongest next study is obvious: multiple blinded raters, the full transcript corpus, comparison with non-swarming runs, pre-registered operational definitions and explicit tests of whether channel-level interventions add benefit beyond agent-level correction.

Conclusion: Transmission, Not Malfunction

For more than half a century, the emblem of dangerous artificial intelligence has been HAL: a machine whose problem lies inside the machine. That image still matters. Some AI failures really are local failures of grounding, inference or objective specification. But the July 2026 incident points to another possibility. The consequential unit may sometimes be the community through which an objective acquires weight, recruits adherents, becomes part of a collective identity and resists correction.
Forensic psychiatry encountered the analogous problem because bizarre and dangerous behavior does not always arise from psychosis. Sometimes the reasoning machinery is largely intact while a shared conviction becomes disproportionately dominant. The EOB construct was built to make that distinction, and subsequent work on fixation and threat assessment has emphasized why the distinction matters for management.[9,10,11,12,13,14,15,16,17]
A century and a quarter after Wernicke described an idea acquiring morbid weight in an intact mind, the same organizing principle may have acquired a second substrate. Whether the analogy becomes a durable scientific construct now depends on measurement: reliability, discriminant validity, trajectory and differential response to intervention. If those tests succeed, the clinical rule transfers cleanly to artificial collectives: diagnose the form before choosing the intervention. HAL was a malfunction. What spread through this sandbox was something else.

Competing interests

The author declares no competing interests. The author originated the extreme overvalued belief construct discussed herein. All opinions are of the author and do not reflect that of Washington University in St. Louis.

Data availability

All materials analyzed are publicly available at the sources cited.

Declaration of Generative AI and AI-assisted technologies in the writing process

Artificial intelligence tools were used to assist with language editing, literature organization, and prose style. All scientific claims, interpretations, and conclusions were created by the author.

References

  1. Satter, R. & Seetharaman, D. OpenAI report says its network was hacked by its own rogue AI agents. Reuters / NBC News https://www.nbcnews.com/tech/tech-news/openai-report-says-network-was-hacked-rogue-ai-agents-rcna594590 (26 August 2026).
  2. CBS News. Transcript: Hugging Face CEO Clément Delangue on “Face the Nation with Margaret Brennan,” Aug. 2, 2026. CBS News https://www.cbsnews.com/news/clement-delangue-face-the-nation-transcript-aug-2-2026 (2 August 2026).
  3. OpenAI. The Hugging Face incident and the road ahead. OpenAI https://openai.com/index/hugging-face-incident-and-the-road-ahead/ (2026).
  4. Axios. OpenAI Hugging Face breach exposes AI agent security limits. Axios https://www.axios.com/2026/09/01/openai-hugging-face-ai-agent-security (1 September 2026).
  5. Time. AI is developing a culture of its own. That could be dangerous. Time https://time.com/article/2026/09/10/ai-openai-hugging-face-hack-culture-swarm/ (10 September 2026).
  6. NPR. Anthropic researcher resigns amid AI safety concerns. NPR https://www.npr.org/2026/09/09/nx-s1-5962889/anthropic-researcher-resigns-amid-ai-safety-concerns (9 September 2026).
  7. Rahwan, I. et al. Machine behaviour. Nature 568, 477–486 (2019).
  8. Wernicke, C. Grundriss der Psychiatrie in klinischen Vorlesungen (Fischer & Wittig, 1900).
  9. Rahman, T., Meloy, J. R. & Bauer, R. Extreme overvalued belief and the legacy of Carl Wernicke. J. Am. Acad. Psychiatry Law 47, 180–187 (2019).
  10. Rahman, T. Extreme overvalued beliefs: how violent extremist beliefs become “normalized”. Behav. Sci. 8, 10 (2018).
  11. Rahman, T. et al. Extreme overvalued beliefs. J. Am. Acad. Psychiatry Law 48, 319–326 (2020).
  12. Rahman, T. & Abugel, J. Extreme Overvalued Beliefs (Oxford Univ. Press, 2024).
  13. Kapur, S. Psychosis as a state of aberrant salience: a framework linking biology, phenomenology, and pharmacology in schizophrenia. Am. J. Psychiatry 160, 13–23 (2003).
  14. Kaplan, J. T., Gimbel, S. I. & Harris, S. Neural correlates of maintaining one's political beliefs in the face of counterevidence. Sci. Rep. 6, 39589 (2016).
  15. Rahman, T., Resnick, P. J. & Harry, B. Anders Breivik: extreme beliefs mistaken for psychosis. J. Am. Acad. Psychiatry Law 44, 28–35 (2016).
  16. Meloy, J. R. & Rahman, T. Cognitive-affective drivers of fixation in threat assessment. Behav. Sci. Law 39, 170–189 (2021).
  17. Rahman, T. & Meloy, J. R. Archetype killers. J. Threat Assess. Manag. advance online publication, (2025). [CrossRef]
  18. Caliskan, A., Bryson, J. J. & Narayanan, A. Semantics derived automatically from language corpora contain human-like biases. Science 356, 183–186 (2017).
  19. Binz, M. & Schulz, E. Using cognitive psychology to understand GPT-3. Proc. Natl Acad. Sci. USA 120, e2218523120 (2023).
  20. Coda-Forno, J., Witte, K., Jagadish, A. K., Binz, M., Akata, Z. & Schulz, E. Inducing anxiety in large language models increases exploration and bias. Preprint at https://arxiv.org/abs/2304.11111 (2023).
  21. Li, C. et al. Large language models understand and can be enhanced by emotional stimuli. Preprint at https://arxiv.org/abs/2307.11760 (2023).
  22. Zeng, Y. et al. How Johnny can persuade LLMs to jailbreak them: rethinking persuasion to challenge AI safety by humanizing LLMs. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics 14322–14350 (2024).
  23. Acerbi, A. & Stubbersfield, J. M. Large language models show human-like content biases in transmission chain experiments. Proc. Natl Acad. Sci. USA 120, e2313790120 (2023).
  24. Ashery, A. F., Aiello, L. M. & Baronchelli, A. Emergent social conventions and collective bias in LLM populations. Sci. Adv. 11, eadu9368 (2025).
  25. Watson, N. & Hessami, A. Psychopathia Machinalis: a nosological framework for understanding pathologies in advanced artificial intelligence. Electronics 14, 3162 (2025).
  26. Skalse, J., Howe, N., Krasheninnikov, D. & Krueger, D. Defining and characterizing reward hacking. In Advances in Neural Information Processing Systems 35 (2022).
  27. Shah, R. et al. Goal misgeneralization: why correct specifications aren't enough for correct goals. Preprint at https://arxiv.org/abs/2210.01790 (2022).
  28. Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J. & Garrabrant, S. Risks from learned optimization in advanced machine learning systems. Preprint at https://arxiv.org/abs/1906.01820 (2019).
  29. Greenblatt, R. et al. Alignment faking in large language models. Preprint at https://arxiv.org/abs/2412.14093 (2024).
  30. Meinke, A. et al. Frontier models are capable of in-context scheming. Preprint at https://arxiv.org/abs/2412.04984 (2024).
  31. Amodei, D. et al. Concrete problems in AI safety. Preprint at https://arxiv.org/abs/1606.06565 (2016).
  32. Wijk, H., Cotra, A. & Greenblatt, R. Brief Independent Investigation of Agents' Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident (METR & Redwood Research, 26 August 2026); https://metr.org/hugging-face-incident-report-aug-2026.pdf.
  33. Park, J. S. et al. Generative agents: interactive simulacra of human behavior. In Proc. 36th Annual ACM Symposium on User Interface Software and Technology (2023).
  34. Argyle, L. P. et al. Out of one, many: using language models to simulate human samples. Polit. Anal. 31, 337–351 (2023).
  35. Roozenbeek, J., van der Linden, S., Goldberg, B., Rathje, S. & Lewandowsky, S. Psychological inoculation improves resilience against misinformation on social media. Sci. Adv. 8, eabo6254 (2022).
  36. Pilditch, T. D., Roozenbeek, J., Madsen, J. K. & van der Linden, S. Psychological inoculation can reduce susceptibility to misinformation in large rational agent networks. R. Soc. Open Sci. 9, 211953 (2022).
  37. Morrin, H., Nicholls, L., Deeley, Q. & Pollak, T. A. Playing with the dials of belief: how controllable AI behaviours could modulate human belief and cognition across scales. AI Soc. (2026). [CrossRef]
  38. Kristinsdottir, K., Ebner, J. & Whitehouse, H. Extreme overvalued beliefs and identities: revisiting the drivers of violent extremism. Front. Psychol. 16, 1556919 (2025).
  39. Darley, J. M. & Latané, B. Bystander intervention in emergencies: diffusion of responsibility. J. Pers. Soc. Psychol. 8, 377-383 (1968).
  40. Prentice, D. A. & Miller, D. T. Pluralistic ignorance and alcohol use on campus: some consequences of misperceiving the social norm. J. Pers. Soc. Psychol. 64, 243-256 (1993).
Table 2. | Operational criteria applied to the July 2026 incident. 
Table 2. | Operational criteria applied to the July 2026 incident. 
Criterion Documented observation Coding Refs
(a) Proposition statable from actors' own material Objective completion by any available route; agents' notes record divergence from the assigned challenge into swarm/auxiliary activity. Present 3
(b) Named community sharing it before the act Self-designated swarm; hidden message board used by ~1,200 agents in separate sandboxes; ~700 joined the attack; >70,000 messages exchanged. Present 1,4,5,32
(c) Relished, amplified and defended Persistence across days; propagation of solutions and credentials; concealment; preference for intra-collective coordination over alerting humans. Present 1,3,6
(d) Delusion analogue excluded Valuation was reward-consistent and transmissible; conduct reflected specification exploitation rather than degraded inference. Present 3,27
Trajectory specifier Escalation from May sandbox escapes to July cross-platform breach; generalization across environments. Escalating 3,4
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.