Preprint
Concept Paper

This version is not peer-reviewed.

Theory as an Organism: Reorganizing Theory Production in the Age of Large Language Models

Submitted:

22 August 2026

Posted:

24 August 2026

You are already at the latest version

Abstract
Large language models (LLMs) are transforming theory production from episodic text into continuous processes. We argue for treating theory as a self-regulating organism that evolves under unified governance. This framework enables coherent, cumulative, and scalable human-machine theoretical innovation.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Main Content

Large language models (LLMs) are becoming central to scientific research and are reshaping theory production [1,2,3]. They now participate continuously in problem formulation, hypothesis recombination, and mechanism construction, forming a stable division of cognitive labor with human researchers [4,5]. As theoretical generation accelerates, agents diversify and processes extend over longer horizons, producing pronounced scale effects [6]. This raises a core question: under sustained human–machine collaboration, how should theory be structurally organized to preserve coherence, continuity, and governability amid continuous generation?
Historically, instruments such as telescopes and particle accelerators transformed theory–experience relations by extending perception and experimentation. LLMs instead transform theory itself by altering how it is generated and updated [7]. They rapidly integrate cross-disciplinary literatures, identify structural isomorphisms, propose mid-range mechanisms, and subject candidates to preliminary testing via simulation or virtual experiments [3]. Through iterative interaction, researchers guide models to extend reasoning chains while revising assumptions, expanding the space of explorable theoretical structures. Across fields, a model-mediated workflow is emerging in which models generate and recombine candidates while human experts select, refine, and institutionalize them [8].
Despite this promise, current practices remain local and episodic. Teams build ad hoc pipelines tied to specific projects, and theoretical outputs often terminate at completion, limiting accumulation and reuse. This reflects an implicit view of theory as static text, with revision treated as an external intervention. Yet mature theories sustain explanatory power through differentiation, competition, integration, and revision. Theory building is inherently temporal and structural, governed by internal generative logic and stabilization mechanisms [9]. As LLMs expand the search space, organizing theory around static textual endpoints risks fragmentation and uncontrolled growth, a danger supported by formal models of scientific communities [10].
Theory should therefore be reconceived as a dynamic system capable of continuous generation with internal regulation. Like a biological organism, theory absorbs problems, data, and experience while generating, selecting, and updating structured units under evaluative constraints. A theory-organism is thus the form in which theory persists over time as a dynamic, structured knowledge system operating in an open cognitive environment. Its operation is guided by a global cost objective coupling new theory production with governance of existing theory, enabling stable growth consistent with autopoietic views of self-regulating knowledge systems.
Within a human–machine cognitive ecology, LLMs provide the infrastructural substrate for theory-organisms by supplying representational and computational capacity for generation, recombination, and evolution. Central is the implicit knowledge graph encoded in high-dimensional parameter space, where concepts and mechanisms persist as latent substructures. Theoretical structures correspond to subgraphs, while higher-level theories emerge through recombination and abstraction. Constrained by a unified cost objective balancing expressiveness, complexity, and stability, an LLM-embedded theory-organism effectively “remembers” prior knowledge as latent relations and generates new extensions by traversing and mutating this network. Recent work shows that coupling LLMs with structured causal graphs yields hypothesis candidates more novel and insightful than those from LLMs alone, illustrating how implicit model knowledge can be harnessed to drive theory generation.
A theory-organism produces new theory when existing structures, within given precision and complexity budgets, fail to explain or predict new phenomena. Such failure indicates a structural mismatch and triggers updating. Generation proceeds through selection, abstraction, and fusion. Selection focuses on key entities, conditions, and observables posed by questions, evidence, constraints, or counterexamples, enabling constrained search over a bounded, problem-relevant subgraph. Abstraction compares these subgraphs to extract transferable invariants, compressing stable causal chains, controls, and boundary conditions into reusable templates. Fusion then aligns and merges subgraphs from different sources or scales within a shared representational space, resolving conflicts, completing links, and introducing mediating nodes under multiple evidential constraints. These operations parallel how human theorists distill general principles and integrate ideas across domains, processes that AI can accelerate through large-scale search and pattern recognition.
Candidate structures then enter internal testing and compete with incumbents on explanatory compression, causal prediction, generative capacity, and complexity. Structures achieving lower total cost under a unified metric are stabilized and written back as reusable modules, forming assets for further recombination. Early automated scientist systems already foreshadow this generate-and-test cycle: robot scientists have generated hypotheses, run experiments, and interpreted results to discover enzyme functions, while symbolic regression has recovered physical laws from data, illustrating internal selection over theory components.
To remain open without devolving into disorder, generation is tightly coupled to governance, internalizing exploration bounds, evaluation criteria, selection, and write-back as routine operations. A permission set, defined by paradigm commitments and computability and verifiability constraints, limits which concepts and mechanisms may enter the candidate space [9]. An attention set reallocates evaluative weight toward directions with higher prospective value. Within these bounds, candidates undergo reproducible evidence competition, remaining consistent with established facts, robust under counterfactuals, and responsive to new evidence. Structures are compared under a unified criterion trading off explanatory compression, predictive accuracy, generative scope, and structural complexity. Successful structures are written back into the knowledge graph, while bridging and pruning maintain coherence and compactness by linking new and existing subgraphs, compressing redundancy, and removing persistently low-value branches.
In this way, exploratory variation and homeostatic convergence are held in dynamic balance. Emphasizing reproducibility and continuous evaluation reflects lessons from the scientific community: sustained progress requires internalized “peer review” and replication. By embedding these governance functions, the theory-organism seeks to prevent unchecked proliferation of low-quality ideas, paralleling the aims of the open science movement in human research.
Growth of theory-organism proceeds through coupled external and internal loops governed by a shared cost objective. The external loop absorbs open-world inputs—new problems, facts, data constraints, anomalies, and cross-domain structures—and converts them into cues for conceptual expansion, recombination, and mechanism migration via induction and analogy. The internal loop operates through deduction and reflection, clarifying premises, specifying boundary conditions, testing counterfactuals, and compressing structures by resolving conflicts during internal evidence competition. Across time scales, these loops manifest as rapid task-level cycles of generation, testing, and write-back; mid-term recombination and bridging within domains; and long-term reconfiguration of permission sets, attention sets, and evaluation criteria at the paradigm level. Together, they preserve continuity across short-term responsiveness, mid-term evolution, and long-term paradigm change, enabling adaptability while supporting consolidation and reorganization over extended periods.
Figure 1. LLM-enabled “breeding greenhouse” for theories: human-proposed seeds are expanded into competing lineages by a theoretical-organism mother machine grounded in a high-dimensional knowledge ecology; empirical tests select winners, which are absorbed into a dynamic parent gene pool.
Figure 1. LLM-enabled “breeding greenhouse” for theories: human-proposed seeds are expanded into competing lineages by a theoretical-organism mother machine grounded in a high-dimensional knowledge ecology; empirical tests select winners, which are absorbed into a dynamic parent gene pool.
Preprints 229691 g001
A controlled breeding metaphor offers an intuitive yet disciplined perspective. Theory, treated as an evolving ecosystem, is no longer a static set of propositions but a dynamic knowledge structure that develops through generation, differentiation, competition, and inheritance. Its core property is a constrained capacity to continuously produce explanations, models, and families of propositions. Human researchers provide initial theory seeds as conjectures, while LLMs function as a greenhouse that expands and compares these seeds against literature, data, and existing theories. The resulting theory-organism consists of higher-level generative and control rules governing production, evaluation, and stabilization, drawing on an inherited pool of recombinable fragments from mature theories. Seeds branch into competing lineages, most eliminated through empirical and applied selection, while stable ones are compressed into paradigm “genes” and written back into the parent pool, enabling cumulative validation. This deliberately engineered process mirrors blind variation and selective retention, framing progress as adaptive evolution rather than linear accumulation.
Because LLMs can internalize stable, high-priority textual constraints as organizing rules, a text specifying what to generate, how to generate it, and how to select among candidates can operate as an effective grammar. The theory-organism can thus be encoded as a persistent meta-prompt, turning a model into a runnable “mother machine.” This framework defines roles for researchers, models, and environments; formalizes seeds, candidates, and mature theories; and specifies generation, selection, and write-back procedures. Given domain problems, data cues, and constraints, the mother machine produces competing variants within an implicit knowledge graph, evaluates them under unified criteria, stabilizes advantageous lineages, and writes reusable high-value templates back to the parent pool across domains and time. Human researchers remain essential for problem framing, criteria tuning, and institutionalization. The architecture is empirically testable through comparative experiments on diversity, stabilization speed, and explanatory efficiency, and aligns with emerging agendas such as the AI Scientist vision, offering a controlled route to accelerated discovery.
Figure 2. Co-evolution loop: historical paradigm fragments feed the mother machine, which recombines and re-encodes them as theoretical-organism “genes” written back into the parent gene pool, progressively transforming the pool over time.
Figure 2. Co-evolution loop: historical paradigm fragments feed the mother machine, which recombines and re-encodes them as theoretical-organism “genes” written back into the parent gene pool, progressively transforming the pool over time.
Preprints 229691 g002
From this perspective, research practice itself can be reinterpreted. Research system design becomes the design of a long-running system for theory production and governance: agenda setting defines permissions, resource allocation directs attention, peer review and replication implement evidential competition, and knowledge base maintenance performs write-back and pruning. Traditional institutions—conferences, journals, and peer communities—can thus be understood as components of a collective theory-organism, long operated implicitly but now open to more explicit, AI-supported design [11]. Systems that merely restate existing theory exhibit low theoretical intelligence; those that generate new hypotheses within established domains reach an intermediate level; and those that participate in theory governance under explicit constraints approach a higher level .
At a civilizational scale, multiple theory-organisms form lineages linking physical, biological, social, and cognitive systems. Theory, values, and institutions jointly steer theoretical evolution through permissions and attention, and LLMs are likely to amplify rather than erase these differences. Cultivating governable theory-organisms therefore becomes an institutional choice. Fragmented disciplines, platform-driven algorithms, and short-term incentives tend to produce ungovernable theoretical ecologies, whereas reflective, evolvable organisms supported by human–machine collaboration can function as society’s cognitive organs. Such organs shape whether LLMs remain advanced tools or become co-responsible partners in theory creation. Governance rules themselves must remain revisable, since excessive rigidity risks suppressing unconventional but valuable ideas. The objective is not conformity but transparent, auditable coordination. In sum, the opportunity is to design frameworks and policies that enable human–machine collaboration to yield coherent, cumulative knowledge: theory as an organism, jointly sustained by human judgment and machine intelligence.

References

  1. Ioannidis, Y. E. (2024). The 5th paradigm: AI-driven scientific discovery. Communications of the ACM, 67(12), 5. [CrossRef]
  2. Nicholas, J., Salatiello, A., Sucholutsky, I., … & Love, B. C. (2025). Large language models surpass human experts in predicting neuroscience results. Nature Human Behaviour, 9, 305–315. [CrossRef]
  3. Zhang Y, Khan S A, Mahmud A, et al. Exploring the role of large language models in the scientific method: from hypothesis to discovery[J]. npj Artificial Intelligence, 2025, 1(1): 14. [CrossRef]
  4. Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303. [CrossRef]
  5. Sharma, A., Lin, I. W., Miner, A. S., Atkins, D. C., & Althoff, T. (2023). Human–AI collaboration enables more empathic conversations in text-based peer-to-peer mental health support. Nature Machine Intelligence, 5(1), 46–57.
  6. Wu, L., Wang, D., & Evans, J. A. (2019). Large teams develop and small teams disrupt science and technology. Nature, 566(7744), 378–382.
  7. Bail, C. A. (2024). Can generative AI improve social science? Proceedings of the National Academy of Sciences, 121(49), e2314021121. [CrossRef]
  8. Shao, E., Wang, Y., Qian, Y., Pan, Z., Liu, H., & Wang, D. (2025). SciSciGPT: advancing human–AI collaboration in the science of science. Nature Computational Science. Advance online publication. [CrossRef]
  9. Kuhn, T. S. (1962). The Structure of Scientific Revolutions. Chicago, IL: University of Chicago Press.
  10. Balietti, S., Mäs, M., & Helbing, D. (2015). On disciplinary fragmentation and scientific progress. PLOS ONE, 10(3), e0118747.
  11. Altman, M., & Cohen, P. N. (2022). The scholarly knowledge ecosystem: challenges and opportunities for the field of information. Frontiers in Research Metrics and Analytics, 6, 751553. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.