Preprint
Review

This version is not peer-reviewed.

From Solo Control to Enterprise Scale Through Agentic AI: A Survey of One-Person Agentic Company (OPAC)

Submitted:

10 August 2026

Posted:

20 August 2026

You are already at the latest version

Abstract
The rapid development of Artificial Intelligence (AI), and especially Agentic AI, is reshaping how work is organized at both the individual and organizational levels. An emerging organizational form—the One-Person Agentic Company (OPAC)—allows a solo or a few founders to orchestrate specialized agent swarms to manage complex business functions through real-time analysis and adaptive coordination. Unlike the traditional one-person company, which offers strong individual control but limited scalability, OPAC extends one-person productivity by combining human leadership with a multi-agent workforce powered by large language models and multimodal models. Despite growing interest, OPAC has emerged as an important research direction, yet the field remains in its early stages. Existing research is fragmented across AI foundation, multi-agent systems, human–computer interaction, and management. This survey provides the first systematic review of research relevant to OPACs. We define OPAC through four organizational characteristics from both technical and management perspectives. We further compare OPACs with traditional companies in terms of organizational composition, cost structure, efficiency, development logic, authority, and risk. We then organize the literature into four thematic areas critical to the technical and managerial design of OPAC. Specifically, we review how OPACs are structured and operate, how agent-related resources are allocated and governed, how such organizations develop through agent and workflow updates, and how reliability, accountability, and human control can be maintained amid increasing agentic delegation. By connecting advances in LLM-based agents and multi-agent systems with management research on organizational design, resource allocation, adaptation, and governance, this survey clarifies the current research landscape of OPACs, highlights critical open challenges, and outlines future directions for scalable and trustworthy human-led agentic enterprises.
Keywords: 
;  ;  ;  

1. Introduction

Entrepreneurship has long been constrained by a fundamental trade-off between control and scale. Traditional one-person companies (OPCs) preserve strong individual control and low organizational complexity, but their productivity is tightly bounded by the founder’s time, skills, and cognitive capacity. In contrast, conventional companies can achieve greater scale through hiring, specialization, and division of labor, yet this expansion typically raises entry barriers, fixed costs, coordination frictions, and managerial overhead. For most solo entrepreneurs, therefore, the path from a small venture to an enterprise-level operation has historically required a shift from individual execution to labor-intensive organization building.
The rapid development of Artificial Intelligence (AI) has improved efficiency and capacity for production at both the individual and organizational levels [1,2]. Recently, the emergence of Agentic AI marks a fundamental shift in business architecture, catalyzing the rise of the One-Person Agentic Company (OPAC) [3]. By orchestrating specialized agent swarms, a solo founder can now manage complex enterprises through real-time analysis and adaptive coordination. This transition from labor-intensive entrepreneurship to automated, scalable systems effectively overcome traditional manpower limits, offering the potential to drastically reduce operational costs, optimize resource allocation, and ultimately reshape the global economy [4]. Differing from traditional OPC which is characterized for ultimate control but extremely limited productivity [5,6,7], OPAC is characterized for its human-led organization of agent systems to tackle tasks or businesses that are way out of one-human productivity. OPAC allows an individual to extend operations across multiple functional areas of a large enterprise, including product development, marketing, finance, customer service, and legal affairs. The scale of this efficiency gain would have been difficult to imagine before the era of AI agents. Within the content creator economy, for instance, the transition from a solo venture to a high-impact enterprise is traditionally bottlenecked by the exhaustive demands of research, content production, and audience interaction. OPAC resolves these constraints by introducing a Multi-Agent System (MAS) [8]—a network of autonomous, interacting specialized digital agents (conceptualized as virtual workers) powered by Multimodal Large Language Models (MLLMs) [9,10]. These agents—ranging from AI screenwriters and musicians to content creators, data analysts, and coding specialists—form a proactive digital workforce capable of autonomously sustaining the full OPAC lifecycle. By doing so, the OPAC unlocks new frontiers of entrepreneurial productivity, transforming individual vision into scalable enterprise.
Figure 1. Illustration of One-Personal Agentic Company (OPAC), which consists of single or few human founder and multiple agentic employees. Supported by advanced technologies of LLMs, agentic AI and human-AI interaction, OPAC has the potential to overcome structural limitations of traditional human-based companies.
Figure 1. Illustration of One-Personal Agentic Company (OPAC), which consists of single or few human founder and multiple agentic employees. Supported by advanced technologies of LLMs, agentic AI and human-AI interaction, OPAC has the potential to overcome structural limitations of traditional human-based companies.
Preprints 227744 g001
OPAC Definition with Characteristics: Although OPAC has attracted increasing attention in recent years, it remains largely conceptual and has not yet been clearly defined in the literature. Defining such a novel and interdisciplinary organizational form is challenging, as it involves both technological and managerial dimensions. From a management perspective, we define OPAC through its key organizational characteristics.
  • A flexible organizational structure composed of a small number of humans and a large number of AI agents. Organizational structure is central to a company’s capacity, as it shapes how resources, roles, and coordination mechanisms are arranged. From a organizational management perspective, continuous adjustment of organizational design is increasingly important for long-term survival [11], making organizational flexibility a critical company capability [12]. Fundamentally, an OPAC is composed of a small number of human operators and a variable number of AI agents. Because AI agents can be added, removed, reconfigured, or recombined more easily than human employees, an OPAC can adjust its organizational structure more rapidly. This flexibility enables it to respond more quickly to changing external environments and to expand or reshape business functions as needed [12,13].
  • Limited human resources but expanded agent-related resources. Resource management is critical to company performance [14,15]. In traditional companies, key resources typically include financial capital, physical assets, technological infrastructure, human labor, and office space. An OPAC, by contrast, operates with minimal human labor but depends heavily on agent-related and computational resources, including tokens, tools, memory, context windows, model access, data, and processing capacity. Whereas traditional companies require substantial investment in physical workspace and on-site infrastructure, OPACs rely primarily on cloud-based platforms for storage, model access, orchestration, and continuous operation. Their operational space is digital rather than physical. Managing these resources requires dedicated technical and organizational mechanisms, including allocation rules, access control, and cost optimization, to ensure efficient and reliable company operations.
  • Faster optimization and organizational iteration. Continuous development and adaptation are essential for a company’s long-term survival. Traditional companies improve organizational performance by training employees [16,17], redesigning incentives, or introducing internal competition to increase productivity [18]. However, these approaches are often slow, costly, and may generate negative psychological side effects [19], which can ultimately harm overall company performance. In traditional organizations, even small-scale structural or capability changes may require substantial time before they are implemented and take effect. By contrast, the agent-based composition of an OPAC makes organizational optimization faster and more modular. Its performance can be improved at multiple levels, including enhancing individual agent capabilities, multi-agent structures, planning and routing mechanisms, and human-agent interaction. The main challenge shifts from managing large-scale human organizational change to solving technical and design problems in agent configuration, evaluation, and coordination. Moreover, OPACs may iterate roles, workflows, and agent compositions with fewer social and emotional frictions due to human psychological cost [20].
  • Trustworthiness as a critical organizational challenge. In conventional enterprises, multiple organizational layers often separate automated recommendations from binding actions, creating checks, reviews, and accountability buffers. An OPAC substantially compresses these layers: a single owner may define agent roles, configure resource access, govern system evolution, and serve even as the only detector [21,22]. This concentration of control, combined with reliance on multi-agent collaboration, introduces multilayer trustworthiness risks [23]. The agentic system may drift from the founder’s intent, misinterpret constraints, propagate errors across agents, or execute harmful actions [24]. Therefore, a core organizational challenge for OPACs is whether a single or very small number of human operators can understand, constrain, and control the broader agentic system [25,26]. Trustworthiness is thus not only a technical reliability issue, but also a condition for robust organizational governance.
Comparison with Traditional Companies: As a novel business model, OPAC differs from traditional companies (TCs) across organizational composition, cost structure, efficiency, development logic, authority, and risk, as summarized in Table 1. Organizationally, TCs are composed primarily of human employees who may use AI as tools, whereas OPACs are led by human owners and operated through AI agents as functional collaborators. This makes OPACs more structurally flexible, since agents and agent teams can be added, reconfigured, or terminated without the organizational frictions associated with human employment. The two forms also differ in resource and cost allocation. While TCs often devote substantial resources to labor, OPACs shift a larger share of operating costs toward agent-related resources, including tokens, memory, model access, tools, and AI infrastructure. In terms of efficiency, OPACs can sustain continuous agent execution once goals and constraints are specified, whereas TCs must account for human working limits, staffing arrangements, and coordination frictions. Their development mechanisms also diverge. Traditional companies often rely on training, incentives, and internal competition, which may produce side effects such as conflict or reduced knowledge sharing. OPACs can iterate more directly through model upgrades, agent reconfiguration, workflow optimization, and improved human-agent interfaces. Finally, authority and risk are organized differently. Because the human owner remains legally and strategically accountable, OPACs tend toward centralized final decision-making, even when operational work is widely delegated. Their risks also shift from human-originated problems, such as agency issues or counterproductive behavior, toward AI-originated risks, including system attacks, misalignment, agent drift, and loss of control.
Contribution of this Survey: Although dedicated studies on OPACs remain limited, a substantial body of related work has emerged across adjacent fields. In AI research, recent advances in LLM-based agents and multi-agent systems [27] have explored how autonomous agents can be equipped with roles, planning capabilities, and interaction protocols to complete complex tasks collaboratively [28,29,30]. Related work on human-agent interaction and human interaction with agent systems further informs how a human founder can communicate with, supervise, and intervene in agent organizations [31,32]. There are also studies examining resource allocation inside agent systems [29] and the methods for updating and iterating multi-agent architectures over time [33,34]. These streams of research are highly relevant to OPACs because they address the technical and organizational conditions under which a single human can govern a scalable agent-based company. Meanwhile, increasing attention has been paid to the risks and trustworthiness of human-agent systems, offering important foundations for understanding how OPACs can maintain reliability, accountability, safety, and human control [35]. However, existing studies remain fragmented across AI, human-computer interaction, multi-agent systems, and management research. To date, neither the AI literature nor the business literature has provided a systematic review of OPAC-related research.
This paper aims to bridge this gap by systematically organizing existing research relevant to OPACs. We review the literature from four perspectives: OPAC organization and operation, OPAC resource management, OPAC development, and trustworthiness risks and safeguards in OPACs. The first perspective examines how agentic organizations are structured and operated, including organizational structure, role and department design, human-AI delegation and interaction, operational planning, and decision-making. The second focuses on the allocation and access of agent-related resources such as tokens, memory and tools. The third reviews how OPACs can be updated and improved over time through agent-level, system-level, and business-level optimization. The fourth analyzes the risks introduced by autonomous agent systems and the safeguards needed to maintain reliability, accountability, and human control. Through this framework, we connect technical advances in agent systems with management questions about organizational design, resource allocation, governance, and long-term development.
Figure 2. Organization of systematic review on OPAC, which covers four major parts: 1) basic organizational and operational designs; 2) resource management strategies; 3) multi-level OPAC development and 4) trustworthiness analysis.
Figure 2. Organization of systematic review on OPAC, which covers four major parts: 1) basic organizational and operational designs; 2) resource management strategies; 3) multi-level OPAC development and 4) trustworthiness analysis.
Preprints 227744 g002

2. OPAC Organization and Operation

This section examines how an OPAC is organized and operated. Unlike conventional firms, an OPAC is staffed primarily by agents under the direction of a single human operator, who retains ultimate responsibility for goals, risk, and external commitments. Organizational and operational design therefore jointly determine whether agents extend the operator’s capacity. We divide the discussion into five parts along two complementary layers. The first layer concerns organization, namely how the system is structurally arranged and functionally divided. It covers two dimensions, agent organization structure and agent role and department design. The second layer concerns operation and governance, namely how work is delegated, coordinated, and controlled in practice. It covers three dimensions, human-AI delegation and interaction, operational planning, and operational decision-making. These five interdependent dimensions define the blueprint according to which an agentic OPC aligns day-to-day agent activity with the operator’s strategic intent. The following sections review major design forms in each dimension, summarize representative instantiations in the literature, and discuss selection strategies under OPC-specific constraints such as limited human attention and the need for controllability and accountability.

2.1. Agent Organization Structure

Organizational structure concerns the macro-level architecture through which agents are arranged into a coordinated system. It specifies relatively stable patterns of connection, control, and collaboration among agents, rather than the detailed responsibilities of individual roles or the runtime procedures for planning and decision-making. In this sense, organizational structure provides the structural context within which agent roles, human delegation, operational planning, and decision-making mechanisms operate. For an OPAC, organizational structure is especially consequential. Since a single human operator retains ultimate accountability, the structure of the agent organization determines how agent capacity is deployed, how complexity is contained, and how much supervisory burden falls on the human. A well-designed structure enables scalable delegation by creating clear coordination paths, interpretable handoff points, and auditable intermediate artifacts. A poorly chosen structure, by contrast, may lead to coordination overhead, opaque agent autonomy, duplicated work, or constant micromanagement. Thus, the central question is not which structure is universally best, but how different structural patterns balance control, adaptability, and human oversight under OPAC constraints. Existing LLM-based multi-agent systems exhibit two broad forms of organizational structure: static structures, in which the main topology and coordination relations are predefined before execution or general operation, and dynamic structures, in which the effective organization is adapted, instantiated, or reconfigured during operation.
Static Structures define the main agent topology in advance. They are useful when the task domain is relatively stable, the coordination logic is known, and the human operator needs predictable control over the system. The most common static forms include centralized, hierarchical, and decentralized structures.
  • Centralized structures, a manager, planner, or orchestrator agent coordinates specialist agents by decomposing tasks, assigning work, monitoring progress, and aggregating results [36]. This structure is widely used in LLM-based MAS because it provides a clear control point and simplifies human supervision [37]. For OPACs, centralized structures are particularly attractive when the human owner needs a compact interface to supervise many agents through a single managerial layer.
  • Hierarchical structures extend centralized control into multiple levels. Instead of a single manager directly coordinating all agents, the system organizes agents into layered planning, execution, review, or compliance units. Such structures are useful for complex tasks that require decomposition across abstraction levels, such as strategic planning followed by operational execution and verification [8,38]. In OPC settings, hierarchy can reduce cognitive load by grouping lower-level agent activity under intermediate controllers or reviewers.
  • Decentralized structures distribute coordination across agents without a strong central controller. Agents interact through peer communication, local negotiation, shared context, or emergent coordination mechanisms [39,40]. These structures can be more adaptive and robust to local uncertainty, but they are harder for a single human operator to monitor. Therefore, in OPACs, decentralized structures are better suited to exploratory or low-risk subtasks unless bounded by explicit logging, evaluation, or escalation mechanisms.
Dynamic Structures focus less on a fixed topology and more on how the organization is instantiated, coordinated, or reconfigured during task execution [41]. They are useful when tasks are open-ended, agent requirements are uncertain, or collaboration patterns must change with runtime feedback. Major forms include pipeline structures, blackboard structures, and task adaptive.
  • Pipeline structures organize agents through staged task flow. Work moves through predefined or dynamically instantiated phases, and each stage produces artifacts that condition later stages. MetaGPT [42] and ChatDev [43] are representative examples in software development, where agents follow structured workflows such as requirement analysis, design, coding, testing, and documentation. Although many pipelines are predefined, they are treated here as dynamic structures because their organizational effect emerges through runtime handoffs, feedback loops, and artifact-based progression.
  • Blackboard structures coordinate agents through a shared information space. Agents read from and write to a common workspace, memory, or artifact pool, allowing them to contribute incrementally without following a strictly linear order [44,45]. This structure supports opportunistic problem solving and iterative refinement, making it useful when different agents contribute partial knowledge to a shared evolving state. For OPACs, blackboard structures can improve transparency if the shared workspace is auditable, but they require mechanisms to prevent inconsistency, stale information, or uncontrolled accumulation of intermediate outputs.
  • Task-adaptive structure allow the agent organization to emerge or adjust through local interactions, feedback, contribution patterns, or environmental signals [41,46,47,48,49]. Compared with fixed organizational configurations, task-adaptive systems offer greater adaptability, but also create higher governance risk. In OPAC contexts, they are valuable for exploration and innovation, but should be paired with monitoring, evaluation, and human escalation mechanisms when used for consequential business operations.
The summary of different agentic organizational structure is shown in Figure 3. Overall, static structures provide predictability, traceability, and easier human supervision, while dynamic structures provide adaptability, scalability, and better fit to uncertain tasks. For OPACs, the design challenge is to combine these properties: using static structures to preserve accountability and control, while introducing dynamic structures where task uncertainty or agent heterogeneity requires flexible coordination.

2.2. Agent Role and Department Design

Agent role and department design concerns the functional composition of an agentic organization: what kinds of agents and departments should exist, what responsibilities and capabilities they hold, and how their boundaries map onto business functions. Unlike organizational structure, which focuses on how agents are arranged and connected, role design focuses on the functional units that populate that structure. A role may specify an agent’s task objective, domain boundary, tool access, expected output, interaction responsibility, and review obligation [50,51]. In an OPAC, this design problem is especially important because the system must translate company functions such as research, product, engineering, marketing, sales, finance, operations, and customer support into executable agent roles or virtual departments. Existing LLM-based MAS studies have three main agent role design strategies, which are pre-defined roles, task-adaptive roles and pre-defined roles with adaptive allocation.
Pre-defined role design defines both the role schema and the main role–agent configuration before execution. Agents are assigned stable identities, responsibilities, and interaction expectations, so collaboration follows a relatively fixed functional division. CAMEL [52] shows how explicit role-playing can guide cooperation between communicative agents. ChatDev [43] models software development as a virtual company with roles such as CEO, CTO, programmer, tester, and reviewer, while MetaGPT [42] encodes Standardized Operating Procedures (SOPs) into workflows with roles such as product manager, architect, project manager, and engineer. Virtual Lab [53] and SiriuS [54] also applied pre-defined role structures for molecular design. These systems show that roles are not merely prompts, but mechanisms for allocating responsibility, structuring communication, and controlling output quality. For OPACs, pre-defined roles provide a basis for stable virtual departments, but may be brittle when tasks require unexpected expertise or temporary cross-functional specialization.
Task-adaptive role configuration dynamically determines which roles, agents, collaboration modes, or tool configurations are needed for a given task. Unlike pre-defined designs, the functional composition of the agent team is adjusted according to task requirements, context, cost, or feedback. AgentVerse [48] introduces expert recruitment, assembling agents for the current goal and adjusting the group based on evaluation feedback. MasRouter [27] formulates MAS routing as a joint problem of collaboration-mode selection, role allocation, and LLM routing, optimizing the overall agent configuration for both performance and cost. Related work also treats role design as a broader team optimization problem involving agent selection, tool configuration, and collaboration-mode design [41,55]. For OPACs, this suggests that departments need not be permanent human-like units; they can instead be task-specific bundles of agents, tools, workflows, and review mechanisms.
Pre-defined roles with dynamic allocation combines stable role schemas with adaptive role–agent matching. Here, the system assumes that certain roles already exist, but dynamically selects which agent or model should fill each role for the current task. Dynamic Role Assignment [56] illustrates this logic in debate-based MAS: before the main debate, candidate agents generate role-specific proposals, peers evaluate them with role- and question-specific criteria, and the best-suited agent is assigned to each debate role. ReSo [57] follows a related adaptive allocation logic through a Dynamic Agent Database, which combines static profiles, such as base model and role prompt, with dynamic profiles tracking reward, cost, and task experience. For OPACs, this hybrid approach preserves the interpretability of stable departmental roles while enabling capability-aware staffing at runtime.

2.3. Human-AI Delegation and Interaction

Human-AI delegation and interaction are central issues in the organizational design of OPACs. Unlike traditional firms, where most business activities are carried out by human employees under established managerial hierarchies, agentic OPACs rely heavily on non-human intelligent agents to perform operational, analytical, communicative, and coordinative tasks. This reliance allows a single human founder to expand organizational capacity beyond individual limits, but it also weakens direct human control over day-to-day operations. As more work is delegated to AI agents, the key managerial question shifts from simply how should work be performed? to which decisions and actions can be delegated, under what conditions, through what interaction mechanisms, and with what forms of oversight?
This section therefore treats delegation and interaction as two complementary governance problems. Delegation defines the boundaries of agent autonomy: what agents may do independently, what requires human approval, and where final authority remains with the human founder. Interaction defines the communicative and interface mechanisms through which human intent, feedback, correction, approval, and learning signals enter the agent organization. Together, they determine how an OPAC balances the benefits of AI-driven operational expansion with the need for human supervision, responsibility, and final accountability [58]. Without deliberate design, agents may over-automate routine work, under-inform the founder about critical decisions, or overwhelm a single operator with fragmented conversations across many agents and tools.
Figure 4. Representative designs for human oversight and human-agent interaction for OPACs. Human oversight encompasses the various modes through which humans are involved in supervising and guiding automated agentic systems. Human–AI interaction characterizes the specific functional roles of humans and the types of information they exchange with the agents.
Figure 4. Representative designs for human oversight and human-agent interaction for OPACs. Human oversight encompasses the various modes through which humans are involved in supervising and guiding automated agentic systems. Human–AI interaction characterizes the specific functional roles of humans and the types of information they exchange with the agents.
Preprints 227744 g004

2.3.1. Human Oversight

Human oversight modes provide a risk-sensitive framework for determining how humans should supervise delegated AI actions. It is particularly important for OPACs, where AI agents may perform operational tasks at scale but should not independently exercise unrestricted authority over financial, legal, strategic, or reputationally sensitive decisions. For low-risk and reversible tasks, human-on-the-loop monitoring may be sufficient, allowing agents to act while humans retain intervention rights. For high-risk, irreversible, legally significant, or externally binding actions, human-in-the-loop approval is necessary before execution. At the highest level, human-in-command ensures that the founder retains control over goals, permissions, system boundaries, and shutdown authority.
Human oversight research shows that keeping a human formally involved does not automatically guarantee effective control over AI systems [59,60]. Humans may misuse automation by over-relying on it, disuse it when trust is too low, or suffer from automation abuse when designers delegate functions to machines without considering human cognitive limits [61,62,63]. Recent work on LLM-based agents and multi-agent systems suggests that human oversight should be designed as a layered control architecture rather than a single approval mechanism. Human-in-the-loop [58,64] oversight refers to cases where human input is constitutive of the final decision or action; it is particularly suitable for high-risk, irreversible, legally binding, or externally visible actions. Magentic-UI [65] operationalizes this mode through co-planning, co-tasking, and action guards, requiring user approval before potentially irreversible agent actions. ResearStudio [66] similarly allows users to pause, edit, and resume deep research agent workflows, while ARIA [67] incorporates human guidance when agents identify knowledge gaps during test-time learning. Human-on-the-loop oversight, by contrast, allows agents to operate autonomously while humans monitor, intervene, or correct behavior when necessary. It is more flexible than human-in-the-loop, which attracts relatively more research attention in literature. For example, OrchVis [68] visualizes hierarchical multi-agent workflows for human supervision while MI9 [69] proposes runtime governance mechanisms including semantic telemetry, authorization monitoring, drift detection, and graduated containment. CowPilot [70] achieves autonomous web navigation while humans supervise the process, provide feedback, and intervene when necessary. SAGA [71] proposes a security architecture for agentic AI systems that maintains human oversight through access-control policies, authorization mechanisms, and auditable constraints on agent actions and interactions. Human-in-command represents the highest level of authority: humans retain control over goals, permission boundaries, shutdown rights, and final accountability. This mode is consistent with meaningful human control theory and regulatory discussions of human oversight, which emphasize traceability, intervention rights, and responsibility chains. For agentic AI-based one-person companies, these three modes can be mapped onto different risk levels: Human-in-the-loop for high-risk or irreversible actions, Human-on-the-loop for routine but monitorable operations, and HIC for strategic, legal, financial, and governance-level authority.

2.3.2. Human-Agent Interaction

While human oversight defines when and how the founder retains authority over delegated actions, human-agent interaction (HAI) concerns the concrete mechanisms through which the founder communicates with, collaborates with, and intervenes in the agent organization. In an OPAC, HAI must support more than one-shot prompting. The founder needs to specify goals and constraints, co-develop plans or artifacts, monitor progress at an appropriate level of abstraction, intervene when necessary, and provide feedback that improves future operation [31,32]. Because multiple agents may operate concurrently, HAI must also aggregate internal agent activity into founder-legible summaries, pending approvals, anomalies, and decision points rather than exposing every raw agent message.
HAI type characterizes the relational role of the human in human–agent interaction: whether the human primarily assigns work, oversees execution, co-builds solutions, or aligns joint activity across agents and tasks. Beyond this functional dimension, each type also implies a characteristic direction of information flow between the human and agent(s). There are four types widely used as an organizing lens in the literature. The first type is delegation, where the human specifies goals, constraints, or process boundaries and expects the agent to carry out substantial execution with limited ongoing co-reasoning [72,73,74]. Information is transmitted downwards from human to agents. The second type is supervision, where the human remains able to monitor, approve, reject, correct, or take over agent behavior while execution proceeds [70,75,76]. This type offers bidirectional information exchange but asymmetric: the agent continuously exposes intermediate states, actions, or proposals to the human, while the human sends evaluative feedback, approvals, corrections, or takeover commands back to the agent. The third type is cooperation, where human and agent contribute as partners to a shared outcome, with interleaved planning, action, and revision rather than one-sided command or monitoring [77,78]. Information flow in this type is reciprocal, iterative and more intensive. Both parties exchange partial plans, observations, revisions, and contributions throughout problem solving, rather than relying on a fixed command–execution or monitor–feedback sequence. The fourth type is coordination, where the interaction focuses on aligning intentions, roles, timing, and dependencies among the human and one or more agents (or among agents under human-specified objectives) [42,79,80]. Human broadcasts objectives and role assignments, agents exchange task-relevant states and dependencies with one another, and alignment information circulates across the system to synchronize collective activity.
HAI functions refer to the structured communicative roles that human–AI interaction must fulfill in agentic systems [31,32]. These functions extend beyond simple prompt–response exchanges and can be organized into four categories. First, HAI needs to support goal and constraint specification, which is treated as a distinct, upstream communicative need [32,80]. Second, in some scenarios, HAI needs to enable collaborative construction, i.e., iterative co-production of plans, analyses, drafts, or code through shared artifacts or a persistent process layer rather than a single command-response turn [31,66,81,82]. Third, HAI should afford intervention and control, for example, pausing, correcting, overriding or even taking over during the operation [83]. Also, effective interaction should support evaluation and organizational learning by capturing human edits, critiques, demonstrations or preference signals so that later runs can be improved [67,84,85]. In the specific context of OPAC, when multiple agents operate simultaneously, the founder cannot participate in every inter-agent message. HAI therefore requires a dedicated aggregation layer that transforms internal coordination into founder-legible communication. Founder-facing aggregation is an important mechanism to make a multi-agent company governable by one person. Systems such as OrchVis [68] expose high-level goals, assigned subtasks, dependencies, and conflicts rather than raw agent dialogues. The founder communicates with the organization of agents through summaries, not with each agent thread in full detail. Wang and Lu [31] argue that a shared process model should be the primary surface for sustained alignment. The founder reads workflow state, decision history, and pending operations, and communicates back through edits to that state rather than through isolated prompts. In practice, founders need high-signal messages for anomalies, threshold breaches, and pending approvals, while routine progress can be compressed.

2.4. Operational Planning

Operational planning refers to the process through which a multi-agent system transforms a collective goal into executable coordinated action [86]. For OPACs, operational planning is not merely abstract reasoning about future actions; it is the operational bridge between business intent and agent execution. It specifies what work should be done, which agent or agent team should do it, how the work should proceed, and when the system should replan or escalate control. From the perspective of MAS, operational planning is affected by two key technical processes, which are task management and agent routing.
Task Analysis is also known as work structuring, which is the beginning of planning. It addresses what sub-problems the overall task contains and how they relate. Existing LLM-based MAS and classical multi-agent planning studies suggest two broad structuring logics: upfront structuring and progressive restructuring. Upfront structuring constructs the work breakdown at the beginning of the planning process. Given a user goal, a central planner or meta-agent decomposes it into subtasks, stages, dependencies, or milestones before execution begins. In this logic, the structure is treated as the initial scaffold for downstream agent assignment and scheduling. HuggingGPT, for example, performs task planning before model selection and execution [87]. MetaGPT and ChatDev encode software work into relatively fixed SOP-like development stages [42,43]. AOP first generates a task decomposition and allocation plan that is later evaluated and refined [30]. Reso [57] generates a task graph for each query before the adaptive construction the agent system. Upfront structuring is suitable when the task has a recognizable workflow or when domain routines can be encoded as templates. Its advantage is clarity and coordination efficiency, but it may be brittle when execution reveals missing information or invalid assumptions. Progressive restructuring treats the initial work structure as provisional. Instead of assuming that the task graph is complete from the beginning, the system updates, expands, or repairs the structure during execution based on feedback, partial results, failures, or changing context. For example, TDAG dynamically decomposes tasks and generates task-specific agents as requirements evolve [88], while DynTaskMAS maintains and updates a dynamic task graph to support asynchronous and parallel multi-agent execution [89]. Recursive frameworks such as ROMA further show how non-atomic tasks can be expanded on demand into dependency-aware subtask graphs and aggregated bottom-up after execution [90]. Progressive restructuring is therefore better suited to long-horizon, uncertain, or open-ended tasks, where planning must remain coupled with execution. ReAcTree [91] proposes a hierarchical task-planning method that decomposes a complex goal into manageable subgoals within a dynamically constructed agent tree. HiPlan [92] introduces expert demonstrations into its hierarchical planning framework to provide guidance.
Agent routing concerns how to determine agent or teams to handle each unit of milestones or dependency for an given task. Existing studies suggest five major assignment logics. First, role- or SOP-based systems assign work according to predefined organizational roles, as in MetaGPT [42] and ChatDev [43]. Second, capability-based routing maps subtasks to agents or models using functional descriptions, skill profiles, or solvability estimates, as in HuggingGPT [87] and AOP [30]. Third, dynamic expert recruitment selects or adjusts the participating agent group according to task state, as in AgentVerse [48] and DyLAN [41]. Fourth, coalition- and optimization-based routing formulate assignment as a constrained allocation or optimization problem rather than a simple role-matching decision, represented by the SMART-LLM [93]. Finally, recent systems move toward feedback-driven and learned routing, where assignment decisions are adapted based on performance, cost, state, or failure signals, as in MasRouter [27], QA-Dragon [94] and CASTER [95]. AOP [30] also extends simple capability matching by adding reward-model evaluation and feedback-based re-planning, thereby turning centralized routing into an adaptive process. Overall, existing agent routing paradigms evolve from static role-based assignment to increasingly adaptive and optimization-driven mechanisms. For OPAC, it must dynamically configure agent teams across milestones, dependencies, and uncertainty states, which requires adaptive, feedback-aware and agile routing mechanisms to balance the specification, efficiency and robustness under real-world operational constraints.

2.5. Operational Decision-Making

After planning specifies what should be done and which agents should execute each part, operational decision-making determines how a multi-agent system makes choices during execution and in real-world business. It concerns the mechanisms through which agents select among competing proposals, accept or reject intermediate outputs, resolve disagreements, decide whether to continue, revise, or terminate a process, and determine when human intervention is required [96,97]. Unlike planning, which primarily structures future action, operational decision-making governs runtime control: it translates agent outputs, feedback signals, evaluations, and environmental observations into concrete decisions. In OPACs, this step is especially important because the system must determine which decisions are sufficiently reliable to adopt, which alternatives should be pursued, and which decisions should remain under the owner’s authority.
Among existing studies, there are the following four representative strategies for operational decision-making:
  • Consensus, voting, or aggregation-based decision making asks agents to produce independent or semi-independent outputs and then combines them through majority vote, ranking, aggregation, or final synthesis [98,99,100]. These strategies ensure straightforward implementation and facilitate debugging, as system behavior can be easily traced back to specific rules. This approach is particularly efficient for tasks with well-defined procedures and limited variability, such as consensus seeking and navigation. However, this strategy may suffer from a lack of adaptability, especially when consensus is difficult to reach, the whole decision-making system may collapse.
  • Centralized managerial decision-making assigns decision authority to a manager, controller, or orchestrator agent. This agent monitors the task state, selects the next action or agent, decides whether intermediate outputs are acceptable, and triggers replanning when progress stalls. This is currently one of the most practical and widely used designs in LLM-MAS, because it gives the system a clear control point. Representative work includes AutoGen [37], Magentic-One [65] and AOP [30]. As a special case, a human can take such a managerial role as well, resulting in a human escalation strategy. Either way, such a centralized managerial strategy fits OPACs well because a one-person company often needs a virtual manager to coordinate specialist agents and maintain operational coherence.
  • Deliberative or debate-based decision-making produces decisions not by a single super agent directly, but through the emergence of structured disagreement, argumentation, and refinement. Multi-agent debate allows several LLM instances to generate candidate answers, read others’ reasoning, critique them, and update their own responses over multiple rounds [48,101]. This strategy is useful when the task has ambiguity or requires judgment, such as strategic planning, product positioning, legal or ethical trade-offs, or evaluating alternative business decisions [102]. Its weakness is cost and instability: debate can improve robustness, but may also create unnecessary token overhead or convergence toward a wrong consensus if the judge or agents are biased.
  • Critic or evaluator-based decision-making separates proposal generation from quality judgment. One or more agents generate outputs, while another agent, reward model, detector, critic, reviewer, or evaluator decides whether the output should be accepted, revised, rejected, or sent back for replanning. ChatEval [103] is a clear example focused on decision-making: it constructs a multi-agent referee team to discuss and evaluate generated texts, improving alignment with human evaluation compared with single-agent scoring. This design is important for OPACs because many operational failures come not from lack of output, but from accepting low-quality output too early. Evaluator-based decision-making provides a mechanism for quality control.

3. OPAC Resource Management

Compared to traditional companies, the fundamental nature of OPAC resource management shifts from physical and labor budgets to computational and cognitive assets, which direct agentic workflow. Therefore, the key to resource management lies in the effective allocation of limited computational budgets (e.g., tokens and memory) and cognitive bottlenecks (e.g., human attention). Given these constraints of OPAC’s internal resources, the resource allocation strategy is fundamentally associated with efficiency concerns: allocation strategies transform constrained resources into efficient utilization, optimizing how computational and cognitive resources are deployed to maximize agentic throughput. Meanwhile, apart from the internal resources, the practical deployment of an OPAC often requires integrating third-party agents from external vendors to access specialized, domain-specific tools (e.g., proprietary financial APIs or cloud server). Therefore, resource access control acts as another critical management mechanism that bridges external resources with the OPAC workflow: unrestricted access to external APIs, proprietary databases, and third-party services incurs prohibitive financial costs and latency bottlenecks. To provide a comprehensive understanding of how OPACs manage their evolving resource landscape, this section is systematically organized into three progressive subsections. First, Section 3.1 illustrates the fundamental computational and cognitive assets unique to the OPAC paradigm, including tokens, memory, and human labor. As shown in Figure 5, it establishes the dependencies and lifecycle of these core resources, along with their management framework. Second, Section 3.2 investigates the efficient allocation strategies for these identified resources. It systematically reviews mechanisms for enhancing token, memory, and labor efficiency, illustrating how OPACs can maximize agentic throughput while avoiding cognitive and computational overhead. Third, Section 3.3 focuses on access control for external resources, examining how OPACs establish tool calling, cross-agent memory sharing, and human-agent interaction (HAI) frameworks to effectively manage the agentic workflow across internal and external resources.

3.1. Fundamental OPAC Resources

In OPAC businesses, fundamental resources shift from traditional physical and labor budgets to computational and cognitive assets that directly steer agentic workflows. The scalability and sustainability of an OPAC are fundamentally determined by the efficient allocation of three core resources:
  • Token. Tokens represent the most fundamental units processed by large language model (LLM) based agents during both input comprehension and output generation. Within the OPAC workflow, tokens serve as the most important computational currency, determining how much contextual knowledge an agent can retain, how deeply it can reason, and how effectively it can coordinate within a multi-agent topology. From a resource management perspective, token consumption is directly tied to operational costs, execution latency, and system scalability. Therefore, unlike traditional enterprises where labor costs and physical infrastructure dominate operational budgets, the OPAC shifts financial and structural dependencies toward model access, GPU capacity, and token consumption. In practice, OPACs deploy heterogeneous agent teams across product development, marketing, finance, and compliance, where token usage scales non-linearly with task complexity, making it the primary bottleneck in sustaining long-horizon business operations.
  • Memory. Unlike traditional companies, where the software system is explicitly managed by human operators, OPAC memory is intrinsically distributed across agentic employees and their organizational topology. In the OPAC workflow, agents are organized into specialized roles with divergent domain boundaries, requiring memory to retain cross-task context, historical decision traces, and evolving preferences. However, the naive accumulation of interaction histories in business activities inevitably triggers high storage budgets and retrieval latency. Therefore, OPAC’s memory resources are not merely about data preservation but about optimizing crucial historical dependencies against computational overhead. In other words, memory serves as the institutional knowledge base that bridges individual agent execution with OPAC’s business purposes.
  • Human Labor. Under the “few human” constraint inherent in OPAC, the founder’s limited attention and decision-making capacity become one of the most critical resource bottlenecks that determine the scale of agent deployment. Specifically, human labor in the OPAC workflow is consumed through two distinct mechanisms: human inputs, which involve the strategic injection of expertise, contextual feedback, and corrective interventions into agentic workflows; and human oversight, which involves monitoring agent behaviors, preventing cognitive overload through adaptive alert filtering, and authorizing high-stakes or irreversible decisions. Strategically allocating this limited cognitive resource to ambiguous decision nodes and critical checkpoints is essential for maintaining the objective alignment between human founders and autonomous agents in open-ended business environments.

3.2. Efficient Allocation for Internal Resources

The architectural shift from traditional labor-intensive firms to a human-led, AI-agent-driven paradigm fundamentally redefines the resource dependencies in OPAC businesses, requiring the efficient management of token consumption, memory space, and human attention. In particular, OPAC’s resource allocation is fundamentally associated with efficiency concerns: optimized allocation transforms constrained resources into efficient utilization, determining how computational and cognitive assets are deployed to maximize agentic throughput. Therefore, achieving allocation efficiency goes beyond mere cost reduction to enable the scalability of autonomous workflows in dynamic market conditions. As shown in Table 2, we systematically investigate how optimized allocation strategies for tokens, memory, and human labor transform these constrained resources into efficient utilization in the OPAC workflow.

3.2.1. Token Efficiency

Efficient token allocation necessitates advanced strategies to compress context, encode latent representations, and structure reasoning paths, ensuring that every allocated token contributes maximally to decision-making. To address the non-linear scaling of token consumption in complex business operations, recent advancements have introduced three primary allocation paradigms that enable efficient utilization of constrained token budgets: explicit context pruning, latent generation, and tree/chain-of-thought structuring.
Context Pruning. Recent context pruning methods address this need across different modalities. For example, SWE-Pruner [104] introduces a self-adaptive, task-aware line-level pruning mechanism for coding agents, dynamically filtering irrelevant code segments based on an explicit goal. Similarly, UTPTrack [108] compresses visual tracking inputs, leveraging token-type awareness and spatial priors, which jointly compress tokens from the search region, the dynamic template, and the static template. To convert multi-turn text histories into compact rendered images, AgentOCR [106] proposes segment-level caching and adaptive compression to reduce token consumption while preserving decision quality. These methods show that removing low-value tokens is an effective approach to managing context growth. However, as the OPAC business becomes more complex, simply discarding observable tokens reaches its limits when handling deep reasoning traces or high-dimensional representations. To address these limitations, latent generation methods encode context into continuous embeddings to reduce complex token dependencies over long contexts.
Latent Generation. Beyond explicit context pruning, latent generation offers a paradigm shift by encoding discrete texts into continuous embedding spaces, thereby bypassing the token overhead of complex textual dependencies. For example, CODI [113] compresses the explicit reasoning into a continuous space via self-distillation, aligning a teacher model’s hidden states with a student’s implicit thought tokens. Notably, SwiReasoning [112] introduces a training-free mechanism that dynamically switches between latent exploration and explicit convergence based on entropy trends. By capping the number of mode transitions, it actively suppresses overthinking and allocates token budgets adaptively across varying problem complexities. Meanwhile, KaVa [111] addresses the supervision gap in continuous reasoning by distilling knowledge from a compressed KV-cache directly into latent thought tokens. While these latent strategies excel at compressing textual context into parameter-efficient embeddings, they inherently lack explicit structural information, such as multi-step reasoning paths. This limits OPAC’s capabilities when navigating complex decision-making, where the tree/chain-of-thought paradigm offers a potential solution.
Tree/Chain-of-Thought. OPAC typically involves complex business decision-making and relies on structured, multi-step reasoning, where the tree/chain-of-thought paradigm provides a highly effective mechanism for token budget optimization. Unlike traditional discrete token generation, recent frameworks leverage hierarchical merging and continuous latent spaces to reduce token consumption while preserving task-critical information. For example, FlashVID [121] implements a training-free, tree-based spatiotemporal token merging mechanism that dynamically consolidates redundant visual features. In parallel, continuous-space methods such as SoftCoT [120] and SemCoT [116] replace explicit reasoning chains with implicit and soft tokens. More specifically, SoftCoT employs a lightweight auxiliary model to generate instance-specific soft thought embeddings, which are linearly projected into the backbone LLM’s representation space. SemCoT further refines this strategy by integrating a customized sentence transformer to enforce strict semantic alignment between compressed implicit tokens and ground-truth reasoning. By adopting tree/chain-structured merging and continuous reasoning representations, OPAC founders can dynamically allocate token resources across heterogeneous agent tasks, scaling complex multimodal analysis and multi-step decision-making, while maintaining computational budgets.

3.2.2. Memory Efficiency

The efficient allocation of memory resources is critical for sustaining long-horizon workflows. In an OPAC, the naive accumulation of interaction histories in business activities creates severe bottlenecks. Recent advances in memory efficiency have evolved from rigid storage allocation to adaptive, compressed paradigms. In the following, we systematically review long-term memory architectures for adaptive retrieval allocation and latent memory mechanisms for continuous state compression.
Long-term Memory. Naive accumulation of interaction histories inevitably triggers high storage budgets, rendering memory efficiency a critical bottleneck in the OPAC’s resource allocation. Recent advances in long-term memory efficiency demonstrate that rigid, linear storage is inherently unsustainable, prompting a shift toward adaptive, multi-stage, and relevance-driven architectures. For instance, LightMem [124] decouples sensory filtering, topic-aware short-term consolidation, and offline sleep-time long-term updates, effectively separating heavy memory maintenance from real-time inference. MemFlow [125] introduces a retrieval mechanism that uses sparse memory activation, ensuring that only the most semantically aligned historical cues, rather than redundant interactions, are loaded into the working context. MemOCR [123] shifts from rigid textual sequence to a visual memory paradigm, achieving adaptive information density by rendering crucial evidence with high visual salience while aggressively downsampling auxiliary details, thereby maintaining reasoning robustness under extreme budgets. While these structured memory frameworks significantly alleviate context overflow and retrieval latency, they still rely on sparse, discrete tokens. To handle high-dimensional and cross-modal dependencies, recent research has explored latent memory mechanisms that compress explicit historical states into dense, parameter-efficient embeddings.
Latent Memory. To address the intrinsic bottlenecks of discrete token accumulation, recent advances in latent memory paradigms demonstrate remarkable efficiency gains across different modalities. For instance, MSA [131] introduces an end-to-end trainable sparse attention architecture with positional encoding, compressing document-level contexts into latent KV caches. Similarly, in streaming multi-modal contexts, HERMES [130] converts the KV cache into a hierarchical latent memory system, leveraging cross-layer smoothing and dynamic position re-indexing to eliminate redundant visual tokens. For agentic workflows, SimpleMem [129] compresses conversational agent interactions by employing semantic density gating and semantic synthesis, distilling verbose multi-turn dialogues into compact, context-independent memory units. These latent approaches fundamentally transform memory from a rigid storage bottleneck into a dynamic, computationally optimized resource. For OPAC, latent memory is pivotal in overcoming context window constraints and mitigating the storage cost of escalating discrete tokens.

3.2.3. Labor Efficiency

Under the “few human” constraint inherent to OPAC, the founder’s cognitive bandwidth represents the most critical bottleneck. Rather than distributing attention across routine tasks, strategic allocation concentrates limited human oversight on high-stakes decision nodes. Since autonomous agents can generate massive amounts of output that easily overwhelm human operators, labor efficiency aims to offload routine execution to agents while deploying human cognitive resources precisely where they yield the highest strategic value. In particular, OPAC’s labor efficiency is achieved by optimizing two fundamental allocation mechanisms: proactive human inputs for strategic intervention, and adaptive human oversight for objective alignment.
Human Inputs. Human inputs refer to the strategic injection of human expertise, contextual feedback, and corrective interventions into agentic workflows. In the OPAC paradigm, human labor shifts from routine execution to the strategic allocation of decision-making capacity, where limited founders’ attention acts as the primary bottleneck. Therefore, efficient human inputs should be allocated precisely to high-stakes, ambiguous, or dynamically shifting decision nodes where agents face inherent limitations. For instance, ARIA [67] introduces a mechanism where agents assess their own uncertainty via structured self-dialogue and proactively targeted human inputs. This ensures that human cognitive resources are expended exclusively on ambiguous tasks. Similarly, CowPilot [70] redefines human inputs as a dynamic, interleaved collaboration. By enabling users to pause, manually correct, and resume agent trajectories, it transforms human input from a static initial prompt into a continuous corrective signal that rescues agents from compounding errors. To seamlessly integrate human input into agent behaviors, AXIS [135] pioneers an API-based paradigm in which humans provide high-level intents and the agent autonomously translates them via API calls.
Human Oversight. As OPAC scales with business complexity, human oversight serves as a critical intervention mechanism for managing diverse agent behaviors and ensuring alignment with business objectives. Recent advancements in human-in-the-loop systems provide foundational architectures for implementing such oversight. In particular, Magentic-UI [65] introduces interactive oversight mechanisms that require explicit human approval for irreversible or high-stakes actions, thereby keeping the human in the loop during critical execution phases. Complementing this interactive approach, SAGA [71] enables users to define strict access policies using cryptographic tokens to prevent unauthorized access to resources. Focusing on a runtime oversight, MI9 [69] proposes a graduated containment mechanism that dynamically routes human-in-the-loop checkpoints based on real-time behavioral drift and context shifts, ensuring precise human intervention when agents exhibit goal drift or attempt to bypass policy constraints during multi-step workflows.

3.3. Access Control for External Resources

Beyond internal resources, the practical deployment of an OPAC often requires external vendors to access specialized, domain-specific resources (e.g., proprietary financial APIs, cloud server, and human expertise) that an OPAC cannot develop on its own. This architectural shift from internal agent workflow to a tool-augmented, memory-shared, and human-in-the-loop implementation fundamentally redefines resource management. While these external resources empower agents to perform complex business functions, their unrestricted access inevitably causes prohibitive API costs, latency bottlenecks, and vulnerabilities. Therefore, resource access control emerges as a pivotal management mechanism. Rather than relying solely on static, predefined rules, an OPAC demands adaptive, fine-grained coordination to balance resource utilization with operational constraints. As shown in Table 3, we systematically examine access control strategies across three dimensions: tool calling, memory sharing, and human-agent interaction (HAI) frameworks.

3.3.1. Tool Calling

Tool calling enables OPAC agents to interact with external environments and execute domain-specific business actions. However, the naive invocation of expansive API inventories often leads to excessive token consumption and severe latency. Therefore, recent advances in tool calling have marked a paradigm shift from naive external retrieval to structured, reusable protocols. Specifically, this evolution can be traced through three progressive mechanisms: API retrieval, which focuses on precise intent-to-API matching; the Model Context Protocol (MCP), which standardizes and governs diverse tool integrations; and skill learning, which abstracts repetitive tool interactions into optimized, agent-native competencies.
API Retrieval. API serves as the fundamental gateway that enables agents to dynamically select and use external tools. The API retrieval generated by agents is necessarily constrained by the OPAC’s resource access control, preventing unrestricted access to expansive API inventories and excessive token consumption. Recent research has systematically investigated the management of API retrieval, enabling precise matching between agent intents and available external APIs. For example, Gorilla [141] introduces a retriever-aware training approach that dynamically adapts to frequently evolving API documentation while mitigating argument hallucination through structured evaluation and semantic parsing. To address inconsistency and redundancy in raw tool documentation, EASYTOOL [138] transforms verbose, heterogeneous API references into concise, unified tool instructions that align with LLMs’ instruction-following capabilities. Meanwhile, ToolGen [136] integrates tool knowledge directly into model parameters via a three-stage training pipeline, eliminating the latency and redundant computation incurred by external API retrieval. While these approaches demonstrate substantial gains in retrieval precision and token efficiency, the differences in external APIs predominantly require model fine-tuning, custom tokenization, or specialized prompt templates. However, OPAC typically operates with heterogeneous agents, underscoring the need for a unified protocol to enable seamless adaptation between heterogeneous agents and external tools.
Model Context Protocol (MCP). The Model Context Protocol (MCP) has emerged as a foundational architecture for unifying tool access, shifting tool integration from external, diverse APIs to a standardized interface. Recent advances demonstrate how MCP can enforce fine-grained access control to alleviate API latency and budgets. For instance, MCP-Zero [142] introduces an active, on-demand tool discovery approach that generates structured tool requests and iteratively constructs cross-domain tool chains through hierarchical semantic routing. To address the scalability of real-world tool access, MCP-Flow [143] equips agents with the structured knowledge needed to expand heterogeneous external resources by systematically collecting, deduplicating, and filtering thousands of diverse MCP servers into high-quality instruction-function call pairs. Notably, HumanTool [144] has extended MCP to human oversight, integrating humans as callable, MCP-style interfaces within AI-led workflows. This allows agents to dynamically request human input, leveraging human authority as a gated resource in access controls. Nevertheless, repeatedly resolving tool schemas and low-level MCP calls remains computationally intensive for long-horizon OPAC workflows. More recently, agent skills have explored transforming repeated tool interactions into reusable, agent-native competencies.
Skill Learning. Beyond static API retrieval and standardized protocol integration, skill learning transforms repeated tool interactions into reusable, agent-native competencies, directly optimizing OPAC’s resource access control. Typically, an agentic skill is a reusable, callable procedural module that encapsulates workflow logic, structurally comprising a routing description for intent-based matching and a skill body containing executable policies, step-by-step instructions, and reference assets Recent advances demonstrate how skill learning abstracts repetitive tool interactions into optimized, reusable modules that inherently regulate access and consumption. For example, SkillReducer [148] introduces a two-stage de-bloating framework that compresses routing descriptions and skill bodies through delta debugging and taxonomy-driven progressive disclosure, improving agent performance via reducing context-window distraction. Similarly, SkillCraft [147] reduces token consumption through test-time skill composition and reuse, enabling agents to automatically compose tool calls into cached, executable skills. From a systemic perspective, SoK [146] maps the skill lifecycle to system-level design patterns, including metadata-driven disclosure, code-as-skill, and self-evolving libraries. For OPAC, skill learning functions as a strategic resource access control mechanism that shifts tool calling from heterogeneous API calls or MCP routing to an on-demand, skill-gated architecture

3.3.2. Memory Sharing

In the multi-agent OPAC topology, memory sharing is essential for cross-task consistency and collaborative reasoning. However, the unconstrained propagation of contextual memory across agents introduces significant challenges, including context overflow, redundant computation, and potential privacy issues. To manage the shared memory resource across agents, existing research has diverged into two primary paradigms: policy-based control, which enforces deterministic access rules and permission structures; and agent-based control, which treats memory sharing as a dynamic, learning-driven coordination problem.
Policy-based Control. Policy-based mechanisms have emerged as a foundational strategy for regulating memory sharing across multi-agent systems. These approaches rely on explicit rules, permission structures, or threshold-driven constraints to determine what information is stored, retrieved, or propagated among agents and users. For instance, Collaborative Memory [150] formalizes multi-user memory sharing through fine-grained read and write policies conditioned on dynamic bipartite access graphs, ensuring that cross-user knowledge transfer strictly adheres to the time-evolving access graph. Similarly, MemIndex [153] introduces an intent-indexed bipartite graph architecture where memory operations are conditioned on user intent urgency, agent workload, and historical performance metrics, ensuring that agents share and retain context only when resource constraints and user intent urgency align. To unify the management of long-term and short-term memory, AgeMem [152] proposes memory control through explicit tool-based read/write policies, which structure the agent’s learning process around policy-aligned reward functions that penalize context overflow and reward high-quality memory maintenance. In the OPAC paradigm, where memory constitutes a constrained computational resource, policy-based control establishes a deterministic baseline for access management and cross-agent consistency. However, purely policy-driven approaches often struggle to adapt to emergent cross-agent dependencies or dynamically shifting task priorities, as they rely on predefined rules.
Agent-based Control. Instead of static policy, agent-based control treats memory sharing as a dynamic, learning-driven coordination problem. Recent advances demonstrate how autonomous agents can continuously determine what to store, share, and retrieve based on contextual relevance, inter-agent dependencies, and downstream task utility. For instance, CoMAM [29] models heterogeneous memory pipelines as sequential Markov decision processes, employing end-to-end reinforcement learning with adaptive credit assignment to co-adapt memory construction and retrieval agents. Similarly, MemSkill [155] utilizes reinforcement learning for skill selection, treating memory operations as an evolvable skill bank, where a lightweight controller selects context-relevant memory routines and a designer iteratively refines the skill set from hard failure cases. In parallel execution settings, Learning to Share (LTS) [154] paradigm introduces a learned memory admission controller that selectively admits intermediate reasoning steps into a global shared memory bank, significantly reducing redundant computation and context over-length. By delegating memory access decisions to self-optimizing agents, agent-based control preserves consistency without intensive human oversight, enabling OPAC’s requirement for flexible, scalable, and cost-aware coordination across heterogeneous agents.

3.3.3. HAI Framework

OPAC founders act not merely as passive overseers but as critical, high-level access control and the ultimate authority in agentic workflows. To seamlessly bridge human operator and agentic workflow, the human-agent interaction (HAI) framework ensures that human cognitive bandwidth is allocated exclusively to high-stakes interventions, without being overwhelmed by low-level agent operations. Recent studies have explored HAI from two distinct perspectives: the human-centered paradigm, which prioritizes structured evaluation and explicit intervention checkpoints; and the shared workspace paradigm, which establishes a unified, context-rich environment for seamless, bidirectional co-planning and co-execution between humans and agents.
Human-centered. The Human-centered paradigm of the HAI framework addresses how OPAC founders can effectively oversee and intervene in multi-agent workflows without being overwhelmed by escalating system complexity. Recent work has highlighted the gap between agent workflows and the cognitive demands placed on human evaluators. Chen et al. [160] identify that the increasing complexity of LLM-powered GUI agents and their backend data transmissions place significant demands on human evaluators’ models. Consequently, they propose a human-centered evaluation framework that integrates risk assessments and in-context consent mechanisms, ensuring that human evaluators can effectively audit agent actions and manage privacy risks without being overwhelmed by system complexity. Similarly, for compound multi-agent systems, VeriLA [159] introduces a human-centered framework designed to verify and interpret agent failures. By designing human-defined agent criteria and training human-aligned verifiers, VeriLA enables human operators to efficiently audit reasoning processes, pinpoint the root causes of task failures, and provide actionable feedback. However, human-centered frameworks inherently rely on step-by-step oversight through intervention checkpoints, which fall short of addressing the seamless bidirectional communication and collaborative planning within the OPAC’s multi-agent topology. As a result, the design of a shared workspace has attracted increasing attention, in which human operators and AI agents can co-evolve in a deeply integrated manner.
Shared Workspace. Unlike traditional isolated interfaces, a shared workspace provides a unified, context-rich environment, where human founders and AI agents co-evolve. Recent advancements in human-agent collaboration highlight the efficacy of shared workspaces in optimizing human decisions and controlling agent access. For instance, Cocoa [81] introduces an interactive system that facilitates interleaved co-planning and co-execution within a document environment. By utilizing interactive plans as a shared representation, Cocoa enables users to seamlessly assign specific workflow steps to agents or themselves. Notably, Lu et al. [161] anticipates user needs by monitoring environmental events and user activities, proactively proposing tasks without explicit instructions. This paradigm significantly minimizes the cognitive burden on human operators. Furthermore, to mitigate risks in untrustworthy scenarios, VeriOS [162] proposes a query-driven human-agent GUI interaction framework. Instead of blindly executing or simply interrupting based on confidence thresholds, VeriOS proactively queries humans for clarification during anomalous or sensitive operations, leveraging the query-answer history to ensure trustworthy task completion.

4. Operational Capability Development in OPACs

In OPAC, development refers to the accumulation and refinement of operational capability through repeated work. Its role is analogous to human-resource development and strategy refinement in traditional companies: employees become more capable, routines become more efficient, and strategic priorities become clearer as the company operates. This development is implemented through agentic employees, multi-agent workflows, and owner-guided business strategy. The aim is not only to correct past errors, but to make future work faster, more reliable, and less dependent on direct owner intervention. We therefore organize OPAC development across three levels of capability growth. At the individual level, it improves how agentic employees execute, revise, retain experience, and externalize skills. At the team level, it improves cooperative and competitive workflows among agents. At the company level, it supports business expansion and goal sharpening by identifying capability gaps and the infrastructure needed for further improvement. Figure 6 summarizes these levels and mechanisms.

4.1. Agentic Employee Development

Agentic employee development concerns how an agentic employee improves through repeated work and converts task execution into stable and reusable capability. We organize this section by the durability and reusability of this development. The discussion moves from execution capacity adaptation and revision capacity enhancement to experience accumulation and skill externalization. This progression shows how task-local adjustment can become persistent individual capability, thus prompting the development of OPAC.

4.1.1. Execution Adaptation

Execution capacity adaptation is the most local form of agentic employee development. It improves how an agent handles an ongoing task before any long-term record or skill is created. This form of adaptation usually operates at test time. The agent changes the execution state of the current run by updating the context, selecting a more suitable decision policy, or revising the action trajectory after new observations. This adaptation improves immediate execution and reduces error propagation, but its effect remains task-bound unless later mechanisms retain the resulting lessons.
Context adjustment updates the evidence, demonstrations, history, and constraints available to the agent during the current run. Example-selection methods choose prior demonstrations or trajectories that are likely to transfer to the current task [163]. Context-compression methods reduce long histories or documents so that limited context is allocated to task-relevant information [109,164]. Retrieval-based methods augment the current context with external evidence when the available information is insufficient [165,166]. The adjusted object is the agent’s information state rather than its decision rule or remaining action sequence.
Inference-policy adjustment changes how the agent selects the next reasoning step, tool call, or action under the current information state. Search-based methods generate and evaluate multiple reasoning or action candidates before selecting the next step, as in Tree of Thoughts [167] and Language Agent Tree Search [168]. Decoding-level and controller-level methods alter candidate selection during test-time execution [169,170,171,172,173]. The adjusted object is the local decision procedure. This differs from context adjustment because the available information may remain fixed. It also differs from trajectory adjustment because it does not necessarily revise the remaining multi-step plan.
Action trajectory adjustment revises the remaining sequence of subgoals or actions after environmental feedback changes the execution state. The agent does not only choose a different next action. It updates how the rest of the task should proceed. ReAct [174] provides the closed-loop structure needed for this adjustment by interleaving reasoning, action, and observation. The same pattern is prominent in agents that operate through the web [175], search [176], and mobile interfaces [177], where later actions depend on the observed effects of earlier ones. The adjusted object is therefore the remaining trajectory rather than only the current context or one-step decision rule.

4.1.2. Validation-Guided Revision

Validation-guided revision extends individual development from adaptation during execution to the repair of a decision or action after validation exposes a failure. The agent interprets the failure signal, identifies the required change, and produces a revised decision or action for further validation. When verifiable evidence is available, revision should be grounded in that evidence to identify the failed requirement and constrain the correction. When such evidence is unavailable or insufficient, the agent should generate an explicit revision hypothesis and test it in the next validation cycle.
Grounding revision signal uses evidence outside the model’s own judgment to decide whether a candidate output should be accepted, rejected, or repaired. The verification can come from different sources, whether a critic evaluates intermediate results against a planning loop [178], a symbolic procedure tests the revision for logical consistency [179], or a tool is run and its output is used as a validation signal [180]. CRITIC [181] directly instantiates this idea by using external tools to validate and revise model outputs. The revision signal can be further strengthened by tying it to task-specific validity criteria, as when a draft query is executed and its result drives the next revision [182,183]. Grounding is preferred because it ties repair to observable outcomes rather than to fluent self-assessment. However, grounded signals are not always available or sufficiently diagnostic. Some tasks provide only a failed outcome, while open-ended artifacts may lack a direct executable validator. In such cases, the agent still needs a way to identify what should be changed.
Generating revision signal addresses cases where direct grounding is absent or insufficient to specify the repair. The agent produces a critique, reflection, or repair plan that turns a failed or unsatisfactory output into a candidate direction for revision. The most direct approach uses the same model to generate an output and evaluate it, then conditions the subsequent revision on that evaluation. Self-Refine [184] runs this as a loop, with the same model generating, critiquing, and revising in turn, while Reflexion [185] keeps the critiques in an episodic buffer so that failure information from earlier trials informs later ones. Free-form self-critique is cheap to elicit, but its value depends on whether the critique identifies repairable errors. Later work therefore makes the critique more structured by anchoring reflection in prior feedback and accumulated experience [186,187], or by combining critique with search and external knowledge [188,189]. Such signals should be treated as revision hypotheses rather than final evidence. They become more reliable when they are checked by tests, tools, external evaluators, or task-specific validity criteria [190].
Overall, revision capacity proceeds from grounded repair when verifiable evidence is available to generated repair hypotheses when such evidence is weak. Its broader developmental value depends on whether the correction or failure pattern can be retained beyond the current task. This leads to experience accumulation, where execution and revision records are preserved for later tasks.

4.1.3. Experience Accumulation

Experience accumulation gives continuity to individual development. In OPAC, recurring tasks often share constraints, formats, or failure modes. Without retained experience, an agent must reconstruct these conditions in each run. Memory provides the main mechanism for this continuity. It records useful information from task execution, retrieves relevant records when a later task arrives, and updates these records as new outcomes reveal what remains useful. During execution, retrieved memory can serve as additional context, decision guidance, or a warning against repeated failures. Agentic employees develop this capability through experience retention, experience abstraction, and experience maintenance. Retention keeps useful records available beyond the current context. Abstraction turns stored records into guidance for later decisions. Maintenance keeps accumulated experience selective and reliable as tasks and operating conditions change.
Experience retention starts with long-term memory, which keeps useful records available across interactions beyond the limited context of a single session. Retention also depends on retrieval, because stored records only become useful when they can be recalled for a relevant later task. Early agent-memory systems show that memory is not a single undifferentiated store. Generative Agents organize observation, memory retrieval, reflection, and planning into an architecture for behavior over time [191]. MemGPT frames long-horizon interaction as virtual context management, where information moves across memory tiers so that an LLM can operate beyond the immediate context window [192]. Later systems focus more directly on memory editing, retention, and retrieval. MemoryBank keeps interaction histories in a memory store and adjusts retention over time, so important memories are reinforced and outdated memories receive lower retrieval priority [193]. Mem0 takes a more explicit editing approach. It extracts candidate facts from a conversation and reconciles them with existing memory through add, update, and delete operations [194,195].
Experience abstraction builds on retained records by turning them into guidance for future decisions. A prior here refers to retrieval-dependent guidance that biases later decisions toward patterns that worked before, while the agent still has to interpret this guidance during execution. One approach condenses recurring patterns across successful trajectories into hints about effective action patterns [196,197]. Another converts reflections and reasoning traces into cautions or templates that signal failure modes and salient task features [198,199,200,201]. Through this abstraction, memory becomes more than an archive of past interactions. It provides retrieval-dependent guidance for future decisions. However, this guidance still remains implicit because it must be interpreted during execution. It is not yet a capability that the agent can directly invoke or compose.
Experience maintenance becomes necessary as memory grows, because redundant and outdated entries accumulate over time and reduce the quality of retrieval. Work on memory orchestration [202] and runtime adaptation [203] shows that accumulated experience must evolve with tasks and operating conditions. Some systems make this maintenance active, as in MemMA [204], which probes its own memory with test questions after each session and repairs the entries that fail. These methods extend experience accumulation from retention to active memory management. Memory governance keeps accumulated experience accessible and reliable, but it does not remove the need to reconstruct recurrent procedures during execution. This gap separates retrieval-dependent experience from reusable skills.

4.1.4. Skill Externalization

Skill externalization is the most reusable form of agentic employee development. Experience accumulation preserves records that can guide later execution, but the agent still needs to interpret those records during each task. Skill externalization goes further by packaging recurrent procedures as callable skill assets. As these assets accumulate, individual development depends on Trajectory distillation, which forms skill assets from past trajectories, and Skill maintenance, which keeps the skill library usable as tasks change.
A skill denotes a reusable procedural unit that can be selected and executed when a task matches its intended use. It usually has two layers. The first is a compact routing description that specifies the task condition under which the skill should be used. The second is a skill body, which contains the executable procedure and the information needed to check its result. During execution, the agent first matches the current task against the routing descriptions in the skill library. Only after a relevant skill is selected does the system expose the corresponding skill body to the agent. This progressive disclosure keeps skill selection lightweight while avoiding the cost of loading many full procedures into context. Voyager [205] illustrates this pattern through executable code skills that are stored and reused across tasks. AutoSkill [206] further treats skills as explicit artifacts extracted from interaction traces and maintained for later use. This progressive disclosure keeps skill selection lightweight while avoiding the cost of loading many full procedures into context [148].
Trajectory distillation provides a direct route through which recurrent behavior becomes a reusable skill asset. While some skills can be predefined from human routines or domain templates, OPAC operation may reveal repeated procedures that are difficult to specify in advance. Trajectory distillation addresses this case by abstracting recurring behavior from past trajectories into a named and self-contained unit. The resulting asset can be retrieved and applied to similar tasks. Existing methods extract skill artifacts from interaction traces [206], distill reinforcement learning trajectories [207], or convert local task lessons into transferable skill units [208]. The resulting asset may be implemented as an executable subagent [209], a code-based routine [210,211], or an entry in an environment-grounded skill repository [212]. Once such assets enter the skill library, the development problem shifts from forming individual skills to keeping the library useful over time.
Skill maintenance is necessary as the skill library grows. Distilled skills may become redundant, obsolete, or poorly matched to new tasks, so the library must be updated rather than only expanded. CoEvoSkills [213] maintains skills through a generate–verify–refine loop. A skill generator revises the skill, and a surrogate verifier checks the revision and returns feedback. SkillX [214] moves maintenance to the skill-library level by merging redundant skills, filtering low-value skills, and updating the library as new tasks appear. Related work further studies skill verification [215], repository management [216,217], and skill optimization without full retraining [218].
Skill externalization completes the transition from task-level improvement to persistent agent-level development. Trajectory distillation creates explicit skill assets from recurring behavior, while skill maintenance keeps the skill library usable as work changes. This allows agentic employees to invoke, compose, and update learned routines across later OPAC tasks.

4.2. Team Workflow Development

Team workflow development concerns how multiple agentic employees coordinate and challenge one another when OPAC tasks exceed the scope of a single role. In a human company, complex work is usually completed through collaboration among employees and through review or competition that tests proposed outputs. OPAC has a similar organizational need, but the roles are performed by agents. A workflow specifies how agents divide subtasks, exchange intermediate artifacts, place review steps, and integrate partial outputs into a usable result. It can be developed in two directions. Cooperative workflow development improves production and integration among agents. Competitive workflow development introduces challenge and verification into the team process.

4.2.1. Cooperative Workflow Development

Cooperative workflow development focuses on the productive organization of agent teams. The cooperation is often represented as a structured process (e.g., communication graphs), where agents are represented as nodes and message channels are represented as edges. It makes cooperation adjustable, since the system can change how agents communicate, which agents are activated, and how intermediate outputs are routed. The cooperative workflow development can be organized into two forms. Workflow refinement keeps the main role decomposition and workflow skeleton fixed, then improves communication, routing, or active-agent selection. Workflow reconfiguration constructs or searches for a new workflow when the current structure no longer fits the task.
Workflow refinement optimizes the topology of existing agentic team workflows. GPTSwarm [33] represents the team as a graph and optimizes its edges with reinforcement learning, pruning links that lower the benchmark score and retaining those that improve it. Later studies extend this graph-based view from edge optimization to communication pruning and active-team adjustment. AgentPrune [219] removes redundant messages from a spatial-temporal communication graph, while AgentDropout [220] eliminates low-contribution agents and links to reduce communication cost. Related methods further refine the collaboration topology through score-guided graph pruning [41,221], joint prompt–topology optimization [222], or gradient-guided edge optimization [34,223]. These methods assume that the basic role decomposition and workflow skeleton remain appropriate. They improve coordination within that structure by adjusting communication, routing, and active-agent selection. When this structural assumption no longer holds, workflow development must move from refinement to reconfiguration.
Workflow reconfiguration is needed when the existing workflow no longer represents the task well. It synthesizes a workflow conditioned on the task instead of working within a fixed structure. Since the workflow design space is large and task-dependent, effective methods must search or synthesize high-performing structures under task-specific constraints. One group of methods performs offline workflow search, synthesizing workflows without a fixed template [224,225] or optimizing and evolving candidates before deployment [226,227]. Another constructs the workflow at run time. FlowReasoner [228], for instance, trains a meta-agent to infer a tailored workflow for each query and refine it from execution feedback. The team is instantiated for each query rather than fixed in advance. Related runtime methods build the architecture and coordination to fit each task [39,229,230,231,232] or keep adapting the workflow as conditions change [233,234,235]. MetaGPT [236] further demonstrates how structured workflows can coordinate specialized agents in software development. AgentVerse [237] supports this view by showing that organized multi-agent groups can outperform a single agent. The practical value of workflow reconfiguration has been demonstrated in scientific discovery [238], embodied operation [239], and recommendation optimization [240], where it maps task requirements to agent composition, execution order, and output integration.

4.2.2. Competitive Workflow Development

Competitive workflow development focuses on reliability and capability-boundary exploration in agent teams. In human companies, internal competition may increase effort, but it can also create conflict, reduce knowledge sharing, or introduce psychological cost. In OPAC, competition can be implemented as structured interaction among agents without the same human incentive burden. The purpose is not rivalry itself, but to expose weak outputs, unsupported assumptions, and untested capability boundaries. Structurally, competitive workflows can use the same process or graph representations as cooperative workflows. The difference lies in what the interaction is designed to transmit. In cooperative workflows, edges mainly pass task outputs and dependencies. In competitive workflows, edges pass critiques, verification results, adversarial tasks, or counterproposals. We distinguish two forms. Peer-challenge places review or verification inside the production process. Self-challenge creates new tasks, opponents, or verification settings to reveal failures that ordinary evaluation may miss.
Peer-challenge workflow develops team reliability by adding challenge relations to a production process. Existing work instantiates this pattern in several debate-based workflows place several agents in a symmetric review loop, where agents compare competing answers, inspect one another’s reasoning, and update their responses across rounds [101]. Dynamic debate workflows further adapt role assignment before the debate so that agents occupy question-specific positions in the review process [56]. Referee-based workflows add an evaluator layer after production, as in ChatEval, where a reviewer team discusses and scores generated outputs before final selection [103]. Solver-verifier workflows create a tighter production-review cycle, where one role proposes a solution and another role constructs tests or formal checks that decide whether the solution should be accepted or sent back for repair [241,242]. Critic-centered workflows strengthen the review node itself by training critics to detect subtle errors from adversarially generated defects or rubric-based feedback [243,244]. Across these variants, peer-challenge develops the workflow by changing the review topology, the feedback path, and the acceptance rule that connects production to revision.
Self-challenge workflow develops a team by generating the pressure under which the team improves. A common form is a challenger, solver, and critic loop. The challenger proposes difficult cases, the solver attempts them, and the critic or verifier filters unreliable challenges and checks the result. This mechanism is useful when predefined evaluations do not cover the weaknesses that appear in operation. The Self-Challenging [245] instantiates this pattern by having an agent discover tool affordances, convert them into verifiable tasks, and train under verifier feedback. SAGE extends the pattern into a persistent challenger-solver-critic configuration, where the challenger raises difficulty, the solver attempts the generated tasks, and the critic filters unreliable ones [246]. Zero-data and self-play methods similarly create learning pressure from internally generated problems or opponents [247,248,249,250]. Guided self-play adds control signals to reduce repetitive or uninformative challenges [251,252,253,254]. A related pattern appears in open-ended research systems that stress-test hypotheses through internal debate and counterproposals [255,256]. The main risk is distributional drift. The team may improve on internally generated challenges without improving on the target task distribution.

4.3. Business Strategy Development

Business strategy development concerns how OPAC directs capability growth at the company level. The business boundary of an OPAC may shift as owner priorities or market conditions change. At this level, the central issue is whether current capabilities can support adjacent work and how broad owner intent can be translated into learnable improvement targets. Business expansion addresses the first issue by examining whether accumulated capabilities can support a new task family. Goal sharpening addresses the second issue by specifying the capability targets and development infrastructure needed for further improvement.

4.3.1. Business Expansion

Business expansion arises when OPAC considers work beyond its current operating scope. This may occur when the owner pursues a new business interest or when external demand makes a related domain more valuable. Expansion is feasible only when the new domain can become repeatable operation. The process begins with Expansion fit, which evaluates whether the domain matches current or attainable operating capacity. When the fit is promising but competence is still incomplete, Capability bootstrapping forms initial capability. Experience reuse further lowers the cost of expansion when prior operation can transfer to the new domain.
Expansion fit asks whether an adjacent task family can become a stable OPAC operation. Benchmarks are useful when they approximate this operational fit. Isolated task success is less informative unless it reflects stable interfaces, business artifacts, and feedback conditions. AgentBench [257], WebArena [258], and OSWorld [259] test whether agents can operate through interactive environments. SWE-bench [260] and SWE-agent [261] evaluate professional software work under repository-level interaction and test execution. GAIA [262], Superglasses [263], and Hibench [264] captures information work that requires reasoning, web use, multimodal evidence, and tool use. ToolBench [265] and τ -bench [266] evaluate tool-mediated service loops under user requests and domain policies. TheAgentCompany [267] is especially relevant to OPAC because it evaluates agents in a simulated workplace where they browse the web, write code, run programs, and communicate with coworkers. These evaluations indicate whether a candidate domain provides the operating conditions needed for persistent capability growth.
Capability bootstrapping becomes necessary when a promising domain lacks the demonstrations, rewards, or routines assumed by supervised training. In this setting, OPAC needs agents to form initial competence before mature supervision is available. AgentEvolver illustrates this idea. Initialized in an unfamiliar environment, it has the agent explore autonomously, generate its own tasks from discovered information, and reuse the resulting experience to guide further exploration without a handcrafted dataset [126]. Other methods use generated data [268,269,270], early self-collected experience [271], or the model’s own judgments of its outputs [272]. These methods reduce the cost of entering a new task family before task-specific training infrastructure is fully available.
Experience reuse determines whether prior OPAC operation can reduce the cost of expansion. The useful unit is a transferable abstraction distilled from past trajectories. ReasoningBank distills generalizable strategies from an agent’s own successes and failures into reusable principles, retrieves the relevant ones when a new task arrives, and folds fresh lessons back in [273]. Forward-learning and experience-lifecycle methods [274,275] and cross-domain memory transfer [276,277] share this objective. They turn prior operation into transferable capability, especially when adjacent tasks share tools or reasoning patterns. Similar reuse appears in GUI operation, embodied agents, medical support, counseling, analysis, and grid-operation systems that carry case histories or prior strategies into related tasks [278,279,280,281,282]. This connects business expansion to the earlier mechanisms of experience accumulation and skill externalization.

4.3.2. Goal Sharpening

Business expansion can identify a plausible growth direction, but it leaves the improvement target underspecified. Owner intent often begins as a broad business preference. Goal sharpening turns this preference into a capability target that can be evaluated and improved. A useful target has operational value and is narrow enough for evidence to show progress. Goal sharpening also determines the evaluation and update infrastructure needed to make that progress learnable.
Target selection focuses on the missing capability that constrains the intended business outcome. The goal is not to optimize an undifferentiated performance metric. It is to locate a learnable capability whose improvement has operational value. TRACE [283] does this by contrasting an agent’s successful and failed runs to identify the specific capability it lacks, then builds a training environment that isolates and rewards that capability. Curriculum-based methods then turn an identified target into graded learning pressure [284,285,286]. A suitable target is not necessarily the most difficult capability. It is the one whose improvement is strategically useful and still supported by enough feedback to be learnable.
Infrastructure setup concerns the situation that once a capability target is selected, reliable improvement depends on infrastructure that produces feedback and applies updates. Data-oriented work improves training data through autonomous curation or self-evolving synthesis [287,288]. Rationale-oriented work bootstraps reasoning traces from the model’s own successful generations, as in STaR [289]. Prompt-optimization work treats prompts as updateable artifacts, as shown by OPRO [290] and PromptBreeder [291]. Pipeline-oriented work optimizes LM programs or agent pipelines against task metrics, as in DSPy [292]. Environment-oriented work synthesizes environments for later learning iterations, so feedback is grounded in real interaction [293,294,295]. Eureka [296] shows that reward design itself can be generated and improved by an LLM. Procedure-oriented work automates the update loop, from agent-driven fine-tuning and evolution pipelines [297,298] to protocols that version each change and allow rollback [299].
Business expansion and goal sharpening give company-level development a strategic direction. Expansion tests whether OPAC capabilities can enter adjacent work. Goal sharpening defines the capability target and the update conditions required for that move.

5. Trustworthiness Risks and Safeguards in OPAC

To systematically investigate the trustworthy concerns in emerging OPAC paradigms, we map existing trustworthiness research onto organizational control problems specific to OPAC, where only one OPAC owner supervises an agentic organization. In particular, OPAC’s trustworthiness stems from a structural asymmetry: a single human owner remains responsible for oversight, authorization, intervention, and external accountability, while operational work is delegated to a larger and increasingly autonomous agent organization. Building upon such insights, the central question is not only whether individual agents are reliable, but also whether effective control can be maintained as delegation scales under a single OPAC owner. To answer this question, we organize the chapter around the first point at which control becomes ineffective in a single-owner OPAC. One set of failures arises on the oversight side, where the owner lacks the calibrated information, independent counter-input, or attention needed to exercise effective judgment. Another arises on the execution side, where delegated agents act on untrusted inputs, unreliable knowledge, or excessive authority before the owner can intervene. We then examine the safeguards needed for single-owner operation according to where they intervene in this control process. Accordingly, Section 5.1 defines this organizational-control lens; Section 5.2 and Section 5.3 review the two risk domains and their recurring failure modes; and Section 5.4 discusses the safeguard families that respond to them. Figure 7 summarizes this structure.

5.1. Trustworthiness as Organizational Control

Trustworthiness depends not only on model-level reliability but also on how technical controls and organizational authority are connected. In other words, risk taxonomies, lifecycle governance, internal auditing, management-system standards, and meaningful-control scholarship all converge on this point [58,300,301,302,303,304,304]. In conventional firms, human governance authority is distributed across managers, specialists, reviewers, and auditors. OPAC concentrates final human governance in a single-owner control loop: agents can perform operational review, filtering, and monitoring as described in Section 2, but the OPAC owner reviews escalated state and evidence, decides whether intervention is needed, authorizes exceptions, and bears accountability for resulting outcomes. We classify risks by the first point in this loop at which control becomes ineffective:
  • Human oversight risks. These risks weaken the feedback available to the owner through poor calibration, distorted advice, or attention overload.
  • Agent execution risks. These risks compromise the operational pipeline through untrusted inputs, unreliable knowledge, or authority that exceeds the intended delegation.
This taxonomy classifies the first point of control failure rather than the ultimate harm. Privacy loss, data-protection violations, and leakage of commercial secrets are therefore treated as cross-cutting consequences that can arise from either domain: compromised inputs, knowledge stores, or delegated actions may expose data directly, while miscalibration or attention failure may allow the exposure to proceed. Failures propagated across agent handoffs are assigned to the point where compromised information first enters the workflow or authority first exceeds its boundary.
The categories are not equally OPAC-specific or equally urgent. Attention capacity is especially distinctive in OPAC because adding human reviewers would relax the one-person premise, making scarce owner attention a structural bottleneck rather than a temporary staffing problem. Input integrity and execution control become higher priorities as agents gain access to sensitive data, external communication, financial authority, or irreversible tools, while knowledge and judgment risks become more urgent when outputs are costly to verify. A risk-proportional ordering therefore depends on consequence, reversibility, exposure, and review cost rather than the order in which categories appear.
The same logic structures the safeguards discussed in Section 5.4. Each targets a different failure point, but none fully reproduces the institutionally independent human redundancy of a conventional firm. Their OPAC-specific role is to connect concentrated authority to inspectable evidence and bounded action.

5.2. Human Oversight Risks

Human oversight risks arise when the OPAC owner remains formally responsible but lacks the information, independent counter-input, or attention needed to exercise effective control. Automation research has long distinguished the presence of a human from effective use, appropriate reliance, and recoverable intervention [61,305]. Human-AI interaction work further shows that interface design and system behavior jointly determine whether users form accurate expectations and can recover from errors [306]. We organize this domain around three failure modes:
  • Explainability and reliance calibration risk. The owner receives signals that appear informative but do not support well-calibrated reliance.
  • Judgment distortion risk. Interaction with the agent weakens the owner’s independent evaluation.
  • Attention capacity risk. The owner’s limited time and monitoring bandwidth become the bottleneck for effective intervention.
These risks can compound: poor calibration makes the owner more susceptible to judgment distortion, and both effects are amplified when attention capacity is exhausted.

5.2.1. Explainability and Reliance Calibration Risk

Explainability and reliance-calibration research addresses related but distinct questions. On the explanation side, methods differ in whether they expose what information shaped a prediction or whether a rationale faithfully reflects the mechanism that produced it [307,308]. On the reliance side, work asks how users decide whether to accept, verify, or override advice when uncertainty signals may be absent, elicited from the model, or communicated through an interface [309,310]. Treating transparency as sufficient can make an interface appear informative while leaving the owner’s reliance poorly calibrated.
Explanation and reasoning faithfulness. Post-hoc methods such as LIME make classifier behavior locally inspectable [307], but human-AI studies show that adding an explanation does not automatically produce complementary team performance [59]. Chain-of-thought research exposes a deeper boundary. Models can produce plausible rationales that omit influential prompt features or rationalize an answer after the fact [308]. Intervention-based work therefore tests faithfulness by removing or unlearning reasoning steps and observing whether final behavior changes [311]. Post-hoc explanation provides access to a representation of behavior, whereas faithfulness analysis asks whether that representation supports causal interpretation.
Uncertainty communication and appropriate reliance. The literature distinguishes three deployment conditions. When an interface exposes no confidence estimate, existing studies operationalize uncertainty by eliciting or approximating an estimate: Tian et al. compare conditional probabilities with verbalized confidence from RLHF language models [312], while Xiong et al. compare verbalization, repeated sampling, and consistency-based aggregation for black-box LLMs [309]. When a numerical or verbal confidence cue is displayed, it may still be miscalibrated; predictive-calibration research documents confidence–correctness gaps and degradation under distribution shift [313,314], while LLM and human-subject studies show overconfident elicited estimates and impaired reliance under misleading confidence cues [309,315,316]. Even informative cues do not guarantee effective joint decisions: confidence displays can improve trust calibration without improving team accuracy [317], explanations can change whether users accept correct and reject incorrect advice [310], and expectation-setting interventions shape users’ mental models before errors occur [318].
An OPAC purchasing workflow may receive a fluent rationale and cited evidence without an explicit uncertainty signal, or it may receive a numerical or verbal estimate whose calibration is unknown. Explanations, provenance, uncertainty cues when available, and the owner’s prior beliefs are therefore distinct inputs rather than confirmations of the same judgment. A conventional firm can distribute interpretation and approval across analysts, reviewers, and decision-makers; in OPAC, calibrated interfaces and agent-based review can support scrutiny, but the same human owner often resolves the remaining uncertainty and authorizes the action.

5.2.2. Judgment Distortion Risk

Judgment distortion concerns whether interaction with an agent preserves the OPAC owner’s independent evaluation. The main literature studies sycophancy: responses that adapt to a user’s expressed position, preference, or desired social outcome rather than maintaining an evidence-grounded assessment. Unlike ordinary hallucination, sycophancy is conditioned on the interlocutor and can therefore reinforce an error precisely because the owner signals commitment to it.
Measurement and interaction effects. SycEval treats sycophancy as a benchmarkable behavior across advice and reasoning settings [319]. ELEPHANT extends the construct from explicit factual agreement to social behaviors such as face preservation and conflict-sensitive affirmation [320]. Argument-driven evaluation further shows that user-provided reasoning can induce stance mirroring, making the interaction history part of the failure mechanism [321]. These studies differ in task design, but collectively move the field from anecdotal agreement toward taxonomies of when and how user-conditioned distortion appears.
Belief reinforcement. A complementary line studies how AI advice and explanations alter human judgment rather than only whether a model agrees with a prompt. Clinical decision-aid experiments show that incorrect AI recommendations can reduce human accuracy even when a person retains final authority [322], while misinformation studies find that deceptive LLM explanations can change beliefs more often than honest explanations [323]. Longitudinal modeling extends these one-shot effects into feedback loops in which validation changes the user’s beliefs and subsequent prompts, which then elicit further validation [324].
Assigning the same advisor agent to draft and critique a market-entry case can turn nominal review into self-confirmation, but this is a design-contingent failure rather than an inherent OPAC property. OPAC can construct counter-evidence through separate evaluator agents, adversarial debate, or isolated sessions; related multi-agent review mechanisms are surveyed in Section 2. Unlike institutional role separation among human professionals, however, this independence is architectural: it must be deliberately configured, and its effectiveness depends on model, context, and evaluation design.

5.2.3. Attention Capacity Risk

Attention capacity risk concerns the amount and timing of supervision that one person can provide. It draws on automation and supervisory-control research on operator reliance, situation awareness, and workload [61,62]. This literature rejects the assumption that adding approval steps or increasing automation monotonically improves control; both strategies can fail when the operator cannot maintain situation awareness or distinguish consequential events from routine traffic.
Supervisory capacity and automation irony. Fan-out models study how many autonomous units one operator can supervise as interaction demand and neglect time change [325]. Dynamic-overload models complement this capacity view by predicting workload fluctuations during supervisory control and evaluating adaptive cueing before overload produces failure [326]. These frameworks make oversight a measurable allocation problem rather than an unlimited human resource. Earlier automation research explains why this capacity can decline as systems become more autonomous: misuse and disuse arise when trust is poorly calibrated [61], while the ironies of automation place humans in charge of rare exceptions after routine practice and system understanding have eroded [305].
Monitoring overload and task allocation. Alert-fatigue research in security operations examines how high-volume monitoring pipelines obscure priority and weaken response quality [327]. Studies of multi-agent supervision further show that explicit alerts redistribute visual attention, while task difficulty affects willingness to accept autonomous behavior, making alert design part of the task-allocation mechanism rather than a neutral notification layer [328]. Meta-analytic evidence on human-AI combinations likewise shows that joint systems are useful under particular task and complementarity conditions, not simply whenever a human is inserted into the loop [63].
An OPAC owner supervising finance, sales, and operations may receive dozens of routine approval requests; if each interrupts the same queue, a high-consequence transaction can arrive amid low-value traffic. Unlike a conventional firm, an OPAC cannot add human reviewers without weakening its defining one-person constraint. Agent reviewers and tiered monitoring queues can filter and prioritize the stream, but the volume of items requiring owner escalation remains the structural bottleneck. Effective oversight therefore depends on selective intervention, workload-aware routing, and preservation of owner competence.

5.3. Agent Execution Risks

Agent execution risks arise when compromised inputs, unreliable knowledge, or excessive action authority alter the OPAC owner’s intended delegation inside the operational pipeline. The owner may be willing to intervene, but an external instruction may already have entered the context, corrupted knowledge may already have shaped a plan, or a tool call may already have changed the external state. An agent execution cycle proceeds through three stages: receiving and parsing external inputs, retrieving and reasoning over knowledge, and invoking tools or producing outputs that change external state. In multi-agent workflows, downstream agents implicitly inherit trust in upstream outputs and delegated authority, so a single failure can amplify along the delegation chain before the owner observes it. Conventional cross-role handoffs can insert structured checks or independent validation; OPAC instead depends on encoding comparable checks into the workflow. Tool-using agents blur the boundaries between these stages and amplify their interactions because a single agent may combine language interpretation, retrieval, memory access, API calls, and delegated authority within one execution trace [329,330]. Here, agent execution risks denote failures inside this pipeline, whereas execution boundaries in Section 5.4.2 denote safeguards placed around it. In OPAC, these actions carry real business consequences: a payment API call completes a financial transaction, an email to a client becomes an external communication, and a database update alters the company’s operational state. We organize this domain around three execution-stage failure modes:
  • Input integrity risk. Untrusted external content alters the instruction hierarchy or plan.
  • Knowledge reliability risk. Generated, retrieved, or remembered state becomes an unreliable basis for action.
  • Execution control risk. Available permissions and tool sequences exceed the intended delegation boundary.

5.3.1. Input Integrity Risk

Input integrity concerns whether external content can alter an agent’s instruction hierarchy or operating assumptions. Prompt injection research distinguishes direct attacks supplied as user instructions from indirect attacks embedded in documents, webpages, emails, or other resources that agents process while pursuing a legitimate task. In agentic settings, the security consequence depends not only on model compliance but also on which tools and data become reachable after the instruction is accepted.
Taxonomies and benchmarks. Liu et al. formalize prompt injection and benchmark attacks and defenses under a common evaluation setting [331]. AgentDojo moves this analysis into dynamic tool-use environments, where attacks and defenses are evaluated against both task utility and security objectives [332]. The former provides an attack and defense vocabulary; the latter shows how injection propagates through realistic agent tasks. Together, they position input integrity as an end-to-end property rather than a prompt-filtering problem alone.
Deployment evidence. Real-world application studies show that indirect prompt injection can enter through retrieved webpages, messages, and other externally supplied content, then redirect integrated applications toward data exfiltration, API manipulation, or task hijacking [333]. Subsequent empirical work classifies how malicious instructions are adapted to agent capabilities and downstream objectives across deployment channels [334]. Together, these studies extend controlled benchmarks toward the source channels that deployed agents actually encounter.
In an OPAC customer-service workflow, a malicious instruction embedded in a customer email can cross from untrusted content into a plan and then reach CRM, refund, or outbound-messaging tools. A conventional organization can assign security personnel to monitor intake channels and investigate incidents; OPAC instead shifts source trust, screening, and escalation into the workflow because the owner cannot continuously inspect every document. Every additional source therefore expands both operational coverage and the untrusted-input boundary.

5.3.2. Knowledge Reliability Risk

Knowledge reliability concerns the state that agents use after input processing. Existing work separates at least three failure mechanisms: generation can produce unsupported content, retrieval can surface adversarial or corrupted evidence, and long-term memory can accumulate stale, conflicting, or low-quality records. In particular, adversarial corruption of an evidence store and non-adversarial degradation of persistent memory can produce similarly grounded-looking outputs but require different safeguards. These mechanisms therefore require different evidence and controls even when they produce the same downstream symptom of an incorrect business decision.
Hallucination and verification. Hallucination surveys organize failures by task, cause, evaluation method, and mitigation strategy [335]. Evaluation work decomposes long-form responses into atomic claims to measure factual precision [336], whereas black-box detection uses disagreement across sampled generations as a signal of unsupported content [337]. Mechanistic detection work asks whether internal hallucination-related signals transfer across domains [338], while reference-checking research examines fabricated or unresolved citations in commercial models and deep-research agents [339]. These streams address different verification objects—claim-level factuality, output consistency, internal signals, and source resolution—so their scope must remain visible to the OPAC owner.
Retrieval and memory poisoning. Retrieval-poisoning research demonstrates that adversarial passages inserted into a corpus can manipulate what dense retrievers return [340]. PoisonedRAG shows that a small number of malicious texts can steer responses from much larger knowledge stores toward attacker-selected answers [341], while AgentPoison extends the attack surface to agent memory and knowledge bases, where retrieved triggers can redirect downstream planning and behavior [342]. This stream differs from ordinary hallucination because the system may faithfully use the evidence it retrieves while the evidence store itself has been corrupted. Provenance, corpus controls, and memory-write policies are therefore central to the security boundary.
Long-term memory lifecycle. Non-adversarial degradation is studied by memory architectures that manage retention, forgetting, consolidation, retrieval, and repair. MemoryBank introduces updating and forgetting mechanisms for long-term interaction [193], while Mem0 and MemMA respectively emphasize scalable consolidation and multi-agent coordination of memory construction, use, and revision [194,204]. These systems are representative memory-management approaches, not established security defenses.
An OPAC due-diligence agent may faithfully summarize a poisoned retrieval corpus, causing an apparently grounded output to inherit adversarial evidence. Independently, the same agent may reuse an obsolete supplier record from persistent memory, producing similar downstream unreliability through non-adversarial decay rather than adversarial corruption. Conventional organizations can distribute fact-checking, records management, and data stewardship across roles; OPAC instead depends more heavily on automated provenance checks and memory maintenance because the same owner cannot continuously audit every evidence store. Persistent organizational memory therefore requires explicit maintenance rules before it can serve as a dependable record for later decisions, linking trustworthiness to the resource-management mechanisms in Section 3.

5.3.3. Execution Control Risk

Execution control concerns the difference between what the OPAC owner intends to delegate and what an agent can technically perform. Static tool access, broad credentials, and composable APIs can allow individually permissible steps to produce an unintended high-consequence sequence. The literature therefore studies both how to evaluate risky tool behavior and how to bound autonomy through permissions, reversibility, and intervention.
Safety evaluation and attack surface. ToolEmu evaluates tool-using agents inside an LM-emulated sandbox so that risky actions can be identified without executing them against real systems [329]. ToolSword benchmarks safety scenarios across the input, execution, and output stages of tool learning [343], while R-Judge tests whether agents recognize safety risks embedded in interaction trajectories [344]. The OWASP agentic-application taxonomy complements these benchmark views by cataloging risks associated with tool misuse, privilege, memory, and chained actions [330]. Together, these works treat execution safety as a system property created by the interaction of model behavior, tool interfaces, state, and authorization.
Bounded autonomy and intervention. Governance frameworks for agentic AI emphasize permission scope, human accountability, action reversibility, and lifecycle controls [345]. Magentic-UI provides a system-level example in which co-planning, action guards, and answer verification place human intervention at selected points in an agent workflow [65]. These streams define what an agent may access, which actions remain reversible, and which consequence levels trigger interception. For example, an OPAC finance agent may read invoices continuously while receiving a transaction-specific credential only after a transfer passes a policy check or owner review. Bounded autonomy therefore does not imply that every action should wait for approval.
Compositional authority and reversibility. Multi-step risk arises when each operation is locally permitted but the sequence crosses an authorization boundary. An agent may be allowed to retrieve customer records, compile a report, and send external emails, while the combination discloses data beyond the intended task. Agentic-risk frameworks therefore emphasize chained actions, permission scope, and reversibility [330,345]. Conventional organizations can separate request, approval, credential custody, and execution across people. OPAC can distribute request initiation, credential custody, and execution across separate agents and policy engines, but ultimate policy authority, exception approval, and accountability remain concentrated in one owner. A locally permitted sequence can therefore cross a business boundary when agent-level controls are insufficiently granular and the sequence does not trigger owner escalation. Yet irreversibility is not a fixed property: an email may be retractable only briefly, a payment may be reversible only below a threshold, and a memory write may persist into later workflows. Because the owner cannot judge every sequence in real time, consequence labels, cumulative authority limits, and rollback windows become operational metadata.

5.4. Safeguards for Single-Owner Operation

The two risk domains do not fail independently. Poisoned evidence can increase owner reliance and enable unsafe action, while attention overload can allow input or knowledge failures to pass through approval gates and reach the external state. The organizational-control lens is therefore end-to-end: safeguards need to interrupt transitions between evidence, judgment, delegation, and action rather than address only isolated failure modes.
The preceding risks cannot be eliminated, but they can be mitigated by safeguards that partially substitute for organizational functions not supplied by separate human roles in OPAC. The mapping is functional rather than one-to-one. We organize these safeguards from top to bottom along the control stack:
  • Governance foundation. This family defines assets, authority, records, and escalation paths at the organizational level.
  • Execution boundaries. This family constrains delegated agent behavior within the operational pipeline, especially around inputs, knowledge, and action authority.
  • Owner-facing safeguards. This family strengthens human judgment at the final control point through better evidence, calibration, and attention allocation.
The central design principle is that safeguards should reduce rather than reproduce the OPAC owner’s cognitive bottleneck. This principle is constrained by trade-offs among security, utility, throughput, and attention, which the closing synthesis makes explicit.

5.4.1. Governance Foundation

The governance foundation makes agents, permissions, consequences, and responsibilities visible before execution. It primarily supports both risk domains by defining the control surface within which the other safeguards operate. Existing frameworks and auditing literature approach this function at different levels. NIST and ISO describe lifecycle governance and management-system requirements [301,303,303], whereas internal-audit and audit-tooling research streams specify how documentation, evaluation, communication, and accountability are connected across development and deployment [302,346]. Critiques of formal human oversight and philosophical accounts of meaningful control add an important boundary: assigning a person responsibility does not ensure that the person has sufficient authority, information, or capacity to intervene [58,304].
Lifecycle governance and auditability. Lifecycle-oriented auditing treats governance as a continuing process rather than a one-time compliance check. End-to-end and ethics-based audit frameworks connect documentation and evaluation across design and deployment, while emphasizing continuous review and allocated accountability [302,347]. Reviewability research expands the audit object from model outputs to the socio-technical decision pipeline, using context-appropriate record keeping to support review of both individual decisions and the process as a whole [348]. Accountability syntheses distinguish the actor, forum, relationship, content and criteria of the account, and resulting consequences [349], while practitioner studies show that organizational structures and communication channels can support or hinder the operationalization of responsible-AI practices [350]. Under OPAC constraints, these streams can be synthesized as a lightweight record connecting each agent and workflow to its owner, data and tool access, evaluated failure modes, and review history.
Lightweight accountability and risk tiering. Minimum Viable Governance adapts governance to resource-constrained organizations by emphasizing a small set of structures that can expand with use [351]. Audit-tooling research adds an operational warning: existing tools concentrate on evaluation and performance analysis while providing weaker support for harm discovery, audit communication, iteration, and accountability workflows [346]. The IMDA agentic-AI framework adds agent-specific factors such as permission scope, domain consequences, human accountability, reversibility, and lifecycle controls [345]. Together, these approaches support consequence tiering and explicit stop authority without assuming a large compliance department. They remain partial substitutes: agent-based review can provide functional redundancy, but an inventory and a named OPAC owner do not reproduce the institutionally independent and professionally accountable review provided by human specialists in a conventional organization.

5.4.2. Execution Boundaries

Execution boundaries continue to operate when the OPAC owner is not examining every input or action. They primarily target input integrity, knowledge reliability, and execution control, while indirectly protecting attention capacity by reducing the volume that requires active review. Their functions include mediating untrusted content, enforcing domain policy outside the model, isolating risky execution, restricting standing authority, and intercepting selected actions. The literature differs mainly in whether control is applied before model reasoning, around the execution environment, or immediately before external state changes.
Input mediation and symbolic policy enforcement. Prompt-injection benchmarks evaluate filters, instruction separation, and other defenses against adaptive attacks [331,332]. Symbolic-guardrail research takes a complementary approach by moving enforceable domain requirements out of unconstrained model judgment and into explicit software checks [352]. Input mediation reduces the chance that untrusted content controls the plan, whereas symbolic enforcement constrains the plan or output even when model reasoning remains imperfect. Neither stream supports a claim of complete prevention; their value lies in defense-in-depth and testable policy behavior.
Sandboxing, least agency, and runtime interception. Sandboxes limit the state and consequences available during evaluation or execution [329]. Least-agency guidance narrows tools, credentials, and action scope so that an agent cannot freely compose all technically available capabilities [330]. Magentic-UI illustrates runtime action guards that suspend selected operations for verification or human approval [65]. These controls are complementary: sandboxing contains effects, scoped permissions reduce reachable effects, and interception governs the remaining high-consequence transitions. Alert-fatigue research cautions that interception must be risk-proportional, because indiscriminate approval requests can transfer execution risk into attention failure [327].

5.4.3. Owner-Facing Safeguards

Owner-facing safeguards improve the evidence, calibration, and attention available at the points where human judgment can change an outcome. They primarily target explainability and calibration, judgment distortion, and attention capacity, while also supporting verification of knowledge and actions. HAI guidelines emphasize expectation setting, feedback, visible system state, and recovery from error [306]. Research on model updates and explanations further shows that higher standalone model performance or additional explanations can still reduce team effectiveness when users cannot anticipate errors or integrate the new behavior [59,60]. The objective is therefore not maximal information display, but appropriate reliance and contestable decisions.
Evidence grounding and reliance calibration. Provenance mechanisms expose the sources used in an output and support later checking, which is particularly important when deep-research agents produce unresolved or fabricated references [339]. LLM uncertainty studies examine how confidence cues can be elicited or approximated [309,312], whereas expectation-setting and confidence-display studies examine how users respond to such cues [315,316,318]. These mechanisms address different questions: provenance indicates where a claim came from, while reliance cues help allocate scrutiny when they are available and sufficiently calibrated. Neither proves that the underlying evidence is correct.
Counter-evidence and anti-sycophancy interaction. Interaction design can reduce judgment distortion before it becomes a strategic feedback loop. Cognitive-forcing functions introduce deliberate friction before users accept AI advice and can reduce overreliance, although they also impose usability costs [353]. Reframing assertions as questions targets the conversational input [354], whereas uncertainty-aware reasoning optimization targets how the model responds under user pressure [355]. These approaches operate at different intervention levels, but treat alternatives, delay, or counter-evidence as resources for limiting confirmation of the owner’s initial position.
Verification-centered attention allocation. Meaningful-oversight research argues for designing human intervention around situations where the person has sufficient information, authority, and opportunity to affect the result [356]. Magentic-UI operationalizes this principle through action guards at selected workflow transitions [65], while alert-fatigue research supplies the counterconstraint that review volume must remain manageable [327]. At the execution-boundary layer, an action guard determines when a transition is intercepted; at the owner-facing layer, the interface determines what evidence and calibration the owner receives at that point. Together, these studies characterize selective verification as a pattern in which routine low-consequence operations remain logged and reviewable, whereas anomalous, irreversible, or high-consequence actions receive active owner attention.
Taken together, governance records, execution boundaries, and owner-facing safeguards connect organizational authority with technical control so that scarce attention is concentrated where it can change the course of agent actions before consequences become irreversible. These layers can also conflict in at least three ways. First, coarse execution restrictions can reduce task utility by blocking legitimate agent actions [332]. Second, because the OPAC owner is the sole human decision-maker, cognitive forcing and repeated verification add interaction cost that can slow decisions and business throughput disproportionately [353]. Third, excessive interception can recreate the attention overload that human involvement was designed to prevent [327]. These trade-offs favor risk-proportional rather than uniformly restrictive safeguards. They neither remove the OPAC owner from accountability nor reproduce the institutionally independent compliance, security, review, and audit functions provided by separate human roles in a conventional organization.

6. Challenges and Future Directions

While OPACs demonstrate significant potential as a new organizational paradigm enabled by multi-agent systems, their scalable deployment and long-term viability in real-world applications depend on addressing a set of structural, technical, and governance challenges. Unlike traditional firms, OPACs concentrate decision authority, organizational memory, and operational coordination within tightly coupled human–agent systems. This architectural shift introduces new risks from various aspects such as infrastructure dependency, cognitive bottlenecks in supervision, objective alignment and so on. In this section, we outline several key open challenges that must be systematically examined to ensure the scalability, robustness, and institutional sustainability of OPACs, highlighting promising directions for future research.
Implementation Barrier and Stable Foundation Models Access. One of the primary open challenges for OPACs lies in lowering the implementation barrier for non-expert users while enabling increasingly personalized and advanced configurations. There have been early attempts from both the academic and industry, providing platforms to build agent operating system (Agent OS) [135]. For example, IBM Watsonx Orchestrate [357], Google’s Agent Development Kit [358] and Nvidia AI Enterprise [359] provide modular infrastructure for coordination, negotiation, and role-based task delegation. Collectively, these initiatives signal rapid movement toward standardization and operational readiness. Although large models and autonomous agents are becoming more capable, deploying and orchestrating them still requires technical knowledge in system integration, workflow design, API management, and security [360]. Future directions should therefore focus on co-evolving hardware and software ecosystems, including lightweight, plug-and-play devices for local or edge deployment, standardized agent orchestration frameworks, and modular software stacks that allow scalable personalization without requiring deep engineering expertise. It is also important to investigate seamless configuration of user (human founder), environment, and agentic AI system. Bridging this gap will be essential for transforming OPAC from a technically demanding prototype into a widely adoptable production paradigm.
Another critical challenge for OPACs is ensuring a stable and reliable supply of foundation models. Many current systems rely heavily on commercial APIs, which introduces risks related to pricing changes, rate limits, service instability, policy shifts, and data governance constraints. As a result, there is a growing trend toward local deployment of smaller foundation models (e.g., lightweight open-source models running on personal devices such as Mac Studio) to enhance stability, safety, and privacy. However, this shift introduces clear trade-offs: only open-source models can be used, model capability may lag behind state-of-the-art proprietary systems, and local inference often suffers from slower execution speed and hardware limitations. Techniques such as model distillation, quantization, and task-specific fine-tuning can partially mitigate these constraints, but balancing autonomy, performance, cost, and capability remains an open research direction for scalable OPAC infrastructure.
Long-period Semantic Drift and Memory Governance. OPACs depend heavily on agent memory mechanisms to retain business context, user preferences, decision histories, and accumulated experience. Unlike traditional organizations, where institutional memory is distributed across teams, documentation, and organizational routines, OPAC memory is largely embedded within model-mediated storage and retrieval processes. Over time, such systems may suffer from drift or degradation due to iterative summarization, shifting contexts, and evolving objectives. reasoning accuracy can drop substantially over longer timescales [361]. In the context of OPAC settings, memory drift can have amplified consequences. Because decision authority is highly centralized in a small number of human supervisors, errors embedded in agent memory may propagate across multiple tasks—affecting pricing strategies, client communication, compliance judgments, or long-term planning. Without redundant human cross-checking structures, gradual misalignment may remain unnoticed until it produces material business risks. Although early research has begun to examine long-term reliability and agent drift phenomena, systematic longitudinal validation remains limited [362]. Future work should carefully explore continuous auditing and correction mechanisms that maintain memory fidelity over time, ensuring that accumulated experience enhances—rather than silently undermines—the reliability and strategic consistency of OPAC operations.
Optimizing Human Control Under Cognitive Ceilings in LLM-Agent Systems. how to optimally allocate and structure human supervisory control under bounded human cognition. Unlike traditional firms, OPACs cannot scale governance by adding managerial layers; instead, a single human founder remains the ultimate decision authority over an MAS. This makes human cognitive bandwidth a structural constraint on scalability, reliability, and risk management. Evaluating this constraint is itself a research problem. Classical work such as the Cummings fan-out model suggests that one operator can supervise four to five autonomous units requiring judgment [325], yet these findings originate from UAV supervision contexts and may not directly generalize to heterogeneous LLM-agent systems. In OPAC settings, supervisory tasks impose uneven cognitive demands. Approving a routine output generated by an agent requires limited mental effort, whereas validating contract clauses or financial decisions demands substantially higher cognitive engagement. Therefore, the effective fan-out limit is not solely determined by the number of agents, but by the distribution and intensity of cognitive load across tasks. Future research should therefore focus on developing optimized supervisory strategies under limited cognitive capacity. This involves modeling cognitive load in agent-mediated workflows and designing control mechanisms that allocate human attention efficiently across tasks of varying complexity. Addressing this problem is essential for ensuring both the scalability and reliability of OPAC systems.
Setting and Aligning Development Objectives. A central challenge for OPAC development is how to define the business objectives that optimization should serve. Existing agent-development methods often begin after the target has already been specified, as in prompt optimization or pipeline optimization [290,292]. This assumption is too strong for OPACs. The owner may know the desired business direction, but that direction still has to be translated into objectives that agents can act on and improve against. If the objective is too vague, agent development lacks a usable feedback signal. If it is reduced to a convenient metric, agents may improve the metric while weakening the business. OPACs, therefore, need mechanisms for turning owner intent into an objective structure that separates non-negotiable constraints from targets that can be optimized.
A consequent challenge is that such objectives must remain valid as the OPAC develops. Owner intent may change in response to market feedback, new risks, or emerging expansion opportunities. Meanwhile, agents may have already accumulated experience, revised routines, or changed workflows under an earlier objective. This creates a risk that old objectives continue to guide new operations through memory, skills, or workflow templates. Future research should therefore examine how OPACs can revise business objectives during operations and whether accumulated agent behavior continues to align with the current objective structure. This requires more than setting a reward or benchmark at the beginning; it requires mechanisms for objective revision, conflict detection, and verification that local agent improvements continue to contribute to OPAC performance.
Governance Boundaries and Human–Agent Authority Challenge. OPACs introduce a structural tension between human authority and agent autonomy. When a large portion of operational decisions is delegated to autonomous agents with only minimal human supervision, questions of responsibility attribution become increasingly complex. If an agent produces financial loss, legal violations, or reputational harm, it remains unclear how liability should be allocated among the human founder, the system designer, the model provider, or the deployment infrastructure. Beyond legal liability, internal authority boundaries also become blurred. Determining which decisions can be safely delegated and which must remain under explicit human control is not merely a technical issue, but a governance problem.
Crucially, defining delegation boundaries is not only about compliance but also about organizational performance. Although agents can efficiently execute standardized, repetitive, and programmable tasks, full automation does not necessarily maximize long-term profitability or trust. Certain domains, such as customer service, complex strategic judgment, and relationship management, may require human-in-the-loop involvement to provide contextual understanding, ethical oversight, emotional intelligence, and authentic interaction. Over-delegation risks loss of accountability and regulatory exposure, while under-delegation limits scalability and efficiency. Existing corporate governance frameworks were designed for human organizational hierarchies and may not directly translate to agent-dominant operational structures. Future research should therefore develop responsibility allocation models, delegation taxonomies, and compliance mechanisms tailored to hybrid human–agent enterprises. Systematic investigation of the timing, degree, and form of human involvement across business processes will be essential to identify optimal human–agent configurations that balance autonomy, accountability, scalability, and profitability in One-person Agentic Companies.

7. Conclusion

This survey has examined the emerging literature on the One-Person Agentic Company (OPAC), a novel organizational form in which a single founder or a small founding team leads a scalable system of AI agents to perform enterprise-level work. We position OPAC as a new business architecture that addresses the long-standing trade-off between individual control and organizational scale. We formalized OPAC from both technical and managerial perspectives and identified four defining characteristics: flexible human-led agent organization, expanded agent-related resources, faster organizational iteration, and trustworthiness as a central governance challenge. We further compared OPACs with traditional companies across organization, cost, efficiency, development, authority, responsibility, and risk. Building on this conceptual foundation, we reviewed OPAC-related studies systematically from four perspectives: organization and operation, resource management, development and iteration, and trustworthiness and safeguards. This survey contributes by providing a unified framework for understanding OPACs, synthesizing fragmented research across AI, multi-agent systems, human-agent interaction, and management. It also identifies key challenges for future work. We hope this survey can provide a foundation for future research on scalable and trustworthy human-led agentic enterprises.

References

  1. Brynjolfsson, E.; Li, D.; Raymond, L. Generative AI at work. Q. J. Econ. 2025, 140, 889–942. [Google Scholar] [CrossRef]
  2. Filippucci, F.; Gal, P.; Jona-Lasinio, C.; Leandro, A.; Nicoletti, G. The impact of Artificial Intelligence on productivity, distribution and growth: Key mechanisms, initial evidence and policy challenges. OECD Artificial Intelligence Papers, 2024. [Google Scholar]
  3. Yu, Z.; Fu, Y.; He, Z.; Huang, Y.; Yiu, L.K.; Fang, M.; Luo, W.; Wang, J. From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company. arXiv 2026, arXiv:2604.22446. [Google Scholar]
  4. Rezazadeh, F.; Bonehgazy, P. Digital Co-Founders: Transforming Imagination into Viable Solo Business via Agentic AI. arXiv 2025, arXiv:2511.09533. [Google Scholar]
  5. Healy, J.; Nicholson, D.; Pekarek, A. Should we take the gig economy seriously? Labour Ind. A J. Soc. Econ. Relat. Work 2017, 27, 232–248. [Google Scholar] [CrossRef]
  6. Vallas, S.; Schor, J.B. What do platforms do? Understanding the gig economy. Annu. Rev. Sociol. 2020, 46, 273–294. [Google Scholar] [CrossRef]
  7. Wu, D.; Huang, J.L. Gig work and gig workers: An integrative review and agenda for future research. J. Organ. Behav. 2024, 45, 183–208. [Google Scholar] [CrossRef]
  8. Wang, Y.; Shen, X.; Han, Y.; Backes, M.; Chen, P.Y.; Ho, T.Y. OrgAgent: Organize Your Multi-Agent System like a Company. arXiv 2026, arXiv:2604.01020. [Google Scholar]
  9. Yin, S.; Fu, C.; Zhao, S.; Li, K.; Sun, X.; Xu, T.; Chen, E. A survey on multimodal large language models. Natl. Sci. Rev. 2024, 11, nwae403. [Google Scholar] [CrossRef] [PubMed]
  10. Yuan, X.; Zhou, L.; Sun, Z.; Zhou, Z.; Lan, J. Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model. In Proceedings of the AAAI Conference on Artificial Intelligence, 2025; pp. 9725–9733. [Google Scholar]
  11. Joseph, J.; Sengul, M. Organization design: Current insights and future research directions. J. Manag. 2025, 51, 249–308. [Google Scholar] [CrossRef]
  12. Volberda, H.W. Toward the flexible form: How to remain vital in hypercompetitive environments. Organ. Sci. 1996, 7, 359–374. [Google Scholar] [CrossRef]
  13. Csaszar, F.A. An efficient frontier in organization design: Organizational structure as a determinant of exploration and exploitation. Organ. Sci. 2013, 24, 1083–1101. [Google Scholar] [CrossRef]
  14. Lovallo, D.; Brown, A.L.; Teece, D.J.; Bardolet, D. Resource re-allocation capabilities in internal capital markets: The value of overcoming inertia. Strateg. Manag. J. 2020, 41, 1365–1380. [Google Scholar] [CrossRef]
  15. Devarakonda, S.V.; Goossen, M.C.; Mulotte, L. The allocation of resource control within the corporate structure: Evidence from post-acquisition patent reassignments. Strateg. Manag. J. 2024. [Google Scholar] [CrossRef]
  16. Aguinis, H.; Kraiger, K. Benefits of training and development for individuals and teams, organizations, and society. Annu. Rev. Psychol. 2009, 60, 451–474. [Google Scholar] [CrossRef] [PubMed]
  17. Huselid, M.A. The impact of human resource management practices on turnover, productivity, and corporate financial performance. Acad. Manag. J. 1995, 38, 635–672. [Google Scholar] [CrossRef]
  18. Lazear, E.P. Performance pay and productivity. Am. Econ. Rev. 2000, 90, 1346–1361. [Google Scholar] [CrossRef]
  19. Flynn, F.J.; Amanatullah, E.T. Psyched up or psyched out? The influence of coactor status on individual performance. Organ. Sci. 2012, 23, 402–415. [Google Scholar] [CrossRef]
  20. Latham, G.P.; Seijts, G.H. The effects of proximal and distal goals on performance on a moderately complex task. Journal of Organizational Behavior: The International Journal of Industrial, Occupational and Organizational Psychology and Behavior 1999, 20, 421–429. [Google Scholar] [CrossRef]
  21. Borch, C. Machine Learning, Knowledge Risk, and Principal-Agent Problems in Automated Trading. Technol. Soc. 2022, 68, 101852. [Google Scholar] [CrossRef]
  22. Bussgang, J.J.; Quinn, T.; Amit, S. Base44: A One-Person AI Company Picks a Path, 2026.
  23. Benbya, H.; Pachidi, S.; Jarvenpaa, S. Special issue editorial: Artificial intelligence in organizations: Implications for information systems research. J. Assoc. Inf. Syst. 2021, 22, 10. [Google Scholar] [CrossRef]
  24. Huang, J.t.; Zhou, J.; Jin, T.; Zhou, X.; Chen, Z.; Wang, W.; Yuan, Y.; Lyu, M.R.; Sap, M. On the resilience of llm-based multi-agent collaboration with faulty agents. arXiv 2024, arXiv:2408.00989. [Google Scholar]
  25. Glikson, E.; Woolley, A.W. Human trust in artificial intelligence: Review of empirical research. Acad. Manag. Ann. 2020, 14, 627–660. [Google Scholar] [CrossRef]
  26. Vanneste, B.S.; Puranam, P. Artificial intelligence, trust, and perceptions of agency. Acad. Manag. Rev. 2024, amr–2022. [Google Scholar] [CrossRef]
  27. Yue, Y.; Zhang, G.; Liu, B.; Wan, G.; Wang, K.; Cheng, D.; Qi, Y. Masrouter: Learning to route llms for multi-agent systems. Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025, Volume 1, 15549–15572. [Google Scholar] [CrossRef]
  28. Ferber, J.; Gutknecht, O.; Michel, F. From agents to organizations: an organizational view of multi-agent systems. In Proceedings of the International workshop on agent-oriented software engineering, 2003; Springer; pp. 214–230. [Google Scholar]
  29. Mao, W.; Liu, H.; Tan, H.; Shi, Y.; Wu, J.; Zhang, A.; Wang, X. Joint Optimization of Multi-agent Memory System. arXiv 2026, arXiv:2603.12631. [Google Scholar]
  30. Li, A.; Xie, Y.; Li, S.; Tsung, F.; Ding, B.; Li, Y. Agent-oriented planning in multi-agent systems. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 19495–19517. [Google Scholar]
  31. Wang, Y.; Lu, Y. Interaction, Process, Infrastructure: A Unified Framework for Human-Agent Collaboration. arXiv 2025, arXiv:2506.11718. [Google Scholar]
  32. Zou, H.P.; Huang, W.C.; Wu, Y.; Chen, Y.; Miao, C.; Nguyen, H.; Zhou, Y.; Zhang, W.; Fang, L.; He, L.; et al. Llm-based human-agent collaboration and interaction systems: A survey. arXiv 2025, arXiv:2505.00753. [Google Scholar]
  33. Zhuge, M.; Wang, W.; Kirsch, L.; Faccio, F.; Khizbullin, D.; Schmidhuber, J. Gptswarm: Language agents as optimizable graphs. In Proceedings of the Forty-first International Conference on Machine Learning, 2024. [Google Scholar]
  34. Hu, Y.; Cai, Y.; Du, Y.; Zhu, X.; Liu, X.; Yu, Z.; Hou, Y.; Tang, S.; Chen, S. Self-evolving multi-agent collaboration networks for software development. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 23007–23039. [Google Scholar]
  35. Raza, S.; Sapkota, R.; Karkee, M.; Emmanouilidis, C. Trism for agentic ai: A review of trust, risk, and security management in llm-based agentic multi-agent systems. AI Open, 2026. [Google Scholar]
  36. Jin, C.; Peng, H.; Zhang, Q.; Tang, Y.; Metaxas, D.N.; Che, T. Two heads are better than one: Test-time scaling of multi-agent collaborative reasoning. arXiv 2025, arXiv:2504.09772. [Google Scholar]
  37. Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. Autogen: Enabling next-gen LLM applications via multi-agent conversations. In Proceedings of the First conference on language modeling, 2024. [Google Scholar]
  38. Zhang, W.; Cui, C.; Zhao, Y.; Hu, R.; Liu, Y.; Zhou, Y.; An, B. Agentorchestra: A hierarchical multi-agent framework for general-purpose task solving. arXiv E-Prints 2025, arXiv–2506. [Google Scholar]
  39. Yang, Y.; Chai, H.; Shao, S.; Song, Y.; Qi, S.; Rui, R.; Zhang, W. Agentnet: Decentralized evolutionary coordination for llm-based multi-agent systems. Adv. Neural Inf. Process. Syst. 2026, 38, 107309–107336. [Google Scholar]
  40. Lu, S.; Shao, J.; Luo, B.; Lin, T. Morphagent: Empowering agents through self-evolving profiles and decentralized collaboration. arXiv 2024, arXiv:2410.15048. [Google Scholar]
  41. Liu, Z.; Zhang, Y.; Li, P.; Liu, Y.; Yang, D. A dynamic LLM-powered agent network for task-oriented agent collaboration. In Proceedings of the First Conference on Language Modeling, 2024. [Google Scholar]
  42. Hong, S.; Zhuge, M.; Chen, J.; Zheng, X.; Cheng, Y.; Wang, J.; Zhang, C.; Wang, Z.; Yau, S.K.S.; Lin, Z.; et al. MetaGPT: Meta programming for a multi-agent collaborative framework. In Proceedings of the The twelfth international conference on learning representations, 2023. [Google Scholar]
  43. Qian, C.; Liu, W.; Liu, H.; Chen, N.; Dang, Y.; Li, J.; Yang, C.; Chen, W.; Su, Y.; Cong, X.; et al. Chatdev: Communicative agents for software development. Proceedings of the Proceedings of the 62nd annual meeting of the association for computational linguistics 2024, volume 1, 15174–15186. [Google Scholar] [CrossRef]
  44. Han, B.; Zhang, S. Exploring advanced llm multi-agent systems based on blackboard architecture. arXiv 2025, arXiv:2507.01701. [Google Scholar]
  45. Salemi, A.; Parmar, M.; Goyal, P.; Song, Y.; Yoon, J.; Zamani, H.; Pfister, T.; Palangi, H. Llm-based multi-agent blackboard system for information discovery in data science. arXiv 2025, arXiv:2510.01285. [Google Scholar]
  46. Li, S.; Liu, Y.; Wen, Q.; Zhang, C.; Pan, S. Assemble your crew: Automatic multi-agent communication topology design via autoregressive graph generation. Proc. Proc. AAAI Conf. Artif. Intell. 2026, Vol. 40, 23142–23150. [Google Scholar] [CrossRef]
  47. Pappu, A.; El, B.; Cao, H.; di Nolfo, C.; Sun, Y.; Cao, M.; Zou, J. Multi-Agent Teams Hold Experts Back. ICML 2026. [Google Scholar] [CrossRef]
  48. Chen, W.; Su, Y.; Zuo, J.; Yang, C.; Yuan, C.; Chan, C.M.; Yu, H.; Lu, Y.; Hung, Y.H.; Qian, C.; et al. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors. In Proceedings of the The Twelfth International Conference on Learning Representations, 2023. [Google Scholar]
  49. Chen, G.; Dong, S.; Shu, Y.; Zhang, G.; Sesay, J.; Karlsson, B.F.; Fu, J.; Shi, Y. Autoagents: A framework for automatic agent generation. arXiv 2023, arXiv:2309.17288. [Google Scholar]
  50. Horling, B.; Lesser, V. A survey of multi-agent organizational paradigms. Knowl. Eng. Rev. 2004, 19, 281–316. [Google Scholar] [CrossRef]
  51. Xu, Y.; Hu, J.; Zhao, Z.; Duan, Z.; Sun, X.; Yang, X. Multiagentesc: A llm-based multi-agent collaboration framework for emotional support conversation. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; pp. 4665–4681. [Google Scholar]
  52. Li, G.; Hammoud, H.; Itani, H.; Khizbullin, D.; Ghanem, B. Camel: Communicative agents for mind exploration of large language model society. Adv. Neural Inf. Process. Syst. 2023, 36, 51991–52008. [Google Scholar] [CrossRef]
  53. Swanson, K.; Wu, W.; Bulaong, N.L.; Pak, J.E.; Zou, J. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature 2025, 646, 716–723. [Google Scholar] [PubMed]
  54. Zhao, W.; Yuksekgonul, M.; Wu, S.; Zou, J. Sirius: Self-improving multi-agent systems via bootstrapped reasoning. Adv. Neural Inf. Process. Syst. 2026, 38, 124475–124504. [Google Scholar]
  55. Masters, C.; Vellanki, A.; Shangguan, J.; Kultys, B.; Gilmore, J.; Moore, A.; Albrecht, S. Orchestrating Human-AI Teams: The Manager Agent as aUnifying Research Challenge. In Proceedings of the Proceedings of the 2025 7th International Conference on Distributed Artificial Intelligence, 2025; pp. 91–107. [Google Scholar]
  56. Zhang, M.; Kim, J.; Xiang, S.; Gao, J.; Cao, C. Dynamic Role Assignment for Multi-Agent Debate. arXiv 2026, arXiv:2601.17152. [Google Scholar]
  57. Zhou, H.; Geng, H.; Xue, X.; Kang, L.; Qin, Y.; Wang, Z.; Yin, Z.; Bai, L. Reso: A reward-driven self-organizing llm-based multi-agent system for reasoning tasks. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; pp. 15990–16009. [Google Scholar]
  58. Green, B. The flaws of policies requiring human oversight of government algorithms. Comput. Law. Secur. Rev. 2022, 45, 105681. [Google Scholar] [CrossRef]
  59. Bansal, G.; Wu, T.; Zhou, J.; Fok, R.; Nushi, B.; Kamar, E.; Ribeiro, M.T.; Weld, D. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the Proceedings of the 2021 CHI conference on human factors in computing systems, 2021; pp. 1–16. [Google Scholar]
  60. Bansal, G.; Nushi, B.; Kamar, E.; Weld, D.S.; Lasecki, W.S.; Horvitz, E. Updates in human-ai teams: Understanding and addressing the performance/compatibility tradeoff. Proc. Proc. AAAI Conf. Artif. Intell. 2019, Vol. 33, 2429–2437. [Google Scholar] [CrossRef]
  61. Parasuraman, R.; Riley, V. Humans and automation: Use, misuse, disuse, abuse. Hum. Factors 1997, 39, 230–253. [Google Scholar] [CrossRef]
  62. Kaber, D.B.; Endsley, M.R. The effects of level of automation and adaptive automation on human performance, situation awareness and workload in a dynamic control task. Theor. Issues Ergon. Sci. 2004, 5, 113–153. [Google Scholar] [CrossRef]
  63. Vaccaro, M.; Almaatouq, A.; Malone, T. When combinations of humans and AI are useful: A systematic review and meta-analysis. Nat. Hum. Behav. 2024, 8, 2293–2303. [Google Scholar] [CrossRef] [PubMed]
  64. Chiodo, M.; Müller, D.; Siewert, P.; Wetherall, J.L.; Yasmine, Z.; Burden, J. Formalising human-in-the-loop: Computational reductions, failure modes, and legal-moral responsibility. arXiv 2025, arXiv:2505.10426. [Google Scholar]
  65. Mozannar, H.; Bansal, G.; Tan, C.; Fourney, A.; Dibia, V.; Chen, J.; Gerrits, J.; Payne, T.; Maldaner, M.K.; Grunde-McLaughlin, M.; et al. Magentic-ui: Towards human-in-the-loop agentic systems. arXiv 2025, arXiv:2507.22358. [Google Scholar]
  66. Yang, L.; Weng, Y. ResearStudio: A Human-intervenable Framework for Building Controllable Deep Research Agents. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2025; pp. 896–905. [Google Scholar]
  67. He, Y.; Li, R.; Chen, A.; Liu, Y.; Chen, Y.; Sui, Y.; Chen, C.; Zhu, Y.; Luo, L.; Yang, F.; et al. Enabling self-improving agents to learn at test time with human-in-the-loop guidance. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, 2025; pp. 1625–1653. [Google Scholar]
  68. Zhou, J. OrchVis: Hierarchical Multi-Agent Orchestration for Human Oversight. arXiv 2025, arXiv:2510.24937. [Google Scholar]
  69. Wang, C.L.; Singhal, T.; Kelkar, A.; Tuo, J. MI9: An Integrated Runtime Governance Framework for Agentic AI. arXiv 2025, arXiv:2508.03858. [Google Scholar]
  70. Huq, F.; Wang, Z.Z.; Xu, F.F.; Ou, T.; Zhou, S.; Bigham, J.P.; Neubig, G. Cowpilot: a framework for autonomous and human-agent collaborative web navigation. In Proceedings of the Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations), 2025; pp. 163–172. [Google Scholar]
  71. Syros, G.; Suri, A.; Ginesin, J.; Nita-Rotaru, C.; Oprea, A. Saga: A security architecture for governing ai agentic systems. arXiv 2025, arXiv:2504.21034. [Google Scholar]
  72. Shao, Y.; Samuel, V.; Jiang, Y.; Yang, J.; Yang, D. Collaborative gym: A framework for enabling and evaluating human-agent collaboration. arXiv 2024, arXiv:2412.15701. [Google Scholar]
  73. Wang, X.; Jiang, Z.; Xiong, Y.; Liu, A. Human-LLM collaboration in generative design for customization. J. Manuf. Syst. 2025, 80, 425–435. [Google Scholar] [CrossRef]
  74. Bai, H.; Zhou, Y.; Cemri, M.; Pan, J.; Suhr, A.; Levine, S.; Kumar, A. Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning. Adv. Neural Inf. Process. Syst. 2024, 37, 12461–12495. [Google Scholar] [CrossRef]
  75. Pan, L.; Li, Y.; Yu, C.; Shi, Y. A human-computer collaborative tool for training a single large language model agent into a network through few examples. arXiv 2024, arXiv:2404.15974. [Google Scholar]
  76. Han, H.; Hwang, S.w.; Samdani, R.; He, Y. Convcodeworld: Benchmarking conversational code generation in reproducible feedback environments. arXiv 2025, arXiv:2502.19852. [Google Scholar]
  77. Zhang, H.; Du, W.; Shan, J.; Zhou, Q.; Du, Y.; Tenenbaum, J.B.; Shu, T.; Gan, C. Building cooperative embodied agents modularly with large language models. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 19373–19401. [Google Scholar]
  78. Chang, M.; Chhablani, G.; Clegg, A.; Dallaire Cote, M.; Desai, R.; Hlavac, M.; Karashchuk, V.; Krantz, J.; Mottaghi, R.; Parashar, P.; et al. Partnr: A benchmark for planning and reasoning in embodied multi-agent tasks. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 65205–65268. [Google Scholar]
  79. Zhang, S.; Wang, X.; Zhang, W.; Chen, Y.; Gao, L.; Wang, D.; Zhang, W.; Wang, X.; Wen, Y. Mutual theory of mind in human-ai collaboration: An empirical study with llm-driven ai agents in a real-time shared workspace task. arXiv 2024, arXiv:2409.08811. [Google Scholar]
  80. Zhang, X.; Deng, Y.; Ren, Z.; Ng, S.K.; Chua, T.S. Ask-before-plan: Proactive language agents for real-world planning. Proc. Find. Assoc. Comput. Linguist. EMNLP 2024, 2024, 10836–10863. [Google Scholar] [CrossRef]
  81. Feng, K.K.; Pu, K.; Latzke, M.; August, T.; Siangliulue, P.; Bragg, J.; Weld, D.S.; Zhang, A.X.; Chang, J.C. Cocoa: Co-planning and co-execution with ai agents. In Proceedings of the Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026; pp. 1–23. [Google Scholar]
  82. Shen, S.Z.; Chen, V.; Gu, K.; Ross, A.; Ma, Z.; Ross, J.; Gu, A.; Si, C.; Chi, W.; Peng, A.; et al. Completion is not Collaboration: Scaling Collaborative Effort with Agents. arXiv 2025, arXiv:2510.25744. [Google Scholar]
  83. Piao, Y.; Min, H.; Su, H.; Zhang, L.; Wang, L.; Yin, Y.; Wu, X.; Xu, Z.; Qu, L.; Li, H.; et al. AgentBay: A Hybrid Interaction Sandbox for Seamless Human-AI Intervention in Agentic Systems. arXiv 2025, arXiv:2512.04367. [Google Scholar]
  84. Gao, G.; Taymanov, A.; Salinas, E.; Mineiro, P.; Misra, D. Aligning llm agents by learning latent preference from user edits. Adv. Neural Inf. Process. Syst. 2024, 37, 136873–136896. [Google Scholar] [CrossRef]
  85. Yang, R.; Ye, F.; Li, J.; Yuan, S.; Tu, Z.; Li, X.; Yang, D.; et al. The lighthouse of language: Enhancing llm agents via critique-guided improvement. Adv. Neural Inf. Process. Syst. 2026, 38, 164647–164678. [Google Scholar]
  86. Torreno, A.; Onaindia, E.; Komenda, A.; Štolba, M. Cooperative multi-agent planning: A survey. ACM Comput. Surv. (CSUR) 2017, 50, 1–32. [Google Scholar] [CrossRef]
  87. Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; Zhuang, Y. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Adv. Neural Inf. Process. Syst. 2023, 36, 38154–38180. [Google Scholar] [CrossRef]
  88. Wang, Y.; Wu, Z.; Yao, J.; Su, J. Tdag: A multi-agent framework based on dynamic task decomposition and agent generation. Neural Netw. 2025, 185, 107200. [Google Scholar] [CrossRef] [PubMed]
  89. Yu, J.; Ding, Y.; Sato, H. Dyntaskmas: A dynamic task graph-driven framework for asynchronous and parallel llm-based multi-agent systems. Proc. Proc. Int. Conf. Autom. Plan. Sched. 2025, Vol. 35, 288–296. [Google Scholar] [CrossRef]
  90. Alzu’bi, S.; Nama, B.; Kaz, A.; Eswaran, A.; Chen, W.; Khetan, S.; Bala, R.; Vu, T.; Oh, S. ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems. arXiv 2026, arXiv:2602.01848. [Google Scholar]
  91. Choi, J.W.; Kim, H.; Ong, H.; Yoon, Y.; Jang, M.; Kim, D.; Kim, J. Reactree: Hierarchical llm agent trees with control flow for long-horizon task planning. arXiv 2025, arXiv:2511.02424. [Google Scholar]
  92. Li, Z.; Chang, Y.; Yu, G.; Le, X. Hiplan: Hierarchical planning for llm-based agents with adaptive global-local guidance. arXiv 2025, arXiv:2508.19076. [Google Scholar]
  93. Kannan, S.S.; Venkatesh, V.L.; Min, B.C. Smart-llm: Smart multi-agent robot task planning using large language models. In Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE, 2024; pp. 12140–12147. [Google Scholar]
  94. Jiang, Z.; Wu, P.; Yuan, X.; Fan, W.; Qing, L. QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering. In Proceedings of the 2025 KDD Cup Workshop for Multimodal Retrieval Augmented Generation, 2026. [Google Scholar]
  95. Liu, S.; Yuan, X.; Chen, T.; Zhan, Z.; Han, Z.; Zheng, D.; Zhang, W.; Cao, S. CASTER: Breaking the Cost-Performance Barrier in Multi-Agent Orchestration via Context-Aware Strategy for Task Efficient Routing. arXiv 2026, arXiv:2601.19793. [Google Scholar]
  96. Jin, W.; Du, H.; Zhao, B.; Tian, X.; Shi, B.; Yang, G. A comprehensive survey on multi-agent cooperative decision-making: Scenarios, approaches, challenges and perspectives. arXiv 2025, arXiv:2503.13415. [Google Scholar]
  97. Wu, P.; Chen, P.Q.; Li, X.; Fan, W.; Li, Q. Datamart-Agent: LLM-Driven Game-Theoretic Agent for Data Marketplace Modeling. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2026, 2026; pp. 32509–32531. [Google Scholar]
  98. Tran, K.T.; Dao, D.; Nguyen, M.D.; Pham, Q.V.; O’Sullivan, B.; Nguyen, H.D. Multi-agent collaboration mechanisms: A survey of llms. arXiv 2025, arXiv:2501.06322. [Google Scholar]
  99. Chen, H.; Ji, W.; Xu, L.; Zhao, S. Multi-agent consensus seeking via large language models. arXiv 2023, arXiv:2310.20151. [Google Scholar]
  100. Zhang, J.; Xu, X.; Zhang, N.; Liu, R.; Hooi, B.; Deng, S. Exploring collaboration mechanisms for llm agents: A social psychology view. arXiv 2023, arXiv:2310.02124. [Google Scholar]
  101. Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J.B.; Mordatch, I. Improving factuality and reasoning in language models through multiagent debate. arXiv 2023, arXiv:2305.14325. [Google Scholar]
  102. Choi, H.K.; Zhu, J.; Li, S. Debate or vote: Which yields better decisions in multi-agent large language models? Adv. Neural Inf. Process. Syst. 2026, 38, 101732–101764. [Google Scholar]
  103. Chan, C.M.; Chen, W.; Su, Y.; Yu, J.; Xue, W.; Zhang, S.; Fu, J.; Liu, Z. Chateval: Towards better llm-based evaluators through multi-agent debate. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 9079–9093. [Google Scholar]
  104. Wang, Y.; Shi, Y.; Yang, M.; Zhang, R.; He, S.; Lian, H.; Chen, Y.; Ye, S.; Cai, K.; Gu, X. SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents. arXiv 2026, arXiv:2601.16746. [Google Scholar]
  105. Kang, M.; Chen, W.N.; Han, D.; Inan, H.A.; Wutschitz, L.; Chen, Y.; Sim, R.; Rajmohan, S. Acon: Optimizing context compression for long-horizon llm agents. arXiv 2025, arXiv:2510.00615. [Google Scholar]
  106. Feng, L.; Yang, F.; Chen, F.; Cheng, X.; Xu, H.; Wan, Z.; Yan, M.; An, B. AgentOCR: Reimagining Agent History via Optical Self-Compression. arXiv 2026, arXiv:2601.04786. [Google Scholar]
  107. Guo, Y.; Yang, W.; Sun, Z.; Ding, N.; Liu, Z.; Lin, Y. Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  108. Wu, H.; Wang, X.; Zhang, J.; Tong, J.; Chen, X.; Lin, J.; Ma, Y.; Shen, X. UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking. arXiv 2026, arXiv:2602.23734. [Google Scholar]
  109. Liang, J.; Han, J.; Li, W.; Wang, X.; Zhang, Z.; Jiang, Z.; Liao, Y.; Li, T.; Huang, Y.; Shen, H.; et al. GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1. 0). arXiv 2026, arXiv:2604.17091. [Google Scholar]
  110. Du, H.; Dong, Y.; Ning, X. Latent thinking optimization: Your latent reasoning language model secretly encodes reward signals in its latent thoughts. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  111. Kuzina, A.; Pióro, M.; Bejnordi, B.E. KaVa: Latent Reasoning via Compressed KV-Cache Distillation. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  112. Shi, D.; Asi, A.; Li, K.; Yuan, X.; Pan, L.; Lee, W.; Xiao, W. SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  113. Shen, Z.; Yan, H.; Zhang, L.; Hu, Z.; Du, Y.; He, Y. Codi: Compressing chain-of-thought into continuous space via self-distillation. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; pp. 677–693. [Google Scholar]
  114. Li, B.; Sun, X.; Liu, J.; Wang, Z.; Wu, J.; Yu, X.; Chen, H.; Barsoum, E.; Chen, M.; Liu, Z. Latent visual reasoning. arXiv 2025, arXiv:2509.24251. [Google Scholar]
  115. Wang, Q.; Shi, Y.; Wang, Y.; Zhang, Y.; Wan, P.; Gai, K.; Ying, X.; Wang, Y. Monet: Reasoning in latent visual space beyond images and language. arXiv 2025, arXiv:2511.21395. [Google Scholar]
  116. He, Y.; Zheng, W.; Zhu, Y.; Zheng, Z.; Su, L.; Vasudevan, S.; Guo, Q.; Hong, L.; Li, J. SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Google Scholar]
  117. Ye, W.; Liang, Y.; Shan, L. Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization. arXiv 2025, arXiv:2510.04182. [Google Scholar]
  118. Kang, H.; Zhang, Y.; Kuang, N.L.; Majamaki, N.; Jaitly, N.; Ma, Y.A.; Qin, L. Ladir: Latent diffusion enhances llms for text reasoning. arXiv 2025, arXiv:2510.04573. [Google Scholar]
  119. Shen, M.; Li, Y.; Chen, L.; Fan, Z.; Li, Y.; Yang, Q. From mind to machine: The rise of manus ai as a fully autonomous digital agent. arXiv 2025, arXiv:2505.02024. [Google Scholar]
  120. Xu, Y.; Guo, X.; Zeng, Z.; Miao, C. Softcot: Soft chain-of-thought for efficient reasoning with llms. Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025, Volume 1, 23336–23351. [Google Scholar] [CrossRef]
  121. Fan, Z.; Chen, K.; Xing, R.; Li, Y.; Jiang, L.; Tian, Z. FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  122. Rafique, M.; Bindschaedler, L. ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents. In Proceedings of the Proceedings of the Sixth European Workshop on Machine Learning and Systems, 2026; pp. 1–12. [Google Scholar]
  123. Shi, Y.; Liu, S.; Yang, Y.; Mao, W.; Chen, Y.; Gu, Q.; Su, H.; Cai, X.; Wang, X.; Zhang, A. MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning. arXiv 2026, arXiv:2601.21468. [Google Scholar]
  124. Fang, J.; Deng, X.; Xu, H.; Jiang, Z.; Tang, Y.; Xu, Z.; Deng, S.; Yao, Y.; Wang, M.; Qiao, S.; et al. LightMem: Lightweight and Efficient Memory-Augmented Generation. arXiv 2025, arXiv:2510.18866. [Google Scholar]
  125. Ji, S.; Chen, X.; Yang, S.; Tao, X.; Wan, P.; Zhao, H. Memflow: Flowing adaptive memory for consistent and efficient long video narratives. arXiv 2025, arXiv:2512.14699. [Google Scholar]
  126. Zhai, Y.; Tao, S.; Chen, C.; Zou, A.; Chen, Z.; Fu, Q.; Mai, S.; Yu, L.; Deng, J.; Cao, Z.; et al. Agentevolver: Towards efficient self-evolving agent system. arXiv 2025, arXiv:2511.10395. [Google Scholar]
  127. Lan, H.; Yu, Y.; Qian, L.; Peng, L.; Wu, J.; Liu, W.; Luan, J.; Bai, T. LightSearcher: Efficient DeepSearch via Experiential Memory. arXiv 2025, arXiv:2512.06653. [Google Scholar]
  128. Zhang, G.; Fu, M.; Yan, S. Memgen: Weaving generative latent memory for self-evolving agents. arXiv 2025, arXiv:2509.24704. [Google Scholar]
  129. Liu, J.; Su, Y.; Xia, P.; Han, S.; Zheng, Z.; Xie, C.; Ding, M.; Yao, H. SimpleMem: Efficient Lifelong Memory for LLM Agents. arXiv 2026, arXiv:2601.02553. [Google Scholar]
  130. Zhang, H.; Yang, S.; Fu, J.; Ng, S.K.; Qiu, X. HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding. arXiv 2026, arXiv:2601.14724. [Google Scholar]
  131. Chen, Y.; Chen, R.; Yi, S.; Zhao, X.; Li, X.; Zhang, J.; Sun, J.; Hu, C.; Han, Y.; Bing, L.; et al. MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens. arXiv 2026, arXiv:2603.23516. [Google Scholar]
  132. Wu, P.; Yu, Z.; Liu, Y.; Wu, C.H.; Zhou, E.; Shen, J. MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
  133. Xiao, C.; Zhang, P.; Han, X.; Xiao, G.; Lin, Y.; Zhang, Z.; Liu, Z.; Sun, M. Infllm: Training-free long-context extrapolation for llms with an efficient context memory. Adv. Neural Inf. Process. Syst. 2024, 37, 119638–119661. [Google Scholar] [CrossRef]
  134. Takerngsaksiri, W.; Pasuksmit, J.; Thongtanunam, P.; Tantithamthavorn, C.; Zhang, R.; Jiang, F.; Li, J.; Cook, E.; Chen, K.; Wu, M. Human-in-the-loop software development agents. In Proceedings of the 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP); IEEE, 2025; pp. 342–352. [Google Scholar]
  135. Lu, J.; Zhang, Z.; Yang, F.; Zhang, J.; Wang, L.; Du, C.; Lin, Q.; Rajmohan, S.; Zhang, D.; Zhang, Q. Axis: Efficient human-agent-computer interaction with api-first llm-based agents. Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025, Volume 1, 7711–7743. [Google Scholar] [CrossRef]
  136. Wang, R.; Han, X.; Ji, L.; Wang, S.; Baldwin, T.; Li, H. ToolGen: Unified Tool Retrieval and Calling via Generation. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  137. Chen, Y.; Yoon, J.; Sachan, D.S.; Wang, Q.; Cohen-Addad, V.; Bateni, M.; Lee, C.Y.; Pfister, T. Re-invoke: Tool invocation rewriting for zero-shot tool retrieval. Proc. Find. Assoc. Comput. Linguist. EMNLP 2024, 2024, 4705–4726. [Google Scholar] [CrossRef]
  138. Yuan, S.; Song, K.; Chen, J.; Tan, X.; Shen, Y.; Ren, K.; Li, D.; Yang, D. Easytool: Enhancing llm-based agents with concise tool instruction. Proceedings of the Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies 2025, Volume 1, 951–972. [Google Scholar] [CrossRef]
  139. Wu, Q.; Liu, W.; Luan, J.; Wang, B. Toolplanner: A tool augmented llm for multi granularity instructions with path planning and feedback. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024; pp. 18315–18339. [Google Scholar]
  140. Qiao, S.; Gui, H.; Lv, C.; Jia, Q.; Chen, H.; Zhang, N. Making language models better tool learners with execution feedback. Proceedings of the Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 2024, Volume 1, 3550–3568. [Google Scholar] [CrossRef]
  141. Patil, S.G.; Zhang, T.; Wang, X.; Gonzalez, J.E. Gorilla: Large language model connected with massive apis. Adv. Neural Inf. Process. Syst. 2024, 37, 126544–126565. [Google Scholar] [CrossRef]
  142. Fei, X.; Zheng, X.; Feng, H. Mcp-zero: Active tool discovery for autonomous llm agents. arXiv 2025, arXiv:2506.01056. [Google Scholar]
  143. Wang, W.; Niu, P.; Xu, Z.; Chen, Z.; Du, J.; Du, Y.; Pang, X.; Huang, K.; Wang, Y.; Yan, Q.; et al. Mcp-flow: Facilitating llm agents to master real-world, diverse and scaling mcp tools. arXiv 2025, arXiv:2510.24284. [Google Scholar]
  144. Tang, Y.; Peng, H.; Zhao, B.; Ding, H.; Song, H.; Wang, T.; Zhong, C.; Gong, J. Human Tool: An MCP-Style Framework for Human-Agent Collaboration. arXiv 2026, arXiv:2602.12953. [Google Scholar]
  145. Chen, J.; Pan, X.; Yu, D.; Song, K.; Wang, X.; Yu, D.; Chen, J. Skills-in-context: Unlocking compositionality in large language models. Proc. Find. Assoc. Comput. Linguist. EMNLP 2024, 2024, 13838–13890. [Google Scholar] [CrossRef]
  146. Jiang, Y.; Li, D.; Deng, H.; Ma, B.; Wang, X.; Wang, Q.; Yu, G. SoK: Agentic Skills–Beyond Tool Use in LLM Agents. arXiv 2026, arXiv:2602.20867. [Google Scholar]
  147. Chen, S.; Gai, J.; Zhou, R.; Zhang, J.; Zhu, T.; Li, J.; Wang, K.; Wang, Z.; Chen, Z.; Kaleb, K.; et al. SkillCraft: Can LLM Agents Learn to Use Tools Skillfully? arXiv 2026, arXiv:2603.00718. [Google Scholar]
  148. Gao, Y.; Li, Z.; Ji, Z.; Ma, P.; Wang, S.; et al. SkillReducer: Optimizing LLM Agent Skills for Token Efficiency. arXiv 2026, arXiv:2603.29919. [Google Scholar]
  149. Sun, Y.; Wei, P.; Hsieh, L.B. Don’t Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG. arXiv 2026, arXiv:2604.14572. [Google Scholar]
  150. Rezazadeh, A.; Li, Z.; Lou, A.; Zhao, Y.; Wei, W.; Bao, Y. Collaborative memory: Multi-user memory sharing in llm agents with dynamic access control. arXiv 2025, arXiv:2505.18279. [Google Scholar]
  151. Ou, T.; Vaduguru, S.; Fried, D. Analyzing Information Sharing and Coordination in Multi-Agent Planning. arXiv 2025, arXiv:2508.12981. [Google Scholar]
  152. Yu, Y.; Yao, L.; Xie, Y.; Tan, Q.; Feng, J.; Li, Y.; Wu, L. Agentic memory: Learning unified long-term and short-term memory management for large language model agents. arXiv 2026, arXiv:2601.01885. [Google Scholar]
  153. Saleh, A.; Tarkoma, S.; Lindgren, A.; Donta, P.K.; Dustdar, S.; Pirttikangas, S.; Lovén, L. Memindex: Agentic event-based distributed memory management for multi-agent systems. ACM Transactions on Autonomous and Adaptive Systems, 2025. [Google Scholar]
  154. Fioresi, J.; Kulkarni, P.P.; Vayani, A.; Wang, S.; Shah, M. Learning to Share: Selective Memory for Efficient Parallel Agentic Systems. arXiv 2026, arXiv:2602.05965. [Google Scholar]
  155. Zhang, H.; Long, Q.; Bao, J.; Feng, T.; Zhang, W.; Yue, H.; Wang, W. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents. arXiv 2026, arXiv:2602.02474. [Google Scholar]
  156. Zhou, H.; Guo, S.; Liu, A.; Yu, Z.; Gong, Z.; Zhao, B.; Chen, Z.; Zhang, M.; Chen, Y.; Li, J.; et al. Memento-skills: Let agents design agents. arXiv 2026, arXiv:2603.18743. [Google Scholar]
  157. Han, D.; Couturier, C.; Diaz, D.M.; Zhang, X.; Rühle, V.; Rajmohan, S. Legomem: Modular procedural memory for multi-agent llm systems for workflow automation. arXiv 2025, arXiv:2510.04851. [Google Scholar]
  158. Mao, W.; Liu, H.; Liu, Z.; Tan, H.; Shi, Y.; Wu, J.; Zhang, A.; Wang, X. Collaborative Multi-Agent Optimization for Personalized Memory System. arXiv 2026, arXiv:2603.12631. [Google Scholar]
  159. Sung, Y.Y.; Kim, H.; Zhang, D. Verila: A human-centered evaluation framework for interpretable verification of llm agent failures. arXiv 2025, arXiv:2503.12651. [Google Scholar]
  160. Chen, C.; Zhang, Z.; Khalilov, I.; Guo, B.; Gebreegziabher, S.A.; Ye, Y.; Xiao, Z.; Yao, Y.; Li, T.; Li, T.J.J. Toward a human-centered evaluation framework for trustworthy llm-powered gui agents. arXiv 2025, arXiv:2504.17934. [Google Scholar]
  161. Lu, Y.; Yang, S.; Qian, C.; Chen, G.; Luo, Q.; Wu, Y.; Wang, H.; Cong, X.; Zhang, Z.; Lin, Y.; et al. Proactive agent: Shifting llm agents from reactive responses to active assistance. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 47431–47457. [Google Scholar]
  162. Wu, Z.; Huang, H.; Lou, X.; Qu, X.; Cheng, P.; Wu, Z.; Liu, W.; Zhang, W.; Wang, J.; Wang, Z.; et al. Verios: Query-driven proactive human-agent-gui interaction for trustworthy os agents. arXiv 2025, arXiv:2509.07553. [Google Scholar]
  163. Wang, R.; Wu, J.; Xia, Y.; Yu, T.; Rossi, R.A.; McAuley, J.; Yao, L. DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer. arXiv 2025, arXiv:2507.23554. [Google Scholar]
  164. Jiang, Z.; Chen, Y.; Wang, S.; Qu, H.; Jindong, Z.; Fan, W.; Qing, L.; Liang, D.; Wang, J. Atomic Intent Reasoning: Bringing LLM Semantics to Industrial Cross-Domain Recommendations. arXiv 2026, arXiv:2606.10357. [Google Scholar]
  165. Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-rag: Learning to retrieve, generate, and critique through self-reflection. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 9112–9141. [Google Scholar]
  166. Yuan, X.; Ning, L.; Ye, Q.; Fan, W.; Li, Q. mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA. In Proceedings of the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2026; pp. 2274–2285. [Google Scholar]
  167. Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; Narasimhan, K. Tree of thoughts: Deliberate problem solving with large language models. Adv. Neural Inf. Process. Syst. 2023, 36, 11809–11822. [Google Scholar] [CrossRef]
  168. Zhou, A.; Yan, K.; Shlapentokh-Rothman, M.; Wang, H.; Wang, Y.X. Language agent tree search unifies reasoning, acting, and planning in language models. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning, 2024; pp. 62138–62160. [Google Scholar]
  169. Bhardwaj, A.; Ong, Y.; Zahid, E.; Shbita, B. Adaptive Decoding via Test-Time Policy Learning for Self-Improving Generation. In Proceedings of the International Conference on Learning Representations, 2026. [Google Scholar]
  170. Tan, L.S.; Chen, J.; Fu, X.; Ma, L.; Huang, J.; Shi, J.; Li, Y.; Wen, L. Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models. arXiv 2026, arXiv:2603.15724. [Google Scholar]
  171. Wu, P.; Li, X. Dynamic Action Space Reinforcement Learning for Optimal Trading Execution. In Proceedings of the Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026; pp. 1883–1891. [Google Scholar]
  172. Wu, P.; Li, X. Timing Optimization in Dynamic Discrete Action Space Lifelong Reinforcement Learning. In Proceedings of the Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems, 2026; pp. 3310–3312. [Google Scholar]
  173. Jiang, Z.; Chen, Y.; Pan, Y.; Hu, Z.; Fan, W.; Li, Q.; Wang, H.; Wang, J.; Ou, W. A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing. arXiv 2026, arXiv:2608.04625. [Google Scholar]
  174. Yao, S.; Zhao, J.; Yu, D.; Shafran, I.; Narasimhan, K.R.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the NeurIPS 2022 Foundation Models for Decision Making Workshop, 2022. [Google Scholar]
  175. Qi, Z.; Liu, X.; Iong, I.L.; Lai, H.; Sun, X.; Sun, J.; Yang, X.; Yang, Y.; Yao, S.; Xu, W.; et al. WEBRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning. In Proceedings of the International Conference on Learning Representations, 2025. [Google Scholar]
  176. Li, J.; Jin, Y.; Liu, D.; Ding, H.; Wu, J.; Chen, D.; Shen, Y.; Qin, Y.; Tai, Y.; Wang, C.; et al. SE-Search: Self-Evolving Search Agent via Memory and Dense Reward. arXiv 2026, arXiv:2603.03293. [Google Scholar]
  177. Wang, Z.; Xu, H.; Wang, J.; Zhang, X.; Yan, M.; Zhang, J.; Huang, F.; Ji, H. Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks. In Proceedings of the Workshop on Scaling Environments for Agents, 2025. [Google Scholar]
  178. Yang, Y.; Liao, Y.; Mei, J.; Wang, B.; Yang, X.; Wen, L.; Zhang, J.; Li, X.; Chen, H.; Shi, B.; et al. SPIRAL: A Closed-Loop Framework for Self-Improving Action World Models via Reflective Planning Agents. arXiv 2026, arXiv:2603.08403. [Google Scholar]
  179. Zhang, X. Stabilizing Iterative Self-Training with Verified Reasoning via Symbolic Recursive Self-Alignment. In Proceedings of the ICLR 2026 Workshop on Logical Reasoning of Large Language Models, 2026. [Google Scholar]
  180. Fang, W.; Lu, Y.; Liu, S.; Wang, J.; Guo, Z.; He, J.; Tu, F.; Xie, Z. Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement. arXiv 2026, arXiv:2604.14989. [Google Scholar]
  181. Gou, Z.; Shao, Z.; Gong, Y.; Yang, Y.; Duan, N.; Chen, W.; et al. Critic: Large language models can self-correct with tool-interactive critiquing. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 57734–57811. [Google Scholar]
  182. Ye, M.; Zhuang, J.; Xu, M.; Zhang, L.; Ke, G.; Cai, H. Draft-Refine-Optimize: Self-Evolved Learning for Natural Language to MongoDB Query Generation. arXiv 2026, arXiv:2604.13045. [Google Scholar]
  183. Pan, C.; Xu, X.; Xu, Y.; Wu, Y.; Li, S.; Chen, J.; He, C.; Wei, J.; Tan, C. Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora. arXiv 2026, arXiv:2604.24819. [Google Scholar]
  184. Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-refine: Iterative refinement with self-feedback. Adv. Neural Inf. Process. Syst. 2023, 36, 46534–46594. [Google Scholar] [CrossRef]
  185. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. Adv. Neural Inf. Process. Syst. 2023, 36, 8634–8652. [Google Scholar] [CrossRef]
  186. Ge, Y.; Romeo, S.; Cai, J.; Sunkara, M.; Zhang, Y. Samule: Self-learning agents enhanced by multi-level reflection. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; pp. 16602–16621. [Google Scholar]
  187. Marc-Antoine, A.; Teinturier, A.; Xing, V.; Viaud, G. Experiential Reflective Learning for Self-Improving LLM Agents. In Proceedings of the ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems, 2026. [Google Scholar]
  188. Tian, Y.; Peng, B.; Song, L.; Jin, L.; Yu, D.; Han, L.; Mi, H.; Yu, D. Toward self-improvement of llms via imagination, searching, and criticizing. Adv. Neural Inf. Process. Syst. 2024, 37, 52723–52748. [Google Scholar] [CrossRef]
  189. Huang, S.; Zhong, W.; Cai, D.; Wan, F.; Wang, C.; Wang, M.; Qiao, M.; Xu, R. Empowering Self-Learning of LLMs: Inner Knowledge Explicitation as a Catalyst. Proc. Proc. AAAI Conf. Artif. Intell. 2025, Vol. 39, 24150–24158. [Google Scholar] [CrossRef]
  190. Valmeekam, K.; Marquez, M.; Kambhampati, S. Can Large Language Models Really Improve by Self-critiquing Their Own Plans? In Proceedings of the NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023. [Google Scholar]
  191. Park, J.S.; O’Brien, J.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the Proceedings of the 36th annual acm symposium on user interface software and technology, 2023; pp. 1–22. [Google Scholar]
  192. Packer, C.; Wooders, S.; Lin, K.; Fang, V.; Patil, S.G.; Stoica, I.; Gonzalez, J.E. MemGPT: Towards LLMs as Operating Systems. arXiv 2023, arXiv:2310.08560. [Google Scholar]
  193. Zhong, W.; Guo, L.; Gao, Q.; Ye, H.; Wang, Y. MemoryBank: Enhancing Large Language Models with Long-Term Memory. Proc. Proc. AAAI Conf. Artif. Intell. 2024, Vol. 38, 19724–19731. [Google Scholar] [CrossRef]
  194. Chhikara, P.; Khant, D.; Aryan, S.; Singh, T.; Yadav, D. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv 2025, arXiv:2504.19413. [Google Scholar]
  195. Wang, Y.; Krotov, D.; Hu, Y.; Gao, Y.; Zhou, W.; McAuley, J.; Gutfreund, D.; Feris, R.; He, Z. M+: Extending MemoryLLM with Scalable Long-Term Memory. In Proceedings of the International Conference on Machine Learning. PMLR, 2025; pp. 63308–63323. [Google Scholar]
  196. Zhu, S.; Wu, W.; Zhou, K.; Wang, S.; Huang, B. Hybrid Self-Evolving Structured Memory for GUI Agents. arXiv 2026, arXiv:2603.10291. [Google Scholar]
  197. Fang, G.; Isahagian, V.; Jayaram, K.R.; Kumar, R.; Muthusamy, V.; Oum, P.; Thomas, G. Trajectory-Informed Memory Generation for Self-Improving Agent Systems. arXiv 2026, arXiv:2603.10600. [Google Scholar]
  198. Tan, Z.; Yan, J.; Hsu, I.H.; Han, R.; Wang, Z.; Le, L.; Song, Y.; Chen, Y.; Palangi, H.; Lee, G.; et al. Prospect and Retrospect: Reflective Memory Management for Long-Term Personalized Dialogue Agents. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025; pp. 8416–8439. [Google Scholar]
  199. Li, Y.; Wu, D.; Boulet, B. A Training-Free Regeneration Paradigm: Contrastive Reflection Memory Guided Self-Verification and Self-Improvement. arXiv 2026, arXiv:2603.20441. [Google Scholar]
  200. Liang, G.; Bei, Y.; Zhou, S.; Qin, Y.; Zhou, H.; Jia, B.; Li, B.; Bu, J. Generalizable Self-Evolving Memory for Automatic Prompt Optimization. arXiv 2026, arXiv:2603.21520. [Google Scholar]
  201. Pan, Z.; Wu, Y.; Hua, J.; Feng, J.; Yan, S.; Deng, B.; Cao, Z.; Ye, J. Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs. In Proceedings of the International Conference on Learning Representations, 2026. [Google Scholar]
  202. Wang, X.; Liao, N.; Wei, S.; Tang, C.; Xiong, F. AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents. arXiv 2026, arXiv:2603.09716. [Google Scholar]
  203. Jiang, Y.; Yan, R.; Peng, Y.; Li, W.; Wang, T.; Fu, F.; Yuan, B. Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics. arXiv 2026, arXiv:2604.07144. [Google Scholar]
  204. Lin, M.; Zhang, Z.; Lu, H.; Liu, H.; Tang, X.; He, Q.; Zhang, X.; Wang, S. MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution. arXiv 2026, arXiv:2603.18718. [Google Scholar]
  205. Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv 2023, arXiv:2305.16291. [Google Scholar]
  206. Yang, Y.; Li, J.; Pan, Q.; Zhan, B.; Cai, Y.; Du, L.; Zhou, J.; Chen, K.; Chen, Q.; Li, X.; et al. AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution. arXiv 2026, arXiv:2603.01145. [Google Scholar]
  207. Xia, P.; Chen, J.; Wang, H.; Liu, J.; Zeng, K.; Wang, Y.; Han, S.; Zhou, Y.; Zhao, X.; Chen, H.; et al. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning. In Proceedings of the ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving, 2026. [Google Scholar]
  208. Ni, J.; Liu, Y.; Liu, X.; Sun, Y.; Zhou, M.; Cheng, P.; Wang, D.; Jiang, X.; Jiang, G. Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills. arXiv 2026, arXiv:2603.25158. [Google Scholar]
  209. Zhang, Z.; Lu, S.; Qian, H.; He, D.; Liu, Z. AgentFactory: A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse. arXiv 2026, arXiv:2603.18000. [Google Scholar]
  210. Tang, H.; Key, D.; Ellis, K. WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment. Adv. Neural Inf. Process. Syst. 2024, 37, 70148–70212. [Google Scholar] [CrossRef]
  211. Ren, J.; Li, H.; Liu, Y.; Li, T.; Liu, Z.; Liang, Y.; Ge, Z.; Wu, C.; Yuan, X.; Liu, D.; et al. A Blueprint for Self-Evolving Coding Agents in Vehicle Aerodynamic Drag Prediction. arXiv 2026, arXiv:2603.21698. [Google Scholar]
  212. Xie, S.; Zhang, Y.; Wang, R.; Chen, X. Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation. arXiv 2026, arXiv:2603.02623. [Google Scholar]
  213. Zhang, H.; Fan, S.; Zou, H.P.; Chen, Y.; Wang, Z.; Zhou, J.; Li, C.; Huang, W.C.; Yao, Y.; Zheng, K.; et al. EvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification. arXiv 2026, arXiv:2604.01687. [Google Scholar]
  214. Wang, C.; Yu, Z.; Xie, X.; Yao, W.; Fang, R.; Qiao, S.; Cao, K.; Zheng, G.; Qi, X.; Zhang, P.; et al. SkillX: Automatically Constructing Skill Knowledge Bases for Agents. arXiv 2026, arXiv:2604.04804. [Google Scholar]
  215. Cheng, Z.; Liu, Z.; Shan, Y.; Wang, X.; Zhu, X.; Ma, Y.; Wang, H.; Guo, Y.; Lin, W.; Wang, Y. Mem2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation. In Proceedings of the ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving, 2026. [Google Scholar]
  216. Wu, X.; Li, Z.; Shi, G.; Duffy, A.; Marques, T.; Olson, M.L.; Zhou, T.; Manocha, D. Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks. arXiv 2026, arXiv:2604.20987. [Google Scholar]
  217. Xia, B.; Hu, M.; Wang, S.; Jin, J.; Jiao, W.; Lu, Y.; Li, K.; Luo, P. Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent. arXiv 2026, arXiv:2603.05578. [Google Scholar]
  218. Tian, Y.; Chen, J.; Zheng, L.; Tao, M.; Zeng, X.; Yin, Z.; Su, H.; Sun, X. Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO. arXiv 2026, arXiv:2604.27488. [Google Scholar]
  219. Zhang, G.; Yue, Y.; Li, Z.; Yun, S.; Wan, G.; Wang, K.; Cheng, D.; Yu, J.; Chen, T. Cut the crap: An economical communication pipeline for llm-based multi-agent systems. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 75389–75428. [Google Scholar]
  220. Wang, Z.; Wang, Y.; Liu, X.; Ding, L.; Zhang, M.; Liu, J.; Zhang, M. Agentdropout: Dynamic agent elimination for token-efficient and high-performance llm-based multi-agent collaboration. Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025, Volume 1, 24013–24035. [Google Scholar] [CrossRef]
  221. Li, B.; Zhao, Z.; Lee, D.H.; Wang, G. Adaptive graph pruning for multi-agent communication. arXiv 2025, arXiv:2506.02951. [Google Scholar]
  222. Zhou, H.; Wan, X.; Sun, R.; Palangi, H.; Iqbal, S.; Vulić, I.; Korhonen, A.; Arık, S.Ö. Multi-agent design: Optimizing agents with better prompts and topologies. arXiv 2025, arXiv:2502.02533. [Google Scholar]
  223. Ma, X.; Lin, C.; Zhang, Y.; Tresp, V.; Ma, Y. Agentic neural networks: Self-evolving multi-agent systems via textual backpropagation. arXiv 2025, arXiv:2506.09046. [Google Scholar]
  224. Zhang, J.; Xiang, J.; Yu, Z.; Teng, F.; Chen, X.; Chen, J.; Zhuge, M.; Cheng, X.; Hong, S.; Wang, J.; et al. Aflow: Automating agentic workflow generation. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 34040–34077. [Google Scholar]
  225. Hu, S.; Lu, C.; Clune, J. Automated design of agentic systems. Proc. Int. Conf. Learn. Represent. 2025, Vol. 2025, 21344–21377. [Google Scholar]
  226. Wang, Y.; Yang, L.; Li, G.; Wang, M.; Aragam, B. Scoreflow: Mastering llm agent workflows via score-based preference optimization. arXiv 2025, arXiv:2502.04306. [Google Scholar]
  227. Zhang, G.; Chen, K.; Wan, G.; Chang, H.; Cheng, H.; Wang, K.; Hu, S.; Bai, L. Evoflow: Evolving diverse agentic workflows on the fly. arXiv 2025, arXiv:2502.07373. [Google Scholar]
  228. Gao, H.; Liu, Y.; He, Y.; Dou, L.; Du, C.; Deng, Z.; Hooi, B.; Lin, M.; Pang, T. Flowreasoner: Reinforcing query-level meta-agents. arXiv 2025, arXiv:2504.15257. [Google Scholar]
  229. Zhang, G.; Niu, L.; Fang, J.; Wang, K.; Bai, L.; Wang, X. Multi-agent Architecture Search via Agentic Supernet. In Proceedings of the International Conference on Machine Learning. PMLR, 2025; pp. 75834–75852. [Google Scholar]
  230. Ye, R.; Tang, S.; Ge, R.; Du, Y.; Yin, Z.; Chen, S.; Shao, J. MAS-GPT: Training LLMs to Build LLM-based Multi-Agent Systems. In Proceedings of the International Conference on Machine Learning. PMLR, 2025; pp. 72063–72090. [Google Scholar]
  231. Huang, J.; Xu, P.; Nan, X.; Luo, W. Co-evolving Agent Architectures and Interpretable Reasoning for Automated Optimization. arXiv 2026, arXiv:2604.17708. [Google Scholar]
  232. Zhang, G.; Yue, Y.; Sun, X.; Wan, G.; Yu, M.; Fang, J.; Wang, K.; Chen, T.; Cheng, D. G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks. In Proceedings of the International Conference on Machine Learning. PMLR, 2025; pp. 76678–76692. [Google Scholar]
  233. Huang, J.; Zhang, Z.; Shi, K.; Ye, Y.; Zhang, C. Evolverouter: Co-evolving routing and prompt for multi-agent question answering. arXiv 2026, arXiv:2604.05149. [Google Scholar]
  234. Feng, X.; Song, X.; Li, L.; Liu, G.; Shao, J. SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents. arXiv 2026, arXiv:2604.07791. [Google Scholar]
  235. He, S.; Wang, R.; Du, Z.; Bai, H.; Cao, Z.; Cheng, Y.; Zheng, B. Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization. arXiv 2026, arXiv:2604.20714. [Google Scholar]
  236. Hong, S.; Zhuge, M.; Chen, J.; Zheng, X.; Cheng, Y.; Wang, J.; Zhang, C.; Yau, S.; Lin, Z.; Zhou, L.; et al. MetaGPT: Meta programming for a multi-agent collaborative framework. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 23247–23275. [Google Scholar]
  237. Chen, W.; Su, Y.; Zuo, J.; Yang, C.; Yuan, C.; Chan, C.M.; Yu, H.; Lu, Y.; Hung, Y.H.; Qian, C.; et al. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 20094–20136. [Google Scholar]
  238. Lu, C.; Lu, C.; Lange, R.T.; Foerster, J.; Clune, J.; Ha, D. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv 2024, arXiv:2408.06292. [Google Scholar]
  239. Huo, D.; Liu, H.; Liu, G.; Qi, D.; Sun, Z.; Gao, M.; He, J.; Yang, Y.; Chang, X.; Xiong, F.; et al. ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents. arXiv 2026, arXiv:2604.10096. [Google Scholar]
  240. Wu, X.; Zhuan, Y.; Wei, R.; Chen, H.; Bai, D.; Liu, J.; Wang, X.; Wang, X.; Wang, L.; Cheng, X. AgenticRecTune: Multi-Agent with Self-Evolving Skillhub for Recommendation System Optimization. arXiv 2026, arXiv:2604.26969. [Google Scholar]
  241. Lin, Z.; Shen, S.; Kulikov, I.; Shang, J.; Weston, J.; Nie, Y. Learning to solve and verify: A self-play framework for code and test generation. arXiv 2025, arXiv:2502.14948. [Google Scholar]
  242. Barone, A.V.M.; Nok, P.T. Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification. arXiv 2026, arXiv:2604.17010. [Google Scholar]
  243. Chen, J.; Zhang, B.; Ma, R.; Wang, P.; Liang, X.; Tu, Z.; Li, X.; Wong, K.Y.K. Spc: Evolving self-play critic via adversarial games for llm reasoning. Adv. Neural Inf. Process. Syst. 2026, 38, 139228–139254. [Google Scholar]
  244. Huang, C.; Chou, S.Y.; Zhang, Z.; Cardie, C. Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text. arXiv 2026, arXiv:2604.20051. [Google Scholar]
  245. Zhou, Y.; Levine, S.; Weston, J.; Li, X.; Sukhbaatar, S. Self-challenging language model agents. Adv. Neural Inf. Process. Syst. 2025, 38, 113959–113991. [Google Scholar]
  246. Peng, Y.; Zhu, X.; Wei, C.; Zeng, N.; Wang, L.; He, Y.T.; Yu, F.R. Sage: Multi-agent self-evolution for llm reasoning. arXiv 2026, arXiv:2603.15255. [Google Scholar]
  247. Huang, C.; Yu, W.; Wang, X.; Zhang, H.; Li, Z.; Li, R.; Huang, J.; Mi, H.; Yu, D. R-Zero: Self-Evolving Reasoning LLM from Zero Data. In Proceedings of the The 5th Workshop on Mathematical Reasoning and AI at NeurIPS, 2025, 2025. [Google Scholar]
  248. Zhao, A.; Wu, Y.; Wu, T.; Xu, Q.; Yue, Y.; Lin, M.; Wang, S.; Wu, Q.; Zheng, Z.; Huang, G. Absolute zero: Reinforced self-play reasoning with zero data. Adv. Neural Inf. Process. Syst. 2026, 38, 105816–105879. [Google Scholar]
  249. Yang, Z.; Shen, W.; Li, C.; Chen, R.; Wan, F.; Yan, M.; Quan, X.; Huang, F. Spell: Self-play reinforcement learning for evolving long-context language models. arXiv 2025, arXiv:2509.23863. [Google Scholar]
  250. Feng, X.; Yin, D.; Feng, X.; Jiang, Y.; Qin, L.; Ye, Y.; Huang, L.; Ma, W.; Li, Q.; Gu, Y.; et al. Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play. arXiv 2026, arXiv:2604.17696. [Google Scholar]
  251. Yu, W.; Liang, Z.; Huang, C.; Panaganti, K.; Fang, T.; Mi, H.; Yu, D. Guided self-evolving llms with minimal human supervision. arXiv 2025, arXiv:2512.02472. [Google Scholar]
  252. Wang, S.; Jiao, Z.; Zhang, Z.; Peng, Y.; Ze, X.; Yang, B.; Wang, W.; Wei, H.; Zhang, L. Socratic-zero: Bootstrapping reasoning via data-free agent co-evolution. arXiv 2025, arXiv:2509.24726. [Google Scholar]
  253. Bailey, L.; Wen, K.; Dong, K.; Hashimoto, T.; Ma, T. Scaling Self-Play with Self-Guidance. arXiv 2026, arXiv:2604.20209. [Google Scholar]
  254. Jana, S.; Sancaktar, C.; Daniš, T.; Martius, G.; Orvieto, A.; Kolev, P. GASP: Guided Asymmetric Self-Play For Coding LLMs. In Proceedings of the ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving, 2026. [Google Scholar]
  255. Gottweis, J.; Weng, W.H.; Daryin, A.; Tu, T.; Palepu, A.; Sirkovic, P.; Myaskovsky, A.; Weissenberger, F.; Rong, K.; Tanno, R.; et al. Towards an AI co-scientist. arXiv 2025, arXiv:2502.18864. [Google Scholar]
  256. Tan, Y.; Zhang, L.; Li, M.; Yu, Y.; Zhong, B.; Zhou, B.; Dong, N.; Hong, L. Self-evolving AI agents for protein discovery and directed evolution. arXiv 2026, arXiv:2603.27303. [Google Scholar]
  257. Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; et al. Agentbench: Evaluating llms as agents. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 52989–53046. [Google Scholar]
  258. Zhou, S.; Xu, F.F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; et al. Webarena: A realistic web environment for building autonomous agents. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 15585–15606. [Google Scholar]
  259. Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T.J.; Cheng, Z.; Shin, D.; Lei, F.; et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Adv. Neural Inf. Process. Syst. 2024, 37, 52040–52094. [Google Scholar] [CrossRef]
  260. Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. Swe-bench: Can language models resolve real-world github issues? Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 54107–54157. [Google Scholar]
  261. Yang, J.; Jimenez, C.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; Press, O. Swe-agent: Agent-computer interfaces enable automated software engineering. Adv. Neural Inf. Process. Syst. 2024, 37, 50528–50652. [Google Scholar] [CrossRef]
  262. Mialon, G.; Fourrier, C.; Wolf, T.; LeCun, Y.; Scialom, T. Gaia: a benchmark for general ai assistants. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 9025–9049. [Google Scholar]
  263. Jiang, Z.; Yuan, X.; Qu, H.; Lin, S.; Liu, K.; Fan, W.; Qing, L. SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 2165–2175. [Google Scholar]
  264. Jiang, Z.; Wu, P.; Liang, Z.; Chen, P.Q.; Yuan, X.; Jia, Y.; Tu, J.; Li, C.; Ng, P.H.F.; Li, Q. HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning. Proceedings of the Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining 2025, V.2, 5505–5515. [Google Scholar] [CrossRef]
  265. Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; Cong, X.; Tang, X.; Qian, B.; et al. Toolllm: Facilitating large language models to master 16000+ real-world apis. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 9695–9717. [Google Scholar]
  266. Yao, S.; Shinn, N.; Razavi, P.; Narasimhan, K.R. τ-bench: A benchmark for Tool-Agent-User interaction in real-world domains. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  267. Xu, F.F.; Song, Y.; Li, B.; Tang, Y.; Jain, K.; Bao, M.; Wang, Z.; Zhou, X.; Guo, Z.; Cao, M.; et al. Theagentcompany: benchmarking llm agents on consequential real world tasks. Adv. Neural Inf. Process. Syst. 2026, 38. [Google Scholar]
  268. Tai, C.; Zheng, Z.; Long, H.; Wu, H.; Long, Z.; Xiang, H.; Shi, R.; Cui, Z.; Zhang, S.; Qiu, G.; et al. Seed2Scale: A Self-Evolving Data Engine for Embodied AI via Small to Large Model Synergy and Multimodal Evaluation. arXiv 2026, arXiv:2603.08260. [Google Scholar]
  269. Li, Z.; Du, H.; Huang, C.; Wu, X.; Yu, L.; He, Y.; Xie, J.; Wu, X.; Liu, Z.; Zhang, J.; et al. Mm-zero: Self-evolving multi-model vision language models from zero data. arXiv 2026, arXiv:2603.09206. [Google Scholar]
  270. Heng, Y.; Jiang, C.; Yang, H.; Zhang, S.; Ye, W. EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations. arXiv 2026, arXiv:2604.18320. [Google Scholar]
  271. Zhang, K.; Chen, X.; Liu, B.; Xue, T.; Liao, Z.; Liu, Z.; Wang, X.; Ning, Y.; Chen, Z.; Fu, X.; et al. Agent learning via early experience. arXiv 2025, arXiv:2510.08558. [Google Scholar]
  272. Wu, Z.; Shi, K.; Zhang, C.; Liao, Z.; Yang, J.; Yang, N.; Peng, Q.; Zhang, L.; Xu, H.; Su, T.; et al. When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning. arXiv 2026, arXiv:2603.21289. [Google Scholar]
  273. Ouyang, S.; Yan, J.; Hsu, I.; Chen, Y.; Jiang, K.; Wang, Z.; Han, R.; Le, L.T.; Daruki, S.; Tang, X.; et al. Reasoningbank: Scaling agent self-evolving with reasoning memory. arXiv 2025, arXiv:2509.25140. [Google Scholar]
  274. Wu, R.; Wang, X.; Mei, J.; Cai, P.; Fu, D.; Yang, C.; Wen, L.; Yang, X.; Shen, Y.; Wang, Y.; et al. Evolver: Self-evolving llm agents through an experience-driven lifecycle. arXiv 2025, arXiv:2510.16079. [Google Scholar]
  275. Cai, Z.; Guo, X.; Pei, Y.; Feng, J.; Su, J.; Chen, J.; Zhang, Y.Q.; Ma, W.Y.; Wang, M.; Zhou, H. Flex: Continuous agent evolution via forward learning from experience. arXiv 2025, arXiv:2511.06449. [Google Scholar]
  276. Yang, Y.; Liu, T.; Zhu, W.B.; Shi, T.; Song, L.; Jia, R. Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks. arXiv 2026, arXiv:2604.11610. [Google Scholar]
  277. Kim, K.; Kang, M.; Kim, T.; Yang, Y.; Ren, M.; Hwang, S.J. Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents. arXiv 2026, arXiv:2604.14004. [Google Scholar]
  278. Lin, Z.; Liu, F.; Yang, Y.; Lyu, J.; Gao, Y.; Liu, Y.; Lu, Z.; Yu, Y.; Yang, M.; Li, J.; et al. UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience. arXiv 2026, arXiv:2603.24533. [Google Scholar]
  279. Wei, B.; Xia, Z.; Liu, D.; Zhou, X.; Lin, Z.; Wang, Y. ELITE: Experiential Learning and Intent-Aware Transfer for Self-improving Embodied Agents. arXiv 2026, arXiv:2603.24018. [Google Scholar]
  280. Wang, C.; Zhao, H.; Chen, Y. InferenceEvolve: Towards Automated Causal Effect Estimators through Self-Evolving AI. arXiv 2026, arXiv:2604.04274. [Google Scholar]
  281. Wang, L.; Huang, M.; Dragut, E. CentaurTA Studio: A Self-Improving Human-Agent Collaboration System for Thematic Analysis. arXiv 2026, arXiv:2604.18589. [Google Scholar]
  282. She, B.; Chen, B.; Guo, L.; Li, F. PFAgent: A Tractable and Self-Evolving Power-Flow Agent for Interactive Grid Analysis. arXiv 2026, arXiv:2604.10846. [Google Scholar]
  283. Kang, H.; Suresh, T.; Saad-Falcon, J.; Mirhoseini, A. TRACE: Capability-Targeted Agentic Training. arXiv 2026, arXiv:2604.05336. [Google Scholar]
  284. Yang, C.; Xiang, Z.; Tang, Y.; Teng, Z.; Huang, C.; Long, F.; Liu, Y.; Su, J. TTCS: Test-Time Curriculum Synthesis for Self-Evolving. arXiv 2026, arXiv:2601.22628. [Google Scholar]
  285. Chen, X.; Lu, J.; Kim, M.; Zhang, D.; Tang, J.; Pich’e, A.; Gontier, N.; Bengio, Y.; Kamalloo, E. Self-evolving curriculum for llm reasoning. arXiv 2025, arXiv:2505.14970. [Google Scholar]
  286. Zhang, E.; Yan, X.; Lin, W.; Zhang, T.; Qianchun, L. Learning like humans: Advancing llm reasoning capabilities via adaptive difficulty curriculum learning and expert-guided self-reformulation. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; pp. 6630–6644. [Google Scholar]
  287. Mi, T.; Shan, D.; Huang, Z.; Qin, Y.; Xie, M.; Qiao, Y.; Liu, Y.; Zhou, C.; Liu, P. Data Darwinism Part II: DataEvolve–AI can Autonomously Evolve Pretraining Data Curation. arXiv 2026, arXiv:2603.14420. [Google Scholar]
  288. Wang, Z.; Zhang, Z.; Li, Y.; Cheng, Y.; Qu, L.; Xu, Z. CoTEvol: Self-Evolving Chain-of-Thoughts for Data Synthesis in Mathematical Reasoning. arXiv 2026, arXiv:2604.14768. [Google Scholar]
  289. Zelikman, E.; Wu, Y.; Mu, J.; Goodman, N. Star: Bootstrapping reasoning with reasoning. Adv. Neural Inf. Process. Syst. 2022, 35, 15476–15488. [Google Scholar] [CrossRef]
  290. Yang, C.; Wang, X.; Lu, Y.; Liu, H.; Le, Q.V.; Zhou, D.; Chen, X. Large language models as optimizers. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 12028–12068. [Google Scholar]
  291. Fernando, C.; Banarse, D.; Michalewski, H.; Osindero, S.; Rocktäschel, T. Promptbreeder: self-referential self-improvement via prompt evolution. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning, 2024; pp. 13481–13544. [Google Scholar]
  292. Khattab, O.; Singhvi, A.; Maheshwari, P.; Zhang, Z.; Santhanam, K.; Haq, S.; Sharma, A.; Joshi, T.; Moazam, H.; Miller, H.; et al. DSPy: compiling declarative language model calls into state-of-the-art pipelines. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 54928–54958. [Google Scholar]
  293. He, T.; Chen, Y.; Jiang, K.; Lee, K.Y.; Zhou, K.; Shao, K.; Wang, S. EE-MCP: Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning. arXiv 2026, arXiv:2604.09815. [Google Scholar]
  294. Dong, G.; Lu, J.; Huang, J.; Zhong, W.; Liu, L.; Huang, S.; Li, Z.; Zhao, Y.; Song, X.; Li, X.; et al. Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence. arXiv 2026, arXiv:2604.18292. [Google Scholar]
  295. Shi, Y.; Liang, Z.; Panaganti, K.; Yu, D.; Yu, W.; Mi, H. Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis. arXiv 2026, arXiv:2605.14392. [Google Scholar]
  296. Ma, Y.J.; Liang, W.; Wang, G.; Huang, D.A.; Bastani, O.; Jayaraman, D.; Zhu, Y.; Fan, J.; et al. Eureka: Human-level reward design via coding large language models. Proc. Int. Conf. Learn. Represent. 2024, Vol. 2024, 26516–26560. [Google Scholar]
  297. Li, Q.; Zhang, Y.; Yang, X.; Yang, X.; Wang, Z.; Liu, W.; Bian, J. FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents. arXiv 2026, arXiv:2603.01712. [Google Scholar]
  298. Zhang, J.; Gu, Y.; Ruan, J.; Song, M.; Peng, Y.; Han, Z.; Xiang, J.; Wang, Z.; Yang, C.; Ouyang, Y.; et al. Harnessing Agentic Evolution. arXiv 2026, arXiv:2605.13821. [Google Scholar]
  299. Zhang, W.; Zhao, Z.; Wen, H.; Wu, Y.; Guo, C.; Yin, M.; An, B.; Wang, M. Autogenesis: A Self-Evolving Agent Protocol. arXiv 2026, arXiv:2604.15034. [Google Scholar]
  300. Weidinger, L.; Uesato, J.; Rauh, M.; Griffin, C.; Huang, P.S.; Mellor, J.; Glaese, A.; Cheng, M.; Balle, B.; Kasirzadeh, A.; et al. Taxonomy of Risks posed by Language Models, New York, NY, USA; 2022; Volume FAccT ’22, pp. 214–229. [Google Scholar]
  301. National Institute of Standards and Technology. Technical Report NIST AI 600-1, NIST Trustworthy and Responsible AI; Artificial intelligence risk management framework: Generative artificial intelligence profile. Gaithersburg, MD, USA, 2024.
  302. Raji, I.D.; Smart, A.; White, R.N.; Mitchell, M.; Gebru, T.; Hutchinson, B.; Smith-Loud, J.; Theron, D.; Barnes, P. Closing the AI accountability gap: defining an end-to-end framework for internal algorithmic auditing. Proceedings of the Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, New York, NY, USA 2020, FAT* ’20, 33–44. [Google Scholar]
  303. International Organization for Standardization. ISO/IEC 42001:2023; Artificial Intelligence – Management System. 2023.
  304. Santoni de Sio, F.; Van den Hoven, J. Meaningful human control over autonomous systems: A philosophical account. Front. Robot. AI 2018, 5, 323836. [Google Scholar] [CrossRef] [PubMed]
  305. Bainbridge, L. Ironies of Automation. Automatica 1983, 19, 775–779. [Google Scholar] [CrossRef]
  306. Amershi, S.; Weld, D.; Vorvoreanu, M.; Fourney, A.; Nushi, B.; Collisson, P.; Suh, J.; Iqbal, S.; Bennett, P.N.; Inkpen, K.; et al. Guidelines for Human-AI Interaction. In Proceedings of the Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, New York, NY, USA, 2019; pp. 1–13. [Google Scholar]
  307. Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. In Proceedings of the Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA, 2016; pp. 1135–1144. [Google Scholar]
  308. Turpin, M.; Michael, J.; Perez, E.; Bowman, S.R. Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting; Red Hook, NY, USA, 2023; Volume NIPS ’23. [Google Scholar]
  309. Xiong, M.; Hu, Z.; Lu, X.; Li, Y.; Fu, J.; He, J.; Hooi, B. Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024; p. 2306.13063. [Google Scholar]
  310. Schemmer, M.; Kuehl, N.; Benz, C.; Bartos, A.; Satzger, G. Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. In Proceedings of the Proceedings of the 28th International Conference on Intelligent User Interfaces, 2023; ACM; pp. 410–422. [Google Scholar]
  311. Tutek, M.; Chaleshtori, F.H.; Marasović, A.; Belinkov, Y. Measuring chain of thought faithfulness by unlearning reasoning steps. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025; pp. 9946–9971. [Google Scholar]
  312. Tian, K.; Mitchell, E.; Zhou, A.; Sharma, A.; Rafailov, R.; Yao, H.; Finn, C.; Manning, C.D. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 2023; pp. 5433–5442. [Google Scholar]
  313. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On Calibration of Modern Neural Networks. In Proceedings of the Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 2017; Proceedings of Machine Learning Research. Vol. 70, pp. 1321–1330. [Google Scholar]
  314. Ovadia, Y.; Fertig, E.; Ren, J.; Nado, Z.; Sculley, D.; Nowozin, S.; Dillon, J.V.; Lakshminarayanan, B.; Snoek, J. Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift. In Proceedings of the Advances in Neural Information Processing Systems, Red Hook, NY, USA, 2019; Vol. 32. [Google Scholar]
  315. Ma, S.; et al. Understanding the Effects of Miscalibrated AI Confidence on User Trust and Reliance in AI-Assisted Decision Making. arXiv 2025, arXiv:2402.07632. [Google Scholar]
  316. Fregosi, C.; Vicente, L.; Campagner, A.; Cabitza, F. Too Sure for Our Own Good: A User Study on AI Confidence and Human Reliance. Proc. Proc. AAAI Conf. Artif. Intell. 2026, Vol. 40, 17445–17453. [Google Scholar] [CrossRef]
  317. Zhang, Y.; Liao, Q.V.; Bellamy, R.K.E. Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-assisted Decision Making. In Proceedings of the Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 2020; ACM; pp. 295–305. [Google Scholar]
  318. Kocielnik, R.; Amershi, S.; Bennett, P.N. Will You Accept an Imperfect AI? Exploring Designs for Adjusting End-User Expectations of AI Systems. In Proceedings of the Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019; pp. 1–14. [Google Scholar]
  319. Fanous, A.; Goldberg, J.; Agarwal, A.; Lin, J.; Zhou, A.; Xu, S.; Bikia, V.; Daneshjou, R.; Koyejo, S. SycEval: Evaluating LLM Sycophancy. Proc. AAAI/ACM Conf. AI Ethics Soc. 2025, 8, 893–900. [Google Scholar] [CrossRef]
  320. Guo, Y.; et al. ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs. In Proceedings of the International Conference on Learning Representations (ICLR), 2026. [Google Scholar]
  321. Kaur, A. Echoes of Agreement: Argument Driven Sycophancy in Large Language Models. Proc. Find. Assoc. Comput. Linguist. EMNLP 2025, 2025, 22803–22812. [Google Scholar] [CrossRef]
  322. Gaube, S.; Suresh, H.; Raue, M.; Merritt, A.; Berkowitz, S.J.; Lermer, E.; Coughlin, J.F.; Guttag, J.V.; Colak, E.; Ghassemi, M. Do as AI Say: Susceptibility in Deployment of Clinical Decision-aids. npj Digit. Med. 2021, 4. [Google Scholar] [CrossRef] [PubMed]
  323. Danry, V.; Pataranutaporn, P.; Groh, M.; Epstein, Z. Deceptive Explanations by Large Language Models Lead People to Change Their Beliefs about Misinformation More Often than Honest Explanations. In Proceedings of the Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025; ACM; pp. 1–31. [Google Scholar]
  324. Chandra, K.; et al. Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians. 2026. [Google Scholar] [CrossRef]
  325. Cummings, M.L.; Nehme, C.E.; Crandall, J.; Mitchell, P. Predicting operator capacity for supervisory control of multiple UAVs. In Innovations in Intelligent Machines-1; Springer, 2007; pp. 11–37. [Google Scholar]
  326. Breslow, L.A.; Gartenberg, D.; McCurry, J.M.; Trafton, J.G. Dynamic Operator Overload: A Model for Predicting Workload During Supervisory Control. IEEE Trans. Hum.-Mach. Syst. 2014, 44, 30–40. [Google Scholar] [CrossRef]
  327. Tariq, S.; Baruwal Chhetri, M.; Nepal, S.; Paris, C. Alert Fatigue in Security Operations Centres: Research Challenges and Opportunities. ACM Comput. Surv. 2025, 57, 1–38. [Google Scholar] [CrossRef]
  328. Bahodi, M.T.; van Berkel, N.; Skov, M.; Merritt, T. Show Me What’s Wrong: Impact of Explicit Alerts on Novice Supervisors of a Multi-Robot Monitoring System. In Proceedings of the Proceedings of the Second International Symposium on Trustworthy Autonomous Systems, 2024; ACM; pp. 1–17. [Google Scholar]
  329. Ruan, Y.; Dong, H.; Wang, A.; Savarese, S.; Hashimoto, T. Identifying the Risks of LM Agents with an LM-Emulated Sandbox. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2023. [Google Scholar]
  330. OWASP Foundation. OWASP Top 10 for Agentic Applications. 2026. [Google Scholar] [CrossRef]
  331. Liu, Y.; Jia, Y.; Geng, R.; Jia, J.; Gong, N.Z. Formalizing and Benchmarking Prompt Injection Attacks and Defenses. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24), Philadelphia, PA, 2024; pp. 1831–1847. [Google Scholar]
  332. Debenedetti, E.; Zhang, J.; Balunović, M.; Beurer-Kellner, L.; Fischer, M.; Tramèr, F. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. 2024, 2406.13352. [Google Scholar]
  333. Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, 2023; ACM; pp. 79–90, [2302.12173. [Google Scholar]
  334. Khodayari, S.; Zhang, X.; Acharya, B.; Pellegrino, G. Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives. arXiv 2026, arXiv:2604.27202. [Google Scholar]
  335. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 2023, 55, 1–38. [Google Scholar] [CrossRef]
  336. Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.t.; Koh, P.W.; Iyyer, M.; Zettlemoyer, L.; Hajishirzi, H. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023; Association for Computational Linguistics; pp. 12076–12100. [Google Scholar]
  337. Manakul, P.; Liusie, A.; Gales, M. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023; Association for Computational Linguistics; pp. 9004–9017. [Google Scholar]
  338. Vaddi, S.; Vaddi, P. Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs. arXiv 2026, arXiv:2604.19765. [Google Scholar]
  339. Rao, D.; Wong, E.; Callison-Burch, C. Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents. arXiv 2026, arXiv:2604.03173. [Google Scholar]
  340. Zhong, Z.; Huang, Z.; Wettig, A.; Chen, D. Poisoning Retrieval Corpora by Injecting Adversarial Passages. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 2023; pp. 13764–13775. [Google Scholar]
  341. Zou, W.; Geng, R.; Wang, B.; Jia, J. PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. In Proceedings of the 34th USENIX Security Symposium (USENIX Security 25), 2025; USENIX Association; pp. 3827–3844. [Google Scholar]
  342. Chen, Z.; Xiang, Z.; Xiao, C.; Song, D.; Li, B. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2024. [Google Scholar]
  343. Ye, J.; Li, S.; Li, G.; Huang, C.; Gao, S.; Wu, Y.; Zhang, Q.; Gui, T.; Huang, X. ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages. Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics 2024, Volume 1, 2181–2211. [Google Scholar] [CrossRef]
  344. Yuan, T.; He, Z.; Dong, L.; Wang, Y.; Zhao, R.; Xia, T.; Xu, L.; Zhou, B.; Li, F.; Zhang, Z.; et al. R-Judge: Benchmarking Safety Risk Awareness for LLM Agents. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, 2024; Association for Computational Linguistics; pp. 1467–1490. [Google Scholar]
  345. Infocomm Media Development Authority (IMDA). Model AI Governance Framework for Agentic AI. 2026. [Google Scholar] [CrossRef]
  346. Ojewale, V.; Steed, R.; Vecchione, B.; Birhane, A.; Raji, I.D. Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling. In Proceedings of the Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025; ACM; pp. 1–29, [2402.17861. [Google Scholar]
  347. Mökander, J.; Floridi, L. Ethics-based auditing to develop trustworthy AI. Minds Mach. 2021, 31, 323–327. [Google Scholar] [CrossRef]
  348. Cobbe, J.; Lee, M.S.A.; Singh, J. Reviewable Automated Decision-Making: A Framework for Accountable Algorithmic Systems. In Proceedings of the Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, 2021; pp. 598–609. [Google Scholar]
  349. Wieringa, M. What to account for when accounting for algorithms: a systematic literature review on algorithmic accountability. In Proceedings of the Proceedings of the 2020 conference on fairness, accountability, and transparency, 2020; pp. 1–18. [Google Scholar]
  350. Rakova, B.; Yang, J.; Cramer, H.; Chowdhury, R. Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. Proc. ACM Hum.-Comput. Interact. 2021, 5, 1–23. [Google Scholar]
  351. van der Meulen, N.; Jewer, J.; Levallet, N. Minimum Viable Governance for Generative AI. Research briefing; MIT Center for Information Systems Research (CISR), 2026. [Google Scholar]
  352. Hong, Y.; She, Y.; Kang, E.; Timperley, C.S.; Kästner, C. Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility. 2026, 2604.15579. [Google Scholar]
  353. Buçinca, Z.; Malaya, M.B.; Gajos, K.Z. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proc. ACM Hum.-Comput. Interact. 2021, 5, 1–21. [Google Scholar]
  354. Dubois, M.; Ududec, C.; Summerfield, C.; Luettgau, L. Ask don’t tell: Reducing sycophancy in large language models. 2026, 2602.23971. [Google Scholar]
  355. Beigi, M.; Shen, Y.; Shojaee, P.; Wang, Q.; Wang, Z.; Reddy, C.; Jin, M.; Huang, L. Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories. 2025, 2509.16742. [Google Scholar]
  356. DeVerna, M.R.; et al. Designing Meaningful Human Oversight in AI. In AI and Ethics; 2026. [Google Scholar]
  357. IBM. Watsonx Orchestrate. [CrossRef]
  358. Google. Google Agent Development Kit. [CrossRef]
  359. Nvidia. Nvidia AI Enterprise. 2026. [Google Scholar] [CrossRef]
  360. Paperclip, A.I. paperclip: The open-source app everyone uses to manage agents at work. 2026. Available online: https://github.com/paperclipai/paperclip.
  361. Lam, C.; Li, J.; Zhang, L.; Zhao, K. Governing evolving memory in llm agents: Risks, mechanisms, and the stability and safety governed memory (ssgm) framework. arXiv 2026, arXiv:2603.11768. [Google Scholar]
  362. Rath, A. Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions. arXiv 2026, arXiv:2601.04170. [Google Scholar]
Figure 3. Representative static and dynamic organizational structures for multi-agent systems.
Figure 3. Representative static and dynamic organizational structures for multi-agent systems.
Preprints 227744 g003
Figure 5. Overview of the OPAC resource lifecycle and management framework. On the left, the lifecycle includes OPAC’s internal resources and external resources governed by access control mechanisms. On the right, we map fundamental OPAC resources (Token, Memory, and Labor) to their efficiency allocation paradigms.
Figure 5. Overview of the OPAC resource lifecycle and management framework. On the left, the lifecycle includes OPAC’s internal resources and external resources governed by access control mechanisms. On the right, we map fundamental OPAC resources (Token, Memory, and Labor) to their efficiency allocation paradigms.
Preprints 227744 g005
Figure 6. Dimensions of OPAC development across individual, team, and company levels.
Figure 6. Dimensions of OPAC development across individual, team, and company levels.
Preprints 227744 g006
Figure 7. OPAC trustworthiness as a two-domain organizational control problem. Human oversight risks degrade the OPAC owner’s ability to interpret, judge, or allocate attention, while agent execution risks affect business action through corrupted inputs, unreliable knowledge, or excessive authority. Governance foundation spans both domains; execution boundaries primarily address the agent execution risks in Section 5.3; and owner-facing safeguards primarily address the human oversight risks in Section 5.2.
Figure 7. OPAC trustworthiness as a two-domain organizational control problem. Human oversight risks degrade the OPAC owner’s ability to interpret, judge, or allocate attention, while agent execution risks affect business action through corrupted inputs, unreliable knowledge, or excessive authority. Governance foundation spans both domains; execution boundaries primarily address the agent execution risks in Section 5.3; and owner-facing safeguards primarily address the human oversight risks in Section 5.2.
Preprints 227744 g007
Table 1. Comparison Between Traditional Company and OPAC.
Table 1. Comparison Between Traditional Company and OPAC.
Dimensions Traditional Company OPAC
Organization
  • Human-centered structure
  • Relatively stable structure
  • Relatively fixed resource allocation rules
  • Human-led agent structure
  • Flexible structure
  • Dynamic resource allocation
Cost
  • High labor and management costs
  • Space costs
  • Higher agent-related costs (computational cost)
  • Less space costs (digital employee)
Efficiency
  • Limited by human work speed and effort
  • Higher communication and collaboration overhead
  • Faster routine and information-heavy execution
  • Lower internal coordination and communication cost
Development
  • Employee training and organizational change
  • Possible negative effects from internal competition
  • Slow and costly adaptation
  • Agent- and system-level optimization
  • Faster modular iteration
Authorities & Responsibilities
  • Distributed managerial authority
  • Shared organizational responsibility
  • Centralized final authority
  • Delegated agent execution
Risk
  • Human-originated risks
  • Organizational incentive problems
  • Agent-originated risks
  • Technical problems
  • Privacy risk
Table 2. Efficient allocation strategies for internal OPAC resources (token, memory, and labor). The columns summarize representative methods, underlying agent backbones, hardware references, and relevant literature.
Table 2. Efficient allocation strategies for internal OPAC resources (token, memory, and labor). The columns summarize representative methods, underlying agent backbones, hardware references, and relevant literature.
Preprints 227744 i001
Table 3. Access control mechanisms for external resources. The columns summarize strategies for tool calling and memory sharing, and human-agent interaction (HAI) paradigms, detailing representative methods, implementation details, and relevant literature.
Table 3. Access control mechanisms for external resources. The columns summarize strategies for tool calling and memory sharing, and human-agent interaction (HAI) paradigms, detailing representative methods, implementation details, and relevant literature.
Preprints 227744 i001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.