Preprint
Article

This version is not peer-reviewed.

Hybrid Multi-Agent Framework for Enterprise Web Application Generation

Submitted:

17 August 2026

Posted:

19 August 2026

You are already at the latest version

Abstract
There are always repeatable functions like building similar CRUD workflows, implementing authentication layers, role based access control and admin interfaces for various projects on an enterprise web application, and different con-ventions can be used by different developers or code-bases. AI-powered coding tools recently emerged that promise to make this possible, but existing studies have shown that automated LLM-based code generation, vibe coding and free-running multi-agent pipelines have difficulty maintaining consistency between relational schemas, backend APIs, and frontend interfaces as applications grow in size, and become a maintenance night-mare. The main idea is that structural code which can be programmed from an application data model does not need to be probabilistic (only truly ambiguous decisions are language-model decisions such as interpreting the relational semantics, or resolving ambiguous requirements). We build this insight into a system called CodeCraft, where a Prisma schema and its Data Model Meta Format (DMMF) representation are the only source of truth for deterministic generators to generate backend APIs, frontend manifests, RBAC structures and database seeders, while a LangGraph six-agent multi-agent orchestration pipeline (from requirements engineering, schema design, orchestration, backend generation, frontend generation, and containerization) can only be used to clarify requirements, validate schema, and make decisions that cannot be made structurally. Generated systems have a consistent schema representation, as opposed to the prompt-driven tools which are used for each individual file. In all the applications it is demonstrated that schema validation slashes the number of automated iterations required to correct schemas to 2-4 and the end-to-end generation generates deploy-able, containerised applications after requirements and schema approval (no manual effort required). Comparative engineering indicates that the development effort is reduced by 70–80%, compared to manual implementation, but this has not yet been substantiated with a controlled external study.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Most of the time in the initial development of a project is consumed solving the same kinds of problems in enterprise software development: CRUD APIs, authentication flows, DTO validation, role-based access control, forms, tables, menu configuration, and deployment set-up, just to mention a few. As this structural work is recreated on each project, even with the same technology stack, teams end up working in various folder structures, with different validation patterns and API designs, leading to ever more complex code bases that become more difficult to maintain and more difficult to be introduced to over time [19].
This repetition has led to the development of AI coding assistants that enable developers to input the app’s requirements in natural language and let AI write the code. They can be useful for rapid prototyping and isolated components, but show to be useful only in the short-term: relational schemas and backend validation differ from the prompt over iterations [16,17] and correction attempts result in repeated instances of the same incorrect fix, that fail to fix the root cause of the underlying inconsistency [4]. This is addressed by multi-agent LLM pipelines, which break down the generation process into planning, coding, and validation tasks, and demonstrated to boost the quality of reasoning and reliability on correction in generation tasks [5,12,13,18]. However, the same cross-file synchronization problems that plague single-model tools (as main elegant, short business logic code is plagued by single-model cross-file synchronization problems) plague present multi-agent software-engineering platforms, but in more agents.
This is one key area that we’ve personally learned from as we evaluated the tools that were available to us ahead of development to support prompts. A new form and possibly new validation rules for a new backend and new constraints for a new database would have to be imported into the tool and synchronized with the old database and forms, if a schema change is applied to an entity that impacts another entity. Much like conversational code generation in general: small, localized prompt corrections have no impact on the previously generated files’ assumptions and unrelated frontend/backend components fall out of synch after a few iterations [16]. The failure is not in the model capacity, larger or more powerful language models will have the same drift in their output but it is a structural failure: if every file is generated using reasoning on a prompt history as opposed to composing from a single authoritative representation of the application, consistency between files is not guaranteed but only probable.
Having noticed that, when the application’s data model can be mechanically translated into structural code, it is not necessary to do so, the authors propose a deterministic-first multi-agent approach to generating enterprise Web applications. In this paper, we suggest a multi-agent approach, where the agents generate enterprise web applications using a deterministic first approach, based on the observation that if structural code can be mechanically generated from an application’s data model [2] then that is not needed. The basic idea is that a Prisma schema and the Data Model Meta Format (DMMF) representation is the one single source of truth that is used to deterministically compile the back-end API, the front-end manifest, RBAC structure and database seeders, and that a 6-agent orchestration pipeline, made up of a Requirements Engineering Agent, Database Schema Design Agent, Orchestrator Agent, Backend API Generator Agent, Frontend Code Generator Agent, and DevOps Containerization Agent, is only used for tasks that truly require reasoning, such as clarifying ambiguous requirements, validating and repairing schema definitions, and resolving relational semantics that can’t be inferred structurally. The resulting pipeline, as shown in Figure 1 consumes the same validated DMMF artifact for both the backend and frontend generation branches of the pipeline, while the various stages of deterministic compilation are isolated from the agentic reasoning stages.
The structural consistency property, which neither can ensure template-only systems, nor can guarantee fully agentic systems [5,18] (due to schema independence), is supported by the application layers in this system that rely on the same DMMF representation as the downstream generator.
Additional benefits other than uniformity are available with this design. First, it is worth noting that the language-model is not designed to make any non-trivially ambiguous decisions, meaning that there is a much smaller space for hallucinations to impact the generated system: the outputs are not source files, but rather configuration artifacts such as PRDs, schema patches, or frontend manifests that can be validated against a schema or compilation step before any code is written, not trusted without question [2,4]. Second, an application generated by the structural generation can be regenerated after schema modification, without re-running all the previous interactions that led to the previous version of the application, but simply recompiling its representation in the DMMF.Second, the structural generation is deterministic, meaning that an application generated by the structural generation can be regenerated after schema modification without re-running all previous interactions that have led to the previous version of the application, but simply recompiling the application representation in the DMMF. In our proposed framework, consistency is a property of the generation architecture rather than an emergent behavior that the agents must maintain, as a fully agentic pipeline, all outputs (structural or otherwise) are produced by a language model and therefore only as consistent as the language model’s context handling. The system architecture is shown in Figure 1.
This work has been deliberately limited in scope. We’re interested in full stack business applications that have a relational data model, have CRUD workflows, have authentication, and have role based access control; this is the class of apps where things tend to repeat from project to project and deterministic compilation may come in handy. We don’t target applications with a niche and experimental front-end, game development, or systems where the main value of the application is a novel and non-relational interaction pattern, as there is not much schema-driven system code to derive in these systems to begin with. In this sense, two constraints are not incidental, but intentional: there can be remaining ambiguities in the requirements that still need to be clarified in the requirements-engineering agent before the generation process can continue; and as accurate as the schema, so accurate can a deterministic generator be.
Let’s now look at three contributions we make. We introduce a schema-first deterministic generation engine that generates both backend APIs, frontend manifests, RBAC logic and backend seed data from one version of the Prisma DMMF representation, without inconsistencies of model re-definition that occur in unconstrained AI-generated systems [1,2]. Secondly, we design an orchestration pipeline for six agents, which has an explicit execution graph and only lets the language model engage in requirements interpretation, schema validation and ambiguity resolution without directly producing code [5]. Second, we explicitly design our six-agent orchestration pipeline in which the language model is only used in requirement interpretation, schema validation and ambiguity resolution, but not in raw code generation [5]. Thirdly, an end-to-end integration of these two components leads to deployable containerized enterprise applications: For the app that we’ve internally tested, we converged the schemas in 2-4 iterations of corrections, and generated schemas manually without any intervention, once the schema was approved along with the requirements. We also qualitatively compare CodeCraft with popular single-LLM and multi-agent code-generation systems, and estimate that the development effort could be reduced by 70–80% compared to manual implementation (pending a controlled external study [3]).

3. Methodology

3.1. System Overview and Design Rationale

In this section we review the system design and the rationale behind it. In this section, the system is described and the reasons for its design explained. The mission of CodeCraft is to go through the natural language requirements and generate a full stack enterprise web application – relational data models, CRUD workflows, authentication, role based access control. The main design feature is to separate the mechanically derivable features of a data model from the interpretation required: A deterministic generation engine derives directly structural code from a schema representation, with a six-agent orchestration pipeline interpreting requirements, validing schemas and resolving ambiguities. The resulting seven staged pipeline from requirements capture, through to containerized deployment is shown in Figure 1. The design of the deterministic engine (Section 3.2) is described, followed by the design of a multi-agent orchestration layer built on top of this deterministic engine (Section 3.3), the rationale for this separation (Section 3.4) and the implications of this separation in terms of technology stack (Section 3.5).
This division had been due to the failure of both of the strategies in the early stages. A fully prompt-driven generation (with evaluation of the actual schema as explained in Section 1) successfully generated isolated components, but left the frontend form, backend validation, and relational schema drifting away from each other after a few iterations of corrections, as there was no shared representation to constrain the relationship between the different layers. An alternative that was only based on templates did not suffer from this syndrome, but was on the other hand too strict: It could not deal with ambiguous requirements, could not resolve relational ambiguities, e.g., which attribute of an associated entity to show in a generated table, could not pick up up if a schema did not validate. Neither the extremes were allowed: too flexible and too consistent. This led to the idea of two complementary tasks – deterministic compilation, and agentic reasoning – to be seen as part of the same pipeline, rather than as mutually exclusive tasks.
The resulting pipeline executes in seven stages, shown in Figure 1. Initialization begins with a conversational exchange in which the Requirements Engineering Agent clarifies scope, pages, roles, and visual direction, producing an approved Product Requirements Document (PRD). Project Setup creates an isolated workspace with standardized template and specification directories. Schema Design passes the PRD to the Database Schema Design Agent, which synthesizes a Prisma schema and validates it iteratively against compilation and relational constraints until it passes cleanly. Seed Generation derives realistic test data directly from the validated schema’s entities and relationships. Backend Generation compiles NestJS controllers, services, DTOs, and RBAC decorators deterministically from the schema’s Data Model Meta Format (DMMF) representation. Frontend Generation produces a JSON manifest describing pages, navigation, and components, which a deterministic compiler then translates into executable Next.js code. Containerization packages the generated backend, frontend, and database into Docker containers and launches a locally accessible preview. The artifact of the previous stage is validated and used by the current stage as its only input, with no means for the current stage to progress to an artifact that is not validated or partially defined.

3.2. Deterministic Generation Engine

The deterministic engine is required because there is a whole class of code that can be generated deterministically (i.e., the structure of this code is so tied to an application’s data model that it can’t be probabilistically generated). The core of this engine is a schema-processing module which converts a Prisma schema into its Data Model Meta Format (DMMF) representation, a structured JSON document that represents entities, their fields, relationships, enums, and constraints. From this representation, the engine attempts to find the junction tables to identify many-to-many relationships, then it traverses the relationships between entities to generate relational traversal metadata, and finally collects all the annotations in the module-level from all the schema models, for example, ///{"modulePath": "...", "moduleName": "..."}, to assign entities to a consistent module hierarchy downstream generators. This is a processed representation which is consumed by each downstream generator, rather than the raw schema text, so that no generator reparses or re-interprets the schema text.
From the above, meaning of fields, relationships and constraints for an entity cannot be different across various parts of the pipeline (backend, frontend, RBAC, seed generation). This directly overcomes concerns raised in previous research of LLM generated systems [1,2] where the schema was reconstructed at each layer—in the case of SDL forms, the database, or the backend DTOs—using different prompts or generators, and there was no common source of truth to compare with.
The common schema representation is not generated per project, but rather built from a library of golden templates that are common to enterprise applications such as tables, forms, layout, modals, RBAC wrappers, and navigation. This is a generic data-table-generator, that is used everywhere: For each entity, the generator doesn’t generate a specific table implementation, it generates a table-configuration-object with the description of the table, its columns, filters, sort, pagination, etc. and relations (also for other entities), and the front-end-renderer is a common object that reads the table-configuration-object and renders the same table (with different content) for each entity. The approach is to design the configuration, and adding a new entity to an application involves only adding a new configuration, which is defined in terms of the schema, and no code change to the tables; so there is no effort required to implement uniformity of the tables; the tables are uniform in structure and appearance.
The frontend compilation works in a deterministic manner, with the sole exception of the language-model reasoning. Using the DMMF representation, the engine produces a high level JSON file called a manifest that lets you define the structure of forms, entity mapping to tables, navigation rules and how to set up the relationships that are displayed on the form; after that, a deterministic compiler converts the manifest into a full Next.js module. The one step in this pipeline that is assigned to a language model is the resolution of relational display ambiguity, which is the case when, with a Doctor entity and a User entity, the generated table needs to display either User.The fullName field or User.email or any other field that makes sense given the meaning of the data, but not based on the structure of the schema. After this ambiguity is eliminated and documented in the manifest, the rest of the compilation from manifest to executable frontend code is completely deterministic.
What makes this limited advisory role both safe and not just convenient is that the language model doesn’t try to create a full React component, rather it attempts to create a manifest, which is a bounded configuration object that can be validated against a schema before any code is compiled from it, but a full React component generated end-to-end by the language model can be difficult to validate, and can contain structural errors without warning. This follows the general idea, also used below in the generating of backend code and RBAC configuration, of using language models to generate artefacts that a deterministic system can execute, instead of direct generation of final code by the language models themselves [2,4].
The idea of the Backend generation is very similar to DMMF, with the difference that for every entity, the controllers, services, DTOs and CRUD endpoints are generated using NestJS, with minimal involvement from the language-model. Every entity is just stored in a specific module and has a set pattern: The module contains a controller with create, paginated list, single record retrieval, update and delete endpoints, as well as a service layer that processes Prisma queries, and DTOs are not separated out from the entity’s field definitions. Therefore, no two modules are structurally different, no matter what they may be representing, when they are generated. In this stage, there’s a loose coupling in the case of some generated endpoint returning an entity along with other records related to it—where the engine uses heuristic rules to determine the depth of the resulting response payload. These heuristics can be used on schemas of moderate relational complexity and can produce very large or very shallow payloads on schemas that have more densely connected schemas; and the schema does not need to be arbitrarily deep to be general enough for these heuristics to apply, so we revisit this as another natural extension point for agentic reasoning in future work in Section 4.5.
There are two other patterns that are less complex than the above, but deterministic: role-based access control and seed data. The two patterns can be extracted from schema and PRD metadata without ambiguity in the relational model. The role-to-endpoint mapping, route guards, and role-visibility rules for navigation menus are automatically created from the role definitions in the PRD, and applied to both the generated backend decorators and the navigation menu configuration on the frontend, helping to prevent inconsistencies between the two layers. Seed data generation also looks at entities and relationships in the validated schema to generate valid seed data, default roles and demo data without having to involve any sort of agent – the constraints in the schema dictate the shape of valid seed data.

3.3. Multi-Agent Orchestration Pipeline

The deterministic engine in Section 3.2 generates structurally consistent code, but it can by construction not interpret natural-language requirements, deal with business-specific relationships not explicitly stated in a schema, recover from a failed generation step and manage the dependencies between pipeline stages. The only things that need to be reasoned, are these tasks, which is why it is not being compiled in the generators, and has an orchestration layer on top of the deterministic one. This layer would separate this “centralist language model” which has been shown by previous research to forget its context and make inconsistent decisions during longer multi-stage executions [18] with exactly one agent for each stage of the pipeline.
Table 1 describes the roles of the 6 agents and their inputs (and outputs). The Requirements Engineering Agent (REA) carries out a structured conversation with the user, which is similar to an interview, and continues to generate the PRD outlined in Section 3.1. The Database Schema Design Agent (DSDA) creates a Prisma schema from the approved PRD and corrects it based on the feedback from the compiler, until it is valid without warnings. Execution is coordinated by the Orchestrator Agent (OOA): it keeps track of the state of the pipeline, routes the right context slice to the agents downstream and enforces the guardrails outlined below. The Backend API Generator Agent (BAGA) and Frontend Code Generator Agent (FCGA) invoke the deterministic generators of Section 3.2 with the validated DMMF and manifest inputs, intervening only in the case of having to disambiguate between relational entities in the manifest. The DevOps Containerization Agent (DCA) packages up the created backend, frontend, and database into Docker containers and keeps the resulting preview instance on the local machine. The output artifact of each agent is just one kind (either PRD, validated schema, manifest, codebase, or container set), and is the only input to the next agent, so no agent needs to be aware of any particular kind of reasoning process used by other agents.
This is implemented as a directed execution graph within LangGraph, with nodes representing the different agents, and edges representing conditional transitions (e.g. from schema design to seed generation only occurs when the DSDA reports a successful compilation). All agents share a single shared project context object that is passed from agent to agent across the graph that references the current PRD, current schema, current DMMF, artefacts created at previous pipeline steps and current pipeline state. The Orchestrator Agent only passes this relevant portion to each node: DMMF and RBAC metadata to the BAGA; and PRD, DMMF and frontend manifest definitions to the FCGA, limiting the amount of information passed to each agent to what each agent needs for its given job.
This graph-based design was motivated by the difficulty of predicting and debugging the order of tool calls and execution permissions in early development when the agents were able to communicate through free-flowing conversation in this domain (Section 2.4). Since the transitions in a directed graph are explicit, and depend on the validated transitions output by the stage instead of on free-form dialogue the agents exchange, there are no problems of coordination and memory-synchronization that could arise from unclear inter-agent communication, as described in the general multi-agent software-engineering literature [18].
In this graph, there are 3 guardrails that limit the execution. The first one is that the permission validation is done before every agent invocation and the agent can only use the tools for generation for which it holds the permission: for instance if there are artifacts of the backend in the shared context, the FCGA will not be able to invoke the backend-generation tools. Second, human-in-the-loop checkpoints are added at two places where the error thrown down the pipeline is the highest: PRD finalization and schema approval, where users have to approve the pipeline explicitly before it goes on to generation. Third, the DSDA’s schema-validation loop has a maximum number of times that it will loop before throwing an exception, if compilation fails because the problem cannot be resolved, the loop does not go on forever, but the programmer is alerted to the problem and can correct it. A similar mechanism exists for the out of scope boundaries defined in the accepted PRD: requests outside the scope defined in the PRD will not be addressed by the generation agents, Orchestrator Agent will reject any such request before it arrives at a generation agent and will not partially process it.
These guardrails are meant to mitigate a failure mode that goes beyond what the agents can do; (a single agent can be right, but if there is no time, resource or scope of action limitation, a globally undesirable outcome can result). For example, if the DSDA is configured to ’silently fail’ (such as retest the query in an infinite loop), then it is not going to improve the accuracy of the schema correction, it will simply not make the schema correction. Likewise, if the scope is enforced at the orchestration level (as opposed to each generation agent), the feature growth reported in iterative conversational development [4,7] cannot occur, as the scope is not reliant on each generation agent remembering and honoring a boundary several generations before.

3.4. Design Principles

The above-mentioned architecture is based on four design principles. The first is schema as single source of truth: The representation of the schema in the Prisma DMMF should represent the source of truth for all generated artefacts, not for different interpretations by different generators or agents. This principle is applied for design reasons: There’s no code path in CodeCraft where it can occur: Each layer’s understanding of the entity is reconstructed based on the same shared DMMF for the entity. This is not allowed in CodeCraft by construction: each layer’s understanding of the entity is rebuilt based on the same common DMMF for the entity.
The second principle is “Determinism over Language-model reasoning for structural code” that determines which structures are passed to the deterministic generators and which to the reasoning agents. The agents are used only for tasks that need to be interpreted and/or need to resolve ambiguity, such as PRD refinement, schema synthesis and relational display disambiguation; the rest of the structural code is generated from schema metadata. The key to the solution in the compiler domain is to remove the compiler work that can be mechanically derived from the compiler and to bring it to a surface area, where hallucination-driven inconsistencies can occur. In the compiler domain, the important part of the solution is to move away from the predictable work that can be easily extracted from the compiler to a surface area where hallucination-driven inconsistencies can enter [1,2,3].
The third one is that the structured configuration artifacts (PRD, schema definition, frontend manifests) are the only means of interfacing with language-model reasoning and deterministic execution, not code generation directly from the agents themselves. The generator is deterministic, all agents that interact with structural generation (like the Frontend Architect feature of the FCGA), and all back-end generation agents, produce a validated configuration object but not a code artifact. This is why agent output can be checked before it has cascading effects: If the error is discovered in a large codebase, it is hard to trace it, but if the error is discovered in a manifest, then the schema can be used to validate without having to dig through the codebase to find the error [2,4].
The fourth rule, “validate before progress”, means that no pipeline stage start to process an artifact from the previous stage until it has been structurally validated. The most obvious is the DSDA’s iterative schema-validation loop: If a schema is not satisfied by the syntax test or the relational constraints test of Prisma, it is not sent to downstream generators for schema correction, but is corrected in place. The bounded-retry guardrail described in Section 3.3 is not a soft guardrail, as the artifact isn’t bound to the next stage until it is validated at that stage and we know it’s correct.

3.5. Implementation and Technology Stack

The implementation and the technology stack are described below.The implementation and the technology stack are described below. The implementation and the technology stack are detailed in this section. It describes how this section is implemented and what technologies are used. The stack of technologies are organized under the framework, and selected specifically for the ability to generate in a schema-driven and deterministic manner. Due to the file routing nature of Next.js, it was selected for the generated frontend as it will map to the route and page structure generated by the FCGA’s manifest, without being configured at runtime. Besides being, for the most part, lighter alternatives to NestJS (such as Express), the generated backend is decided for due to the method the BAGA generates new modules for its entities – the modules are isolated units, so the NestJS decorator-based modular architecture fits the bill.
This schema-first approach to the design of the whole pipeline is built upon Prisma ORM, where all generators in Section 3.2 receive the structured input, and there are no generators without this abstraction layer that are deterministic. Most of the target features of the framework, such as transactional consistency, relational data models and RBAC hierarchies were obvious when choosing the relational database target, and so was MySQL. The orchestration layer was implemented by LangGraph instead of sequential API calls, as it was required to store and share persistent state between agents that are dependent on each other, and also to control the flow of execution as a graph between steps of agent reasoning, as described in Section 3.3; Gemini API was chosen for reasoning steps because it was required to follow instructions and as part of the reasoning process, share context, which is required to be persistent, between steps of reasoning.
Using two separate job queues which don’t block, and perform generation requests concurrently, Redis and BullMQ can be used to decouple long running generation tasks (schema validation, frontend compilation, dependency installation, container builds), from the main backend process. The reason for choosing Docker for deployment is that generated applications require isolated, reproducible runtimes: The DCA packages are made up of backend, frontend, and database components, all of which are deployed to machines with no manual configuration; these packages are then pushed to a container registry for deployment in the cloud without having to be rebuilt.

4. Results & Discussion

4.1. End-to-End Generation Outcome

We test the end-to-end performance of the CodeCraft by running a number of generation applications we’ve created ourselves across typical enterprise areas, such as an e-commerce catalog and order-management application. In a typical application, like an e-commerce platform that includes product catalog, category management, and order workflows, a developer writes the description of the target application in natural language, and the pipeline passes that description to the REA, which refines it into an approved PRD, passes it to the DSDA, where it is used to synthesize and validate the corresponding schema, passed to the BAGA, which compiles the backend code base from that schema, and passed to the FCGA, which compiles the frontend code base from that schema, and finally to the DCA, which bundles the result into a working application available via a local preview URL, without requiring any code-level intervention. A typical screen shot of one of the generated applications is shown in Figure 4 with category filtering, search and a product grid created entirely from an application manifest generated based on the frontend, as described in Section 3.2.
This outcome is an illustration of the specific claim which led to the design of the CodeCraft; once the PRD and schema are accepted by the user, the following five steps are performed automatically at the code level: all the downstream artifacts are automatically compiled from the approved schema, or are created by agents whose output is automatically validated prior to use. This is a difference that is very distinct from the prompt-based tools described in Section 2.3, which would need manual edits at every iteration as the schema is added inconsistently in the files as part of a similarly sized application. This doesn’t imply that this business logic (as opposed to CRUD, RBAC, and the standard workflows) is incorrect, it simply means that this business logic was not tested for this evaluation, as stated in Section 1, it was observed when the pipeline was completed.

4.2. Observed Pipeline Behavior

The iterative schema-validation loop typically converged to a schema that passed the compilation and relational constraint checking within two to four passes of the “correction” pass in all of the applications tested. The tested applications are summarized in Table 2 with the convergence iterations plotted against schema size for each of them in Figure 5. Schemas that were simpler, with fewer interdependent entities, tended to fall around the low end of this range, and schemas that were more “complex” (more foreign keys, more many-to-many junction tables) sometimes needed more iterations, and occasionally more manual fine-tuning of the prompt description of the relevant entities.
This is a range of properties which the tested applications exhibit, and not guaranteed for other schemas that have more than around twenty entities, or have denser connections between them.
Once the PRD and schema stages passed the validations, the process of generation was finished without any manual effort for most of the tested applications, as expected, based on the description given in Section 4.1. The failures that did occur tended to be one of three types: unexpected schema relationships not covered by the rules in the DSDA; large enough manifests to cause problems in the context handling of the manifest-planning reasoning step; dependency conflicts when building Docker images for the containerization step. None of these failure modes needed to be changed in the code previously generated and none of them required changing the code of the affected stage, which is the design principle of the framework of validating the artifacts before they propagate downstream (Section 3.4).
These observations are made on applications developed by the CodeCraft’s developers, not on a held out benchmark set, a third party test suite, or on applications developed by third parties who are not involved with the CodeCraft. The above convergence and failure-rate statistics should therefore be seen as representative of performance under the conditions in which the CodeCraft was developed and tested and not as a claim to generality for arbitrary enterprise schemas and to unfamiliar users, which we shall address directly in Section 4.5, and which we believe to be the primary target for future evaluation.

4.3. Comparative Positioning

In order to put the CodeCraft in perspective, we qualitatively compare it with four representative tools in four categories that we mentioned in Section 2: GitHub Copilot, v0 are context- and prompt-based single-model assistants, Bolt.new is a full prompt-driven code generator and AgentCoder is a representative multi-agent code generation system. The five aspects of the model, which we consider to be important for enterprise application generation, are compared here summarized in Table 3: the first aspect is input model, either schema-driven or just prompt-based; the second is the generation architecture; the third is the explicit agent-coordination rules; the fourth is the enterprise readiness, and the fifth is traceability of actions. This table is reflective of our own qualitative analysis of published capabilities and architecture of each tool, and not a result of any common benchmark task, common scoring rubric, or controlled experiment conducted under identical conditions, which we mention below.
In fact the only point with which the comparison in this paper makes its case is that the difference between the tools compared is that our tool separates schema-derived production of structure from agent-based reasoning, and this separation is what makes the consistency properties mentioned in Section 4.1 and Section 4.2 possible even if they are not proved.
We do this not to imply that this is experimental proof of superior, but as a positioning device. No common test applications were utilized in Section 4.1 and Section 4.2 and there was no common set of tasks or acceptance criteria defined across all the different tools; the labels of “Yes/No” and the qualitative labels in Table 3 are based upon our understanding of the architecture of each tool, documented in their respective literature; and the results are not measured. The comparison would have to be based on the same application specifications and examine them using the same application evaluation rubric, in a controlled study, of which we have not performed here, but which we recognize as a necessary future study in Section 5.3.

4.4. Estimated Effort Reduction

If such pipeline behaviors are observed in Section 4.1 and Section 4.2, then the pipeline framework is estimated to save around 70–80% repetitive development effort as compared to equivalent enterprise modules developed manually. It was estimated by comparing the time it takes to get the application up and running from start to end, depending on the size of the application, ranging from a few minutes for smaller ones, up to a few hours for larger applications with multiple modules and workflows, which is dominated by schema complexity, frontend manifest size, Docker build time and LLM response latency, and we estimated a few minutes to an hour for similar applications, where we would handcraft the CRUD modules, RBAC integration, web dashboards and deployment configuration.
It is a preliminary engineering estimate, not a statistically-proven result. It has not been undertaken in a controlled study with external developers, it ignores the fact that developers may be faster or slower than others and may differ in their familiarity with the target stack, and it was not tested with a different set of manual implementations that were withheld from the project, but rather implemented by developers who were not the creators of the framework. This estimate is made because it’s a realistic one based on our experience with the effort saved during the development and testing process, but we feel that closing the gap between this estimate and a validated measurement is the #1 priority for future work, which will be discussed in Section 5.3.

4.5. Limitations

The current implementation is based on a fixed technology stack (NestJS backend generation and Next.js frontend generation), which is chosen for the reasons mentioned above, but this limits the deterministic generation engine to not be able to currently generate other common enterprise stacks like Express.js, FastAPI or Vue.js applications, even if the schema is complex.
The endpoints generated are still relational (as described in Section 3.2), meaning that heuristic rules are still used as a temporary workaround for full schema-aware reasoning. We have not tested the heuristics in schemas with a large number of entities or schemas that have a large number of connected entities, where a fixed heuristic is unlikely to be generalizing.The schemas used were not very densely connected, and we have not tested the heuristics in such schemas, or in schemas with many entities. The next natural step is to find a dedicated include-planning agent, where it is explicitly dealt with with a query what relations it needs, which will remove this final heuristic aspect, as described in Section 5.3.
While the DCA can create applications and push them to a container registry, deploy them to AWS ECS or EC2, and verify the deployment, the deployment to AWS is not fully automated in this work as the configuration and the verification steps to complete the deployment were manual procedures.
Due to the external language model (Gemini API) provided, the PRD, schema interpretation or a PRD generated by an agent stage with the same input in natural language may differ from one run to the next. This non-determinism is only present in the reasoning phases: once an artifact (a validated schema or an approved manifest) is created, the deterministic generators of Section 3.2 always generate the same structure for a given prompt, but what is compiled is not deterministic across repeated runs of the system [2].
Last, and most important, all the findings in this section are internal, not from an external user study, but by the testers of the framework. This is the constraint that we feel is most important and must be overcome before these results can be considered final; indeed we feel that all claims in this section, in conjunction with the quantitative and qualitative, require this constraint to be met: the convergence rates of DSDA, the characterization of failure modes, the comparative positioning in Section 4.3, and the estimate of effort-reduction in Section 4.4.

5. Conclusions

In order to solve the problem of structural code generation inconsistency of enterprises that uses language models for end-to-end generation, this paper put forward a deterministic-first multi-agent framework that decouples the structural code generation from the reasoning of agents. The main concept is that code that can be mechanically generated based on the data model for an application (CRUD APIs, DTOs, RBAC decorators, frontend forms and tables) should be compiled deterministically from a single schema representation, whereas a 6-agent orchestration pipeline should be considered for tasks that are truly interpretative, such as clarifying requirements, verifying schema design, and resolving relational ambiguity that can’t be determined in a fixed manner.
It took just a couple of automated schema corrections after the requirements and schema acceptance to deploy and containerize enterprise applications and validated schema in all internal tests. In contrast to existing methods and our own, which separately considered each application layer in turn by individual agents or prompts, the generated applications were not vulnerable to the cross-layer drift we and other works identified as a chronic failure mode of purely prompt-driven and unconstrained multi-agent generation [1,2,4,7].
Overall, this research suggests that deterministic compilation and agentic reasoning are not conflicting methods for software generation using AI, as one method can be used when the other is not, and vice versa. It’s a subtle shift in how a developer’s job is done when using this kind of system, as it’s not about writing and re-syncing boilerplate by hand anymore, it’s about being able to precisely define requirements, architecture and scope, and then relying on an automated pipeline to get the rest.
This scope isn’t a coincidence. The framework is designed for enterprise applications with relational data and CRUD operations, and it wasn’t designed with highly experimental frontends, or with structurally sparse domains, in mind, both of which have very little to gain from a schema-driven engine. There are two unclosed restrictions in that range. Even the relational response depth is still heuristic and not schema-aware control and all the results reported here are from internal tests and not external studies. This is the estimated effort reduction of 70–80% to be interpreted as a preliminary reduction and not validated.
The next logical step is a controlled examination between the tools qualitatively described in Section 4.3 and manual development, in which the time required for development and architectural consistency are measured against a common task set and rubric, with the tools used by developers who have not used the framework. In addition, the following extensions are highlighted: support for additional backend and frontend frameworks, replace the heuristic relation-depth control with a dedicated include-planning agent and run the pipeline against schemas larger than those that were tested here. Each deals with a restriction mentioned previously in this paper, not a new one.

References

  1. “The Hidden Pitfalls of Using LLMs in Software Development,” Metabob. [Online]. Available: https://metabob.com/blogs/hidden-pitfalls-of-using-llms-in-sw. [Accessed: 22-May-2026].
  2. J. Meaden, “COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models,” arXiv, pp. 2–8, 2025.
  3. R. Rawal, “Benchmarking Correctness and Security in Multi-Turn Code Generation,” arXiv, 2025.
  4. S. Bird, “Why 1 in 5 Vibe-Coded Applications Fail Basic Security Tests,” Blott, 5 September 2025. [Online]. Available: https://www.blott.com/blog/post/why-1-in-5-vibe-coded-applications-fail-basic-security-tests. [Accessed: 26-May-2026].
  5. D. Huang, “AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation,” arXiv, 2024.
  6. J. Bossle, “Why AI Coding hits limits as adoption grows.” [Online]. Available: https://www.knowis.com/blog/why-ai-coding-hits-limits-as-adoption-grows. [Accessed: 26-May-2026].
  7. A. A. Abbassi, “ReCatcher: Towards LLMs Regression Testing for Code Generation,” arXiv, 25 July 2025.
  8. X. Wang, “CodeAct: Your LLM Agent Acts Better when Generating Code,” Apple Machine Learning Research, 2024.
  9. X. Wang, “Executable Code Actions Elicit Better LLM Agents,” in International Conference on Machine Learning, 2024.
  10. L. Sbeitan, “Agentic SDLC in practice: the rise of autonomous software delivery,” PwC.
  11. S. Teyssier, “A framework for integrating AI into platform engineering,” 3 December 2025. [Online]. Available: https://platformengineering.org/blog/a-framework-for-integrating-ai-into-platform-engineering. [Accessed: 26-May-2026].
  12. K. Mao, “Blueprint2Code: a multi-agent pipeline for reliable code generation via blueprint planning and repair,” Frontiers in Artificial Intelligence, 2025. [CrossRef]
  13. M. A. Islam, “MapCoder: Multi-Agent Code Generation for Competitive Problem Solving,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024.
  14. M. Kapur, “Vibe, then verify: SonarQube 2025 year in review,” 8 January 2026. [Online]. Available: https://www.sonarsource.com/blog/sonarqube-2025-year-in-review. [Accessed: 26-May-2026].
  15. “Architectural Modernization,” vFunction. [Online]. Available: https://vfunction.com/about/. [Accessed: 26-May-2026].
  16. A. Sarkar, “Vibe coding: programming through conversation with artificial intelligence,” arXiv, 2025.
  17. Nzall, “What is `vibe coding’ and what are the strong points of this methodology?,” Software Engineering Stack Exchange, 2025.
  18. J. He, “LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead,” arXiv, 2024.
  19. L. Rosenfeld, Information Architecture: For the Web and Beyond. USA: O’Reilly Media, 2015.
  20. X. Shen, “Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction,” arXiv, 2025.
  21. S. Author et al., Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging. arXiv preprint arXiv:2505.02133, 2025.
  22. X. Zhang et al., A Unified Debugging Approach via LLM-Based Multi-Agent Synergy. arXiv preprint arXiv:2404.17153, 2024.
  23. J. Doe et al., Testing and Enhancing Multi-Agent Systems for Robust Code Generation. arXiv preprint arXiv:2510.10460, 2025.
  24. A. Smith et al., Build AI agents with fine-tuned open models. Fireworks RFT, 2024.
  25. Y. Chen et al., AutoSafeCoder: A Multi-Agent Framework for Securing LLM Code Generation through Static Analysis and Fuzz Testing. arXiv preprint arXiv:2409.10737, 2024.
  26. AI Agent Architecture Patterns in 2025: The Powerful Way Multi-Agent Architectures Scale. Nexa AI, 2025. [Online]. Available: https://nexaitech.com/multi-ai-agent-architecutre-patterns-for-scale/.
  27. Z. Ma et al., Multi-Agent Systems Integration in Enterprise Environments Using Web Services. IGI Global, 2006.
  28. A. Garcia et al., Engineering multi-agent systems with aspects and patterns. Journal of the Brazilian Computer Society, 2002. [CrossRef]
  29. Z. Chen et al., MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents. arXiv preprint arXiv:2503.01935, 2025.
Figure 1. The figure below shows the system architecture. Below is shown the architecture of a system. The proposed architecture of the framework is presented below. The blue nodes represent probabilistic reasoning agents (LLM), the orange nodes represent deterministic reasoning (compilers) without LLM, and the yellow document shapes are artifacts that are checked before being used in the downstream processes. The validated schema.prisma/DMMF representation (green cylinder) is only read by both the backend (BAGA) and frontend (FCGA) generation branches. The Orchestrator Agent (OOA, purple) controls the routing of the stages, permission guardrails and retry limits and human-in-loop checkpoints (dashed control edges).
Figure 1. The figure below shows the system architecture. Below is shown the architecture of a system. The proposed architecture of the framework is presented below. The blue nodes represent probabilistic reasoning agents (LLM), the orange nodes represent deterministic reasoning (compilers) without LLM, and the yellow document shapes are artifacts that are checked before being used in the downstream processes. The validated schema.prisma/DMMF representation (green cylinder) is only read by both the backend (BAGA) and frontend (FCGA) generation branches. The Orchestrator Agent (OOA, purple) controls the routing of the stages, permission guardrails and retry limits and human-in-loop checkpoints (dashed control edges).
Preprints 228790 g001
Figure 3. The graph that describes the execution of the orchestrator. The Agent stages (blue), deterministic stages (orange) and typed transitions (bold caps) are connected, and the yellow diamonds are human-in-loop checkpoints. It will then go into a bounded automated repair loop, and only be shown to the user when it fails at the actual docker build stage, which will be retried within a bounded time.
Figure 3. The graph that describes the execution of the orchestrator. The Agent stages (blue), deterministic stages (orange) and typed transitions (bold caps) are connected, and the yellow diamonds are human-in-loop checkpoints. It will then go into a bounded automated repair loop, and only be shown to the user when it fails at the actual docker build stage, which will be retried within a bounded time.
Preprints 228790 g003
Figure 4. Sample screen shot of an e-commerce application created. The product grid, category filter and search interface was completely created from the approved frontend manifest, and the supporting API, RBAC guards and seed data was created from a similar DMMF representation. None of the components were handwritten or edited.
Figure 4. Sample screen shot of an e-commerce application created. The product grid, category filter and search interface was completely created from the approved frontend manifest, and the supporting API, RBAC guards and seed data was created from a similar DMMF representation. None of the components were handwritten or edited.
Preprints 228790 g004
Figure 5. The convergence of schema-validation values for the different applications and sizes of schema for the DSDA loop is shown.
Figure 5. The convergence of schema-validation values for the different applications and sizes of schema for the DSDA loop is shown.
Preprints 228790 g005
Table 1. Agent roles, inputs, core processing, and outputs.
Table 1. Agent roles, inputs, core processing, and outputs.
Agent Role Input Core Processing Output
REA Requirements Engineering Agent User prompts, conversation history Conducts a structured stakeholder interview; clarifies scope, roles, pages, and visual direction Approved PRD
DSDA Database Schema Design Agent Approved PRD Synthesizes entities and relationships into a Prisma schema; iteratively validates and repairs compilation errors Validated schema.prisma and DMMF
OOA Orchestrator Agent Shared project context, agent outputs Routes workflow stages; enforces permission and scope guardrails; manages retry limits and human-in-the-loop checkpoints Updated project state
BAGA Backend API Generator Agent DMMF, PRD context Invokes deterministic generators for NestJS modules, controllers, services, DTOs, CRUD endpoints, and RBAC decorators; resolves relation-depth heuristics Complete backend codebase
FCGA Frontend Code Generator Agent PRD, DMMF, frontend manifest Plans routes, pages, navigation, and component blocks; resolves relational display ambiguity; compiles the approved manifest Complete frontend codebase
DCA DevOps Containerization Agent Generated code artifacts Builds Docker images; installs dependencies; configures inter-service networking; launches the local preview Dockerized, deployable application
Table 2. Tested Applications and Pipeline behaviour.
Table 2. Tested Applications and Pipeline behaviour.
Additionally, we can see the following application domain, entity, relation, DSDA, time and manual.Furthermore, we can see the following application domain, entity, relation, DSDA, time and manual.
iter. (min) interv.
HR onboarding & leave 5 4 2 9 No
Clinic appointments 7 6 2 18 No
Inventory & suppliers 8 9 2 22 No
Project & task tracking 11 12 3 30 No
E-commerce catalog & orders 12 14 3 38 No
Learning management 17 21 4 95 Yesa
Helpdesk ticketing 20 26 4 160 No
aManual refinement of the prompt describing the two entities related by a junction-table.
Table 3. Qualitative comparison of code-generation approaches
Table 3. Qualitative comparison of code-generation approaches
Feature GitHub Copilot v0 by Vercel Bolt.new AgentCoder CodeCraft
Model None (context based) None (prompt based) None (prompt based) None (prompt based) Formal Prisma schema
Architecture, Single-file LLM, LLM (with focus on generation), Full-stack probabilistic LLM, Multi-agent probabilistic Deterministic generation + constrained agentic reasoning
Implicit agent-coordination rules No No Yes (2 roles) Yes (3 roles)
Use of action traceability No No No No Yes (typed artifact per stage)
RAC ready (RBAC, deployment) No No No Yes
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.