Submitted:
17 August 2026
Posted:
19 August 2026
You are already at the latest version
Abstract
There are always repeatable functions like building similar CRUD workflows, implementing authentication layers, role based access control and admin interfaces for various projects on an enterprise web application, and different con-ventions can be used by different developers or code-bases. AI-powered coding tools recently emerged that promise to make this possible, but existing studies have shown that automated LLM-based code generation, vibe coding and free-running multi-agent pipelines have difficulty maintaining consistency between relational schemas, backend APIs, and frontend interfaces as applications grow in size, and become a maintenance night-mare. The main idea is that structural code which can be programmed from an application data model does not need to be probabilistic (only truly ambiguous decisions are language-model decisions such as interpreting the relational semantics, or resolving ambiguous requirements). We build this insight into a system called CodeCraft, where a Prisma schema and its Data Model Meta Format (DMMF) representation are the only source of truth for deterministic generators to generate backend APIs, frontend manifests, RBAC structures and database seeders, while a LangGraph six-agent multi-agent orchestration pipeline (from requirements engineering, schema design, orchestration, backend generation, frontend generation, and containerization) can only be used to clarify requirements, validate schema, and make decisions that cannot be made structurally. Generated systems have a consistent schema representation, as opposed to the prompt-driven tools which are used for each individual file. In all the applications it is demonstrated that schema validation slashes the number of automated iterations required to correct schemas to 2-4 and the end-to-end generation generates deploy-able, containerised applications after requirements and schema approval (no manual effort required). Comparative engineering indicates that the development effort is reduced by 70–80%, compared to manual implementation, but this has not yet been substantiated with a controlled external study.
Keywords:
code generation
; software architecture
; multi-agent systems
; langgraph
; deterministic scaffolding
; LLM agents
1. Introduction
Most of the time in the initial development of a project is consumed solving the same kinds of problems in enterprise software development: CRUD APIs, authentication flows, DTO validation, role-based access control, forms, tables, menu configuration, and deployment set-up, just to mention a few. As this structural work is recreated on each project, even with the same technology stack, teams end up working in various folder structures, with different validation patterns and API designs, leading to ever more complex code bases that become more difficult to maintain and more difficult to be introduced to over time [19].
This repetition has led to the development of AI coding assistants that enable developers to input the app’s requirements in natural language and let AI write the code. They can be useful for rapid prototyping and isolated components, but show to be useful only in the short-term: relational schemas and backend validation differ from the prompt over iterations [16,17] and correction attempts result in repeated instances of the same incorrect fix, that fail to fix the root cause of the underlying inconsistency [4]. This is addressed by multi-agent LLM pipelines, which break down the generation process into planning, coding, and validation tasks, and demonstrated to boost the quality of reasoning and reliability on correction in generation tasks [5,12,13,18]. However, the same cross-file synchronization problems that plague single-model tools (as main elegant, short business logic code is plagued by single-model cross-file synchronization problems) plague present multi-agent software-engineering platforms, but in more agents.
This is one key area that we’ve personally learned from as we evaluated the tools that were available to us ahead of development to support prompts. A new form and possibly new validation rules for a new backend and new constraints for a new database would have to be imported into the tool and synchronized with the old database and forms, if a schema change is applied to an entity that impacts another entity. Much like conversational code generation in general: small, localized prompt corrections have no impact on the previously generated files’ assumptions and unrelated frontend/backend components fall out of synch after a few iterations [16]. The failure is not in the model capacity, larger or more powerful language models will have the same drift in their output but it is a structural failure: if every file is generated using reasoning on a prompt history as opposed to composing from a single authoritative representation of the application, consistency between files is not guaranteed but only probable.
Having noticed that, when the application’s data model can be mechanically translated into structural code, it is not necessary to do so, the authors propose a deterministic-first multi-agent approach to generating enterprise Web applications. In this paper, we suggest a multi-agent approach, where the agents generate enterprise web applications using a deterministic first approach, based on the observation that if structural code can be mechanically generated from an application’s data model [2] then that is not needed. The basic idea is that a Prisma schema and the Data Model Meta Format (DMMF) representation is the one single source of truth that is used to deterministically compile the back-end API, the front-end manifest, RBAC structure and database seeders, and that a 6-agent orchestration pipeline, made up of a Requirements Engineering Agent, Database Schema Design Agent, Orchestrator Agent, Backend API Generator Agent, Frontend Code Generator Agent, and DevOps Containerization Agent, is only used for tasks that truly require reasoning, such as clarifying ambiguous requirements, validating and repairing schema definitions, and resolving relational semantics that can’t be inferred structurally. The resulting pipeline, as shown in Figure 1 consumes the same validated DMMF artifact for both the backend and frontend generation branches of the pipeline, while the various stages of deterministic compilation are isolated from the agentic reasoning stages.
The structural consistency property, which neither can ensure template-only systems, nor can guarantee fully agentic systems [5,18] (due to schema independence), is supported by the application layers in this system that rely on the same DMMF representation as the downstream generator.
Additional benefits other than uniformity are available with this design. First, it is worth noting that the language-model is not designed to make any non-trivially ambiguous decisions, meaning that there is a much smaller space for hallucinations to impact the generated system: the outputs are not source files, but rather configuration artifacts such as PRDs, schema patches, or frontend manifests that can be validated against a schema or compilation step before any code is written, not trusted without question [2,4]. Second, an application generated by the structural generation can be regenerated after schema modification, without re-running all the previous interactions that led to the previous version of the application, but simply recompiling its representation in the DMMF.Second, the structural generation is deterministic, meaning that an application generated by the structural generation can be regenerated after schema modification without re-running all previous interactions that have led to the previous version of the application, but simply recompiling the application representation in the DMMF. In our proposed framework, consistency is a property of the generation architecture rather than an emergent behavior that the agents must maintain, as a fully agentic pipeline, all outputs (structural or otherwise) are produced by a language model and therefore only as consistent as the language model’s context handling. The system architecture is shown in Figure 1.
This work has been deliberately limited in scope. We’re interested in full stack business applications that have a relational data model, have CRUD workflows, have authentication, and have role based access control; this is the class of apps where things tend to repeat from project to project and deterministic compilation may come in handy. We don’t target applications with a niche and experimental front-end, game development, or systems where the main value of the application is a novel and non-relational interaction pattern, as there is not much schema-driven system code to derive in these systems to begin with. In this sense, two constraints are not incidental, but intentional: there can be remaining ambiguities in the requirements that still need to be clarified in the requirements-engineering agent before the generation process can continue; and as accurate as the schema, so accurate can a deterministic generator be.
Let’s now look at three contributions we make. We introduce a schema-first deterministic generation engine that generates both backend APIs, frontend manifests, RBAC logic and backend seed data from one version of the Prisma DMMF representation, without inconsistencies of model re-definition that occur in unconstrained AI-generated systems [1,2]. Secondly, we design an orchestration pipeline for six agents, which has an explicit execution graph and only lets the language model engage in requirements interpretation, schema validation and ambiguity resolution without directly producing code [5]. Second, we explicitly design our six-agent orchestration pipeline in which the language model is only used in requirement interpretation, schema validation and ambiguity resolution, but not in raw code generation [5]. Thirdly, an end-to-end integration of these two components leads to deployable containerized enterprise applications: For the app that we’ve internally tested, we converged the schemas in 2-4 iterations of corrections, and generated schemas manually without any intervention, once the schema was approved along with the requirements. We also qualitatively compare CodeCraft with popular single-LLM and multi-agent code-generation systems, and estimate that the development effort could be reduced by 70–80% compared to manual implementation (pending a controlled external study [3]).
2. Related Work
This section focuses on automated and agent-based code generation. Previous research in automated code generation has aimed at reducing the amount of manual effort for code development while keeping the generated code correct for an application with a large structure, but early enterprise automation frameworks have tended to focus on optimizing their delivery pipelines and integrating AI tools into the existing software development workflows rather than ensuring consistency in the architecture of the systems generated by them [10,11]. Later, research has focused on pipelines with agents which divide this process into specific tasks. Multi-agent reasoning is shown to be more reliable than one prompt generating agent, as Blueprint2Code added a cycle of planning-and-repair, where one agent would provide a structural blueprint and another agent would provide and repair the code based on this blueprint [12]. AgentCoder extended this trend by enabling iterative loops between specialized coding agents, demonstrating that code verification steps can further improve code quality when embedded directly into the code generation process, compared to a single agent [5]. In contrast, the team behind MapCoder had another type of problem domain for competitive programming (planner and solver), and found that the separation of concerns works well beyond the enterprise domain of work, and in fact performs better than a monolithic generation [13]. This means that the separation of concerns principle applies to everything, even with things we do outside of work. Our approach is along similar lines of role specialization: a schema is validated by a particular agent, a particular agent orchestrates the process, a particular agent generates the backend code and a particular agent generates the frontend code, but differs from these systems in one way: there is no restriction on which parts of the generation process are delegated to agents and which are derived deterministically: the entire pipeline structural + non-structural code is an agent responsibility.
2.1. LLM Based Software Generation Limitations
There are restrictions on LLM-Based Software Generation. There have been other works that have demonstrated that LLM-generated software is not reliable when it comes to cross-file and cross-layer coordination, and as agent-based decomposition has improved the quality of reasoning, it is expected that more will follow. While both single model and an ensemble of agents have been investigated, in both cases, inconsistencies between architectural layers increase with the complexity of the project [1,2]. This is the same effect that is mentioned in our Introduction: when a schema change was cascaded through the files, this interface, created correctly in isolation, was no longer synced with the backend validation and database constraints. ReCatcher reports more data that this is not a problem with first generation products, but a recurring one: common modifications to a code base generated by an LLM frequently lead to regression bugs, meaning that fixing the code can just as easily cause it to break as fix it [7]. Another team of researchers (SonarQube, Oracle) has validated this concern from a different angle: AI-generated code can pass straightforward functional tests, but can also have more subtle flaws that will become apparent as maintainability issues later in the development process [14]. Legacy system modernization is again discussed in the context of architectural modernization, where one important finding is that the price to pay for structural coherence rises with the development and evolution of the system, an observation that is also pertinent to AI-generated systems that, after some generations, experience a structural drift [15]. Combined, this volume suggests that the benefit of multi-agent decomposition when it comes to reasoning reliability is not, alone, sufficient to solve another, more basic, failure mode: lack of a single authoritative structural representation on which all the layers of generation can be checked.
2.2. Vibe Coding/Prompt-Driven Development
It also inspired another approach that uses `vibe coding,’ enabling developers to convey intent in natural language and having the model fill in the implementation for them [16,17]. The literature consistently demonstrates that conversational generation helps to quickly prototype, particularly for developers who lack full-stack experience, and is suitable for isolated components that are generated in chunks [17]. Some CRUD interfaces and standalone components can be easily generated from prompts as we have already had some experience with tools in this category in our own early evaluations, as described in the Introduction. The same literature can also give a consistent answer to the question of where it fails—when an application must be maintained over a long time, architectural consistency or coordination of interdependent modules [4,16]. A common practical experience referred to in the literature is the one in which the developer knows the architecture of the application, and is able to do manual checks at each step of the coding process (which is not the case in the enterprise environment that is the focus of this paper, where the explicit motivation for automation is the elimination of manual checks as early in the coding process as possible). As explained in the section on AI-conversation driven programming, it is impossible to detect, and correct, that there is an inconsistency between the previous events of the conversation and the intended change in behavior of the model; such an inconsistency is local [16]. It’s the same failure mode as described in the Introduction in which another file on the front or back-end was not kept in sync. CodeCraft differs from this literature by considering prompts as high-level requirement inputs for a bounded requirements-engineering process, which are not instructions-level commands for structuring the code (as vibe coding tools are, by design).
Figure 2.
A single schema representation; deterministic compilation. No LLM is used to convert the Prisma schema fragment to the DMMF and use the DMMF to generate four artifacts: controller, service, DTOs, CRUD endpoints, along with the frontend manifest entry (routes, table columns, form fields), RBAC guard decorators, and the foreign-key–ordered database seed entry, which are then compiled into a NestJS backend module. As the input for all 4 layers is identical, they are all synced by construction: if one layer has an artifact, all 4 layers have the same artifact.
Figure 2.
A single schema representation; deterministic compilation. No LLM is used to convert the Prisma schema fragment to the DMMF and use the DMMF to generate four artifacts: controller, service, DTOs, CRUD endpoints, along with the frontend manifest entry (routes, table columns, form fields), RBAC guard decorators, and the foreign-key–ordered database seed entry, which are then compiled into a NestJS backend module. As the input for all 4 layers is identical, they are all synced by construction: if one layer has an artifact, all 4 layers have the same artifact.

2.3. Multi-Agent Systems in Software Engineering
Apart from the pipelines mentioned above in Section 2.1, there is also a broader literature on multi-agent software engineering that has looked into the relationship between task decomposition and the reliability of the resultant whole system. Several works demonstrated that splitting the tasks of the generation process into different stages among different agents to achieve better generations, instead of relying on a single prompting strategy to handle the whole generation process, is preferable [5,12,18]. This decomposition principle directly influences how the responsibility for CodeCraft is divided among the various stages, where each of the stages is assigned to a separate Responsibility, and not all Responsibilities to be carried out in one stage. A wider range of multi-agent software engineering systems, however, shows that most challenges of coordination overhead, memory synchronization and unreliable inter-agent communication yet to be solved, and that most solutions on the horizon are specialization.But a much wider survey of multi-agent software engineering systems reveals that the problems of coordination overhead, memory synchronization, and unreliable communication between agents do not disappear with specialization, and remain open problems [18]. For instance, this problem has haunted us since we began experimenting with a version of our orchestration layer that enables agents to coordinate through an open-ended, conversational exchange, execution order, permission to call tools, transitions in the workflow etc. etc. and are thus hard to predict and hard to debug. This observation led to a design decision, which is reflected in CodeCraft, that coordination is mediated by an explicit execution graph with typed transitions between stages (Figure 3) and not by free-form communication between agents, which means that the same coordination failures that are presented in the broader multi-agent literature, e.g., [18], cannot occur due to miscommunication between agents.
2.4. Deterministic Generation, Information Architecture, and Structural Consistency
There is a substructure of Determination Creation, Information Architecture, and Structural Consistency. Deterministic generation, information architecture and structural consistency are presented in the subsections respectively.The subsections are: Deterministic generation, Information architecture, and Structural consistency respectively. To provide the conceptual foundations of the central design principle of this paper, one strand of research, albeit peripheral to AI-assisted development, is outlined below. It has been demonstrated that, at scale, information architecture evolves more easily when a system’s navigation structure and data relationships are created by several different parties, while the application logic follows an information structure model shared by all of those parties [19]. The main idea of this research is similar to the failure mode of sections 2.2 and 2.3, that is, the independently produced application layers break apart as the system expands, essentially because there is no single representation that restricts the ways in which the application layers can intersect. In full on-demand generation systems, the idea of a “more direct ancestor” arises because maintaining file synchronization becomes a challenge as the system gets larger, because of the absence of a common structural restriction in the system other than the prompt history of the previous generation [2]. This separation will be immediately put to use in our central design decision: that we’re generating structural code on a deterministically single Prisma DMMF representation, rather than generating the code based on language-model reasoning. The same analysis also suggests that the semantics of the initial schema design [2] is the reason why our approach does not treat the creation of the schema as a deterministic task - the iterative validation loop of the Database Schema Design Agent makes sure that the schema that the downstream generators are to compile from is structurally sound prior to actual generation.
Overall, the evidence gathered from the various studies summarized in this section indicates progress in both aspects, namely multi-agent decomposition for ensuring reliability of the AI reasoning during programming code generation (both in terms of correct code generation and in terms of consistency with the specifications) in the one hand, and research on prompt-driven coding and generated-code quality in the other hand, in that reasoning-based coding is not a viable method to ensure architectural consistency at scale. However, the two findings are not explored together in a system in which they do not compete: Schema-based deterministic compilation of the structural code requires consistency, the interpretation of requirements, schema validation and resolution of ambiguities are performed by multi-agent reasoning, in which deterministic code generation is not enough. This is what we are here to fill.
3. Methodology
3.1. System Overview and Design Rationale
In this section we review the system design and the rationale behind it. In this section, the system is described and the reasons for its design explained. The mission of CodeCraft is to go through the natural language requirements and generate a full stack enterprise web application – relational data models, CRUD workflows, authentication, role based access control. The main design feature is to separate the mechanically derivable features of a data model from the interpretation required: A deterministic generation engine derives directly structural code from a schema representation, with a six-agent orchestration pipeline interpreting requirements, validing schemas and resolving ambiguities. The resulting seven staged pipeline from requirements capture, through to containerized deployment is shown in Figure 1. The design of the deterministic engine (Section 3.2) is described, followed by the design of a multi-agent orchestration layer built on top of this deterministic engine (Section 3.3), the rationale for this separation (Section 3.4) and the implications of this separation in terms of technology stack (Section 3.5).
This division had been due to the failure of both of the strategies in the early stages. A fully prompt-driven generation (with evaluation of the actual schema as explained in Section 1) successfully generated isolated components, but left the frontend form, backend validation, and relational schema drifting away from each other after a few iterations of corrections, as there was no shared representation to constrain the relationship between the different layers. An alternative that was only based on templates did not suffer from this syndrome, but was on the other hand too strict: It could not deal with ambiguous requirements, could not resolve relational ambiguities, e.g., which attribute of an associated entity to show in a generated table, could not pick up up if a schema did not validate. Neither the extremes were allowed: too flexible and too consistent. This led to the idea of two complementary tasks – deterministic compilation, and agentic reasoning – to be seen as part of the same pipeline, rather than as mutually exclusive tasks.
The resulting pipeline executes in seven stages, shown in Figure 1. Initialization begins with a conversational exchange in which the Requirements Engineering Agent clarifies scope, pages, roles, and visual direction, producing an approved Product Requirements Document (PRD). Project Setup creates an isolated workspace with standardized template and specification directories. Schema Design passes the PRD to the Database Schema Design Agent, which synthesizes a Prisma schema and validates it iteratively against compilation and relational constraints until it passes cleanly. Seed Generation derives realistic test data directly from the validated schema’s entities and relationships. Backend Generation compiles NestJS controllers, services, DTOs, and RBAC decorators deterministically from the schema’s Data Model Meta Format (DMMF) representation. Frontend Generation produces a JSON manifest describing pages, navigation, and components, which a deterministic compiler then translates into executable Next.js code. Containerization packages the generated backend, frontend, and database into Docker containers and launches a locally accessible preview. The artifact of the previous stage is validated and used by the current stage as its only input, with no means for the current stage to progress to an artifact that is not validated or partially defined.
3.2. Deterministic Generation Engine
The deterministic engine is required because there is a whole class of code that can be generated deterministically (i.e., the structure of this code is so tied to an application’s data model that it can’t be probabilistically generated). The core of this engine is a schema-processing module which converts a Prisma schema into its Data Model Meta Format (DMMF) representation, a structured JSON document that represents entities, their fields, relationships, enums, and constraints. From this representation, the engine attempts to find the junction tables to identify many-to-many relationships, then it traverses the relationships between entities to generate relational traversal metadata, and finally collects all the annotations in the module-level from all the schema models, for example, ///{"modulePath": "...", "moduleName": "..."}, to assign entities to a consistent module hierarchy downstream generators. This is a processed representation which is consumed by each downstream generator, rather than the raw schema text, so that no generator reparses or re-interprets the schema text.
From the above, meaning of fields, relationships and constraints for an entity cannot be different across various parts of the pipeline (backend, frontend, RBAC, seed generation). This directly overcomes concerns raised in previous research of LLM generated systems [1,2] where the schema was reconstructed at each layer—in the case of SDL forms, the database, or the backend DTOs—using different prompts or generators, and there was no common source of truth to compare with.
The common schema representation is not generated per project, but rather built from a library of golden templates that are common to enterprise applications such as tables, forms, layout, modals, RBAC wrappers, and navigation. This is a generic data-table-generator, that is used everywhere: For each entity, the generator doesn’t generate a specific table implementation, it generates a table-configuration-object with the description of the table, its columns, filters, sort, pagination, etc. and relations (also for other entities), and the front-end-renderer is a common object that reads the table-configuration-object and renders the same table (with different content) for each entity. The approach is to design the configuration, and adding a new entity to an application involves only adding a new configuration, which is defined in terms of the schema, and no code change to the tables; so there is no effort required to implement uniformity of the tables; the tables are uniform in structure and appearance.
The frontend compilation works in a deterministic manner, with the sole exception of the language-model reasoning. Using the DMMF representation, the engine produces a high level JSON file called a manifest that lets you define the structure of forms, entity mapping to tables, navigation rules and how to set up the relationships that are displayed on the form; after that, a deterministic compiler converts the manifest into a full Next.js module. The one step in this pipeline that is assigned to a language model is the resolution of relational display ambiguity, which is the case when, with a Doctor entity and a User entity, the generated table needs to display either User.The fullName field or User.email or any other field that makes sense given the meaning of the data, but not based on the structure of the schema. After this ambiguity is eliminated and documented in the manifest, the rest of the compilation from manifest to executable frontend code is completely deterministic.
What makes this limited advisory role both safe and not just convenient is that the language model doesn’t try to create a full React component, rather it attempts to create a manifest, which is a bounded configuration object that can be validated against a schema before any code is compiled from it, but a full React component generated end-to-end by the language model can be difficult to validate, and can contain structural errors without warning. This follows the general idea, also used below in the generating of backend code and RBAC configuration, of using language models to generate artefacts that a deterministic system can execute, instead of direct generation of final code by the language models themselves [2,4].
The idea of the Backend generation is very similar to DMMF, with the difference that for every entity, the controllers, services, DTOs and CRUD endpoints are generated using NestJS, with minimal involvement from the language-model. Every entity is just stored in a specific module and has a set pattern: The module contains a controller with create, paginated list, single record retrieval, update and delete endpoints, as well as a service layer that processes Prisma queries, and DTOs are not separated out from the entity’s field definitions. Therefore, no two modules are structurally different, no matter what they may be representing, when they are generated. In this stage, there’s a loose coupling in the case of some generated endpoint returning an entity along with other records related to it—where the engine uses heuristic rules to determine the depth of the resulting response payload. These heuristics can be used on schemas of moderate relational complexity and can produce very large or very shallow payloads on schemas that have more densely connected schemas; and the schema does not need to be arbitrarily deep to be general enough for these heuristics to apply, so we revisit this as another natural extension point for agentic reasoning in future work in Section 4.5.
There are two other patterns that are less complex than the above, but deterministic: role-based access control and seed data. The two patterns can be extracted from schema and PRD metadata without ambiguity in the relational model. The role-to-endpoint mapping, route guards, and role-visibility rules for navigation menus are automatically created from the role definitions in the PRD, and applied to both the generated backend decorators and the navigation menu configuration on the frontend, helping to prevent inconsistencies between the two layers. Seed data generation also looks at entities and relationships in the validated schema to generate valid seed data, default roles and demo data without having to involve any sort of agent – the constraints in the schema dictate the shape of valid seed data.
3.3. Multi-Agent Orchestration Pipeline
The deterministic engine in Section 3.2 generates structurally consistent code, but it can by construction not interpret natural-language requirements, deal with business-specific relationships not explicitly stated in a schema, recover from a failed generation step and manage the dependencies between pipeline stages. The only things that need to be reasoned, are these tasks, which is why it is not being compiled in the generators, and has an orchestration layer on top of the deterministic one. This layer would separate this “centralist language model” which has been shown by previous research to forget its context and make inconsistent decisions during longer multi-stage executions [18] with exactly one agent for each stage of the pipeline.
Table 1 describes the roles of the 6 agents and their inputs (and outputs). The Requirements Engineering Agent (REA) carries out a structured conversation with the user, which is similar to an interview, and continues to generate the PRD outlined in Section 3.1. The Database Schema Design Agent (DSDA) creates a Prisma schema from the approved PRD and corrects it based on the feedback from the compiler, until it is valid without warnings. Execution is coordinated by the Orchestrator Agent (OOA): it keeps track of the state of the pipeline, routes the right context slice to the agents downstream and enforces the guardrails outlined below. The Backend API Generator Agent (BAGA) and Frontend Code Generator Agent (FCGA) invoke the deterministic generators of Section 3.2 with the validated DMMF and manifest inputs, intervening only in the case of having to disambiguate between relational entities in the manifest. The DevOps Containerization Agent (DCA) packages up the created backend, frontend, and database into Docker containers and keeps the resulting preview instance on the local machine. The output artifact of each agent is just one kind (either PRD, validated schema, manifest, codebase, or container set), and is the only input to the next agent, so no agent needs to be aware of any particular kind of reasoning process used by other agents.
This is implemented as a directed execution graph within LangGraph, with nodes representing the different agents, and edges representing conditional transitions (e.g. from schema design to seed generation only occurs when the DSDA reports a successful compilation). All agents share a single shared project context object that is passed from agent to agent across the graph that references the current PRD, current schema, current DMMF, artefacts created at previous pipeline steps and current pipeline state. The Orchestrator Agent only passes this relevant portion to each node: DMMF and RBAC metadata to the BAGA; and PRD, DMMF and frontend manifest definitions to the FCGA, limiting the amount of information passed to each agent to what each agent needs for its given job.
This graph-based design was motivated by the difficulty of predicting and debugging the order of tool calls and execution permissions in early development when the agents were able to communicate through free-flowing conversation in this domain (Section 2.4). Since the transitions in a directed graph are explicit, and depend on the validated transitions output by the stage instead of on free-form dialogue the agents exchange, there are no problems of coordination and memory-synchronization that could arise from unclear inter-agent communication, as described in the general multi-agent software-engineering literature [18].
In this graph, there are 3 guardrails that limit the execution. The first one is that the permission validation is done before every agent invocation and the agent can only use the tools for generation for which it holds the permission: for instance if there are artifacts of the backend in the shared context, the FCGA will not be able to invoke the backend-generation tools. Second, human-in-the-loop checkpoints are added at two places where the error thrown down the pipeline is the highest: PRD finalization and schema approval, where users have to approve the pipeline explicitly before it goes on to generation. Third, the DSDA’s schema-validation loop has a maximum number of times that it will loop before throwing an exception, if compilation fails because the problem cannot be resolved, the loop does not go on forever, but the programmer is alerted to the problem and can correct it. A similar mechanism exists for the out of scope boundaries defined in the accepted PRD: requests outside the scope defined in the PRD will not be addressed by the generation agents, Orchestrator Agent will reject any such request before it arrives at a generation agent and will not partially process it.
These guardrails are meant to mitigate a failure mode that goes beyond what the agents can do; (a single agent can be right, but if there is no time, resource or scope of action limitation, a globally undesirable outcome can result). For example, if the DSDA is configured to ’silently fail’ (such as retest the query in an infinite loop), then it is not going to improve the accuracy of the schema correction, it will simply not make the schema correction. Likewise, if the scope is enforced at the orchestration level (as opposed to each generation agent), the feature growth reported in iterative conversational development [4,7] cannot occur, as the scope is not reliant on each generation agent remembering and honoring a boundary several generations before.
3.4. Design Principles
The above-mentioned architecture is based on four design principles. The first is schema as single source of truth: The representation of the schema in the Prisma DMMF should represent the source of truth for all generated artefacts, not for different interpretations by different generators or agents. This principle is applied for design reasons: There’s no code path in CodeCraft where it can occur: Each layer’s understanding of the entity is reconstructed based on the same shared DMMF for the entity. This is not allowed in CodeCraft by construction: each layer’s understanding of the entity is rebuilt based on the same common DMMF for the entity.
The second principle is “Determinism over Language-model reasoning for structural code” that determines which structures are passed to the deterministic generators and which to the reasoning agents. The agents are used only for tasks that need to be interpreted and/or need to resolve ambiguity, such as PRD refinement, schema synthesis and relational display disambiguation; the rest of the structural code is generated from schema metadata. The key to the solution in the compiler domain is to remove the compiler work that can be mechanically derived from the compiler and to bring it to a surface area, where hallucination-driven inconsistencies can occur. In the compiler domain, the important part of the solution is to move away from the predictable work that can be easily extracted from the compiler to a surface area where hallucination-driven inconsistencies can enter [1,2,3].
The third one is that the structured configuration artifacts (PRD, schema definition, frontend manifests) are the only means of interfacing with language-model reasoning and deterministic execution, not code generation directly from the agents themselves. The generator is deterministic, all agents that interact with structural generation (like the Frontend Architect feature of the FCGA), and all back-end generation agents, produce a validated configuration object but not a code artifact. This is why agent output can be checked before it has cascading effects: If the error is discovered in a large codebase, it is hard to trace it, but if the error is discovered in a manifest, then the schema can be used to validate without having to dig through the codebase to find the error [2,4].
The fourth rule, “validate before progress”, means that no pipeline stage start to process an artifact from the previous stage until it has been structurally validated. The most obvious is the DSDA’s iterative schema-validation loop: If a schema is not satisfied by the syntax test or the relational constraints test of Prisma, it is not sent to downstream generators for schema correction, but is corrected in place. The bounded-retry guardrail described in Section 3.3 is not a soft guardrail, as the artifact isn’t bound to the next stage until it is validated at that stage and we know it’s correct.
3.5. Implementation and Technology Stack
The implementation and the technology stack are described below.The implementation and the technology stack are described below. The implementation and the technology stack are detailed in this section. It describes how this section is implemented and what technologies are used. The stack of technologies are organized under the framework, and selected specifically for the ability to generate in a schema-driven and deterministic manner. Due to the file routing nature of Next.js, it was selected for the generated frontend as it will map to the route and page structure generated by the FCGA’s manifest, without being configured at runtime. Besides being, for the most part, lighter alternatives to NestJS (such as Express), the generated backend is decided for due to the method the BAGA generates new modules for its entities – the modules are isolated units, so the NestJS decorator-based modular architecture fits the bill.
This schema-first approach to the design of the whole pipeline is built upon Prisma ORM, where all generators in Section 3.2 receive the structured input, and there are no generators without this abstraction layer that are deterministic. Most of the target features of the framework, such as transactional consistency, relational data models and RBAC hierarchies were obvious when choosing the relational database target, and so was MySQL. The orchestration layer was implemented by LangGraph instead of sequential API calls, as it was required to store and share persistent state between agents that are dependent on each other, and also to control the flow of execution as a graph between steps of agent reasoning, as described in Section 3.3; Gemini API was chosen for reasoning steps because it was required to follow instructions and as part of the reasoning process, share context, which is required to be persistent, between steps of reasoning.
Using two separate job queues which don’t block, and perform generation requests concurrently, Redis and BullMQ can be used to decouple long running generation tasks (schema validation, frontend compilation, dependency installation, container builds), from the main backend process. The reason for choosing Docker for deployment is that generated applications require isolated, reproducible runtimes: The DCA packages are made up of backend, frontend, and database components, all of which are deployed to machines with no manual configuration; these packages are then pushed to a container registry for deployment in the cloud without having to be rebuilt.
4. Results & Discussion
4.1. End-to-End Generation Outcome
We test the end-to-end performance of the CodeCraft by running a number of generation applications we’ve created ourselves across typical enterprise areas, such as an e-commerce catalog and order-management application. In a typical application, like an e-commerce platform that includes product catalog, category management, and order workflows, a developer writes the description of the target application in natural language, and the pipeline passes that description to the REA, which refines it into an approved PRD, passes it to the DSDA, where it is used to synthesize and validate the corresponding schema, passed to the BAGA, which compiles the backend code base from that schema, and passed to the FCGA, which compiles the frontend code base from that schema, and finally to the DCA, which bundles the result into a working application available via a local preview URL, without requiring any code-level intervention. A typical screen shot of one of the generated applications is shown in Figure 4 with category filtering, search and a product grid created entirely from an application manifest generated based on the frontend, as described in Section 3.2.
This outcome is an illustration of the specific claim which led to the design of the CodeCraft; once the PRD and schema are accepted by the user, the following five steps are performed automatically at the code level: all the downstream artifacts are automatically compiled from the approved schema, or are created by agents whose output is automatically validated prior to use. This is a difference that is very distinct from the prompt-based tools described in Section 2.3, which would need manual edits at every iteration as the schema is added inconsistently in the files as part of a similarly sized application. This doesn’t imply that this business logic (as opposed to CRUD, RBAC, and the standard workflows) is incorrect, it simply means that this business logic was not tested for this evaluation, as stated in Section 1, it was observed when the pipeline was completed.
4.2. Observed Pipeline Behavior
The iterative schema-validation loop typically converged to a schema that passed the compilation and relational constraint checking within two to four passes of the “correction” pass in all of the applications tested. The tested applications are summarized in Table 2 with the convergence iterations plotted against schema size for each of them in Figure 5. Schemas that were simpler, with fewer interdependent entities, tended to fall around the low end of this range, and schemas that were more “complex” (more foreign keys, more many-to-many junction tables) sometimes needed more iterations, and occasionally more manual fine-tuning of the prompt description of the relevant entities.
This is a range of properties which the tested applications exhibit, and not guaranteed for other schemas that have more than around twenty entities, or have denser connections between them.
Once the PRD and schema stages passed the validations, the process of generation was finished without any manual effort for most of the tested applications, as expected, based on the description given in Section 4.1. The failures that did occur tended to be one of three types: unexpected schema relationships not covered by the rules in the DSDA; large enough manifests to cause problems in the context handling of the manifest-planning reasoning step; dependency conflicts when building Docker images for the containerization step. None of these failure modes needed to be changed in the code previously generated and none of them required changing the code of the affected stage, which is the design principle of the framework of validating the artifacts before they propagate downstream (Section 3.4).
These observations are made on applications developed by the CodeCraft’s developers, not on a held out benchmark set, a third party test suite, or on applications developed by third parties who are not involved with the CodeCraft. The above convergence and failure-rate statistics should therefore be seen as representative of performance under the conditions in which the CodeCraft was developed and tested and not as a claim to generality for arbitrary enterprise schemas and to unfamiliar users, which we shall address directly in Section 4.5, and which we believe to be the primary target for future evaluation.
4.3. Comparative Positioning
In order to put the CodeCraft in perspective, we qualitatively compare it with four representative tools in four categories that we mentioned in Section 2: GitHub Copilot, v0 are context- and prompt-based single-model assistants, Bolt.new is a full prompt-driven code generator and AgentCoder is a representative multi-agent code generation system. The five aspects of the model, which we consider to be important for enterprise application generation, are compared here summarized in Table 3: the first aspect is input model, either schema-driven or just prompt-based; the second is the generation architecture; the third is the explicit agent-coordination rules; the fourth is the enterprise readiness, and the fifth is traceability of actions. This table is reflective of our own qualitative analysis of published capabilities and architecture of each tool, and not a result of any common benchmark task, common scoring rubric, or controlled experiment conducted under identical conditions, which we mention below.
In fact the only point with which the comparison in this paper makes its case is that the difference between the tools compared is that our tool separates schema-derived production of structure from agent-based reasoning, and this separation is what makes the consistency properties mentioned in Section 4.1 and Section 4.2 possible even if they are not proved.
We do this not to imply that this is experimental proof of superior, but as a positioning device. No common test applications were utilized in Section 4.1 and Section 4.2 and there was no common set of tasks or acceptance criteria defined across all the different tools; the labels of “Yes/No” and the qualitative labels in Table 3 are based upon our understanding of the architecture of each tool, documented in their respective literature; and the results are not measured. The comparison would have to be based on the same application specifications and examine them using the same application evaluation rubric, in a controlled study, of which we have not performed here, but which we recognize as a necessary future study in Section 5.3.
4.4. Estimated Effort Reduction
If such pipeline behaviors are observed in Section 4.1 and Section 4.2, then the pipeline framework is estimated to save around 70–80% repetitive development effort as compared to equivalent enterprise modules developed manually. It was estimated by comparing the time it takes to get the application up and running from start to end, depending on the size of the application, ranging from a few minutes for smaller ones, up to a few hours for larger applications with multiple modules and workflows, which is dominated by schema complexity, frontend manifest size, Docker build time and LLM response latency, and we estimated a few minutes to an hour for similar applications, where we would handcraft the CRUD modules, RBAC integration, web dashboards and deployment configuration.
It is a preliminary engineering estimate, not a statistically-proven result. It has not been undertaken in a controlled study with external developers, it ignores the fact that developers may be faster or slower than others and may differ in their familiarity with the target stack, and it was not tested with a different set of manual implementations that were withheld from the project, but rather implemented by developers who were not the creators of the framework. This estimate is made because it’s a realistic one based on our experience with the effort saved during the development and testing process, but we feel that closing the gap between this estimate and a validated measurement is the #1 priority for future work, which will be discussed in Section 5.3.
4.5. Limitations
The current implementation is based on a fixed technology stack (NestJS backend generation and Next.js frontend generation), which is chosen for the reasons mentioned above, but this limits the deterministic generation engine to not be able to currently generate other common enterprise stacks like Express.js, FastAPI or Vue.js applications, even if the schema is complex.
The endpoints generated are still relational (as described in Section 3.2), meaning that heuristic rules are still used as a temporary workaround for full schema-aware reasoning. We have not tested the heuristics in schemas with a large number of entities or schemas that have a large number of connected entities, where a fixed heuristic is unlikely to be generalizing.The schemas used were not very densely connected, and we have not tested the heuristics in such schemas, or in schemas with many entities. The next natural step is to find a dedicated include-planning agent, where it is explicitly dealt with with a query what relations it needs, which will remove this final heuristic aspect, as described in Section 5.3.
While the DCA can create applications and push them to a container registry, deploy them to AWS ECS or EC2, and verify the deployment, the deployment to AWS is not fully automated in this work as the configuration and the verification steps to complete the deployment were manual procedures.
Due to the external language model (Gemini API) provided, the PRD, schema interpretation or a PRD generated by an agent stage with the same input in natural language may differ from one run to the next. This non-determinism is only present in the reasoning phases: once an artifact (a validated schema or an approved manifest) is created, the deterministic generators of Section 3.2 always generate the same structure for a given prompt, but what is compiled is not deterministic across repeated runs of the system [2].
Last, and most important, all the findings in this section are internal, not from an external user study, but by the testers of the framework. This is the constraint that we feel is most important and must be overcome before these results can be considered final; indeed we feel that all claims in this section, in conjunction with the quantitative and qualitative, require this constraint to be met: the convergence rates of DSDA, the characterization of failure modes, the comparative positioning in Section 4.3, and the estimate of effort-reduction in Section 4.4.
5. Conclusions
In order to solve the problem of structural code generation inconsistency of enterprises that uses language models for end-to-end generation, this paper put forward a deterministic-first multi-agent framework that decouples the structural code generation from the reasoning of agents. The main concept is that code that can be mechanically generated based on the data model for an application (CRUD APIs, DTOs, RBAC decorators, frontend forms and tables) should be compiled deterministically from a single schema representation, whereas a 6-agent orchestration pipeline should be considered for tasks that are truly interpretative, such as clarifying requirements, verifying schema design, and resolving relational ambiguity that can’t be determined in a fixed manner.
It took just a couple of automated schema corrections after the requirements and schema acceptance to deploy and containerize enterprise applications and validated schema in all internal tests. In contrast to existing methods and our own, which separately considered each application layer in turn by individual agents or prompts, the generated applications were not vulnerable to the cross-layer drift we and other works identified as a chronic failure mode of purely prompt-driven and unconstrained multi-agent generation [1,2,4,7].
Overall, this research suggests that deterministic compilation and agentic reasoning are not conflicting methods for software generation using AI, as one method can be used when the other is not, and vice versa. It’s a subtle shift in how a developer’s job is done when using this kind of system, as it’s not about writing and re-syncing boilerplate by hand anymore, it’s about being able to precisely define requirements, architecture and scope, and then relying on an automated pipeline to get the rest.
This scope isn’t a coincidence. The framework is designed for enterprise applications with relational data and CRUD operations, and it wasn’t designed with highly experimental frontends, or with structurally sparse domains, in mind, both of which have very little to gain from a schema-driven engine. There are two unclosed restrictions in that range. Even the relational response depth is still heuristic and not schema-aware control and all the results reported here are from internal tests and not external studies. This is the estimated effort reduction of 70–80% to be interpreted as a preliminary reduction and not validated.
The next logical step is a controlled examination between the tools qualitatively described in Section 4.3 and manual development, in which the time required for development and architectural consistency are measured against a common task set and rubric, with the tools used by developers who have not used the framework. In addition, the following extensions are highlighted: support for additional backend and frontend frameworks, replace the heuristic relation-depth control with a dedicated include-planning agent and run the pipeline against schemas larger than those that were tested here. Each deals with a restriction mentioned previously in this paper, not a new one.
References
- “The Hidden Pitfalls of Using LLMs in Software Development,” Metabob. [Online]. Available: https://metabob.com/blogs/hidden-pitfalls-of-using-llms-in-sw. [Accessed: 22-May-2026].
- J. Meaden, “COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models,” arXiv, pp. 2–8, 2025.
- R. Rawal, “Benchmarking Correctness and Security in Multi-Turn Code Generation,” arXiv, 2025.
- S. Bird, “Why 1 in 5 Vibe-Coded Applications Fail Basic Security Tests,” Blott, 5 September 2025. [Online]. Available: https://www.blott.com/blog/post/why-1-in-5-vibe-coded-applications-fail-basic-security-tests. [Accessed: 26-May-2026].
- D. Huang, “AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation,” arXiv, 2024.
- J. Bossle, “Why AI Coding hits limits as adoption grows.” [Online]. Available: https://www.knowis.com/blog/why-ai-coding-hits-limits-as-adoption-grows. [Accessed: 26-May-2026].
- A. A. Abbassi, “ReCatcher: Towards LLMs Regression Testing for Code Generation,” arXiv, 25 July 2025.
- X. Wang, “CodeAct: Your LLM Agent Acts Better when Generating Code,” Apple Machine Learning Research, 2024.
- X. Wang, “Executable Code Actions Elicit Better LLM Agents,” in International Conference on Machine Learning, 2024.
- L. Sbeitan, “Agentic SDLC in practice: the rise of autonomous software delivery,” PwC.
- S. Teyssier, “A framework for integrating AI into platform engineering,” 3 December 2025. [Online]. Available: https://platformengineering.org/blog/a-framework-for-integrating-ai-into-platform-engineering. [Accessed: 26-May-2026].
- K. Mao, “Blueprint2Code: a multi-agent pipeline for reliable code generation via blueprint planning and repair,” Frontiers in Artificial Intelligence, 2025. [CrossRef]
- M. A. Islam, “MapCoder: Multi-Agent Code Generation for Competitive Problem Solving,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024.
- M. Kapur, “Vibe, then verify: SonarQube 2025 year in review,” 8 January 2026. [Online]. Available: https://www.sonarsource.com/blog/sonarqube-2025-year-in-review. [Accessed: 26-May-2026].
- “Architectural Modernization,” vFunction. [Online]. Available: https://vfunction.com/about/. [Accessed: 26-May-2026].
- A. Sarkar, “Vibe coding: programming through conversation with artificial intelligence,” arXiv, 2025.
- Nzall, “What is `vibe coding’ and what are the strong points of this methodology?,” Software Engineering Stack Exchange, 2025.
- J. He, “LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead,” arXiv, 2024.
- L. Rosenfeld, Information Architecture: For the Web and Beyond. USA: O’Reilly Media, 2015.
- X. Shen, “Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction,” arXiv, 2025.
- S. Author et al., Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging. arXiv preprint arXiv:2505.02133, 2025.
- X. Zhang et al., A Unified Debugging Approach via LLM-Based Multi-Agent Synergy. arXiv preprint arXiv:2404.17153, 2024.
- J. Doe et al., Testing and Enhancing Multi-Agent Systems for Robust Code Generation. arXiv preprint arXiv:2510.10460, 2025.
- A. Smith et al., Build AI agents with fine-tuned open models. Fireworks RFT, 2024.
- Y. Chen et al., AutoSafeCoder: A Multi-Agent Framework for Securing LLM Code Generation through Static Analysis and Fuzz Testing. arXiv preprint arXiv:2409.10737, 2024.
- AI Agent Architecture Patterns in 2025: The Powerful Way Multi-Agent Architectures Scale. Nexa AI, 2025. [Online]. Available: https://nexaitech.com/multi-ai-agent-architecutre-patterns-for-scale/.
- Z. Ma et al., Multi-Agent Systems Integration in Enterprise Environments Using Web Services. IGI Global, 2006.
- A. Garcia et al., Engineering multi-agent systems with aspects and patterns. Journal of the Brazilian Computer Society, 2002. [CrossRef]
- Z. Chen et al., MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents. arXiv preprint arXiv:2503.01935, 2025.
Figure 1.
The figure below shows the system architecture. Below is shown the architecture of a system. The proposed architecture of the framework is presented below. The blue nodes represent probabilistic reasoning agents (LLM), the orange nodes represent deterministic reasoning (compilers) without LLM, and the yellow document shapes are artifacts that are checked before being used in the downstream processes. The validated schema.prisma/DMMF representation (green cylinder) is only read by both the backend (BAGA) and frontend (FCGA) generation branches. The Orchestrator Agent (OOA, purple) controls the routing of the stages, permission guardrails and retry limits and human-in-loop checkpoints (dashed control edges).
Figure 1.
The figure below shows the system architecture. Below is shown the architecture of a system. The proposed architecture of the framework is presented below. The blue nodes represent probabilistic reasoning agents (LLM), the orange nodes represent deterministic reasoning (compilers) without LLM, and the yellow document shapes are artifacts that are checked before being used in the downstream processes. The validated schema.prisma/DMMF representation (green cylinder) is only read by both the backend (BAGA) and frontend (FCGA) generation branches. The Orchestrator Agent (OOA, purple) controls the routing of the stages, permission guardrails and retry limits and human-in-loop checkpoints (dashed control edges).

Figure 3.
The graph that describes the execution of the orchestrator. The Agent stages (blue), deterministic stages (orange) and typed transitions (bold caps) are connected, and the yellow diamonds are human-in-loop checkpoints. It will then go into a bounded automated repair loop, and only be shown to the user when it fails at the actual docker build stage, which will be retried within a bounded time.
Figure 3.
The graph that describes the execution of the orchestrator. The Agent stages (blue), deterministic stages (orange) and typed transitions (bold caps) are connected, and the yellow diamonds are human-in-loop checkpoints. It will then go into a bounded automated repair loop, and only be shown to the user when it fails at the actual docker build stage, which will be retried within a bounded time.

Figure 4.
Sample screen shot of an e-commerce application created. The product grid, category filter and search interface was completely created from the approved frontend manifest, and the supporting API, RBAC guards and seed data was created from a similar DMMF representation. None of the components were handwritten or edited.
Figure 4.
Sample screen shot of an e-commerce application created. The product grid, category filter and search interface was completely created from the approved frontend manifest, and the supporting API, RBAC guards and seed data was created from a similar DMMF representation. None of the components were handwritten or edited.

Figure 5.
The convergence of schema-validation values for the different applications and sizes of schema for the DSDA loop is shown.
Figure 5.
The convergence of schema-validation values for the different applications and sizes of schema for the DSDA loop is shown.

Table 1.
Agent roles, inputs, core processing, and outputs.
| Agent | Role | Input | Core Processing | Output |
|---|---|---|---|---|
| REA | Requirements Engineering Agent | User prompts, conversation history | Conducts a structured stakeholder interview; clarifies scope, roles, pages, and visual direction | Approved PRD |
| DSDA | Database Schema Design Agent | Approved PRD | Synthesizes entities and relationships into a Prisma schema; iteratively validates and repairs compilation errors | Validated schema.prisma and DMMF |
| OOA | Orchestrator Agent | Shared project context, agent outputs | Routes workflow stages; enforces permission and scope guardrails; manages retry limits and human-in-the-loop checkpoints | Updated project state |
| BAGA | Backend API Generator Agent | DMMF, PRD context | Invokes deterministic generators for NestJS modules, controllers, services, DTOs, CRUD endpoints, and RBAC decorators; resolves relation-depth heuristics | Complete backend codebase |
| FCGA | Frontend Code Generator Agent | PRD, DMMF, frontend manifest | Plans routes, pages, navigation, and component blocks; resolves relational display ambiguity; compiles the approved manifest | Complete frontend codebase |
| DCA | DevOps Containerization Agent | Generated code artifacts | Builds Docker images; installs dependencies; configures inter-service networking; launches the local preview | Dockerized, deployable application |
Table 2.
Tested Applications and Pipeline behaviour.
| Additionally, we can see the following application domain, entity, relation, DSDA, time and manual.Furthermore, we can see the following application domain, entity, relation, DSDA, time and manual. | |||||
|---|---|---|---|---|---|
| iter. | (min) | interv. | |||
| HR onboarding & leave | 5 | 4 | 2 | 9 | No |
| Clinic appointments | 7 | 6 | 2 | 18 | No |
| Inventory & suppliers | 8 | 9 | 2 | 22 | No |
| Project & task tracking | 11 | 12 | 3 | 30 | No |
| E-commerce catalog & orders | 12 | 14 | 3 | 38 | No |
| Learning management | 17 | 21 | 4 | 95 | Yesa |
| Helpdesk ticketing | 20 | 26 | 4 | 160 | No |
| aManual refinement of the prompt describing the two entities related by a junction-table. | |||||
Table 3.
Qualitative comparison of code-generation approaches
| Feature | GitHub Copilot | v0 by Vercel | Bolt.new | AgentCoder | CodeCraft |
|---|---|---|---|---|---|
| Model | None (context based) | None (prompt based) | None (prompt based) | None (prompt based) | Formal Prisma schema |
| Architecture, Single-file LLM, LLM (with focus on generation), Full-stack probabilistic LLM, Multi-agent probabilistic | Deterministic generation + constrained agentic reasoning | ||||
| Implicit agent-coordination rules | No | No | Yes (2 roles) | Yes (3 roles) | |
| Use of action traceability | No | No | No | No | Yes (typed artifact per stage) |
| RAC ready (RBAC, deployment) | No | No | No | Yes |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.