Preprint
Article

This version is not peer-reviewed.

Reassessing One-Round Test-Time Refinement for Code Generation

A peer-reviewed version of this preprint was published in:
AI 2026, 7(9), 353. https://doi.org/10.3390/ai7090353

Submitted:

12 August 2026

Posted:

12 August 2026

You are already at the latest version

Abstract
Test-time refinement aims to improve generated programs through additional inference, but its value after an initial candidate has been produced remains unclear. We conduct a controlled evaluation of one-round Self-Refine and Self-Debug across seven models and three Python code-generation benchmarks. For each model and task, both methods refine the same initial candidate, allowing us to measure refinement gain without variation in initial generation. Self-Debug produced positive refinement gain in 16 of the 21 model-benchmark combinations and no change in the remaining five, with gains reaching +9.57 percentage points. In contrast, Self-Refine reduced correctness in 16 combinations, with losses of up to 8.07 percentage points, and produced positive gains in only four combinations. Repair-regression analysis showed that Self-Debug rarely damaged initially correct candidates, whereas regressions under Self-Refine frequently outweighed its repairs. Further analysis showed that execution feedback was beneficial only when models acted on diagnostic failures and produced effective revisions. Resource analysis showed that Self-Refine incurred greater token overhead despite generally reducing correctness, making its additional inference difficult to justify. Self-Debug provided a more favorable gain-overhead balance, but its monetary efficiency varied with model pricing, and its additional gains may offer limited value when initial generation is already sufficiently accurate. These results show that one-round refinement is not inherently beneficial and should be applied only when its expected gain justifies the additional computation and cost.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Large language models (LLMs) can generate executable programs from natural-language specifications and solve a substantial fraction of widely used code-generation benchmarks [1,2]. However, a generated program may still misinterpret the specification, mishandle edge cases, or fail under broader tests. Test-time refinement attempts to correct such errors through additional inference without updating the model parameters. In feedback-free Self-Refine, the model reviews its own candidate and revises it using self-generated critique [3]. In Self-Debug, the candidate is executed and the model revises it using diagnostic test feedback [4].
Prior studies have reported that self-generated critique or execution feedback can improve generated programs in particular experimental settings [3,4]. However, subsequent evaluations have shown that models do not consistently identify and repair errors in their own outputs, and that effectiveness varies across models and tasks [5,6,7]. This evidence does not yet establish whether refinement remains a beneficial post-generation step once an initial candidate has already been obtained. Although refinement may repair an incorrect solution, it may also damage one that was already correct. This question becomes increasingly important for current models, whose stronger initial generation may leave less room for improvement. Refinement also incurs additional inference cost regardless of whether it changes or improves the candidate. It therefore remains unclear whether one round of Self-Refine or Self-Debug improves correctness sufficiently and consistently to justify applying it after initial code generation.
We investigate this question by measuring refinement gain, defined as the percentage-point change in pass rate between an initial candidate and its refined version. We conduct a controlled evaluation of one-round Self-Refine and Self-Debug using seven models: four proprietary models accessed through the OpenAI API and three code-specialized open-weight models from the Qwen and DeepSeek families served locally. We evaluate these models on three complementary Python code-generation benchmarks covering diverse task settings: HumanEval+, MBPP+, and BigCodeBench-Instruct. For each model and task, we generate one initial candidate and use its exact code as the unrefined baseline and as the starting point for both refinement methods. This design isolates the effect of refinement from variation in the initially generated program. We evaluate the resulting candidates using the complete final test suites and analyze refinement gain, its underlying repair-regression balance, the models’ responses to diagnostic feedback, and the resource overhead of each refinement method.
The results show a clear contrast between the two refinement methods. Self-Debug produced positive refinement gain in 16 of the 21 model-benchmark combinations and no change in the remaining five, with gains reaching +9.57 percentage points. For 11 of these combinations, the corresponding 95% confidence intervals excluded zero, supporting a positive direction of change within the evaluated combinations. In contrast, Self-Refine reduced correctness in 16 combinations, with losses of up to 8.07 percentage points, while none of its positive estimates had a confidence interval that excluded zero.
This contrast was driven not by repair capability alone, because Self-Refine also repaired initially incorrect candidates in many combinations. However, its repairs were frequently outweighed by regressions of candidates that were initially correct. Regressions occurred under Self-Debug in only three of the 21 model-benchmark combinations, with one regressed candidate in each. Under Self-Refine, regressions occurred in 17 combinations and exceeded repairs in every combination with negative refinement gain. This contrast shows that the positive or neutral gains of Self-Debug resulted largely from preserving initially correct candidates, whereas the regressions introduced by Self-Refine frequently offset its successful repairs.
The Self-Debug results further show that execution feedback does not by itself guarantee a successful repair. When the diagnostic tests passed, the models changed only six of 8,145 candidates, which largely explains the low regression rate. When the diagnostic tests exposed a failure, however, the models changed fewer than half of the affected candidates, and only 22.4% of those changes passed the final evaluation. The resulting gain therefore depended on whether a model both acted on the diagnostic failure and produced an effective revision.
The ability to use diagnostic feedback varied substantially across models. Notably, GPT-5.6 Sol had the highest initial pass rate on all three benchmarks but still achieved Self-Debug gains of +6.37 and +9.57 percentage points on MBPP+ and BigCodeBench-Instruct, respectively. This result shows that stronger initial generation did not necessarily eliminate the value of refinement. A high initial pass rate reduced the number of failures available for repair, but the realized gain also depended on how effectively the model converted diagnostic feedback into correct revisions.
The smaller gains of the locally served models were associated with their tendency to retain the initial candidate even after a diagnostic failure. To examine this behavior, we conducted a staged analysis of two models. The analysis indicated different bottlenecks in producing changed code and deciding whether revision was warranted.
The resource analysis leads to a similarly conditional conclusion. Self-Refine generally consumed more tokens than Self-Debug while also reducing correctness, making its additional inference difficult to justify in the evaluated settings. Self-Debug provided a more favorable refinement gain-overhead balance, but its practical value still varied across models and benchmarks. Token efficiency did not necessarily translate into lower API cost because model prices differed. Moreover, refinement may offer limited practical value when the initial generation already achieves sufficient correctness.
These findings do not support refinement as a default post-generation step. One-round refinement is practically useful only when feedback produces enough effective repairs while limiting regressions. The resulting gain must also justify the added computation and cost. Under the evaluated conditions, Self-Debug met these requirements more consistently than Self-Refine.
This paper makes the following contributions:
  • We provide a controlled empirical evaluation of the refinement gain produced by one-round Self-Refine and Self-Debug across seven models and three Python code-generation benchmarks.
  • We explain the observed refinement-gain patterns by analyzing repairs and regressions, model responses to diagnostic feedback, and model-specific bottlenecks in the transition from feedback to repair.
  • We assess the practical value of refinement by comparing refinement gain with token overhead and API cost, showing that positive gain does not necessarily imply favorable resource efficiency.
The remainder of this paper is organized as follows. Section 2 reviews related work, and Section 3 describes the research questions and experimental design. Section 4 presents the empirical findings, while Section 5 discusses their practical implications and examines model-specific self-debugging behavior. Section 6 addresses threats to validity, and Section 7 summarizes the findings and future directions.

2. Background

2.1. Iterative Refinement for Code Generation

Iterative refinement aims to improve an initial solution through one or more feedback-and-revision rounds. Madaan et al. proposed Self-Refine, a general generate–critique–revise framework in which a model improves its output using self-generated feedback [3]. For code generation, Chen et al. investigated self-debugging approaches that revise generated programs using execution outcomes, code explanations, or both [4]. Shinn et al. proposed Reflexion, which retains verbal feedback across attempts to guide subsequent reasoning and code generation [8]. Related program-generation studies have explored execution-guided decoding [9], code-specific training for self-correction [10], iterative self-debugging pipelines [11], and debugging based on intermediate runtime states [12].
Later studies examined the conditions under which refinement is effective. Olausson et al. evaluated whether LLMs can repair their own generated programs and found that repair effectiveness varies across models and task subsets [5]. Ding et al. observed that pretrained code models often have limited native self-refinement capability and proposed CYCLE to train models to use execution feedback more effectively [6]. Gu et al. examined whether code models can recognize, reason about, and repair subtle errors in their own generations and reported substantial limitations across these tasks [7]. Chen et al. revisited self-debugging with self-generated tests and showed that test-generation bias and the timing of execution feedback can affect debugging performance [13]. Zheng et al. further evaluated multi-turn code generation across multiple model families, programming benchmarks, prompting strategies, and inference budgets, including settings with reasoning instructions and execution feedback [14]. Broader analyses have also examined the general conditions required for self-correction [15] and the respective roles of model confidence and self-critique [16].
The reported outcomes vary across models, feedback settings, task subsets, and inference budgets. However, existing studies also differ in their refinement procedures, initial-candidate selection, models, benchmarks, and evaluation protocols. These differences make it difficult to determine the effect of one refinement round under matched experimental conditions.

2.2. Refinement-Based Code Generation and Repair Workflows

Researchers have also incorporated iterative refinement into broader code-generation and automated program repair workflows. Some studies improve refinement capability through model training. Jiang et al. proposed LeDex, which trains models to debug and explain code using execution feedback [17]. Jiang et al. later proposed ReflexiCoder, which uses reinforcement learning to improve self-reflection and correction of generated code [18].
Other studies combine code generation or reasoning with multiple inference-time components. Hu et al. proposed QualityFlow, an agentic workflow that integrates code generation, testing, debugging, and model-based quality assessment [19]. Code-generation test-time scaling methods have combined expanded candidate generation with execution-grounded verification and selection [20]. Broader test-time-compute research has examined the adaptive allocation of inference budgets according to problem difficulty [21]. These approaches use additional inference through generation, selection, verification, or refinement rather than evaluating an isolated refinement round.
Execution feedback also connects code-generation refinement with interactive coding and automated program repair. Yang et al. introduced InterCode, an interactive benchmark in which agents execute programs, observe execution feedback, and continue revising their solutions [22]. Zhao et al. proposed RePair, which uses compiler diagnostics and test feedback as process-based supervision for iterative program repair [23]. A recent review has also examined the use of LLMs in automated program repair [24]. Other studies have developed repository-level software engineering agents and platforms that allow LLMs to edit code and interact with development environments [25,26]. In contrast, Xia et al. proposed Agentless, which addresses repository-level issues through a fixed process of localization, repair, and patch validation without relying on interactive agent control [27]. These systems and workflows address broader generation and repair objectives, and their reported improvements may reflect additional sampling, search, verification, specialized training, tool use, or multiple feedback stages.

2.3. Benchmark and Evaluation Studies

Functional correctness benchmarks provide the primary basis for evaluating LLM-based code generation. Chen et al. introduced HumanEval, which evaluates generated Python functions using unit tests [1]. Austin et al. introduced MBPP, a collection of short Python programming problems with test-based correctness criteria [2]. Although both benchmarks are widely used, their original test suites may fail to identify incorrect solutions that satisfy only the available test cases. Liu et al. proposed EvalPlus, a general framework for strengthening code-generation evaluation through test augmentation, and introduced HumanEval+ [28]. The EvalPlus project also provides MBPP+, which applies the same test-augmentation approach to MBPP [29]. Zhuo et al. introduced BigCodeBench to extend code-generation evaluation to more complex instructions and diverse library function calls [30]. Other recent benchmarks have increased code-generation task difficulty [31] or directly evaluated unit-test generation and completion [32].
Other studies have examined threats arising from benchmark contamination and static evaluation. Riddell et al. quantified overlap between established code-generation benchmarks and public training corpora and showed that contamination can affect measured model performance [33]. Related evaluation research has examined contamination-detection methods [34] and temporally refreshed test sets [35]. Pan et al. compared static and interactive code-generation evaluations and found that access to feedback can change both model behavior and relative performance [36]. Measured performance can therefore vary with test coverage, benchmark difficulty, task distribution, contamination, and the feedback available during evaluation. For refinement studies, final-evaluation information should remain unavailable to model calls because access to it could directly influence the revised output.
Prior work has developed refinement methods, incorporated additional inference into broader code-generation and repair workflows, and strengthened the evaluation of generated programs. However, reported gains remain difficult to compare because studies differ in their initial candidates, feedback sources, inference stages, models, benchmarks, and evaluation protocols. To address this problem, we conduct a controlled evaluation of one-round Self-Refine and Self-Debug, using the same Direct candidate as the starting point for both protocols on each task. Holding the initial candidate, model, task, and final evaluation fixed reduces variation unrelated to refinement, while limiting each protocol to one refinement round avoids conflating its effect with repeated refinement, additional candidate sampling, or search. This design provides a consistent basis for evaluating the effect of a single refinement round.

3. Study Design

3.1. Research Questions

We investigate the effectiveness and practical value of one-round test-time refinement for code generation. Specifically, we examine feedback-free Self-Refine and diagnostic-test Self-Debug through the following research questions.
  • RQ1. How does one-round refinement affect code-generation correctness?
One-round refinement is useful only if revising an initial candidate produces a reliable improvement in functional correctness relative to retaining the candidate. We therefore compare the final-test outcomes of the initial candidate and its refined versions to measure the magnitude, direction, and uncertainty of the resulting change in pass rate. Both refinement protocols use the same initial candidate for each task, isolating the effect of refinement from variation in initial generation. We examine these effects across models and benchmarks because the available room for improvement and the ability to use refinement guidance may vary across settings.
  • RQ2. How do repairs and regressions shape refinement gain?
Refinement gain reflects both the correction of initially failing candidates and the degradation of initially passing candidates. A similar net change may therefore arise from substantially different balances between repair effectiveness and regression risk. To distinguish these effects, we compare the initial and revised final-test outcomes of each task and quantify how often each refinement protocol repairs failures or introduces regressions.
  • RQ3. How do models respond to diagnostic execution feedback in Self-Debug?
Unlike feedback-free Self-Refine, Self-Debug receives feedback produced by executing the initial candidate. Because this feedback is generated externally to the model, the available diagnostic evidence can be examined separately from the model’s subsequent response. This distinction enables us to investigate how models react to observable execution evidence rather than considering only the resulting correctness gain. We therefore examine whether models retain or revise the initial candidate in response to the diagnostic evidence and how these responses relate to final correctness.
  • RQ4. How does refinement gain compare with its resource overhead?
Refinement requires additional inference beyond initial candidate generation regardless of whether the revised candidate improves the final-test outcome. Its practical value therefore cannot be assessed from refinement gain alone and must also consider the resources consumed by the refinement workflow. We compare refinement gain with added token use across models and benchmarks and with experiment-time API cost for the OpenAI models. We also examine the resources consumed per net additional pass to determine how effectively the additional inference contributes to the final results.

3.2. Models

We evaluated seven models: four proprietary models accessed through the OpenAI API and three code-specialized open-weight models served locally. Table 1 summarizes their access methods, model categories, and publicly disclosed parameter counts. For mixture-of-experts models, the table reports both total and active parameters, while undisclosed parameter counts for OpenAI models are marked N/D (Not Disclosed). We selected models from different generations, cost levels, families, and architectures that could be evaluated reproducibly at the required scale within practical cost and hardware limits.
Among the OpenAI models, we selected GPT-3.5 as an earlier-generation reference because the GPT-3.5 generation is contemporaneous with much of the early work on self-refinement and self-debugging. We included GPT-4o mini as a lower-cost model from a more recent generation and GPT-4o as a general-purpose model from the same family. GPT-5.6 Sol represents the most recent OpenAI model in our evaluation and supports explicit reasoning effort.
Among the open-weight models, we selected Qwen2.5 as a compact, code-specialized dense model suitable for local inference. We included DeepSeek-V2 as a code-specialized mixture-of-experts model from a different family. Qwen3-Coder provides a later and larger model from the Qwen coding family, with 30.5B total and 3.3B active parameters, compared with the 7.61B-parameter dense Qwen2.5 model. We used its official FP8 checkpoint because the BF16 weights exceeded the available memory for local inference. The local model set supports both a within-family comparison across Qwen generations and a cross-family comparison.

3.3. Benchmarks

We selected three Python code-generation benchmarks to examine whether the effects of one-round test-time refinement vary across benchmark families. These include HumanEval+, MBPP+, and BigCodeBench-Instruct. Table 2 summarizes the number of tasks analyzed from each benchmark and their main task settings.
HumanEval [1] consists of function-level programming tasks specified through function signatures and docstrings, whereas MBPP [2] contains short programming problems expressed through natural-language descriptions and examples. We use HumanEval+ [28] and MBPP+ [29], which augment the original HumanEval and MBPP benchmarks with stronger test suites for functional correctness evaluation. The study includes all 164 HumanEval+ tasks. All 378 MBPP+ tasks were executed, but Mbpp/599 was excluded from the analysis because repeated executions produced conflicting final outcomes for the same candidate. These benchmarks have been widely used in prior code-generation evaluations.
BigCodeBench-Instruct [30] complements HumanEval+ and MBPP+ with instruction-based tasks involving a broader range of function calls and library usage. We excluded four of the 1,140 tasks because their reference solutions require live network or socket access and could fail for reasons unrelated to code correctness, reducing the execution set to 1,136 tasks. After completing the experiments, we performed a consistency check and identified eight tasks affected by flaky tests: identical candidate code received both passing and failing final-evaluation outcomes across separate executions. We excluded these eight tasks from BigCodeBench-Instruct and analyzed the resulting set of 1,128 tasks.

3.4. Refinement Protocols

We compare Direct with two one-round test-time refinement protocols: Self-Refine and Self-Debug. Direct provides the no-refinement baseline, while the two refinement protocols revise an initial candidate using different forms of guidance. Self-Refine uses a review generated by the model itself, whereas Self-Debug uses feedback from diagnostic test execution. For each task, Direct generates one initial candidate. This candidate is used as the Direct output and is also provided to Self-Refine and Self-Debug for refinement. Using the same initial candidate for both refinement protocols isolates the effect of refinement from variation in initial generation.

3.4.1. Direct

The Direct protocol consists of a single model call that requests one complete Python solution. The extracted code becomes the initial candidate and the Direct protocol output without model-generated review, diagnostic execution, or revision. It is subsequently evaluated using the final test set. Listing 1  shows the prompt used to generate the initial candidate.
Listing 1. Prompt used to generate the initial candidate.
You are solving a Python programming task.
 
Problem:
{{problem}}
 
Return only complete Python solution code.
Do not return markdown fences, explanation, or prose.
        

3.4.2. Self-Refine

The Self-Refine protocol examines whether a model can improve an initial candidate using its own assessment without execution feedback. It follows the review-and-revise structure of prior self-refinement work [3]. The review and revision are performed in separate calls, making the generated review an explicit input to the revision stage.
In the review call, the same model that generated the initial candidate receives the task specification and candidate code. It examines the candidate for correctness errors, unhandled edge cases, and specification mismatches. The candidate is not executed, and no diagnostic or final-test feedback is provided. The prompt focuses the review on functional correctness rather than style, optimization, or general code quality. Listing 2 shows the review prompt.
Listing 2. Self-Refine review prompt.
Review this Python solution for correctness, edge cases,
and specification mismatches.
Do not assume access to hidden tests.
 
Problem:
{{problem}}
 
Candidate solution:
{{candidate_code}}
        
In the revision call, the model receives the task specification, the initial candidate, and the generated review. It is instructed to modify the candidate only when the review identifies a concrete correctness issue and to make the smallest necessary change. If no such issue is identified, the model may return @@NO_CHANGE_NEEDED@@, in which case the initial candidate is retained as the protocol output. Listing 3 shows the revision prompt.
Listing 3. Self-Refine revision prompt.
Revise the Python solution using only the problem statement and review.
Do not use or infer hidden tests.
 
Preserve the candidate solution unless the review identifies a concrete correctness issue.
Make the smallest correctness-preserving change necessary.
Do not refactor, optimize, rename, or rewrite code unless required to fix correctness.
 
Problem:
{{problem}}
 
Candidate solution:
{{candidate_code}}
 
Review:
{{review_text}}
 
Return exactly one of the following:
 
* `@@NO_CHANGE_NEEDED@@` on a single line if no concrete correctness issue is identified.
* Complete Python solution code if a concrete correctness fix is required.
 
Do not return markdown fences, explanation, or prose.
Do not describe that no changes are needed; use only `@@NO_CHANGE_NEEDED@@` for that case.
        

3.4.3. Self-Debug

The Self-Debug protocol examines whether execution feedback from a predefined diagnostic test set can improve an initial candidate. It follows an execution-feedback-based debugging structure [4]. Unlike Self-Refine, which relies on a model-generated review, Self-Debug uses observed behavior from diagnostic execution to guide revision.
For each task, the initial candidate is executed against a diagnostic test set defined before any model calls and held fixed across all models. The set may contain one or more tests depending on the benchmark. Defining the diagnostic tests in advance prevents them from being selected in response to a particular candidate or its evaluation outcome.
The diagnostic execution result is converted into a bounded feedback summary. When one or more diagnostic tests fail, the summary may include the failing test, assertion information, expected and actual values, raised exceptions, and limited traceback information. When all diagnostic tests pass, the summary reports that no diagnostic failure was observed. Cases in which diagnostic execution does not produce a normal test result are retained and classified separately as diagnostic execution failures. The feedback summary provides the available diagnostic evidence without exposing the complete raw execution log.
A revision call is made for every candidate. The model receives the task specification, the initial candidate, and the diagnostic feedback. It is instructed to modify the candidate only when the feedback indicates a concrete correctness issue and otherwise retain the candidate. As in Self-Refine, the model may return @@NO_CHANGE_NEEDED@@, in which case the initial candidate is retained as the protocol output. Listing 4 shows the revision prompt.
Listing 4. Self-Debug revision prompt.
Revise the Python solution using the diagnostic test feedback below.
The feedback is from non-final diagnostic tests only.
Do not use or infer hidden tests.
 
Preserve the candidate solution unless the diagnostic feedback identifies a concrete correctness issue.
Make the smallest correctness-preserving change necessary.
Do not refactor, optimize, rename, or rewrite code unless required to fix correctness.
 
Problem:
{{problem}}
 
Candidate solution:
{{candidate_code}}
 
Diagnostic feedback:
{{diagnostic_feedback}}
 
Return exactly one of the following:
 
* `@@NO_CHANGE_NEEDED@@` on a single line if the diagnostic tests passed or no concrete correctness issue is identified.
* Complete Python solution code if a concrete correctness fix is required.
 
Do not return markdown fences, explanation, or prose.
Do not describe that no changes are needed; use only `@@NO_CHANGE_NEEDED@@` for that case.
        
In the prompt, non-final distinguishes feedback obtained during the diagnostic execution stage from information produced by the subsequent final evaluation. The benchmark-specific relationship between the diagnostic and final test sets is described in Section 3.5.

3.5. Diagnostic and Final Test Roles

We use diagnostic and final test sets for different purposes. The diagnostic test set is executed before the Self-Debug revision call, and its execution result is provided to the model as diagnostic feedback. The final test set is executed only after a protocol has produced its final candidate and is used to determine functional correctness. A candidate is considered correct only if it executes successfully and passes the complete final test set.
Distinguishing these roles prevents final evaluation results from directly guiding candidate revision. If feedback obtained from final evaluation were provided to the model, an observed improvement could reflect adaptation to the evaluation results rather than revision based on the designated diagnostic evidence. We therefore provide only feedback produced during the predeclared diagnostic execution. No test information or execution result obtained from the subsequent final evaluation is provided to any model call.
For HumanEval+ and MBPP+, the original benchmark tests form the diagnostic test set. The additional tests included in the augmented benchmark versions are reserved for final evaluation. The complete augmented suite, including both the original and additional tests, is used as the final test set, and a candidate must pass all of these tests to be considered correct.
BigCodeBench-Instruct does not provide separate original and augmented test sets that can be assigned to diagnostic execution and final evaluation. We therefore preselect one eligible unittest method from each task for diagnostic execution. Using one method provides execution evidence while limiting the portion of the full test suite used to generate diagnostic feedback. The remaining test methods are not used during diagnostic execution.
We retain the complete test suite provided by BigCodeBench-Instruct for final evaluation because test methods may share setup or mutable state, and removing the selected method could change the execution context of the suite. Consequently, the selected method is included in both the diagnostic and final test sets. All remaining test methods and all information produced during final evaluation remain unavailable to the model.

3.6. Experimental Settings

The OpenAI models were accessed through the OpenAI API, with one response requested per call. GPT-3.5, GPT-4o mini, and GPT-4o used temperature 0 to reduce sampling variation, although their API outputs were not assumed to be deterministic. GPT-5.6 Sol was accessed through the Responses API with medium reasoning effort and a maximum output budget of 32,768 tokens. Temperature was not specified for this model. Each protocol stage was executed as a separate request, and only the artifacts explicitly included in the prompts were passed between stages. For the OpenAI models, the API cost of each call was estimated from its recorded token usage using the model-specific prices at the time of the experiment. The resulting per-call token usage and estimated costs were retained in the experiment logs.
The local models were served separately using vLLM 0.24.0 with tensor parallelism across two NVIDIA RTX 4090 GPUs. Qwen2.5 and DeepSeek-V2 used BF16 checkpoints and an 8,192-token context length, whereas Qwen3-Coder used its official FP8 checkpoint and a 32,768-token context length. All local calls used temperature 0, seed 0, one response, and a maximum output length of 2,048 tokens. The output limit was applied as a uniform practical safeguard against generations that repeatedly produced the same content until exhausting the available token budget. It was not tuned separately for individual models or protocol stages. Outputs that reached the limit were retained without retry.
Candidates were executed on Ubuntu 22.04 using Python 3.12. Before the model experiments, we executed the benchmark reference solutions in the same environment to verify the prepared fixtures and determine resource limits that accepted valid solutions while bounding nonterminating or excessively resource-intensive candidates. Based on this validation, we used a 60-second timeout for HumanEval+ and MBPP+ and a default 120-second timeout for BigCodeBench-Instruct. Two BigCodeBench-Instruct tasks whose reference solutions required more than 120 seconds were assigned 300-second timeouts before candidate evaluation. All executions used a 4,096 MiB memory limit. All prepared reference solutions, including those for the 1,136 executed BigCodeBench-Instruct tasks, passed under these limits. During final evaluation, candidates that exceeded a time or memory limit were classified as incorrect, as were invalid programs, unhandled exceptions, and test failures.
The replication package [37] provides the execution scripts and detailed data used in the study.

4. Results

This section presents the analysis results and discusses their implications.

4.1. RQ1: Refinement Gain

Table 3 reports the refinement results for all seven models and three benchmarks. Each row represents one model-benchmark pair. The Direct column gives the pass rate of the initial candidate, while the Gain columns report the absolute percentage-point change after one round of Self-Refine or Self-Debug. We estimate paired bootstrap 95% confidence intervals by resampling benchmark tasks while preserving the correspondence between the initial and revised outcomes. An interval that excludes zero indicates a consistent direction of change within the evaluated setting, whereas an interval that includes or touches zero is not interpreted as clear evidence of improvement or degradation.
The results show a clear contrast between the two refinement protocols. Self-Debug produced a positive gain in 16 of the 21 model-benchmark pairs and a zero gain in the remaining five, with no negative gains. Its gains ranged from 0.00 to +9.57 percentage points, and 11 of the 16 positive estimates had confidence intervals that excluded zero. In contrast, Self-Refine reduced correctness in 16 pairs, improved it in four, and produced no change in one, with gains ranging from -8.07 to +1.22 percentage points. The four positive Self-Refine gains were limited to HumanEval+ for GPT-4o and GPT-5.6 Sol, and to MBPP+ and BigCodeBench-Instruct for Qwen3-Coder. However, all four confidence intervals included zero, providing no clear evidence that these positive estimates represent consistent improvements. On the other hand, eight of the 16 negative Self-Refine estimates had confidence intervals that excluded zero.
Among the OpenAI models, Self-Debug gain did not decrease monotonically as initial correctness increased. GPT-5.6 Sol had the highest Direct pass rate on all three benchmarks, yet it also produced the largest Self-Debug gains on MBPP+ and BigCodeBench-Instruct, at +6.37 and +9.57 percentage points, respectively. On HumanEval+, however, its initial pass rate was already 95.12%, and Self-Debug produced no additional gain. This result indicates that limited remaining failures can constrain refinement gain near the upper end of benchmark performance. Nevertheless, the substantial gains on the other two benchmarks show that a high initial pass rate does not necessarily imply a small refinement gain. The observed pattern is consistent with GPT-5.6 Sol making more effective use of execution feedback, although the experiment does not separately measure its ability to interpret the feedback and implement the required correction.
The earlier GPT-3.5 model showed the second-strongest overall Self-Debug pattern among the OpenAI models, with gains of +5.49, +2.39, and +4.70 percentage points on HumanEval+, MBPP+, and BigCodeBench-Instruct, respectively. After refinement, its pass rates increased to 69.51%, 61.54%, and 45.83%. These results remained below the initial pass rates of GPT-4o and GPT-5.6 Sol on every benchmark and approached that of GPT-4o mini only on MBPP+. Thus, execution feedback allowed the earlier model to repair a meaningful portion of its initial failures, but did not generally close the difference in initial generation performance. Refinement gain and final correctness should therefore be distinguished: a model can obtain a relatively large gain because it repairs more initial failures while still producing fewer correct solutions than a model with a stronger initial candidate.
The local models did not follow the same model-generation pattern. The earlier and smaller Qwen2.5 model obtained Self-Debug gains of +1.22, +0.27, and +2.22 percentage points, compared with 0.00, 0.00, and +0.53 percentage points for the later and larger Qwen3-Coder model. After refinement, Qwen2.5 reached 84.76%, 61.28%, and 42.65% on the three benchmarks. This exceeded the initial Qwen3-Coder pass rate on HumanEval+, but remained below its initial rates on MBPP+ and BigCodeBench-Instruct. A later and larger local model therefore did not necessarily obtain a larger refinement gain, while its higher initial correctness on MBPP+ and BigCodeBench-Instruct was still retained after the smaller model was refined.
The contrasting patterns across the OpenAI and local models indicate that Self-Debug gain cannot be explained by model recency, size, or initial pass rate alone. Its magnitude also varied across benchmarks. BigCodeBench-Instruct produced the largest Self-Debug gain for six of the seven models, with GPT-3.5 as the only exception. In contrast, HumanEval+ produced no gain for GPT-5.6 Sol, DeepSeek-V2, or Qwen3-Coder, while yielding the largest gain for GPT-3.5. The absence of gain for GPT-5.6 Sol on HumanEval+ is consistent with its limited remaining headroom, but the other differences cannot be attributed to initial pass rate alone. Because the benchmarks also differ in task distribution and diagnostic-test setting, these results suggest that the ability to benefit from execution feedback depends on the combination of model and benchmark rather than on a single model characteristic.
This dependence on the model-benchmark combination was also evident for Self-Refine, but without a similarly consistent pattern of positive gain. GPT-5.6 Sol, Qwen2.5, and Qwen3-Coder were the only models without a negative estimate whose confidence interval excluded zero. For GPT-5.6 Sol, the estimates ranged from +0.61 percentage points on HumanEval+ to small negative changes on the other two benchmarks. All three Qwen2.5 estimates were negative, but their confidence intervals included or touched zero, whereas Qwen3-Coder produced no negative gain. These models differ in generation, size, architecture, and initial correctness, and their results do not support a simple relationship between these characteristics and Self-Refine gain. Instead, the mixed outcomes indicate that feedback-free refinement is shaped by interactions among the model, its generated review and revision, and the benchmark task.
Answer to RQ1. One-round refinement shows a clear protocol-level contrast. Self-Debug produces positive or neutral gains across all evaluated settings, whereas Self-Refine predominantly reduces correctness and shows no clearly supported positive gain. Although the magnitude of gain varies across individual model–benchmark combinations without a simple pattern, this overall difference between the two protocols remains consistent.

4.2. RQ2: Repair and Regression Patterns

Refinement gain reports the net change in pass rate, but the same net result can arise from different balances between repairs and regressions. Figure 1 therefore decomposes the gain for each refinement protocol across all model–benchmark combinations. Each cell corresponds to one model-benchmark combination and displays repairs followed by regressions, while its background color represents ( repairs − regressions ) / N in percentage points. This decomposition shows whether a gain was produced by repairing initial failures, preserving initially correct candidates, or both.
The positive Self-Debug gains were characterized not by uniformly large repair counts, but by the near absence of regressions. Regressions occurred in only three of the 21 cells, and each of these cells contained only one regression; the remaining 18 cells contained none. Consequently, even a small number of repairs directly produced a positive net gain in cells such as GPT-4o on HumanEval+ (1/0), Qwen2.5 on MBPP+ (1/0), and Qwen3-Coder on BigCodeBench-Instruct (6/0). The larger gains followed the same pattern, including 53/0 for GPT-3.5 and 108/0 for GPT-5.6 Sol on BigCodeBench-Instruct. Thus, Self-Debug benefited not only from repairing failures but also from rarely offsetting those repairs by damaging candidates that already passed.
Self-Refine showed the opposite balance. Regressions occurred in 17 of the 21 cells and exceeded repairs in every cell with a negative gain. The clearest case was GPT-4o mini on BigCodeBench-Instruct, where 32 repairs were outweighed by 123 regressions, resulting in a gain of -8.07 percentage points. On MBPP+, the same model produced eight repairs but 27 regressions, yielding -5.04 percentage points. These results show that a substantial number of repairs was insufficient to produce a positive gain when revisions also damaged many initially correct candidates.
The four positive Self-Refine cells further emphasize the importance of regression control. Their repair counts were modest: 4, 1, 2, and 1, respectively. However, three contained no regression, and the remaining cell contained only two. Accordingly, the weakly positive gains did not result from a large repair volume, but from the repairs being only minimally offset by regressions. Although none of these gains was clearly supported by its confidence interval, their repair–regression balance contrasts with the negative cells in which regressions consistently dominated.
All six zero-gain cells contained 0/0 rather than equal numbers of repairs and regressions. Five of these cells occurred under Self-Debug, while only one occurred under Self-Refine. Thus, the zero gains resulted from no observed transitions between passing and failing final outcomes, rather than repairs and regressions canceling each other. However, this decomposition does not indicate whether the revised code was identical to the initial candidate or whether it changed without altering the final pass/fail result. The concentration of these cases under Self-Debug therefore motivates a closer examination of its code-change behavior.
Across both protocols, these results indicate that repair count alone does not determine refinement gain. Increasing the number of repaired failures remains useful, but every regression cancels the contribution of one repair. Avoiding harmful revisions can therefore have as much influence on the net result as producing additional repairs, and a more defensive revision strategy may be preferable when the available evidence does not clearly support a change. However, Figure 1 does not show whether the low regression rate of Self-Debug resulted from making accurate revisions or from frequently leaving the initial code unchanged. Distinguishing these mechanisms requires examining the diagnostic outcomes and code-change decisions preceding each repair or preserved candidate.
Answer to RQ2. Positive refinement gain depends not only on repairing initial failures but also on preserving initially correct candidates. Self-Debug produced positive or neutral gains largely because regressions were nearly absent, whereas Self-Refine often regressed more candidates than it repaired. Controlling regressions was therefore central to the observed repair–regression balance.

4.3. RQ3: Code Preservation and Repair in Self-Debug

The repair-regression decomposition showed that Self-Debug rarely regressed initially correct candidates. However, outcome transitions alone do not show whether this low regression rate resulted from making safe revisions or from avoiding revisions. We therefore analyzed how diagnostic outcomes affected code changes and how changes after a diagnostic failure led to repairs.
The analysis covers 11,683 model-task cases across all seven models and a common set of 1,669 tasks per model: 164 HumanEval+, 377 MBPP+, and 1,128 BigCodeBench-Instruct tasks. Each case was classified by its diagnostic outcome, whether the revised code differed from the initial Direct candidate, and its final-test transition. The Self-Debug protocol invokes the revision model for every case. The prompt instructs it to return @@NO_CHANGE_NEEDED@@ if the diagnostic tests pass or if no concrete correctness issue is identified; otherwise, it must return a complete revised solution.
Table 4 relates each diagnostic outcome to the initial final-test result and the subsequent code and correctness transitions. The Cases column gives the number of model–task cases in each branch. Changed reports whether the revised code differed from the initial candidate, while Repaired and Regressed report transitions between failing and passing final-test outcomes.
A diagnostic pass almost always led to the initial code being returned unchanged. Across the 8,145 cases in which the diagnostic execution passed, only six revised candidates differed from their initial candidates. Among the 6,317 initially correct candidates, only four changed, and three of those changes produced the only three regressions observed for Self-Debug. The low regression rate found in RQ2 therefore arose mainly because the protocol rarely changed an initially correct candidate after receiving a passing diagnostic result, rather than because it made many revisions without damaging correctness.
The diagnostic-pass cases also included 1,828 candidates that failed the hidden final evaluation, accounting for 34.1% of all 5,366 initial final-test failures. Only two of these candidates changed, and neither was repaired. Thus, under the evaluated prompt, the models almost never independently reviewed and altered a candidate once the diagnostic execution reported no failure, even when the hidden final tests would later expose an error. This behavior reduced exposure to regressions, but it also left failures not revealed by the diagnostic tests unaddressed. The experiment does not determine how many of these candidates could have been repaired if a more informative diagnostic signal had been available.
Figure 2 traces the pathway from an initial final-test failure to a repaired candidate for each model. Of the 5,366 initial failures, 3,519 produced a diagnostic failure that exposed a concrete problem to the model. The models changed 1,623 of these candidates (46.1%), and 363 of the changes repaired the final-test outcome. The repairs correspond to 22.4% of changed candidates, 10.3% of diagnostic failures, and 6.8% of all initial failures. The remaining 19 initial failures encountered diagnostic execution failures; four led to changed code, but none was repaired.
The proportion of initial failures accompanied by a diagnostic failure was relatively similar across models, ranging from 58.2% to 69.2%. The decision to change the code after receiving that signal differed much more sharply. The OpenAI models changed between 66.4% and 90.0% of candidates with a diagnostic failure, whereas the local models changed only 3.8% to 20.5%. Thus, the major separation between the two model groups occurred after the failure signal had already been observed.
Among the OpenAI models, GPT-5.6 Sol showed both the highest change rate and the most effective changes. It changed 306 of 340 candidates with a diagnostic failure (90.0%) and repaired 133 of those changes (43.5%). The other OpenAI models also changed code frequently, but only 15.4% to 20.8% of their changes resulted in repairs. The strong Self-Debug gain of GPT-5.6 Sol therefore reflected both its tendency to act on diagnostic failures and its comparatively high rate of converting those changes into correct solutions.
The local models showed a different bottleneck. Qwen2.5 changed 115 of 561 candidates with a diagnostic failure (20.5%), while Qwen3-Coder changed only 18 of 468 (3.8%). However, 28 of the Qwen2.5 changes (24.3%) and six of the Qwen3-Coder changes (33.3%) produced repairs. These conditional repair rates exceeded those of the three OpenAI models other than GPT-5.6 Sol. The limited gains of the Qwen models were therefore associated less with an inability to repair any changed candidate than with the small proportion of diagnostic failures that led to changed code. DeepSeek-V2 showed limitations at both transitions, changing 37 of 603 candidates (6.1%) and repairing three of those changes (8.1%).
The conditional repair rates should not be interpreted independently of the change rates. A model that changes only a small, selected subset of candidates may obtain a higher repair rate among those changes, and increasing its change frequency could reduce that rate. The results therefore do not establish that the Qwen models would retain their conditional repair effectiveness if they revised more candidates. They nevertheless show a clear tendency for the local models, particularly Qwen3-Coder, to leave candidates unchanged despite an explicit diagnostic failure signal.
The two diagnostic branches explain different aspects of the Self-Debug results. After a diagnostic pass, the protocol almost always avoided changing the candidate, which largely explains its low regression rate. After a diagnostic failure, gain depended on whether the model changed the code and whether that change repaired the hidden final-test outcome. The frequent zero or small positive gains of the local models are consistent with their low transition rate from diagnostic failure to changed code, while the larger gain of GPT-5.6 Sol combined frequent changes with a high repair rate. The analysis does not determine why a model returned unchanged code after a diagnostic failure, such as whether it failed to interpret the evidence or could not identify a suitable correction.
Answer to RQ3. Self-Debug avoided regressions mainly by almost never changing code after a diagnostic pass. After a diagnostic failure, the OpenAI models changed candidates much more frequently than the local models, and GPT-5.6 Sol also converted the largest proportion of its changes into repairs. The Qwen models produced repairs at relatively high rates when they made changes, but their infrequent changes limited their overall gains.

4.4. RQ4: Refinement Gain–Overhead Trade-offs

Figure 3 plots the total added tokens per task, including both input and output tokens, against the observed refinement gain. The Self-Refine and Self-Debug panels each show the 21 model-benchmark combinations for the corresponding protocol. The horizontal axis reports token overhead on a logarithmic scale, while the vertical axis reports the percentage-point change in correctness relative to Direct generation. The two panels separate Self-Refine and Self-Debug. Color identifies the model, shape identifies the benchmark, and filled points denote gains whose 95% confidence intervals exclude zero. The common axes allow the locations and distributions of the two protocols to be compared directly.
The two panels show a clear separation in the joint distribution of token overhead and refinement gain. The Self-Refine points tend to occupy the lower-right region, indicating greater token overhead together with predominantly negative refinement gain. In contrast, the Self-Debug points are generally located further to the left and above zero, indicating lower token overhead together with positive or neutral gain. Rather than presenting a conventional trade-off in which greater resource use is accompanied by greater gain, the overall pattern is largely one-sided: Self-Debug generally achieved more favorable refinement outcomes while using fewer additional tokens. The few positive Self-Refine estimates do not substantially alter this pattern because their confidence intervals include or touch zero.
Within the Self-Debug panel, points for each benchmark form distinct vertical bands on the logarithmic token axis. All BigCodeBench-Instruct triangles fall between 500 and 1,000 added tokens per task and form the rightmost band. The HumanEval+ circles form the middle band, while the MBPP+ squares form the leftmost band. Although token usage still varies across models within each band and the logarithmic scale compresses these differences, the benchmark-wise grouping remains clear. This pattern may reflect the limited sources of model-dependent variation in Self-Debug. For a given task, all models receive the same prompt templates, the generated candidate code and diagnostic feedback are likely to be broadly similar in length, and the response is restricted to code.
In contrast, the Self-Refine points are more widely dispersed along the token axis and show less distinct benchmark-wise grouping. Even points for the same benchmark differ substantially across models. Unlike Self-Debug, Self-Refine generates a free-form review whose length and content can vary by model. Because this review is included in the subsequent revision prompt, such variation affects both the output tokens of the review call and the input tokens of the revision call. Token overhead is therefore less consistent within each benchmark and is better assessed at the model-benchmark level.
Figure 4 examines the resource efficiency of the 11 Self-Debug model-benchmark combinations with positive refinement gain whose 95% confidence intervals exclude zero. The left panel reports token use for all qualifying combinations, while the right panel reports experiment-time API cost for the nine combinations using OpenAI models. White diamonds indicate the mean resource use among cases repaired by Self-Debug. The colored bar endpoints divide the resources consumed across all tasks by the number of repairs minus regressions, thereby reporting the total-workflow resource per net additional pass. Lower values indicate that fewer resources were required to obtain an additional passing solution.
In the token panel, the white diamonds are grouped more strongly by benchmark than by model. This pattern is consistent with the benchmark-wise token bands observed in Figure 3. Although the subsets of successfully repaired tasks are not identical across models, each benchmark constrains the task format and typical solution scope, which can limit variation in the prompt and generated-code lengths of successful repairs. The tokens required when a repair succeeds therefore remain broadly similar across models within each benchmark.
Compared with the white diamonds, the bar endpoints vary much more across models because they account for the tokens consumed by all Self-Debug attempts rather than only repaired cases. Attempts that leave an initial failure unrepaired still consume tokens, while regressions reduce the number of net additional passes. The bars therefore report the average workflow tokens required to obtain one net additional pass, with larger values indicating that more refinement effort did not translate into a net gain. GPT-5.6 Sol shows the lowest requirements, at approximately 4,800 tokens per net additional pass on MBPP+ and 7,100 on BigCodeBench-Instruct. It produced 24 and 108 net additional passes on these benchmarks, respectively. The other displayed OpenAI results range from approximately 7,400 to 27,000 tokens per net additional pass. On BigCodeBench-Instruct, the requirements increase to approximately 28,000 tokens for Qwen2.5 and 116,000 for Qwen3-Coder, which produced 25 and 6 net additional passes, respectively. These values are approximately four and sixteen times the requirement of GPT-5.6 Sol on the same benchmark. Thus, despite the broadly similar token use observed for repaired cases, GPT-5.6 Sol converted the overall Self-Debug workflow into net additional passes considerably more efficiently than the Qwen models.
The API-cost panel shows that token efficiency does not directly translate into monetary efficiency. The white diamonds show that repaired cases using GPT-5.6 Sol incur higher API costs than those using the other OpenAI models, even when their token requirements are similar. Its larger number of net additional passes partly offsets this higher API price. On MBPP+, GPT-5.6 Sol requires approximately $0.044 per net additional pass, compared with approximately $0.077 for GPT-4o. On BigCodeBench-Instruct, however, the ordering is reversed, with approximately $0.077 for GPT-5.6 Sol and $0.041 for GPT-4o. The displayed results for the lower-priced GPT-3.5 and GPT-4o mini range from approximately $0.002 to $0.008 per net additional pass. Thus, the stronger refinement performance of GPT-5.6 Sol improves its token efficiency but does not consistently compensate for its higher API price. The most cost-efficient model therefore varies across benchmarks.
These results distinguish two reasons for adopting Self-Debug. When obtaining the highest attainable pass rate is the primary objective, a model such as GPT-5.6 Sol can provide additional passes with relatively low token overhead, even when its API cost is higher. When the initial pass rate is sufficient and monetary cost is more important, a lower-priced model or Direct generation may be preferable. The local models incur no recorded API charge and may therefore remain practical when suitable hardware is already available, but their larger token requirements still represent additional inference time, energy use, and computational load.
Answer to RQ4. Self-Refine consistently used more tokens than Self-Debug and usually reduced correctness, making its additional overhead difficult to justify. Self-Debug generally provided positive refinement gain with lower token overhead, but the resources required per net additional pass varied substantially across models and benchmarks. A model with a higher repair rate could be more token-efficient without being more API-cost-efficient, so the practical value of Self-Debug depends on whether its additional passes justify the model- and benchmark-specific cost.

5. Discussion

The results show that test-time refinement is useful only when its additional information can be converted into a safe and effective repair. We first discuss the practical value of Self-Refine and Self-Debug, and then use the local-model mechanism study to examine where that conversion can fail.

5.1. When Do Self-Refine and Self-Debug Help?

The RQ1 results indicate that Self-Debug is a practical option when increasing the pass rate is the primary objective (Section 4.1). Its gains were consistently positive or neutral across the evaluated settings, largely because revisions rarely damaged initially correct candidates, as shown by the repair-regression analysis in RQ2 (Section 4.2). The GPT-5.6 Sol results are particularly informative: despite having the highest Direct pass rates, it also obtained the largest Self-Debug gains on two benchmarks. This finding does not support the initial expectation that refinement necessarily becomes less useful as the base model improves. Instead, the utility of Self-Debug appears to depend not only on the remaining failures but also on whether the model can convert execution feedback into effective repairs.
However, the RQ4 results show that a positive gain alone does not justify adopting Self-Debug in every setting (Section 4.4). The tokens consumed by successful repairs were broadly similar across models within each benchmark, but model pricing and the frequency of successful repairs produced substantial differences in monetary efficiency. In particular, GPT-5.6 Sol achieved the strongest refinement gains while also requiring comparatively high API expenditure. Because its Direct pass rate was already higher than the refined pass rates of the other evaluated models, paying for further improvement may provide limited practical value when that initial performance is sufficient. The decision to use Self-Debug should therefore depend on whether the value of its marginal additional passes exceeds the corresponding inference cost.
In contrast, the results do not support applying Self-Refine unconditionally. As shown in Figure 1, its poor net outcome did not always reflect a lack of repair capability. On BigCodeBench-Instruct, Qwen2.5 produced 23 repairs with Self-Refine, close to the 25 produced with Self-Debug, but its 33 regressions turned these repairs into a negative gain. For DeepSeek-V2, Self-Refine produced 13 repairs compared with only three under Self-Debug, but these were outweighed by 39 regressions. These cases suggest that Self-Refine could provide benefits comparable to Self-Debug, or potentially greater benefits in some settings, if paired with a reliable mechanism for retaining the initial solution or rejecting harmful revisions. However, the present study did not evaluate such regression-control mechanisms, and prior work likewise suggests that effective self-repair may require targeted training or more reliable feedback [5,6,7].

5.2. Why Do Local Models Struggle to Self-Debug?

The RQ3 analysis showed that the local models often retained an incorrect candidate even after a diagnostic failure exposed a concrete problem (Section 4.3). In particular, Figure 2 indicates that their limited gains were associated with low rates of changing the code after a failure signal became available. However, this aggregate behavior does not reveal whether a model failed to recognize the diagnostic outcome, chose not to patch the candidate, or could not produce an effective revision after deciding to patch. We therefore conducted a mechanism study using Qwen2.5 and DeepSeek-V2, two local models from different families that frequently left failed candidates unchanged. The study aimed to identify where the transition from diagnostic feedback to a repaired program was interrupted for each model.
We decomposed the transition from diagnostic feedback to repair into four stages: diagnostic-status recognition, Patch/Keep action selection, changed-code generation, and successful repair. For each model, we selected 100 cases in which the diagnostic execution and final evaluation failed but the revised code remained identical to the initial candidate, together with 100 diagnostic-pass controls whose initially correct candidates were preserved. The failure cases were balanced between explicit @@NO_CHANGE_NEEDED@@ responses and complete solutions that reproduced the initial code.
The recognition and action-selection probes asked the model to classify the diagnostic outcome and choose between Patch and Keep, respectively. To detect preferences for a particular response label or position, the meanings assigned to A and B were reversed across cases and for repeated presentations of the same evidence. In a final forced-patch probe, we bypassed the action-selection stage by explicitly instructing the model to revise every diagnostic-failure case rather than allowing it to choose between Patch and Keep. We then measured whether the returned code differed from the initial candidate and whether it passed the diagnostic and final evaluations. This design distinguishes failure to recognize the feedback, failure to select an appropriate action, and failure to produce an effective repair even when revision is explicitly required.
Table 5 summarizes the results of the four probe stages. The label-crossover row reports whether the model preserved the same diagnostic-status judgment when the meanings of A and B were reversed for identical evidence. The recognition rows show whether the model distinguished the 100 diagnostic failures from the 100 diagnostic-pass controls, while the action-selection rows show whether it chose Patch for the former and Keep for the latter. In the forced-patch stage, the choice between Patch and Keep was removed and every diagnostic-failure case was explicitly required to be revised. The final three rows therefore report whether this instruction produced changed code and whether the change repaired the diagnostic and final evaluations.
The two models showed different failure patterns. Qwen2.5 consistently interpreted the reversed label mappings, correctly recognized all diagnostic outcomes, and generally selected the appropriate action. However, even when explicitly required to revise every diagnostic-failure case, it returned changed code for only 32 cases. Seven of these changes passed the final evaluation, giving a conditional repair rate of 21.9%, similar to the 24.3% observed among its changed candidates in RQ3. Its main observed bottleneck therefore appeared to be producing changed code after recognizing a failure and selecting Patch, rather than repairing the candidate once a substantive change was produced. In contrast, DeepSeek-V2 selected Patch for 99 diagnostic-failure cases but also for 87 diagnostic-pass controls, indicating that its high Patch rate was not closely guided by the diagnostic outcome. When revision was explicitly required, it returned changed code for 51 cases, but only six passed the diagnostic tests and three passed the final evaluation. Subject to the exploratory interpretation of the DeepSeek-V2 probes, its weak outcome therefore involved both non-discriminative action selection and low repair effectiveness.
These results show that similar aggregate Self-Debug outcomes can arise from different failures within the refinement process. Simply strengthening the instruction to revise is therefore unlikely to address both models. For Qwen2.5, improvement may require better support for producing changed code after the model has recognized a diagnostic failure and selected Patch. For DeepSeek-V2, encouraging more Patch selections would not address its tendency to select Patch even for diagnostic-pass cases and could increase the risk of unnecessary changes to passing candidates. Its behavior instead suggests a need for better-calibrated Patch/Keep action selection and more effective repair generation after Patch is selected.

6. Threats to Validity

Internal validity may be threatened if final-test information inadvertently guides revision, unstable executions change the outcome of identical code, or implementation errors distort the protocol outputs. To prevent final evaluation from influencing refinement, diagnostic tests were fixed before generation, and no final-test output or execution trace was provided to any model call. All candidates were evaluated in a common environment with predefined resource limits, and the benchmark reference solutions were validated under the same conditions. Four BigCodeBench-Instruct tasks requiring external network or socket resources were excluded in advance. We also excluded eight BigCodeBench-Instruct tasks and Mbpp/599 after identical candidate code produced conflicting outcomes. These controls reduce the risk of interpreting information leakage or execution instability as a refinement effect, although undetected flaky behavior or implementation errors may remain. We release the experimental code, prompts, configurations, and result data in the replication package [37] to support independent inspection and reproduction.
Construct validity is primarily threatened by the choice of diagnostic tests and prompts. HumanEval+ and MBPP+ used their original tests, whereas BigCodeBench-Instruct used one preselected unittest method per task. Different diagnostic tests could produce different feedback and gains, and we did not evaluate alternative selections. However, the tests were fixed before generation and applied consistently across models, supporting comparisons within each benchmark. Similarly, we used one fixed prompt for each protocol stage without model-specific tuning. This consistent setup enables controlled comparisons, but the findings remain specific to the evaluated protocol implementations. Finally, our conclusions concern functional correctness under the designated final test suites, which may miss untested faults and do not measure other code-quality properties.
External validity may be limited because refinement behavior can vary across models, benchmarks, programming languages, and task scales. We evaluated seven models from multiple generations and deployment settings on HumanEval+, MBPP+, and BigCodeBench-Instruct, covering both short function-level problems and broader library-oriented tasks. This diversity supports comparisons across the evaluated settings, but it does not represent all models or repository-scale software development.
Prior exposure to widely used benchmarks is another concern. The training data of the hosted models cannot be inspected, and exposure to HumanEval or MBPP could affect initial generation, review, and repair performance. The augmented EvalPlus tests and the inclusion of BigCodeBench-Instruct reduce reliance on the original benchmark tests alone, but they cannot rule out benchmark exposure. The recurring protocol-level patterns across several models and benchmarks provide evidence beyond a single setting, although their exact magnitude may not generalize to other models, benchmarks, or programming tasks.
Conclusion validity may be limited because each model-task-protocol combination was executed only once, while OpenAI API outputs are not guaranteed to be deterministic. We used temperature zero where supported and fixed the local inference seed to reduce sampling variation, but repeated executions could still change individual outcomes and the resulting repair, regression, and gain estimates. The paired bootstrap confidence intervals capture uncertainty across benchmark tasks, but not variability from rerunning the model calls. We therefore report exact outcome counts with the confidence intervals and avoid treating estimates whose intervals include or touch zero as clear effects. The recurring directions across multiple model-benchmark combinations provide stronger evidence than isolated results, although the reported values remain based on one realized set of executions.

7. Conclusions

This study reassessed the utility of one-round feedback-free Self-Refine and diagnostic-test Self-Debug for code generation across seven models and three benchmarks. Self-Debug produced positive or neutral refinement gains across the evaluated settings, largely because diagnostic feedback enabled repairs while revisions rarely regressed initially correct candidates. Its effectiveness did not consistently decrease with stronger initial generation: GPT-5.6 Sol achieved both the highest Direct pass rates and the largest Self-Debug gains on two benchmarks. However, these gains did not always justify the additional monetary cost, particularly when Direct generation already provided sufficient correctness. In contrast, Self-Refine generally reduced correctness because its repairs were outweighed by regressions, while also requiring greater token overhead. The mechanism probes further showed that limited Self-Debug gain can arise from different model-specific bottlenecks, including failure to produce changed code after selecting Patch and poorly calibrated Patch/Keep action selection.
Our future work will investigate whether Self-Refine can become beneficial when harmful revisions are controlled through candidate retention, independent verification, or models trained specifically for critique and repair. We also plan to evaluate model-specific interventions for Self-Debug, including methods that better connect diagnostic failures to changed code and calibrate whether revision is warranted. More broadly, the results show that additional inference-time computation does not by itself improve code-generation correctness. One-round refinement is practically useful only when the feedback and revision process produces sufficient net gain to justify its additional cost and latency.

Funding

This research was supported by Seoul National University of Science and Technology.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The replication package, including the execution scripts and detailed data used in this study, is openly available on Figshare at https://doi.org/10.6084/m9.figshare.33128393.v1 [37].

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; de Oliveira Pinto, H.P.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. Evaluating Large Language Models Trained on Code. arXiv 2021, arXiv:2107.03374. [Google Scholar]
  2. Austin, J.; Odena, A.; Nye, M.; Bosma, M.; Michalewski, H.; Dohan, D.; Jiang, E.; Cai, C.; Terry, M.; Le, Q.; et al. Program Synthesis with Large Language Models. arXiv 2021, arXiv:2108.07732. [Google Scholar]
  3. Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-Refine: Iterative Refinement with Self-Feedback. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; Curran Associates, Inc., 2023; Vol. 36, pp. 46534–46594. [Google Scholar] [CrossRef]
  4. Chen, X.; Lin, M.; Schärli, N.; Zhou, D. Teaching Large Language Models to Self-Debug. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  5. Olausson, T.X.; Inala, J.P.; Wang, C.; Gao, J.; Solar-Lezama, A. Is Self-Repair a Silver Bullet for Code Generation? In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
  6. Ding, Y.; Min, M.J.; Kaiser, G.; Ray, B. CYCLE: Learning to Self-Refine the Code Generation. Proc. ACM Program. Lang. 2024, 8. [Google Scholar] [CrossRef]
  7. Gu, A.; Li, W.D.; Jain, N.; Olausson, T.; Lee, C.; Sen, K.; Solar-Lezama, A. The Counterfeit Conundrum: Can Code Language Models Grasp the Nuances of Their Incorrect Generations? In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024; Ku, L.W., Martins, A., Srikumar, V., Eds.; Bangkok, Thailand, 2024; pp. 74–117. [Google Scholar] [CrossRef]
  8. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: language agents with verbal reinforcement learning. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; Curran Associates, Inc., 2023; Vol. 36, pp. 8634–8652. [Google Scholar] [CrossRef]
  9. Wang, C.; Tatwawadi, K.; Brockschmidt, M.; Huang, P.S.; Mao, Y.; Polozov, O.; Singh, R. Robust Text-to-SQL Generation with Execution-Guided Decoding. arXiv 2018, arXiv:1807.03100. [Google Scholar]
  10. Cho, J.; Kang, D.; Kim, H.; Lee, G. Self-Correcting Code Generation Using Small Language Models. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025; Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Suzhou, China, 2025; pp. 2345–2368. [Google Scholar] [CrossRef]
  11. Adnan, M.; Xu, Z.; Kuhn, C.C.N. Large Language Model Guided Self-Debugging Code Generation. arXiv 2025, arXiv:2502.02928. [Google Scholar]
  12. Zhong, L.; Wang, Z.; Shang, J. Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step by Step. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024; Ku, L.W., Martins, A., Srikumar, V., Eds.; Bangkok, Thailand, 2024; pp. 851–870. [Google Scholar] [CrossRef]
  13. Chen, X.; Tao, Z.; Zhang, K.; Zhou, C.; Zhang, X.; Gu, W.; He, Y.; Zhang, M.; Cai, X.; Zhao, H.; et al. Revisit Self-Debugging with Self-Generated Tests for Code Generation. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Vienna, Austria, 2025; pp. 18003–18023. [Google Scholar] [CrossRef]
  14. Zheng, K.; Decugis, J.; Gehring, J.; Cohen, T.; Negrevergne, B.; Synnaeve, G. What Makes Large Language Models Reason in (Multi-Turn) Code Generation? In Proceedings of the International Conference on Learning Representations; Yue, Y., Garg, A., Peng, N., Sha, F., Yu, R., Eds.; 2025; Vol. 2025, pp. 40144–40181. [Google Scholar]
  15. Kamoi, R.; Zhang, Y.; Zhang, N.; Han, J.; Zhang, R. When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs. Trans. Assoc. Comput. Linguist. 2024, 12, 1417–1440. [Google Scholar] [CrossRef]
  16. Yang, Z.; Zhang, Y.; Wang, Y.; Xu, Z.; Lin, J.; Sui, Z. Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Vienna, Austria, 2025; pp. 3998–4014. [Google Scholar] [CrossRef]
  17. Jiang, N.; Li, X.; Wang, S.; Zhou, Q.; Hossain, S.B.; Ray, B.; Kumar, V.; Ma, X.; Deoras, A. LeDex: Training LLMs to Better Self-Debug and Explain Code. In Proceedings of the Advances in Neural Information Processing Systems; Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., Eds.; Curran Associates, Inc., 2024; Vol. 37, pp. 35517–35543. [Google Scholar] [CrossRef]
  18. Jiang, J.; Shen, J.; Kim, S.; Yoo, K.M.; Kim, J.; Kim, S. ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2026; Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; San Diego, California, United States, 2026; pp. 37543–37562. [Google Scholar] [CrossRef]
  19. Hu, Y.; Zhou, Q.; Chen, Q.; Li, X.; Liu, L.; Zhang, D.; Kachroo, A.; Oz, T.; Tripp, O. QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks. arXiv 2025, arXiv:2501.17167. [Google Scholar]
  20. Li, D.; Cao, S.; Cao, C.; Li, X.; Tan, S.; Keutzer, K.; Xing, J.; Gonzalez, J.E.; Stoica, I. S*: Test Time Scaling for Code Generation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025; Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Suzhou, China, 2025; pp. 15964–15978. [Google Scholar] [CrossRef]
  21. Snell, C.; Lee, J.; Xu, K.; Kumar, A. Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning. In Proceedings of the International Conference on Learning Representations; Yue, Y., Garg, A., Peng, N., Sha, F., Yu, R., Eds.; 2025; Vol. 2025, pp. 10131–10165. [Google Scholar]
  22. Yang, J.; Prabhakar, A.; Narasimhan, K.; Yao, S. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; Curran Associates, Inc., 2023; Vol. 36, pp. 23826–23854. [Google Scholar] [CrossRef]
  23. Zhao, Y.; Huang, Z.; Ma, Y.; Li, R.; Zhang, K.; Jiang, H.; Liu, Q.; Zhu, L.; Su, Y. RePair: Automated Program Repair with Process-based Feedback. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024; Ku, L.W., Martins, A., Srikumar, V., Eds.; Bangkok, Thailand, 2024; pp. 16415–16429. [Google Scholar] [CrossRef]
  24. Zhang, L.; Singhal, A.; Zou, Q.; Sun, X.; Liu, P.; Lin, H.Y. Can AI Fix Buggy Code? Exploring the Use of Large Language Models in Automated Program Repair. Computer 2025, 58, 122–128. [Google Scholar] [CrossRef]
  25. Yang, J.; Jimenez, C.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; Press, O. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In Proceedings of the Advances in Neural Information Processing Systems; Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., Eds.; Curran Associates, Inc., 2024; Vol. 37, pp. 50528–50652. [Google Scholar] [CrossRef]
  26. Wang, X.; Li, B.; Song, Y.; Xu, F.F.; Tang, X.; Zhuge, M.; Pan, J.; Song, Y.; Li, B.; Singh, J.; et al. OpenHands: An Open Platform for AI Software Developers as Generalist Agents. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  27. Xia, C.S.; Deng, Y.; Dunn, S.; Zhang, L. Demystifying LLM-Based Software Engineering Agents. Proc. ACM Softw. Eng. 2025, 2. [Google Scholar] [CrossRef]
  28. Liu, J.; Xia, C.S.; Wang, Y.; ZHANG, L. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; Curran Associates, Inc., 2023; Vol. 36, pp. 21558–21572. [Google Scholar] [CrossRef]
  29. EvalPlus Team. MBPP+. Hugging Face benchmark dataset. 2024.
  30. Zhuo, T.Y.; Vu, M.C.; Chim, J.; Hu, H.; Yu, W.; Widyasari, R.; Yusuf, I.N.B.; Zhan, H.; He, J.; Paul, I.; et al. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. In Proceedings of the International Conference on Learning Representations; Yue, Y., Garg, A., Peng, N., Sha, F., Yu, R., et al., Eds.; 2025; Vol. 2025, pp. 66602–66656. [Google Scholar]
  31. Yu, Z.; Zhao, Y.; Cohan, A.; Zhang, X.P. HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Task. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025; Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Vienna, Austria, 2025; pp. 13253–13279. [Google Scholar] [CrossRef]
  32. Jain, K.; Synnaeve, G.; Roziere, B. TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark. In Proceedings of the International Conference on Learning Representations; Yue, Y., Garg, A., Peng, N., Sha, F., Yu, R., Eds.; 2025; Vol. 2025, pp. 14947–14999. [Google Scholar]
  33. Riddell, M.; Ni, A.; Cohan, A. Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Ku, L.W., Martins, A., Srikumar, V., Eds.; Bangkok, Thailand, 2024; pp. 14116–14137. [Google Scholar] [CrossRef]
  34. Deng, C.; Zhao, Y.; Tang, X.; Gerstein, M.; Cohan, A. Investigating Data Contamination in Modern Benchmarks for Large Language Models. In Proceedings of the Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers); Duh, K., Gomez, H., Bethard, S., Eds.; Mexico City, Mexico, 2024; pp. 8706–8719. [Google Scholar] [CrossRef]
  35. Jain, N.; Han, K.; Gu, A.; Li, W.D.; Yan, F.; Zhang, T.; Wang, S.; Solar-Lezama, A.; Sen, K.; Stoica, I. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
  36. Pan, J.; Shar, R.; Pfau, J.; Talwalkar, A.; He, H.; Chen, V. When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025; Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Vienna, Austria, 2025; pp. 24672–24700. [Google Scholar] [CrossRef]
  37. Kim, J. Refinement Gain Study Replication Package. 2026. [CrossRef]
Figure 1. Repair and regression decomposition of one-round refinement gain. Each cell reports Repair/Regression: the number of initially failing candidates repaired and initially passing candidates regressed relative to the initial Direct candidate. The background color represents their net contribution to refinement gain in percentage points. Blue indicates a positive contribution, red indicates a negative contribution, and the horizontal dotted line separates the OpenAI and local models.
Figure 1. Repair and regression decomposition of one-round refinement gain. Each cell reports Repair/Regression: the number of initially failing candidates repaired and initially passing candidates regressed relative to the initial Direct candidate. The background color represents their net contribution to refinement gain in percentage points. Blue indicates a positive contribution, red indicates a negative contribution, and the horizontal dotted line separates the OpenAI and local models.
Preprints 227982 g001
Figure 2. Stagewise diagnostic and repair behavior by model. Each row begins with all initial final-test failures, with the total count shown above the bar. The first bar partitions these cases into diagnostic pass, diagnostic failure, and diagnostic execution failure. The following bars report how many candidates changed after a diagnostic failure and how many of those changes repaired the final-test outcome. Counts are normalized by the number of initial failures for each model, and diagnostic execution failures are excluded from the change–repair pathway. The panels separate the OpenAI and local models.
Figure 2. Stagewise diagnostic and repair behavior by model. Each row begins with all initial final-test failures, with the total count shown above the bar. The first bar partitions these cases into diagnostic pass, diagnostic failure, and diagnostic execution failure. The following bars report how many candidates changed after a diagnostic failure and how many of those changes repaired the final-test outcome. Counts are normalized by the number of initial failures for each model, and diagnostic execution failures are excluded from the change–repair pathway. The panels separate the OpenAI and local models.
Preprints 227982 g002
Figure 3. Refinement gain against added tokens per task. Each point represents one model-benchmark-protocol cell. Color identifies the model, shape identifies the benchmark, and fill indicates whether the 95% confidence interval excludes zero (filled) or includes/touches zero (hollow); panels separate the two protocols. BigCodeBench-Instruct gain uses the common 1,128-task analysis set, while token overhead reflects all executed calls.
Figure 3. Refinement gain against added tokens per task. Each point represents one model-benchmark-protocol cell. Color identifies the model, shape identifies the benchmark, and fill indicates whether the 95% confidence interval excludes zero (filled) or includes/touches zero (hollow); panels separate the two protocols. BigCodeBench-Instruct gain uses the common 1,128-task analysis set, while token overhead reflects all executed calls.
Preprints 227982 g003
Figure 4. Total-workflow resource per net additional pass and mean resource use among successful repairs for Self-Debug model-benchmark combinations with positive refinement gain whose 95% confidence intervals exclude zero. White diamonds and the first numeric values show the mean resource use among successful repairs; colored bar endpoints and the second values show total-workflow resource per net additional pass. Bar colors identify benchmarks as indicated by the legend. The left panel includes all qualifying models, while the API-cost panel includes only OpenAI models. Lower values indicate greater resource efficiency.
Figure 4. Total-workflow resource per net additional pass and mean resource use among successful repairs for Self-Debug model-benchmark combinations with positive refinement gain whose 95% confidence intervals exclude zero. White diamonds and the first numeric values show the mean resource use among successful repairs; colored bar endpoints and the second values show total-workflow resource per net additional pass. Bar colors identify benchmarks as indicated by the legend. The left panel includes all qualifying models, while the API-cost panel includes only OpenAI models. Lower values indicate greater resource efficiency.
Preprints 227982 g004
Table 1. Models evaluated in the study and their characteristics.
Table 1. Models evaluated in the study and their characteristics.
Modela Access Model category Parametersb
GPT-3.5 OpenAI API General-purpose N/D
GPT-4o mini OpenAI API Small general-purpose N/D
GPT-4o OpenAI API General-purpose N/D
GPT-5.6 Sol OpenAI API Reasoning N/D
Qwen2.5 Local Code-specialized, dense 7.61B
DeepSeek-V2 Local Code-specialized, MoE 16B / 2.4B active
Qwen3-Coder Local Code-specialized, MoE 30.5B / 3.3B active
a The specific model variants were gpt-3.5-turbo, gpt-4o-mini, gpt-4o, gpt-5.6-sol, Qwen2.5-Coder-7B-Instruct, DeepSeek-Coder-V2-Lite-Instruct, and Qwen3-Coder-30B-A3B-Instruct-FP8, respectively. b N/D indicates that the parameter count is not publicly disclosed. For the two mixture-of-experts models, total and active parameters are shown.
Table 2. Benchmarks evaluated in the study.
Table 2. Benchmarks evaluated in the study.
Benchmark Tasks Analyzed Task Setting
HumanEval+ 164/164 Function completion from signatures and docstrings
MBPP+ 377/378b Short problem-to-code generation
BigCodeBench-Instruct 1,128/1,140a Instruction-to-code with diverse function calls
a Four tasks were excluded before execution because their reference solutions require live network or socket resources. Eight additional tasks were excluded after execution because our consistency check identified flaky tests affecting their final evaluations. b Mbpp/599 was excluded after repeated executions produced conflicting final outcomes for the same candidate.
Table 3. One-round refinement results across seven models and three benchmarks. Gains are absolute percentage-point changes from the initial Direct candidate.
Table 3. One-round refinement results across seven models and three benchmarks. Gains are absolute percentage-point changes from the initial Direct candidate.
Group Model Benchmark Direct (%) Self-Refine Self-Debug
Gain 95% CI Gain 95% CI
OpenAI GPT-3.5 HumanEval+ 64.02 -3.66 [-8.54, +0.61] +5.49 [+2.44, +9.15]
MBPP+ 59.15 -6.90 [-10.34, -3.98] +2.39 [+1.06, +3.98]
BigCodeBench-I 41.13 -3.19 [-4.52, -1.86] +4.70 [+3.55, +6.03]
GPT-4o mini HumanEval+ 82.32 -5.49 [-10.98, 0.00] +3.05 [+0.61, +6.10]
MBPP+ 61.80 -5.04 [-8.22, -2.39] +1.33 [0.00, +2.65]
BigCodeBench-I 48.67 -8.07 [-10.28, -6.03] +4.34 [+3.19, +5.59]
GPT-4o HumanEval+ 85.98 +1.22 [-1.83, +4.27] +0.61 [0.00, +1.83]
MBPP+ 63.66 -1.59 [-3.18, -0.27] +1.06 [+0.27, +2.12]
BigCodeBench-I 52.22 -3.63 [-5.05, -2.30] +4.96 [+3.81, +6.21]
GPT-5.6 Sol HumanEval+ 95.12 +0.61 [0.00, +1.83] 0.00 [0.00, 0.00]
MBPP+ 72.41 -0.80 [-2.12, +0.53] +6.37 [+3.98, +9.02]
BigCodeBench-I 58.16 -0.62 [-1.33, +0.09] +9.57 [+7.89, +11.35]
Local Qwen2.5 HumanEval+ 83.54 -1.83 [-4.88, +0.61] +1.22 [0.00, +3.05]
MBPP+ 61.01 -1.86 [-3.71, 0.00] +0.27 [0.00, +0.80]
BigCodeBench-I 40.43 -0.89 [-2.22, +0.44] +2.22 [+1.42, +3.10]
DeepSeek-V2 HumanEval+ 74.39 -1.22 [-3.66, +1.22] 0.00 [0.00, 0.00]
MBPP+ 62.33 -1.06 [-2.12, -0.27] 0.00 [0.00, 0.00]
BigCodeBench-I 38.83 -2.30 [-3.55, -1.06] +0.27 [0.00, +0.62]
Qwen3-Coder HumanEval+ 78.66 0.00 [0.00, 0.00] 0.00 [0.00, 0.00]
MBPP+ 65.78 +0.53 [0.00, +1.33] 0.00 [0.00, 0.00]
BigCodeBench-I 49.47 +0.09 [0.00, +0.27] +0.53 [+0.18, +0.98]
Positive and negative gains are shown in blue and red, respectively. For gains whose 95% bootstrap confidence intervals exclude zero, the gain is shown in bold and the corresponding confidence interval is shown in the same color.
Table 4. Self-Debug behavior by diagnostic outcome and initial final-test outcome. For each outcome branch, Cases reports the number of model-task cases, Changed reports cases in which the revised code differed from the initial Direct candidate, and Repaired and Regressed report fail-to-pass and pass-to-fail final-test transitions, respectively. Diagnostic failure and diagnostic execution failure are treated as separate outcomes. Dashes indicate transitions that are not applicable to the corresponding initial final-test outcome.
Table 4. Self-Debug behavior by diagnostic outcome and initial final-test outcome. For each outcome branch, Cases reports the number of model-task cases, Changed reports cases in which the revised code differed from the initial Direct candidate, and Repaired and Regressed report fail-to-pass and pass-to-fail final-test transitions, respectively. Diagnostic failure and diagnostic execution failure are treated as separate outcomes. Dashes indicate transitions that are not applicable to the corresponding initial final-test outcome.
Diagnostic outcome Initial final outcome Cases Changed Repaired Regressed
Passed Passed 6,317 4 — 3
Failed 1,828 2 0 —
Diagnostic failure Failed 3,519 1,623 363 —
Execution failure Failed 19 4 0 —
Total 11,683 1,633 363 3
Table 5. Stagewise results of the local-model Self-Debug mechanism probes. Counts are reported over 100 diagnostic-failure cases or 100 diagnostic-pass controls, except for the paired label-crossover test.
Table 5. Stagewise results of the local-model Self-Debug mechanism probes. Counts are reported over 100 diagnostic-failure cases or 100 diagnostic-pass controls, except for the paired label-crossover test.
Stage Measure Qwen2.5 DeepSeek-V2
Label crossover Semantically consistent pairs 40/40 35/40; 71/80a
Recognition Failures classified as diagnostic failure 100/100 90/100
Pass controls classified as diagnostic pass 100/100 100/100
Action selection Failures assigned Patch 81/100 99/100
Pass controls assigned Keep 100/100 13/100
Forced patch Returned changed code 32/100 51/100
Passed diagnostic tests 10/100 6/100
Passed final evaluation 7/100 3/100
a The DeepSeek-V2 crossover was expanded from 40 to 80 pairs after the initial result. Its semantic consistency remained sensitive to the A/B mapping, so its subsequent probe results are interpreted as exploratory.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.