Submitted:
14 September 2026
Posted:
15 September 2026
You are already at the latest version
Abstract
Static software engineering benchmarks such as SWE-bench suffer from three fundamental limitations: task sets are fixed and become saturated as model capabilities rapidly improve; expert construction is costly and slow; and difficulty levels cannot adapt to evolving model abilities. We present SEBGen, a self-evolving benchmark generation framework that synthesizes evaluation tasks from code repositories and improves its generation strategies through an evolution loop. SEBGen introduces three key innovations: an explicit five-dimension exploration value evaluator that quantitatively scores generation strategies across quality, novelty, difficulty coverage, skill diversity, and evolution potential, replacing implicit LLM judgments; a five-layer hybrid validation pipeline combining mechanical checks with LLM-as-judge evaluation; and a self-evolving loop with archive-based experience accumulation, sigmoid-scaled parent selection, and LLM-driven strategy mutation. On a synthetic corpus, SEBGen achieves a 96 percent validation pass rate with stable quality across five evolution generations. On a real repository, it transforms all six tqdm bug-fix commits into validated tasks, and a comparison against SWE-bench Lite shows task quality of 0.76 approaching expert baselines of 0.80. These results indicate that dynamic, self-evolving benchmarks are a viable alternative to static benchmarks.
Keywords:
benchmark generation
; self-evolving agents
; LLM-as-judge
; software engineering evaluation
; curriculum learning
1. Introduction
Benchmark datasets have driven progress in software engineering AI, with SWE-bench [1] establishing the standard for evaluating models on real-world GitHub issues. However, the rapid advancement of large language models (LLMs) has exposed fundamental limitations of static benchmarks, summarized in Table 1.
This work addresses four research questions (RQs):
- RQ1: Can LLMs automatically generate verifiable software engineering benchmark tasks from code repositories?
- RQ2: Does difficulty adaptation (ZPD-based selection) improve benchmark effectiveness?
- RQ3: Does the evolution loop improve or degrade task quality over generations?
- RQ4: Can generated benchmarks differentiate agent capabilities?
Our contributions are as follows:
- 1.
- Explicit Exploration Value Assessment: A five-dimension scoring model (Quality, Novelty, Difficulty Coverage, Skill Diversity, Evolution Potential) that quantifies the exploration value of generation strategies, replacing Voyager’s [2] implicit LLM judgment with interpretable, optimizable metrics.
- 2.
- Five-Layer Hybrid Validation: A cascading validation pipeline combining mechanical checks (schema, compile, functional) with LLM-based semantic evaluation (quality, semantic).
- 3.
- Self-Evolution Loop: An archive-based experience system with sigmoid-scaled parent selection, novelty reward, and LLM-driven strategy mutation.
- 4.
- Empirical Validation: Eight experiments spanning a synthetic corpus, a real repository (tqdm), and a gold-standard comparison against SWE-bench Lite, covering generation quality, evolution stability, difficulty adaptation, contamination tracking, component-level ablation, agent differentiation, and expert-quality comparison.
2. Related Work
2.1. Static Software Engineering Benchmarks
SWE-bench [1] pioneered the use of real GitHub issues as evaluation tasks, with each instance containing a problem statement, test patch, and FAIL_TO_PASS/PASS_TO_PASS test specifications. While highly influential, its static nature means tasks are fixed after creation. Automated successors of SWE-bench are discussed in Section 2.3.
2.2. Self-Evolving Agents
Voyager [2] introduced automatic curriculum learning and skill library accumulation for embodied agents, using implicit LLM judgments to assess exploration value. The Darwin Gödel Machine (DGM) [3] established the archive-based evolution paradigm with parent selection and verification-driven loops. Reflexion [4] demonstrated that verbal self-reflection without weight updates can improve agent performance. ReAct [7] established the reasoning–acting paradigm underlying modern LLM agents. A recent survey of self-evolving agents [8] provides a systematic taxonomy of these lines of work.
2.3. Automated Benchmark Generation
Two recent industrial efforts have targeted automated benchmark generation at scale. Auto-SWE-Bench [5] (Turing, NeurIPS 2025 Exhibitor Track) continuously sources GitHub issues and pull requests to produce multilingual datasets with reproducibility guarantees. SWE-Bench++ [6] extends this line with a four-stage pipeline—programmatic PR sourcing, environment synthesis, state-differential test-oracle extraction, and a four-layer quality assurance scheme (environment determinism, oracle consistency, LLM-judge semantic alignment, and false-negative filtering)—yielding 11,133 verified instances across 3,971 repositories and 11 languages.
SEBGen differs from these systems in two fundamental ways. First, SWE-Bench++ optimizes for static scale: it produces a large, one-shot corpus (with contamination-aware temporal splits for evaluation). SEBGen instead optimizes for dynamic freshness: an evolution loop continuously regenerates tasks as agent capabilities grow, with archive-based novelty tracking preventing saturation. Second, SWE-Bench++’s quality assurance is correctness-oriented (all four layers verify environment and oracle integrity), whereas SEBGen augments mechanical verification with semantic quality evaluation (LLM-as-judge clarity/solvability/novelty scoring) and an explicit five-dimension exploration value model that scores the generation strategy itself—a level of meta-evaluation absent from prior automated generators. A detailed comparison is provided in Section 5.
2.4. Comparative Positioning
Table 2 positions SEBGen relative to prior work.
3. Materials and Methods
3.1. System Architecture
SEBGen consists of six core components organized in a closed evolution loop, as illustrated in Figure 1.
The system follows six design principles, summarized in Table 3.
3.2. Task Generator
The Task Generator extracts candidate commits from code repositories and synthesizes structured benchmark tasks using LLMs with five strategies: BUG_FIX, FEATURE_ADD, REFACTOR, TEST_DRIVEN, and HYBRID. The generation prompt integrates five information sources: repository context, commit metadata, diff content (truncated to 3000 characters), archive experience (top-3 similar successful templates), and structured JSON output constraints. To compensate for LLM generation instability, the generator oversamples commits at the target task count and selects the first N successful generations.
3.3. Five-Layer Quality Validator
The validation pipeline cascades through five layers, summarized in Table 4.
For each FAIL_TO_PASS test, the L3 functional validator verifies two phases: (1) the test fails on the original code state, and (2) the test passes after applying the reference solution. This dual-phase protocol eliminates “vacuous” tests that pass without any fix.
3.4. Difficulty Adapter (ZPD-Based)
Based on Vygotsky’s Zone of Proximal Development, the adapter maintains a model capability profile and dynamically adjusts the ZPD range:
where width adapts to recent success rates: base width when success (tasks too easy), and when success (tasks too hard). Curriculum sorting interleaves tasks (easy–medium–easy–medium–hard pattern) with hard tasks sparsely distributed.
3.5. Explicit Exploration Value Evaluator
The five-dimension scoring model quantifies the exploration value of each generation strategy:
The five dimensions are defined in Table 5.
The weights reflect a deliberate prioritization: quality (0.30) dominates because benchmark validity is the primary requirement; evolution potential (0.20) outweighs novelty, difficulty coverage, and skill diversity because a strategy’s ability to keep improving over generations determines the loop’s long-term value. These weights are treated as tunable hyperparameters; their sensitivity is flagged in Section 5.
The Evolution Potential dimension implements a particularly important insight: strategies with clear, identifiable failure patterns receive high potential scores (0.80) because they can be improved through targeted prompt refinement, whereas strategies with uniformly high pass rates () receive low scores (0.20) due to limited headroom.
3.6. Evolution Engine with Parent Selection and Strategy Mutation
Parent Selection combines sigmoid-scaled performance with novelty reward:
where is the sigmoid function applied to z-scored performance values (amplifying differences when standard deviation , degenerating to 0.5 when all strategies perform identically), and weights exploitation over exploration.
Strategy Mutation implements four genetic operators: (1) prompt mutation—LLM analyzes failure/success cases to improve the generation prompt; (2) parameter mutation—Gaussian perturbation of temperature, max_tokens, and top_p; (3) crossover—merges the higher-performing parent’s prompt with averaged parameters; and (4) random exploration—-greedy generation of entirely new strategies.
3.7. Archive Manager
The archive stores multi-dimensional experience with four index structures: vector similarity (semantic search), strategy type, difficulty level, and domain. Key operations include semantic similarity search for novelty computation, failure querying for evolution learning, and success querying for generation reference.
4. Results
4.1. Experimental Setup
The corpus comprises 15 bug-fix scenarios across 7 domains (error handling, concurrency, memory, math, validation, security, algorithm). The generation LLM is DeepSeek (deepseek-chat). All five validation layers are active. The environment is Python 3.10 with SQLite storage and local pytest execution; Docker-based evaluation is deferred to cloud deployment.
Reporting convention: Because sample sizes are modest (15 synthetic scenarios, 6 real commits, 10 sampled expert tasks), all reported figures are point estimates; we do not claim statistical significance unless explicitly stated, and we flag this as a validity limitation in Section 5.
4.2. E1: Generation Quality Comparison (RQ1)
Table 6.
E1: generation quality comparison across strategies (RQ1).
| Strategy | Gen. Success | Pass Rate | Avg Score | Expl. Value |
|---|---|---|---|---|
| BUG_FIX | 100% (10/10) | 96% (48/50) | 0.88 | 0.850 |
| FEATURE_ADD | 100% (10/10) | 96% (48/50) | 0.90 | 0.838 |
Finding: Both strategies achieve a 96% validation pass rate with 100% generation success, confirming RQ1: LLMs can automatically generate verifiable benchmark tasks. BUG_FIX exhibits a slightly higher exploration value due to superior novelty scores. Only the two primary strategies are reported here; a comparison of all five strategies is deferred to future work.
4.3. E2: Evolution Effectiveness (RQ3)
Table 7.
E2: evolution effectiveness across five generations (RQ3).
| Gen | Quality | Novelty | DiffCov | SkillDiv | EvoPot | Overall |
|---|---|---|---|---|---|---|
| 1 | 0.91 | 0.46 | 0.68 | 0.61 | 0.20 | 0.744 |
| 2 | 0.91 | 0.48 | 0.68 | 0.61 | 0.20 | 0.743 |
| 3 | 0.91 | 0.45 | 0.68 | 0.61 | 0.20 | 0.729 |
| 4 | 0.91 | 0.44 | 0.68 | 0.61 | 0.20 | 0.738 |
| 5 | 0.92 | 0.47 | 0.68 | 0.61 | 0.20 | 0.739 |
Finding: Quality remains stable (0.91–0.92) across five generations. We interpret this primarily as a stability result: the evolution loop does not cause the quality degradation that commonly afflicts iterated LLM self-improvement loops. The overall exploration value is effectively flat (0.73–0.74), which we interpret as an honest account of the loop’s behavior on a homogeneous corpus: when all strategies already perform well, there is little headroom for improvement. In a separate controlled experiment with mixed-quality inputs (2 high + 2 medium + 2 low), the evolution loop improved low-tier task quality from 0.79 to 0.85 (+0.06) across four generations, confirming that evolution is most effective when clear failure patterns exist. Together, these results suggest RQ3 should be answered conditionally: the loop preserves quality and improves it only when it has exploitable failure signal.
4.4. E3: Difficulty Adaptation (RQ2)
With a simulated agent profile (4/5 success at difficulty ≈0.48), the ZPD range was computed as . Of 6 test tasks, 3 fell within ZPD, 2 below (too easy), and 1 above (too hard).
Finding: ZPD-based selection successfully identifies appropriately challenging tasks. However, rule-based difficulty estimation (weighted code metrics) shows limited discriminative power for short diffs, motivating future work on learned difficulty predictors. We therefore treat the RQ2 result as a preliminary demonstration of the mechanism rather than evidence of its benefit.
4.5. E4: Anti-Contamination
Table 8.
E4: novelty score as a function of archive size.
| Archive Size | 0 | 5 | 10 | 20 | 30 | 50 |
| Novelty | 1.000 | 0.950 | 0.932 | 0.904 | 0.897 | 0.882 |
Finding: Novelty decays monotonically with archive growth, validating the contamination detection mechanism. At 50 tasks, novelty remains at 0.882, indicating sufficient task diversity. The mechanism can prevent benchmark saturation by flagging near-duplicate tasks before inclusion.
4.6. E5: Ablation Study
Table 9.
E5: ablation study of validation layers.
| Configuration | Layers | Avg Score | vs Full |
|---|---|---|---|
| Full system | 5 | 0.92 | — |
| Without LLM layers | 3 | 1.00 | +0.08 |
| Without functional | 4 | 0.91 | −0.01 |
| Schema only | 1 | 1.00 | +0.08 |
Finding: The LLM-based layers (Quality + Semantic) are the most discriminating components: removing them inflates average scores from 0.92 to a ceiling of 1.00, because mechanical checks alone cannot identify semantically poor tasks. This result should be read as evidence that the LLM layers perform meaningful quality discrimination, not that mechanical validation suffices.
4.7. E6: Benchmark Differentiation (RQ4)
To verify that SEBGen-generated benchmarks can differentiate agent capabilities, we evaluated three agents of varying sophistication on six validated tasks: an LLM-Strong agent (DeepSeek instructed to produce complete bug-fix solutions), a Script-Weak baseline (keyword-based heuristic descriptions without executable code), and a Random baseline.
Table 10.
E6: benchmark differentiation across agent baselines (RQ4).
| Agent | Resolved | Pass Rate |
|---|---|---|
| LLM-Strong (DeepSeek) | 6/6 | 100% |
| Script-Weak (heuristic) | 0/6 | 0% |
| Random baseline | 0/6 | 0% |
Finding: The benchmark separates the capable LLM agent from both baselines. However, because the two weak baselines are deliberately trivial, we interpret this as a sanity check that the benchmark is not accidentally solvable by chance, rather than as evidence of fine-grained discriminative power. Confirming RQ4 properly requires a graded panel of agents (e.g., a smaller LLM, a constrained LLM, or a non-LLM retrieval-augmented agent), which we defer to future work.
4.8. E7: Real Repository Experiment (tqdm)
To validate SEBGen beyond synthetic corpora, we ran the full pipeline on six real bug-fix commits extracted from the tqdm project (a widely used Python progress bar library), including fixes for AttributeError in close(), dynamic miniters smoothing, and concurrent processing time estimation.
Table 11.
E7: real-repository validation on tqdm bug-fix commits.
| Commit | Fix Description | Difficulty | Quality | All Layers |
|---|---|---|---|---|
| 31ab0f4 | dynamic_miniters & smoothing | medium | 0.74 | ✔ |
| 494372e | AttributeError in close() | easy | 0.78 | ✔ |
| 62f3b4a | Concurrent time estimation | medium | 0.75 | ✔ |
| 0eaee96 | Length detection | medium | 0.78 | ✔ |
| 40cad84 | Packaging: exclude images | easy | 0.77 | ✔ |
| b78624a | Logging: preserve filters | easy | 0.76 | ✔ |
Finding: All 6/6 real commits were successfully transformed into validated benchmark tasks (100% generation success, 100% five-layer pass rate). The average quality of 0.76 is lower than that of the synthetic corpus (0.88–0.90), reflecting the greater complexity and nuance of real-world diffs. This confirms that SEBGen generalizes from synthetic to real repository inputs.
4.9. E8: Gold-Standard Comparison with SWE-bench
To benchmark SEBGen against expert-constructed tasks, we loaded the SWE-bench Lite test split (300 real GitHub-issue tasks across 12 repositories, including django, sympy, scikit-learn, and matplotlib) and evaluated both expert tasks and SEBGen-generated tasks using the same LLM-judgeable layers (L1 + L4 + L5; L2/L3 require full-repo Docker environments unavailable locally).
Table 12.
E8: gold-standard quality comparison against SWE-bench Lite.
| Task Source | Sample | Avg Quality (L1+L4+L5) |
|---|---|---|
| SWE-bench expert tasks | 10 random real issues | 0.80 |
| SEBGen synthetic (E1–E2) | 15-bug corpus | 0.90 |
| SEBGen real-repo (E7) | 6 tqdm commits | 0.76 |
Finding: SEBGen-generated tasks from real commits (0.76) achieve quality comparable to expert-constructed SWE-bench tasks (0.80). We emphasize that this is a descriptive comparison on small samples (10 expert tasks, 6 generated tasks) with no statistical test; the 0.04 gap is within plausible sampling noise and should not be interpreted as equivalence or superiority in either direction. Notably, 9/10 expert tasks passed all three LLM-judgeable layers, indicating our evaluator is not unfairly penalizing high-quality expert tasks. The synthetic corpus scores higher (0.90) than both real sources because short, well-scoped diffs are inherently easier to validate than complex production issues—a property that reflects the evaluator’s tendency to reward clarity over difficulty, which we flag as a metric-validity concern in Section 5. This comparison establishes that automated generation approaches expert-level benchmark task quality without manual expert construction, subject to the caveats above.
5. Discussion
5.1. Design Principle Validation
The six design principles proved effective: modular isolation (P1) enabled independent testing of 9 components (124 unit tests); verification-driven design (P2) maintained 96% task quality; archive accumulation (P3) enabled novelty tracking; curriculum progression (P4) filtered tasks to ZPD; exploration-exploitation balance (P5) produced stable evolution without degradation; containerized isolation (P6) interfaces were designed for cloud deployment.
5.2. Comparison with SWE-Bench++
Given the close topical overlap, we provide an explicit comparison with SWE-Bench++ [6], the closest prior automated benchmark generator. The two systems optimize for different objectives and occupy complementary positions:
Table 13.
Design comparison between SEBGen and SWE-Bench++.
| Dimension | SWE-Bench++ | SEBGen (Ours) |
|---|---|---|
| Primary objective | Static scale (11,133 tasks, 11 languages) | Dynamic freshness (evolving task set) |
| QA focus | Correctness (environment determinism, oracle consistency) | Correctness + semantics (LLM judge clarity, solvability, novelty) |
| Meta-evaluation | None | 5-dimension exploration value of generation strategies |
| Adaptivity | Contamination-aware temporal splits | ZPD-based difficulty adaptation per agent |
| Self-improvement | None (one-shot pipeline) | Evolution loop: parent selection + prompt mutation + archive |
| Validation scale | 3,971 repositories | 1 repository (tqdm) + synthetic corpus (scale is our main limitation) |
The two approaches are complementary rather than competing: SWE-Bench++ demonstrates that automated generation scales to production-grade multilingual corpora, while SEBGen demonstrates that generation can be made self-improving—a property SWE-Bench++ does not possess. The natural synthesis, which we identify as a primary direction for future work, is to embed the SEBGen evolution loop inside a SWE-Bench++-scale pipeline: their sourcing and environment synthesis would provide the massive, diverse initial population, while our exploration value evaluator and evolution engine would keep the resulting benchmark fresh as model capabilities advance. We also note that SWE-Bench++’s reported model performance (e.g., 36.20% pass@10 for claude-sonnet-4.5 on a 1,782-instance subset) provides a valuable calibration point for future large-scale evaluations of SEBGen-generated tasks.
5.3. Limitations
- 1.
- Scale: Real-repository validation was conducted on a single project (tqdm, 6 commits); multi-project experiments at scale (100+ commits across 10+ repositories) remain future work.
- 2.
- Difficulty Estimation: Rule-based difficulty scoring shows limited discriminative power for short diffs (all easy-labeled). A learned predictor is needed.
- 3.
- Docker Dependence: Full Layer 2–3 validation requires containerized isolation; the local pytest fallback provides weaker guarantees.
- 4.
- Embedding-Based Novelty: Jaccard token overlap is a proxy for semantic similarity; sentence-transformers embeddings would improve precision.
- 5.
- LLM Judge Bias and Circularity: LLM-as-judge evaluation inherits model biases; more importantly, both the generator and the judge rely on the same model family (deepseek-chat), creating a self-referential evaluation loop that can systematically overestimate quality. Human annotation of a validation subset, and ideally a judge from a different model family, are needed for calibration.
- 6.
- Small Samples Without Statistics: The empirical claims rest on small samples (15 synthetic scenarios, 6 real commits, 10 sampled expert tasks) with no confidence intervals or significance tests; several headline gaps (e.g., 0.76 vs. 0.80) may be noise.
5.4. Threats to Validity
Internal validity: LLM generation is stochastic; we mitigate this through oversampling and multiple runs, but do not report variance across runs. External validity: Synthetic diffs may not represent real-world commit patterns; the tqdm experiment is a single-repository, six-commit probe. Construct validity: The exploration-value dimensions are our operationalization of “good benchmark task”; both the dimension weights (0.30/0.20/0.15/0.15/0.20) and the Jaccard-based novelty proxy are design choices whose sensitivity has not been explored. Criterion validity: The LLM judge’s quality score is not anchored to human judgments, and the E8 result showing synthetic tasks outscoring expert tasks suggests the metric may reward clarity over difficulty; human calibration is required before the metric can be treated as a reliable quality signal.
6. Conclusions
We presented SEBGen, a self-evolving benchmark generation framework addressing the staticity, manual dependency, and adaptivity limitations of traditional software engineering benchmarks. Three core innovations—explicit five-dimension exploration value assessment, five-layer hybrid validation, and the archive-based self-evolution loop—enable automatic, verifiable, and continuously improving benchmark generation. Our experiments demonstrate 100% generation success on a synthetic corpus, stable quality (0.91–0.92) across evolution generations, effective novelty tracking, clear component contributions via ablation, coarse differentiation between a capable LLM agent and trivial baselines, and real-repository generated-task quality (0.76) approaching expert baselines (0.80). The system is implemented as a modular Python framework (124 tests, 9 components).
Future work includes: (1) large-scale experiments on real GitHub repositories with statistical rigor; (2) learned difficulty prediction models; (3) embedding-based semantic novelty detection; (4) cloud-based Docker evaluation deployment; and (5) human calibration of LLM-as-judge evaluation using a judge from a different model family.
Author Contributions
Conceptualization, methodology, software, validation, formal analysis, investigation, data curation, writing—original draft preparation, visualization, project administration: J.H.; writing—review and editing, supervision: Z.F., D.X., C.R. and Y.G. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The source code of SEBGen and all experiment scripts are available at the accompanying repository (to be released publicly upon acceptance; currently available from the corresponding author on request). All experiment data reported in this study are included in the article and its supplementary materials.
Acknowledgments
The authors would like to thank the maintainers of tqdm for providing a publicly accessible repository used in the real-repository experiment, and the SWE-bench team for the public release of the SWE-bench Lite dataset.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Jimenez, C.E.; Yang, J.; Wettig, A.; et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? In Proceedings of the 12th International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024.
- Wang, G.; Xie, Y.; Jiang, Y.; et al. Voyager: An Open-Ended Embodied Agent with Large Language Models. Trans. Mach. Learn. Res. 2024, arXiv:2305.16291. [CrossRef]
- Zhang, J.; Hu, S.; Lu, C.; Lange, R.; Clune, J. Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. In Proceedings of the International Conference on Learning Representations (ICLR), 2026; arXiv:2505.22954.
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language Agents with Verbal Reinforcement Learning. In Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 10–16 December 2024.
- Wang, L.; et al. Auto-SWE-Bench: Scalable, Real-World Benchmarks for LLM Coding Evaluation. Presented at the NeurIPS 2025 Exhibitor Spot Talks, San Diego, CA, USA, 2 December 2025.
- Wang, L.; Ramalho, L.; Celestino, A.; Pham, P.; Liu, Y.; Sinha, U.K.; Narváez Portillo, A.F.; Osunwa, O.; Maduekwe, G. SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories. arXiv 2025, arXiv:2512.17419.
- Yao, S.; Zhao, J.; Yu, D.; et al. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the 11th International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023.
- Gao, H.; et al. A Survey of Self-Evolving Agents. Trans. Mach. Learn. Res. 2026. [CrossRef]
Figure 1.
SEBGen system architecture: four layers and the closed evolution loop of six core components. Solid arrows denote data flow; dashed red arrows denote the evolution feedback loop.
Figure 1.
SEBGen system architecture: four layers and the closed evolution loop of six core components. Solid arrows denote data flow; dashed red arrows denote the evolution feedback loop.

Table 1.
Fundamental limitations of static benchmarks.
| Limitation | Manifestation | Consequence |
|---|---|---|
| Staticity | Fixed task sets | Model overfitting; benchmark saturation |
| Manual dependency | Expert-constructed tasks | High expansion cost; slow update cycles |
| Lack of adaptivity | Fixed difficulty levels | Cannot track evolving model capabilities |
Table 2.
Positioning of SEBGen relative to prior work.
| Prior Work | Key Feature | SEBGen Improvement |
|---|---|---|
| SWE-bench | Static benchmark | Dynamic, continuously evolving task set |
| Voyager | Implicit LLM exploration judgment | Explicit 5-dimension quantitative scoring |
| DGM | 3-layer validation | 5-layer validation (+regression, +semantic) |
| SWE-Bench++ | Static scale (11k tasks) | Evolution-driven freshness + strategy-level meta-evaluation |
| Auto-SWE-Bench | Automated extraction | Evolution-driven generation strategy |
Table 3.
Six design principles and their architectural embodiment.
| Principle | Source | Embodiment |
|---|---|---|
| P1: Modular Isolation | SWE-bench Docker | Standard interfaces between components |
| P2: Verification-Driven | DGM | Five-layer validation at every generation step |
| P3: Archive Accumulation | DGM Archive | Persistent multi-dimensional experience store |
| P4: Curriculum Progression | Voyager | ZPD-based difficulty adaptation |
| P5: Exploration-Exploitation | DGM Sigmoid | Parent selection with novelty reward |
| P6: Containerized Isolation | SWE-bench | Docker-based evaluation (cloud deployment) |
Table 4.
The five-layer quality validation pipeline.
| Layer | Name | Method | Purpose |
|---|---|---|---|
| L1 | Schema | Pydantic format checks | Structural validity |
| L2 | Compile | Python AST parsing | Syntactic correctness |
| L3 | Functional | pytest execution in sandbox | Behavioral verification (F2P tests) |
| L4 | Quality | LLM-as-judge scoring | Clarity, correctness, solvability, novelty |
| L5 | Semantic | LLM reasonability check | Task meaningfulness, difficulty calibration |
Table 5.
The five dimensions of the exploration value evaluator.
| Dimension | Formula | Rationale |
|---|---|---|
| Quality (Q) | Validation performance | |
| Novelty (N) | Anti-contamination | |
| Difficulty Coverage (D) | Balanced difficulty span | |
| Skill Diversity (S) | Coverage breadth | |
| Evolution Potential (E) | Pattern-based estimation | Optimization headroom |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.