Submitted:
11 June 2026
Posted:
12 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We identify and empirically demonstrate the “Semantic-Syntax Trade-off” in modern reasoning LLMs, showing how strict Grammar-Constrained Decoding disrupts Chain-of-Thought planning and induces semantic dead-ends in niche DSLs.
- We present a novel Dual-Phase Cascaded Framework that implements Optimistic Bypassing. By preserving unconstrained reasoning from a failed draft and using it to guide a strict GCD fallback, our architecture effectively utilizes the symbolic grammar as a localized syntax repair engine.
- We construct and openly release a comprehensive benchmark dataset comprising 100 natural language to MiniZinc constraint programming tasks of varying complexities.
- We conduct a rigorous empirical evaluation using a metric, scored by a strict semantic LLM judge and native compiler. Our results demonstrate that the dual-phase “think-then-constrain” approach significantly outperforms zero-shot, pure few-shot, and pure GCD baselines for unseen DSL code generation.
2. Previous Work
3. Methodology
3.1. In-Context Symbolic Grounding
3.2. Phase 1: Unconstrained Reasoning and Optimistic Bypassing
- 1.
- Syntactic Gate: A fast Context-Free Grammar parser (e.g., Lark) verifies that the generated string perfectly conforms to the EBNF structure.
- 2.
- Semantic Gate: The draft is passed to the native DSL compiler in type-checking mode. Because the compiler evaluates the code globally, it instantly flags any type mismatches, uninitialized variables, or invalid operator overloads.
3.3. Phase 2: Grammar-Constrained Syntax Repair
3.4. Algorithmic Formalization
| Algorithm 1:Dual-Phase Cascaded Neurosymbolic Generation |
![]() |
4. Empirical Evaluation
4.1. Benchmark Dataset Construction
4.2. Experimental Setup and Baselines
- Baseline 1: Zero-Shot Autoregressive
- Baseline 2: One-Shot (Unconstrained)
- Baseline 3: One-Shot (GCD Only)
- Proposed Method: Dual-Phase Cascaded GCD
4.3. Evaluation Metrics: and Automated Semantic Judging
- 1.
- Syntactic Gate: The generated string must be successfully parsed by the formal Lark CFG parser, verifying absolute structural compliance.
- 2.
- Semantic Compiler Gate: The code is passed to the native MiniZinc compiler using the –model-check-only flag. This deterministic check instantly rejects uninitialized variables, out-of-bounds array accesses, and semantic type mismatches.
- 3.
- Functional Intent Gate (LLM-as-a-Judge): Because code can compile perfectly but fail to solve the requested problem, we implement a strict semantic judge. We employ a larger, independent local model (Qwen3.5) via Ollama and prompted with a highly strict CoT evaluation rubric. The judge compares the generated code against the golden solution from the benchmark, analyzing variable mapping, constraint logic, and optimization directions. The judge returns a normalized score between and ; a candidate is only marked as successful if it achieves a score of .
4.4. Results and Discussion
| Method | pass@1 | pass@3 | pass@5 | |||
|---|---|---|---|---|---|---|
| Zero-Shot | 8.0 | +687.5% | 18.0 | +283.3% | 21.0 | +261.9% |
| One-Shot (No GCD) | 47.0 | +34.0% | 62.0 | +11.3% | 68.0 | +11.8% |
| One-Shot (GCD Only) | 8.0 | +687.5% | 8.0 | +762.5% | 9.0 | +744.4% |
| Dual-Phase (Proposed) | 63.0 | - | 69.0 | - | 76.0 | - |
4.4.1. The Semantic Collapse of Strict GCD
4.4.2. Unconstrained Reasoning vs. Zero-Shot Pre-training
4.4.3. The Efficacy of the Dual-Phase Architecture
4.4.4. Scaling with Sampling Diversity ()
5. Conclusions and Future Work
5.1. Limitations and Threats to Validity
5.2. Future Research Directions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H.P.d.O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. Evaluating Large Language Models Trained on Code. arXiv.org 2021.
- Rozière, B.; Gehring, J.; Gloeckle, F.; Sootla, S.; Gat, I.; Tan, X.E.; Adi, Y.; Liu, J.; Sauvestre, R.; Remez, T.; et al. Code Llama: Open Foundation Models for Code. arXiv.org 2023.
- Nethercote, N.; Stuckey, P.J.; Becket, R.; Brand, S.; Duck, G.J.; Tack, G. MiniZinc: Towards a Standard CP Modelling Language. In Proceedings of the Principles and Practice of Constraint Programming – CP 2007; Bessière, C., Ed., Berlin, Heidelberg, 2007; pp. 529–543. [CrossRef]
- Poesia, G.; Polozov, A.; Le, V.; Tiwari, A.; Soares, G.; Meek, C.; Gulwani, S. Synchromesh: Reliable Code Generation from Pre-trained Language Models. In Proceedings of the International Conference on Learning Representations, 2022.
- Willard, B.T.; Louf, R. Efficient Guided Generation for Large Language Models. arXiv.org 2023. [CrossRef]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; brian ichter.; Xia, F.; Chi, E.H.; Le, Q.V.; Zhou, D. Chain of Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A.H.; Agarwal, A.; Belgrave, D.; Cho, K., Eds., 2022.
- Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.L.; Cao, Y.; Narasimhan, K.R. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Proceedings of the Thirty-seventh Conference on Neural Information Processing Systems, 2023.
- Xue, T.; Li, X.; Azim, T.; Smirnov, R.; Yu, J.; Sadrieh, A.; Pahlavan, B. Multi-Programming Language Ensemble for Code Generation in Large Language Model. ArXiv 2024, abs/2409.04114.
- Sarker, L.; Downing, M.; Desai, A.; Bultan, T. Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation, 2026. arXiv:2404.01535 [cs.SE].
- Liang, Q.; Zhang, Z.; Sun, Z.; Lin, Z.; Luo, Q.; Xiao, Y.; Chen, Y.; Zhang, Y.; Zhang, H.; Zhang, L.; et al. Grammar-Based Code Representation: Is It a Worthy Pursuit for LLMs? In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025; Che, W.; Nabende, J.; Shutova, E.; Pilehvar, M.T., Eds., Vienna, Austria, 2025; pp. 15640–15653. [CrossRef]
- Thakur, S.; Ahmad, B.; Fan, Z.; Pearce, H.; Tan, B.; Karri, R.; Dolan-Gavitt, B.; Garg, S. Benchmarking Large Language Models for Automated Verilog RTL Code Generation. In Proceedings of the 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2023, pp. 1–6. ISSN: 1558-1101. [CrossRef]
- Thakur, S.; Ahmad, B.; Pearce, H.; Tan, B.; Dolan-Gavitt, B.; Karri, R.; Garg, S. VeriGen: A Large Language Model for Verilog Code Generation. ACM Transactions on Design Automation of Electronic Systems 2024, 29, 46:1–46:31. [CrossRef]
- Bassamzadeh, N.; Methani, C. A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation, 2024. arXiv:2407.02742 [cs.SE]. [CrossRef]
- Fu, D.J.; Gupta, A.; Councilman, A.; Grove, D.; Wang, Y.X.; Adve, V. SLMFix: Leveraging Small Language Models for Error Fixing with Reinforcement Learning, 2025. arXiv:2511.19422 [cs.SE]. [CrossRef]
- Delgado, D.; Burgueño, L.; Clarisó, R. A framework for assessing the capabilities of code generation of constraint domain-specific languages with large language models. Journal of Systems and Software 2026, 238, 112871. [CrossRef]
- Shen, D.; Chen, X.; Wang, C.; Sen, K.; Song, D. Benchmarking Language Models for Code Syntax Understanding. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2022.
- Mora, F.; Wong, J.; Lepe, H.; Bhatia, S.; Elmaaroufi, K.; Varghese, G.; Gonzalez, J.; Polgreen, E.; Seshia, S.A. Synthetic Programming Elicitation for Text-to-Code in Very Low-Resource Programming and Formal Languages. Advances in Neural Information Processing Systems 37 2024.
- Gao, M.; Zhao, J.; Lin, Z.; Ding, W.; Hou, X.; Feng, Y.; Li, C.; Guo, M. AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMs. 2024 IEEE 42nd International Conference on Computer Design (ICCD) 2024, pp. 162–169.
- Lu, Y.; Liu, S.; Zhang, Q.; Xie, Z. RTLLM: An Open-Source Benchmark for Design RTL Generation with Large Language Model. 2024 29th Asia and South Pacific Design Automation Conference (ASP-DAC) 2023, pp. 722–727.
- Geng, S.; Josifoski, M.; Peyrard, M.; West, R. Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; Bouamor, H.; Pino, J.; Bali, K., Eds., Singapore, 2023; pp. 10932–10952. [CrossRef]
- Park, K.; Wang, J.; Berg-Kirkpatrick, T.; Polikarpova, N.; D’ Antoni, L. Grammar-Aligned Decoding. In Proceedings of the Advances in Neural Information Processing Systems; Globerson, A.; Mackey, L.; Belgrave, D.; Fan, A.; Paquet, U.; Tomczak, J.; Zhang, C., Eds. Curran Associates, Inc., 2024, Vol. 37, pp. 24547–24568. [CrossRef]
- Wen, H.; Zhu, Y.; Liu, C.; Ren, X.; Du, W.; Yan, M. Fixing Function-Level Code Generation Errors for Foundation Large Language Models 2025. arXiv:2409.00676 [cs.SE]. [CrossRef]
- Liu, C.; Bao, X.; Zhang, H.; Zhang, N.; Hu, H.; Zhang, X.; Yan, M. Guiding ChatGPT for Better Code Generation: An Empirical Study. In Proceedings of the 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), 2024, pp. 102–113. ISSN: 2640-7574. [CrossRef]
- Albinhassan, M.; Madhyastha, P.; Russo, A. $\texttt{SEM-CTRL}$: Semantically Controlled Decoding. Transactions on Machine Learning Research 2025.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
