Submitted:
29 July 2026
Posted:
29 July 2026
Read the latest preprint version here
Abstract
We tackle the problem of job shop scheduling. Exact methods are limited to small instances, and traditional metaheuristics do not guarantee feasibility at each iteration. In this research, we propose a new two phases hybrid metaheuristic. The first phase uses Gradient Descent on a convex energy function to quickly construct a feasible solution by fixing the operations sequence. We provide a mathematical proof of convergence for this phase to a feasible solution. The second phase applies an Iterated Local Search to explore solutions space and minimize the makespan.This separation of objectives guarantees convergence of initial phase and simplifies parameter tuning. Experiments on standard FT and LA benchmarks show that EGD-ILS achieves competitive results with reduced computation time.
Keywords:
job shop scheduling
; energy-based model
; gradient descent
; hybrid optimization
; convergence
1. Introduction
The JSSP is a central problem in operations research and production management. It consists of assigning a set of n jobs on a set of m machines. Each job consists of m operations that must be processed in a given order. An operation has a processing time and it has a predefined machine . Our goal is to achieve the minimum makespan. Precedence and resource constraints must be satisfied.
This problem is NP-hard. Exact approaches such as Integer Linear Programming or Constraint Programming fail when the size exceeds 10x10. Metaheuristics such as Genetic Algorithms, Simulated Annealing, or Particle Swarm Optimization provide good results but require fine parameter tuning and spend much time in infeasible regions of the search space [1].
Since the late 1980s, Hopfield networks were applied to JSSP by encoding each decision as a binary variable indicating if job j is processed on operation o at time t [2]. This leads to neurons with T is the number of time steps. With Deep Learning, GNN (Graph Neural Networks) and RL (Reinforcement Learning) methods achieved strong results but require GPU and large datasets [3]. Recently, hybrid approaches have emerged to combine the advantages of different methods. The approach involves breaking down the problem into two parts: 1. Find a feasible solution, 2. Optimize the objective function [4].
We propose EGD-ILS, a hybrid metaheuristics that uses the energy-based philosophy but optimize directly the continuous vector of start times . A differentiable energy measures the magnitude of violations. Minimizing via gradient descent repairs an infeasible solution until an admissible schedule is obtained. Iterated Local Search explores admissible solutions to find optimal or near optimal solution. Our main contributions are threefold:
- 1.
- We prove mathematically that by fixing the operation sequence, the feasibility search becomes a convex optimization problem that converges to E=0.
- 2.
- We propose a simple and efficient architecture that combines the speed of Gradient Descent for feasibility and the power of Iterated Local Search for optimization.
- 3.
- Experimental validation on FT and LA benchmarks.
- 4.
- Our work is the first work that provides a theoretical convergence guarantee for the feasibility phase of a hybrid JSSP solver based on gradient methods.
2. Review and Related Work
The use of neural-inspired energy functions for scheduling dates back to the late 1980. In a pioneering study, Foo and Takefuji [2] formulated JSSP as a minimization problem on a Hopfield network. Their objective combined penalty terms for constraint violations with a term related to the makespan. A key drawback reported was the difficulty in balancing the different penalty weights, which often led to poor convergence behaviour [4]. Van den Bout [5] introduced Mean Field Annealing but this approach suffers from the problem of combinatorial explosion and local minima..
Recent research has focused on energy formulations using modern optimization strategies. for example is the method developed by Peng and others, using an augmented Lagrangian framework which incorporates any constraint violations within the Lagrangian. This approach employed a gradient technique to minimize Lagrangian and then applied a heuristic to achieve feasibility. This approach works well for medium instances only, but it still depends on choice of the correct penalty [6].
To enhance the robustness level, Wang and Zheng proposed a metaheuristic in two phases. In the first one, the simulated annealing is utilized to minimize the compound function of the makespan and the violations of constraints. However, once again, some admissible solution is found, the second phase is based on Tabu search and aimed to minimize the makespan [3,7,8].
In contrast to the methods above, the proposed EGD-ILS does not rely on stochastic search or data-driven training in the first phase. Instead, we show that by committing to a specific operation order upfront, the feasibility problem becomes convex. This allows us to use deterministic Gradient Descent with a proven convergence property. To our knowledge, this explicit convexity argument and its use for guaranteed feasibility in JSSP have not been explored before.
3. Methodology: The EGD-ILS Approach
3.1. Problem Formulation
We developed a mathematical statement of the JSSP problem [9]. Without loss of generality, Let’s use the instance Table 1 as example highlight the mathematical model used.
is the starting time of operation . is the resource used. is the processing time . H : Sum of all processing times .
We introduce the Boolean variable:
The goal is: minimize subject to :
▪ Starting Time constraints : equations
▪ Precedence Constraints : equations
▪ Machine Constraints : equations
3.2. Phase 1: Feasible Construction by Gradient Descent with Fixed Sequence
3.2.1. Starting Energy
3.2.2. Precedence Energy
3.2.3. Machine Conflict Energy
The key idea of EGD-ILS is to fix the operation sequence in phase 1. This transforms the disjunctive machine constraint into a linear constraint. Without any loss of generality, consider that is a fixed operation sequence. Machine Constraints becomes:
must be non negative.
We define the Machine Conflict Energy as:
3.2.4. Resolution Algorithm
We define the energy function to minimize:
| Algorithm 1 Energy Descent for JSSP |
|
3.2.5. Mathematical Proof of Convergence
Lemma 1.
The energy function , which was defined above for a fixed order π, is convex and L-smooth.
Proof.
has the form . is convex and 2-smooth. . . Z is an affine function of S. The composition is convex and L-smooth.
The function is a sum of convex and smooth functions. Therefore is convex and L-smooth [10]. □
Lemma 2.
For any fixed order π, there exists such that .
Proof.
By definition of the JSSP, for any order there exists at least one feasible schedule. For this schedule, all violations are . As , is the global minimum. □
Theorem 1.
For any step , converges to such that .
Proof.
From Lemma 1, we can see that E is convex and L-smooth. From Lemma 2, we know that the set of minima is non-empty. Moreover is L-Lipschitz. We apply the classical convergence theorem of gradient descent for convex functions. ([11], Sec. 9.1) □
3.3. Phase 2: Optimization by Iterated Local Search
The ILS explores the set of admissible schedules by using neighbours structure and deciding to transit or not to the candidate solution.
| Algorithm 2 Optimization by ILS |
|
Denote permutes two operations. is right shift insertion. uses left shift insertion. is inversion blocks between two positions [12]. uses random combination of the four neighbours .
An initial temperature of is given. The parameter the speed of convergence.
4. Experimental Results
The RPD is: where is the makespan found by the EGD-ELS; and is the optimal makespan.
EGD-ILS delivers high-quality solutions on FT and LA benchmarks with reduced computation time. It achieves RPD < 3.9%, solving FT06 and LA05 to LA15 optimally.
5. Conclusions
We proposed a straightforward energy-based approach for JSSP that’s simple and fast. EGD-ILS shows strong performance on FT and LA benchmarks with RPD < 3.9% , solving FT06 and LA05-LA15 optimally with reduced computation time.
References
- Pinedo, M.L. Scheduling: Theory, Algorithms, and Systems, 5 ed.; Springer: New York, 2016. [Google Scholar]
- Foo, Y.P.S.; Takefuji, Y. Stochastic neural networks for solving job-shop scheduling. I. Problem representation. In Proceedings of the Proceedings of the IEEE International Conference on Neural Networks; IEEE, 1988; Vol. 2, pp. 275–282. [Google Scholar]
- Zhang, C.; et al. Learning to Share in Multi-Agent Reinforcement Learning. In Proceedings of the Proceedings of the 39th International Conference on Machine Learning (ICML), 2022. [Google Scholar]
- Garey, M.R.; Johnson, D.S. Computers and Intractability: A Guide to the Theory of NP-Completeness; W. H. Freeman: San Francisco, 1979. [Google Scholar]
- Van den Bout, D.; Miller, T. A traveling salesman approach to job shop scheduling. In Proceedings of the International Joint Conference on Neural Networks (IJCNN), 1990. [Google Scholar]
- Peng, B.; Lü, Z.; Chiang, T.C. A new Lagrangian dual method for the job shop scheduling problem. Eur. J. Oper. Res. 2016, 255, 356–366. [Google Scholar]
- Wang, C.; Zheng, H.F. A Two-Stage Approach for Job Shop Scheduling. In Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics (SMC); IEEE, 1998. [Google Scholar]
- Alet, F.; et al. Neural Constraint Satisfaction for Combinatorial Optimization. Tech report, DeepMind, 2024. [Google Scholar]
- Nohair, L.; El Adraoui, A.; Namir, A. Solving non-delay job-shop scheduling problems by a new matrix heuristic. Procedia Comput. Sci. 2022, 198, 410–416. [Google Scholar] [CrossRef]
- Nesterov, Y. Introductory Lectures on Convex Optimization: A Basic Course. In Applied Optimization; Kluwer Academic Publishers: Boston, 2004; Vol. 87. [Google Scholar]
- Boyd, S.; Vandenberghe, L. Convex Optimization; Cambridge University Press: New York, 2004. [Google Scholar]
- Belabid, J.; Aqil, S.; Allali, K. Solving Permutation Flow Shop Scheduling Problem with Sequence-Independent Setup Time. J. Appl. Math. 2020, 2020, 7132469. [Google Scholar] [CrossRef]
Table 1.
2 jobs, 3 machines H=53.
| Jobs | Operation 1 | Operation 2 | Operation 3 |
|---|---|---|---|
| M2(8) | M1(12) | M3(3) | |
| M3(11) | M2(5) | M2(14) |
Table 2.
Simulation results on FT benchmarks.
| Instance | Size | RPD% | Max iterations | ||||
|---|---|---|---|---|---|---|---|
| FT06 | 6*6 | 55 | 55 | 0 | 2000 | 2000 | 0,7 |
| FT10 | 10*10 | 930 | 957 | 2,90 | 10000 | 1800 | 0,9 |
| FT20 | 20*5 | 1165 | 1197 | 2,75 | 10000 | 1800 | 0,9 |
Table 3.
Simulation results on LA instances.
| Instance | Size | RPD% | Max iterations | ||||
|---|---|---|---|---|---|---|---|
| LA 01 | 10X5 | 666 | 666 | 0 | 3000 | 1400 | 0,95 |
| LA 02 | 10X5 | 655 | 667 | 1,83 | 5000 | 1600 | 0,95 |
| LA 03 | 10X5 | 597 | 604 | 1,17 | 5000 | 1400 | 0,7 |
| LA 04 | 10X5 | 590 | 598 | 1,36 | 10000 | 1600 | 0,95 |
| LA 05 | 10X5 | 593 | 593 | 0 | 3000 | 1600 | 0,95 |
| LA 06 | 15X5 | 926 | 926 | 0 | 3000 | 1600 | 0,95 |
| LA 07 | 15X5 | 890 | 890 | 0 | 3000 | 1600 | 0,95 |
| LA 08 | 15X5 | 863 | 863 | 0 | 3000 | 1600 | 0,95 |
| LA 09 | 15X5 | 951 | 951 | 0 | 3000 | 1600 | 0,95 |
| LA 10 | 15X5 | 958 | 958 | 0 | 3000 | 1600 | 0,95 |
| LA 11 | 20X5 | 1222 | 1222 | 0 | 3000 | 1600 | 0,95 |
| LA 12 | 20X5 | 1039 | 1039 | 0 | 3000 | 1600 | 0,95 |
| LA 13 | 20X5 | 1150 | 1150 | 0 | 3000 | 1600 | 0,95 |
| LA 14 | 20X5 | 1292 | 1292 | 0 | 3000 | 1600 | 0,95 |
| LA 15 | 20X5 | 1207 | 1207 | 0 | 5000 | 1600 | 0,95 |
| LA 16 | 10X10 | 945 | 982 | 3,91 | 5000 | 1600 | 0,95 |
| LA 17 | 10X10 | 784 | 793 | 1,15 | 10000 | 2000 | 0,7 |
| LA 18 | 10X10 | 848 | 861 | 1,53 | 5000 | 1600 | 0,95 |
| LA 19 | 10X10 | 842 | 875 | 3,92 | 10000 | 2000 | 0,7 |
| LA 20 | 10X10 | 902 | 914 | 1,33 | 10000 | 2000 | 0,7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.