Preprint
Article

This version is not peer-reviewed.

Quasi-Random Sampling-Enhanced Metaheuristic Algorithms for CAMD-Based Solvent Selection in Octacosanol Extraction

Submitted:

07 July 2026

Posted:

08 July 2026

You are already at the latest version

Abstract
This study develops a computer-aided molecular design (CAMD) framework to select solvents for extracting octacosanol from multicomponent sugarcane wax. The solvent-selection problem is formulated as a mixed-integer nonlinear programming problem in which candidate solvents are generated from UNIFAC functional groups and evaluated using the net distribution coefficient and net solvent selectivity. Four metaheuristic solvers—ant-colony optimization (ACO), simulated annealing (SA), efficient ant-colony optimization (EACO), and efficient simulated annealing (ESA)—are compared to assess solvent quality and computational efficiency. EACO and ESA incorporate Hammersley sequence sampling to improve multidimensional sampling uniformity relative to conventional pseudo-random sampling. The results show that all four solvers consistently identify the same highest-ranked candidate solvent. ACO and EACO require substantially fewer objective-function evaluations than SA and ESA, while SA-based methods provide greater diversity among lower-ranked candidates. EACO reduces the computational cost relative to ACO while retaining the same top-ranked solvents, whereas ESA provides only a modest efficiency improvement over SA. Overall, the proposed CAMD framework reliably identifies promising solvents for nutraceutical extraction, and quasi-random sampling improves the efficiency of metaheuristic solvent design.
Keywords: 
;  ;  ;  

1. Introduction

Natural products often occur as complex mixtures, making the isolation of target compounds an important objective because of their specific applications. In industry, extraction is one of the most widely used separation operations, and solvent selection is critical to its effectiveness. In an extraction system, the solvent should selectively attract the desired compound into the extract phase while leaving impurities in the raffinate phase. For a given system, solvents are generally selected using one of three approaches [1,2]:
(a)
Experimental screening of candidate solvents for the system,
(b)
Database screening based on physical, chemical, and related properties, and
(c)
Computer-aided molecular design (CAMD) to generate optimal solvents.
The first two approaches may not identify the best solvent and can lead to incomplete evaluation. Although experimental screening provides reliable results, it is often expensive and time-consuming. In contrast, CAMD uses thermodynamic models to design solvents tailored to target performance requirements.[1] In CAMD, molecules are generated by combining available functional groups to ensure that the resulting structures meet desired performance metrics. Because the number of possible group combinations and attachment patterns is extremely large, combinatorial optimization is central to this problem.
Papadimitriou and Steiglitz [3] define a combinatorial optimization problem as one that seeks an optimal solution, such as an integer, permutation, or graph structure, from a finite or countably infinite set. Common examples include the traveling salesman and scheduling problems.[4] CAMD can also be formulated as a combinatorial optimization problem. Metaheuristic algorithms, including ant colony optimization (ACO), evolutionary computation, genetic algorithms, and simulated annealing (SA), use random number generation during the search process, and their performance strongly depends on the quality of the sampling. In this work, we enhance SA and ACO by replacing the pseudo-random numbers used in conventional implementations with quasi-Monte Carlo samples generated through Hammersley sequence sampling. The resulting algorithms, efficient ant-colony optimization (EACO) and efficient simulated annealing (ESA), preserve k-dimensional sampling uniformity.
The optimization problem considered in this study focuses on extracting nutraceuticals from multicomponent sugarcane wax. Nutraceuticals are health-promoting compounds derived from food sources such as animal and plant waxes, as well as apple and grape seeds, and they provide medicinal benefits beyond basic nutrition.[5] As sugarcane is produced on a large scale worldwide, extracting nutraceuticals from sugarcane waste has become an important separation operation.[6] Sugarcane wax contains various long-chain fatty compounds, including alcohols, esters, aldehydes, acids, and saturated hydrocarbons. Among these, saturated alcohols known as policosanols are important nutraceuticals. Policosanols typically contain 18–34 carbon atoms and a hydroxyl group (-OH) on the alpha carbon. Octacosanol, a 28-carbon policosanol, is particularly promising because of its reported health benefits, including immune-supporting, anti-Parkinsonian, and anti-fatigue effects.[7,8] Therefore, this study uses CAMD to identify optimal solvents for extracting octacosanol from the multicomponent mixture present in sugarcane wax.
To provide context, the next section discusses the importance of octacosanol and reviews reported extraction studies for obtaining it from sugarcane wax. This is followed by a brief overview of CAMD, molecular construction, and the group contribution method. The formulation and development of the solvent selection optimization problem for extracting octacosanol from multicomponent sugarcane wax are then presented. Section 3 describes the algorithms used in this study, beginning with ACO and SA, followed by previously reported enhancements to these methods. It then introduces Hammersley sequence sampling (HSS), explains its advantages, and describes its incorporation into ACO and SA to develop EACO and ESA. The algorithmic frameworks and parameter values used in this work are also provided. Section 4 presents the solvents generated by the four solvers and compares the efficiency of EACO and ESA with that of ACO and SA, respectively. Finally, the study concludes by identifying the best solvents for extracting octacosanol from sugarcane wax and comparing the four solvers for this process-systems optimization problem.

2. The Solvent Selection Problem

Policosanols offer several health benefits, including lipid regulation, anti-atherosclerotic effects, and support in managing obesity, etc.[8,9] Among them, octacosanol is a particularly promising nutraceutical. Numerous studies have attempted to extract policosanols or octacosanol-rich fractions from sugarcane using methods such as Soxhlet extraction, supercritical CO2 (99.99%), and subcritical dimethyl ether extraction.[10,11,12,13]
However, in most cases, solvent selection has been ad-hoc, guided largely by prior experience. This approach risks overlooking solvents with superior performance, especially, since experimentally screening a large solvent set is costly and time-consuming. A systematic, performance-based solvent selection strategy is therefore essential.
Solvent choice in extraction depends on properties such as the distribution coefficient, selectivity, recoverability, and environmental safety.[1,5] In this study, we focus on distribution coefficient and solvent selectivity.
  • Distribution coefficient is the ratio of the desired solute’s concentration in the solvent phase to its concentration in the original mixture.
  • Solvent selectivity measures how effectively a solvent attracts the desired solute relative to undesired components.
These metrics form the basis of the solvent selection optimization problem developed in the following sections.

2.1. Computer-Aided Molecular Design

The group contribution method estimates molecular properties—melting and boiling points, density, viscosity, and activity coefficients (γ)—from the functional groups that make up a molecule. As distribution coefficient and selectivity can be determined using activity coefficients, this method is central to our analysis.
CAMD reverses this logic: instead of predicting properties from a molecule, it constructs molecules from available groups to optimize performance metrics.[1] CAMD has been widely applied in designing solvents for separation processes, refrigerants, polymers, pharmaceuticals, and more.[1,2,14,15,16,17] Recent advances even integrate machine learning to generate molecules with drug-like characteristics.[18]
In CAMD, molecules are built by selecting:
  • N1: the total number of functional groups
  • N2(i): the index of each group (i = 1,…,N1)
Feasible molecular structures must satisfy chemical bonding rules — the octet rule for covalent systems.[5] Thus, N1 and N2(i) serve as decision variables, while bonding constraints ensure chemically valid structures.
Figure 1 outlines the CAMD workflow for identifying optimal solvents for sugarcane wax extraction. The process begins by identifying mixture components and estimating their activity coefficients, followed by formulating and solving the solvent-performance optimization problem.

2.2. Group Contribution Method: UNIFAC

The group contribution method (GCM) predicts thermodynamic properties by decomposing molecules into functional groups and summing their contributions.[19] This approach is especially valuable when experimental data are unavailable, enabling property estimation across chemical engineering, molecular design, and process simulation.
Among the many properties GCM can predict, activity coefficients are particularly important for mixture behavior. They quantify deviations from ideality, with γ = 1 representing ideal solutions. For solvent design, the most widely used GCM for estimating activity coefficients is UNIFAC, which provides reliable predictions for non-ideal liquid mixturesFor molecule i present in a solution, its activity coefficient, γ i , is defined as shown in Equation 1.[19]
U γ i = l n   γ i l n   γ i C + l n   γ i R = 0
Equation 1 is the standard UNIFAC equation to estimate the activity coefficient. The estimation of γ i is done in two parts namely, the combinatorial part ( γ i C ) and the residual part ( γ i R ). Such a distinction allows us to understand the effects of size and shape, and the group energy interactions. The combinatorial part takes care of the effects caused by size and shape while on the other hand, the residual part takes care of the energy interactions among the groups of molecules present in the solution.
Firstly, the combinatorial part is estimated as shown in Equation 2(a-f). [19]
l n γ i C = ln Φ i x i b + 5 q i ln θ i Φ i + l i Φ i x i b j x j b l j
where
Φ i = r i x i b j N C r j x j b (2b), θ i = q i x i b j N C q j x j b (2c), l i = 5 r i q i r i 1 (2d)
where Φ i and θ i are the molecular volume fraction and the molecular surface area fraction of molecule i, ranging from 1 to NC molecules. This is followed by the estimation of van der Waals’ volume, r i , and surface area, q i , using Equation 2(e) and Equation 2(f), respectively.
r i = g N g ν g ( i ) R g (2e), q i = g N g ν g ( i ) Q g (2f)
where R g and Q g are van der Waals’ group volume and group surface area values ranging from i to N g groups, and ν g ( i ) is the total number of functional groups present in the molecule. The values of R g and Q g are obtained from the literature for this study.
While the estimation of the residual part is based on the Equation 3(a-e). [19]
l n γ i R = k v k i ln Γ k ln Γ k i
where Γ k is defined as the residual activity coefficient of group k in a solution, shown in Equation 3(b). Similarly, Γ k i is defined as the residual activity coefficient of group k in a solution containing only molecules of type i.
l n Γ k = Q k 1 l n m Θ m Ψ m k m Θ m Ψ k m n Θ n Ψ n m
where m and n represent all the functional groups in the molecule, Θ m is the surface fraction of group m , Equation 3c, which is further expressed in terms of mole fraction of groups present in the bulk phase, Equation 3d, and lastly, the temperature-dependent parameter as per UNIFAC is Ψ m n , Equation 3e.[19]
Θ m = Q m X m m Q m X m
X m = i j ν m ( i ) x j 1 j 1 n ν m ( i ) x j
Ψ m n = e a m n T
In Equation 3e, a m n is called the group interaction parameter (represents interaction of group m with group n). It is defined as the difference in the energy level between group m and group n or between two of the main group m. In the latter case, the interaction parameter is zero because of the same energy levels of both the groups. These values are obtained from the literature for every group. [19]
In the next subsection, we discuss the development of the optimization problem for solvent selection used in this study along with the parameters important to build the optimization problem.

2.3. Solvent Selection Optimization Problem

This section formulates the solvent-selection optimization problem for extracting a desired compound from a multicomponent mixture. Solvent performance is evaluated using the distribution coefficient ( m ) and solvent selectivity ( β ), both estimated using the UNIFAC group contribution method. For a binary mixture containing desired solute B, undesired solute A, and solvent S, these quantities are defined by Equations 4(a) and 4(b), respectively.[5]
m B = C B ,   S C B ,   A = γ B A γ B S
β B = m B m A = C B , S C A , S = γ S A γ B S
where Cx,y represents the concentration of x present in y. However, in the present study, the system is a multicomponent. Therefore, a recent study has reported the definitions of m and β in a better way to estimate these based on the mole fractions ( x A i ) of every undesired solute (Ai, i=1,2,.., N) present in the multicomponent system. Equation 5(a) and Equation 5(b) shows the net distribution coefficient ( m n e t ) and the net solvent selectivity ( β n e t ), respectively, as presented in the literature.[5]
m n e t = γ B X γ B S
β n e t = i = 1 N x A i β i
where γ B X = i = 1 N x A i γ B A i i = 1 N x A i , and β i is the solvent selectivity value obtained between desired solute B and the undesired solute Ai. The system of undesired solutes is represented as X. That is the reason the activity coefficient ( γ B X ) is defined between the desired solute B and every undesired solute Ai by weighing against their corresponding mole fraction.
Next, in this study, the multicomponent system is built based on earlier experimental studies. Sugarcane wax is a complex matrix of compounds. Therefore, we considered a few of the most abundant compounds reported in the literature to be present in this multicomponent mixture, while the desired solute remains to be octacosanol.[13] Figure 2 shows the system of compounds considered in this study along with their corresponding mole fractions as reported in the literature.
Solvent generation requires a defined set of UNIFAC groups. In this study, candidate molecules are constructed by selecting up to 12 groups from the 20 groups listed in Table 1, which were obtained from Fredenslund et al.[19] Each group is indexed from 1 to 20. A molecule may contain multiple instances of the same group, but the final structure must satisfy the octet rule. This feasibility constraint is expressed mathematically in Equation 6.[2]
i N 1 b i = 2 ( N 1 1 )
where b i determines the number of groups that can be attached to the group index i.
Finally, our objective is to maximize the net distribution coefficient of the solvent generated by any of the 20 groups listed in Table 1, so that it can isolate octacosanol with a net solvent selectivity ( β n e t , m i n ) greater than or equal to the set value from the ten undesired solutes present in the mixture. Based on these, the optimization problem is defined as shown in Equation 7.
m a x N 1 , N 2 ( i )   m n e t
s . t .                   β n e t β n e t , m i n
U γ i = 0
i N 1 b i = 2 ( N 1 1 )
2 N 1 12
1 N 2 ( i )   20     i   ϵ   [ 1 , 2 , , N 1 ]
Equation 7 shows that the decision variables, N 1 and N 2 ( i ) , are discrete, and the objective function and constraints, which are evaluated using UNIFAC, are nonlinear; therefore, the problem is a mixed-integer nonlinear programming problem. The combinatorial complexity arises because the groups in Table 1 can be connected in many feasible ways. To address this challenge, metaheuristic algorithms such as ACO and SA, and their efficient variants ESA and EACO, are used to identify optimal solvent candidates. The next section presents the four solvers used in this work to solve this optimization problem.

3. Algorithmic Framework

As mentioned, ACO and SA are two metaheuristic approaches for solving combinatorial optimization problems. Firstly, we briefly discussed the two conventional algorithms in this section, followed by a discussion of Hammersley sequence sampling. Then, the implementation of the HSS technique in ACO and SA to develop EACO and ESA, respectively, is discussed. Furthermore, the efficiency of these algorithms compared with their counterparts reported in the literature is discussed. In addition, the frameworks of both EACO and ESA are also presented along with the algorithm parameters used in this study.

3.1. Ant-Colony Optimization

Ant-colony optimization (ACO) is a nature-inspired metaheuristic based on the foraging behaviour of ants, introduced by Dorigo for solving combinatorial optimization problems.[23] In nature, ants deposit pheromones as they return from food sources, allowing other ants to follow promising paths. Because shorter paths are traversed more frequently, pheromone accumulates faster on them, gradually guiding the colony toward the shortest route.[23,24] Similarly, in ACO, artificial ants construct solutions using pheromone information and, when available, heuristic information. At iteration iter, the probability of selecting solution component j from state I, ( P i j i t e r ) is given by Equation 8.[23,24]
P i j i t e r = τ i j α ( i t e r ) η i j β i t e r τ i j α ( i t e r ) η i j β
where and are the pheromone value and edge desirability for edge ij, respectively, at iteration iter, while and determine the influence of pheromone and heuristic information.[24] After all ants construct solutions, pheromone trails are updated according to solution quality, with additional pheromone assigned to edges in better solutions, as shown in Equation 9.[24]
τ i j i t e r + 1 = ρ τ i j i t e r + τ i j i t e r
where ρ is the pheromone evaporation factor, and τ i j is the pheromone addition for edge ij.
The artificial pheromone along the path acts as a communication bridge among artificial ants. In the less favorable solution routes, the artificial pheromone decays gradually, whereas in the better solution routes, the pheromone concentration increases, as observed in Equation 9. Eventually, iterative updating of pheromone values based on information gained from earlier iterations encourages the artificial ants towards the optimal solution in the solution space.
The solution construction in this study is based on incremental manner i.e., an ant, first, probabilistically chooses one of the solutions from the archive based on the Equation 10.[24]
P i = ω i j = 1 K ω j
where ω i is the weight associated with solution i, i is the rank of the solution in the archive, K is the size of solution archive. The weight is defined as shown in Equation 11.[24]
ω i = e x p ( i 1 2 2 q K 2 ) q K 2 π
where q is the algorithmic parameter of ACO, and mean of the Gaussian function is 1. However, in this study, the optimization problem is a mixed integer non-linear programming problem. Therefore, using a set of probability density functions, a multimodal one-dimensional density function is obtained. To estimate the multimodal density function, the following equation, Equation 12, is used.[25]
G j x = i = 1 K ω i exp x μ i 2 2 σ i 2 σ i 2 π ,   i , j 1 , , K
where G j x is the Gaussian kernel, ω i is the weight associated with individual Gaussian function, μ i and σ i are the vector means of the individual solution components and standard deviations, respectively.

3.2. Simulated Annealing

Simulated annealing (SA) is a probabilistic metaheuristic for finding near-optimal solutions to optimization problems. It is inspired by metallurgical annealing, in which a material is heated and then slowly cooled so that its particles settle into a low-energy state. Similarly, SA gradually reduces a control parameter, analogous to temperature, to guide the search toward lower objective-function values while allowing occasional uphill moves. If cooling is sufficiently slow, the system can approach thermal equilibrium at each temperature level, described by the Boltzmann distribution. At temperature T, the probability of occupying energy level Ei is given by Equation 13.[26]
P T = e x p E i k B T Z
where 1/Z is the normalization factor, and kB is the Boltzmann’s constant (1.3806 x 10-23 J/K).
In SA, the objective function represents the system energy. Starting from an initial configuration, the algorithm generates a candidate configuration by perturbing the current one. If the candidate has a lower objective value, the move is accepted. If it has a higher objective value, the move may still be accepted with probability according to the Metropolis criterion, given in Equation 14.[26,27]
Accept   the   move   if   P r a c c e p t A i j = e x p E T   i f   E = E j E i 0 1                                       o t h e r w i s e   i , j move
where A i j is the probability of accepting a transition from state i to state j. At high temperatures, SA accepts more uphill moves, improving exploration of the solution space. As temperature decreases, the probability of accepting such moves declines, allowing the algorithm to converge while reducing the risk of becoming trapped in local optima.
ACO and SA differ substantially in their search mechanisms. SA follows a single-solution trajectory, exploring the search space through successive perturbations of one candidate solution; therefore, complex problems may require many iterations, and escape from local optima depends on the Metropolis criterion.[26,28] In contrast, ACO uses multiple ants to construct solutions in parallel and share information through pheromone trails, which improves exploration and reduces the likelihood of premature trapping.[23] However, as a population-based heuristic, ACO may still miss some high-quality regions of the search space.

3.3. Hammersley Sequence Sampling

Sampling strongly affects optimization efficiency. Monte Carlo sampling uses independently generated pseudo-random numbers over the sampling domain; these samples are statistically unbiased and useful for estimating probabilities and objective-function values in complex systems.[27] However, pseudo-random samples can cluster in some regions while leaving others unexplored, as shown in Figure 3. This reduces sampling efficiency and often requires many samples to adequately cover the search space.[27,29] Quasi-random methods, such as Hammersley sequence sampling (HSS), reduce this clustering by distributing samples more uniformly. HSS, developed by Kalagnanam and Diwekar based on Hammersley points [30], provides better multidimensional uniformity than Monte Carlo methods and other methods, such as Latin hypercube sampling, as illustrated in Figure 3.[27]
The next section explains how HSS is incorporated into ACO and SA to develop variants that preserve k-dimensional sampling uniformity.

3.4. Efficient Ant-Colony Optimization (EACO)

ACO can suffer from premature convergence when excessive pheromone accumulation reduces population diversity and limits exploration. Several variants have been developed to improve its performance, including the Elitist Ant System, Rank-Based Ant System, Max-Min Ant System, Ant Colony System, hypercube-based ACO, and hybrids with local search methods such as tabu search.[4] Efficient Ant-Colony Optimization (EACO) improves ACO by replacing selected pseudo-random numbers with quasi-random numbers generated using Hammersley sequence sampling (HSS).[24,29] Reported studies show that EACO improves computational efficiency by 3%–71% for selected problems.[24] In this study, HSS is used only during initialization, where k-dimensional uniformity is most important for generating the initial ant population, as shown in Figure 4.
In both ACO and EACO, 30 artificial ants are used, with a pheromone decay factor of 0.75 and a solution archive of size 1500 × 13. The first 12 columns store the group indices used to generate the solvent molecule, and the last column stores the objective-function value. The algorithm terminates when the absolute difference between successive objective-function values falls below 10-6 or when the maximum limit of 20,000 iterations is reached. Because solvent generation involves multiple constraints, the objective function is formulated as a penalty function. Following the oracle penalty method reported by Kalagnanam and Diwekar,[30] the objective function is transformed into an equality constraint of the form OF−Ω, where Ω is the oracle value and is set to 3. Further details on the oracle penalty method are available in Diwekar and Gebreslassie.[24]

3.5. Efficient Simulated Annealing (ESA)

SA uses two probabilities: the acceptance probability, A i j , and the generational probability, G i j .[27] The acceptance probability determines whether a candidate configuration is accepted, whereas the generational probability controls how new configurations are sampled near the current solution. SA efficiency depends strongly on the generational probability.[27] In conventional SA, these probabilities are generated using uniformly distributed pseudo-random numbers, which can cluster sampled moves and require longer Markov chains. Several SA variants, including fast SA and hybrid SA, modify the generational probability, cooling schedule, or combine SA with methods such as differential evolution and genetic algorithms.[31,32,33] However, these approaches do not directly address sampling uniformity. Kim and Diwekar improved configurational sampling by incorporating HSS into SA, leading to efficient simulated annealing (ESA).[27] As shown in Figure 5, ESA replaces pseudo-random numbers with HSS to generate configuration moves, thereby improving k-dimensional uniformity and reducing the Markov chain length required at each temperature level. Studies report that ESA is approximately 30%–54% more efficient than conventional SA in terms of total moves.[27]
In this study, the initial SA temperature is 50 and is reduced by a cooling factor of 0.85 until it reaches the freezing temperature of 10-7. The Markov chain length is 1000 for both SA and ESA. SA uses pseudo-random numbers for all four generational probabilities and the acceptance probability. In ESA, the four generational probabilities are generated using HSS, while the acceptance probability remains pseudo-random. Both algorithms terminate when the temperature falls below the freezing temperature or when the candidate solvent remains unchanged for 10 consecutive moves across different samples.

4. Results and Discussions

This study evaluates the efficiency of the proposed metaheuristic algorithms for solving the CAMD problem of selecting solvents to extract octacosanol from a multicomponent sugarcane wax mixture. In Equation 13, the minimum net solvent selectivity value (βnet, min) is set to 1.2, and the top 20 candidate solvents are generated using ACO, SA, EACO, and ESA.

4.1. Candidate Solvents Using ACO and SA

Based on the algorithm parameters mentioned in the methodology section, the candidate solvents generated using ACO and SA are listed in Table 2 and Table 3, respectively.
As shown in Table 2 and Table 3 and Figure 6, the top nine solvents generated by ACO and SA are identical. However, SA identified better candidates among the bottom 11 solvents. Despite this improvement, SA required substantially more objective function evaluations than ACO: 1,271,876 compared with 153,330. ACO takes 12% of the computational cost as compared to SA. This higher computational cost arises because SA produces one solution per run and must be repeated with different initial values to obtain multiple solutions, whereas ACO generates all 20 solutions in a single run.

4.2. Candidate Solvents Using EACO and ESA

The same optimization problem was then solved using EACO and ESA. Table 4 lists the candidate solvents obtained from EACO for the problem defined in Equation 7. Several optimal solvents generated by EACO were also identified by conventional ACO, as shown in Figure 7. The top 15 solvents are identical for EACO and ACO, while EACO produced better candidates among the remaining five. EACO required 131,550 objective function evaluations to obtain these solvents, and reducing the computational requirements by 14.2 %.
Similarly, Table 5 lists the optimal solvents obtained by solving the solvent-selection optimization problem with ESA. Similar to Figure 7, Figure 8 compares the net distribution coefficient values for the solvents generated by SA and ESA, excluding the top five candidates. ESA again outperformed SA by identifying five better solvents among the lower-ranked 20 candidates. ESA required 1,264,327 objective function evaluations, which is 0.59% fewer than SA.
Figure 9 presents solvents obtained by ESA and EACO. This figure shows the same trend as the comparison of SA and ACO in Figure 6. The top 9 solvents found using EACO are same as that of found in ESA case, while the bottom 11 are better solvents found by ESA as compared EACO. However, EACO takes 10.4% of the computational cost as compared to ESA.

5. Conclusion

This study developed and evaluated a CAMD-based framework for selecting solvents to extract octacosanol from multicomponent sugarcane wax. The solvent-selection problem was formulated as a mixed-integer nonlinear programming problem using UNIFAC-based estimates of the net distribution coefficient and net solvent selectivity. Four metaheuristic solvers—ACO, SA, EACO, and ESA—were compared to assess both solvent quality and computational efficiency. Across all methods, the highest-ranked candidates were consistent, indicating that the CAMD formulation reliably identifies strong solvent structures for the target extraction. Chemically, the best candidate solvent, ethane, achieved the highest net distribution coefficient, and other solvents, aldehyde- and carboxylic-acid-containing structures, also satisfied the selectivity constraint and appeared among the top candidates. Incorporating the Hammersley sequence sampling improved search performance by distributing samples more uniformly across the solution space. EACO reduced the number of objective function evaluations compared with ACO while retaining the same top-ranked solvents and improving some lower-ranked candidates. ESA also identified better lower-ranked candidates than SA, although its reduction in computational cost was modest. Overall, ACO- and EACO-based approaches were substantially more computationally efficient than SA-based approaches for generating a list of candidate solvents, whereas SA and ESA provided improved diversity among lower-ranked solutions. These results show that quasi-random sampling can enhance metaheuristic solvent design and that the proposed CAMD framework is robust for complex nutraceutical extraction problems.

References

  1. Kim K-J, Diwekar UM, Joback KG. Greener solvent selection under uncertainty. ACS Symposium Series, vol. 819, Washington, DC; American Chemical Society; 1999; 2002, p. 224–37.
  2. Kim KJ, Diwekar UM. Efficient combinatorial optimization under uncertainty. 2. Application to stochastic solvent selection. Ind Eng Chem Res 2002;41:1285–96. [CrossRef]
  3. Papadimitriou C, Steiglitz K. Combinatorial Optimization: Algorithms and Complexity. vol. 32. New York: DOVER PUBLICATIONS, INC.; 1982. [CrossRef]
  4. Blum C. Ant colony optimization: Introduction and recent trends. Phys Life Rev 2005;2:353–73. [CrossRef]
  5. Nistala VS, Bhartiya S, Juvekar V, Diwekar UM. Designing Solvents to Extract Nutraceuticals from Multicomponent Sugar Cane Wax Using a Computer-Aided Molecular Design (CAMD) Approach. Ind Eng Chem Res 2026;65:11655–70. [CrossRef]
  6. Food and Agriculture Organization of the United Nations. Food and Agriculture Organization of the United Nations (2025) – with major processing by Our World in Data. “Sugar cane production – UN FAO” 2026.
  7. Sharma R, Matsuzaka T, Kaushik MK, Sugasawa T, Ohno H, Wang Y, et al. Octacosanol and policosanol prevent high-fat diet-induced obesity and metabolic disorders by activating brown adipose tissue and improving liver metabolism. Sci Rep 2019;9. [CrossRef]
  8. Shen J, Luo F, Lin Q. Policosanol: Extraction and biological functions. J Funct Foods 2019;57:351–60. [CrossRef]
  9. Marinangeli CPF, Jones PJH, Kassis AN, Eskin MNA. Policosanols as nutraceuticals: Fact or fiction. Crit Rev Food Sci Nutr 2010;50:259–67. [CrossRef]
  10. Asikin Y, Takahashi M, Hirose N, Hou DX, Takara K, Wada K. Wax, policosanol, and long-chain aldehydes of different sugarcane (Saccharum officinarum L.) cultivars. European Journal of Lipid Science and Technology 2012;114:583–91. [CrossRef]
  11. Kamchonemenukool S, Ho CT, Boonnoun P, Li S, Pan MH, Klangpetch W, et al. High Levels of Policosanols and Phytosterols from Sugar Mill Waste by Subcritical Liquefied Dimethyl Ether. Foods 2022;11. [CrossRef]
  12. Irmak S, Dunford NT, Milligan J. Policosanol contents of beeswax, sugar cane and wheat extracts. Food Chem 2006;95:312–8. [CrossRef]
  13. Attard TM, McElroy CR, Rezende CA, Polikarpov I, Clark JH, Hunt AJ. Sugarcane waste as a valuable source of lipophilic molecules. Ind Crops Prod 2015;76:95–103. [CrossRef]
  14. Cignitti S, Rodriguez-Donis I, Abildskov J, You X, Shcherbakova N, Gerbaud V. CAMD for entrainer screening of extractive distillation process based on new thermodynamic criteria. Chemical Engineering Research and Design 2019;147:721–33. [CrossRef]
  15. Salazar J, Diwekar U, Joback K, Berger AH, Bhown AS. Solvent selection for post-combustion CO2 capture. Energy Procedia 2013;37:257–64. [CrossRef]
  16. Trevizo C, Daniel D, Nirmalakhandan N. Screening alternative degreasing solvents using multivariate analysis. Environ Sci Technol 2000;34:2587–95. [CrossRef]
  17. Kim KJ, Diwekar UM. Integrated solvent selection and recycling for continuous processes. Ind Eng Chem Res 2002;41:4479–88. [CrossRef]
  18. Button A, Merk D, Hiss JA, Schneider G. Automated de novo molecular design by hybrid machine intelligence and rule-driven chemical synthesis. Nat Mach Intell 2019;1:307–15. [CrossRef]
  19. Fredenslund A, Gmehling J, Rasmussen P. Vapor-liquid equilibria using UNIFAC : a group contribution method. 1st ed. Amsterdam: Elsevier Scientific Pub. Co.; 1977.
  20. Stodola P, Michenka K, Nohel J, Rybanský M. Hybrid algorithm based on ant colony optimization and simulated annealing applied to the dynamic traveling salesman problem. Entropy 2020;22. [CrossRef]
  21. Wu Y, Wang H, Li M, Tan H, Wang D, Sheng M. The Adaptive Two-Stage Ant Colony Simulated Annealing Algorithm for Solving The Traveling Salesman Problem. RAIRO - Operations Research 2025;59:1199–213. [CrossRef]
  22. Anggraeni DAF, Dianutami VR, Tyasnurita R. Investigation of Simulated Annealing and Ant Colony optimization to Solve Delivery Routing Problem in Surabaya, Indonesia. Procedia Comput. Sci., vol. 234, Elsevier B.V.; 2024, p. 592–601. [CrossRef]
  23. Dorigo M, Stu¨tzle T. Ant Colony Optimization. The MIT Press; 2004.
  24. Diwekar U, Gebreslassie B. Efficient ant colony optimization (EACO) algorithm for deterministic optimization. International Journal of Swarm Intelligence and Evolutionary Computation 2016;1. [CrossRef]
  25. Socha K, Dorigo M. Ant colony optimization for continuous domains. Eur J Oper Res 2008;185:1155–73. [CrossRef]
  26. Aarts E, Korst J. Simulated Annealing and Boltzmann Machines: A Stochastic Approach to Combinatorial Optimization and Neural Computing. Essex, Great Britain: John Wiley & Sons Ltd.; 1989.
  27. Kim KJ, Diwekar UM. Hammersley stochastic annealing: Efficiency improvement for combinatorial optimization under uncertainty. IIE Transactions (Institute of Industrial Engineers) 2002;34:761–77. [CrossRef]
  28. van Laarhoven P, Aarts E. Simulated Annealing: Theory and Applications. Springer-Science+ Business Media, B. V. ; 1992.
  29. Gebreslassie BH, Diwekar UM. Efficient ant colony optimization for computer aided molecular design: Case study solvent selection problem. Comput Chem Eng 2015;78:1–9. [CrossRef]
  30. Kalagnanam JR, Diwekar UM. An efficient sampling technique for off-line quality control. Technometrics 1997;39:308–19. [CrossRef]
  31. Szu H, Hartley R. Fast simulated annealing. Phys Lett A 1987;122:157–62. [CrossRef]
  32. Azizi N, Zolfaghari S, Liang M. Hybrid simulated annealing with memory: An evolution-based diversification approach. Int J Prod Res 2010;48:5455–80. [CrossRef]
  33. Yu X, Liu Z, Wu X, Wang X. A hybrid differential evolution and simulated annealing algorithm for global optimization. Journal of Intelligent and Fuzzy Systems 2021;41:1375–91. [CrossRef]
Figure 1. Schematic flowchart of solvent generation using CAMD.
Figure 1. Schematic flowchart of solvent generation using CAMD.
Preprints 222115 g001
Figure 2. Multicomponent system present in sugarcane wax and their mole fractions.
Figure 2. Multicomponent system present in sugarcane wax and their mole fractions.
Preprints 222115 g002
Figure 3. 100 sample points ranging from 0 to 1 generated using Monte Carlo and Hammersley sequence sampling techniques.
Figure 3. 100 sample points ranging from 0 to 1 generated using Monte Carlo and Hammersley sequence sampling techniques.
Preprints 222115 g003
Figure 4. The working mechanism of ACO algorithm with HSS initialization. Note OF: objective function, DV: decision variable.
Figure 4. The working mechanism of ACO algorithm with HSS initialization. Note OF: objective function, DV: decision variable.
Preprints 222115 g004
Figure 5. The working mechanism of SA algorithm with HSS initialization.
Figure 5. The working mechanism of SA algorithm with HSS initialization.
Preprints 222115 g005
Figure 6. Comparison of net distribution coefficient values of solvents generated using SA and ACO, except the top five.
Figure 6. Comparison of net distribution coefficient values of solvents generated using SA and ACO, except the top five.
Preprints 222115 g006
Figure 7. Comparison of net distribution coefficient values of solvents generated using ACO and EACO, except the top five.
Figure 7. Comparison of net distribution coefficient values of solvents generated using ACO and EACO, except the top five.
Preprints 222115 g007
Figure 8. Comparison of net distribution coefficient values of solvents generated using SA and ESA, except the top five.
Figure 8. Comparison of net distribution coefficient values of solvents generated using SA and ESA, except the top five.
Preprints 222115 g008
Figure 9. Comparison of net distribution coefficient values of solvents generated using ESA and EACO, except the top five.
Figure 9. Comparison of net distribution coefficient values of solvents generated using ESA and EACO, except the top five.
Preprints 222115 g009
Table 1. Various UNIFAC groups ranging from general hydrocarbons to aromatic (groups starting with “A”), and a few other functional groups such as hydroxyl, carboxylic acid, aldehyde, and ketone.
Table 1. Various UNIFAC groups ranging from general hydrocarbons to aromatic (groups starting with “A”), and a few other functional groups such as hydroxyl, carboxylic acid, aldehyde, and ketone.
i N 2 ( i ) i N 2 ( i ) i N 2 ( i ) i N 2 ( i ) i N 2 ( i )
1 CH3 2 CH2< 3 –CH< 4 >C< 5 H2O
6 CH2=CH– 7 –CH=CH– 8 –CH=C< 9 CH2=C< 10 –OH
11 ACH 12 AC 13 ACCH3 14 ACCH2 15 ACCH
16 ACOH 17 CH3–CO– 18 –CH2–CO– 19 –CHO 20 –COOH
Table 2. List of candidate solvents generated using ACO.
Table 2. List of candidate solvents generated using ACO.
Sno. UNIFAC groups in the candidate solvent m n e t β n e t
1 2 CH3 529.49 1.69
2 1 CH3, 1 CHO 168.32 7.29
3 1 CH3, 1 CH2, 1 CHO 64.58 5.41
4 2 CH3, 1 CH, 1 CHO 30.89 4.34
5 1 CH3, 2 CH2, 1 CHO 30.71 4.34
6 2 CH3, 1 CH2, 1 CH, 1 CHO 17.43 3.65
7 1 CH3, 3 CH2, 1 CHO 17.36 3.65
8 3 CH3, 1 C, 1 CHO 15.71 3.67
9 3 CH3, 1 CH2, 1 C, 1 CHO 10.35 3.20
10 2 CH3, 1 CH=C, 1 CHO 4.03 3.17
11 2 CH3, 1 CH, 1 COOH 3.408 3.11
12 1 CH3, 2 CH2, 1 COOH 3.401 3.11
13 1 CH3, 1 CH2, 1 COOH 3.07 4.29
14 2 CH3, 1 CH2, 1 CH, 1 COOH 3.03 2.43
15 1 CH3, 3 CH2, 1 COOH 3.03 2.43
16 3 CH3, 1 C, 1 COOH 2.84 2.42
17 2 CH3, 2 CH2, 1 CH, 1 COOH 2.57 2.03
18 3 CH3, 1 CH2, 3 CH, 1 CHO, 1 COOH 2.45 3.54
19 1 CH3, 3 CH2, 1 CH, 1 CHO, 1 COOH 2.23 4.38
20 4 CH3, 1 CH2, 2 CH, 1 CH=C, 2 CHO 1.48 2.97
Table 3. List of candidate solvents generated using SA.
Table 3. List of candidate solvents generated using SA.
Sno. UNIFAC groups in the candidate solvent m n e t β n e t
1 2 CH3 529.49 1.69
2 1 CH3, 1 CHO 168.32 7.29
3 1 CH3, 1 CH2, 1 CHO 64.58 5.41
4 2 CH3, 1 CH, 1 CHO 30.89 4.34
5 1 CH3, 2 CH2, 1 CHO 30.71 4.34
6 2 CH3, 1 CH2, 1 CH, 1 CHO 17.43 3.65
7 1 CH3, 3 CH2, 1 CHO 17.36 3.65
8 3 CH3, 1 C, 1 CHO 15.71 3.67
9 3 CH3, 2 CH, 1 CHO 11.19 3.18
10 2 CH3, 2 CH2, 1 CH, 1 CHO 11.15 3.18
11 1 CH3, 4 CH2, 1 CHO 11.12 3.18
12 3 CH3, 1 CH2, 1 C, 1 CHO 10.35 3.20
13 3 CH3, 1 CH2, 2 CH, 1 CHO 7.84 2.83
14 2 CH3, 3 CH2, 1 CH, 1 CHO 7.82 2.83
15 1 CH3, 5 CH2, 1 CHO 7.80 2.83
16 2 CH3, 1 CH=C, 1 CHO 4.03 3.17
17 2 CH3, 1 CH2, 1 CH=C, 1 CHO 3.73 2.81
18 2 CH3, 1 CH, 1 COOH 3.408 3.11
19 1 CH3, 2 CH2, 1 COOH 3.401 3.11
20 1 CH3, 1 CH2, 1 COOH 3.07 4.29
Table 4. List of candidate solvents generated using EACO.
Table 4. List of candidate solvents generated using EACO.
Sno. UNIFAC groups in the candidate solvent m n e t β n e t
1 2 CH3 529.49 1.69
2 1 CH3, 1 CHO 168.32 7.29
3 1 CH3, 1 CH2, 1 CHO 64.58 5.41
4 2 CH3, 1 CH, 1 CHO 30.89 4.34
5 1 CH3, 2 CH2, 1 CHO 30.71 4.34
6 2 CH3, 1 CH2, 1 CH, 1 CHO 17.43 3.65
7 1 CH3, 3 CH2, 1 CHO 17.36 3.65
8 3 CH3, 1 C, 1 CHO 15.71 3.67
9 3 CH3, 1 CH2, 1 C, 1 CHO 10.35 3.20
10 2 CH3, 1 CH=C, 1 CHO 4.03 3.17
11 2 CH3, 1 CH, 1 COOH 3.408 3.11
12 1 CH3, 2 CH2, 1 COOH 3.401 3.11
13 2 CH3, 2 CH2, 2 CH, 2 CHO 3.07 3.78
14 1 CH3, 1 CH2, 1 COOH 3.07 4.29
15 2 CH3, 1 CH2, 1 CH, 1 COOH 3.03 2.43
16 1 CH3, 3 CH2, 1 COOH 3.03 2.43
17 2 CH3, 1 CH2, 2 CH, 2 CHO 3.02 4.08
18 3 CH3, 1 C, 1 COOH 2.84 2.42
19 2 CH3, 1 CH2, 1 C, 2 CHO 2.75 4.51
20 3 CH3, 2 CH, 1 COOH 2.57 2.03
Table 5. List of candidate solvents generated using ESA.
Table 5. List of candidate solvents generated using ESA.
Sno. UNIFAC groups in the candidate solvent m n e t β n e t
1 2 CH3 529.49 1.69
2 1 CH3, 1 CHO 168.32 7.29
3 1 CH3, 1 CH2, 1 CHO 64.58 5.41
4 2 CH3, 1 CH, 1 CHO 30.89 4.34
5 1 CH3, 2 CH2, 1 CHO 30.71 4.34
6 2 CH3, 1 CH2, 1 CH, 1 CHO 17.43 3.65
7 1 CH3, 3 CH2, 1 CHO 17.36 3.65
8 3 CH3, 1 C, 1 CHO 15.71 3.67
9 3 CH3, 2 CH, 1 CHO 11.19 3.18
10 2 CH3, 2 CH2, 1 CH, 1 CHO 11.15 3.18
11 1 CH3, 4 CH2, 1 CHO 11.12 3.18
12 3 CH3, 1 CH2, 1 C, 1 CHO 10.35 3.20
13 3 CH3, 1 CH2, 2 CH, 1 CHO 7.84 2.83
14 2 CH3, 3 CH2, 1 CH, 1 CHO 7.82 2.83
15 1 CH3, 5 CH2, 1 CHO 7.80 2.83
16 4 CH3, 2 CH2, 1 C, 1 CHO 7.42 2.85
17 3 CH3, 2 CH2, 1 C, 1 CHO 7.39 2.85
18 4 CH3, 3 CH, 1 CHO 5.88 2.57
19 3 CH3, 2 CH2, 2 CH, 1 CHO 5.87 2.57
20 2 CH3, 4 CH2, 1 CH, 1 CHO 5.86 2.57
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings