2.1. Computational Prototype and Input Databases
The meteorological database serves as the first layer of information in the proposed automated PV sizing approach, as the FE-oriented objective function depends directly on the hourly representation of the solar resource and ambient operating conditions. Hourly data were obtained from the NASA POWER database through its API service for the 2016–2025 period. The extracted variables were ALLSKY_SFC_SW_DWN, used as global horizontal irradiance (GHI) for PV performance calculations; T2M, corresponding to ambient temperature at 2 m; and RH2M, corresponding to relative humidity at 2 m. These variables provide the minimum meteorological information required to estimate the plane-of-array irradiance, determine the cell temperature, and calculate the hourly AC energy output of each candidate PV configuration. The use of a single corrected physical irradiance series also ensures that all candidate designs are evaluated under identical solar-resource conditions.
The case study corresponds to the installed PV system in La Tebaida, Quindio, Colombia, located at latitude N, longitude W, and an elevation of approximately 1500 m a.s.l. This geographic location was used to query the NASA POWER data and to define the solar-geometry calculations used throughout the PV performance model. At this location, the raw dataset contains 87,672 hourly records, corresponding to a continuous ten-year UTC series that was converted to the America/Bogota working time zone for the subsequent PV simulation. The preprocessing stage first verified the temporal continuity of the hourly index to prevent missing or duplicated timestamps from propagating into the energy calculation. Physical consistency checks were then applied to the irradiance series, including detection of negative and non-physical values, as well as night-time GHI inconsistencies based on the solar zenith angle. When required, irradiance values outside physically admissible ranges were clipped or corrected before calculating the direct, diffuse, and plane-of-array irradiance components.
The cleaned meteorological database was used consistently in all stages of the computational workflow, including irradiance transposition, PV performance modeling, local validation, ENFICC calculation, and automated sizing. No statistical, data-driven, forecasted, or alternative irradiance series was used as an input to the optimization process. Therefore, the FE-oriented objective function was evaluated exclusively using the corrected NASA POWER physical irradiance series. This convention avoids inconsistencies that would arise if different solar-resource series were used for validation, energy conversion, and sizing. It also makes the comparison among candidate PV configurations dependent on design decisions and physical constraints rather than on changes in the irradiance input.
The available installation surface was represented as a rectangular area of 32 m × 33 m. A 1.0 m perimeter margin was excluded from this area to account for practical spacing and accessibility requirements, yielding a net usable area of approximately 969.8 m
2. This net area was imposed as a geometric constraint during the evaluation of candidate PV layouts. The resulting spatial representation is consistent with prototype-level PV sizing studies in which the optimal configuration is constrained by land availability, module dimensions, tilt angle, azimuth angle, and inverter compatibility [
14,
15]. Thus, the sizing process combines meteorological, electrical, and geometric information within a single reproducible computational framework.
The second information layer corresponds to the photovoltaic module database used to define the discrete component alternatives available to the automated PV sizing approach. This catalog was derived from the California Energy Commission (CEC) database and filtered to retain commercially relevant crystalline-silicon modules compatible with the technical ranges required for ENFICC-oriented PV-system modeling. The filtering criteria included monocrystalline technology, nominal STC power above 350 W, feasible module area, complete electrical parameters, a negative temperature coefficient of maximum power, and exclusion of building-integrated photovoltaic modules. After filtering, 977 modules remained from the original source database. Eight additional high-power bifacial modules representative of the 2024-2025 commercial market were manually incorporated to extend the upper range of available module ratings. The resulting expanded database contains 985 PV modules from 58 manufacturers, with an availability period spanning 2019-2024.
The technical attributes retained from the module database were selected based on their roles in the PV performance model and the feasibility constraints of the automated sizing procedure. For each candidate module, the workflow used the nominal power at standard test conditions, , the voltage and current at the maximum power point, and , the open-circuit voltage, , the physical module area, and the estimated conversion efficiency. These variables determine the DC capacity of each candidate layout, the string-level voltage and current limits, the land-use constraint, and the expected energy yield under the corrected NASA POWER irradiance series. Therefore, the sizing procedure does not rely on an abstract or continuous module model, but on a traceable set of commercially available component alternatives. This design choice is relevant because the best achievable ENFICC value is constrained by the technical diversity and parameter boundaries of the component database. Consequently, expanding and documenting the module catalog contributes directly to the reproducibility of the proposed sizing methodology.
Figure 1 summarizes the numerical span of the filtered and expanded PV module database. The range plot is included immediately after the module-database description to make explicit the module search space available to the automated sizing procedure before the optimization results are presented. The expanded database covers STC powers from 350.0 to 680.0 W, module areas from 1.63 to 2.98 m2, and estimated efficiencies from 13.89% to 23.04%. This range allows the decision variables to span compact high-efficiency modules, intermediate alternatives, and larger high-power modules under the same land, electrical, and inverter-compatibility constraints. As a result, the selected module is not imposed a priori, but emerges from the interaction between the component database, the corrected physical irradiance input, the feasibility constraints, and the ENFICC-oriented objective function.
Figure 1.
Operating ranges of the photovoltaic module database used in the genetic algorithm. The curves represent the nominal power, maximum power point voltage, and estimated efficiency for each module in the dataset.
Figure 1.
Operating ranges of the photovoltaic module database used in the genetic algorithm. The curves represent the nominal power, maximum power point voltage, and estimated efficiency for each module in the dataset.
The third information layer corresponds to the inverter database, which defines the discrete set of AC conversion technologies available to the automated PV sizing approach. This catalog was constructed from the Sandia inverter database and subsequently filtered according to AC power rating, maximum DC voltage limit, MPPT operating window, maximum DC current capability, and parameter-validity criteria. The resulting database contains 222 inverter entries from 40 manufacturers, with an availability period spanning 2018–2024. Five additional commercial inverter models with high market penetration in Colombia were included to improve the representativeness of the search space for local PV applications. Among these models, the Solis 60K-LV-5G was retained because it matches the inverter used in the La Tebaida validation case and is therefore a relevant feasible alternative among the candidate configurations. The final inverter database covers nominal AC power ratings from 15 to 200 kW, which is suitable for the PV system scale evaluated in this study.
The technical attributes retained from the inverter database were selected based on their roles in the electrical-feasibility constraints and the AC energy-conversion model. For each candidate inverter, the workflow used the nominal AC power, , the maximum admissible DC voltage, , the MPPT voltage range, and the maximum DC current. These parameters determine whether the voltage and current produced by a candidate module-string configuration are compatible with the selected inverter. They also define the admissible DC/AC ratio, the onset of AC clipping, and the maximum AC power that each candidate plant layout can inject. Therefore, the automated sizing procedure does not select an inverter from an abstract capacity interval, but from a traceable database of technically feasible commercial devices. This formulation ensures that the sizing process simultaneously accounts for energy yield, electrical compatibility, and component-level implementation constraints.
Figure 2 presents the numerical range of the filtered and expanded inverter database. The range plot is included immediately after the inverter-database description to make explicit the electrical search space available to the automated sizing procedure. The database spans nominal AC powers from 15.0 to 200.0 kW and DC power ratings from 15.46 to 206.26 kW. In addition, the MPPT voltage windows cover a broad operating region, allowing the sizing procedure to evaluate different string lengths and module-inverter pairings under realistic voltage constraints. This inverter parameterization is relevant because compatibility must be verified before calculating hourly AC generation and the resulting FE-oriented objective value. Consequently, the selected inverter emerges from the interaction between the module database, the corrected NASA POWER irradiance input, the electrical constraints, and the objective of maximizing firm energy.
Figure 2.
Operating ranges of the inverter module database used in the genetic algorithm. The curves represent the nominal AC power, maximum DC voltage, and maximum MPPT current for each module in the dataset.
Figure 2.
Operating ranges of the inverter module database used in the genetic algorithm. The curves represent the nominal AC power, maximum DC voltage, and maximum MPPT current for each module in the dataset.
2.2. Photovoltaic and Firm Energy Modeling Framework
The proposed methodology represents the photovoltaic (PV) plant as a sequential conversion process in which the available solar resource is transformed into hourly AC generation. Subsequently, the Firm-Energy (FE) indicator results from the simulated PV generation profile and the applicable regulatory constraints. The modeling sequence characterizes the atmospheric conditions associated with the hourly irradiance input, estimates the irradiance incident on the PV plane, converts this irradiance into DC and AC power, and evaluates the resulting monthly energy profile through a critical-month criterion. This formulation allows the automated PV sizing approach to compare candidate configurations based on their contribution during the most restrictive month, rather than only maximizing annual energy production.
The first step in the sequential conversion estimates the clearness index,
, which quantifies the fraction of extraterrestrial solar irradiance that reaches the site on a horizontal plane. This index is calculated as the ratio between the global horizontal irradiance and the extraterrestrial normal irradiance projected onto the horizontal plane:
where
t denotes the hourly simulation step,
the global horizontal irradiance, and
the solar zenith angle. The term
corresponds to the extraterrestrial normal irradiance associated with the day of the year corresponding to each hourly record. Thus, hourly variables such as
and
vary within the day, whereas the extraterrestrial correction varies with the day index
n. The hour-day relation
maps each hourly timestamp to its corresponding day of the year. The extraterrestrial normal irradiance is calculated using the eccentricity correction associated with the annual variation in the Earth–Sun distance:
being
the solar constant and
n the day of the year. The coefficient
approximates the annual eccentricity effect, and the denominator 365 defines the yearly periodicity used in the implementation. Consequently, Equation (1) compares the measured or estimated irradiance at the site with the extraterrestrial irradiance available under the same solar-position conditions.
Given
, the global horizontal irradiance decomposes into diffuse horizontal irradiance
and direct normal irradiance
. This decomposition allows tilted PV modules to receive direct, diffuse, and ground-reflected irradiance components with different geometric factors according to the Erbs correlation from the clearness index [
17]:
The irradiance components are then obtained as:
with
denoting the empirical diffuse-fraction model. The decomposition is applied only under daylight conditions to avoid numerical instabilities when
approaches zero, hence mitigating errors in
and
that propagate to the plane-of-array irradiance and, consequently, to monthly AC generation. This effect may be significant during cloudy periods, which can determine the critical-month value used in the FE-oriented objective function.
The irradiance incident on the tilted PV plane is estimated as the sum of beam, sky-diffuse, and ground-reflected components. In this implementation, an isotropic sky transposition model is adopted. The resulting plane-of-array irradiance,
, is given by:
where
estimates the beam transposition ratio from the incidence angle on the module surface
and the solar zenith angle
.
corresponds to the module tilt angle. The terms
and
represent the isotropic sky-view and ground-view factors, respectively. The parameter
denotes ground albedo. Here,
is adopted when site-specific albedo measurements are unavailable to represent a generic ground-reflection condition. Since
denotes the irradiance received by the module plane, this stage links array orientation, tilt, and topology with the monthly energy profile used by the automated PV sizing approach.
Numerical and physical safeguards were applied before converting irradiance into electrical power. Negative irradiance values were clipped to zero, and irradiance decomposition was applied only when . The beam transposition ratio was evaluated only under valid daylight and front-side incidence conditions; otherwise, the beam contribution was set to zero. Non-finite values of , , , , or the incidence-angle modifier were replaced by zero. In addition, monthly FE calculations retained only months satisfying the minimum data-coverage threshold defined in the implementation. In this work, months with at least 99% hourly coverage were retained for the critical-month calculation, thereby preventing incomplete edge months from controlling the FE estimate.
It is worth noting that the irradiance received by the module is not converted into electrical power under ideal thermal conditions. Part of the incident energy increases the module temperature, and higher cell temperature reduces the power output of crystalline-silicon PV modules. Hence, the cell temperature is estimated using the Faiman model:
being
the cell temperature, the ambient temperature, and the wind speed, respectively. The coefficients
and
represent the thermal exchange between the module and the surrounding environment. The former captures the basic heat-loss mechanism, whereas the latter accounts for the additional cooling effect of wind. When local wind measurements are unavailable, the model assumes
to maintain numerical consistency.
In addition to thermal effects, the proposed PV generation model also considers optical losses caused by non-normal solar incidence, which are parameterized by the incidence-angle modifier
. Particularly, the formulation from the American Society of Heating, Refrigerating and Air-Conditioning Engineers (ASHRAE) is adopted:
with coefficient
approximating the increase in reflection losses as the incidence angle increases. A default
is used when module-specific optical characterization is unavailable. The optical correction is applied only for valid front-side incidence conditions and is clipped to non-negative values to avoid non-physical optical gains or losses at extreme incidence angles. Although early-morning and late-afternoon hours usually have lower irradiance, their cumulative contribution may affect the monthly energy profile. Therefore, including
improves the physical consistency of the FE-oriented PV performance model.
Once the irradiance, temperature, and optical effects are defined, the DC power produced by the PV array is estimated as:
being
the installed DC capacity under Standard Test Conditions,
the reference irradiance, and
the reference cell temperature. The coefficient
stands for the module power temperature coefficient, obtained from the manufacturer’s datasheet or module database. Since the temperature coefficient is typically negative for crystalline-silicon modules, increases in cell temperature reduce
. The term
represents DC-side losses, including wiring, mismatch, and other losses not explicitly resolved in the prototype implementation. In this formulation,
and
are both expressed in watts.
Subsequently, a central calculation stage converts the DC power into net AC power through the inverter system. This stage accounts for inverter conversion efficiency, AC clipping, AC-side losses, and transformer losses:
where
corresponds to the DC power assigned to inverter
j,
denotes the inverter efficiency,
stands for the nominal AC rating of inverter
j, and
indicates the number of inverters in the plant. The minimum operator
prevents inverter output from exceeding either the converted DC power or the nominal AC rating. The loss factors
and
represent AC-side and transformer losses, respectively. Here, constant inverter efficiency is used as a reduced-order approximation when a detailed inverter efficiency curve is unavailable to preserve the main effects of inverter conversion and AC clipping while maintaining computational efficiency during the automated sizing process. If manufacturer-specific efficiency curves are available,
is replaced by a load-dependent inverter model.
2.3. Genetic Algorithm Formulation
The PV sizing problem involves continuous (e.g., tilt angle, azimuth angle, DC/AC ratio), discrete (e.g., number of series and parallel components), and categorical (e.g., module selection, inverter selection, topology) decision variables. Therefore, each candidate solution encodes the selected PV module from the expanded CEC-based database, the selected inverter from the expanded SANDIA-based database, the array topology, the DC/AC ratio, the installation angles, and the feasible number of modules and inverters. Candidate configurations are evaluated using the equivalent daily energy of the worst-performing month as the fitness function, as described in Equation (10). Infeasible solutions, including those that violate land-use or component-compatibility constraints, receive a near-zero fitness value. Although other population-based metaheuristics, such as particle swarm optimization (PSO), differential evolution (DE), and NSGA-II, have demonstrated strong performance in renewable-energy optimization [
11,
13], the present work selects the Genetic Algorithm (GA) as the solver for the PV sizing problem because its chromosome representation naturally accommodates mixed discrete–continuous decision variables and discrete catalogue-based component selection without requiring gradient information, making it well suited to the proposed prototype.
Figure 3 summarizes the architecture of the implemented GA. The workflow begins with the admissible design space defined by the corrected NASA POWER physical irradiance series, the filtered equipment databases, the land-use constraint, and the ENFICC regulatory requirements. Each chromosome represents a complete PV plant configuration and is evaluated through the irradiance-processing chain, PV performance equations, electrical compatibility checks, and the ENFICC objective function. The evolutionary cycle then applies selection, crossover, mutation, elitism, and replacement until the stopping criterion is reached. The GA used a population size of
individuals evolving over
generations. Parent selection was performed using tournament selection, with
parents participating in each mating cycle, while elitism preserved the best
individuals at every generation. New candidate solutions were generated through uniform crossover (
) and random mutation applied to the mixed discrete–continuous chromosome encoding the PV system design variables. The fitness function corresponded to the equivalent daily energy of the worst-performing month, whereas infeasible individuals violating land-use constraints or PV module–inverter compatibility were assigned a penalized objective value, guiding the evolutionary search toward feasible and high-performing photovoltaic configurations.
As illustrated in
Figure 3, the GA operates on a mixed discrete–continuous search space. The discrete genes are associated with the selection of the PV module, inverter model, and topology, whereas the bounded integer and continuous genes define the number of inverters, the number of series-connected modules, the number of parallel strings, the tilt angle, and the azimuth angle. For each individual, the optimization routine computes the PV energy yield, verifies the electrical and geometric feasibility of the design, applies the regulatory limit and the secondary-data correction factor, and assigns the final fitness value according to the ENFICC-oriented objective function.