Preprint
Article

This version is not peer-reviewed.

Exploring the Reaction-Guided Polymer Universe with RxnChainer: Digital Synthesis, Retrosynthesis and Beyond

Submitted:

17 August 2026

Posted:

18 August 2026

You are already at the latest version

Abstract
Polymers enable countless modern technologies, yet vast regions of their chemical space remain unexplored. Traditional polymer discovery relies on chemical intuition, ingenuity, and experience (with a healthy dose of serendipity), yet it fails to leverage millions of potentially accessible and synthesizable polymer structures. Here, we present RxnChainer, a digital methodology integrating virtual polymer generation, retrosynthetic analysis, and post-polymerization modification to systematically explore reaction-guided polymer space. Using molecular entries from the Toxic Substances Control Act (TSCA) and ChEMBL databases and RxnChainer, we generated over 289 million hypothetical polymers across 44 reaction chains spanning 30 polymer classes, including polyamides, polyimides, polyesters, and polyethers. Comparison with known polymers from PolyInfo indicates that many sampled reaction-guided structures are structurally distinct from the known-polymer comparison set. We demonstrate the methodology’s versatility through automated retrosynthetic planning for 30,000 polyesters and targeted functionalization via four post-polymerization modification pathways incorporating vinyl and nitrile pendant groups. The resulting datasets enable downstream tasks such as property-driven screening, application-specific design, and training of generative models.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Polymers are essential materials that underpin modern technology, from biomedical devices and food packaging to energy storage and electronic systems [1,2,3,4,5,6]. They exhibit remarkable chemical diversity, ranging from simple homopolymers and copolymers to complex polymer blends, composites, and formulations [7,8,9,10]. This extraordinary chemical diversity, coupled with the vast array of possible architectures, creates a polymer universe that is theoretically infinite [11]. Yet the vast majority of this space remains unexplored, and critically, it is unclear which regions are synthetically accessible and what property ranges are achievable. This knowledge gap represents both a fundamental challenge and an enormous opportunity for accelerated materials discovery.
To meet urgent contemporary challenges, industries spanning energy, environment, and manufacturing are exploring the replacement of existing materials with safer, higher-performing, and environmentally friendly alternatives [12]. Examples include novel polymers for membrane separation, fuel cells, food packaging, additive manufacturing, and high-energy-density capacitors [13,14,15,16,17,18,19]. Systematic and efficient search of reaction-guided polymer candidate space offers a pathway to identify candidates optimized for specific applications based on targeted property criteria.
Traditional polymer design relies heavily on expert intuition and trial-and-error experimentation, leading to lengthy development cycles and limited exploration of chemical space [20]. A fundamental challenge is identifying designs that are both synthetically accessible and exhibit the target properties required for specific applications. Meanwhile, emerging virtual design paradigms that leverage artificial intelligence (AI) methods have enabled accelerated polymer discovery and design [21]. Given the vast scale of the polymer chemical space, these AI-driven approaches may be effective to systematically explore reaction-guided candidate regions and guide the identification of promising candidates.
Among existing computational approaches, virtual forward synthesis (VFS) has generated hypothetical polymer structures by digitally reacting molecular entries from databases such as ChEMBL, ZINC-15, and eMolecules via known polymerization pathways [22,23,24]. Resources including Open Macromolecular Genome (OMG), SMIPOLY, PolyVERSE, and PolyUniverse provide access to millions of such polymers [19,25,26,27], with select designs validated through experimental synthesis [19,28,29]. While laboratory synthesis requires optimization of reaction conditions depending on the specific monomer pair, the computational generation of the polymers derived from reaction-compatible molecular entries via established polymerization chemistries provides a valuable starting point for experimental exploration. In addition, computational retrosynthesis offers a complementary capability, identifying commercially available precursors for target polymer structures and streamlining synthetic planning [30]. Beyond these generation and retrosynthesis capabilities, post-polymerization modification (PPM) provides an additional route to chemical diversity, enabling tuning of properties such as mechanical strength, thermal stability, or chemical resistance without redesigning base polymer architectures [31,32,33,34,35].
Here, we introduce RxnChainer, a unified computational methodology that integrates polymer generation, retrosynthetic analysis, and post-polymerization modification within a single algorithmic framework for systematically exploring reaction-guided polymer space. The methodology, shown in Figure 1, operates on a core algorithm that filters reactants based on compatibility criteria and applies reaction chains to generate products, supporting three complementary applications: forward synthesis for polymer generation, retrosynthesis for identifying synthetic precursors, and post-polymerization modification for expanding chemical diversity. Using molecular entries from the Toxic Substances Control Act (TSCA) and ChEMBL databases, we generated over 289 million hypothetical polymers spanning 44 reaction chains across 30 polymer classes [22,36]. Comparison with the PolyInfo comparison subset indicates that many sampled generated polymers are structurally distinct under the molecular representations evaluated here, suggesting that RxnChainer can expand the set of candidate polymer chemistries available for further investigation [37].
A downstream task undertaken here is to determine property ranges accessible by such a comprehensive chemical space exploration. Machine learning models trained on existing polymer datasets predict key properties of the generated polymers, including bandgap, glass transition temperature, and dielectric constant, revealing broad property ranges across polymer families. Diversity analysis of the generated polymers through Tanimoto similarity distributions and UMAP projection maps the structural landscape relative to the PolyInfo comparison subset and highlights regions with varying degrees of structural overlap. To demonstrate practical utility, the generated polymers were screened to identify polymer candidates that meet desired target property criteria as attempted before [13,18,19]. Interestingly, many previously designed polymers emerge from this screening, providing retrospective validation of the procedure adopted here. Additional downstream applications of the generated library of polymers would be to train generative models for reaction-guided polymer design, including language model pretraining tasks as have been done recently [38,39,40].

2. Methods

  • Molecular Databases.
In this work, we utilized the Toxic Substances Control Act (TSCA) inventory, a comprehensive list of chemical substances manufactured or processed in the United States [36]. The TSCA inventory is maintained by the Environmental Protection Agency (EPA) and contains information on the chemical identity, production volume, and use of each substance. Our TSCA inventory dataset (May, 2024 version) contains more than 70 thousand entries.
In addition to the TSCA inventory, we also utilized the ChEMBL database to identify potential reactant molecules for polymerization reactions [22]. The ChEMBL database is maintained by the European Bioinformatics Institute (EBI) and contains information on bioactive compounds, including their chemical structures, biological activities, and pharmacological properties. The ChEMBL database is widely used in drug discovery and development research, as it provides a wealth of information on the chemical properties and biological activities of compounds. We used the database to identify potential reactant molecules for polymerization reactions by filtering for compounds that contain specific functional groups or structural motifs that are compatible with the polymerization reaction chains we implemented. Our ChEMBL database dataset (2015 version) contains more than 1.2 million molecular entries. Both databases contain molecular structures represented in the Simplified Molecular Input Line Entry System (SMILES) format and identification numbers for each molecule.
Molecular entries from TSCA and ChEMBL were screened separately for compatibility with each polymerization reaction. For example, step-growth polymerization routes required the corresponding difunctional reactant patterns, such as AA and BB precursor structures. Only compatible entries were subjected to reaction execution. For each molecular reaction, a single generated product representation was retained, and only products containing exactly two wildcard polymer connection points ([*]) were retained as polymer repeat units. No additional category-specific standardization was applied to salts, charged species, isotopologues, stereochemical variants, or ambiguous database entries beyond the source representations and the reaction-specific filtering procedures described below.
  • Filtering of reactant molecules.
A crucial step in generating polymers through the RxnChainer methodology was identifying compatible molecules from molecular databases. We utilized the RDKit library and SMARTS patterns for substructure searches and screened for compatible molecules []. Simpler methods for filtering compatible molecules were also implemented, such as string-based search. Additionally, the RxnChainer methodology also incorporated capabilities to search for molecules with a user-specified Gasteiger charge difference for a specified SMILES arbitrary target specification (SMARTS) pattern. In this work, this capability was used for AA-BB polycondensation reactions as a heuristic filter for a similar electronic environment within each difunctional reactant. Specifically, for the absolute Gasteiger-charge differences between the two A-type reacting functional groups in an AA reactant, and separately between the two B-type reacting functional groups in the BB reactant, a threshold of less than 0.001 was used. These Gasteiger charges were computed using RDKit, and this criterion was used to select reactive sites with similar computed electronic environments rather than to infer reaction kinetics. Because the threshold is strict, it retains difunctional precursors whose two reactive sites are electronically near-equivalent, such as symmetric diacids, diols, and diamines, and excludes precursors in which the two reactive groups are in substantially different electronic environments. This filter is therefore applied only after the reaction-specific SMARTS compatibility screening and acts as an additional restriction on the AA-BB precursor pool.
  • Polymerization rules through reaction chains.
Polymerization reactions were executed within the RxnChainer methodology through the use of reaction SMARTS, which serve as the foundation for defining reaction rules, a functionality supported by RDKit [41,42]. Once a polymerization reaction was chosen for implementation, we applied the reaction rules by obtaining general templates of the reactions from the literature, textbooks, and other sources to form idealized polymer repeat units. These templates involve reacting functional groups and spacer groups with generalized reactions. We then implemented the general template within the RxnChainer methodology by encoding the reaction rules as reaction SMARTS. The number of reaction SMARTS needed to encode a polymerization reaction depends on the specific polymerization mechanism. The list of all reaction chains implemented in this work is provided in the Supplementary Information. Throughout this work, a reaction chain refers to an implemented polymerization workflow encoded by one or more sequential reaction SMARTS steps; a reaction template refers to the SMARTS-level rule describing an individual reaction step; a reaction route refers to the precursor/reactant classes combined within a reaction chain, such as diacid + diol; and a polymer class refers to the resulting polymer family, such as polyester or polyamide.
Generated polymer structures are represented as idealized repeat units in polymer SMILES (PSMILES) format, with wildcard atoms ([*]) indicating the polymer connection points. End groups and finite-chain degrees of polymerization are not explicitly represented. For AA+BB polycondensation routes, the generated structure corresponds to the idealized alternating repeat unit derived from the two difunctional precursor components. Stereochemical and regiochemical information is retained only when explicitly encoded in the input molecular representation and reaction SMARTS; when such outcomes are not specified, the generated structure should be interpreted as a connectivity-level representation rather than a complete stereochemical or regiochemical specification. AA+BB routes are represented assuming ideal 1:1 stoichiometry between the two difunctional precursors, and the repeat unit is written in a single directional orientation; deviations from ideal stoichiometry and alternative head-to-tail arrangements are not enumerated.
  • Calculation of Synthetic Accessibility Score.
The synthetic accessibility score (SA score) quantifies the ease of synthesizing a molecule and is computed from known molecular fragment data and a complexity penalty [43]. In this work, we used the SA score algorithm implemented in the RDKit library as an approximate repeat-unit complexity descriptor for hydrogen-capped polymer repeat units. To calculate the SA score for polymers, we first converted the polymer SMILES to monomer SMILES by replacing the dangling bonds ([*]) with a hydrogen ([H]) atom and then calculated the SA score of the resulting structure. This approach provides only a relative structural-complexity metric and should not be interpreted as proof of polymer synthesizability, achievable molecular weight, selectivity, conversion, purity, or processability.
  • Similarity Calculation.
The Tanimoto similarity metrics implemented in the RDKit library were used to compare the hypothetical polymers with polymer repeat units from the PolyInfo comparison subset, which measures the overlap between the molecular fingerprints of two compounds, and other structural similarity indices [44]. Morgan fingerprints were implemented for fast similarity calculations.
  • RxnChainer Modules.
RxnChainer methodology for polymer repeat-unit generation, retrosynthesis, and post-polymerization modification developed in this work has been implemented in PolymRizeTM [45], a standardized software platform for polymer informatics.

3. Results and Discussion

3.1. Overview of the RxnChainer Methodology

The RxnChainer workflow (Figure 1a) begins with a pool of reactants and a defined reaction chain. For each reaction chain, the candidate reactants are filtered based on compatibility criteria, including reaction-specific SMARTS patterns that identify compatible functional groups; additionally, Gasteiger charge difference was used for screening reactants for AA-BB polycondensation reactions [41,46,47]. Compatible reactants then proceed through the reaction chain, which may consist of multiple reaction steps, each specified by a reaction SMARTS pattern. The reaction step may input the products of a previous reaction step as reactants for the subsequent step. This sequential approach enables combining multiple reactions for complex polymerization or product generation tasks.
New hypothetical polymer designs can be generated using this approach by providing a list of reacting molecules (monomers) and a polymerization reaction template, as illustrated in Figure 1b. In the filtering stage, the input monomers are screened against the defined compatibility criteria to ensure only suitable monomers proceed to the polymerization stage. In the polymerization stage, the specified reaction chain is applied to the filtered monomers to generate hypothetical polymer structures. This approach can be used for various polymerization methods, including polyaddition, polycondensation, and ring-opening polymerization, as demonstrated later in this work.
Beyond forward synthesis, the RxnChainer methodology also enables retrosynthesis of hypothetical or known polymers. Given a target polymer structure and a set of reaction chains, this approach identifies the polymer class and then applies class-specific inverse reactions to obtain plausible reactants. The methodology further supports post-polymerization modifications for expanding chemical diversity. Starting from a base polymer, functional groups of the base polymer are identified during filtering to select compatible modification reactions from a predefined library. The group-specific reactions are then applied to the base polymers to generate a set of modified polymers. Additional technical details on the implementation and specific templates used in each stage of the RxnChainer methodology are provided in the Methods and Supplementary Information sections.

3.2. Generation of Hypothetical Polymers

In this study, we generated polymer designs utilizing the TSCA inventory [36] and the ChEMBL database [22] via RxnChainer. The TSCA inventory comprises a comprehensive list of chemical substances manufactured or processed in the United States, making it a valuable resource for identifying potential monomers for polymer synthesis. ChEMBL was selected for its extensive collection of bioactive molecules that can also serve as potential monomers for polymerization reactions. Considering molecular entries from these databases provides a broad pool of putative precursor molecules for reaction-guided polymer generation, but database membership does not by itself establish supplier availability or experimental polymerizability.
We successfully implemented 44 reaction chains using the RxnChainer methodology. The 44 reaction chains include 22 polycondensation, 13 polyaddition, and 9 ring-opening polymerization reactions. As shown in Figure 2a, the generation workflow begins with 70,000 molecules from TSCA and 1.2 million molecules from ChEMBL. Following filtering for monomers compatible with the reaction chains, 14,470 TSCA molecules (20.7%) and 293,875 ChEMBL molecules (24.5%) were retained as suitable reactants. The number of polymers generated for each class, grouped by their types and templates, are detailed in Table 1. Notably, the polycondensation reactions yield the highest number of hypothetical polymers due to the multiplicative potential arising from two distinct reactants, followed by polyaddition and ring-opening polymerization. A comprehensive list of the implemented polymerization reaction templates, including the reaction steps for representative polymers from each class, is available in the Supplementary Information.
A comparative analysis of the total number of unique polymers generated across the 30 polymer classes established in this work with those produced in previous studies, including the Open Macromolecular Genome (OMG) and SMIPOLY, is presented in Figure 2b. The comparison suggests that RxnChainer covers a broad range of polymer classes relative to previously reported generated datasets. However, these comparisons are not normalized by source database, reaction coverage, deduplication procedure, or polymer-class definitions, and should therefore be interpreted as a qualitative comparison of generated chemical-space coverage rather than a direct benchmark. For instance, polysulfonate is the most populated class in this work, accounting for over 2.5 × 10 8 generated structures. Similarly, for polyether, polyester, and polyamide—classes with significant industrial relevance—this work generates 10 7 to 10 8 structures. The enhanced coverage is particularly pronounced for condensation polymers, where the multiplicative nature of two-reactant systems (e.g., diol + dichloride for polyether, diacid + diamine for polyamide) enables combinatorial expansion of the design space. Even for less-explored classes such as polytriazole, polyoxadiazole, and polythiourethane, RxnChainer generates provides substantial coverage. Although the generated library is large, the class counts primarily reflect database composition, functional-group filtering, and reaction-template/reactant-pool combinatorics rather than experimental prevalence or demonstrated chemical feasibility. In particular, the polysulfonate class dominates the generated library because the SuFEx-based reaction chain retains a large number of compatible primary-amine-derived precursor entries and diol entries, leading to extensive two-reactant combinatorial pairing. This dominance should not be interpreted as evidence that polysulfonates are more commonly synthesized or more broadly adopted than polymer classes such as polyesters, polyamides, or polyethers; rather, polysulfonates remain specialty polymers in experimental practice.

3.3. Novelty and Synthetic Feasibility of Generated Polymers

To investigate the novelty of the polymers, we conducted a comparative analysis of the generated polymers from the top 16 polymer classes, randomly sampling a maximum of 10,000 TSCA-derived generated polymers from each class for computational efficiency. We calculated the Tanimoto similarity matrix for the selected polymers and 14,263 unique polymer repeat units from the PolyInfo comparison subset [37] using Morgan fingerprints [44,48]. The similarity metrics for each polymer class are presented in Figure 3a. For each sampled generated polymer, Tanimoto similarities were computed against all 14,263 polymer repeat units in the PolyInfo comparison subset and averaged; these per-polymer values were then averaged within each generated polymer class to obtain the class-level mean similarity reported in Figure 3a. For nearly all polymer classes in the sampled generated set, at least one Morgan-fingerprint match with Tanimoto similarity = 1.0 was identified within the PolyInfo comparison subset. This is expected because some molecular entries from TSCA and ChEMBL correspond to precursors of previously reported polymers. For four polymer classes, namely polysulfonate, polynorcantharimide, polythiourethane, and polytriazole, no Morgan-fingerprint matches with Tanimoto similarity = 1.0 were identified within the PolyInfo comparison subset. However, the absence of such a match should not be interpreted as proof that these polymers are previously unknown, because PolyInfo is not an exhaustive representation of all experimentally synthesized polymers.
The absence of PolyInfo matches for polysulfonates further indicates that this dominant generated class is poorly represented in the experimentally realized polymer comparison set, and therefore aggregate novelty or diversity statistics should be interpreted at the class level rather than as conclusions dominated by the polysulfonate class. On average, the sampled generated polymers show low similarity to the PolyInfo comparison subset, as reflected by the low average Tanimoto similarity values and their associated standard deviations. As a stricter representation-level comparison, we also performed exact canonical-SMILES matching after independent RDKit canonicalization of the generated and PolyInfo repeat units; these results are reported separately in Table S5 of the Supplementary Information.
Additional sensitivity analyses are provided in the Supplementary Information. Table S7 evaluates the robustness of class-level nearest-neighbor similarity trends to fingerprint representation using Morgan fingerprints with radii 2 and 3 and MACCS structural keys. Table S8 evaluates the sensitivity of the whole-PolyInfo mean similarity reported in Figure 3a to five independent random samples, while Table S9 compares the whole-database mean with top-k nearest-neighbor similarity metrics.
Next, we assessed the complexity of the repeat units of the generated polymers using the Synthetic Accessibility (SA) score implemented in RDKit. The metric evaluates molecules based on fragment similarity and imposes penalties based on the complexity of well-known fragments [43]. In our approach, we converted the PSMILES representation to molecular SMILES by replacing the dangling bonds ([*]) with hydrogen atoms. Because the RDKit SA score was developed for small molecules, this hydrogen-capped repeat-unit calculation should be interpreted only as an approximate repeat-unit complexity descriptor, not as a direct measure of polymer synthesizability.
Because the generated and PolyInfo comparison sets have different sample sizes and class compositions, Figure 3b is shown as a normalized density distribution rather than raw counts. The generated-polymer distribution contains 511,161 TSCA-derived generated repeat units, while the PolyInfo comparison distribution contains 14,263 unique polymer repeat units. The class composition of the generated set is provided in the Supplementary Information.
The distribution plots of the SA scores for the generated polymers and the PolyInfo comparison subset are depicted in >Figure 3b. The results indicate that the vast majority of the generated polymer repeat units possess an SA score ranging from 2 to 6. In comparison, the PolyInfo comparison subset has SA scores in the range of 2.5 to 7. There is substantial overlap between the SA-score distributions of generated and PolyInfo repeat units, indicating that many generated repeat units have fragment-complexity scores comparable to those in the PolyInfo comparison subset. However, this comparison should not be interpreted as evidence of experimental synthesizability, because the SA score does not account for polymerization feasibility, functional-group tolerance, side reactions, stoichiometric balance, catalyst requirements, achievable molecular weight, purification, or processability. This comparison provides a relative repeat-unit complexity reference for prioritization, but it does not establish experimental synthesizability without additional route-specific feasibility analysis and experimental validation.

3.4. Retrosynthesis of Hypothetical Polymers

The capabilities of RxnChainer also extend to retrosynthetic analysis of polymers, enabling the deconstruction of complex polymer structures into their constituent monomers or reactants. This functionality is particularly valuable for researchers aiming to determine the routes to synthesize existing polymers or to identify plausible precursor candidates to create new polymers. The overview of the retrosynthesis workflow is illustrated in Figure 4a. The process begins with the classification of the input polymer via substructure matching using the RxnChainer’s filtering capabilities. This approach mirrors established retrosynthetic strategies in polymer chemistry, where the polymer backbone and functional groups guide the selection of likely synthetic routes. Once the polymer class is determined, RxnChainer applies a set of inverse reaction templates specific to that class to deconstruct the input polymer into its constituent monomers or reactants. If a polymer class cannot be determined, the workflow does not yield a result for the input polymer. Multiple retrosynthetic pathways may exist for a single polymer, reflecting the diversity of synthetic strategies available in polymer science.
In this study, we demonstrate the retrosynthetic capabilities of RxnChainer for polyesters, providing proof-of-concept examples for the class as illustrated in Figure 4b. The retrosynthetic evaluation reported here is restricted to the polyester class. Extending the analysis to other polymer classes, including polysulfonates, requires class-specific inverse reaction templates and separate validation, and is therefore left to future work. First, the apparent monomer is identified by detecting the repeating unit of the input polymer structure. Next, the polymer class of the apparent monomer is determined through substructure matching in the filtering stage. Once the class is established, RxnChainer applies one or more pre-defined inverse reaction chains specific to the identified polymer class to generate the constituent monomers. For polyesters, three inverse reaction chains, each corresponding to an independent retrosynthetic route, have been implemented: the diacid plus diol route, the diacid chloride plus diol route, and the dimethyl ester plus diol route, shown in Figure 4c, d, e. Each polyester can, in principle, be synthesized via any of these three pathways; the choice of which route to pursue in practice depends on the availability of the corresponding monomers or starting materials. For each pathway, reactant molecules are generated by applying the corresponding inverse reaction chain, enabling comprehensive identification of the monomeric precursors.
To assess the accuracy of the retrosynthetic workflow of RxnChainer, 30,000 polyester polymers were randomly selected from the pool of polymers generated specifically from the TSCA inventory. The workflow was then applied to these polymers to recover their constituent monomers. For validation, the recovered precursor structures were canonicalized and compared with the canonicalized original precursor structures used during forward generation. This comparison evaluates algorithmic recovery under the encoded inverse reaction templates and does not imply supplier availability, neutralization-state validation, or experimental feasibility of the recovered precursors. As shown in Figure 4f, on average, the workflow correctly matched the original monomers for 95% of the selected polyester polymers, with the accuracy varying depending on the specific reaction pathway. Approximately 5% of the polyesters did not lead to the generation of the original monomers primarily because of the presence of multiple ester linkages in the polymers, making the heuristics-based approach for retrosynthesis prone to errors. Regardless, the high accuracy demonstrates the effectiveness of RxnChainer’s retrosynthetic capabilities in identifying plausible monomeric precursors for target polymers. The results also indicate that heuristic inverse-template approaches based on reaction SMARTS have limitations, especially for polymers containing multiple similar disconnection sites or chemically ambiguous linkages.

3.5. Post Polymerization Modification

Post-polymerization modification (PPM) is another powerful approach to modify the properties of polymers. This is particularly useful when desired properties of the polymer are not immediately achievable through the initial polymerization reaction. Many existing polymers, such as ion exchange membranes, are modified to achieve the desired transport properties using PPM [49].
RxnChainer, as a general capability, can also be used to implement in-silico PPM to modify known or generated polymers. In the RxnChainer PPM workflow, illustrated in Figure 5a, we first identify the class and pendant group of the input polymer using substructure search in the filtering stage. Once the pendant group is identified, a set of pre-defined reaction chains is applied to modify the polymer. This is performed by using a set of group-specific reaction templates that are designed to modify the polymer by adding functional groups or changing the polymer architecture, similar to laboratory experiments.
As a proof of concept, we have defined reaction chains for polyolefins or vinyl polymers to modify the polymer with 4 different functional groups. A primary advantage of the reaction chains is that they are designed to be compatible with vinyl polymer backbone and can be applied to any vinyl polymer. Figure 5b shows an example of a vinyl polymer containing nitrile and vinyl pendant groups. The PPM workflow identifies the compatible pendant groups and applies the corresponding reaction chains to modify the polymer. The nitrile group is modified using a cycloaddition reaction to form a tetrazole group. Similarly, the vinyl group is modified using thiol-ene, epoxidation, or bromination reactions. These transformations may alter properties associated with polarity, reactivity, and intermolecular interactions; however, such effects are structure-dependent and are treated here as qualitative hypotheses unless supported by the ML-predicted property changes discussed below. The generated polymer structures after functionalization are shown in Figure 5c. These polymers can be further modified by applying the PPM workflow iteratively.
Figure 5d shows the effects of the post-polymerization modification on the properties of a base vinyl polymer. The results show that the post-polymerization modification significantly alters the bandgap and dielectric constant of the polymer as predicted using a Polymer Genome-based machine learning model, [50] implemented in PolymRize [45]. The modification performed by epoxidation (polymer 5) simultaneously increases the bandgap and dielectric constant of the polymer, which is a desirable for dielectric materials. Similarly, the modification performed by the thiol–ene reaction (polymer 2) significantly increases the predicted dielectric constant of the polymer; however, because the bandgap remains relatively low, this modification would require additional screening before being considered for dielectric applications. These results illustrate that the PPM workflow can generate modified repeat-unit candidates with altered predicted properties. However, the transformations are template-based computational modifications and require experimental validation of conversion, selectivity, reaction conditions, and polymer-substrate compatibility before practical synthesizability can be established.
The workflow is designed to be flexible and can be extended to other polymer classes by defining new reaction templates. In addition, the workflow can be applied iteratively to further modify the properties of a modified polymer. This approach, combined with an optimization algorithm such as Bayesian optimization or a genetic algorithm, could be used to search for candidate modified polymer chemistries for subsequent experimental evaluation.

3.6. Downstream Applications for Polymer Discovery

The library of hypothetical polymers generated in this work using the Rxnchainer methodology enables multiple downstream applications towards polymer discovery. These applications leverage the diversity, estimated repeat-unit complexity, and predicted properties associated with the generated polymer structures to accelerate both computational and experimental research efforts. We highlight a few use cases below.

3.6.1. I. Accessible Property Values of Generated Polymers

We predicted several key properties for the newly generated polymers, including bandgap, glass transition temperature (Tg), and dielectric constant. To do this, we used machine learning models trained on extensive datasets of known polymers with experimentally measured properties. The accuracy of these models has been validated in prior studies for polymers within the chemical domain represented by their training data [25,38,50]. Property predictions for generated polymers that differ substantially from the training distribution constitute extrapolation beyond the applicability domain of the models and carry substantially higher uncertainty. We applied these models to a randomly selected subset of 10 million polymers from our generated library, covering various polymer classes.
As shown in Figure 6, a wide range of property values is covered by different polymer chemistries, and they exhibit substantial variation across the generated polymer classes. Bandgap values range from approximately 0 eV to nearly 8 eV, with the majority of polymers falling within the 2-5 eV range. Glass transition temperatures (Tg) span from below 200 K to above 600 K, with a concentration of values between 300 K and 500 K. Notably, polyimides and polyethers exhibit higher Tg distributions relative to other polymer classes, which is consistent for polyimides as they are a polymer class known for their thermal stability. However, polyethers are an exception, as the ether linkage is flexible, resulting in a lower Tg; this increase is due to stiff substructures such as aromatic and fused rings in the backbone of the generated polyethers, arising from the spacer groups in the reaction-compatible precursor entries. Dielectric constants vary widely from approximately 2 to over 8 at 60 Hz, with polyamides displaying a higher dielectric constant distribution compared to other polymer classes, which can be attributed to their polar chemical structure. These predicted trends are qualitatively consistent with expected behavior for several polymer classes; however, they should be interpreted as screening-level estimates, particularly for generated structures that are far from the training distribution.

3.6.2. II. Chemical Space Exploration and Diversity Analysis

The systematically generated library of hypothetical polymers also enables quantitative analysis of accessible polymer chemical space. We computed Uniform Manifold Approximation and Projections (UMAPs) using Morgan fingerprints (see Figure S1), which compare 10,000 generated polymers randomly selected from the top 16 polymer classes against 14,263 unique polymer repeat units from the PolyInfo comparison subset. The results suggest that while many of the generated polymers overlap with known space, a significant portion represent entirely novel chemistries warranting investigation. The UMAP projections also reveal regions of high structural novelty where property predictions may be less certain, guiding prioritization of experimental characterization to fill data gaps. Conversely, regions of high similarity to known polymers provide opportunities for targeted optimization of established chemistries. This dual capability supports both exploratory research into uncharted chemical spaces and iterative improvement of existing materials.

3.6.3. III. Property-Driven Screening: Example of Dielectrics Design

The availability of 289 million reaction-guided polymer structures with predicted properties enables high-throughput virtual screening for application-specific design. This capability supports the systematic exploration of the structure-property relationships across vast chemical spaces that would be impractical to synthesize and test experimentally. For instance, polymers for high-energy-density capacitors, a critical application for advancing energy storage technologies, must satisfy multiple property constraints simultaneously. High glass transition temperatures ( T g > 150 °C) ensure thermal stability, wide bandgaps ( E g > 4 eV) enable higher breakdown fields, and elevated dielectric constants ( ϵ > 3.5 at 60 Hz) increase energy storage capacity. Since energy density scales as U = 1 2 ϵ 0 ϵ r E 2 , both high dielectric constant and breakdown strength are critical [19].
When applying these screening criteria to our virtual library, numerous candidates passed the property thresholds. The polymers discussed below and listed in Table 2, however, were not drawn from this screened set. As a retrospective case study, we added the corresponding literature precursor molecules for selected recently reported all-organic dielectric polymers to the RxnChainer workflow and reconstructed their repeat-unit structures (Figure 7). Among these rediscovered polymers, polynorbornene-based dielectrics including POFNB, o-POFNB, and PONB-2Me-5Cl exhibit glass transition temperatures ranging from 459 to 517 K and bandgaps between 4.39 and 5.0 eV, with measured energy densities reaching 5.7 to 8.3 J/cc [51,52,53]. Polyimides such as HBPDA-PACM and PI-oxo-iso achieve even higher T g values (527–531 K) and exceptionally wide bandgaps up to 6.32 eV, with energy densities of 2.1 to 5.65 J/cc [19,54]. Polysulfate P6 demonstrates a T g of 551 K, a bandgap of 3.8 eV, a dielectric constant of 3.4 at 473 K, and an energy density of 6.37 J/cc [55]. The polyethersulfone sAI-DG_p1 achieves the highest T g of 573 K with a bandgap of 4.0 eV, dielectric constant of 3.7 at 473 K, and energy density of 6.2 J/cc [56]. This retrospective reconstruction demonstrates that, when supplied with the appropriate literature precursor molecules, RxnChainer can generate repeat-unit structures corresponding to experimentally reported dielectric polymers. This result should not be interpreted as evidence that these polymers were present in the original TSCA/ChEMBL-generated library or as independent validation of the property-prediction models.

3.6.4. IV. Training Data for Generative Models

Beyond screening existing designs, the generated datasets provide training data for polymer-specific generative models. Large-scale polymer datasets enable pretraining of chemical language models such as polyBERT, [38] polyT5 [40], and polyBART [39], which learn representations of polymer structures and chemical space from SMILES strings. The diversity of polymer classes (30 classes spanning addition, condensation, and ring-opening polymerizations) and structural variety within each class provides comprehensive coverage of polymer chemistry grammar. These pretraining models can subsequently be fine-tuned for specific tasks including property prediction, inverse design, and reaction outcome prediction. The systematic generation approach ensures that training data encompasses both common polymer architectures and novel structural motifs, improving model generalization.

4. Conclusions

RxnChainer provides a unified computational methodology integrating polymer generation, retrosynthesis, and post-polymerization modification for systematically exploring reaction-guided polymer space. Using molecule entries from the TSCA and ChEMBL databases, we generated over 289 million hypothetical polymers across 44 reaction chains spanning 30 polymer classes. Comparison with the PolyInfo comparison subset indicates that many sampled generated polymers are structurally distinct under the molecular representations evaluated here, suggesting that the workflow can expand the set of candidate chemistries available for further evaluation. Retrosynthetic analysis of 30,000 generated polyesters achieved 95% recovery of the original precursor molecules in an internal consistency test of the polyester inverse-template workflow. Post-polymerization modification pathways targeting compatible vinyl and nitrile pendant groups were demonstrated using tetrazole-forming cycloaddition, thiol–ene, epoxidation, and bromination templates, providing candidate modified repeat units for subsequent property screening and experimental evaluation. The generated library enables multiple downstream applications. Machine learning property predictions reveal wide accessible ranges—bandgaps from 0 to 8 eV, glass transition temperatures from 200 to 600 K, and dielectric constants from 2 to 8—supporting high-throughput virtual screening. In a retrospective case study, RxnChainer reconstructed repeat-unit polymers corresponding to experimentally reported high-performance dielectric polymers after the associated literature precursor molecules were supplied to the workflow. The library also provides comprehensive training data for polymer-specific language models. Prospective experimental validation of the generated candidates, including synthesis of selected structures, evaluation against supplier-verified precursor sets, and extension of the retrosynthetic analysis to additional polymer classes, remains an important direction for future work. These results demonstrate RxnChainer’s potential to connect computational polymer design with subsequent experimental evaluation.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org. List of polymerization reaction templates implemented in RxnChainer, and additional figures and tables (PDF).

Author Contributions

Conceptualization: R.R., S.S., A.M., Data generation and analysis: S.S., A.M., Visualization: S.S., A.M., C.K., Original draft: S.S., A.M., Review & editing: S.S., A.M., C.K., R.G., R.R.

Data Availability Statement

The datasets used to generate the polymers are available from the official TSCA inventory at https://www.epa.gov/tsca-inventory/ and the ChEMBL database at https://www.ebi.ac.uk/chembl/. The known polymer datasets used in this study were derived from the publicly accessible PolyInfo database (https://polymer.nims.go.jp). To support non-commercial evaluation and reproducibility of the analyses reported in this manuscript, a representative subset of 10 million generated polymer structures used in the property-distribution analyses has been released at https://doi.org/10.5281/zenodo.21878520. Polymer structure representations, polymer class labels, reaction routes, and source database identifiers, where available, are included in the released dataset.

Conflicts of Interest

RxnChainer has been implemented in PolymRizeTM, a standardized polymer informatics platform by Matmerize, Inc. The authors are employees, affiliates or founders of Matmerize, Inc., which owns intellectual property rights related to the PolymRizeTM platform, including proprietary technology and, potentially, further patent rights.

References

  1. Siracusa, V.; Rocculi, P.; Romani, S.; Rosa, M.D. Biodegradable polymers for food packaging: A review. Trends Food Sci. Technol. 2008, 19, 634–643. [Google Scholar] [CrossRef]
  2. Cao, Y.; Uhrich, K.E. Biodegradable and biocompatible polymers for electronic applications: A review. J. Bioact. Compat. Polym. 2019, 34, 3–15. [Google Scholar] [CrossRef]
  3. Ulery, B.D.; Nair, L.S.; Laurencin, C.T. Biomedical applications of biodegradable polymers. J. Polym. Sci. Part B Polym. Phys. 2011, 49, 832–864. [Google Scholar] [CrossRef] [PubMed]
  4. Wright, W. Polymers in aerospace applications. Mater. Des. 1991, 12, 222–227. [Google Scholar] [CrossRef]
  5. Ghori, S.W.; Siakeng, R.; Rasheed, M.; Saba, N.; Jawaid, M. 2 - The role of advanced polymer materials in aerospace. In Sustainable Composites for Aerospace Applications; Jawaid, M., Thariq, M., Eds.; Woodhead Publishing Series in Composites Science and Engineering; Woodhead Publishing, 2018; pp. 19–34. [Google Scholar] [CrossRef]
  6. Rathod, V.T.; Kumar, J.S.; Jain, A. Polymer and Ceramic Nanocomposites for Aerospace Applications. Appl. Nanosci. 2017, 7, 519–548. [Google Scholar] [CrossRef]
  7. González-Campos, J.B.; Luna-Bárcenas, G.; Zárate-Triviño, D.G.; Mendoza-Galván, A.; Prokhorov, E.; Villaseñor-Ortega, F.; Sanchez, I.C. Polymer States and Properties. In Handbook of Polymer Synthesis, Characterization, and Processing; John Wiley & Sons, Ltd, 2013; Volume 2, pp. 15–39. http://arxiv.org/abs/https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781118480793.ch2.
  8. Pfaendner, R. Polymer Additives. In Handbook of Polymer Synthesis, Characterization, and Processing; John Wiley & Sons, Ltd, 2013; Volume 11, pp. 225–247. http://arxiv.org/abs/https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781118480793.ch11.
  9. Robledo-Ortíz, J.R.; Fuentes-Talavera, F.J.; González-Núñez, R.; Silva-Guzmán, J.A. Wood and Natural Fiber-Based Composites (NFCs). In Handbook of Polymer Synthesis, Characterization, and Processing; John Wiley & Sons, Ltd, 2013; Volume 26, pp. 493–503. http://arxiv.org/abs/https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781118480793.ch26.
  10. Sánchez-Valdes, S.; Ramos-De Valle, L.F.; Manero, O. Polymer Blends. In Handbook of Polymer Synthesis, Characterization, and Processing; John Wiley & Sons, Ltd, 2013; Volume 27, pp. 505–517. https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781118480793.ch27. [CrossRef]
  11. Saldívar-Guerra, E.; Vivaldo-Lima, E. Introduction to Polymers and Polymer Types. In Handbook of Polymer Synthesis, Characterization, and Processing; John Wiley & Sons, Ltd, 2013; Volume 1, pp. 1–14. https://onlinelibrary.wiley.com/doi/pdf/10.1002/9781118480793.ch1. [CrossRef]
  12. Pillai, C.K.S. Challenges for Natural Monomers and Polymers: Novel Design Strategies and Engineering to Develop Advanced Polymers. Des. Monomers Polym. 2010, 13, 87–121. [Google Scholar] [CrossRef]
  13. Ma, R.; Baldwin, A.F.; Wang, C.; Offenbach, I.; Cakmak, M.; Ramprasad, R.; Sotzing, G.A. Rationally Designed Polyimides for High-Energy Density Capacitor Applications. ACS Appl. Mater. Interfaces 2014, 6, 10445–10451. [Google Scholar] [CrossRef] [PubMed]
  14. Fazio, A.; Caroleo, M.C.; Cione, E.; Plastina, P. Novel Acrylic Polymers for Food Packaging: Synthesis and Antioxidant Properties. Food Packag. Shelf Life 2017, 11, 84–90. [Google Scholar] [CrossRef]
  15. Ilinitch, O.; Semin, G.; Chertova, M.; Zamaraev, K. Novel Polymeric Membranes for Separation of Hydrocarbons. J. Membr. Sci. 1992, 66, 1–8. [Google Scholar] [CrossRef]
  16. Kerres, J.; Ullrich, A.; Meier, F.; Häring, T. Synthesis and Characterization of Novel Acid–Base Polymer Blends for Application in Membrane Fuel Cells. Solid State Ion. 1999, 125, 243–249. [Google Scholar] [CrossRef]
  17. Wang, Y.; Ding, Y.; Yu, K.; Dong, G. Innovative polymer-based composite materials in additive manufacturing: A review of methods, materials, and applications. Polym. Compos. 2024, 45, 15389–15420. https://4spepublications.onlinelibrary.wiley.com/doi/pdf/10.1002/pc.28854. [CrossRef]
  18. Wang, Y.; Zhou, X.; Chen, Q.; Chu, B.; Zhang, Q. Recent development of high energy density polymers for dielectric capacitors. IEEE Trans. Dielectr. Electr. Insul. 2010, 17, 1036–1042. [Google Scholar] [CrossRef]
  19. Gurnani, R.; Shukla, S.; Kamal, D.; Wu, C.; Hao, J.; Kuenneth, C.; Aklujkar, P.; Khomane, A.; Daniels, R.; Deshmukh, A.A.; et al. AI-assisted Discovery of High-Temperature Dielectrics for Energy Storage. Nat. Commun. 2024, 15, 6107. [Google Scholar] [CrossRef] [PubMed]
  20. Dangayach, R.; Jeong, N.; Demirel, E.; Uzal, N.; Fung, V.; Chen, Y. Machine Learning-Aided Inverse Design and Discovery of Novel Polymeric Materials for Membrane Separation. Environ. Sci. Technol. 2025, 59, 993–1012. [Google Scholar] [CrossRef] [PubMed]
  21. Ramprasad, R.; Batra, R.; Pilania, G.; Mannodi-Kanakkithodi, A.; Kim, C. Machine Learning in Materials Informatics: Recent Applications and Prospects. npj Comput. Mater. 2017, 3, 54. [Google Scholar] [CrossRef]
  22. Gaulton, A.; Bellis, L.J.; Bento, A.P.; Chambers, J.; Davies, M.; Hersey, A.; Light, Y.; McGlinchey, S.; Michalovich, D.; Al-Lazikani, B.; et al. ChEMBL: A large-scale bioactivity database for drug discovery. Nucleic Acids Res. 2011, 40, D1100–D1107. https://academic.oup.com/nar/article-pdf/40/D1/D1100/16955876/gkr777.pdf. [CrossRef] [PubMed]
  23. Sterling, T.; Irwin, J.J. ZINC15 – Ligand discovery for everyone. J. Chem. Inf. Model. 2015, 55, 2324–2337. [Google Scholar] [CrossRef] [PubMed]
  24. eMolecules Database. 2025. https://www.emolecules.com (accessed on 2025-08-28).
  25. Kim, S.; Schroeder, C.M.; Jackson, N.E. Open Macromolecular Genome: Generative Design of Synthetically Accessible Polymers. ACS Polym. Au 2023, 3, 318–330. [Google Scholar] [CrossRef] [PubMed]
  26. Ohno, M.; Hayashi, Y.; Zhang, Q.; Kaneko, Y.; Yoshida, R. SMiPoly: Generation of a Synthesizable Polymer Virtual Library Using Rule-Based Polymerization Reactions. J. Chem. Inf. Model. 2023, 63, 5539–5548. [Google Scholar] [CrossRef] [PubMed]
  27. Yue, T.; He, J.; Li, Y. Polyuniverse: Generation of a Large-Scale Polymer Library Using Rule-Based Polymerization Reactions for Polymer Informatics. Digit. Discov. 2024, 3, 2465–2478. [Google Scholar] [CrossRef]
  28. Kern, J.; Su, Y.L.; Gutekunst, W.R.; Ramprasad, R. An Informatics Framework for the Design of Sustainable, Chemically Recyclable, Synthetically Accessible, and Durable Polymers. npj Comput. Mater. 2025, 11, 182. [Google Scholar] [CrossRef]
  29. Phan, B.K.; Kim, C.; Nistane, J.; Xiong, W.; Chen, H.; Jang, W.J.; Gholami, F.; Su, Y.; Qi, J.; Lively, R.; et al. AI-assisted design of chemically recyclable polymers for food packaging. arXiv 2025, arXiv:2511.04704. [Google Scholar] [CrossRef]
  30. Chen, L.; Kern, J.; Lightstone, J.P.; Ramprasad, R. Data-Assisted Polymer Retrosynthesis Planning. Appl. Phys. Rev. 2021, 8, 031405. http://arxiv.org/abs/https://pubs.aip.org/aip/apr/article-pdf/doi/10.1063/5.0052962/14581497/031405\_1\_online.pdf. [CrossRef]
  31. Gauthier, M.; Gibson, M.; Klok, H.A. Synthesis of Functional Polymers by Post-Polymerization Modification. Angew. Chem. Int. Ed. 2009, 48, 48–58. http://arxiv.org/abs/https://onlinelibrary.wiley.com/doi/pdf/10.1002/anie.200801951. [CrossRef]
  32. Urban, V.M.; Machado, A.L.; Vergani, C.E.; Giampaolo, E.T.; Pavarina, A.C.; de Almeida, F.G.; Cass, Q.B. Effect of Water-Bath Post-Polymerization on the Mechanical Properties, Degree of Conversion, and Leaching of Residual Compounds of Hard Chairside Reline Resins. Dent. Mater. 2009, 25, 662–671. [Google Scholar] [CrossRef] [PubMed]
  33. Farmer, T.J.; Comerford, J.W.; Pellis, A.; Robert, T. Post-polymerization modification of bio-based polymers: Maximizing the high functionality of polymers derived from biomass. Polym. Int. 2018, 67, 775–789. https://scijournals.onlinelibrary.wiley.com/doi/pdf/10.1002/pi.5573. [CrossRef]
  34. Lv, A.; Cui, Y.; Du, F.S.; Li, Z.C. Thermally Degradable Polyesters with Tunable Degradation Temperatures via Postpolymerization Modification and Intramolecular Cyclization. Macromolecules 2016, 49, 8449–8458. [Google Scholar] [CrossRef]
  35. Agar, S.; Baysak, E.; Hizal, G.; Tunca, U.; Durmaz, H. An emerging post-polymerization modification technique: The promise of thiol-para-fluoro click reaction. J. Polym. Sci. Part A Polym. Chem. 2018, 56, 1181–1198. https://onlinelibrary.wiley.com/doi/pdf/10.1002/pola.29004. [CrossRef]
  36. U.S. Environmental Protection Agency. TSCA Chemical Substance Inventory. 2025. https://www.epa.gov/tsca-inventory (accessed on 2025-08-28).
  37. Otsuka, S.; Kuwajima, I.; Hosoya, J.; Xu, Y.; Yamazaki, M. PoLyInfo: Polymer Database for Polymeric Materials Design. In Proceedings of the 2011 International Conference on Emerging Intelligent Data and Web Technologies, 2011; pp. 22–29. [Google Scholar] [CrossRef]
  38. Kuenneth, C.; Ramprasad, R. polyBERT: A Chemical Language Model to Enable Fully Machine-Driven Ultrafast Polymer Informatics. 14, 4099. [CrossRef] [PubMed]
  39. Savit, A.; Sahu, H.; Shukla, S.; Xiong, W.; Ramprasad, R. polyBART: A Chemical Linguist for Polymer Property Prediction and Generative Design. arXiv 2025, arXiv:2506.04233. [Google Scholar] [CrossRef]
  40. Sahu, H.; Xiong, W.; Savit, A.; Shukla, S.S.; Ramprasad, R. An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design. arXiv 2025, arXiv:2510.18860. [Google Scholar] [CrossRef]
  41. Daylight Chemical Information Systems, I.;Weininger, D. SMILES, a chemical language and information system. 3. SMARTS: A language for describing molecular patterns. Journal of Chemical Information and Computer Sciences 1989, 29, 97–101. [CrossRef]
  42. Weininger, D. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 1988, 28, 31–36. [Google Scholar] [CrossRef]
  43. Ertl, P.; Schuffenhauer, A. Estimation of Synthetic Accessibility Score of Drug-like Molecules Based on Molecular Complexity and Fragment Contributions. J. Cheminformatics 2009, 1, 8. [Google Scholar] [CrossRef] [PubMed]
  44. Willett, P.; Barnard, J.M.; Downs, G.M. Chemical Similarity Searching. J. Chem. Inf. Comput. Sci. 1998, 38, 983–996. [Google Scholar] [CrossRef]
  45. Matmerize, I. PolymRize 0.24.0 (July 2025) Release, 2025.
  46. Landrum, G. RDKit: Open-source cheminformatics. 2006. http://www.rdkit.org (accessed on 2025-08-28).
  47. Gasteiger, J.; Marsili, M. Iterative partial equalization of orbital electronegativity—a rapid access to atomic charges. Tetrahedron 1980, 36, 3219–3228. [Google Scholar] [CrossRef]
  48. Rogers, D.; Hahn, M. Extended-Connectivity Fingerprints. J. Chem. Inf. Model. 2010, 50, 742–754. [Google Scholar] [CrossRef] [PubMed]
  49. Chen, J.; Yan, J.; Yang, W.; Xu, Y.; Chi, R.; Zhang, Q.; Yan, Y. Cobaltocenium-Containing Poly(Carbazole)s towards Alkaline-Stable Anion Exchange Membranes via Post-Polymerization Modification. Sustain. Energy Fuels 2024, 8, 4767–4771. [Google Scholar] [CrossRef]
  50. Kim, C.; Chandrasekaran, A.; Huan, T.D.; Das, D.; Ramprasad, R. Polymer Genome: A Data-Powered Polymer Informatics Platform for Property Predictions. J. Phys. Chem. C 2018, 122, 17575–17585. [Google Scholar] [CrossRef]
  51. Wu, C.; Deshmukh, A.A.; Li, Z.; Chen, L.; Alamri, A.; Wang, Y.; Ramprasad, R.; Sotzing, G.A.; Cao, Y. Flexible Temperature-Invariant Polymer Dielectrics with Large Bandgap. Adv. Mater. 2020, 32, 2000499. https://advanced.onlinelibrary.wiley.com/doi/pdf/10.1002/adma.202000499. [CrossRef]
  52. Deshmukh, A.A.; Wu, C.; Yassin, O.; Mishra, A.; Chen, L.; Alamri, A.; Li, Z.; Zhou, J.; Mutlu, Z.; Sotzing, M.; et al. Flexible polyolefin dielectric by strategic design of organic modules for harsh condition electrification. Energy Environ. Sci. 2022, 15, 1307–1314. [Google Scholar] [CrossRef]
  53. Wang, R.; Zhu, Y.; Fu, J.; Yang, M.; Ran, Z.; Li, J.; Li, M.; Hu, J.; He, J.; Li, Q. Designing Tailored Combinations of Structural Units in Polymer Dielectrics for High-Temperature Capacitive Energy Storage. 14, 2406. [CrossRef] [PubMed]
  54. Huang, W.; Wan, B.; Yang, X.; Cheng, M.; Zhang, Y.; Li, Y.; Wu, C.; Dang, Z.M.; Zha, J.W. Alicyclic Polyimide With Multiple Breakdown Self-Healing Based on Gas-Condensation Phase Validation for High Temperature Capacitive Energy Storage. Adv. Mater. 2024, 36, 2410927. https://advanced.onlinelibrary.wiley.com/doi/pdf/10.1002/adma.202410927. [CrossRef]
  55. Li, H.; Zheng, H.; Yue, T.; Xie, Z.; Yu, S.; Zhou, J.; Kapri, T.; Wang, Y.; Cao, Z.; Zhao, H.; et al. Machine Learning-Accelerated Discovery of Heat-Resistant Polysulfates for Electrostatic Energy Storage. 10, 90–100. [CrossRef]
  56. Ren, W.; Tong, H.; Cao, S.; Zhao, S.; Yang, M.; Li, X.; Pan, J.; Sun, N.; Xiao, Y.; Xu, E.; et al. Semi-Alicyclic Dipolar Glass Dielectric Polymer Capacitors for Superior High-Temperature Capacitive Energy Storage. Adv. Mater. 2025, 37, e05296. https://advanced.onlinelibrary.wiley.com/doi/pdf/10.1002/adma.202505296. [CrossRef]
Figure 1. (a) Overview of the RxnChainer methodology, which processes a pool of input reactants by filtering and producing the output products based on a reaction chain containing multiple reaction steps. (b) Flow diagrams showing different stages for applications of RxnChainer, including the generation of hypothetical polymers, the retrosynthesis of hypothetical or known polymers, and the post-polymerization modifications of a base polymer.
Figure 1. (a) Overview of the RxnChainer methodology, which processes a pool of input reactants by filtering and producing the output products based on a reaction chain containing multiple reaction steps. (b) Flow diagrams showing different stages for applications of RxnChainer, including the generation of hypothetical polymers, the retrosynthesis of hypothetical or known polymers, and the post-polymerization modifications of a base polymer.
Preprints 228798 g001
Figure 2. (a) Overview of generated polymer repeat units using RxnChainer. We considered 70,000 molecules from the TSCA inventory and 1.2 million molecules from the ChEMBL database. After filtering for reaction-compatible precursor molecules, 14,470 molecules from TSCA and 293,875 molecules from ChEMBL were retained. These precursor molecules were used to generate 289 million hypothetical polymer repeat units spanning 30 polymer classes through 44 reaction chains. (b) Class-level comparison of generated polymer counts with previous datasets. Because the compared datasets differ in source molecules, reaction coverage, deduplication rules, and polymer-class definitions, this comparison should be interpreted as a qualitative coverage comparison rather than a normalized benchmark.
Figure 2. (a) Overview of generated polymer repeat units using RxnChainer. We considered 70,000 molecules from the TSCA inventory and 1.2 million molecules from the ChEMBL database. After filtering for reaction-compatible precursor molecules, 14,470 molecules from TSCA and 293,875 molecules from ChEMBL were retained. These precursor molecules were used to generate 289 million hypothetical polymer repeat units spanning 30 polymer classes through 44 reaction chains. (b) Class-level comparison of generated polymer counts with previous datasets. Because the compared datasets differ in source molecules, reaction coverage, deduplication rules, and polymer-class definitions, this comparison should be interpreted as a qualitative coverage comparison rather than a normalized benchmark.
Preprints 228798 g002
Figure 3. (a) Morgan-fingerprint Tanimoto similarity scores between sampled generated polymer repeat units and 14,263 unique polymer repeat units from the PolyInfo comparison subset, grouped by the top 16 generated polymer classes. The low average similarity scores indicate that the sampled generated structures are, on average, structurally distinct from the PolyInfo comparison subset under this representation. A Tanimoto similarity of 1.0 indicates identical Morgan fingerprints under the representation used here; it should not be interpreted as an exact canonical-SMILES match. (b) Normalized SA-score density distributions for 511,161 TSCA-derived generated polymer repeat units and 14,263 unique polymer repeat units from the PolyInfo comparison subset. The class composition of the generated set is reported in the Supplementary Information. The SA score is used here as an approximate repeat-unit complexity descriptor and should not be interpreted as direct evidence of experimental synthesizability.
Figure 3. (a) Morgan-fingerprint Tanimoto similarity scores between sampled generated polymer repeat units and 14,263 unique polymer repeat units from the PolyInfo comparison subset, grouped by the top 16 generated polymer classes. The low average similarity scores indicate that the sampled generated structures are, on average, structurally distinct from the PolyInfo comparison subset under this representation. A Tanimoto similarity of 1.0 indicates identical Morgan fingerprints under the representation used here; it should not be interpreted as an exact canonical-SMILES match. (b) Normalized SA-score density distributions for 511,161 TSCA-derived generated polymer repeat units and 14,263 unique polymer repeat units from the PolyInfo comparison subset. The class composition of the generated set is reported in the Supplementary Information. The SA score is used here as an approximate repeat-unit complexity descriptor and should not be interpreted as direct evidence of experimental synthesizability.
Preprints 228798 g003
Figure 4. (a) Schematic overview of the retrosynthesis pipeline used to identify plausible constituent reactants from a target polymer. The structures in this overview panel are schematic placeholders and are not intended to represent actual polymerizable monomers. The polymer class is first identified by the RxnChainer, and the monomers are identified by inverse reaction chains. (b) Identification of the apparent monomer and polymer class. (c, d, e) Determination of the constituent monomers identified by different inverse reaction pathways for the identified polymer class. (f) Accuracy of the retrosynthesis pipeline in identifying the monomers for 30,000 polyester polymers generated from known monomers. Panels (b–e) show representative retrosynthetic disconnections and recovered precursor structures from the inverse-template workflow. Recovered precursor structures were canonicalized before comparison with the original forward-generation precursors. The displayed structures are computational representations used to assess algorithmic recovery and should not be interpreted as supplier-validated or experimentally verified monomers.
Figure 4. (a) Schematic overview of the retrosynthesis pipeline used to identify plausible constituent reactants from a target polymer. The structures in this overview panel are schematic placeholders and are not intended to represent actual polymerizable monomers. The polymer class is first identified by the RxnChainer, and the monomers are identified by inverse reaction chains. (b) Identification of the apparent monomer and polymer class. (c, d, e) Determination of the constituent monomers identified by different inverse reaction pathways for the identified polymer class. (f) Accuracy of the retrosynthesis pipeline in identifying the monomers for 30,000 polyester polymers generated from known monomers. Panels (b–e) show representative retrosynthetic disconnections and recovered precursor structures from the inverse-template workflow. Recovered precursor structures were canonicalized before comparison with the original forward-generation precursors. The displayed structures are computational representations used to assess algorithmic recovery and should not be interpreted as supplier-validated or experimentally verified monomers.
Preprints 228798 g004
Figure 5. (a) Workflow to perform post-polymerization modification (PPM). First, the pendant groups of the base polymers are identified, then compatible functionalization reaction chains are applied to generate the modified polymers. The generated polymer can be further functionalized iteratively. (b) Example functionalizations performed on a base polymer containing nitrile and vinyl pendant groups. (c) Generated polymers after functionalizations via different groups. (d) Predicted bandgap and dielectric constant of the modified polymers compared to the base polymer. Figure 5 is intended as a schematic representation of idealized post-polymerization modification templates. Reagents, reaction conditions, substrate-specific selectivity, stereochemical outcomes, and regioisomeric outcomes are not assigned unless explicitly encoded in the corresponding reaction template. The transformations assume full modification of the targeted pendant group unless otherwise specified; incomplete conversion, side reactions, and polymer-substrate compatibility are not modeled.
Figure 5. (a) Workflow to perform post-polymerization modification (PPM). First, the pendant groups of the base polymers are identified, then compatible functionalization reaction chains are applied to generate the modified polymers. The generated polymer can be further functionalized iteratively. (b) Example functionalizations performed on a base polymer containing nitrile and vinyl pendant groups. (c) Generated polymers after functionalizations via different groups. (d) Predicted bandgap and dielectric constant of the modified polymers compared to the base polymer. Figure 5 is intended as a schematic representation of idealized post-polymerization modification templates. Reagents, reaction conditions, substrate-specific selectivity, stereochemical outcomes, and regioisomeric outcomes are not assigned unless explicitly encoded in the corresponding reaction template. The transformations assume full modification of the targeted pendant group unless otherwise specified; incomplete conversion, side reactions, and polymer-substrate compatibility are not modeled.
Preprints 228798 g005
Figure 6. (a) Bandgap, (b) Glass Transition Temperature (Tg), and (c) Dielectric Constant distributions of randomly selected 10 million generated polymers predicted using the PolymRizeTM ML models. The histograms show the number of polymers (y-axis, log scale) versus the predicted property values (x-axis). Different colors represent different polymer classes grouped by a block of 6. Dielectric constant is shown for frequency 60 Hz. Vertical dashed bars represent the validity cutoff for predictions outside physically meaningful ranges.
Figure 6. (a) Bandgap, (b) Glass Transition Temperature (Tg), and (c) Dielectric Constant distributions of randomly selected 10 million generated polymers predicted using the PolymRizeTM ML models. The histograms show the number of polymers (y-axis, log scale) versus the predicted property values (x-axis). Different colors represent different polymer classes grouped by a block of 6. Dielectric constant is shown for frequency 60 Hz. Vertical dashed bars represent the validity cutoff for predictions outside physically meaningful ranges.
Preprints 228798 g006
Figure 7. Structures of recently synthesized all-organic homopolymer dielectrics rediscovered in the RxnChainer-generated library, spanning polynorbornenes, polyimides, polysulfates, and polyethersulfones. Corresponding ML-predicted and experimentally measured properties are listed in Table 2.
Figure 7. Structures of recently synthesized all-organic homopolymer dielectrics rediscovered in the RxnChainer-generated library, spanning polynorbornenes, polyimides, polysulfates, and polyethersulfones. Corresponding ML-predicted and experimentally measured properties are listed in Table 2.
Preprints 228798 g007
Table 1. Number of hypothetical polymer repeat units generated using RxnChainer, grouped by reaction type, polymer class, reaction route, and precursor molecules from the TSCA inventory and ChEMBL database.The polymer classes are grouped by their reaction types (polyaddition, polycondensation, and ring-opening). The reaction routes describe the precursor/reactant classes used within each reaction chain. The number of corresponding precursor/reactant molecules for each class is shown in parentheses. A dash indicates that no successfully generated products were obtained for that route from the corresponding database under the implemented filtering and reaction rules.
Table 1. Number of hypothetical polymer repeat units generated using RxnChainer, grouped by reaction type, polymer class, reaction route, and precursor molecules from the TSCA inventory and ChEMBL database.The polymer classes are grouped by their reaction types (polyaddition, polycondensation, and ring-opening). The reaction routes describe the precursor/reactant classes used within each reaction chain. The number of corresponding precursor/reactant molecules for each class is shown in parentheses. A dash indicates that no successfully generated products were obtained for that route from the corresponding database under the implemented filtering and reaction rules.
Polymer Class Reaction Route TSCA ChEMBL
Reaction Type: Addition
Polyacetylene Acetylene 109 (109) 4,796 (4,796)
Polyaldehyde Aldehyde 819 (819) 7,117 (7,117)
Polyamic acid Dianhydride + Diamine 4,913 (17+289) 10,971 (3+3,657)
Polyamide Heterocumulene 74 (74) 10 (10)
Polydiene Diene 22 (22) 26 (26)
Polyethersulfone Vinylsulfonylethanol - 18 (18)
Polyhemiacetal Cyclic Acetal 19 (19) 267 (267)
Polymaleimide Maleimide 16 (16) 142 (142)
Polyoxazolidone Diepoxide + Diisocyanate 1,300 (25+52) 58 (29+2)
Polysuccinimide Dimaleimide + Diamine 972 (9+108) 4,748 (2+2,374)
Polyurea Diamine + Diisocyanate 9,176 (37+248) 5,012 (2+2,506)
Polyurethane Diol + Diisocyanate 26,159 (37+707) 15,730 (2+7,865)
Polyolefin Vinyl Monomer 924 (924) 13,004 (13,004)
Reaction Type: Condensation
Polyamide Diacid + Diamine 56,165 (239+235) 3,661,596 (2,316+1,581)
Polyamide Diacid Chloride + Diamine 8,526 (29+294) -
Polyamide Dimethyl Ester + Diamine 27,636 (94+294) 4,184,510 (1,355+1,581)
Polyester Diacid + Diol 112,540 (170+662) 4,924,920 (840+5,863)
Polyester Diacid Chloride + Diol 18,536 (28+662) -
Polyester Dimethyl Ester + Diol 38,340 (60+639) 2,157,805 (373+5,785)
Polyether Dichloride + Diol 112,200 (200+561) 19,097,736 (4,123+4,632)
Polyether Dibromide + Diol 36,465 (65+561) 2,427,168 (524+4,632)
Polyether Diiodide + Diol 7,293 (13+561) 226,968 (49+4,632)
Polyimide Dianhydride + Diamine 4,913 (17+289) 10,971 (3+3,657)
Polynorcantharimide Difuran + Dimaleimide 108 (12+9) 618 (309+2)
Polyoxadiazole Diacid Chloride + Dihydrazide 135 (27+5) -
Polyoxadiazole Diacid + Dihydrazine 230 (230+1) 2,125 (2,125+1)
Polyoxime Dialdehyde + Diaminoxy - 294 (147+2)
Polypyrazoline Tetrazole-Alkene - 5 (5)
Polysulfates Diacid 888 (888) 11,294 (11,294)
Polysulfonate Bis(sulfonyl fluoride) (from amine precursor) + Diol 1,185,954 (2,114+561) 252,564,432 (54,526+4,632)
Polythioether Diene + Dithiol 4,991 (161+31) 61,920 (860+72)
Polythioether Diyne + Dithiol 403 (13+31) 12,456 (173+72)
Polythioether Dibromo + Dithiol 2,449 (79+31) 141,636 (1,914+74)
Polythiourethane Diisocyanate + Dithiol 1,581 (51+1) 148 (2+74)
Polytriazole Diazide + Dialkyne 130 (10+13) -
Reaction Type: Ring Opening Polymerization (ROP)
Polyamide Lactam 178 (178) 56,293 (56,293)
Polycarbonate Cyclic Carbonate 15 (15) 133 (133)
Polycycloalkene Cyclic Alkene - 12,502 (12,502)
Polyester Lactone 451 (451) 40,131 (40,131)
Polyether Cyclic Ether 558 (558) 14,280 (14,280)
Polythiocane Cyclic Thiocane - 9 (9)
Polythioester Cyclic Thioester 6 (6) 3,999 (3,999)
Polythioether Cyclic Thioether 9 (9) 4,291 (4,291)
Polythionoester Cyclic Thionoester 67 (67) 2,401 (2,401)
Total 1,665,253 287,640,070
Table 2. Retrospective reconstruction of recently synthesized all-organic dielectric polymers using RxnChainer after adding the corresponding literature precursor molecules to the workflow. For each property, the first line shows the ML-predicted screening-level value and the second line (in italics) shows the experimentally measured literature value. Dielectric constants predicted at 60 Hz and experimental dielectric constants measured at 1 kHz and/or elevated temperature are not directly comparable; measurement conditions are noted in parentheses where available.
Table 2. Retrospective reconstruction of recently synthesized all-organic dielectric polymers using RxnChainer after adding the corresponding literature precursor molecules to the workflow. For each property, the first line shows the ML-predicted screening-level value and the second line (in italics) shows the experimentally measured literature value. Dielectric constants predicted at 60 Hz and experimental dielectric constants measured at 1 kHz and/or elevated temperature are not directly comparable; measurement conditions are noted in parentheses where available.
Polymer Class T g
(K)
Bandgap
(eV)
ϵ Ref.
POFNB Polynorbornene
(Polyalkene)
453.5
459
4.433
4.9
2.543
2.5a
[51]
o-POFNB Polynorbornene
(Polyalkene)
503.9
517
4.402
5.0
2.892
2.8b
[52]
PONB-2Me-5Cl Polynorbornene
(Polyalkene)
506.7
505
4.659
4.39
2.909
3.0b
[53]
HBPDA-PACM Alicyclic
polyimide
496.9
527
5.616
6.32
2.764
2.8a
[54]
PI-oxo-iso Polyimide 511.5
531
3.308
4.3
3.395
[19]
Polysulfate P6 Polysulfate 538.0
551
3.462
3.8
2.913
3.4b
[55]
sAI-DG_p1 Polyethersulfone 493.7
573
3.893
4.0
2.678
3.7b
[56]
a Measured at 423 K, 1 kHz. b Measured at 473 K, 1 kHz.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.