Preprint
Technical Note

This version is not peer-reviewed.

Evaluating Multi-Objective Genetic Search for Healthcare Data Anonymization Using Synthetic Records

Submitted:

24 September 2026

Posted:

28 September 2026

You are already at the latest version

Abstract
Anonymizing sensitive healthcare data while preserving data utility is always a tradeoff between suppression and generalization. In this experiment, we employ a genetic algorithm to search the anonymization policy space using synthetic healthcare-style data. Each policy candidate specifies different levels of generalization for quasi-identifiers such as age, a five-digit numeric location code, and sex, while any equivalence class that does not satisfy the constraints of k-anonymity or distinct ℓ-diversity is suppressed. The GA experiment is carried out under constraints of k = 5, ℓ = 2, and a 30% suppression limit. There are 32 possible anonymization policies, and the search space is intentionally kept small so that the results of the GA-based search can be compared directly with exhaustive search to identify its accuracy, gaps, and limitations. Across 60 seeded runs using three different table sizes, the GA search recovered the exact Pareto set in 59 runs while evaluating a median of 28–29 policies. These results show that the GA was capable of reliably recovering the Pareto-optimal solutions in this small synthetic setting, but GA still ended up evaluating most of the policies. The efficiency advantage of GA has yet to be evaluated in the future on large real world healthcare data with substantially larger policy spaces to establish if GA indeed offers more clinical utility with less compute compared to other traditional methods.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

Generalization and suppression are the primary tools used on real word data to create equivalence-classes of any size. One generalization benchmark is k-anonymity parameter, that means a combination of quasi-identifiers occurs in at least k records [1]. ℓ-diversity applies to sensitive attributes of class, means the equivalence class sensitive attributes should have l distinct values. [2]. Neither criterion by itself is a general guarantee against re-identification or attribute disclosure.
The intended goal of this research is to establish whether multi-objective genetic search can find anonymization policies that meet privacy constraints while preserving information for specific healthcare analyses. Before tackling that question, we need an implementation whose behavior can be checked against a known answer. This note provides such a pilot experiment. It leverages a concrete policy encoding and class-suppression rule to perform a reproducible comparison between a genetic search and exhaustive enumeration in a 32-policy space. The search builds on a nondominated sorting genetic algorithm techinique described here [3] and utilizes healthcare constraints to expand its scope.

2. Policy and Evaluation

2.1. Synthetic Records and Candidate Policies

Each artificial record has an integer age in [ 18 , 89 ] , a five-digit numeric location code, a sex value in { F , M } , and one of four diagnosis labels. The records are generated randomly using a fixed seed so that the same synthetic dataset can be reproduced in later runs. Age, location, sex, and diagnosis are generated independently, and the diagnosis labels are not intended to represent actual medical conditions.
A chromosome g = ( a , p , s ) specifies one global transformation, with a , p ∈ { 0 , 1 , 2 , 3 } and s ∈ { 0 , 1 } . Age levels are exact age, a 10-year interval, a 20-year interval, and a wildcard. Location levels are the exact five digits, the first three digits, the first digit, and a wildcard. Sex is either exact or a wildcard. Consequently there are 4 · 4 · 2 = 32 possible chromosomes. The same transformation is applied to every row; the method does not perform local recoding.
For each transformed quasi-identifier value q, let C g ( q ) be the corresponding equivalence class. The released set R g contains a class in full only when
| C g ( q ) | ≥ k and { d i : i ∈ C g ( q ) } ≥ ℓ ,
where d i is the unmodified diagnosis label. Otherwise the whole class is suppressed. The second condition is distinctℓ-diversity; it places no lower bound on the frequency of the less common diagnosis values.

2.2. Objectives and Feasibility

For the three hierarchy levels, define a normalized generalization-depth cost
H ( g ) = 1 3 a 3 + p 3 + s , S ( g ) = 1 − | R g | n .
The search minimizes H and S separately. We call a policy feasible when R g is nonempty and S ( g ) ≤ 0.30 after applying Eq. (1) with k = 5 and ℓ = 2 . Among feasible policies, g 1 dominates g 2 when both costs for g 1 are no greater and at least one is strictly smaller. Feasible policies dominate infeasible ones; infeasible policies are ordered by excess suppression, with an additional penalty for empty output.
H measures how much the data has been generalized based on the hierarchy levels used. Privacy is handled separately through the k-anonymity and ℓ-diversity constraints. Therefore, the Pareto front shows the tradeoff between generalization and suppression, not a direct tradeoff between patient privacy and clinical usefulness.

2.3. Search and Exact Reference

The search starts with the fully generalized and fully specific policies, then samples distinct random policies until it has a population of 12. It ranks policies by nondomination and crowding distance, selects parents by tournaments, chooses each offspring level from one of its parents, and mutates each level with probability 0.25 to another allowable level. Elitist replacement keeps 12 distinct policies from parents and offspring. Evaluation results are cached by chromosome. After 12 generations the algorithm reports the nondominated feasible policies among all chromosomes it evaluated, including those no longer in the population. All stochastic choices use a specified search seed.
To independently check the GA results, we also evaluate all 32 possible policies using exhaustive search and compare them with the policies found by the GA. This is practical for the small search space used here, but it would become expensive as the number of possible policies grows. The implementation also verifies the k-anonymity and distinct-ℓ-diversity requirements on the final output before saving the results to a CSV file.

3. Experimental Protocol and Results

We ran 20 paired seeds at each of n ∈ { 160 , 240 , 480 } generated rows. Replicate i ∈ { 0 , … , 19 } uses data seed 2026 + i and search seed 7 + i . The search settings, k, ℓ, and suppression threshold are identical in all 60 runs. A run “matches” only if the set of policies returned by the search equals the exhaustive feasible Pareto set. The experiment script records each seed, exact and search front sizes, policies evaluated, and the feasible solution selected by minimizing H and then S.
Table 1. Results over 20 generated datasets at each size. Search settings (population 12, generations 12, mutation probability 0.25 per gene) are identical across sizes. Evaluation range counts distinct chromosomes scored by the genetic search, out of 32 possible. Reported medians concern the selected candidate, not clinical performance.
Table 1. Results over 20 generated datasets at each size. Search settings (population 12, generations 12, mutation probability 0.25 per gene) are identical across sizes. Evaluation range counts distinct chromosomes scored by the genetic search, out of 32 possible. Reported medians concern the selected candidate, not clinical performance.
Rows Exact matches Median eval. Eval. range Median H Median S
160 19/20 29 25–31 0.4444 0.0000
240 20/20 29 26–31 0.3333 0.2125
480 20/20 28 26–30 0.2222 0.2094
For the default n = 240 run (data seed 2026, search seed 7), exhaustive enumeration finds two nondominated policies. Policy ( 3 , 1 , 0 ) has H = 4 / 9 and no suppressed rows. Policy ( 1 , 2 , 0 ) has H = 1 / 3 and suppresses 46 of 240 rows ( S ≈ 0.1917 ), leaving 194 rows in 28 retained classes. The fully generalized ( 3 , 3 , 1 ) policy retains every row but has H = 1 . The genetic search evaluates 29 of 32 policies and finds both exact-front policies in this run.
The only mismatch occurred for n = 160 with data seed 2033 and search seed 14. In this run, the genetic search evaluated 29 policies and returned two nondominated solutions. But exhaustive search found that policy ( 3 , 1 , 0 ) was better than both and had not been evaluated by the GA. This shows that the GA can miss the true Pareto front when it does not explore the full search space. Although the GA matched the exact result in 59 of 60 runs, it still ended up exhausting most of the 32 possible policies, so these results do not show an efficiency advantage over exhaustive search. We also did not measure runtime or compare the GA against random search.

4. Limitations and Next Experiments

The synthetic inputs do not resemble the joint distribution of real healthcare data. Because diagnosis is sampled independently and H only counts generalization steps; thus, H does not tell us whether a medical analysis would still be accurate after anonymization. The fixed shallow hierarchies and small feature set also make the search unusually easy. We have not evaluated uncertainty from different hierarchy designs, population settings and mutation rates as those are left for future iteration.
The released groups meet the implemented k and distinct-ℓ checks, but those checks do not imply compliance with any privacy law or safety for patient-data release. Distinct ℓ-diversity can permit highly skewed sensitive values. The software does not implement t-closeness, differential privacy, privacy accounting, or a data governance workflow. These are research directions requiring separate definitions and tests, not features of the present method.
A useful next study would introduce documented generalization hierarchies with substantially more policies; task-specific utility measures evaluated on properly governed data or realistic synthetic data with known correlations; threat-based privacy checks; and matched-budget comparisons against random search and established anonymization methods. Repeated seeds and runtime measurements would be needed before making any claim of improved efficiency. Any study using real healthcare records would require appropriate permission, safeguards, and disclosure review.

5. Conclusion

This experiment makes a proposed search approach executable and tests it against an exact answer on generated data. In 59 of 60 runs the genetic search recovered the full reference front, but it evaluated nearly the entire 32-policy space. The current evidence therefore supports reproducibility of a small implementation and identifies its limitations. The broader research question—whether genetic search helps retain useful information in large healthcare datasets under credible privacy constraints—remains open.

Code and Reproducibility

All code, the sweep script that reproduces the 60 runs summarized in Table 1, and per-run outputs are available at https://github.com/UsamaMehboob/healthcare-privacy-ga.

References

  1. L. Sweeney. k-anonymity: A model for protecting privacy. International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, 10(5):557–570, 2002. https://dataprivacylab.org/dataprivacy/projects/kanonymity/kanonymity.html.
  2. A. Machanavajjhala, J. Gehrke, D. Kifer, and M. Venkitasubramaniam. ℓ-diversity: Privacy beyond k-anonymity. In Proceedings of the 22nd IEEE International Conference on Data Engineering, 2006. https://www.cs.cornell.edu/johannes/papers/2006/2006-icde-publishing.pdf.
  3. K. Deb, A. Pratap, S. Agarwal, and T. Meyarivan. A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2):182–197, 2002. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.