Submitted:
24 September 2026
Posted:
28 September 2026
You are already at the latest version
Abstract
Anonymizing sensitive healthcare data while preserving data utility is always a tradeoff between suppression and generalization. In this experiment, we employ a genetic algorithm to search the anonymization policy space using synthetic healthcare-style data. Each policy candidate specifies different levels of generalization for quasi-identifiers such as age, a five-digit numeric location code, and sex, while any equivalence class that does not satisfy the constraints of k-anonymity or distinct ℓ-diversity is suppressed. The GA experiment is carried out under constraints of k = 5, ℓ = 2, and a 30% suppression limit. There are 32 possible anonymization policies, and the search space is intentionally kept small so that the results of the GA-based search can be compared directly with exhaustive search to identify its accuracy, gaps, and limitations. Across 60 seeded runs using three different table sizes, the GA search recovered the exact Pareto set in 59 runs while evaluating a median of 28–29 policies. These results show that the GA was capable of reliably recovering the Pareto-optimal solutions in this small synthetic setting, but GA still ended up evaluating most of the policies. The efficiency advantage of GA has yet to be evaluated in the future on large real world healthcare data with substantially larger policy spaces to establish if GA indeed offers more clinical utility with less compute compared to other traditional methods.
Keywords:
healthcare data anonymization
; multi-objective genetic search
; synthetic healthcare data
; k-anonymity
; ℓ-diversity
; generalization
; suppression
; Pareto optimization
1. Introduction
Generalization and suppression are the primary tools used on real word data to create equivalence-classes of any size. One generalization benchmark is k-anonymity parameter, that means a combination of quasi-identifiers occurs in at least k records [1]. ℓ-diversity applies to sensitive attributes of class, means the equivalence class sensitive attributes should have l distinct values. [2]. Neither criterion by itself is a general guarantee against re-identification or attribute disclosure.
The intended goal of this research is to establish whether multi-objective genetic search can find anonymization policies that meet privacy constraints while preserving information for specific healthcare analyses. Before tackling that question, we need an implementation whose behavior can be checked against a known answer. This note provides such a pilot experiment. It leverages a concrete policy encoding and class-suppression rule to perform a reproducible comparison between a genetic search and exhaustive enumeration in a 32-policy space. The search builds on a nondominated sorting genetic algorithm techinique described here [3] and utilizes healthcare constraints to expand its scope.
2. Policy and Evaluation
2.1. Synthetic Records and Candidate Policies
Each artificial record has an integer age in , a five-digit numeric location code, a sex value in , and one of four diagnosis labels. The records are generated randomly using a fixed seed so that the same synthetic dataset can be reproduced in later runs. Age, location, sex, and diagnosis are generated independently, and the diagnosis labels are not intended to represent actual medical conditions.
A chromosome specifies one global transformation, with and . Age levels are exact age, a 10-year interval, a 20-year interval, and a wildcard. Location levels are the exact five digits, the first three digits, the first digit, and a wildcard. Sex is either exact or a wildcard. Consequently there are possible chromosomes. The same transformation is applied to every row; the method does not perform local recoding.
For each transformed quasi-identifier value q, let be the corresponding equivalence class. The released set contains a class in full only when
where is the unmodified diagnosis label. Otherwise the whole class is suppressed. The second condition is distinctℓ-diversity; it places no lower bound on the frequency of the less common diagnosis values.
2.2. Objectives and Feasibility
For the three hierarchy levels, define a normalized generalization-depth cost
The search minimizes H and S separately. We call a policy feasible when is nonempty and after applying Eq. (1) with and . Among feasible policies, dominates when both costs for are no greater and at least one is strictly smaller. Feasible policies dominate infeasible ones; infeasible policies are ordered by excess suppression, with an additional penalty for empty output.
H measures how much the data has been generalized based on the hierarchy levels used. Privacy is handled separately through the k-anonymity and ℓ-diversity constraints. Therefore, the Pareto front shows the tradeoff between generalization and suppression, not a direct tradeoff between patient privacy and clinical usefulness.
2.3. Search and Exact Reference
The search starts with the fully generalized and fully specific policies, then samples distinct random policies until it has a population of 12. It ranks policies by nondomination and crowding distance, selects parents by tournaments, chooses each offspring level from one of its parents, and mutates each level with probability 0.25 to another allowable level. Elitist replacement keeps 12 distinct policies from parents and offspring. Evaluation results are cached by chromosome. After 12 generations the algorithm reports the nondominated feasible policies among all chromosomes it evaluated, including those no longer in the population. All stochastic choices use a specified search seed.
To independently check the GA results, we also evaluate all 32 possible policies using exhaustive search and compare them with the policies found by the GA. This is practical for the small search space used here, but it would become expensive as the number of possible policies grows. The implementation also verifies the k-anonymity and distinct-ℓ-diversity requirements on the final output before saving the results to a CSV file.
3. Experimental Protocol and Results
We ran 20 paired seeds at each of generated rows. Replicate uses data seed and search seed . The search settings, k, ℓ, and suppression threshold are identical in all 60 runs. A run “matches” only if the set of policies returned by the search equals the exhaustive feasible Pareto set. The experiment script records each seed, exact and search front sizes, policies evaluated, and the feasible solution selected by minimizing H and then S.
Table 1.
Results over 20 generated datasets at each size. Search settings (population 12, generations 12, mutation probability 0.25 per gene) are identical across sizes. Evaluation range counts distinct chromosomes scored by the genetic search, out of 32 possible. Reported medians concern the selected candidate, not clinical performance.
Table 1.
Results over 20 generated datasets at each size. Search settings (population 12, generations 12, mutation probability 0.25 per gene) are identical across sizes. Evaluation range counts distinct chromosomes scored by the genetic search, out of 32 possible. Reported medians concern the selected candidate, not clinical performance.
| Rows | Exact matches | Median eval. | Eval. range | Median H | Median S |
|---|---|---|---|---|---|
| 160 | 19/20 | 29 | 25–31 | 0.4444 | 0.0000 |
| 240 | 20/20 | 29 | 26–31 | 0.3333 | 0.2125 |
| 480 | 20/20 | 28 | 26–30 | 0.2222 | 0.2094 |
For the default run (data seed 2026, search seed 7), exhaustive enumeration finds two nondominated policies. Policy has and no suppressed rows. Policy has and suppresses 46 of 240 rows (), leaving 194 rows in 28 retained classes. The fully generalized policy retains every row but has . The genetic search evaluates 29 of 32 policies and finds both exact-front policies in this run.
The only mismatch occurred for with data seed 2033 and search seed 14. In this run, the genetic search evaluated 29 policies and returned two nondominated solutions. But exhaustive search found that policy was better than both and had not been evaluated by the GA. This shows that the GA can miss the true Pareto front when it does not explore the full search space. Although the GA matched the exact result in 59 of 60 runs, it still ended up exhausting most of the 32 possible policies, so these results do not show an efficiency advantage over exhaustive search. We also did not measure runtime or compare the GA against random search.
4. Limitations and Next Experiments
The synthetic inputs do not resemble the joint distribution of real healthcare data. Because diagnosis is sampled independently and H only counts generalization steps; thus, H does not tell us whether a medical analysis would still be accurate after anonymization. The fixed shallow hierarchies and small feature set also make the search unusually easy. We have not evaluated uncertainty from different hierarchy designs, population settings and mutation rates as those are left for future iteration.
The released groups meet the implemented k and distinct-ℓ checks, but those checks do not imply compliance with any privacy law or safety for patient-data release. Distinct ℓ-diversity can permit highly skewed sensitive values. The software does not implement t-closeness, differential privacy, privacy accounting, or a data governance workflow. These are research directions requiring separate definitions and tests, not features of the present method.
A useful next study would introduce documented generalization hierarchies with substantially more policies; task-specific utility measures evaluated on properly governed data or realistic synthetic data with known correlations; threat-based privacy checks; and matched-budget comparisons against random search and established anonymization methods. Repeated seeds and runtime measurements would be needed before making any claim of improved efficiency. Any study using real healthcare records would require appropriate permission, safeguards, and disclosure review.
5. Conclusion
This experiment makes a proposed search approach executable and tests it against an exact answer on generated data. In 59 of 60 runs the genetic search recovered the full reference front, but it evaluated nearly the entire 32-policy space. The current evidence therefore supports reproducibility of a small implementation and identifies its limitations. The broader research question—whether genetic search helps retain useful information in large healthcare datasets under credible privacy constraints—remains open.
Code and Reproducibility
All code, the sweep script that reproduces the 60 runs summarized in Table 1, and per-run outputs are available at https://github.com/UsamaMehboob/healthcare-privacy-ga.
References
- L. Sweeney. k-anonymity: A model for protecting privacy. International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, 10(5):557–570, 2002. https://dataprivacylab.org/dataprivacy/projects/kanonymity/kanonymity.html.
- A. Machanavajjhala, J. Gehrke, D. Kifer, and M. Venkitasubramaniam. ℓ-diversity: Privacy beyond k-anonymity. In Proceedings of the 22nd IEEE International Conference on Data Engineering, 2006. https://www.cs.cornell.edu/johannes/papers/2006/2006-icde-publishing.pdf.
- K. Deb, A. Pratap, S. Agarwal, and T. Meyarivan. A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2):182–197, 2002. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.