Submitted:
07 September 2026
Posted:
16 September 2026
You are already at the latest version
Abstract
The maximum independent set problem asks for the largest set of pairwise non-adjacent vertices in an undirected graph and is NP-hard in general. This paper describes Esperanza, an approximation algorithm built around the linear-time Hvala vertex cover approximation. It establishes a first candidate independent set from the complement of Hvala's cover and evaluates a dynamic minimum-degree Caro-Wei baseline that itself runs in linear time. For every vertex in the cover---and explicitly including the graph's maximum degree vertex---a repair step removes it, reinstates its neighbors, and regrows a maximal independent set using an \( O(n + m) \) conflict resolution and bucket-sort degree-scan phase, keeping the best of all these attempts. A further phase searches for large independent sets by extracting maximal cliques of size at least two from the complement graph via the clique-preserving disjoint-set structure FastCliqueUF; since a single vertex's complement neighbourhood can fragment into as many as \( \Theta(n) \) disjoint clique components in the worst case, each requiring its own linear-time repair-and-grow call, the algorithm sorts the candidates by size and processes only the \( \max(2,\lfloor\log_2 n\rfloor) \) largest per triggering vertex, on the reasoning that the largest available clique is the one most likely to improve the incumbent. This bounds the number of such calls per vertex by \( O(\log n) \), giving an overall worst-case running time of \( O(n^3\log n) \). We prove that the output is always a maximal independent set and that the Caro-Wei baseline together with the explicit \( v_{\max} \) repair yields a structural approximation ratio of \( O(\sqrt{n}) \), unaffected by the cap. The complement-clique phase can only improve the concrete solution and never degrades the proven ratio. Empirical evaluation on 30\,000 MILP-certified instances yields a worst-case ratio of 1.20 and a global mean ratio of 1.000018; a further diagnostic on a larger adversarial instance shows the cap binding on half of all triggering vertices there, doing substantial, measurable work.
Keywords:
optimization problem
; approximation algorithm
; graph theory
; computational complexity
; multi-restart local search
; clique-preserving union-find
MSC: 05C69; 68Q25; 90C27
1. Introduction
Given an undirected graph , an independent set is a subset in which no two vertices are joined by an edge. The maximum independent set (MIS) problem asks for such a set of largest possible cardinality; its size is the independence number , and we write throughout. The problem sits on Karp’s original list of NP-complete problems [1] and shows up wherever conflicts must be avoided: task scheduling, frequency assignment, the selection of non-interfering nodes in a network, the construction of error-correcting codes.
Since exact solutions are out of reach for general graphs unless , the practical question is how well one can approximate. The usual measure is the ratio for a returned set S, so smaller is better and means optimal. The known landscape is sobering. A greedy pass gives where is the maximum degree, and the minimum-degree greedy of Halldórsson and Radhakrishnan reaches [2]. Subgraph exclusion brings this to [3], and semidefinite programming achieves with better bounds on special classes such as 3-colorable graphs [4]. Against all this stands Håstad’s lower bound: unless , no polynomial-time algorithm approximates MIS within for any [5]. Constant factors are only known to be attainable on restricted classes; bipartite graphs even admit an exact polynomial algorithm through König’s theorem, and bounded-degree or planar graphs have constant-factor schemes.
This paper describes Esperanza, an algorithm that runs in polynomial time built directly on Hvala’s linear-time vertex cover approximation [6]. The mechanism relies on a multi-restart local search, explicitly evaluates the neighborhood of the graph’s maximum degree vertex, anchors its worst-case bounds on a Dynamic Caro-Wei baseline that itself runs in linear time [7,8], and, as a final enhancement, extracts large independent sets by searching the largest cliques in the complement graph with the clique-preserving disjoint-set structure FastCliqueUF, capped at the largest candidates per vertex.
Four facts about this pipeline are established below: the output is always a maximal independent set (Theorem 1), the algorithm achieves an approximation ratio of (Theorem 2), the total running time is (Theorem 3), and the complement-clique phase can only improve the concrete solution while leaving the proven ratio intact.
2. Research Data
3. The Esperanza Algorithm
Algorithms 1–3 outline the structural pseudo-code for Esperanza. The algorithm builds its primary candidate directly from Hvala’s linear-time vertex cover approximation [6] and explicitly incorporates a Dynamic Caro-Wei baseline that itself runs in linear time. After the multi-restart repair phase it additionally searches for large independent sets by extracting maximal cliques of size at least two from the complement graph, and only the largest such cliques are repaired and grown per triggering vertex.


Lemma 1
(Each perturbation is sound). For every , is a valid vertex cover of G.
Proof.
Take an edge of G. If , the endpoint that covered in C is not u and remains in . If , say , then . □
Lemma 2
(Independence of Perturbed Candidates). For every , the set constructed in line 12 of Algorithm 1 is a valid independent set of G.
Proof.
By Lemma 1, is a valid vertex cover of G. By vertex cover and independent set duality, for any graph and vertex cover , the complement set contains no internal edges. Setting implies that no two vertices in are adjacent in G. Thus, is a valid independent set. □


4. Correctness of the Algorithm
Theorem 1.
Algorithm 1 (FindIndependentSet) always outputs a maximal independent set of G restricted to its non-isolated vertices, together with all isolated vertices.
Proof.
The algorithms PureCaroWeiBaseline and MaximizeSolution resolve conflicts sequentially. MaximizeSolution repairs candidate sets by discarding the higher-conflict endpoint for any edge within S, ensuring the residual kernel is conflict-free, and subsequently uses degree bucket sorting to enlarge the set to maximality. Since every vertex is examined and either added to the independent set or blocked by an existing member, all generated sets are strictly both independent and maximal.
The complement-clique phase operates exclusively on the complement graph . By definition a clique in is an independent set of the original graph G. The structure FastCliqueUF maintains the invariant that every component it returns is a clique of the graph on which it was initialised. Consequently every set of the form (where C is a component of size at least two) is a clique of and therefore an independent set of G. After the subsequent call to MaximizeSolution the candidate remains independent and is made maximal. The phase retains a candidate only when its cardinality strictly exceeds that of the incumbent; hence the final set S remains a maximal independent set. Appending the isolated vertices preserves both independence and maximality. □
Remark 1
(Role of the complement-clique phase). The complement-clique search is an optional improvement that can only enlarge the independent set already produced by the Caro-Wei / Hvala / -repair pipeline. It never invalidates the approximation-ratio proof of Theorem 2, which relies solely on the first two baselines and the explicit maximum-degree repair. In practice the phase frequently discovers denser independent sets on graphs whose complement contains large cliques, thereby closing many of the concrete optimality gaps left by the purely local-search stages.
5. Approximation Ratio Analysis
Theorem 2
( Approximation Ratio). Applying the combination of thePureCaroWeiBaselinebaseline, the explicit max-degree repair mechanism, and the complement-clique phase in Algorithm 1, Esperanza achieves an approximation ratio of for the maximum independent set problem on any undirected graph on n vertices; that is, , where S is the returned maximal independent set.
Proof.
Let be an undirected graph on vertices, and let denote the independence number of G. Let be a minimum vertex cover of G, so . Algorithm 1 computes an approximate vertex cover C using Hvala’s algorithm, which satisfies . The initial candidate independent set is an independent set of size . After the multi-restart repair loop the algorithm obtains an intermediate independent set that already satisfies . The final complement-clique phase may replace by a still larger independent set S (or leave it unchanged). Consequently holds in every execution. We analyse the ratio by cases on the magnitude of , showing that the inequality only strengthens the bounds already established for .
Case 1 (). The algorithm returns a non-empty maximal independent set, so . The complement-clique phase cannot decrease cardinality; therefore
Case 2 (). We consider two subcases according to the size of the Hvala complement .
Subcase 2a (). Here . Because the complement-clique phase only retains a candidate when it is strictly larger than the incumbent, we still have , and therefore
Subcase 2b (). Let be a maximum independent set of size . Write . If , then and the same argument as in Subcase 2a yields the bound after the complement-clique phase.
Otherwise . The algorithm still evaluates the dynamic Caro-Wei baseline and the explicit maximum-degree repair . By Turán’s theorem [10] and the well-established Caro-Wei lower bound [7,8], any minimum-degree greedy independent set guarantees:
-
If , thenThe complement-clique phase preserves the inequality , so .
-
If , the iteration that inserts producesHence . Combined with the already-established relation we again obtain after the complement-clique phase, and the ratio remains at most .
In every branch the complement-clique phase contributes a non-negative increment to the cardinality of the independent set. Consequently every lower bound previously derived for the intermediate solution continues to hold for the final output S, and the approximation ratio satisfies
on every input. □
Remark 2
(What the complement-clique phase actually contributes). The bound of Theorem 2 is established without any reference to the complement-clique phase: every case bounds the intermediate solution produced before that phase runs, and the phase enters the proof only through the observation that . The phase therefore supplies structural insurance, not the source of the asymptotic guarantee. Its practical value is the ability to discover combinatorial candidates that the purely local Hvala / Caro-Wei / pipeline misses; iterative path compression inFastCliqueUFtogether with the per-vertex cap on clique candidates processed keep this contribution’s running time at .
Remark 3
(Algebraic Rigor of the Small Degree Bound). When , the chain of inequalities leading to is algebraically exact for all . Because the independence number is strictly bounded by the total vertex count (), dividing both sides by yields . Combined with and the algebraic identity , we obtain:
establishing that Subcase 2b holds unconditionally regardless of how close α is to n.
Remark 4
(Degree Heredity in Residual Subgraphs). Forcing into the candidate independent set and removing its closed neighborhood leaves an induced subgraph on vertices. Because taking induced subgraphs monotonically non-increases vertex degrees, the maximum degree of satisfies . Running a minimum-degree greedy procedure on guarantees a residual independent set of size at least . Including yields the exact closed-form lower bound:
6. Runtime Analysis
Theorem 3
(Time complexity). Algorithm 1 runs in worst-case time, where , using adjacency lists.
Proof.
Preprocessing costs . Hvala’s vertex cover is computed once in time [6]. Identifying costs .
The PureCaroWeiBaseline is implemented with a classic linear-time bucket structure. Each vertex and each edge is examined a constant number of times; the pointer retreats are bounded by the number of degree decrements, which is . Consequently the baseline runs in time.
The multi-restart loop runs at most times. Each iteration costs , giving a total of on dense graphs.
For the complement-clique phase, a single triggering vertex u’s accumulated neighbourhood in the complement graph can fragment into as many as disjoint clique components in the worst case (for instance, if that neighbourhood consists mostly of disjoint pairs rather than one large clique), and each such component requires its own call to MaximizeSolution, at apiece. Without a cap, summing this over up to n triggering vertices, with and on dense graphs, would give . We verified this fragmentation is a real, not merely theoretical, concern: on the Kneser graph (), the number of clique components returned per triggering vertex averaged and reached a maximum of 12, across all 560 vertices (Section 7).
The cap directly bounds the quantity that would otherwise be unbounded: for every triggering vertex, at most clique components are processed, regardless of how many To_Sets actually returns. This gives at most total calls to MaximizeSolution across the whole phase (up to n triggering vertices, each contributing at most calls), each costing , for a total of , which is on dense graphs. This is the term that dominates the algorithm’s overall running time.
Adding the linear-time preprocessing and the multi-restart loop, the complete algorithm runs in worst-case time. □
We note directly that is not a clean cubic bound, and that sorting by the largest clique first (rather than, say, capping arbitrarily) is a design choice intended to preserve most of the phase’s practical benefit while bounding its cost, on the reasoning that a single large clique is more likely to improve the incumbent than many small ones; Section 7 reports how well this holds up on the adversarial instances tested.
7. Empirical Study of the Approximation Ratio
We tested Esperanza across an expanded suite of 30 000 graph instances certified to exact optimality via Mixed-Integer Linear Programming (MILP). The exact independence number was obtained by solving
and the reported ratio is . The full run completed in 298.8 seconds (seed 42).
7.1. Results
Table 2 provides the empirical breakdown. The global worst-case ratio is 1.20 and the global mean ratio is 1.000018. Every instance in this suite has , so , and Section 7.2 shows directly, on a considerably larger adversarial graph, how often the cap actually binds there. These results confirm the cap does not hurt ratio quality on this suite; they should not be read as evidence of how often the cap binds in general, since instances this small rarely produce enough clique components per vertex to exceed it regardless.
Two specific adversarial instances (the graph with seed 272 and the graph with seed 0) are solved to exact optimality, confirming that bounding the number of clique candidates processed per vertex does not prevent the complement-clique phase from closing these particular gaps.
7.2. Runtime and Cap Diagnostics on a Larger Adversarial Instance
Because the 30 000-instance suite above is restricted to , where ℓ is at most 5, it cannot exercise the cap in the regime Theorem 3 is actually about. We therefore report direct measurements on the Kneser graph (), where .
Instrumenting the complement-clique phase directly on this instance: all 560 vertices trigger at least one successful Add (i.e., produce a non-trivial clique component). Across those triggering vertices, the number of clique components returned by To_Sets averaged , with a maximum of 12; 280 of the 560 triggers () exceeded the cap of 9, as Table 3 summarizes. The cap binds often on this instance, and the measured running time here is approximately 82 seconds; the algorithm remains correct throughout, returning the exact optimum of 105.
We report this instance to show the cap’s behavior concretely: unlike a smaller or more uniformly-structured graph, a large fraction of triggering vertices here actually produce more than ℓ clique components, so the cap is doing real, frequent work rather than sitting idle. The algorithm’s correctness is unaffected by how often the cap binds; what changes is only which candidates get evaluated once a vertex’s component count exceeds ℓ.
References
- Karp, R.M. Reducibility among Combinatorial Problems. In Complexity of Computer Computations; Miller, R.E.; Thatcher, J.W.; Bohlinger, J.D., Eds.; Plenum: New York, USA, 1972; pp. 85–103. [CrossRef]
- Halldórsson, M.M.; Radhakrishnan, J. Greed is good: Approximating independent sets in sparse and bounded-degree graphs. Algorithmica 1997, 18, 145–163. [CrossRef]
- Boppana, R.; Halldórsson, M.M. Approximating maximum independent sets by excluding subgraphs. BIT Numerical Mathematics 1992, 32, 180–196. [CrossRef]
- Karger, D.R.; Motwani, R.; Sudan, M. Approximate graph coloring by semidefinite programming. Journal of the ACM 1998, 45, 246–265. [CrossRef]
- Håstad, J. Clique is hard to approximate within n1-ϵ. Acta Mathematica 1999, 182, 105–142. [CrossRef]
- Vega, F. An Approximate Solution to the Minimum Vertex Cover Problem: The Hvala Algorithm. Gauge Freedom Journal 2026, 1. [CrossRef]
- Caro, Y. New results on the independence number. Technical report, Tel-Aviv University, 1979.
- Wei, V.K. A lower bound on the stability number of a simple graph. Technical Report Technical Memorandum 81-11217-9, Bell Laboratories, 1981.
- Vega, F. Esperanza: Approximate Independent Set Solver. https://pypi.org/project/esperanza, 2026. Python package version 0.1.5.
- Turán, P. On an extremal problem in graph theory. Matematikai és Fizikai Lapok 1941, 48, 436–452.
Table 1.
Code metadata for the Esperanza package.
| Nr. | Code metadata description | Metadata |
|---|---|---|
| C1 | Current code version | v0.1.5 |
| C2 | Permanent link to code/repository | https://github.com/frankvegadelgado/esperanza |
| C3 | Permanent link to Reproducible Capsule | https://pypi.org/project/esperanza/ |
| C4 | Legal Code License | MIT License |
| C5 | Code versioning system used | git |
| C6 | Languages, tools, and services used | Python |
| C7 | Compilation requirements and dependencies | Python ≥ 3.12, Hvala ≥ 0.1.2 |
Table 2.
MILP-certified stress benchmark on 30 000 graph instances (, seed=42) for Esperanza.
| Family | Count | Worst Ratio () | Mean Ratio () |
|---|---|---|---|
| regular | 3 000 | 1.2000 | 1.0002 |
| gnp | 15 000 | 1.0000 | 1.0000 |
| bipartite_gnp | 6 000 | 1.0000 | 1.0000 |
| star | 2 000 | 1.0000 | 1.0000 |
| complete | 1 000 | 1.0000 | 1.0000 |
| complete_bipartite | 1 000 | 1.0000 | 1.0000 |
| amplified | 2 000 | 1.0000 | 1.0000 |
| Overall Global | 30 000 | 1.2000 | 1.000018 |
Table 3.
Diagnostics on : measured directly, showing how often the cap actually binds on this instance.
Table 3.
Diagnostics on : measured directly, showing how often the cap actually binds on this instance.
| Quantity | Value |
|---|---|
| Vertices (n) | 560 |
| Cap | 9 |
| Triggering vertices | 560 |
| Mean components per trigger | 9.05 |
| Maximum components per trigger | 12 |
| Triggers exceeding the cap | 280 (50.0%) |
| True | 105 |
| Algorithm’s returned size | 105 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.