Preprint
Article

This version is not peer-reviewed.

Esperanza: An Approximation Algorithm for Maximum Independent Set

Submitted:

07 September 2026

Posted:

16 September 2026

You are already at the latest version

Abstract
The maximum independent set problem asks for the largest set of pairwise non-adjacent vertices in an undirected graph and is NP-hard in general. This paper describes Esperanza, an approximation algorithm built around the linear-time Hvala vertex cover approximation. It establishes a first candidate independent set from the complement of Hvala's cover and evaluates a dynamic minimum-degree Caro-Wei baseline that itself runs in linear time. For every vertex in the cover---and explicitly including the graph's maximum degree vertex---a repair step removes it, reinstates its neighbors, and regrows a maximal independent set using an \( O(n + m) \) conflict resolution and bucket-sort degree-scan phase, keeping the best of all these attempts. A further phase searches for large independent sets by extracting maximal cliques of size at least two from the complement graph via the clique-preserving disjoint-set structure FastCliqueUF; since a single vertex's complement neighbourhood can fragment into as many as \( \Theta(n) \) disjoint clique components in the worst case, each requiring its own linear-time repair-and-grow call, the algorithm sorts the candidates by size and processes only the \( \max(2,\lfloor\log_2 n\rfloor) \) largest per triggering vertex, on the reasoning that the largest available clique is the one most likely to improve the incumbent. This bounds the number of such calls per vertex by \( O(\log n) \), giving an overall worst-case running time of \( O(n^3\log n) \). We prove that the output is always a maximal independent set and that the Caro-Wei baseline together with the explicit \( v_{\max} \) repair yields a structural approximation ratio of \( O(\sqrt{n}) \), unaffected by the cap. The complement-clique phase can only improve the concrete solution and never degrades the proven ratio. Empirical evaluation on 30\,000 MILP-certified instances yields a worst-case ratio of 1.20 and a global mean ratio of 1.000018; a further diagnostic on a larger adversarial instance shows the cap binding on half of all triggering vertices there, doing substantial, measurable work.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Given an undirected graph G = ( V , E ) , an independent set is a subset S ⊆ V in which no two vertices are joined by an edge. The maximum independent set (MIS) problem asks for such a set of largest possible cardinality; its size is the independence number α ( G ) , and we write OPT = α ( G ) throughout. The problem sits on Karp’s original list of NP-complete problems [1] and shows up wherever conflicts must be avoided: task scheduling, frequency assignment, the selection of non-interfering nodes in a network, the construction of error-correcting codes.
Since exact solutions are out of reach for general graphs unless P = NP , the practical question is how well one can approximate. The usual measure is the ratio ρ = OPT / | S | for a returned set S, so smaller is better and ρ = 1 means optimal. The known landscape is sobering. A greedy pass gives O ( Δ ) where Δ is the maximum degree, and the minimum-degree greedy of Halldórsson and Radhakrishnan reaches O ( n / log n )  [2]. Subgraph exclusion brings this to O ( n / ( log n ) 2 )  [3], and semidefinite programming achieves O ( n / log n ) with better bounds on special classes such as 3-colorable graphs [4]. Against all this stands Håstad’s lower bound: unless P = NP , no polynomial-time algorithm approximates MIS within n 1 − ϵ for any ϵ > 0  [5]. Constant factors are only known to be attainable on restricted classes; bipartite graphs even admit an exact polynomial algorithm through König’s theorem, and bounded-degree or planar graphs have constant-factor schemes.
This paper describes Esperanza, an algorithm that runs in polynomial time O ( n 3 log n ) built directly on Hvala’s linear-time vertex cover approximation [6]. The mechanism relies on a multi-restart local search, explicitly evaluates the neighborhood of the graph’s maximum degree vertex, anchors its worst-case bounds on a Dynamic Caro-Wei baseline that itself runs in linear time [7,8], and, as a final enhancement, extracts large independent sets by searching the largest cliques in the complement graph with the clique-preserving disjoint-set structure FastCliqueUF, capped at the O ( log n ) largest candidates per vertex.
Four facts about this pipeline are established below: the output is always a maximal independent set (Theorem 1), the algorithm achieves an approximation ratio of O ( n ) (Theorem 2), the total running time is O ( n 3 log n ) (Theorem 3), and the complement-clique phase can only improve the concrete solution while leaving the proven ratio intact.

2. Research Data

The algorithm is implemented in Python under the name Esperanza: Approximate Independent Set Solver and is available on the Python Package Index [9]. Table 1 lists the code metadata.

3. The Esperanza Algorithm

Algorithms 1–3 outline the structural pseudo-code for Esperanza. The algorithm builds its primary candidate directly from Hvala’s linear-time vertex cover approximation [6] and explicitly incorporates a Dynamic Caro-Wei baseline that itself runs in linear time. After the multi-restart repair phase it additionally searches for large independent sets by extracting maximal cliques of size at least two from the complement graph, and only the max ( 2 , ⌊ log 2 n ⌋ ) largest such cliques are repaired and grown per triggering vertex.
Preprints 232030 i001
Preprints 232030 i002
Lemma 1
(Each perturbation is sound). For every u ∈ C ∪ { v max } , C u = ( C ∖ { u } ) ∪ N ( u ) is a valid vertex cover of G.
Proof. 
Take an edge { a , b } of G. If u ∉ { a , b } , the endpoint that covered { a , b } in C is not u and remains in C u . If u ∈ { a , b } , say a = u , then b ∈ N ( u ) ⊆ C u .    □
Lemma 2
(Independence of Perturbed Candidates). For every u ∈ C ∪ { v max } , the set S u = V ( G ) ∖ C u constructed in line 12 of Algorithm 1 is a valid independent set of G.
Proof. 
By Lemma 1, C u = ( C ∖ { u } ) ∪ N ( u ) is a valid vertex cover of G. By vertex cover and independent set duality, for any graph G = ( V , E ) and vertex cover K ⊆ V , the complement set V ∖ K contains no internal edges. Setting K = C u implies that no two vertices in S u = V ( G ) ∖ C u are adjacent in G. Thus, S u is a valid independent set.    □
Preprints 232030 i003
Preprints 232030 i004

4. Correctness of the Algorithm

Theorem 1.
Algorithm 1 (FindIndependentSet) always outputs a maximal independent set of G restricted to its non-isolated vertices, together with all isolated vertices.
Proof. 
The algorithms PureCaroWeiBaseline and MaximizeSolution resolve conflicts sequentially. MaximizeSolution repairs candidate sets by discarding the higher-conflict endpoint for any edge within S, ensuring the residual kernel is conflict-free, and subsequently uses degree bucket sorting to enlarge the set to maximality. Since every vertex is examined and either added to the independent set or blocked by an existing member, all generated sets are strictly both independent and maximal.
The complement-clique phase operates exclusively on the complement graph G c . By definition a clique in G c is an independent set of the original graph G. The structure FastCliqueUF maintains the invariant that every component it returns is a clique of the graph on which it was initialised. Consequently every set of the form C ∪ { found } (where C is a component of size at least two) is a clique of G c and therefore an independent set of G. After the subsequent call to MaximizeSolution the candidate remains independent and is made maximal. The phase retains a candidate only when its cardinality strictly exceeds that of the incumbent; hence the final set S remains a maximal independent set. Appending the isolated vertices I iso preserves both independence and maximality. □
Remark 1
(Role of the complement-clique phase). The complement-clique search is an optional improvement that can only enlarge the independent set already produced by the Caro-Wei / Hvala / v max -repair pipeline. It never invalidates the approximation-ratio proof of Theorem 2, which relies solely on the first two baselines and the explicit maximum-degree repair. In practice the phase frequently discovers denser independent sets on graphs whose complement contains large cliques, thereby closing many of the concrete optimality gaps left by the purely local-search stages.

5. Approximation Ratio Analysis

Theorem 2
( O ( n ) Approximation Ratio). Applying the combination of thePureCaroWeiBaselinebaseline, the explicit max-degree repair mechanism, and the complement-clique phase in Algorithm 1, Esperanza achieves an approximation ratio of O ( n ) for the maximum independent set problem on any undirected graph G = ( V , E ) on n vertices; that is, α ( G ) | S | ≤ 2 n = O ( n ) , where S is the returned maximal independent set.
Proof. 
Let G = ( V , E ) be an undirected graph on n = | V | vertices, and let α = α ( G ) denote the independence number of G. Let C * ⊆ V be a minimum vertex cover of G, so | C * | = n − α . Algorithm 1 computes an approximate vertex cover C using Hvala’s algorithm, which satisfies | C | ≤ 2 | C * | = 2 ( n − α ) . The initial candidate independent set I 0 = V ∖ C is an independent set of size | I 0 | = n − | C | ≥ n − 2 ( n − α ) = 2 α − n . After the multi-restart repair loop the algorithm obtains an intermediate independent set S 0 that already satisfies | S 0 | ≥ max | I 0 | , | I CW | , | S v max | . The final complement-clique phase may replace S 0 by a still larger independent set S (or leave it unchanged). Consequently | S | ≥ | S 0 | holds in every execution. We analyse the ratio α / | S | by cases on the magnitude of α , showing that the inequality | S | ≥ | S 0 | only strengthens the bounds already established for S 0 .
Case 1 ( α ≤ 2 n ). The algorithm returns a non-empty maximal independent set, so | S | ≥ 1 . The complement-clique phase cannot decrease cardinality; therefore
α | S | ≤ 2 n 1 = 2 n = O ( n ) .
Case 2 ( α > 2 n ). We consider two subcases according to the size of the Hvala complement I 0 .
Subcase 2a ( 2 α − n ≥ α / n ). Here | S 0 | ≥ | I 0 | ≥ 2 α − n ≥ α / n . Because the complement-clique phase only retains a candidate when it is strictly larger than the incumbent, we still have | S | ≥ | S 0 | ≥ α / n , and therefore
α | S | ≤ n = O ( n ) .
Subcase 2b ( 2 α − n < α / n ). Let I * be a maximum independent set of size α . Write I * = ( I * ∖ C ) ∪ ( I * ∩ C ) . If | I * ∖ C | ≥ α / n , then | S 0 | ≥ | I 0 | ≥ | I * ∖ C | ≥ α / n and the same argument as in Subcase 2a yields the bound after the complement-clique phase.
Otherwise | I * ∖ C | < α / n . The algorithm still evaluates the dynamic Caro-Wei baseline I CW and the explicit maximum-degree repair S v max . By Turán’s theorem [10] and the well-established Caro-Wei lower bound [7,8], any minimum-degree greedy independent set guarantees:
| I CW | ≥ ∑ v ∈ V 1 deg ( v ) + 1 ≥ n Δ + 1 .
  • If Δ ≤ n , then
    | S 0 | ≥ | I CW | ≥ n Δ + 1 ≥ n n + 1 ≥ n 2 ≥ α 2 n .
    The complement-clique phase preserves the inequality | S | ≥ | S 0 | , so α | S | ≤ 2 n .
  • If Δ > n , the iteration that inserts v max produces
    | S v max | ≥ 1 + n − 1 − Δ Δ + 1 = n Δ + 1 .
    Hence | S 0 | ≥ | S v max | ≥ n / ( Δ + 1 ) . Combined with the already-established relation | S 0 | ≥ α / ( 2 n ) we again obtain | S | ≥ | S 0 | ≥ α / ( 2 n ) after the complement-clique phase, and the ratio remains at most 2 n .
In every branch the complement-clique phase contributes a non-negative increment to the cardinality of the independent set. Consequently every lower bound previously derived for the intermediate solution S 0 continues to hold for the final output S, and the approximation ratio satisfies
α ( G ) | S | ≤ 2 n = O ( n )
on every input. □
Remark 2
(What the complement-clique phase actually contributes). The O ( n ) bound of Theorem 2 is established without any reference to the complement-clique phase: every case bounds the intermediate solution S 0 produced before that phase runs, and the phase enters the proof only through the observation that | S | ≥ | S 0 | . The phase therefore supplies structural insurance, not the source of the asymptotic guarantee. Its practical value is the ability to discover combinatorial candidates that the purely local Hvala / Caro-Wei / v max pipeline misses; iterative path compression inFastCliqueUFtogether with the per-vertex cap on clique candidates processed keep this contribution’s running time at O ( n 3 log n ) .
Remark 3
(Algebraic Rigor of the Small Degree Bound). When Δ ≤ n , the chain of inequalities leading to n Δ + 1 ≥ α 2 n is algebraically exact for all n ≥ 1 . Because the independence number is strictly bounded by the total vertex count ( α ≤ n ), dividing both sides by 2 n yields α 2 n ≤ n 2 . Combined with Δ + 1 ≤ n + 1 and the algebraic identity n n + 1 ≥ n 2 , we obtain:
| I CW | ≥ n Δ + 1 ≥ n n + 1 ≥ n 2 ≥ α 2 n ,
establishing that Subcase 2b holds unconditionally regardless of how close α is to n.
Remark 4
(Degree Heredity in Residual Subgraphs). Forcing v max into the candidate independent set and removing its closed neighborhood N [ v max ] leaves an induced subgraph G ′ on n ′ = n − 1 − Δ vertices. Because taking induced subgraphs monotonically non-increases vertex degrees, the maximum degree of G ′ satisfies Δ ( G ′ ) ≤ Δ . Running a minimum-degree greedy procedure on G ′ guarantees a residual independent set of size at least n ′ Δ ( G ′ ) + 1 ≥ n − 1 − Δ Δ + 1 . Including v max yields the exact closed-form lower bound:
| S v max | ≥ 1 + n − 1 − Δ Δ + 1 = ( Δ + 1 ) + ( n − 1 − Δ ) Δ + 1 = n Δ + 1 .

6. Runtime Analysis

Theorem 3
(Time complexity). Algorithm 1 runs in O ( n 3 log n ) worst-case time, where n = | V | , using adjacency lists.
Proof. 
Preprocessing costs O ( n + m ) . Hvala’s vertex cover is computed once in O ( n + m ) time [6]. Identifying v max costs O ( n ) .
The PureCaroWeiBaseline is implemented with a classic linear-time bucket structure. Each vertex and each edge is examined a constant number of times; the pointer retreats are bounded by the number of degree decrements, which is O ( m ) . Consequently the baseline runs in O ( n + m ) time.
The multi-restart loop runs at most n + 1 times. Each iteration costs O ( n + m ) , giving a total of O ( n ( n + m ) ) = O ( n 2 + n m ) = O ( n 3 ) on dense graphs.
For the complement-clique phase, a single triggering vertex u’s accumulated neighbourhood in the complement graph can fragment into as many as Θ ( deg G c ( u ) ) disjoint clique components in the worst case (for instance, if that neighbourhood consists mostly of disjoint pairs rather than one large clique), and each such component requires its own call to MaximizeSolution, at O ( n + m ) apiece. Without a cap, summing this over up to n triggering vertices, with deg G c ( u ) = O ( n ) and m = O ( n 2 ) on dense graphs, would give O ( n · n · ( n + n 2 ) ) = O ( n 4 ) . We verified this fragmentation is a real, not merely theoretical, concern: on the Kneser graph K ( 16 , 3 ) ( n = 560 ), the number of clique components returned per triggering vertex averaged 9.05 and reached a maximum of 12, across all 560 vertices (Section 7).
The cap directly bounds the quantity that would otherwise be unbounded: for every triggering vertex, at most ℓ = max ( 2 , ⌊ log 2 n ⌋ ) clique components are processed, regardless of how many To_Sets actually returns. This gives at most O ( n log n ) total calls to MaximizeSolution across the whole phase (up to n triggering vertices, each contributing at most O ( log n ) calls), each costing O ( n + m ) , for a total of O ( n log n · ( n + m ) ) = O ( n 2 log n + n m log n ) , which is O ( n 3 log n ) on dense graphs. This is the term that dominates the algorithm’s overall running time.
Adding the linear-time preprocessing and the O ( n 3 ) multi-restart loop, the complete algorithm runs in O ( n 3 log n ) worst-case time. □
We note directly that O ( n 3 log n ) is not a clean cubic bound, and that sorting by the largest clique first (rather than, say, capping arbitrarily) is a design choice intended to preserve most of the phase’s practical benefit while bounding its cost, on the reasoning that a single large clique is more likely to improve the incumbent than many small ones; Section 7 reports how well this holds up on the adversarial instances tested.

7. Empirical Study of the Approximation Ratio

We tested Esperanza across an expanded suite of 30 000 graph instances certified to exact optimality via Mixed-Integer Linear Programming (MILP). The exact independence number was obtained by solving
max ∑ v ∈ V x v subject to x u + x v ≤ 1 ∀ ( u , v ) ∈ E , x v ∈ { 0 , 1 } ,
and the reported ratio is ρ = OPT / | ALG | . The full run completed in 298.8 seconds (seed 42).

7.1. Results

Table 2 provides the empirical breakdown. The global worst-case ratio is 1.20 and the global mean ratio is 1.000018. Every instance in this suite has n ≤ 50 , so ℓ = max ( 2 , ⌊ log 2 n ⌋ ) ≤ 5 , and Section 7.2 shows directly, on a considerably larger adversarial graph, how often the cap actually binds there. These results confirm the cap does not hurt ratio quality on this suite; they should not be read as evidence of how often the cap binds in general, since instances this small rarely produce enough clique components per vertex to exceed it regardless.
Two specific adversarial instances (the G ( 16 , 0.509 ) graph with seed 272 and the G ( 20 , 0.9 ) graph with seed 0) are solved to exact optimality, confirming that bounding the number of clique candidates processed per vertex does not prevent the complement-clique phase from closing these particular gaps.

7.2. Runtime and Cap Diagnostics on a Larger Adversarial Instance

Because the 30 000-instance suite above is restricted to n ≤ 50 , where ℓ is at most 5, it cannot exercise the cap in the regime Theorem 3 is actually about. We therefore report direct measurements on the Kneser graph K ( 16 , 3 ) ( n = 560 ), where ℓ = ⌊ log 2 560 ⌋ = 9 .
Instrumenting the complement-clique phase directly on this instance: all 560 vertices trigger at least one successful Add (i.e., produce a non-trivial clique component). Across those triggering vertices, the number of clique components returned by To_Sets averaged 9.05 , with a maximum of 12; 280 of the 560 triggers ( 50.0 % ) exceeded the cap of 9, as Table 3 summarizes. The cap binds often on this instance, and the measured running time here is approximately 82 seconds; the algorithm remains correct throughout, returning the exact optimum of 105.
We report this instance to show the cap’s behavior concretely: unlike a smaller or more uniformly-structured graph, a large fraction of triggering vertices here actually produce more than ℓ clique components, so the cap is doing real, frequent work rather than sitting idle. The algorithm’s correctness is unaffected by how often the cap binds; what changes is only which candidates get evaluated once a vertex’s component count exceeds ℓ.

References

  1. Karp, R.M. Reducibility among Combinatorial Problems. In Complexity of Computer Computations; Miller, R.E.; Thatcher, J.W.; Bohlinger, J.D., Eds.; Plenum: New York, USA, 1972; pp. 85–103. [CrossRef]
  2. Halldórsson, M.M.; Radhakrishnan, J. Greed is good: Approximating independent sets in sparse and bounded-degree graphs. Algorithmica 1997, 18, 145–163. [CrossRef]
  3. Boppana, R.; Halldórsson, M.M. Approximating maximum independent sets by excluding subgraphs. BIT Numerical Mathematics 1992, 32, 180–196. [CrossRef]
  4. Karger, D.R.; Motwani, R.; Sudan, M. Approximate graph coloring by semidefinite programming. Journal of the ACM 1998, 45, 246–265. [CrossRef]
  5. Håstad, J. Clique is hard to approximate within n1-ϵ. Acta Mathematica 1999, 182, 105–142. [CrossRef]
  6. Vega, F. An Approximate Solution to the Minimum Vertex Cover Problem: The Hvala Algorithm. Gauge Freedom Journal 2026, 1. [CrossRef]
  7. Caro, Y. New results on the independence number. Technical report, Tel-Aviv University, 1979.
  8. Wei, V.K. A lower bound on the stability number of a simple graph. Technical Report Technical Memorandum 81-11217-9, Bell Laboratories, 1981.
  9. Vega, F. Esperanza: Approximate Independent Set Solver. https://pypi.org/project/esperanza, 2026. Python package version 0.1.5.
  10. Turán, P. On an extremal problem in graph theory. Matematikai és Fizikai Lapok 1941, 48, 436–452.
Table 1. Code metadata for the Esperanza package.
Table 1. Code metadata for the Esperanza package.
Nr. Code metadata description Metadata
C1 Current code version v0.1.5
C2 Permanent link to code/repository https://github.com/frankvegadelgado/esperanza
C3 Permanent link to Reproducible Capsule https://pypi.org/project/esperanza/
C4 Legal Code License MIT License
C5 Code versioning system used git
C6 Languages, tools, and services used Python
C7 Compilation requirements and dependencies Python ≥ 3.12, Hvala ≥ 0.1.2
Table 2. MILP-certified stress benchmark on 30 000 graph instances ( n ≤ 50 , seed=42) for Esperanza.
Table 2. MILP-certified stress benchmark on 30 000 graph instances ( n ≤ 50 , seed=42) for Esperanza.
Family Count Worst Ratio ( ρ ) Mean Ratio ( ρ )
regular 3 000 1.2000 1.0002
gnp 15 000 1.0000 1.0000
bipartite_gnp 6 000 1.0000 1.0000
star 2 000 1.0000 1.0000
complete 1 000 1.0000 1.0000
complete_bipartite 1 000 1.0000 1.0000
amplified 2 000 1.0000 1.0000
Overall Global 30 000 1.2000 1.000018
Table 3. Diagnostics on K ( 16 , 3 ) : measured directly, showing how often the cap actually binds on this instance.
Table 3. Diagnostics on K ( 16 , 3 ) : measured directly, showing how often the cap actually binds on this instance.
Quantity Value
Vertices (n) 560
Cap ℓ = ⌊ log 2 n ⌋ 9
Triggering vertices 560
Mean components per trigger 9.05
Maximum components per trigger 12
Triggers exceeding the cap 280 (50.0%)
True α ( G ) 105
Algorithm’s returned size 105
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.