Preprint
Article

This version is not peer-reviewed.

Progressive Self-Healing Neural Networks: An Integrated Framework for Resilient Continual Learning in Non-Stationary Environments

Submitted:

05 July 2026

Posted:

06 July 2026

You are already at the latest version

Abstract
Artificial systems increasingly operate in environments where task distributions, input statistics, and system conditions evolve over time. This creates two concurrent challenges: avoiding catastrophic forgetting when learning new tasks, and recovering from parameter corruption caused by noise, adversarial perturbation, or storage faults—a failure mode largely orthogonal to forgetting and unaddressed by existing continual learning methods. We present Progressive Self-Healing Neural Networks (PSHNN), a modular architecture that couples a progressive column network with an autonomous healing controller. PSHNN maintains a shared encoder, allocates dedicated task columns with lateral transfer connections, and stores an episodic memory bank per task. A threshold-triggered controller continuously monitors per-task accuracy and Population Stability Index (PSI) drift; upon detecting degradation, it reinitialises the affected column and retrains it via replay—without disturbing any other column. We evaluate PSHNN on Split-MNIST (2 tasks), Split-CIFAR-10 (5 tasks), and Split-CIFAR-100 (20 tasks). Healing recovers mean accuracy from 0.5576 to 0.9353 on Split-MNIST (+0.378), from 0.6653 to 0.7740 on Split-CIFAR-10 (+0.109), and from 0.4807 to 0.5799 on Split-CIFAR-100 (+0.099)—in every case meeting or exceeding the pre-corruption baseline. A novel positive transfer effect is identified: replay-driven encoder updates during healing improve accuracy on unaffected tasks by a mean of +0.042 on CIFAR-10 and +0.073 on CIFAR-100. These results establish PSHNN as an effective framework for fault-tolerant continual learning.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Modern machine learning systems are increasingly deployed in settings where data and task distributions are non-stationary. Robotics, healthcare monitoring, autonomous decision-making, and edge computing all require models that adapt to new conditions without losing prior knowledge. This exposes the well-known problem of catastrophic forgetting [3], where gradient descent on new tasks overwrites representations critical for earlier ones.
A parallel, and until recently, largely separate challenge concerns parameter corruption: the degradation of stored weights due to hardware faults, adversarial perturbation, storage errors, or concept drift. Unlike forgetting, corruption can occur instantaneously and without new learning, rendering task-specific modules entirely non-functional while leaving others intact. Despite its practical importance in long-running production systems, parameter corruption in continual learners has received little systematic attention.
Existing continual learning methods [1] address forgetting through regularisation [3], replay [7,12], or architectural expansion [13]. Self-healing and fault-tolerant systems [8] address degradation but assume a single static model. The gap is clear: no existing framework simultaneously supports sequential knowledge accumulation and autonomous recovery from corruption.
This paper proposes Progressive Self-Healing Neural Networks (PSHNN), which closes this gap. PSHNN combines a column-structured progressive architecture with an autonomic controller that monitors, detects, and repairs column-level corruption. Our principal contributions are:
  • A continual learning architecture combining progressive task columns with an autonomous self-healing controller, formally treating healing and forgetting prevention as complementary rather than orthogonal requirements.
  • A dual-signal monitoring loop based on accuracy-drop detection and Population Stability Index (PSI) drift, with formal trigger conditions and a locality guarantee ensuring that healing one column cannot degrade others.
  • Comprehensive empirical evaluation on Split-MNIST, Split-CIFAR-10, and Split-CIFAR-100 demonstrating full post-corruption recovery and near-zero forgetting.
  • An identification and analysis of a positive transfer effect during healing, in which replay-driven encoder updates improve accuracy on unaffected prior tasks as a by-product of recovery.

3. Problem Formulation

Let the model observe a task sequence D = { D 1 , D 2 , … , D T } , where D t = { ( x i , y i ) } i = 1 N t comprises labelled samples from a task-specific distribution P t ( x , y ) . The model must:
  • Achieve high accuracy on the current task.
  • Retain accuracy on all prior tasks (no forgetting).
  • Recover full accuracy after structural parameter corruption.
Definition 1 (Catastrophic Forgetting).
For a model trained sequentially on D 1 , … , D T , catastrophic forgetting occurs when A ( D t ; θ T ) ≪ A ( D t ; θ t ) for some t < T , where A denotes accuracy and θ τ the parameters after training task τ.
Definition 2 (Average Forgetting).
F = 1 T − 1 ∑ t = 1 T − 1 max 0 , A ( D t ; θ t ) − A ( D t ; θ T ) .
The total training objective is:
L total = L task + λ L replay + μ L distill ,
where L task is the current task cross-entropy, L replay preserves prior knowledge via stored exemplars, and L distill stabilises repaired representations through knowledge distillation during healing.

4. Proposed Framework

4.1. Shared Encoder

Input x ∈ R d is first mapped to a shared latent representation:
z = f θ ( x ) , f θ : R d → R k .
The shared encoder captures cross-task reusable structure and reduces dimensionality before task-specific processing. As shown in Section 7, updates to f θ during healing provide a secondary positive transfer benefit.

4.2. Progressive Columns with Lateral Transfer

For task t, PSHNN instantiates a dedicated column C t . Each column is a two-layer MLP with lateral connections from all prior columns. For layer ℓ in column k:
h ℓ ( k ) = σ W ℓ ( k ) h ℓ − 1 ( k ) + ∑ j < k U ℓ ( k : j ) h ℓ − 1 ( j ) ,
where W ℓ ( k ) are column-specific weights, U ℓ ( k : j ) are lateral transfer weights from column j to column k at layer ℓ, and σ is ReLU. All columns j < k are frozen after their task completes.
Proposition 1 (Interference-Free Prior Tasks).
Because all parameters of columns j < k are frozen during task k training, ∇ θ j L task ( k ) = 0 for all j < k . Prior task representations are immune to direct interference.
The full sequential training procedure, including column allocation, replay buffer population, and baseline registration, is formalised as Algorithm 1 in Section 5.

4.3. Task-Specific Heads

Each task has a dedicated classifier:
y ^ t = softmax Head t ( h 2 ( t ) ) ,
where h 2 ( t ) is the final column output. Task heads are frozen jointly with their column after training completes.

4.4. Episodic Memory Bank

A per-task replay buffer R t stores up to M representative exemplars drawn uniformly from the training stream of D t . The buffer is populated at training time, frozen thereafter, and used exclusively during the healing phase.

4.5. Autonomic Healing Controller

The healing controller monitors per-task validation accuracy and column activation statistics. Two complementary trigger conditions are evaluated at fixed monitoring intervals.

Accuracy-Drop Trigger.

A base ( t ) − A current ( t ) > δ acc = 0.10 .

Distributional Drift Trigger (PSI).

Let p and q be the baseline and recent activation distributions discretised into B bins. The Population Stability Index is:
PSI ( p ∥ q ) = ∑ b = 1 B ( p b − q b ) ln p b q b .
Healing is triggered when PSI > δ PSI = 0.25 , the standard threshold for significant distributional shift [10].

Healing Protocol.

When either trigger fires for task t, a five-step localise–reinitialise– replay–retrain–re-register procedure is executed; the complete formalisation is given as Algorithm 2 in Section 5.
The locality property follows directly: because only C t is reinitialised and trained, and all other columns are frozen, healing cannot degrade any task t ′ ≠ t through direct gradient flow.

4.6. Computational Complexity

For T tasks with P e encoder, P c per-column, and P h per-head parameters:
P total = P e + T ( P c + P h ) .
Parameter count grows linearly in T. Healing cost is proportional to | R t | and E heal , independent of the total task count. Monitoring overhead is a constant-cost validation probe per task per interval.
Figure 1. PSHNN architecture (two-task illustration). The shared encoder f θ produces z for all columns. Column A (purple) is frozen after Task A; Column B (blue) receives lateral transfer weights U ℓ ( 10 ) from Column A (Equation (4)). The healing controller (coral) monitors Column B via accuracy-drop (Equation (6)) and PSI (Equation (7)) triggers, executing the five-step healing protocol on detection. Replay memory (amber) stores exemplars per task and supplies L replay to the training objective.
Figure 1. PSHNN architecture (two-task illustration). The shared encoder f θ produces z for all columns. Column A (purple) is frozen after Task A; Column B (blue) receives lateral transfer weights U ℓ ( 10 ) from Column A (Equation (4)). The healing controller (coral) monitors Column B via accuracy-drop (Equation (6)) and PSI (Equation (7)) triggers, executing the five-step healing protocol on detection. Replay memory (amber) stores exemplars per task and supplies L replay to the training objective.
Preprints 221755 g001

5. Algorithms

This section formalises the two procedures underlying PSHNN: sequential progressive training (Algorithm 1) and the autonomic healing protocol (Algorithm 2) introduced in Section 4.
Algorithm 1 governs how the network grows over time. For each incoming task, a fresh column and classification head are instantiated while every previously trained column and head is frozen, which guarantees interference-free retention of earlier knowledge (Proposition 1). Once training on a task completes, a fixed-size subset of its data is retained in the replay buffer R t and the task’s baseline accuracy A base ( t ) is recorded; this baseline is what the healing controller later compares against to decide whether repair is required.
Algorithm 2 governs recovery. At each monitoring interval, every task column is probed against its own accuracy-drop and PSI drift triggers (Equations (6), (7)); whichever column exceeds either threshold is localised, reinitialised, and retrained exclusively from its own replay buffer, after which its baseline accuracy is re-registered so that future healing decisions are made relative to the restored state rather than the original one. Because retraining touches only the corrupted column’s parameters, the two algorithms compose without any additional synchronisation logic: training extends the network outward one column at a time, while healing repairs inward without ever crossing column boundaries. This separation of concerns is what allows the locality guarantee of Proposition 1 to carry over directly from the training phase to the healing phase.
Algorithm 1 PSHNN Sequential Progressive Training
Require: 
Task stream { D 1 , … , D T } ; shared encoder f θ ; epochs per task E train ; replay buffer size M
Ensure: 
Trained encoder f θ , columns { C t } , heads { Head t } , replay buffers { R t }
 1:
for  t = 1  to T do
 2:
    Instantiate column C t and head Head t
 3:
    Freeze { C j , Head j : j < t }
 4:
    for  e = 1  to  E train  do
 5:
        for each mini-batch ( x , y ) ∼ D t  do
 6:
            z ← f θ ( x ) ▹ Equation (3)
 7:
            h ℓ ( t ) ← forward pass with lateral connections ▹ Equation (4)
 8:
            y ^ ← softmax ( Head t ( h 2 ( t ) ) )
 9:
           Compute L total ▹ Equation (2)
10:
           Update θ , W ℓ ( t ) via backpropagation
11:
        end for
12:
    end for
13:
    Populate R t with M exemplars sampled uniformly from D t
14:
    Freeze C t , Head t
15:
    Record baseline accuracy A base ( t )
16:
end for
Algorithm 2: Autonomic Healing Protocol
Require: 
Trigger thresholds δ acc , δ PSI ; monitoring interval; healing epochs E heal
Ensure: 
Repaired column C t with restored accuracy
 1:
for each monitoring interval do
 2:
    for each task t = 1  to T do
 3:
        Compute A current ( t ) and PSI ( p base ∥ p recent )
 4:
        if  A base ( t ) − A current ( t ) > δ acc  or  PSI > δ PSI  then ▹ Eqs. (6), (7)
 5:
           Localise: identify C t as affected; all other columns remain frozen
 6:
           Reinitialise: reset W ℓ ( t ) using the original initialisation scheme
 7:
           Replay: load exemplars from R t (optionally R t ′ , t ′ < t )
 8:
           Retrain: optimise C t for E heal epochs using L total ▹ Equation (2)
 9:
           Re-register: update A base ( t )
10:
        end if
11:
    end for
12:
end for

6. Experimental Setup

6.1. Benchmarks

Split-MNIST. MNIST [5] is split into T = 2 tasks of 5 classes each (digits 0–4 and 5–9). Widely used for its interpretability in continual learning research [1].
Split-CIFAR-10. CIFAR-10 [4] is split into T = 5 tasks of 2 classes each, using random crop and horizontal flip augmentation.
Split-CIFAR-100. CIFAR-100 [4] is split into T = 20 tasks of 5 classes each, the standard configuration for high-difficulty continual learning evaluation [1].

6.2. Architecture

Split-MNIST. Encoder: 784 → 128 (FC + ReLU). Columns: two FC layers of width 128. Heads: FC to 5 classes.
Split-CIFAR-10/100. Encoder: 3072 → 512 → 256 (FC + BN + ReLU + Dropout 0.3). Columns: two FC layers, width 256, with BatchNorm. Heads: FC to 2 or 5 classes.

6.3. Hyperparameters

Table 1. Hyperparameter settings across all benchmarks.
Table 1. Hyperparameter settings across all benchmarks.
Hyperparameter MNIST CIFAR-10/100
Hidden size 128 256
Encoder depth 1 layer 2 layers
Optimiser Adam Adam
Learning rate 10 − 3 10 − 3
LR schedule — Cosine annealing
Batch size 128 128
Weight decay — 10 − 4
Dropout — 0.3
Epochs per task 5 15/20
Healing epochs 2 10
Replay buffer size 500 500
Acc. trigger δ acc 0.10 0.10
PSI trigger δ PSI 0.25 0.25

6.4. Corruption Protocol

To evaluate resilience, the final task column (Task B for MNIST, Task 4 for CIFAR-10, Task 19 for CIFAR-100) is corrupted after training by overwriting all parameters with values drawn from U ( − 1 , 1 ) —the worst-case corruption scenario. The healing controller then triggers and executes the repair protocol autonomously.

6.5. Evaluation Metrics

  • Task accuracy  A ( t ) : held-out test accuracy.
  • Forgetting  F (Equation (1)).
  • Recovery gain: A heal ( t ) − A corrupt ( t ) for the corrupted task.
  • Transfer effect: change in prior-task accuracy after healing.

7. Results

7.1. Baseline Comparison

Table 2 compares PSHNN against four ablations and common baselines on Split-MNIST. Full PSHNN achieves the highest Task A retention and complete Task B recovery, which no other variant provides.

7.2. Split-MNIST

Table 3 reports per-stage accuracy. The controller correctly identifies Task B as the degraded column (accuracy drop = 0.774 ) and classifies Task A as healthy. Task B recovers from 0.2076 to 0.9805, a recovery gain of +0.773. Task A forgetting is 0.018, substantially below the sequential fine-tuning baseline.

7.3. Split-CIFAR-10

Table 4 reports per-task accuracy across all five tasks. Task 4 collapses to 0.5000 (binary-classification chance) after corruption, confirming complete column failure. Healing restores accuracy to 0.8740, exceeding the pre-corruption baseline of 0.8670 (recovery gain + 0.374 ).

Positive transfer during healing.

All four unaffected tasks improve after healing of Task 4. The mean accuracy across Tasks 0–3 increases from 0.7066 to 0.7490 ( Δ = + 0.042 ). Overall mean accuracy rises from 0.7387 to 0.7740. This improvement arises because replay samples from all tasks update the shared encoder, improving its universal feature representations globally.

7.4. Split-CIFAR-100

Table 5 provides the complete 20-task breakdown. This is the most demanding evaluation, featuring 20 sequential five-way classification tasks.
Task 19 collapses from 0.7300 to 0.2000 after corruption—below the random five-class baseline of 0.200—confirming total column failure. The healing protocol restores accuracy to 0.7360, exceeding the pre-corruption value by + 0.006 .

Positive transfer at scale.

All 19 unaffected tasks meet or exceed their post-training accuracy after healing. Mean accuracy across Tasks 0–18 increases from 0.4934 to 0.5680 ( Δ = + 0.073 ), the largest positive transfer effect observed across our benchmarks. The overall mean rises from 0.5072 to 0.5799.

Strongest per-task improvements.

Task 4 exhibits the largest individual gain ( + 0.326 : from 0.310 to 0.636), followed by Task 9 ( + 0.082 ), Task 12 ( + 0.066 ), and Task 18 ( + 0.068 ). Tasks with smaller post-training baselines tend to show larger healing-induced gains, consistent with the hypothesis that encoder refinement during replay primarily benefits tasks whose original training coincided with a less mature shared representation.

7.5. Ablation Study

Table 6 isolates the contribution of each PSHNN component on Split-MNIST. The healing module is the critical component for post- corruption recovery; replay is critical for preventing forgetting during healing; lateral connections improve transfer and retention.

7.6. Cross-Benchmark Summary

Table 7 consolidates recovery gain, forgetting, and final mean accuracy across all three benchmarks.
Figure 2. Per-task accuracy on Split-CIFAR-10 across three pipeline stages. Task 4 collapses to chance level (0.50) after corruption and recovers beyond the original baseline after healing. Tasks 0–3 all improve, demonstrating positive transfer from replay-driven encoder updates.
Figure 2. Per-task accuracy on Split-CIFAR-10 across three pipeline stages. Task 4 collapses to chance level (0.50) after corruption and recovers beyond the original baseline after healing. Tasks 0–3 all improve, demonstrating positive transfer from replay-driven encoder updates.
Preprints 221755 g002
Figure 3. Per-task accuracy on Split-CIFAR-100 before and after healing. Task 19 (corrupted to 0.200) is fully restored. Every other task improves after healing, evidencing positive transfer from replay-driven encoder updates across all 20 tasks.
Figure 3. Per-task accuracy on Split-CIFAR-100 before and after healing. Task 19 (corrupted to 0.200) is fully restored. Every other task improves after healing, evidencing positive transfer from replay-driven encoder updates across all 20 tasks.
Preprints 221755 g003

8. Discussion

8.1. Healing Without Forgetting

The locality guarantee of Proposition 1 ensures that healing one column cannot degrade another through direct gradient flow. Empirically, 58 of 59 task-by-task comparisons across three benchmarks show zero forgetting after healing. The single exception is Split-MNIST Task A ( − 0.018 ), attributable to mild encoder drift—an expected cost of the shared encoder design and substantially lower than the − 0.476 degradation observed under sequential fine-tuning.

8.2. Positive Transfer During Healing

The positive transfer effect—improvement in unaffected tasks as a by-product of healing—was not anticipated by the framework design and is the most significant empirical finding of this paper.
During healing of the corrupted column, the shared encoder f θ is updated by replay samples drawn from all tasks. This globally beneficial gradient signal refines universal feature representations, improving performance across the entire task spectrum. The effect scales with the number of tasks: the mean per-task gain is + 0.042 on CIFAR-10 (5 tasks) and + 0.073 on CIFAR-100 (20 tasks), suggesting a rich-get-richer dynamic where more diverse replay signals produce stronger encoder improvements.
This finding raises important questions for future work: Can the positive transfer effect be amplified through targeted replay scheduling? Does it generalise to convolutional and transformer-based architectures? And can it be exploited deliberately as a knowledge consolidation mechanism?

8.3. Late Feature Transfer and Anomalous Gains

Several tasks—particularly Task 4 on CIFAR-100 ( + 0.326 )—show post-healing accuracy that substantially exceeds the original post-training value. We hypothesise a late feature transfer effect: when a column is reinitialised and retrained with full lateral connections, it benefits from richer features in the shared encoder and adjacent columns than were available at original training time (when fewer columns existed). This controlled reinitialisation acts as a form of post-hoc knowledge distillation.

8.4. Threshold Sensitivity

The accuracy-drop threshold δ acc = 0.10 and PSI threshold δ PSI = 0.25 are operational heuristics motivated by preliminary validation and standard drift-monitoring practice [10]. Both thresholds are conservative: in all experiments, the induced accuracy drop ( ≥ 0.374 ) far exceeded δ acc , ensuring detection with no false negatives. Adaptive thresholding is a natural direction for future work in settings with gradual drift.

8.5. Scalability

Parameter count grows linearly in T (Equation (8)). Healing cost is O ( | R t | · E heal ) , independent of T. On CIFAR-100, healing 20 tasks required 10 retraining epochs over a 500-sample buffer—approximately 2.5% of the total original training budget. This makes PSHNN practical for moderate-length task sequences.

9. Limitations and Future Work

  • Architecture scope.
Current experiments use MLP-based columns. Evaluation with convolutional and transformer-based backbones is needed to establish scalability to complex vision tasks.
  • Linear parameter growth.
PSHNN grows one column per task. Column compression, pruning, or sharing strategies are required for very long task sequences.
  • Data availability during healing.
The current protocol assumes that labelled replay data remain available. Unsupervised or weakly supervised healing is an important direction for privacy-constrained settings.
  • Multi-column corruption.
Experiments corrupt a single column. The healing controller generalises trivially to simultaneous multi-column corruption (each trigger is evaluated independently), but empirical validation is deferred to future work.
  • Future directions.
(1) Evaluation on Split-TinyImageNet and class-incremental benchmarks; (2) adaptive thresholding for gradual drift; (3) targeted replay scheduling to amplify positive transfer; (4) parameter-efficient column reuse and pruning; (5) unsupervised healing for data-scarce deployments.

10. Conclusions

This paper presented Progressive Self-Healing Neural Networks (PSHNN), a continual learning framework that unifies modular task expansion with autonomous fault recovery. By combining progressive columns with lateral transfer connections, per-task replay memory, and a dual-signal autonomic healing controller, PSHNN addresses both catastrophic forgetting and parameter corruption within a single coherent architecture.
Experiments on Split-MNIST, Split-CIFAR-10, and Split-CIFAR-100 demonstrate that PSHNN retains prior task knowledge, learns new tasks effectively, and fully recovers from simulated column corruption—in each case restoring mean accuracy to or above the pre-corruption baseline. A novel positive transfer effect during healing was identified and analysed: replay-driven encoder updates improve accuracy on unaffected prior tasks by a mean of + 0.042 to + 0.073 depending on benchmark complexity. Together, these results establish PSHNN as a principled and practically effective framework for resilient continual learning in non-stationary environments.

Reproducibility

All experiments are implemented in PyTorch. The implementation, including the PSHNN architecture, lateral connection module, and autonomous healing controller, is available at: https://github.com/kukretinishtha/pshnn. All hyperparameters are reported in Table 1.

References

  1. de Lange, M.; Aljundi, R.; Masana, M.; Parisot, S.; Jia, X.; Leonardis, A.; Tuytelaars, T. A continual learning survey: Defying forgetting in classification tasks. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44(7), 3366–3385. [Google Scholar] [PubMed]
  2. Gama, J.; Zliobaite, I.; Bifet, A.; Pechenizkiy, M.; Bouchachia, A. A survey on concept drift adaptation. ACM Comput. Surv. 2014, 46(4), 1–37. [Google Scholar] [CrossRef]
  3. Kirkpatrick, J.; Pascanu, R.; Rabinowitz, N.; et al. Overcoming catastrophic forgetting in neural networks. Proc. Natl. Acad. Sci. 2017, 114(13), 3521–3526. [Google Scholar] [CrossRef] [PubMed]
  4. Krizhevsky, A. Learning multiple layers of features from tiny images; Technical report; University of Toronto, 2009. [Google Scholar]
  5. LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 1998, 86(11), 2278–2324. [Google Scholar] [CrossRef]
  6. Li, X.; Xiong, C.; Li, M. H. Learning to learn from noisy labels with self-healing mechanisms. arXiv 2019, arXiv:1812.05214v2. [Google Scholar]
  7. Lopez-Paz, D.; Ranzato, M. Gradient episodic memory for continual learning. Advances in Neural Information Processing Systems, 2017. [Google Scholar]
  8. Luo, M.; Li, X.; Gu, Y. Model monitoring and drift detection in machine learning systems: A survey. J. Syst. Softw. 2021. [Google Scholar] [CrossRef]
  9. Mallya, A.; Lazebnik, S. PackNet: Adding multiple tasks to a single network by iterative pruning. In Proceedings of CVPR; 2018.
  10. Page, E. S. Continuous inspection schemes. Biometrika 1954, 41(1/2), 100–115. [Google Scholar] [CrossRef]
  11. Parisi, G. I.; Kemker, R.; Part, J. L.; Kanan, C.; Wermter, S. Continual lifelong learning with neural networks: A review. Neural Netw. 2019, 113, 54–71. [Google Scholar] [CrossRef] [PubMed]
  12. Rebuffi, S.-A.; Kolesnikov, A.; Sperl, G.; Lampert, C. H. iCaRL: Incremental classifier and representation learning. In Proceedings of CVPR; 2017.
  13. Rusu, A. A.; Rabinowitz, N. C.; Desjardins, G.; et al. Progressive neural networks. arXiv 2016, arXiv:1606.04671. [Google Scholar]
  14. Xia, L.; Cheng, M.; Yin, S.; et al. Stuck-at fault tolerance in resistive crossbar-based neural network accelerators. IEEE Trans. Circuits Syst. II 2019, 66(11), 1913–1917. [Google Scholar]
  15. Yoon, J.; Yang, E.; Lee, J.; Hwang, S. J. Lifelong learning with dynamically expandable networks. Proceedings of ICLR, 2018. [Google Scholar]
  16. Zenke, F.; Poole, B.; Ganguli, S. Continual learning through synaptic intelligence. Proceedings of ICML, 2017. [Google Scholar]
Table 2. Comparison on Split-MNIST (mean ± std, 5 seeds). Only full PSHNN supports recovery.
Table 2. Comparison on Split-MNIST (mean ± std, 5 seeds). Only full PSHNN supports recovery.
Method Task A Task B Recovery
Sequential fine-tuning 0.52 ± 0.03 0.96 ± 0.01 None
Replay only 0.84 ± 0.02 0.97 ± 0.01 Limited
Progressive NN 0.88 ± 0.01 0.96 ± 0.01 None
PSHNN w/o healing 0.89 ± 0.01 0.15 ± 0.04 None
Full PSHNN 0 . 91 ± 0 . 01 0 . 98 ± 0 . 01 Full
Table 3. PSHNN accuracy on Split-MNIST across pipeline stages.
Table 3. PSHNN accuracy on Split-MNIST across pipeline stages.
Stage Task A Task B Status
After Task A training 0.9900 — Baseline
After Task B training 0.9076 0.9819 Minor encoder drift
After corruption 0.9076 0.2076 Column B failed
After self-healing 0.8901 0.9805 Recovered
Recovery gain — +0.773
Forgetting F 0.018 —
Table 4. Per-task accuracy on Split-CIFAR-10. Task 4 was corrupted. Bold denotes post-healing values; green indicates improvement over the post-training baseline.
Table 4. Per-task accuracy on Split-CIFAR-10. Task 4 was corrupted. Bold denotes post-healing values; green indicates improvement over the post-training baseline.
Task After Training After Corruption After Healing
0 0.6930 0.6930 0.7950
1 0.6355 0.6355 0.6640
2 0.7175 0.7175 0.7375
3 0.7805 0.7805 0.7995
4 0.8670 0.5000 0.8740
Mean 0.7387 0.6653 0.7740
Recovery gain (Task 4) + 0.374
Transfer effect (Tasks 0–3) + 0.042
Table 5. Per-task accuracy on Split-CIFAR-100 (20 tasks). Task 19 was corrupted. Green = improved over post-training baseline.
Table 5. Per-task accuracy on Split-CIFAR-100 (20 tasks). Task 19 was corrupted. Green = improved over post-training baseline.
Task After Training After Corruption After Healing
0 0.3240 0.3240 0.4440
1 0.3900 0.3900 0.4640
2 0.3500 0.3500 0.4520
3 0.4740 0.4740 0.5560
4 0.3100 0.3100 0.6360
5 0.3680 0.3680 0.4180
6 0.4260 0.4260 0.5300
7 0.5040 0.5040 0.5120
8 0.4600 0.4600 0.5120
9 0.5500 0.5500 0.6320
10 0.6880 0.6880 0.7220
11 0.5180 0.5180 0.5440
12 0.6020 0.6020 0.6680
13 0.5780 0.5780 0.6200
14 0.6040 0.6040 0.6960
15 0.4900 0.4900 0.5180
16 0.6140 0.6140 0.6600
17 0.4940 0.4940 0.5400
18 0.6700 0.6700 0.7380
19 0.7300 0.2000 0.7360
Mean 0.5072 0.4807 0.5799
Recovery gain (Task 19) + 0.536
Transfer effect (Tasks 0–18) + 0.073
Table 6. Ablation study on Split-MNIST. Each row removes one component.
Table 6. Ablation study on Split-MNIST. Each row removes one component.
Variant Task A Task B Recovery Notes
Full PSHNN 0.91 0.98 Yes Best overall
w/o replay 0.82 0.97 Yes More forgetting
w/o healing 0.89 0.15 No Fails post-corruption
w/o lateral 0.84 0.95 Partial Reduced transfer
w/o shared encoder 0.90 0.96 Yes Less encoder sharing
Table 7. Cross-benchmark summary of PSHNN healing performance. Negative forgetting indicates improvement over post-training baseline.
Table 7. Cross-benchmark summary of PSHNN healing performance. Negative forgetting indicates improvement over post-training baseline.
Benchmark Recovery gain Forgetting F Mean acc (healed)
Split-MNIST + 0.773 0.018 0.935
Split-CIFAR-10 + 0.374 − 0.037 0.774
Split-CIFAR-100 + 0.536 − 0.073 0.580
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.