Submitted:
19 August 2025
Posted:
19 August 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
- (1)
- Formalization: We introduce a rigorous view of website activity as a nonlinear dynamical system, enabling the use of chaos-theoretic tools for security analysis.
- (2)
- Prototype implementation: We design and implement a detection pipeline that computes Lyapunov exponents and entropy measures from browser traces, requiring no access to code or network payloads.
- (3)
- Empirical evidence: We demonstrate that these chaos-based metrics significantly separate benign from malicious domains, opening a new pathway for dynamic web security beyond lexical and code-based heuristics.
2. Related Work and Background
3. Threat Model and Assumptions
4. System as a Discrete Dynamical Model
5. Methodology and Illustrative Example
5.1. Algorithmic Procedure
- (1)
- Feature collection. Runtime characteristics such as CPU usage, memory allocation, DOM size, and the number of active scripts are recorded at regular intervals. This produces a multidimensional time series representing the system trajectory.
- (2)
- Chaos quantification. Two trajectories are generated by replaying the same link with slightly different initialization delays. Their divergence is used to estimate the largest Lyapunov exponent (sensitivity to initial conditions) and a tractable entropy measure H (information complexity).
- (3)
- Verdict derivation. The pair is evaluated against simple thresholds to produce a classification into Safe, Caution, or Unsafe. This provides an interpretable signal that can be integrated into security pipelines.
| Algorithm 1:Link Safety Analyzer using Lyapunov Exponent and Entropy |
|
5.2. Illustrative Example
5.3. Malicious Link Example
- ✗Critical error: the domain could not be resolved.
- ▲Final verdict:, link classified as UNSAFE.
6. Results
6.1. Dataset Description
6.2. Metrics Distribution
7. Why Entropy Signals Maliciousness
7.1. The Entropy We Use: Finite-Time Divergence Entropy
7.2. Link to Chaos Theory: KS Entropy and Lyapunov Exponents
- If : average separation grows exponentially (), indicating positive Lyapunov behavior and nonzero information production — a hallmark of chaotic dynamics.
- If : average separation contracts (), indicating stable dynamics with zero KS entropy in the limit.
7.3. Why Malicious Links Drive
7.4. Machine-Learning Perspective: Uncertainty, Predictability, and Complexity
- (1)
- Predictive Uncertainty. Consider a forecaster that predicts the next state increment . Higher entropy rate corresponds to larger irreducible predictive error (higher conditional variance), which aligns with traces being harder to predict and thus riskier.
- (2)
- Information Bottleneck / Minimum Description Length (MDL). Sequences with higher entropy rate require longer codes (higher stochastic complexity). From an anomaly-detection viewpoint, such sequences are penalized by MDL and naturally flagged as atypical. Our H captures this compressibility gap operationally through divergence growth.
- (3)
- Ensemble Disagreement as Epistemic Signal. In practice, two nearby runs act like a tiny ensemble under perturbation. Persistent growth of disagreement (increasing ) indicates model mismatch and non-smooth dynamics, correlating with adversarial manipulation; contraction suggests regularity and predictability.
7.5. Decision-Theoretic Mapping of H to Verdicts
7.6. Practical Notes and Caveats
- Finite-time effects. is a finite-window estimator; short windows or heavy noise can blur the sign. We mitigate via smoothing, robust regression for the Lyapunov fit, and repetition.
- Normalization. Because uses a ratio , it is invariant to absolute state scaling; negative values arise naturally when perturbations decay.
- Confounders. Highly interactive yet benign dashboards may induce H close to zero or mildly positive. We combine H with the largest Lyapunov exponent and conservative thresholds to reduce false positives.
8. Comparative Analysis with Sequential Deep Learning Approaches
9. Conclusion
10. Future Work
Data Availability Statement
Conflicts of Interest
Appendix A *
Appendix Description of the Implementation
- (1)
- Launches a headless Chrome browser and records runtime telemetry, including CPU usage, memory allocation, Document Object Model (DOM) size, and number of active scripts.
- (2)
- Executes the same URL twice with slight initialization perturbations to generate paired trajectories.
- (3)
- Computes the largest Lyapunov exponent () from the divergence of the trajectories, indicating sensitivity to initial conditions.
- (4)
- Estimates a finite-time divergence entropy (H), serving as a proxy for Kolmogorov–Sinai entropy and quantifying unpredictability in execution traces.
- (5)
- Produces a unified visualization showing divergence curves, computed metrics, and a safety verdict.
Appendix Classification Output
- SAFE: and (stable and predictable behavior),
- CAUTION: borderline cases near zero values,
- UNSAFE: or (unstable and malicious behavior).
Appendix Intended Use
References
- A. Herzberg and A. Jbara. Security and identification indicators for browsers against spoofing and phishing attacks. ACM Transactions on Internet Technology, 8(4):1–36, 2008.
- PhishTank Community. PhishTank: Online Phishing Database. Available at: https://www.phishtank.com, accessed 2025.
- J. Ma, L. K. J. Ma, L. K. Saul, S. Savage, and G. M. Voelker. Beyond blacklists: Learning to detect malicious websites from suspicious URLs. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 347–356, 2009.
- M. Antonakakis, R. M. Antonakakis, R. Perdisci, Y. Nadji, N. Vasiloglou, S. Abu-Nimeh, W. Lee, and D. Dagon. From throw-away traffic to bots: Detecting the rise of DGA-based malware. In Proceedings of the 21st USENIX Security Symposium (USENIX Security 12), pages 49–64, 2012.
- R. Perdisci, W. R. Perdisci, W. Lee, P. Rittenmeyer, and C. Kruegel. Automating the analysis of malicious websites. In Proceedings of the 17th ACM Conference on Computer and Communications Security (CCS), pages 437–448, 2010.
- C. Seifert, R. C. Seifert, R. Steenson, and I. Welch. Watching the watchers: Detecting evasive malicious web pages. In Proceedings of the 1st USENIX Workshop on Large-Scale Exploits and Emergent Threats (LEET), 2008.
- M. Cova, C. M. Cova, C. Kruegel, and G. Vigna. Detection and analysis of drive-by-download attacks and malicious JavaScript code. In Proceedings of the 19th International Conference on World Wide Web (WWW), pages 281–290, 2010.
- E. Ott. Chaos in Dynamical Systems. Cambridge University Press, 2nd edition, 2002.
- B. Krumnow, H. Jonker, and S. Karsch. Analysing and strengthening OpenWPM’s reliability. arXiv preprint arXiv:2205.08890, 2022. [CrossRef]
- V. Le Pochat, T. V. Le Pochat, T. Van Goethem, S. Tajalizadehkhoob, M. Korczyński, and W. Joosen. Tranco: A research-oriented top sites ranking hardened against manipulation. In Proceedings of the Network and Distributed System Security Symposium (NDSS), 2019.
- abuse.ch. URLhaus: Malware URL exchange. https://urlhaus.abuse.ch, accessed 2025.
- PhishTank. PhishTank API information. https://phishtank.org/api_info.php, accessed 2025.
- H. Ghalechyan, A. Shahverdyan, A. Exposito-Jimenez, D. Gulin, R. Musleh, J. Palanca, M. A. Fokkar, T. Peltonen, A. Vihavainen, and A. Gyrard. Phishing URL detection with neural networks: an empirical study. Scientific Reports, 14:25134, 2024.
- C. Çatal, B. Giray, and A. A. Aydın. Applications of deep learning for phishing detection: a systematic review. Knowledge and Information Systems, 64(6):1457–1500, 2022.
- H. Bensaoud, J. K. Kalita, and Y. Bensaoud. A survey of malware detection using deep learning. arXiv preprint arXiv:2407.19153, 2024.
- F. Casino, D. Hurley-Smith, J. Hernandez-Castro, and C. Patsakis. Not on my watch: ransomware detection through classification of high-entropy file segments. Journal of Cybersecurity, 11(1):tyaf009, 2025.
- M. Williams et al. Entropy-based network traffic analysis for efficient ransomware detection. TechRxiv preprint, 2024.
- M. T. Hoang and M. Özlük. A simple approach for global asymptotic stability of a malware model via Lyapunov functions. Mathematical Foundations of Computing, 7(4):559–574, 2024.
- K. V. Nithya, S. Das, and A. Abraham. Delayed dynamics analysis of SEI2RS malware propagation models in networks. Computer Networks, 248:110481, 2024.
- S. Englehardt and A. Narayanan. OpenWPM: An automated platform for web privacy measurement. Technical Report, Princeton University, 2015.
- S. Gopali, A. S. S. Gopali, A. S. Namin, F. Abri, and K. S. Jones. The Performance of Sequential Deep Learning Models in Detecting Phishing Websites Using Contextual Features of URLs. arXiv:2404.09802, 2024.
- L. Herrmann, M. Granz, and T. Landgraf. Chaotic dynamics are intrinsic to neural network training with SGD. Advances in Neural Information Processing Systems, 35:5219–5229, 2022.
- B. Chang, L. Meng, E. Haber, F. Tung, and D. Begert. Multi-level residual networks from dynamical systems view. arXiv preprint arXiv:1710.10348, 2017.
- L. S. Pontryagin. Mathematical Theory of Optimal Processes. Routledge, 2018.
- Z. Rafik and A. Humberto Salas. Chaotic dynamics and zero distribution: implications and applications in control theory for Yitang Zhang’s Landau Siegel zero theorem. European Physical Journal Plus, 139:217, 2024. [CrossRef]



| Aspect | Sequential Deep Learning (e.g., Gopali et al. [21]) | Chaos-Based Dynamics (this work) |
|---|---|---|
| Input signal | URL tokens and lexical features only | Runtime telemetry (CPU, memory, DOM events, scripts) |
| Learning style | End-to-end black-box classifiers (LSTM, BiLSTM, TCN, MHA) | Physics-inspired indicators: Lyapunov exponent and entropy H, with interpretable thresholds |
| Evasion resistance | Vulnerable to URL obfuscation, shorteners, and polymorphic generation | Resilient to runtime evasions (delays, packed scripts, redirect chains); detects instability in execution traces |
| Interpretability | Limited: model scores are opaque | High: SAFE; UNSAFE |
| Operational cost | Very fast inference (no rendering required) | Requires sandboxed rendering (60s window), but lightweight and content-agnostic |
| Novelty | Extends NLP methods to URL text | First to operationalize chaos theory (Lyapunov + entropy) for cybersecurity threat detection |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).