Fractional optimization methods and fractal activation functions have recently emerged as two independent directions for improving neural network training. Fractional optimizers extend first-order optimization through fractional derivatives and memory effects, whereas fractal activations introduce multi-scale nonlinear representations based on self-similar Weierstrass- and Blancmange-type functions. Despite their common mathematical motivation, both approaches have largely been studied independently. Here, we investigate their interaction within a unified experimental framework. We evaluate fractional optimizer families on Ackley and Himmelblau benchmark surfaces, in standard form and with additive Weierstrass-type perturbations, and then in feed-forward neural networks with conventional and fractal activations on ten OpenML classification datasets. The comparison includes standard methods, Herrera-style optimizers, explicit and adaptive memory-based fractional optimizers, and representative literature methods. Results are analysed by accuracy, convergence, cost, robustness, and optimizer–activation interaction. Overall, fractional optimization and fractal activations show useful but selective pairings. Herrera-style fractional scaling performs well with selected fractal activations in network training, while Grünwald–Letnikov memory is most relevant on perturbed surfaces. Adaptive memory improves plain memory substitution in several cases, supporting controlled fractional memory as a promising direction rather than a universal replacement.