Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Probability and Statistics

Ercan Gürvit

Abstract: Modern non-stationary signal denoisers increasingly replace a global wavelet threshold by a \emph{detect-then-act} (gated) rule that first localises an artefact in the time--scale plane and then suppresses it selectively. Such gates couple sub-bands, and it has been unclear how to quantify, in a basis-intrinsic and interpretable way, \emph{how much} of the denoising performance is produced by that coupling. We introduce the \emph{interaction budget} $I=1-\sum_j S_j$ and prove that it equals the normalised $L^2$-distance of the performance functional to the space of band-additive functions; hence $I=0$ for any band-diagonal operator and $I>0$ exactly when the gate couples bands. We then (i) give an exact pairwise identity for a binary gate, a two-point lower bound, and a sharp iff-condition on the gate; (ii) identify a coupling-strength functional with the leading-order law $I(\kappa)=(\mathcal C/V_0)\kappa^2+o(\kappa^2)$; and (iii) bound the second-order HDMR/EMPR truncation error, which vanishes for single-trigger gates. Finally, for coloured noise---where the sub-band factors are correlated and $I$ alone conflates the two effects---an EMPR product-support decomposition \emph{separates} operator-induced from correlation-induced interaction, with a correlation-invariant operator signature and a permutation-based estimator. Controlled experiments confirm all results.

Article
Computer Science and Mathematics
Probability and Statistics

Daniel Rodriguez

Abstract: Background: With the proliferation of repeated measures data from wearable technology and apps, researchers need diverse methods to assess change, beyond common approaches such as Generalized Estimating Equations and Latent Growth Curve Modeling. One such method is Area Under the Curve (AUC). The purpose of this study was to assess the efficacy of AUC with a larger number of repeated measures using simulated data. Methods: We generated two samples of 21 hypothetical cyclists with 30 and 100 repeated measures of performance data (speed, time, and power) based on a Strava app segment, and calculated AUC using the trapezoid rule and definite integrals with the best fitting linear and a six-degree (sextic) polynomial. We then assessed the relations between time, speed, and power with all three calculations using bivariate correlations, multiple regression analysis, and a mediation analysis whereby power was hypothesized to predict a reduction in time indirectly through speed. Results: There was little difference in relative performance comparing the three calculation methods in the two samples. However, there was a difference in mediation results when comparing the two samples with full mediation when using 100 replications and only partial mediation with 30 replications. Conclusions: The results of this study suggest that Area Under the Curve may be a viable method for assessing change when dealing with datasets including many repeated measures such as those acquired from wearable technology and apps. However, AUC may be more sensitive to performance changes with a larger number of replications. Future studies should assess AUC with larger numbers of repeated measures, additional equations, and non-simulated data.

Article
Computer Science and Mathematics
Probability and Statistics

Ben Magabashele Malope

,

Retius Chifurira

,

Temesgen Zewotir

,

Knowledge Chinhamu

Abstract: Macroeconomic relationships in African economies are often nonlinear and heterogeneous, limiting the adequacy of conventional linear models for analysing and forecasting GDP growth and inflation. This study applies a semiparametric Panel Generalized Additive Model (Panel GAM) to quarterly panel data for 53 African economies over 2005Q1–2025Q4, estimating smooth nonlinear effects while accounting for temporal variation and country-specific heterogeneity. The estimated smooth terms indicate statistically significant nonlinear relationships for both outcomes. The model explains a larger share of the variation in the inverse hyperbolic sine-transformed Consumer Price Index than in GDP growth (adjusted R² = 0.671; deviance explained = 67.7% versus adjusted R² = 0.236; deviance explained = 25.1%), suggesting that the selected macroeconomic variables are more strongly associated with inflation dynamics than with economic growth during the study period. Out-of-sample results indicate that the framework can generate forecasts for both outcomes, while providing interpretable smooth functions. These results may inform policy analysis by highlighting the relevance of nonlinear responses and cross-country heterogeneity, particularly for inflation. The study contributes an interpretable semiparametric panel approach that complements Panel Vector Autoregressive and Panel Multi-Output Gaussian Process Regression models.

Article
Computer Science and Mathematics
Probability and Statistics

Stefano Barone

,

Santo Orlando

,

Antonino Paladino

Abstract: Introduction. Forest fires are complex phenomena causing considerable damage to the environment, habitat destruction, soil erosion, greenhouse gas emissions, and biodiversity loss. They are increasing globally, with extreme events becoming more frequent and destructive. Understanding their root causes and influencing factors is crucial. Methods. This work focuses on analyzing data of forest fires that occurred in the period 2010-2023 in Sicily, an Italian region and big island with special orographic characteristics and substantial agricultural and forestry-pastoral activities. The methods concern a careful extraction of data by using QGIS software and official databases and their appropriate statistical analysis. Results. A definition of forest fire risk, coherent with the literature, is here formulated, and a risk ranking and classification of the Sicilian municipalities is so obtained. Risk factors are elicited by expert advice, and their significance is determined via multiple regression analysis with a transformed dependent variable. Conclusions. The work shows an optimal balancing between ecological perspective and operational risk management. Forest fire data collection empowerment is highlighted, such as fire-starting location and total damage caused by each fire event. The study allows optimally distributing the regional budget for forest fire prevention among the municipalities.

Article
Computer Science and Mathematics
Probability and Statistics

Harry Vite-Cevallos

,

Omar Ruiz-Barzola

,

Purificación Galindo-Villardón

Abstract: Understanding tourist behaviour requires analytical frameworks capable of capturing both symmetric relationships among motivational constructs and asymmetric causal effects on behavioural intentions. Conventional segmentation approaches, particularly those based on structural equation modelling, primarily estimate directional relationships and provide limited insight into the multivariate interaction structures underlying tourist decision-making. To address this limitation, this study proposes an integrated analytical framework that combines GH-Biplot, Co-Inertia Analysis (COIA), STATICO, and Neutrosophic Psychology to analyse tourist motivations under uncertainty. The proposed framework was applied to a sample of 400 tourists participating in poverty-reducing tourism research. Measurement models were validated using confirmatory factor analysis and Partial Least Squares Structural Equation Modelling (PLS-SEM), while the proposed multivariate approach was employed to identify latent symmetric and asymmetric structures linking behavioural intentions, motivational constructs, and personal values. Results show that biospheric values exhibit the strongest association with intentions to participate in poverty-reducing tourism. More importantly, the proposed framework reveals multivariate relationships and behavioural patterns that remain hidden when conventional asymmetric causal models are applied independently. The incorporation of neutrosophic psychology further extends the analysis by explicitly representing indeterminacy in tourist motivations through truth, falsity, and indeterminacy components. The study contributes by introducing a novel analytical framework that integrates complementary multivariate techniques to improve tourism segmentation, enhance the interpretation of complex behavioural relationships, and support evidence-based decision-making for sustainable tourism management under uncertainty.

Article
Computer Science and Mathematics
Probability and Statistics

Yi Chen

,

Rufeng Tang

,

Yuqiang Li

,

Niansheng Tang

Abstract: In satellite and space debris laser ranging, photon-counting time-of-flight sequences exhibit spatio-temporal echo coherence and deterministic orbital constraints. We propose an unsupervised framework exploiting this spatial-kinematic coupling to extract weak returns under high background noise. First, a fuzzy clustering regression employs a dynamic energy functional and cross-entropy-regularized photon attribution, guided by target motion priors. To enhance low signal-to-noise ratio sensitivity, we introduce a low-gradient sampling strategy that theoretically guarantees a signal-to-background ratio exceeding 1/2. Furthermore, a dual-stream autoencoder fuses orbital kinematic parameters and multi-scale echo densities via noise-adaptive latent gating. Validation on 88 satellite and 24 debris datasets achieves F1-scores of 0.73 and 0.93, respectively. The sampling strategy reduces computational latency by ∼5.2% with no loss in tracking precision. This label-free, physically grounded approach enables robust weak signal detection in ground-based photon-counting lidar.

Article
Computer Science and Mathematics
Probability and Statistics

A. Hilaire Nzokem

Abstract: This paper identifies and characterizes the Background Driving L\'evy Process (BDLP) associated with the Generalized Tempered Stable (GTS) distribution, a flexible seven-parameter family of infinitely divisible distributions with applications in physics and quantitative finance. We show that the corresponding BDLP is a finite-variation, infinite-activity Type B L\'evy process and derive its cumulants.

Article
Computer Science and Mathematics
Probability and Statistics

Ryan A. Peterson

,

Sarah M. Bird

,

Logan M. Harris

,

Patrick J. Breheny

,

Joseph E. Cavanaugh

Abstract: The concept of ranked sparsity, originally introduced in the context of penalized regression, arises in modeling applications when an expected disparity exists in the quality of information between different feature sets. Its presence can cause traditional and modern model selection methods to fail because such procedures commonly presume “covariate equipoise” — that each potential parameter is equally worthy of entering into the final model. However, this presumption does not always hold, especially in the presence of derived variables or with highly disparate feature sets (i.e., multi-modal data). For instance, when all possible interactions are considered as candidate predictors, the sheer number of them grossly inflates the number of false discoveries, resulting in unnecessarily complex and difficult-to-interpret models with many (truly spurious) interactions. In this work, we motivate a ranked sparsity extension to the Bayesian Information Criterion (RBIC) that requires a stronger level of evidence in order to allow certain variables (e.g. interactions vs main effects and genetic vs clinical covariates) into a model. We compare the performance of RBIC relative to competing methods for selecting polynomials and interactions in a simulation study and in two applications, showing that stepwise selection guided by RBIC produces better-predicting, more transparent models (with fewer false interactions) compared to existing alternatives.

Article
Computer Science and Mathematics
Probability and Statistics

Justice Yaw Effah

,

Gifty Duah

,

Eric Nyarko

,

Miriam Appiah

,

Natasha Adjoa Anderson

Abstract: Obesity is a major problem worldwide, particularly in Latin countries such as Mexico, Peru and Colombia. This study aims to evaluate five models (XGBoost, Random Forest, Multinomial Logistic Regression, Linear Discriminant Analysis (LDA) and Support Vector Machines (SVM)) from the statistical and machine learning fields to determine which of them is the most effective in classification of obesity levels in a multivariate classification model. The research uses a dataset of 2111 records consisting of 17 physiological and behavioral attributes which had to be preprocessed, such as one-hot encoding and multicollinearity filtering, and a powerful 10-fold cross validation methodology. It is shown that the results of ensemble methods are much better than those of the traditional parametric ones. XGBoost achieved the highest classification accuracy (97.54%), and an F_1-score of 0.9747, while Random Forest achieved the 2nd highest classification accuracy (95.64%). For all comparisons, McNemar's pairwise significance tests validated the superiority of XGBoost (with p<0.0001). Feature importance analysis was used to determine the most significant features as follows: weight, height, age and number of vegetables consumed per week (FCVC). While the study showcases the power of gradient boosting for capturing non-linear interactions in health data, the exclusive use of synthetic data (77% from SMOTE) poses drawbacks as it may overestimate performance and limit generalizability. The proposed ensemble architectures should be further evaluated on larger, all-nonsynthetic datasets to be clinically generalizable in the future.

Article
Computer Science and Mathematics
Probability and Statistics

C.S. Withers

Abstract: I give estimates of low bias for functions of moments. Let \( F(x) \)be a distribution on \( R^s \). Let \( F_n(x) \) be the empirical distribution of a random sample of size \( n \) from\( F(x) \). Given a functional \( F(x) \), \( E\ T(F_n) \)estimates \( T(F) \)with bias \( \sim n^{-1} \). (The bias is zero for a mean, but this is the exception.) The jackknife and bootstrap estimates only reduce this bias to \( \sim n^{-2} \), and are computationally intensive. I review the main two analytic methods to obtain an estimate of \( T(F) \) of bias \( \sim n^{-k} \)for \( k\leq 4 \)in terms of the functional derivatives of \( T(F) \). I give a chain rule for these derivatives when \( T(F)=g(U(F)) \) and \( g:R^q\rightarrow R \)is any given smooth function with finite partial derivatives at \( U(F)\in R^q \). I apply this to give an estimate of \( T(F) \) of bias \( \sim n^{-k} \)for \( k\leq 4 \), in terms of the derivatives of \( g \)and \( U(F) \). Examples include moment estimates and maximum likelihood estimates.

Article
Computer Science and Mathematics
Probability and Statistics

Paul A. Quaye

,

Andrew A. Neath

Abstract: There has been significant interest in the quantification of statistical evidence. This paper presents a statistical philosophy focused on quantifying statistical evidence derived from data, rather than following the traditional approach of presenting statistical methodology. Our focus will be on the reasoning underlying the methods, offering a fresh development of several familiar statistical approaches. We will also illustrate how a simple example requires considerably more detail. Our paper explores additional questions that analysts should pursue. We examine issues related to the determination of sample size, stopping rules, multiple hypotheses, and how effect sizes impact the quantification of statistical evidence.

Article
Computer Science and Mathematics
Probability and Statistics

Demetris Koutsoyiannis

Abstract: A novel axiomatic foundation of entropy has recently been proposed overcoming the limitations of classical and information-theoretic entropy foundations and eventually unifying probabilistic and physical entropy. A new set of postulates leads to a rigorous, uncertainty-based definition of entropy consistent with the principle of maximum entropy. Entropy is thus a purely stochastic concept quantifying uncertainty, thereby completing Kolmogorov’s probability system. Applied to gas thermodynamics, the new framework reproduces classical results and derives, rather than assumes, the laws of thermodynamics. In atmospheric applications, entropy maximization yields an isothermal state as the molecular equilibrium. Gravitation does not alter the isothermal state but distinguishes it from the isentropic one of macroscopic air parcels, whose motion drives the atmosphere away from equilibrium. Radiatively active gases, through interactions with shortwave and longwave radiation, sustain non-equilibrium vertical profiles of the atmospheric variables. Combined with the Stefan-Boltzmann law, these mechanisms provide a simple, parsimonious and coherent explanation of observed atmospheric behaviours and the climatic system. The framework highlights thermodynamics as emergent from stochastics, offering new insights into molecular uncertainty, emergence of macroscopic structures and radiation in shaping Earth’s climate. It also suggests a broader stochastic view of nature and atmospheric processes.

Article
Computer Science and Mathematics
Probability and Statistics

Sthitadhi Das

Abstract: Classical moment theory constitutes one of the fundamental pillars of probability and statistics, providing quantitative measures of location, dispersion, skewness, and higher-order characteristics of probability distributions. However, many contemporary problems in statistics, machine learning, finance, engineering, and data science involve nonlinear transformations of random variables, for which ordinary moments may not adequately capture the underlying stochastic behavior. This motivates the study of functional moments, defined as expectations of powers of transformed random variables. This chapter develops a comprehensive theoretical framework for functional moments of Gaussian random variables and introduces the concept of the Functional Moment Generating Function (FMGF), which serves as a unified generating mechanism for all functional moments. Beginning with a general formulation of functional moments, exact infinite-series representations are derived using Taylor expansions and the central moments of Gaussian distributions. These representations lead naturally to operator formulations, asymptotic approximations, Jensen-type inequalities, derivative-based bounds, Hölder inequalities, Lyapunov inequalities, and Lipschitz-type estimates. The chapter further introduces functional cumulants through the logarithm of the FMGF and establishes their relationship with transformed stochastic processes. Several numerical investigations involving polynomial, exponential, logarithmic, trigonometric, and logistic transformations are presented to illustrate the theoretical results. Comparative studies between exact functional moments and Taylor-series approximations demonstrate the accuracy and computational efficiency of the proposed framework. To highlight practical relevance, five real-world applications are examined, including uncertainty propagation in machine learning activation functions, reliability assessment in engineering systems, expected utility analysis in financial mathematics, nonlinear signal processing, and environmental risk modelling. Numerical examples, simulation studies, tables, and graphical illustrations are provided throughout the chapter. The proposed framework extends classical moment theory to a substantially broader setting and provides a unified methodology for studying nonlinear transformations of Gaussian random variables. Several future research directions are discussed, including multivariate functional moments, dependence-aware functional moment generating functions, functional cumulant theory, high-dimensional asymptotics, and applications in modern statistical learning.

Article
Computer Science and Mathematics
Probability and Statistics

Mohan D. Pant

,

Aditya Chakraborty

,

Jovanna A. Tracz

Abstract: Healthcare continuous data often deviate from normality, which can substantially increase the risk of making invalid inferences, given that many inferential statistical procedures rely on normality assumption. To obviate this issue, we propose a new family of non-normal distributions based on a linear combination of the quantile functions of standard logistic and uniform (0, 1) distributions. This new family of non-normal distributions is characterized by using the methods of L-moments, conventional moments, and percentiles. Its performance is compared among the three methods in the context of parameter estimation and data modeling. The results of Monte Carlo simulation and bootstrapping techniques indicate that the L-moment-based estimates of parameters of L-skewness and L-kurtosis are substantially less biased than their percentile-based estimates of left-right tail-weight ratio (a measure of skewness) and tail-weight factor (a measure of kurtosis), which in turn are superior to their moment-based counterparts of skewness and kurtosis, especially for small sample sizes and higher-order moments. On the other hand, the data modeling results indicate that the percentile-based fits of the proposed distributions provide slightly better approximations to real-world healthcare data than their L-moment-based counterparts, whereas both percentile- and L-moment-based methods are superior to their conventional moment-based counterparts.

Article
Computer Science and Mathematics
Probability and Statistics

Dimitri Volchenkov

Abstract: Institutional distrust is treated here not as a low value of trust but as a positive social disposition, the settled expectation that formal procedures and official explanations no longer carry their stated public meaning. The paper studies the consolidation of that disposition as a threshold event. Building on a stochastic trust-phase model, it applies the same multiplicative-noise mechanism to a delegitimating assertion, so the bounded state variable is the probability of adopting institutional distrust. A logit transformation maps the inherited nonlinear diffusion exactly onto Brownian motion with drift and yields closed-form first-passage formulas for the crossing of operational distrust thresholds. The endpoints of the bounded variable are limiting consolidated regimes rather than finite-time targets, so observable institutional failure is threshold passage and not literal absorption at zero trust. The drift-to-turbulence ratio fixes the shape of the crossing probabilities and the noise scale fixes the time scale. The same coordinate measures the distance between social layers facing one assertion. In United States partisan survey data this inter-layer logit distance is large and, on consolidated assertions, stationary, the empirical signature of a completed passage, while valence assertions reset with the change of incumbent.

Article
Computer Science and Mathematics
Probability and Statistics

Tristan Guillaume

Abstract:

Let \(X = \left( X_{t} \right)_{0 \leq t \leq T}\) be a real-valued continuous process. For a threshold \(a\), the sub-threshold time set \[E_{T}(a) = \{ t \in \lbrack 0,T\rbrack:X_{t} \leq a\}\] encodes several different threshold observables. The most elementary one is the cumulative occupation time \[A_{T}(a) = \int_{0}^{T}\mathbf{1}_{\{ X_{t} \leq a\}}\, dt.\] For a regular one-dimensional diffusion, the classical occupation density formula gives \[A_{T}(a) = \int_{- \infty}^{a}\frac{L_{T}^{y}(X)}{\sigma^{2}(y)}\, dy,\] and hence \[\frac{\partial A_{T}}{\partial a}(a) = \frac{L_{T}^{a}(X)}{\sigma^{2}(a)}.\] Thus additive threshold occupation admits a local-time sensitivity calculus. In the terminology of barrier contracts, this additive clock is the cumulative, non-resetting Parisian clock, also called the Parasian clock. The purpose of this paper is to contrast this additive/Parasian regime with the behavior of resetting Parisian burst functionals. The connected components of \(E_{T}(a)\) represent sub-threshold episodes. We study in particular the longest burst \[M_{T}(a) = \sup\{|I|:I\text{ is a connected component of }E_{T}(a)\}.\] While \(A_{T}\) is locally controlled by local time, \(M_{T}\) is governed by the connectivity of the sub-threshold time set. We prove that \(M_{T}\) is monotone, that its supremum is attained, and that the weak-sublevel version is right-continuous with left limits, while the strict-sublevel version is its left-continuous regularization. The jump at a level is the increase in the maximal connected-component length produced by adjoining the level set. This gives a deterministic càdlàg/càglàd calculus for longest-burst profiles. For regular one-dimensional diffusions, this yields a sharp structural contrast. At deterministic levels which are almost surely not local-extreme values, the weak and strict longest bursts agree almost surely. Whenever the path has a unique interior maximum, the level-indexed longest-burst profile has a positive jump at the maximum level and is therefore not absolutely continuous. Brownian motion satisfies this criterion almost surely. We further identify the deterministic mechanism behind this instability: small threshold increases may fill short temporal bridges and merge large sub-threshold components. Finally, we show that the longest burst is exactly a one-sided continuous Parisian functional. This yields an exact Laplace-transform representation of its Brownian law through the Chesney--Jeanblanc-Picqué--Yor [1] Parisian transform, and an excursion-measure formulation in which local time enters only as the Itô excursion intensity. We also discuss smoothed burst statistics, moving thresholds, and diffusion examples. The paper is intended as a threshold-sensitivity comparison: local time controls cumulative Parasian occupation, whereas resetting Parisian burst observables are controlled by component mergers and excursion structure.

Article
Computer Science and Mathematics
Probability and Statistics

David Arango-Londoño

,

Delia Ortega-Lenis

,

Mauricio A. Mazo-Lopera

,

Paula Moraga

Abstract: Evaluating joint predictive performance for multivariate hydroclimatic models requires metrics that simultaneously assess marginal accuracy and cross-variable dependence recovery. Existing metrics – the Energy Score, Variogram Score, and their derivatives – do not adapt to the structural complexity of the residual correlation matrix, treating a single correlated pair identically to a fully dense dependence structure. We propose two novel metric families: Metric~E (E-CVWMD: Enhanced Coefficient-of-Variation Weighted Marginal-Dependence) and Metric~E2 (E-CVWMD-Pairwise), designed for mixed-type multivariate responses combining continuous and binary outcomes within a cross-validation framework. We position Metrics~E and~E2 as diagnostic ranking tools for comparing competing models rather than as strictly proper scoring rules, and we provide a strictly proper Log-Loss variant (E-LL / E2-LL) for applications that require the full properness guarantee. Metric~E assigns variable-level weights proportional to the coefficient of variation (CV) of each outcome on the training partition, and adaptively calibrates the marginal-dependence trade-off parameter $\alpha^*$ via a global distance-correlation test. Metric~E2 refines this by replacing the global test with a pairwise Spearman screening index $\hat{\pi}$ – the proportion of variable pairs with significant residual correlation – which maps linearly to $\alpha^*(\hat{\pi}) = 1 - \hat{\pi}/2 \in [0.5, 1]$. Applied to the validation of a Generalized Multivariate Functional Additive Mixed Model (GMFAMM) on 62 Valle del Cauca meteorological stations ($N_{\text{test}} \approx 31\,663$), the naive significance-based index saturates ($\hat{\pi} = 1.0$) at this large sample size – every pair, including correlations as small as $|\hat{\rho}_s| \approx 0.01$, is flagged ``significant'' – which is precisely the sample-size sensitivity we address. Under the effect-size screening ($|\hat{\rho}_s| \geq 0.05$), three negligibly correlated pairs are excluded, yielding $\hat{\pi} = 0.70$ and $\alpha^*_{E2} = 0.65$, a better-calibrated weight than Metric~E's $\alpha^*_E \approx 0.797$ under the same data. A large-scale simulation study with 37,440 model evaluations confirms that Metric~E inverts the correct ranking at correlation levels $\rho \geq 0.40$ (CDR = 0\%), while E2 maintains correct discrimination in 14 of 15 simulation conditions (M1 vs. M3). We also delimit the metrics' scope: E2 degrades under near-saturated uniform dependence – a regime in which the strictly proper Energy Score remains preferable – and the pairwise index is sensitive to sample size, for which we provide an effect-size-based variant. An R package (mvmetrics v0.2.0, https://darango2025.github.io/mvmetrics) implementing both metrics, the Log-Loss variant, alternative weighting schemes, and the effect-size screening is publicly available.

Article
Computer Science and Mathematics
Probability and Statistics

Rihab Ahmed Abed

,

Wafaa A. Ashour

,

Nooruldeen A. Noori

Abstract: This paper proposes a new four-parameter statistical distribution based on the neutrosophic Gompertz (NGo-G) family and the extended Nadarajah-Haghighi distribution in neutrosophic logic called neutrosophic Gompertz Nadarajah-Haghighi (NGoNH) distribution to deal with uncertain and indeterminate data, known as neutrosophic data. The basic distribution functions and some properties are derived and the parameters of the proposed distribution are estimated using three different methods. To compare the performance of the different estimation methods, numerical simulations are performed using the evaluation criteria: MSE, RMSE, and bias. NGoNH distribution is applied to real data set representing the monthly minimum and maximum temperatures in Lahore, Pakistan, for the period 2016-2020. To verify the consistency of the data with the proposed distribution, the properties of the neutrosophic data are tested and the three components (True, Indeterminate, False) are plotted. In addition, the performance of the proposed distribution is compared with six other neutrosophic distributions using some information criteria and some goodness-of-fit tests. The extent to which the data fit the proposed distribution is plotted to demonstrate its effectiveness. The results indicate that the proposed distribution provides a better fit to neutrosophic data than other distributions, which enhances its effectiveness in analyzing data with uncertainty.

Article
Computer Science and Mathematics
Probability and Statistics

Maksym Luz

,

Mikhail Moklyachuk

Abstract: We consider the problem of optimal linear estimation of the functional ANξ = ∑Nk=0 a(k)ξ(k) which depends on the unknown values ξ(k), k = 0,1,...,N, of a stochastic sequence with harmonizable symmetric α-stable nth increments. The derived estimates are based on observations at points m ∈ Z \ {0,1,2,...,N}. The cases of observations without noise, with harmonizable symmetric α-stable noise and with noise having harmonizable symmetric α-stable increments are studied. The classical solutions as well as the minimax robust ones are obtained.

Article
Computer Science and Mathematics
Probability and Statistics

Francisco Novoa-Muñoz

Abstract: The Poisson–Three-Parameter Lindley (PTPL) distribution constitutes a flexible Poisson mixture model for overdispersed count data, encompassing several classical count distributions as special or limiting cases. Despite its growing use in applied contexts, no formal goodness-of-fit test specifically designed for this distribution is currently available. In this paper, we propose and study a new goodness-of-fit test for the PTPL model based on a Cramér-von Mises type distance between the empirical and theoretical probability generating functions (PGFs). For polynomial weight functions, the test statistic admits an explicit closed-form representation; in practice, it is computed efficiently via numerical quadrature. The null distribution of the statistic is approximated via parametric bootstrap. We establish theoretical properties of the proposed procedure, including consistency against fixed alternatives and the validity of the bootstrap approximation. Monte Carlo simulations with sample sizes n ∈ {50,100,150,200,500} for size evaluation and n = 250 for power comparisons, and weight exponents a ∈ {0,1,2}, show that the empirical size is well controlled at both the 5% and 10% nominal levels, and that the test exhibits competitive power against Poisson, Negative Binomial, COM-Poisson, and Zero-Inflated Poisson alternatives. Areal data application to five overdispersed count datasets further illustrates the practical utility of the method.

of 30

Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings