Submitted:
25 September 2026
Posted:
29 September 2026
You are already at the latest version
Abstract
Let \(S(x)\) denote the sum of the decimal digits of a positive integer \(x\). For a prime \(p\) and an exponent \(n \geq 1\), we study \(\widetilde{A}_p(n) = \frac{S(p^n)}{\lfloor n\log_{10}p\rfloor+1}\), the average value of the decimal digits of \(p^n\). We computed this statistic for the first fifty primes, from 2 through 229, and for exponents up to \(n = 8000\), giving 400,000 prime–exponent observations. Over this range, the prefix means move toward and remain close to \(9/2\), while the empirical prefix standard deviations decrease substantially as larger exponents are included. These observations motivate two conjectures: that the Cesàro mean of the normalized digit-sum statistic tends to \(9/2\), and that its prefix standard deviation tends to zero. The results are empirical and concern only this scalar statistic; they do not establish normality, independence, mixing, or uniform distribution of the decimal digits of prime powers.
Keywords:
digit sums
; prime powers
; computational number theory
; decimal expansions
; Cesàro means
; empirical concentration
1. Introduction
For a positive integer x, let denote the sum of its decimal digits. Digit sums arise naturally in number theory and in questions about the statistical behaviour of deterministic sequences. Even when a sequence is generated by a simple arithmetic rule, its decimal expansion can display complicated behaviour, and proving strong distributional results for such expansions is often difficult.
In this paper, we study the powers
of a fixed prime p. Rather than trying to describe the full decimal expansion of , we look at a simpler statistic: its digit sum divided by its number of decimal digits. Writing
we define
Thus is simply the arithmetic mean of the decimal digits of .
A natural benchmark comes from a simple random-digit model. A digit chosen uniformly from
has expected value
This suggests using as a reference value for the normalized digit sums of prime powers.
The comparison is only heuristic. The decimal digits of are deterministic and satisfy arithmetic restrictions. For example, terminal digits are constrained by congruences modulo powers of 10, while the leading digit is not distributed like an interior digit. Also, an average digit close to is much weaker than uniform digit frequencies, independence, normality, mixing, or equidistribution of digit blocks. We therefore use the random-digit model only as a benchmark for the particular statistic studied here.
The main question is:
How does the normalized digit sum behave as the exponent grows, and how stable is its aggregate behaviour across different primes?
To investigate this, we computed the statistic for the first fifty primes,
and for exponents
We examine the individual values of , their prefix means, and their empirical prefix standard deviations.
The computations show three main finite-range features:
- For the primes tested, the prefix means of move toward values close to as the cutoff increases.
- The empirical prefix standard deviations decrease substantially over the sampled cutoffs, so later values of fluctuate on a smaller scale than the early values.
- The primes and show especially pronounced early transients, suggesting that arithmetic relationships with the decimal base can strongly affect short-range behaviour.
These observations lead to conjectures about the limiting prefix mean and the asymptotic concentration of . The purpose of the paper is exploratory: to record the finite-range numerical evidence, state the questions suggested by it, and keep those conjectures separate from stronger claims about the full decimal expansions of prime powers.
2. Literature and Positioning
Digit-sum functions have been studied extensively in analytic and probabilistic number theory. Classical work considers their average behaviour, distribution, and asymptotic properties, while later work also studies digit sums along structured sequences.
For general background on sum-of-digits behaviour, see [1,2,3]. Results for structured sequences, including polynomial sequences, can be found in [4,5].
The present paper focuses specifically on powers. Digit sums of sequences such as
have been studied through questions such as lower bounds, unboundedness, and dependence on the base. These questions are closely related to the present one, but they do not directly determine the behaviour of the normalized quantity
For example, knowing that grows or is unbounded does not by itself describe its size relative to the number of digits of .
Relevant results on powers include lower bounds for the decimal digit sum of and more general bounds for digital sums of powers in arbitrary bases; see [6,7].
Classical two-base results also impose strong restrictions on sparse digit representations. Senge and Straus proved a finiteness theorem for integers having bounded digit sums in two multiplicatively independent bases, and Stewart later obtained an effective quantitative form of this phenomenon; see [8,9]. Applied to powers, these results give important growth information for digital sums, but they do not determine the normalized average-digit statistic or its Cesàro behaviour considered here.
There is also a separate connection with leading digits. For powers , the fractional parts
control the leading decimal digits. When is irrational, these fractional parts are equidistributed modulo 1, giving the usual logarithmic leading-digit distribution. This is related to Benford’s law, but it does not imply that all digits of are uniformly distributed or independent.
Any fixed number of leading digits contributes only
to the normalized total digit sum, so leading-digit results do not settle the prefix-mean or concentration questions considered here.
The relevant equidistribution principle goes back to Weyl [10]; for Benford’s law and the significant-digit phenomenon, see [11,12].
The contribution of this paper is computational. We study
for the first fifty primes and for exponents up to 8000. The aim is to record the finite-range behaviour of the individual values, prefix means, and prefix standard deviations, and to formulate the asymptotic questions suggested by those computations. The results concern this normalized digit-sum statistic only and are not intended as claims about normality, uniform digit frequencies, or independence of the full decimal expansions.
3. Definitions and Methodology
3.1. Digit Sum and Average-Digit Statistic
For a positive integer x, let denote the sum of its decimal digits. For example,
Let p be a prime and let . The exact number of decimal digits of is
Definition 1.
For a prime p and an integer , define thenormalized digit-sum statistic
Thus is the arithmetic mean of the decimal digits of .
Remark 1
(Exact and asymptotic normalizations). A closely related normalization is
Since
the two normalizations differ by .
Indeed,
Because
and
we obtain
All computations in this paper use the exact normalization .
3.2. A Random-Digit Benchmark
As a simple benchmark, suppose that decimal digits are uniformly distributed on
A digit D drawn from this distribution satisfies
and
If an integer had L independent uniformly distributed digits, then the mean of those digits would have expectation
and variance
This model is not a literal model for . In particular, an actual L-digit integer cannot have leading digit 0. If the leading digit is instead taken uniformly from
while the remaining digits are uniform on
then the expected average digit becomes
So the leading-digit correction is and becomes negligible as L grows.
Prime powers also satisfy deterministic digit constraints. Their terminal digits are restricted by residues modulo powers of 10, digit sums satisfy
and carries introduce dependence among digits. We therefore use the random-digit model only as a heuristic benchmark for the scale and central value of .
3.3. Experimental Range
We computed the statistic for the first fifty primes,
and for exponents
Aggregate statistics were recorded at
This range lets us compare early transients with later behaviour across many primes, but it is not large enough by itself to establish an asymptotic law.
3.4. Computational Procedure
All computations were carried out in Python using arbitrary-precision integer arithmetic. For each prime p, the powers were generated iteratively using
For each exponent n, we converted to its decimal representation, summed its digits, and recorded
For each prime p and cutoff N, the prefix mean was computed as
The prefix standard deviation was computed using the sample convention
This is the convention used for the values in Table 1. The plots were generated with matplotlib.
3.5. AI-Assisted Manuscript Preparation
OpenAI ChatGPT was used during manuscript preparation for language editing, organizational feedback, and assistance in reviewing presentation and code-related explanations. The author is responsible for the final text, mathematical statements, computations, citations, and conclusions.
3.6. Reproducibility and Data Availability
The computations use only arbitrary-precision integer arithmetic, decimal conversion, digit summation, and the formulas above. For each tested prime p, we generated the powers iteratively and computed
directly from the resulting integers.
3.7. Interpretive Scope
The quantities studied here are empirical statistics of a deterministic sequence.
A prefix mean near does not imply pointwise convergence of , and it does not imply uniform digit frequencies, normality, independence, or mixing.
Likewise, a decreasing prefix standard deviation does not determine its exact asymptotic rate. Part of the decrease is naturally expected because averages over
digits, so later terms involve longer decimal expansions.
The conjectures below are therefore stated only for the normalized digit-sum statistic, not for the full decimal digit process.
4. Experimental Results
We now examine the exact normalized digit-sum statistic
Across the computed range, three features stand out:
- 1.
- The individual values fluctuate around a level close to , with noticeably larger fluctuations at small exponents.
- 2.
- For each tested prime, the prefix meanmoves toward a value close to as N increases.
- 3.
- The empirical prefix standard deviation decreases substantially over the sampled cutoffs, although the data do not determine its precise asymptotic rate.
These statements describe only the finite range that was computed. The asymptotic questions suggested by the data are stated separately as conjectures.
4.1. Behaviour of the Sequence
We begin with the individual values of .
Figure 1 shows the statistic for several small primes. The fluctuations are large at small exponents. As n increases, their visible scale becomes smaller, while the values continue to lie around a level close to .
This is consistent with the fact that is an average over
decimal digits. In the simple random-digit model, the variance of an average over L digits is proportional to . So a smaller fluctuation scale at larger exponents is natural under that heuristic, even though the actual digits of are deterministic and dependent.
Figure 1 does not suggest pointwise convergence of to . The statistic continues to oscillate, which makes its prefix averages the more natural object to study.
4.2. Prefix Means
For each prime p, define
Figure 2 shows how changes with N for several representative primes. After the largest early variations, the curves move toward and remain close to over the computed range.
Figure 3 compares the prefix means at selected cutoffs, including and , where the early behaviour is more pronounced. By , the displayed values are closely clustered around the benchmark. This is suggestive, but the figure alone cannot establish convergence or determine the limiting value.
The data lead naturally to the following conjecture.
Conjecture 1
(Prefix-Mean Conjecture). For every prime p,
Equivalently, in the notation above,
This is only a statement about the Cesàro mean of the normalized digit-sum statistic. It does not imply that converges pointwise, nor does it imply uniform digit frequencies, normality of , independence of digits, or equidistribution of finite digit blocks.
The value comes from the mean of the simple uniform-digit benchmark. Agreement with it should therefore be read as agreement with a first-moment heuristic, not as evidence that the full decimal expansions behave like independent random strings.
4.3. Prefix Dispersion and Concentration
We next consider the sample standard deviation of the first N values,
For every tested prime, this quantity decreases substantially across the sampled cutoffs. So, within the computed range, longer prefixes are less dispersed than shorter ones.
Table 1 gives representative values for , , and .
There is a simple heuristic for this decay. Under the independent uniform-digit model, an average over digits has variance
Since
this gives the rough scale
Averaging these termwise variances over the prefix gives
This suggests the heuristic scale
This argument is only a benchmark. The digits of are not independent, and the sequence is deterministic and nonstationary. Carries, congruence restrictions, and dependence between successive powers may change the constant or even the detailed asymptotic behaviour. The present computations do not determine whether
is the true leading-order scale.
What the data do show is a clear downward trend in prefix dispersion across the sampled cutoffs. They do not show monotonic decrease for every integer N, since local increases may occur between those checkpoints.
This leads to the second conjecture.
Conjecture 2
(Prefix-Concentration Conjecture). For every prime p,
Equivalently,
The stronger rate prediction
for some prime-dependent constant is a plausible heuristic from the random-digit model, but the present data are not enough to state it as a formal conjecture.
Figure 4 shows the observed decrease at the sampled values of N.
4.4. Arithmetic Structure and Pronounced Early Transients
Two primes show especially pronounced early behaviour in the present computations: and . They illustrate two different ways in which arithmetic structure can affect finite-range digit-sum statistics.
4.4.0.1. The prime .
The powers of 2 have relatively few decimal digits at small and moderate exponents:
So each early value of is an average over a shorter digit string than the corresponding statistic for a larger prime. This gives the early terms a larger fluctuation scale and more influence on prefix statistics.
The prime 2 is also distinguished by the decimal base. Since
the primes 2 and 5 are exactly the prime divisors of the base. Unlike primes coprime to 10, they are not units modulo , so their powers do not have the same purely periodic residue behaviour modulo .
In particular, and are not invertible modulo , and their residue sequences therefore differ from the unit-group behaviour of primes coprime to 10. This arithmetic distinction does not determine the normalized digit sum by itself, but it explains why 2 and 5 should be treated separately when base-10 structure is discussed.
In the present data, has a particularly large early transient. The prime is structurally special as well, although it is not as pronounced an outlier in the displayed prefix statistics.
4.4.0.2. The prime .
The case has a different source of structure. Since
the binomial theorem gives
When every coefficient is smaller than 100, these terms occupy nonoverlapping blocks of two decimal places. In this carry-free range, the decimal expansion of can be read directly from the binomial coefficients, inserting leading zeros inside two-digit blocks when needed.
This description applies for , since
throughout that range.
For example,
so
The zero-separated block structure is therefore a direct consequence of the binomial expansion.
For larger exponents, some binomial coefficients exceed 99, so adjacent two-digit blocks interact through carries and the simple block description breaks down. Even so, the relation
gives a natural explanation for the unusually structured early expansions and the pronounced transient seen in the data.
4.4.0.3. Periodic terminal digits for primes coprime to 10.
For any prime and any fixed integer , the sequence
is purely periodic in n. So the final k decimal digits of satisfy deterministic periodic constraints.
These constraints rule out a literal model of independent random decimal strings. However, any fixed number of terminal digits contributes only
to the normalized digit average. Strong structure in finitely many terminal positions therefore does not by itself contradict the prefix-mean conjecture.
These examples suggest that arithmetic relationships with the base can have a substantial effect on short- and medium-range behaviour. A systematic classification of such transients would require a more precise criterion than the informal idea of an “exceptional prime,” and is left for future work.
5. Discussion
The numerical evidence supports two deliberately narrow claims about finite-range behaviour: stabilization of prefix means near the uniform-digit benchmark and decreasing prefix dispersion over the sampled cutoffs. Neither observation is evidence that the decimal digits of form an independent or stationary process. The arithmetic examples above show that deterministic structure remains visible, especially at small exponents and in positions constrained by congruences.
The Prefix-Mean Conjecture is weaker than the natural pointwise heuristic
which is already a natural open benchmark in the case [6], and which would imply . The present computations do not suggest claiming this stronger limit: the individual values continue to oscillate, and only the Cesàro behaviour is formulated as a conjecture here. Similarly, the Prefix-Concentration Conjecture concerns the average squared deviation across a growing prefix; it does not assert a limiting distribution or a central limit theorem.
The principal limitation is computational range. Although prime–exponent observations provide a broad finite sample, an asymptotic statement cannot be inferred from this alone. The heuristic scale is therefore treated only as a benchmark for future theoretical or computational work. A useful next step would be to test substantially larger exponent ranges and other bases while separating base-dependent arithmetic effects from statistics that persist across bases.
6. Conclusion
We studied the normalized decimal digit sum
for the first fifty primes and for exponents up to .
Over this range, the values of fluctuate around a level close to . Their prefix means move toward and remain close to this benchmark, while the empirical prefix standard deviations decrease substantially across the sampled cutoffs.
This leads to two asymptotic conjectures:
- 1.
- for every prime p,converges to ;
- 2.
- for every prime p,converges to 0.
The benchmark comes from the simple uniform-digit model. The same heuristic suggests a possible dispersion scale of
since the number of digits of grows linearly with n. The present computations do not determine whether this is the true asymptotic rate.
These results concern only the normalized digit-sum statistic. They do not imply pointwise convergence, normality, independence, mixing, uniform digit frequencies, or equidistribution of finite digit blocks. The decimal expansions of prime powers retain clear arithmetic structure through congruence restrictions, terminal-digit periodicity, carries, and base-related effects such as those seen for , , and .
Several questions remain. The most immediate are whether the two conjectures hold for every prime, what the true fluctuation rate is, and how finite-range behaviour depends on arithmetic relationships between p and the base. It would also be natural to study the same statistic in a general base b, to compare prime and composite bases of exponentiation, and to investigate forms such as
where the decimal base is built directly into the arithmetic structure.
The computations here provide numerical evidence for a stable pattern, but the main questions are still theoretical. They give a natural starting point for understanding why this behaviour appears and how far it extends.
Data Availability Statement
The source code, numerical data, validation reports, and reproducibility materials associated with this study are openly archived at https://doi.org/10.5281/zenodo.21702236.
References
- Bush, L.E. An Asymptotic Formula for the Average Sum of the Digits of Integers. Am. Math. Mon. 1940, 47, 154–156. [Google Scholar] [CrossRef]
- Delange, H. Sur la fonction sommatoire de la fonction “somme des chiffres”. L’Enseignement Mathématique 1975, 21, 31–47. [Google Scholar]
- Drmota, M.; Gajdosik, J. The Distribution of the Sum-of-Digits Function. J. De Théorie Des. Nr. De Bordx. 1998, 10, 17–32. [Google Scholar] [CrossRef]
- Drmota, M.; Mauduit, C.; Rivat, J. The Sum-of-Digits Function of Polynomial Sequences. J. Lond. Math. Soc. 2011, 84, 81–102. [Google Scholar] [CrossRef]
- Gelfond, A.O. Sur les nombres qui ont des propriétés additives et multiplicatives données. Acta Arith. 1968, 13, 259–265. [Google Scholar] [CrossRef]
- Radcliffe, D.G. The Growth of Digital Sums of Powers of 2, 2016. arXiv arXiv:math.
- Radcliffe, D.G. Elementary Bounds on Digital Sums of Powers, Factorials, and LCMs, 2025; arXiv.
- Senge, H.G.; Straus, E.G. PV-Numbers and Sets of Multiplicity. Period. Math. Hung. 1973, 3, 93–100. [Google Scholar] [CrossRef]
- Stewart, C.L. On the Representation of an Integer in Two Different Bases. J. Für Die Reine Und Angew. Math. 1980, 319, 63–72. [Google Scholar] [CrossRef]
- Weyl, H. Über die Gleichverteilung von Zahlen mod Eins. Math. Ann. 1916, 77, 313–352. [Google Scholar] [CrossRef]
- Benford, F. The Law of Anomalous Numbers. Proc. Am. Philos. Soc. 1938, 78, 551–572. [Google Scholar]
- Hill, T.P. The Significant-Digit Phenomenon. Am. Math. Mon. 1995, 102, 322–327. [Google Scholar] [CrossRef]
Figure 1.
Values of the normalized digit-sum statistic for selected primes. The dashed line marks the heuristic benchmark .
Figure 1.
Values of the normalized digit-sum statistic for selected primes. The dashed line marks the heuristic benchmark .

Figure 2.
Prefix mean of the normalized digit-sum statistic for several small primes. The dashed line marks .
Figure 2.
Prefix mean of the normalized digit-sum statistic for several small primes. The dashed line marks .

Figure 3.
Prefix mean over for selected primes and cutoffs. The displayed values are closely clustered around at the larger tested cutoffs.
Figure 3.
Prefix mean over for selected primes and cutoffs. The displayed values are closely clustered around at the larger tested cutoffs.

Figure 4.
Sample standard deviation as a function of the prefix cutoff N for selected primes. The plotted checkpoints show a substantial decrease over the tested range.
Figure 4.
Sample standard deviation as a function of the prefix cutoff N for selected primes. The plotted checkpoints show a substantial decrease over the tested range.

Table 1.
Sample standard deviation of for representative primes, using the exact digit count .
| N | |||
|---|---|---|---|
| 50 | |||
| 250 | |||
| 500 | |||
| 1000 | |||
| 2000 | |||
| 4000 | |||
| 8000 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.