Preprint
Article

This version is not peer-reviewed.

Prior Density Selection and Statistical Models

Submitted:

10 August 2026

Posted:

12 August 2026

You are already at the latest version

Abstract
In the application of Bayesian methods, the choice of a prior density for defined model parameters is often a practical issue. In scientific and medical applications, it requires careful justification. In many settings, the properties of the statistical likelihood function and the overall statistical model, within which the model parameters are defined, are useful in justifying the chosen prior density. In some settings this is extended to matching Bayesian and frequentist model results. In this paper, prior density selections reflecting specific model related and matching properties are reviewed. It is further shown that if these considerations are extended to include the matching of likelihood and posterior density higher order derivatives, they provide new exclusionary rules for the selection of prior densities. In addition, the class of related acceptable prior densities are non-informative on a scale related to the Pratt-Arrow utility-based measure of relative risk. Multivariate extensions are briefly considered.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Mathematical and population-based probability models are often used to summarize, simplify and make interpretable scientifically relevant patterns in data [1]. In some settings, pre-existing prior information for the parameters in these models, typically interpreted as population characteristics, can also be expressed on a probability scale. This information can be incorporated into the statistical analysis via Bayesian statistical models [2].
The subjective Bayesian approach to statistical modeling uses a probability density p ( θ ) to represent and describe existing baseline beliefs regarding the unknown and unobservable model parameters θ . After assuming a population-based probability model and collecting the data x , the beliefs represented by p ( θ ) are updated to include this new information via the posterior density p ( θ | x ) , using the likelihood function L ( θ | x ) and application of Bayes theorem. Once an experiment is conducted, p ( θ | x ) provides the new, updated probability weighting for each value of θ . Note that from the Bayesian perspective, there is no assumption of a “true” value for θ to formally test or estimate, rather p ( θ | x ) gives a probability weight for each possible θ value. The researcher can interpret this directly, or compare distinct values of θ using measures such as the posterior odds ratio or the Bayes factor [3].
The likelihood function, can be written:
L ( θ | x ) = k i = 1 n f ( x i | θ )
where f ( x i | θ ) is the probability model or density for the i t h independent subject response x i and the constant k emphasizes the fact that the likelihood is a function of θ , not necessarily a density for θ . The likelihood function is the key source of information to be drawn from a given model-data combination. Often the mode of the likelihood function, the maximum likelihood estimator, is the basis of frequentist inference. A function of the local curvature of the log-likelihood function about its maximum value provides an estimate of sampling variation [4].
Many applications of Bayesian methods view the selection of0 p ( θ ) as primarily subjective in nature, reflecting the ideas and beliefs of the individual researcher [5]. From this perspective there are no limitations on the shape or entropy of p ( θ ) apart from the need to select a function that technically integrates to one and has no negative values. Indeed, a common choice is a non-informative density, a flat prior which may not technically integrate to one [6]. In some model settings, the need to employ Monte Carlo based methods to estimate marginal posterior elements supports the selection of convenience priors to aid in convergence [7].
Other justifications of p ( θ ) are more formal and less subjective. For example, some argue for the use of meta-analysis and a choice of p ( θ ) that reflects existing scientific understanding of the possible values for θ [8]. The choice of subjective prior may also be guided by the need to respect specific aspects of the statistical model and related inferential approach.
For example, the idea of matching priors, choosing prior densities for θ that deliberately structure the posterior density in a manner that yields tail areas similar to those underlying frequentist p-value tail area calculations. This limits the choice of p ( θ ) to a specific form [9].
The selection of prior density may also be guided by the need to maintain specific properties of the likelihood function and related model structure, for example likelihood invariance or model symmetry [10,11]. Note that the model parameter θ is defined and takes its meaning from the context of the statistical model.
Here, prior density selections reflecting specific model related and matching properties are reviewed. It is shown that when these considerations are extended to include the matching of likelihood and posterior density higher order derivatives, they provide new exclusionary rules for the selection of prior densities. In addition, the class of related acceptable prior densities are non-informative on a scale related to the Pratt-Arrow utility-based measure of relative risk

2. Bayesian Inference

As noted, the Bayesian approach or perspective is based on the joint posterior density p ( θ | x ) which can be expressed as:
p ( θ | x ) = c p ( θ ) L ( θ | x )
or
log p ( θ | x ) = c ' + log p ( θ ) + log L ( θ | x )
where p ( θ ) is the prior density, L ( θ | x ) the likelihood function and c the constant of integration. All three functions of θ can be viewed as weighting the model parameter space, with prior and posterior densities restricted to a probability scale.
Once the joint posterior density p ( θ | x ) is obtained, integration is employed in the Bayesian setting to obtain marginal posterior densities for any given θ i . For example,
p ( θ 1 | x ) = p ( θ | x ) d θ 2 d θ 3 d θ p
gives the marginal posterior for θ 1 alone. The central region of this density is a Bayesian credible region which can be used for estimation regarding θ 1 . With the advent of modern methods for numerical integration [12], accurate calculations in many Bayesian settings are available. Posterior odds ratios, Bayes factors and predictive distributions can be used for inference [3].
Note that, in general, re-scaling variables can sometimes induce information into the Bayesian setting, affecting nominally non-informative priors and making them informative. It is assumed here that once the appropriate scaling is found for the parameters θ in the likelihood function, the likelihood is expressed in terms of this parameter scaling and prior densities expressed in a similar manner.

3. Bayesian and Frequentist Similarities

Limitations on the choice of subjective prior can arise. The idea of matching priors, choosing prior densities for θ that deliberately structure the posterior density in a manner that yields tail areas similar to those underlying frequentist p-value tail area calculations, limits the choice of prior density. While similarities between Bayesian and frequentist statistics exist in large samples due to the presence of central limit theorems along with specific assumptions, they also occur in smaller samples where specific model structures are a key consideration. In these settings, limitations on the choice of subjective priors arise.

4. Matching Prior Selection

As noted above, in large samples and subject to regularity conditions [4], as n ,
θ m l e ( x ) N ( θ , I 1 ( θ | x ) )
where θ m l e ( x ) is the maximum likelihood estimate and I 1 ( θ | x ) the Fisher information.
From the related Bayesian perspective, with a non-informative prior independent of the sample size n , and subject to regularity conditions, as n ,
θ | x N ( θ m l e ( x ) , J 1 ( θ | x ) )
with J ( θ | x ) the observed Fisher information with J 1 ( θ | x ) I 1 ( θ | x ) in probability.
Note that these results both have multivariate versions that are very similar:
θ m l e ( x ) N p ( θ , I 1 ( θ | x ) )
θ | x N p ( θ m l e ( x ) , J 1 ( θ | x ) )
This close technical relationship between large sample Bayesian and frequentist inferential probability has motivated investigations regarding which choice of p ( θ ) leads to equivalence in tail area probabilities and central regions related to inference [9]. Examples include [13,14,15]. Usually this involves matching tail areas of the sampling and posterior distributions to the first or second order of approximation.
Results can be formulated in terms of the accuracy of the Bayesian tail area in frequentist terms. The prior density p ( θ ) is defined to be an r t h order probability matching prior if its use satisfies the condition:
p θ i | x ( θ i θ i ; 1 α ) = 1 α + o ( n r / 2 )
where p θ | x ( θ i θ ) is the cumulative distribution function of the marginal posterior density for θ i with a frequentist probability coverage of 1 α with an relative error of o ( n r / 2 ) . Typically, accuracy at the r = 1 ,   2 level are considered.
So-called matching priors for θ = ( θ 1 , θ 2 , . . . , θ p ) that satisfy this type of condition for say θ 1 are often of the form:
p ( θ ) g ( θ 2 , . . . , θ p ) [ i 11 ( θ ) ] 1 / 2
where i 11 ( θ ) is the first entry in the Fisher information, g ( ) is an arbitrary smooth function, and the Fisher information has a specific assumed structure, for example the correlations along the first row or column of the Fisher information are assumed to be zero, ie. θ 1 is orthogonal to ( θ 2 , . . . , θ p ) . Note that this prior is not unique.
The overlap of Bayesian and frequentist perspectives show a conceptual link going beyond tail area similarity, especially where the prior is independent of the sample size n . The overlap links the informative or learning aspect of the Bayesian method to the large sample minimal sufficiency of the m . l . e . and the likelihood based score function, in units given by the Fisher information and is further discussed below.

5. Fiducial or Transferred Probability

It is interesting when discussing the Bayesian and frequentist overlap in large samples to consider briefly the concept of fiducial probability, the idea that in certain settings, often in smaller samples, probability intervals derived on the sample space can be transferred or matched to intervals on the parameter space without use of a prior density and Bayes theorem, relying instead on the structure of the statistical model or a matching condition.
While this approach can be linked to the likelihood concept, it often reflects the importance of symmetry in defining a statistical model.
The fiducial approach is primarily justified by the inversion of significance tests, namely the region of the sample space not included in a rejection region for a test of the null hypothesis H 0 : θ = θ 0 with Type I error = α , based on minimal sufficient statistics, and viewed as a region of supported values for θ . The level of this supported region is 1 α and this level is justified by repeating the significance test for each possible value of θ , keeping those values of θ (usually in the initial form of a pivotal quantity) that are not rejected by the significance test.
The supported region is sometimes viewed as the central region of a fiducial or confidence distribution. This distribution is built using the tail area of the minimal sufficient statistic distribution, taking into account the pivotal quantity structure that underlies most significance tests (the minimal sufficient statistic is standardized to give a pivotal quantity) with this pivotal quantity having a known distribution. Note that in exponential family probability models the likelihood statistics are the minimal sufficient statistics.
The fiducial distribution can easily be derived in one dimension. Remember that a probability density f ( x ) = ( d / d x ) F ( x ) where F ( x ) = P ( X x ) . In significance testing, the tail area ( 1 F ( x ) ) is required and it follows that f ( x ) = d / d x ( 1 F ( x ) ) . To include the pivotal structure, this is usually written in terms of T ( x , θ ) , where typically T ( x , θ ) = ( S ( x ) θ ) / V a r ( S ( x ) ) and S ( x ) is the minimal sufficient statistic. If | d T ( x , θ ) / d x | = | d T ( x , θ ) / d θ | is assumed to hold, then there is a symmetry between S ( x ) and θ (they are interchangeable) within T ( x , θ ) , and d / d T ( 1 F ( T ( x , θ ) ) is equivalent to d / d θ ( 1 F ( T ( x , θ ) ) )  .
On this basis it can be claimed that the probability distribution can be transferred from the sample space to the parameter space through the symmetry of the pivotal, without assuming a prior density for θ and assuming S ( x ) is an observed value. Note that the likelihood ratio itself can be used as a pivotal quantity, so aspects of the likelihood function, apart from the likelihood statistics, may be involved.
More practically, the significance test requires both an observed value of S ( x ) and a null value for θ , say θ 0 . If the rejection region of the test is well defined (easiest in one dimension with a scalar θ ) then the probability of the rejection region will alter as θ 0 alters. As θ 0 goes from to (the defined parameter space) the probability of the rejection region based on the observed minimal sufficient statistic or likelihood statistic will range from 0 to 1 with each value greater than zero. So the linking of the θ 0 values to the respective coverage of the rejection region can be viewed as defining a distribution for θ .
There are serious issues regarding how this relates to formal probability theory, especially in multivariate settings where integrating out unwanted θ i values is required and ordered mappings are difficult to define [16]. The fiducial “transfer” is often not stable in terms of probability and cannot necessarily be interpreted as a Bayesian posterior density [17].
This issue arises since the mapping from the sample space to parameter space through the pivotal is often not unique [18], as typically T ( x , θ ) = g [ ( S ( x ) θ ) / V a r T ( x ) ] and the inverse function g 1 ( x ) are not 1 1 . The procedure is also subject to selection bias as the observed p-value is a component of the construct. But this can be viewed as a potentially positive aspect; a frequentist approach reflecting important aspects of the observed data, rather than only average, pre-planned frequentist concepts.
Where this mapping is 1 1 , for example by imposing a group theoretic perspective (ie. symmetry) on the model as a first step in model definition and conditioning formally on the observed data (usually in the form of residuals), the fiducial argument has better support. This is accomplished in the work of Fraser’s [19] structural inference and applies primarily to location scale models.
There have been recent attempts to generalize the fiducial concept [20], and the relative weighting of the parameter space given by the fiducial approach may be of interest in some settings, as it reflects a frequentist inferential procedure conditional on aspects of the observed data.

6. Location-Scale Statistical Models

Note that even in settings with small sample size, outside of improper non-informative prior selection, isomorphic mappings exist between Bayesian and conditional frequentist methods in linear models, when multivariate symmetry is imposed, yielding pivotal inversions and related probability mappings from the sample space to the parameter space that give identical Bayesian and frequentist inference regions for θ . The symmetry in question typically requires the selection of a right invariant Haar prior density from the Bayesian perspective and imposition of a specific type of symmetry as a basic property of the linear model from both frequentist and Bayesian perspectives [18].
In the pivotal quantity based setting above, if applying a standard Bayesian linear model with a right invariant Haar prior density to preserve a certain type of model symmetry, the conditional frequentist model, for example Fraser’s structural inference [19], is mathematically identical (isomorphic) to the Bayesian model, in small sample linear models. Namely the central regions estimating θ will be identical. A formal proof of this connection was developed [21].
To state this more practically, in the linear model y N ( X β , σ 2 I ) , the conditional distribution of the least squares estimator b = ( X ' X ) 1 X ' y , conditional on the observed residuals ( y X b ) , is isomorphic (in p dimensions) to the Bayesian posterior distribution of β (again conditional on the observed data), if the prior for ( β , σ ) is the right invariant Haar prior. This choice preserves a symmetry of the linear model under the location-scale group action and the induced group acting on the parameter space in the Bayesian setting. Note that a justification for choosing such a prior can be the imposition of model symmetry itself.
These direct small sample connections between the Bayesian and frequentist settings are more limited in generalized linear models using exponential family densities where mean and variance elements are not independent and the resulting residuals harder to interpret. These models are inherently non-linear and not easily prone to group theoretic representation. The likelihood function is the basis of inference in such settings, often using large sample properties of the maximum likelihood estimate [4].
In large samples in general, due to the central limit theorem, this type of pivotal symmetry often gives similar frequentist and Bayesian results, since in large samples, with non-informative priors and additional assumptions, the Bayesian central limit theorem and the frequentist central limit theorem for the maximum likelihood estimate often overlap [22] and the connection between Bayesian credible regions and frequentist confidence regions is direct for the overall joint distributions. Note that inference for individual parameters θ i may still differ due to integration of the posterior density and related shrinkage effects.

7. Jeffreys Prior: Preserving Likelihood-Based Invariance

When discussing Bayesian prior selection in general it is useful to consider the Jeffreys prior. This is closely linked in concept to the likelihood function, but within a Bayesian setting. It is primarily justified as preserving the invariance of the likelihood function under rescaling of the parameter space, in the context of the posterior density [10]. However, it is not based on prior existing information, as it is a function of the sample size and likelihood function for the upcoming experiment.
If the Bayesian asymptotic result θ N ( θ m l e ( x ) , J 1 ( θ | x ) ) can be assumed to hold, the Jeffreys prior here can be viewed as the inverse of the estimated Cramer-Rao bound J ( θ | x ) 1 / 2 = ( J 1 ( θ | x ) ) 1 / 2 . Assuming that I ( θ | x ) is a positive definite symmetric (Hermitian) matrix, this is essentially the inverse of the local curvature of the log likelihood function and will behave locally as the inverse of asymptotic variation; it will be relatively flat where the likelihood is pronounced and non-informative in the sense that it preserves the local shape of the likelihood function about its mode as a key element of the shape of the posterior density, even under transformation of scale θ g ( θ ) .
In the multivariate setting, the Jeffreys prior is assumed to be p ( θ ) = | I ( θ | x ) | 1 / 2 [2]. This has the same basic interpretation regarding invariance, but as a determinant. The determinant of a square semi-positive definite matrix A can be written in several forms useful for statistical application.
In general, for a positive, semi-definite matrix A , the following properties hold:
t r ( A ) = j λ j ( A )
log d e t ( A ) = log j λ j ( A ) = j log λ j ( A )
where λ j ( A ) are the eigenvalues of A ,   t r ( A ) its trace. The multivariate Jeffreys prior can therefore be expressed:
log p ( θ ) = ( 1 / 2 ) log | d e t ( I ( θ | x ) ) |
= ( 1 / 2 ) log j λ j ( I ( θ | x ) )
= ( 1 / 2 ) j log λ j ( I ( θ | x ) )
where I ( θ | x ) is a positive-definite symmetric matrix. This links the eigenvalues of the information matrix to the invariance of the posterior density in regard to the scaling of the likelihood function.

8. Prior Density and Matching Information as a Non-Informative Scale

The likelihood function summarizes the information present in the model-data probability structure describing possible outcomes of the experiment in question. The need to consider likelihood properties in relation to prior density selection follows as the model parameter θ takes its meaning within the context of the probability model and related likelihood function.
If the goal of a non-informative prior is to ensure that the information regarding θ in the posterior density is dominated by the statistical information in the likelihood function, then requiring the second order derivative based condition:
2 ln p ( θ | x ) θ 2 | θ m l e = 2 ln L ( θ | x ) θ 2 | θ m l e + c
matches the curvature of the log-likelihood function about its mode to the curvature of the resulting log posterior density, up to an additive constant. This assumption indirectly imposes restrictions on the form and scale of p ( θ ) .
Definition 1: For a scalar parameter θ , define the concept of posterior information as the local curvature of the log-posterior about its mode:
2 ln p ( θ | x ) θ 2 = 2 θ 2 [ ln c + ln p ( θ ) + ln L ( θ | x ) ]
= 2 θ 2 ln p ( θ ) + 2 θ 2 ln L ( θ | x )
= 2 θ 2 ln p ( θ ) + J ( θ | x )
where J ( θ ) is the observed Fisher information J ( θ | x ) = 2 θ 2 ln L ( θ | x ) .
Definition 2: Let p ( θ | x ) be a unimodal posterior density function. A matching restriction on derivatives resulting in similar information is defined as:
2 ln p ( θ | x ) θ 2 | θ = θ m l e = 2 ln L ( θ | x ) θ 2 | θ = θ m l e + c
where c is a constant, not a function of θ , and θ m l e (x) is the mode of the likelihood. This implies that the local shape of the posterior density is given by the shape of the likelihood function.
Theorem 1: If the condition defined above holds, the following second order condition is true for the prior density:
2 θ 2 ln p ( θ ) | θ = θ m l e = c
where c is a constant.
Proof: The definition of the posterior density gives:
2 ln p ( θ | x ) θ 2 = 2 θ 2 [ ln c + ln p ( θ ) + ln L ( θ | x ) ] = 2 θ 2 ln p ( θ ) + 2 ln L ( θ | x ) θ 2
Assuming similar second order derivatives:
2 ln p ( θ | x ) θ 2 = 2 ln L ( θ | x ) θ 2 + c
and it follows that:
2 θ 2 ln p ( θ ) = c
Note that this implies a similar condition for all higher order derivatives.
For example,
3 ln p ( θ | x ) θ 3 = θ 2 ln p ( θ | x ) θ 2 = θ 2 ln L ( θ | x ) θ 2 + C = 3 ln L ( θ | x ) θ 3
and more generally, for m > 2 :
m ln p ( θ | x ) θ m = m ln L ( θ | x ) θ m
and for the prior:
m ln p ( θ ) θ m = 0
The second order condition above can be used to derive a general form for related priors.
Corollary 1: The family of non-excluded priors when matching higher order log-likelihood and log-posterior derivatives are probability densities of the form:
ln p ( θ ) = a θ 2 + b θ + d
p ( θ ) = exp a θ 2 + b θ + d
where d is a rescaling of the norming constant, a and b are hyper-parameters. Some of these can be zero.
This defines a family of possible priors which includes the normal distribution. A reasonable restriction on the parameters a , b , d is to require the prior to be well defined within the context of large sample regularity conditions.
Examples of acceptable prior densities in this context are all prior densities p ( θ ) such that the second derivative of ln p ( θ ) is a constant. The normal, exponential and uniform distributions satisfy this restriction. Technically, so do flat or highly non-informative priors on restricted support in the parameter space.
Theorem 2: Priors that do not meet the matching information restriction above include:
p ( θ ) exp ( θ m ) , m > 2
p ( θ ) θ m , m > 2
p ( θ ) sin ( g ( θ ) )
p ( θ ) cos ( exp ( θ 3 ) )
p ( θ ) a m θ m ; m 3
Proof: Follows from applying the previous result.
Prior densities with third or higher order non-zero derivatives are ruled out in this setting. Interestingly, if a selected prior density is based on a continuous but non-differentiable function, for example a fractal prior, which may occur if belief is assessed by physical measurement of individual responses, for example observed EEG assessments, rather than simple betting behavior, then according to the Weierstrass theorem this can be approximated by a higher order polynomial which is differentiable, but would not be acceptable here if the order of the polynomial were 3 .

9. Similarities to the A R A Utility Based Measure

It is interesting to note that probability as a measure that processes information is not unique. In regard to providing a measure of belief for a specific set of θ values, the probability measure p ( θ ) is equivalent to c p ( θ ) , where c is any constant, since this measure gives the same set of preferences in regard to possible θ values as p ( θ ) . This is clearly seen on a logarithmic scale. The use of a probability scale ( c = 1 ) is a choice with many advantages, but it is only one choice among many to measure relative levels of belief in relation to information for θ .
The use of Bayesian probability is very similar to the broader concept of utility where outcomes are ranked in terms of relative perceived satisfaction [5]. Here, weighting the perceived probabilities of occurrence for possible values for θ are of interest. Interestingly, utility is itself an ordinal, not a cardinal or absolute scale [23]. It is more similar to likelihood in interpretation when using probability measures to assess information. Note that some of the technical limitations of a utility related theory when a model uses many individual model parameters θ i in the statistical model are adjusted for in the Bayesian setting by the technical requirement of exchangeability [2].
The form of non-excluded information similar priors given above supports a scale of non-informativity related to utility. Taking the information similar prior density to be of the general form p ( θ ) = exp ( a θ 2 + b θ + d ) , this satisfies the following general condition:
p ( θ ) p ' ( θ ) = k
or
θ ln p ' ( θ ) = k
a restriction on the form of the prior that results from the initial information prior restriction. In this sense, information similar prior densities are constant and non-informative on this scale.
Interestingly, the utility-based Pratt-Arrow A R A relative risk aversion measure [24] is similar in structure and can be defined as:
A R A ( u ) = d d u log [ d d u I ( u ) ] = d 2 d u 2 I ( u ) d d u I ( u )
where I ( u ) is a function representing utility and I ' ( u ) > 0 for all u .
If the A R A measure of relative risk is constant, a similar condition is obtained:
A R A ( u ) = d 2 d u 2 I ( u ) d d u I ( u ) = K
A constant value for the A R A measure implies a stabilization in terms of risk avoidance behavior, interpreted in relation to
p ( θ ) p ' ( θ ) = k
choosing a prior density that avoids differences between the statistical information obtained using the second-order properties of the posterior and the second-order properties of the likelihood function.
Note that when employing the likelihood function and model structure to guide the selection of p ( θ ) a consideration of sample size is necessary. The use and interpretation of frequentist sampling distributions to define likelihood functions requires caveats regarding large sample effects, including central limit theorems. For example, distributions such as the gamma and Weibull distributions, once standardized, excluded in small samples, are not excluded in large samples.

10. Kullback-Liebler Distance

Kullback-Liebler distance [25] is a measure of the distance between two probability densities p and q . It is closely related to measures of entropy. It can be written:
D ( p , q ) = log p ( x ) q ( x ) p ( x ) d x
= E p log p ( x ) q ( x )
It can be used here to assess the effects of different prior densities on the resulting posterior density. It is especially useful if the densities in question are members of the exponential family p ( x | θ ) . Note that the difference between two densities in the exponential family defined by the K-L distance is typically called the deviance.
For example, taking a common likelihood and two prior densities p 1 and p 2 , the distance between the two resulting posterior densities p and q is given by
D ( p ( θ | x ) , q ( θ | x ) ) = log p ( θ | x ) q ( θ | x ) p ( θ | x ) d θ
= log c 1 p 1 ( θ ) L ( θ | x ) c 2 p 2 ( θ ) L ( θ | x ) p ( θ | x ) d θ
= log c 1 p 1 ( θ ) c 2 p 2 ( θ ) p ( θ | x ) d θ
= log ( c 1 / c 2 ) + log p 1 ( θ ) p 2 ( θ ) p ( θ | x ) d θ
= C + E θ | x log p 1 ( θ ) log p 2 ( θ )
where E θ | x [ ] is taken with regard to p ( θ | x ) = c p 1 ( θ ) L ( θ | x ) . This implies the following theorem.
Theorem 3: Given two distinct prior densities and a common likelihood, the distance between the resulting posterior densities is a function of the difference in the posterior mean rate of change of the difference in the l o g prior mean values as a function of θ , assuming regularity conditions support the interchange of differentiation and integration. For the case of two priors satisfying Definition 2, the distance is a function of the first two posterior moments.
Proof: Consider two information similar priors and a common likelihood function. The common likelihood function cancels out on the l o g scale. Using the above result, the Kullback-Liebler distance between the posterior densities is given by:
D ( p ( θ | x ) , q ( θ | x ) ) = C + E θ | x log p 1 ( θ ) log p 2 ( θ )
= C + E θ | x a 1 θ 2 + b 1 θ + d 1 a 2 θ 2 + b 2 θ + d 2
= C + E θ | x ( a 1 a 2 ) θ 2 + ( b 1 b 2 ) θ + ( d 1 d 2 )
= C ' + ( a 1 a 2 ) E θ | x ( θ 2 ) + ( b 1 b 2 ) E θ | x ( θ )

11. Multiparameter Extensions

The considerations above can be extended to multivariate models and are briefly mentioned here. In multiparameter settings, where θ = ( θ 1 , . . . , θ p ) , the matching of second derivates can be expressed:
2 ln p ( θ | x ) θ j 2 | θ m l e = 2 ln L ( θ | x ) θ j 2 | θ m l e + c , j = 1 , . . . , p
The related prior density restrictions can be generalized via partial derivatives:
θ 1 ln p ' ( θ ) = k 1
θ 2 ln p ' ( θ ) = k 2
θ p ln p ' ( θ ) = k p .
where k i are constants. This is a specific form of non-informativity in higher dimensions.
If symmetry or independence is useful, p ( θ )   = j = 1 p p j ( θ i ) can be assumed. If the cross derivatives can be set equal to zero, multiparameter priors can be taken with the general form:
ln p ( θ ) = a j θ j 2 + b j θ j + d j
This rules out multivariate prior distributions with factors of θ j that are higher order polynomials ( k > 2 ) or transcendental functions such as:
p ( θ ) ( sin ( θ 1 ) , . . . , sin ( θ p ) )
or
p ( θ ) ( exp ( exp ( θ 1 ) ) , . . . , exp ( exp ( θ p ) ) ) .
The resulting posterior density in the large sample vector parameter setting is a multivariate distribution and can be expressed generally as:
p ( θ | x ) = c exp a i θ i 2 + b i θ i + d i L ( θ | x )

12. Discussion

Many applications of Bayesian methods view the selection of p ( θ ) as primarily subjective in nature. The parameter θ is typically defined within the context and constraints of the probability model, inferential approach and related likelihood function. The choice of prior however, requires justification within the scientific or medical context being studied.
When model-based restrictions on the prior are assumed, for example in relation to model symmetry or likelihood properties, the set of prior densities available for subjective choice reflects the context of the science and related statistical model that defines θ .
In particular, if the posterior information, as characterized by the local curvature of the posterior density p ( θ | x ) , is assumed proportional to the local curvature of the log-likelihood function L ( θ | x ) , a type of matching condition, then exclusionary restrictions affect the form of the prior density p ( θ ) that is available to the researcher.
The class of prior densities available to the researcher within this matching restriction are non-informative in a unique manner, reflecting a structure similar to that found in the A R A utility based measure related to the assessment of relative risk.
These considerations can be extended to multivariate statistical models, along with prior robustness, which can be examined applying the Kullback-Liebler measure of distance.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

Data Availability Statements are available in the section “MDPI Research Data Policies” at https://www.mdpi.com/ethics.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Brimacombe, M. Likelihood Methods in Biology and Ecology: A Modern Approach to Statistics; CRC Press, 2019. [Google Scholar]
  2. Bernardo, J.M.; Smith, A.F.M. Bayesian Theory; John Wiley and Sons Inc: New York, NY, 1994. [Google Scholar]
  3. Congden, P. Applied Bayesian Modelling; John Wiley and Sons: NJ, 2003. [Google Scholar]
  4. Casella, G.; Berger, R.L. Statistical Inference, 2nd ed.; Duxbury, CA, 2002. [Google Scholar]
  5. de Finetti, B. Logical foundations and measurement of subjective probability. Acta Psychol. 1967, 34, 129–145. [Google Scholar]
  6. O’Hagan, A.; Forster, J. Bayesian Inference. In Kendall’s Advanced Theory of Statistics, 2nd ed.; John Wiley and Sons: New York, 2010. [Google Scholar]
  7. Smith, A.F.M. Bayesian computational methods. Phil. Trans. Roy. Soc. Lond. A 1991, 337, 369–386. [Google Scholar] [CrossRef]
  8. Boos, D.D.; Monahan, J.F. Bootstrap Methods Using Prior Information. Biometrika 1986, Vol. 73(No.1), 77–83. [Google Scholar] [CrossRef]
  9. Scriccolo, C. Probability matching priors: a review. J. Ital. Stat. Soc. 1999, 8, 83. [Google Scholar] [CrossRef]
  10. Jeffreys, H. Probability and Scientific Method. Proc. R. Soc. 1934, A, 146, 9–16. [Google Scholar] [CrossRef]
  11. Stone, M. Right Haar Measure for Convergence in Probability to Quasi Posterior Distributions. Ann. Math. Stat. 1965, Vol. 36(No. 2), 440–453. [Google Scholar] [CrossRef]
  12. Robert, C.P.; Casella, G. Introducing Monte Carlo Methods with R; Springer: New York, 2010. [Google Scholar]
  13. Welch, B.L.; Peers, H.W. On formulae for confidence points based on intervals of weighted likelihoods. J. Roy. Stat. Soc. B 1963, 27, 1–8. [Google Scholar]
  14. Mukerjee, R.; Dey, D.K. Frequentist validity of posterior quantiles in the presence of a nuisance parameter: Higher order asymptotics. Biometrika 1993, 80, 499–505. [Google Scholar] [CrossRef]
  15. Tibshirani, R. Noninformative priors for one parameter of many. Biometrika 1989, 76, 604–608. [Google Scholar] [CrossRef]
  16. Tukey, J.W. Some examples with fiducial relevance. Ann. Math. Stat. 1957, 28(No.3), 687–695. [Google Scholar] [CrossRef]
  17. Geisser, S.; Cornfield, J. Posterior distributions for multivariate normal parameters. J. Roy. Stat. Soc. B 1963, 25, 368–376. [Google Scholar] [CrossRef]
  18. Helland, I.S. Statistical inference under symmetry. Intern. Stat. Rev. 2004, 72, 409–422. [Google Scholar] [CrossRef]
  19. Fraser, D.A.S. Necessary and adaptive inference. J. Amer. Stat. Assoc. 1976, 71. [Google Scholar] [CrossRef]
  20. Hannig, J. On generalized fiducial inference. Stat. Sin. 2009, 19(No. 2), 491–544. [Google Scholar]
  21. Bondar, J.V. A conditional confidence principle. Ann. Stat. 1977, 5(No.5), 881–891. [Google Scholar] [CrossRef]
  22. Johnstone, I.M. High Dimensional Bernstein-von Mises: Simple Examples. Borrowing Strength: Theory Powering Applications -- A Festschrift for Lawrence D. Brown. In Institute of Mathematical Statistics; 2010. [Google Scholar] [CrossRef] [PubMed]
  23. Varian, H.R. Microeconomic Analysis, 2nd ed.; W.W. Norton & Company, Inc.: New York, 1984. [Google Scholar]
  24. Pratt, J. Risk aversion in the small and in the large. Econometrica 1964, 32, 122–36. [Google Scholar] [CrossRef]
  25. Kullback, S. Information Theory and Statistics; J. Wiley and Sons: New York, 1959. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings