The goal of a Bayes factor is to provide a measure of evidence based on how belief in a hypothesis changes after observing the data. See [
4] for an in-depth discussion. To motivate the derivation of a Bayes factor, suppose we observe
, or equivalently observe
. Supposed that under the hypothesized model,
. The development of a Bayes factor requires an alternative hypothesis stated as a predictive model.
Before the creation of an alternative model, let’s think about the interpretation of
. Since we set
and
, then
is on a standard deviation scale. A parameter defined in such a way corresponds to Cohen’s effect size. A scale commonly used for interpreting Cohen’s effect size is displayed below.
The reason we need to know typical values for
is that we need to model
under the alternative hypothesis. We will define the effect size for our alternative model using a Bayesian prior probability distribution. We write
where
can be chosen to correspond to the expected effect size. If
is relatively small, the alternative is modelling a small effect size. If
is relatively large, the alternative is modelling a large effect size. Since the data distribution is modelled as
then the predictive distribution under the alternative becomes
Let
represent Bayesian prior probabilities on the respective hypotheses / models. That is,
where
and
denote the competing models. Using Bayes Rule, we write the posterior probability on model
as
The above probability statement leads to
Let
denote the prior odds. Let
denote the posterior odds. Then we define the Bayes Factor as
We will write
as predictive densities evaluated at the observed
. To make the connection to a likelihood ratio statistic, we will use the notation for the likelihood function. Let
Then
can be represented as
and
can be represented as
where
represents the prior on
under the alternative model
. Now we can write
Note how the BF
01 depends on the data only through the likelihood function. Thus, a Bayes factor provides a measure of evidence which satisfies the likelihood principle. Rather than using
for the denominator, as in the definition of a likelihood ratio statistic, the Bayes factor uses a weighted average of the likelihood function, where the weights are determined by prior distribution used in defining the alternative model. Since maximizing the denominator for the Bayes factor leads to the likelihood statistic, we see that the likelihood ratio statistic overstates the evidence against the hypothesized model. Both evidence measures are formed as ratios of the likelihood function, so both can be interpreted on the same scale.
Refer back to Fisher’s scale for interpreting evidence strength displayed in
Table 1 and
Table 2. Instead of Fisher’s scale, Jeffreys’ scale is more often used for interpreting strength of evidence measured by a Bayes factor [
3]. As with the p-value and the likelihood ratio statistic, small values of BF
01 represent evidence against the hypothesized model.
Table 6.
Jeffreys’ evidence scale.
|
to 1 |
not worth a bare mention |
|
to
|
positive |
|
to
|
strong |
|
to
|
very strong |
Jeffreys’ evidence scale is more conservative than Fisher’s scale. We can take this difference in interpretation as a reminder that there is no unique answer for scaling a measure of evidence strength.