Preprint
Communication

This version is not peer-reviewed.

Assumptions, Borderline Cases, and Misapplications of the GRIM Test in Forensic Metascience

Submitted:

24 July 2026

Posted:

27 July 2026

You are already at the latest version

Abstract
Background: The Granularity-Related Inconsistency of Means (GRIM) test is a simple arithmetic method for detecting impossible means in integer-valued data. If observations are whole numbers and the sample size is known, a reported mean must be compatible with an integer total divided by that sample size. Purpose: This commentary summarises the assumptions required for valid GRIM use, explains borderline failures, and illustrates common misapplications using PubPeer and journal-debate examples. Conclusion: GRIM is strongest when applied to verified integer instruments, but it can mislead when used on age, continuous scales, composite scores, ambiguous sample sizes, or unclear rounding precision.
Keywords: 
;  ;  

1. Introduction

The GRIM test, introduced by Brown and Heathers, identifies means that cannot arise from integer-valued or discrete whole-number scores. For example, if 10 participants answer a 1-5 Likert item, the total score must be an integer, and the mean must be that integer divided by 10. A reported mean that corresponds to no possible integer total is arithmetically impossible under the stated conditions.
GRIM has become part of the forensic metascience toolkit, especially in post-publication review. It is often combined with p-value checks, standard deviations, effect sizes, and internal consistency. However, GRIM is not a general fraud detector. It is valid only when its assumptions are satisfied. If the variable is not integer-valued, if the reported mean is based on a transformed or composite score, or if the wrong sample size is used, GRIM can produce false-positive claims.

2. The GRIM and GRIMMER Tests

For a sample of N participants with integer-valued observations x_i, the sample mean is M = S/N, where S is the sum of all observations and must be an integer. If M is reported to D decimal places, standard rounding implies that the true mean lies in:
[M - 0.5 x 10^-D, M + 0.5 x 10^-D)
Multiplying by N gives the feasible interval for the integer total:
[(M - 0.5 x 10^-D)N, (M + 0.5 x 10^-D)N)
If at least one integer lies inside this half-open interval, the mean passes GRIM. If no integer lies inside, the mean fails.
GRIMMER extends this logic to standard deviations. After identifying integer totals compatible with the mean, GRIMMER asks whether the reported standard deviation is compatible with any possible integer dataset. This requires constraints on the sum of squared scores and is more sensitive to rounding, formula choice, and reporting precision than GRIM alone.

3. Core Assumptions and Violations

GRIM depends on five practical assumptions that should be stated before a result is interpreted.
Assumption Requirement Common violation
A1. Integer-scored data The variable was recorded as whole numbers Age in decimals, VAS pain, time, gestational age
A2. Mean of raw scores The mean is a direct average of integer scores Composite scores, proportions, normalised scales
A3. Correct sample size and usually less than 100 The relevant N is known Missing data, subgroup Ns, inconsistent tables
A4. Known precision Decimal places are unambiguous 3.4 versus 3.40; omitted trailing zeros
A5. Rounding convention Standard rounding is assumed Banker’s rounding or software-specific rules
The most serious problem is A1's failure. If the variable is continuous or quasi-continuous, a GRIM failure has no evidential value. Failures of A3 and A4 usually create uncertainty rather than definitive invalidity: the result should be reported as conditional on the assumed sample size and precision.

A1: Integer-Scored Data

The first common false-positive route is a violation of A1. In one PubPeer-style EMDR depression example, some variables were plausible GRIM targets, whereas others were not. Checklist-type symptom totals can be appropriate because they are usually scored as integers. For instance, a TSC-40 pretreatment mean of 56.47 with n = 10 gives the feasible total interval [564.65, 564.75), which contains no integer and therefore fails GRIM under the integer-score assumption.
Quality-of-life indices are more problematic. If a QLI score was transformed, averaged, weighted, or computed from non-integer components, GRIM is not applicable. An apparent failure would not show that the mean is impossible; it would show that the analyst may have imposed an integer scale on a variable that was not analysed as an integer total. Age creates the same problem. PubPeer comments often treat age as whole years, but age may be calculated from date of birth or recorded in decimal years. Age-based GRIM failures should therefore be written as conditional claims: valid if age was recorded in completed years, inconclusive if fractional ages were used.

A2: The “Smell of Sex” Debate

A different problem appears in the debate over the “Smell of Sex” studies. Wisman and Shrira reported experiments suggesting that men could process olfactory signals of women’s sexual arousal. The claim attracted scrutiny because the samples were small and the effects appeared large. Sakaluk’s commentary applied GRIM-style checks, SPRITE reconstruction, and broader statistical forensics. The initial GRIM logic treated the attractiveness outcomes as if each participant contributed one integer Likert rating. Under that reconstruction, several cells appeared arithmetically impossible.
The authors later replied that this used the wrong unit of aggregation. The reported attractiveness outcome was not a single raw integer rating per person. It was a composite formed from two or three integer ratings per odour, such as pleasantness and sexiness. Therefore, the relevant GRIM denominator was not simply N, but 2N or 3N, depending on the experiment. Once the composite structure was considered, some apparently impossible means became compatible with the scoring procedure. This is primarily an A2 violation: the items may be integers, but the reported mean was not the simple average of one raw integer score per participant. Composite scoring changes the granularity of the data and, therefore, changes the correct GRIM denominator.

A3. Correct Sample Size and the Loss of Power When N Is Large

GRIM also requires that the correct sample size is used for the specific mean being tested. This may be the full group size, but it may also be a smaller denominator if some participants had missing data or if the value comes from a subgroup. Using the wrong N can create either false-positive or false-negative results.
A second issue is that GRIM becomes less informative as N increases. For a mean reported to two decimal places, the feasible sum interval has a width N × 0.01 . When N = 100 , this interval has width 1; when N > 100 It is wider than 1. As a result, almost every reported two-decimal mean will contain at least one integer total and therefore pass GRIM automatically. This does not prove that the value is correct; it only means that GRIM no longer has enough granularity to detect the problem.
For example, if N = 100 and the reported mean is 3.47, the feasible total-score interval is:
[ ( 3.47 0.005 ) × 100 , ( 3.47 + 0.005 ) × 100 ) [ 346.5 , 347.5 )
This interval contains the integer 347, so the value passes GRIM. The same will happen for essentially any two-decimal mean when N 100 . Thus, large sample sizes increase the risk of false negatives: erroneous or fabricated means may pass simply because the reporting precision is too coarse relative to the sample size.
More generally, for means reported to D decimal places, GRIM loses power when:
N 10 D
So for one decimal place, GRIM becomes weak at N 10 ; for two decimal places, at N 100 ; and for three decimal places, at N 1000 .

A4. Known Decimal Precision

GRIM also depends on knowing how many decimal places were actually reported. This seems simple, but it can be ambiguous. A value printed as 3.4 may mean that the authors rounded to one decimal place, or it may be a displayed version of 3.40 with the trailing zero omitted. These two interpretations produce different GRIM intervals.
For example, with N = 16 , a reported mean of 3.4 to one decimal place has a wide feasible interval:
[ ( 3.4 0.05 ) × 16 , ( 3.4 + 0.05 ) × 16 ) [ 53.6 , 55.2 )
which contains integers and may pass. But if the intended value was 3.40 to two decimal places, the interval becomes:
[ 54.32 , 54.48 )
which contains no integer and fails. Therefore, when trailing zeros are omitted, GRIM results should be reported as conditional on the assumed precision.

A5. Rounding Convention

The standard GRIM test assumes conventional rounding, where values exactly halfway between two reported numbers are rounded upward. However, some software uses banker’s rounding, also called round-half-to-even. Under banker’s rounding, a value ending exactly in 5 may round up or down depending on whether the nearest even digit is above or below.

4. Borderline Failures

The GRIM interval is half-open: the lower bound is included, but the upper bound is excluded. This matters because a value exactly at the upper rounding boundary would normally round to the next higher reported mean. For example, [167.60, 168.00) contains no integer. The integer 168 is exactly the excluded upper bound, while 167 is below the lower bound. This is a true GRIM failure, but it can be labelled a borderline failure because it lies immediately next to a feasible integer value. Borderline does not mean 'pass'; it means 'failure close to the boundary'.
A PubPeer example is the trial Effects of relaxation on self-esteem of patients with cancer. The familial self-esteem subscale was described as the sum of eight binary items, with 40 participants per arm. The reported means 5.67 and 5.77 imply feasible total-score intervals of [226.60, 227.00) and [230.60, 231.00). In both cases, the nearest integer lies exactly at the excluded upper bound, so both are GRIM-inconsistent but best described as borderline failures. The PubPeer comment described these values as impossible under the stated scoring scheme; the article was later retracted after concerns about data and p-values.

5. Practical Recommendations

Practical Recommendations
  • Verify the scale first. Apply GRIM only to variables that are clearly integer-scored, such as Likert items, symptom totals, counts, or binary-item sums.
  • Do not assume age is integer-valued. Age, gestational age, disease duration, and time variables may be recorded in decimals. Treat GRIM failures for these variables as conditional.
  • Check how the score was constructed. Composite or averaged scales may require a denominator such as 2 N , 3 N , or k N , not simply N .
  • Use the correct sample size. The relevant N is the number contributing to that specific mean, not necessarily the total randomised sample.
  • Remember that large  N  weakens GRIM. With two decimal places, GRIM loses power around N 100 , increasing false negatives.
  • Treat borderline cases as failures with caution. If the only integer lies at the excluded upper bound, the result fails GRIM but should be labelled borderline.
  • Do not overinterpret one failed GRIM test. GRIM shows arithmetic inconsistency under the assumptions; it does not, by itself, prove fabrication.
  • Report results conditionally. Use phrasing such as: “This means is GRIM-inconsistent if the variable was integer-scored and N is correct.”
  • Combine GRIM with other checks. Stronger evidence comes from converging anomalies: incorrect p-values, implausible SDs, duplicated values, implausible effects, or inconsistent sample sizes.

6. Conclusions

GRIM is powerful because it converts a reporting claim into a simple arithmetic question: could these numbers exist? When applied to verified integer instruments, it provides strong evidence of reporting error or data irregularity. But its simplicity is also its danger. Misunderstanding the scale, denominator, precision, or rounding convention can turn GRIM from a forensic tool into a source of false accusations. The best use of GRIM is therefore not merely computational but methodological: verify the assumptions first, then interpret the arithmetic.

References

  1. Brown, N. J. L.; Heathers, J. A. J. The GRIM Test: A Simple Technique Detects Numerous Anomalies in the Reporting of Results in Psychology. In Social Psychological and Personality Science; 2017. [Google Scholar] [CrossRef]
  2. Anaya, J. The GRIMMER test: A method for testing the validity of reported measures of variability. PeerJ Preprints. 2016. Available online: https://peerj.com/preprints/2400/.
  3. Allard, A. Analytic-GRIMMER: A new way of testing the possibility of standard deviations. 2018. Available online: https://aurelienallard.netlify.app/post/anaytic-grimmer-possibility-standard-deviations/.
  4. Wisman, A.; Shrira, I. Sexual Chemosignals: Evidence that Men Process Olfactory Signals of Women’s Sexual Arousal. In Archives of Sexual Behavior; 2020. [Google Scholar] [CrossRef] [PubMed]
  5. Sakaluk, J. K. Commentary on the sexual chemosignals studies. In Archives of Sexual Behavior; 2020. [Google Scholar] [CrossRef] [PubMed]
  6. Wisman, A.; Shrira, I. Additional Notes of Caution: A Reply to Sakaluk (2020). In Archives of Sexual Behavior; 2022. [Google Scholar] [CrossRef] [PubMed]
  7. Harorani, M.; et al. Effects of relaxation on self-esteem of patients with cancer: a randomized clinical trial. In Supportive Care in Cancer; 2020. [Google Scholar] [CrossRef] [PubMed]
  8. PubPeer record for Harorani et al. self-esteem trial. Available online: https://pubpeer.com/publications/97CD1B9FD075CACD50B0E0760A66FF.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings