Submitted:
22 August 2016
Posted:
23 August 2016
You are already at the latest version
Abstract
In this data deposit, I describe a dataset that is the result of content mining 167,318 published articles for statistical test results. As a result of this content mining, 688,112 results from 50,845 articles were extracted. In order to provide a comprehensive set of data, the statistical results are supplemented with metadata from the article they originate from. The dataset is provided in a comma separated file (CSV) in long-format. For each of the 688,112 results, 20 variables are included, of which seven are article metadata and 13 pertain to the individual statistical results (e.g., reported and recalculated p-value). A five-pronged approach was taken to generate the dataset: (i) collect journal lists, (ii) spider journal pages for articles, (iii) download articles, (iv) add article metadata, and (v) mine articles for statistical results.
Keywords:
nhst
; p-values
; apa
; content mining
; tdm
; errors
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.