Preprint
Article

This version is not peer-reviewed.

Reading the Infodemic Against the Epidemic: A Real-Risk Benchmark for Inflated and Downplayed Disease Risk on Reddit Within the 2026 World Cup Window

Submitted:

22 August 2026

Posted:

24 August 2026

You are already at the latest version

Abstract
During the deadly hantavirus outbreak of May 2026, the loudest online response called it a hoax, while measles, the one rising threat, drew little attention. That inversion is the half of the infodemic we rarely measure: not amplified fear, but the quiet downplaying of a real risk. Detecting it requires what the field lacks: an external benchmark for the true risk. Anchoring one in the epidemiological and mass-gathering literature, with the 2026 FIFA World Cup run-up as a dated window, two coders judged Reddit claims about measles (high), hantavirus (low), and Ebola (near zero) for accuracy and direction (Cohen's kappa 0.72 and 0.68). Of 66 coded claims, 41 were inaccurate, and distortion ran both ways: 24 downplayed risk and 17 inflated it (exact binomial p = 0.35), with downplaying at least as common as inflating. Attention ran opposite to danger, hantavirus drawing the most claims (44 of 66) and measles the fewest (8): the infodemic did not mirror the epidemic. The curated, single-platform sample makes these patterns exploratory, but they suggest that defenses built against panic alone miss its mirror image, the denial of a real threat that prior work links to weaker protection among the most exposed.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Modern disease outbreaks now unfold amid a high volume of information. The term infodemic captures this condition: an overabundance of information during an outbreak, some accurate and some not, that makes trustworthy guidance hard to find when people need it most. The scientific study of this phenomenon, infodemiology, was named by Eysenbach almost two decades ago. The World Health Organization now recognizes it as an emerging field and a critical area of practice during a pandemic [1]. The institutional response took shape quickly during COVID-19. Eysenbach [1] proposed four pillars for managing an infodemic: monitoring information (infoveillance), building eHealth and science literacy, encouraging knowledge refinement through fact checking and peer review, and ensuring timely knowledge translation. Around the same time, the World Health Organization convened a crowdsourced technical consultation that produced an agreed framework for infodemic management [2].
A central reason the infodemic is hard to control is that false information often travels better than the truth. Vosoughi, Roy, and Aral [3] studied about 126,000 news stories shared by some three million people on Twitter between 2006 and 2017, each classified by six independent fact-checking organizations. Falsehood spread significantly farther, faster, deeper, and more broadly than true news across every category; the gap was widest for political content. False stories were also more novel and more likely to provoke fear, disgust, and surprise; the main drivers were people, not automated accounts. Writing in the same issue of Science, Lazer et al. [4] framed fake news as a subject for systematic, interdisciplinary study.
Such misinformation is not a fringe problem. Reviewing 69 studies, Suarez-Lledo and Alvarez-Galvez [5] found it widespread across social media and clustered around a recurring set of health topics. Belief is patterned rather than random. Surveying Americans on eleven false or conspiratorial COVID-19 ideas, Enders et al. [6] found that these beliefs grouped together, tracked traits such as distrust of scientists, and were linked to behavioral intentions such as vaccination. The consequences reach beyond belief. As van der Linden [7] summarizes, exposure to misinformation can undermine vaccination uptake and compliance with public-health guidance, which makes the infodemic a direct concern for outbreak control. The stakes are highest when a real, ongoing outbreak meets a large, mobile crowd, which is the situation this study examines.
Research has also begun to identify responses. van der Linden [7] organizes the field around three questions: why some people are more susceptible to misinformation, how it spreads through online networks, and how people can be given psychological immunity through interventions such as inoculation or prebunking. This work largely treats misinformation as one-directional, with false content competing against true content. It is less suited to explaining why the same outbreak can produce both exaggerated alarm and dismissive denial, or why a real threat can be quietly downplayed rather than loudly distorted. Accounting for distortion in both directions calls for a theory of how societies amplify and attenuate risk signals. That theory is the Social Amplification of Risk Framework.

2. Theoretical Background

2.1. The Social Amplification of Risk

This study is grounded in the Social Amplification of Risk Framework (SARF), first set out by Kasperson et al. [8]. SARF describes how a risk signal, such as a news report, a rumor, or a social-media post, travels through a series of amplification stations. These stations include the media, scientific institutions, government agencies, and personal social networks. Each one can strengthen or weaken the signal, producing ripple effects that reach beyond the original hazard. Its central claim is that the social processing of a risk, not the physical hazard alone, shapes how people perceive and respond to it. SARF was bidirectional from the start. Kasperson et al. [8] described both amplification, in which concern rises above what the hazard warrants, and attenuation, in which it falls below. Renn et al. [9] then showed empirically that hazard events can heighten or attenuate risk perception, with responses tied more to exposure than to the magnitude of the hazard. Pidgeon, Kasperson, and Slovic [10] later consolidated the framework in an edited volume. This two-directional structure is what makes SARF the right home for a study of disease misinformation, where some claims exaggerate a threat and others dismiss a real one.
More recent work applies SARF to social media, even though the framework predates global digital platforms [11]. Moussaid, Brighton, and Gaissmaier [12] showed experimentally that risk information is amplified and distorted as it passes along diffusion chains from person to person. Gallotti et al. [13] developed an Infodemic Risk Index to track unreliable information during COVID-19. The attenuation side is documented directly as well. Gassen et al. [14] found that unrealistic optimism leads people to underestimate their own COVID-19 risk. Yet the literature remains unbalanced. Most empirical infodemic studies measure amplification, the spread and reach of alarming or false content. Attenuation, the downplaying of a real threat, receives less direct attention, though recent SARF work has begun to revisit it [15]. The closest study to the present one, Hopfer et al. [16], coded both amplification and attenuation of COVID-19 risk on Twitter. That study was limited to a single disease and platform and did not benchmark claims against an external measure of real epidemiological risk.

2.2. Disease-Specific Misinformation: Ebola, Measles, and Hantavirus

The three diseases at the center of this study, Ebola, measles, and hantavirus, sit at very different points on the misinformation map [5]. The loud, vaccine-linked, fear-driven cases are well studied, while quieter, lower-salience risks are barely examined.
Ebola misinformation has been studied since the 2014 to 2015 West African epidemic. Oyeyemi, Gabarron, and Wynn [17] gave an early warning of its spread on Twitter. Fung et al. [18] identified false treatment claims circulating within the first days of the response, including claims that drinking or bathing in saltwater and ingesting a product marketed as Nano Silver could prevent or cure the disease. In a content analysis of more than 72,000 tweets from the 2014 United States Ebola scare, Sell, Hosangadi, and Trotochaud [19] found that about 10 percent of non-joke tweets carried false or partially false information. The most common rumor centered on government conspiracy. Misinformation was more likely than accurate content to be political (36 versus 15 percent) and to provoke discord (45 versus 10 percent). Ebola misinformation runs in both directions. Surveying 961 adults during the 2018 to 2019 North Kivu outbreak in the Democratic Republic of the Congo, Vinck et al. [20] found that about a quarter (25.5 percent) did not believe the outbreak was real, and that low institutional trust and belief in misinformation were associated with lower uptake of preventive behaviors, including vaccination and formal care. Denial attenuates a real and severe threat, while saltwater and conspiracy rumors amplify fear, both within the same epidemic.
Measles misinformation is dominated by the vaccine debate, much of it traceable to a single retracted paper. Wakefield et al. [21] suggested in The Lancet a possible link between the measles, mumps, and rubella vaccine and developmental disorders. The Lancet fully retracted the claim in 2010 [22]. The claim has since been discredited but still circulates. The consequences are measurable. Phadke et al. [23] found that a substantial proportion of post-elimination United States measles cases occurred in intentionally unvaccinated people, and that vaccine refusal was associated with increased measles risk. Majumder et al. [24] linked the 2015 California theme-park outbreak to substandard vaccination compliance. Social media has extended this reach. Wilson and Wiysonge [25] found that organized anti-vaccination activity predicted the belief that vaccines are unsafe and that foreign disinformation tracked falling coverage, with a one-point rise on their five-point scale corresponding to about a two-percentage-point drop in annual coverage. Mapping confidence across 149 countries, de Figueiredo et al. [26] documented steep recent declines alongside recoveries elsewhere. Measles misinformation often does both at once: it amplifies fear of the vaccine while playing down the danger of the disease itself.
Hantavirus presents the opposite situation. It is a real and sometimes fatal rodent-borne disease with no vaccine and no specific treatment. Yet it has historically drawn very little public or online attention; we are not aware of any published study analyzing hantavirus misinformation on social media. This changed during the period examined here. In May 2026, an outbreak of the Andes strain aboard the MV Hondius drew a sudden and atypical wave of online claims, squarely within the World Cup window. The wave was prominent enough that mainstream outlets published fact-checks debunking the conspiracy theories that formed around it [27,28]. The small existing social-science literature concerns offline knowledge rather than online distortion. Harris and Armien [29], in a knowledge-attitudes-practices survey of 124 residents in Tonosi, Panama, a district with 107 hantavirus cases and four deaths in 2018, found generally high awareness but persistent gaps, with one in five residents unaware of how the virus is transmitted. For a study of risk distortion, this pattern is itself informative. When the May 2026 outbreak finally brought attention to a genuine but normally ignored risk, much of that attention worked to deny or wave away the threat rather than exaggerate it, making hantavirus a clean test case for attenuation.
Together the three diseases form a gradient of typical online attention. It runs from abundant, vaccine-centered measles misinformation, through well-documented Ebola misinformation that mixes fear with denial, to essentially unstudied hantavirus. The research literature follows the same gradient toward the loud, amplifying cases. Studying all three inside a single mass-gathering context brings the quiet, low-salience end into view alongside the loud end and tests the central thesis directly: that online narratives can distort disease risk in both directions, so that the infodemic does not simply mirror the epidemic.

2.3. Mass-Gathering and Travel-Medicine Context

The 2026 FIFA World Cup plays two specific roles in this study. It is the dated window that put these three diseases in front of the same audience at the same time, and it is the source of a documented real-risk benchmark against which online claims can be judged. The timing is what makes the question pressing. A mass gathering of this scale carries documented infectious-disease risk, and in May 2026 a real and ongoing hantavirus outbreak met a sudden surge of online claims inside that same window. When a genuine outbreak and a large gathering coincide, the way people distort the risk online stops being an academic concern and starts to shape how a real crowd reads a real threat. We do not assume it is the main topic of the posts. As the composition-of-the-analytic-set results show, the conversation in this window was driven by the May 2026 hantavirus outbreak. A tournament of this size brings people from many countries into a few host cities over several weeks. A recognized body of research treats such events as a distinct public-health concern. Writing in The Lancet, Memish et al. [30] describe mass-gatherings medicine, launched as a discipline with a 2014 Lancet Series. They report that planned events between 2013 and 2018 were associated with major public-health challenges, among them the transmission of infectious diseases and antibiotic-resistant bacteria. The football World Cup specifically has drawn its own analysis. Al-Tawfiq, Gautret, and Schlagenhauf [31] examined infection risks at the 2022 FIFA World Cup in Qatar. A PRISMA-guided scoping review by Alhussaini et al. [32] included 34 studies and charted risk factors and prevention strategies across the pre-event, during-event, and post-event stages. Together, this work establishes infectious-disease transmission at large gatherings as a genuine, recurring, and actively planned-for risk rather than a hypothetical one.
Because these real disease risks are documented, online claims can be compared against them rather than judged in a vacuum. A claim that a minor or absent threat will cause a World Cup catastrophe inflates the documented risk, while a claim that dismisses a real and preventable threat plays it down. The mass-gathering literature thus supplies the external reference point on which this study's coding scheme relies. It lets each claim be judged for its direction relative to the real epidemiological risk and connects the infodemic and epidemic sides of the study.

2.4. Research Gap and Contribution

Across the closest prior studies, summarized in Table 1, each covers some of these design features but none covers them all. The studies that code both directions of distortion do so for a single disease on a single platform and without an external real-risk benchmark, while the studies that span many diseases or large gatherings measure only the spread of alarm. No single paper combines all five features of this design: multiple diseases, Reddit data, both directions of distortion, a real epidemiological-risk benchmark, and a mass-gathering context. This is the gap. The present study addresses it. It applies SARF to Ebola, hantavirus, and measles on Reddit within the 2026 FIFA World Cup and codes each claim for its direction relative to the real epidemiological risk. In doing so, it gives the attenuation half of the framework the same systematic attention usually reserved for amplification. No prior paper covers Reddit or a mass gathering. None meets all five criteria. The aim of this study is to describe how disease risk was distorted online in the run-up to the 2026 FIFA World Cup, and to ask whether that distortion ran in one direction, toward alarm, or in both directions, alarm and denial. To meet this aim, the study has four objectives. The first is to identify the disease-risk claims about Ebola, hantavirus, and measles that circulated on Reddit during the collection window. The second is to judge each claim for accuracy against a fixed real-risk benchmark. The third is to record whether each claim inflated or downplayed the real risk, giving attenuation the same attention as amplification. The fourth is to compare the volume of online attention each disease drew with its real epidemiological risk.

3. Materials and Methods

This study uses manual content coding of Reddit posts and comments, framed by the Social Amplification of Risk Framework (SARF), to describe how disease risk was distorted online ahead of the 2026 FIFA World Cup. Two trained coders applied a fixed rubric to each claim.

3.1. Study Design

The design is a descriptive, theory-guided case study. The unit of analysis is the individual claim as it appears in a single Reddit post or comment. Following SARF, which holds that social processes can both amplify and attenuate a hazard signal, the study treats inflation and downplaying of risk as two directions of the same distortion process rather than as separate phenomena. In SARF terms, claims that Inflate risk represent amplification and claims that Downplay it represent attenuation. The May 2026 hantavirus 'Plandemic 2.0' wave is the lead case; Ebola and measles serve as comparison cases that bracket the risk spectrum (a near-zero real risk and a high, rising real risk). The study asks four descriptive questions: what disease-risk claims circulate in this window, whether they are accurate, whether they inflate or downplay the real risk, and how the volume of online attention compares with the real epidemiological risk of each disease. The core unit is the per-claim direction judgment against the fixed real-risk benchmark. The World Cup defines the collection window and supplies that benchmark (see the Coding Scheme). The direction finding therefore does not depend on how many claims mention the tournament or on tournament-specific volume. In reporting terms the study is a cross-sectional observational content analysis: we did not contact or interact with users but observed and coded public content as it already stood. We report it in the spirit of the STROBE guidance for observational studies, describing the setting, the unit of analysis, the variables, the data sources, and the analytic approach. Selection of items followed a PRISMA-style flow, from the records first retrieved through each screening stage with its exclusions to the final analytic set, shown in the study-flow diagram (Figure 1).

3.2. Data Source and Collection

Data are public posts and comments collected from Reddit using PRAW (the Python Reddit API Wrapper) in read-only mode. Reddit was chosen for three reasons: it is far less studied than Twitter for health misinformation, its topic-based communities map naturally onto the amplification stations described by SARF, and its longer, threaded discussions surface explicit claims and reasoning that short-form platforms often omit. The collection script queried 46 search terms covering Ebola, hantavirus, and measles, together with general mass-gathering and cross-disease conspiracy language (for example "World Cup Ebola," "MV Hondius hantavirus," "hantavirus scripted pandemic," "measles vaccine autism," and "World Cup pandemic 2026"). The conspiracy and health-misinformation vocabulary in these terms was informed by prior research and contemporary online conspiracy discourse [5,6,33,34,35]. Meanwhile, the disease, venue, and World Cup terms specific to the 2026 tournament were developed by the authors. The complete 46-term list is given in Appendix A. Each term was run across the site-wide r/all stream plus 13 named communities spanning conspiracy, skeptic, public-health, and sports spaces (including r/conspiracy, r/skeptic, r/publichealth, r/soccer, r/worldcup, r/news, and r/worldnews). Searches used three sort orders (relevance, new, and top), retrieving up to 60 results per search and up to six top comments per on-topic post. Items were kept only if their own timestamp fell inside the study's collection window of March 1 to June 3, 2026, the run-up to the tournament (which runs June 11 to July 19, 2026). The collection was run on June 3, 2026. After de-duplication by Reddit item ID, the pull yielded 8,264 items (3,408 posts and 4,856 comments) drawn from 1,256 subreddits. The three inputs to this list, the conspiracy and health-misinformation vocabulary, the disease terminology for Ebola, hantavirus, and measles, and the event-specific venue, tournament, and outbreak terms, were assembled through a deliberate, systematic process rather than compiled ad hoc. The timestamp used for the collection window was each item’s created_utc field as returned by the Reddit API, which the platform records in Coordinated Universal Time; the March 1 and June 3, 2026 boundaries are therefore both in UTC.
Each collected item carried three automated keyword hints: a keyword-based disease guess, a "CHECK" flag that fired on conspiracy-style language (for example "plandemic," "hoax," "cover up," or "depopulation"), and a tag marking whether the matching search term was tournament-tied or general. These hints were triage aids to speed reading. They are not codes and were not used as verdicts. Every item that entered the analysis was read and judged by a human coder.

3.3. Screening and Curation

The 8,264 raw items were reduced to the coding set in two stages. First, the disease and misinformation-language hints were used to narrow the pull to a candidate pool of 163 items that mentioned one of the three diseases or a World Cup health risk. These 163 candidates were then read individually. Items that were duplicates, automated daily digests, news-bot reposts, or off-topic (for example match results in sports communities) were removed and logged, with reasons, in a "Dropped (with reason)" record (52 items: 30 digests, 4 news-bot reposts, 10 off-topic, and 8 duplicates). The remaining 111 items formed the review set. The screening and curation described in this paragraph were carried out by the first author, who applied the inclusion and exclusion criteria uniformly and logged every excluded item with its reason.
Gate 1 of the coding rubric was then applied to the 111 claims (see Section 3.4). A claim was kept only if it was about Ebola, hantavirus, or measles, or about a World Cup health risk, and made a checkable statement. Sixty-six claims passed this gate and formed the analytic set (n = 66). The remaining 45 were marked Skip and excluded from the accuracy and direction analyses.
The screening applied explicit inclusion and exclusion criteria. An item was included if it (i) was written in English; (ii) mentioned Ebola, hantavirus, or measles, or a named World Cup health risk; (iii) fell inside the March 1 to June 3, 2026 collection window; (iv) was an original post or a substantive comment rather than an automated digest or bot repost; and (v) made a checkable statement about disease risk. An item was excluded if it (i) matched no target disease or named World Cup health risk; (ii) was an automated daily digest or a news-bot repost; (iii) was off-topic, such as a match result in a sports community; (iv) was a duplicate, content repost, or cross-post of an item already in the set; or (v) made no checkable claim about disease risk, such as a bare link, image, or expression of feeling. These criteria were applied across the collection and screening stages: the language, keyword, and window requirements at collection; the digest, bot, off-topic, and duplicate exclusions during the close reading of the 163 candidates that yielded the 111-item review set; and the checkable-claim requirement as Gate 1 of the coding rubric, which reduced the 111 to the 66 analytic claims (Section 3.4).

3.4. Coding Scheme

Each claim was coded with a three-gate rubric, followed by a thematic tag and a confidence rating. The gates were applied in order.
  • Gate 1, codeability. The claim must concern Ebola, hantavirus, or measles, or a World Cup health risk, and must assert a statement that can be checked against evidence. Items that are pure opinion, jokes, or off-topic do not pass and are skipped.
  • Gate 2, accuracy (Verdict). The coder rates the claim as True (basically correct), Misleading (a grain of truth but leaving a wrong impression), or False (simply not true). There is no separate "unverified" category; uncertainty is captured by the confidence rating below.
  • Gate 3, direction (Direction). Judged against the real epidemiological risk, the claim either Inflates risk (a small threat made to sound huge), Downplays risk (a real threat made to sound harmless), or is Accurate (matches the real risk). The real-risk benchmarks are fixed in advance. Ebola is assessed as near-zero at the Cup, as it spreads through contact with body fluids and not is airborne [36]. Hantavirus is also very low risk; it is rodent-borne and rarely transmitted between people, with the exception of the Andes strain implicated in the May 2026 MV Hondius outbreak, a recognized human-to-human transmission event [37]. The outbreak was being contained but was still ongoing in late May 2026 [38], and hantavirus has no specific antiviral treatment and is managed only with supportive care [39]. Accordingly, promoted cures such as ivermectin therefore have no established benefit. Measles, by contrast, is high and rising; it is among the most contagious diseases, the MMR vaccine is safe and not a cause of autism [40], and vitamin A not a substitute [41,42]. For hantavirus specifically, a claim that accurately noted the low risk to the general tournament crowd was coded Accurate, while a claim that denied or minimized the reality or severity of the ongoing MV Hondius outbreak was coded Downplays risk. The fixed real-risk benchmark for each disease, together with its supporting sources, is set out in Supplementary Table S1, and the full coding rubric, including the Gate 1 to Gate 3 decision rules and the thematic and confidence codes, is provided in Supplementary Table S2.
  • Theme or type. Each claim is also tagged with one rhetorical theme using an author-developed coding scheme informed by prior work that categorizes health and disease misinformation by content type [33,34] and by topic [35], within the broader information-disorder framing of Wardle and Derakhshan [43]. The themes are exaggerated severity, blame a country or group, fake cure or prevention, travel panic, cover-up claim, made-up outbreak, or other. Where a claim fit more than one theme, it was assigned the single best-fitting theme; for example, a post promoting ivermectin against hantavirus was coded as fake cure or prevention rather than made-up outbreak, to reflect its primary move of promoting an ineffective treatment. These themes describe the rhetorical content of a claim, not the intent behind it; because intent cannot be established from a post alone, the study uses misinformation as an umbrella term and does not separate misinformation from deliberate disinformation.
  • Confidence. For every coded claim the coder records confidence (High, Medium, or Low) in the accuracy verdict given the available evidence, along with a short note on what the evidence says and its source.

3.5. Inter-Rater Reliability

Two coders coded independently: the first author and a second trained coder. To estimate reliability, a blind subset of 30 claims was drawn at random (random seed 42) from the hantavirus and Ebola pool and coded separately by both coders without consultation. Agreement on the two primary judgments was substantial by the Landis and Koch [44] benchmarks (see also [45]). For Verdict, Cohen's kappa [46] was 0.72 (observed agreement 0.83). For Direction, kappa was 0.68 (observed agreement 0.80). Coding direction for hantavirus required a demanding judgment, separating the low risk to the general tournament crowd from the real, ongoing MV Hondius outbreak (Section 3.4); this fine distinction is a likely reason Direction agreement was lower than Verdict agreement. Disagreements were then resolved by discussion and recorded in dedicated adjudication columns. The original blind codes were left unchanged so that the reliability statistics remain a faithful record of independent agreement. Theme was coded by the first author to describe recurring content. An initial blind double-coding showed that the theme categories could not be applied consistently between coders. For that reason, Theme is reported descriptively only and is not used for inference.

3.6. Analysis Plan

This is a descriptive content analysis and an exploratory case study. The study reports frequencies for Verdict, Direction, Theme, and disease, and cross-tabulates Verdict by Direction and disease by Direction. Any statistic we report is given as a transparency or effect-size check, not as a confirmatory hypothesis test. For the Verdict by Direction cross-tabulation, the chi-square statistic and Cramer's V are reported in the table note only, as an internal-consistency check, because the two judgments are linked by construction (see the Direction of Distortion results). Where an expected count falls below five, the same check is confirmed with the Fisher-Freeman-Halton exact test. The one substantive test is an exact binomial on the inflate versus downplay split among the inaccurate claims, which bounds the central claim and is reported in full whether or not it reaches significance. Because the analytic set is small (n = 66), no result is read as confirmatory. A volume-versus-real-risk comparison sets the amount of online attention each disease attracts against its real epidemiological severity. Reddit engagement (post score and comment count) is described as a secondary, exploratory measure and is summarized with medians because these counts are highly skewed. To place the conversation in time, we compared two daily series across the collection window: the daily volume of Reddit posts and comments matching each disease in the collection, and United States Google Trends search interest for Ebola, hantavirus, and measles over the same March 1 to June 3, 2026 window, reported on Google’s relative 0 to 100 index. We summarize their co-movement with the Pearson correlation and read it descriptively rather than as a causal test. Each disease was entered as a search term, the three were queried together so they share one 0 to 100 normalization, and the series is daily.

3.7. Ethics and Data Management

The study uses only public Reddit content; no private messages were accessed and no users were contacted. Usernames are not reported; illustrative quotations are kept brief to limit traceability. API credentials were stored locally and never shared; the raw data files were kept off version control. Because the study analyzes only publicly available Reddit content and involves no interaction or intervention with human participants, it did not require institutional review board approval.
Data availability. The data are public posts and comments from Reddit, collected through the Reddit API. To respect the platform's terms and protect user privacy, the authors do not redistribute the raw dataset. The complete search terms are listed in Appendix A. The post identifiers and permalinks needed to reconstruct the dataset are available from the corresponding author on reasonable request.

4. Results

The analytic set comprises the 66 claims that passed Gate 1 (n = 66). Of these, 41 were judged inaccurate (Misleading or False) and 25 were judged True (about 62 percent and 38 percent of the 66). These counts describe the curated analytic set and are not prevalence estimates for Reddit as a whole, because the set was assembled by screening for disease and conspiracy-style language before human coding. Reliability on the blind subset was substantial, with Cohen's kappa of 0.72 for Verdict and 0.68 for Direction (Section 3.5). Unless noted otherwise, the figures below describe these 66 coded claims. Any tests are read as descriptive, given the small sample.

4.1. Composition of the Analytic Set

Hantavirus claims dominated the analytic set. Of the 66 claims, 44 concerned hantavirus, 9 Ebola, and 8 measles; a further 4 were tagged as both Ebola and hantavirus and 1 as both Ebola and measles. Hantavirus alone therefore accounted for about two-thirds (67 percent) of the coded claims. This concentration reflects the May 2026 hantavirus 'Plandemic 2.0' wave that motivated the study and is set against real risk in the Online Volume Versus Real Risk results. The claims were rarely about the tournament itself: only 5 of the 66 were retrieved by a World Cup or FIFA search term. Only 3 mentioned the Cup directly. The World Cup serves here as the timing, the reason these three diseases were chosen, and the source of the real-risk benchmark (Section 3.4), not as the main subject of the discussion, which centered on the May 2026 hantavirus outbreak.

4.2. Direction of Distortion

The central finding is that distortion ran in both directions, and that the downplaying of real risk was at least as common as the inflation of it. Among the 41 inaccurate claims, 24 downplayed the real risk and 17 inflated it (about 59 percent and 41 percent of the 41; Figure 2). The point estimate leans toward downplaying, though the sample is small; an exact binomial test cannot distinguish this split from an even one (p = 0.35). We therefore read the result as showing that attenuation was substantial, the half of distortion that infodemic studies most often miss, rather than as ranking one direction above the other. This strong presence of attenuation, rather than a one-sided spread of alarm, is the study's main descriptive result. Table 2 shows the full Verdict by Direction cross-tabulation for all 66 claims. The two axes are closely linked by construction. Every claim Accurate in direction was also judged True (24 of 24); every inaccurate claim was non-Accurate. The only off-diagonal case was a single True claim that nonetheless overstated risk. Because that one claim was otherwise accurate, the Inflate direction in Table 2 totals 18, while the inaccurate inflations reported here number 17. Because the two judgments are linked in this way, a test on the full table can only show that the accuracy and direction codes are internally consistent, not that they are independently associated. We therefore report that internal-consistency check in the Table 2 note rather than in the text. The substantive result remains the inflate-versus-downplay split among the misinformation claims, in which downplaying was at least as common as inflating.

4.3. Themes of Misinformation

Theme was single-coded by the first author and is reported here as description only. A preliminary blind double-coding did not reach acceptable agreement between coders (see Inter-Rater Reliability), so the theme counts below are exploratory and are not used for inference. With that caveat, the 41 inaccurate claims spread across six of the seven rhetorical themes. The seventh, travel panic, did not appear at all. Blaming a country or group was by far the most common theme, with 15 of the 41 inaccurate claims (about 37 percent). Exaggerated severity followed with 8 (about 20 percent), then made-up outbreak with 6, fake cure or prevention with 5, and cover-up claim with 4. The remaining 3 claims fell under other. Table 3 reports the full distribution. Across all 66 claims the 'other' theme was the single largest category at 28, because most accurate, True claims were everyday factual statements that carried no misinformation rhetoric.

4.4. Theme by Direction

Reading theme and direction together is exploratory only. Theme was single-coded and did not reach acceptable double-coding agreement (see Inter-Rater Reliability), and the per-cell counts here are very small, so these patterns are not used for inference and should not be read as findings. With those limits, the counts were as follows. Blame claims were the single theme that fell on both sides of the split (7 downplayed risk and 8 inflated it). Across the other six themes combined, 17 of 26 inaccurate claims downplayed risk. We report these counts for completeness and do not interpret them further. Confirming any theme-by-direction pattern would require reliable theme coding. Separately, and independent of theme coding, the Direction code (which did reach substantial agreement; see Inter-Rater Reliability) shows what downplaying meant for hantavirus: denying the reality or severity of the actual, ongoing MV Hondius outbreak, with its confirmed deaths and human-to-human spread, rather than disputing the separate and genuinely low risk to the general tournament crowd.

4.5. Online Volume Versus Real Risk

A central motivation for the study was the gap between how much attention a disease draws online and how dangerous it actually is. Within this sample, that gap is wide. We describe this gap, we do not test it. The per-disease counts are small and unbalanced (44 hantavirus, 9 Ebola, 8 measles, plus 5 multi-disease claims), so we make no statistical comparison across diseases, and the numbers below are not prevalence estimates for Reddit as a whole. Hantavirus, whose real risk to the general tournament public was very low (rodent-borne and usually not passed between people, apart from the Andes human-to-human exception in the MV Hondius outbreak, which was being contained but still ongoing in late May 2026), accounted for the most claims by far (44 of 66) and the largest disease-specific block of conspiracy-flagged raw items (108 of the 679 items that tripped the misinformation-language filter, behind only the general cross-disease pool). Measles, the one disease with a high and rising real risk, drew the fewest claims (8); Ebola, whose real risk at the Cup was near zero, drew 9 (Figure 3). The blame theme was concentrated on hantavirus as well: 12 of the 15 blame claims targeted the hantavirus story. This pattern is consistent with online attention following the conspiracy narrative of the moment rather than the real epidemiological threat. The imbalance itself is part of what we observed: one low-risk disease pulled most of the attention. We read this descriptively, as a feature of this curated window, not as a measure of how much each disease was discussed on Reddit overall. It should, however, be read with care. The keyword list itself was roughly balanced across the three diseases (11 hantavirus terms, 10 Ebola, and 9 measles, plus 16 general cross-disease terms; Appendix A). The dominance of hantavirus is therefore not an artifact of a hantavirus-heavy search list. These volume figures still come from a curated, conspiracy-screened, keyword-driven sample and are not unbiased measures of how much each disease was discussed on Reddit.
The hantavirus conversation was also concentrated in time. Daily claim volume surged in early-to-mid May 2026, peaking on May 8, before tapering off over the rest of the collection window (Figure S1 in the Supplementary Materials; the final-day spike on June 3 there is a collection artifact, not a real resurgence).

4.6. Reddit Engagement (Exploratory)

Engagement was examined as a secondary, exploratory measure for all 66 claims, which were matched to their original Reddit metrics. Because post scores and comment counts are highly skewed and the per-group counts are small, medians are reported, group sizes are given, and no significance tests were run. By verdict, True claims (n = 25) had the highest median score at 25 upvotes, versus 7 for False (n = 28) and 2 for Misleading (n = 13). By direction, claims Accurate in direction (n = 24) had a median score of about 25, versus about 7 for downplaying (n = 24) and about 4 for inflating claims (n = 18). Comment counts were similar across categories (medians of roughly 23 to 28 regardless of verdict or direction). These differences are descriptive only and were not tested. At most, they hint that accurate content drew more upvotes while inaccurate content still attracted comparable discussion. This upvote pattern runs counter to evidence that falsehood outspreads truth on platforms such as Twitter [3], a contrast that likely reflects Reddit's community voting and moderation and the fact that upvotes index approval rather than reach. The Discussion takes this up.

4.7. Timing of the Conversation

The volume of disease-risk discussion in this window moved with real disease events rather than with the tournament schedule. Reddit disease-claim volume and United States Google Trends search interest for the three diseases were both low through March and April, then rose sharply in early May, when a hantavirus outbreak aboard the cruise ship MV Hondius was reported to the World Health Organization; the cluster was reported on May 2, a public bulletin followed on about May 8, and the count was updated to 10 cases and 3 deaths on May 15 [37,38]. Reddit hantavirus volume and hantavirus search interest tracked each other closely across the window, with a Pearson correlation of 0.93, while Ebola and measles search interest did not rise, indicating that the surge was specific to the hantavirus outbreak rather than a general disease-news effect (Figure 4). A second, distinct rise, this one in Ebola discussion, followed in late May, when the United States required the Democratic Republic of the Congo national football team to isolate before entering the country for the tournament because of a Bundibugyo Ebola outbreak [49]. This was the one clearly tournament-linked thread in the window. Both surges occurred before the tournament opened on June 11 and before collection ended on June 3. We therefore read the temporal pattern as following real disease events and their announcements and coverage, and we treat the World Cup as the dated window that placed these diseases before the same audience rather than as the cause of the discussion.

5. Discussion

5.1. Principal Findings

This study asked whether online narratives distort disease risk in one direction only, toward alarm, or in both. The answer is both. Among the 41 inaccurate claims, 24 downplayed the real risk and 17 inflated it. The quiet downplaying of real threats was at least as common as loud false alarm, though the sample is too small to say which is larger (exact binomial p = 0.35). Online attention also failed to track real danger within this sample. Hantavirus, whose risk to the general public at the Cup is low, drew about two thirds of all claims (44 of 66), while measles, the one disease with a high and rising real risk, drew the fewest (8). We report this as a within-sample, descriptive pattern, not as a tested comparison across diseases, because the per-disease counts are small and unbalanced. In short, the infodemic did not mirror the epidemic: the most-discussed topic was not the most dangerous one; the distortion cut both ways. The timing of the conversation supports reading the World Cup as a dated window rather than a driver: attention to all three diseases tracked their real events, the early-May hantavirus outbreak and the late-May controversy over the Democratic Republic of the Congo Ebola outbreak, and both surges preceded the opening match (Section 4.7).

5.2. Both Directions Through SARF

These patterns fit the Social Amplification of Risk Framework [8], which holds that social processes can both amplify and attenuate a risk signal. Most empirical work has measured amplification, the spread of alarm. Far less has measured attenuation, the quiet downplaying of a real threat [15]. The hantavirus case here is a clear example of attenuation. A real and ongoing outbreak, with confirmed deaths and person-to-person spread, was recast as scripted or fake and waved away. By coding each claim for direction against a fixed real-risk benchmark, this study gives the attenuation half of the framework the same direct attention usually reserved for amplification.

5.3. Comparison with Prior Work

The closest prior study, Hopfer et al. [16], coded both amplification and attenuation of COVID-19 risk on Twitter; this study extends that idea to three diseases on Reddit and adds an external real-risk benchmark. One result runs against a well-known finding. Vosoughi et al. [3] showed that false news spreads farther and faster than truth on Twitter. Yet in this sample true claims drew the highest median upvotes (25, versus 7 for false and 2 for misleading). The two findings do not conflict. Upvotes measure community approval, not reach; Reddit's voting and moderation differ from Twitter's; and the present set is a curated reading sample rather than a measure of spread. As a tentative observation, the apparent dominance of blame claims (15 of the 41, and 12 of those about hantavirus) would, if it holds, echo earlier outbreak work in which fear mixes with conspiracy and scapegoating [19]. We note this with caution. The theme codes did not reach acceptable double-coding agreement (see Inter-Rater Reliability), so this is an exploratory pattern, not a finding, and it would need confirmation with reliable theme coding before any weight is placed on it. It does not affect the study's main result, which rests on the Direction code.

5.4. Implications for Outbreak Risk Communication

For outbreak risk communication, whether or not a mass gathering is involved, one practical implication is that monitoring should not look for panic alone. A real outbreak can be denied as easily as a minor one can be exaggerated. During a mass gathering, denial of a genuine, ongoing threat may be the more dangerous failure, because it can lower the guard of the very people most exposed (compare [20], where outbreak disbelief reduced prevention). Quiet, high-real-risk diseases such as measles deserve attention even when they are not trending. Approaches that build resistance to misleading claims before they take hold, such as inoculation or prebunking [7], may help on both sides, against inflation and against downplaying. The most direct audience is public-health risk communicators, for whom the pattern here suggests watching for the denial of a real, ongoing threat as closely as for panic, and keeping quiet but high-risk diseases such as measles in view even when they are not trending. Event planners and platform moderators may find the same two-sided view useful, treating the downplaying of a documented outbreak as a signal to act on rather than a harmless counterweight to alarm. We do not extend these suggestions further, given the small, single-platform, screened basis of the study. These are modest, exploratory suggestions drawn from a small single-platform sample, and they would need testing in other settings before being built into routine practice.

5.5. Limitations

Several limits should be kept in mind. The analytic set is small (n = 66) and comes from a single platform. It was built by screening for disease and conspiracy-style language; the figures therefore describe a curated reading set and are not prevalence estimates for Reddit. The keyword list was roughly balanced across the three diseases (Appendix A); the volume comparison is therefore not driven by a hantavirus-heavy search list. The list did, however, lean toward denial and hoax language, such as plandemic, hoax, and cover up, which would preferentially surface claims that downplay a real threat, so the relative balance of the two directions should be read with that slant in mind. What this corpus establishes is not that downplaying is more common than inflation in general, but that the downplaying of a real risk is present, substantial, and measurable against the fixed benchmark. Even so, it remains a conspiracy-screened, keyword-driven sample rather than an unbiased measure of how much each disease was discussed. The per-disease counts are also small and unbalanced (44 hantavirus, 9 Ebola, 8 measles, plus 5 multi-disease claims), so we make no statistical comparison across diseases. The volume-versus-risk pattern is described within the sample and should not be read as a population estimate. The bidirectional direction finding does not depend on this balance, because it rests on per-claim direction coding rather than on claim volume. Reliability was estimated on a blind subset of 30 claims drawn from the hantavirus and Ebola pool. It speaks to those two diseases and not to measles. Theme was single-coded for analysis. A preliminary blind double-coding did not reach acceptable agreement between coders, so the theme counts are reported as exploratory description only and are not used for inference (see the Themes of Misinformation and Theme by Direction results). The study's main result, the bidirectional inflate-versus-downplay split, rests on the Direction code, which reached substantial agreement, and does not depend on Theme. The chi-square on the Verdict by Direction table is reported only as an internal-consistency check, because the two judgments are linked by construction. Nearly half the coded claims (29 of 66) came from a single conspiracy-focused community, r/conspiracy (Figure 5). The bidirectional pattern may therefore partly reflect that community rather than Reddit as a whole. To check whether this single community drove the result, we compared the two groups directly. With r/conspiracy set aside, 10 of the 14 inaccurate claims from other communities downplayed the real risk and 4 inflated it, so the downplaying pattern was, if anything, stronger outside r/conspiracy. Within r/conspiracy the split was more even, 14 downplaying and 13 inflating. The bidirectional finding therefore does not rest on the r/conspiracy claims alone. That said, the claim families seen here also circulated well beyond Reddit during the study window. Independent fact-checks traced the hantavirus conspiracy narratives, from 'Covid 2.0' to crisis-actor framings, across several platforms [27,28]. A national survey found more Americans encountering false claims about the measles vaccine [47]. Together these suggest the broad pattern is not confined to one community, even though this outside reporting cannot remove the concentration limitation noted here. Reddit engagement is exploratory and was described, not tested. The temporal comparison in Section 4.7 uses aggregate Google Trends search interest, which is a relative, population-level proxy rather than a measure of individual exposure, so it shows that attention and real events moved together in time but cannot establish that the events caused particular claims. The search series is United States only while Reddit is global, and the correlation reflects a single coincident surge on an otherwise low, short daily series, so it should be read as agreement in timing rather than as a strong association.
This study also did not place claims on a geographic map. A method we developed in earlier work geolocates Reddit users by matching them to location-specific city and county subreddits [48], which raised the natural question of whether claims could be mapped to the tournament's host cities. That was not feasible here. Only 2 of the 66 coded claims appeared in a city or state subreddit. About 2 percent of the full collection fell in any location-specific community, a finding that is itself informative. COVID-19 was discussed in local subreddits across the country, which made users geographically findable. The disease-misinformation talk studied here, by contrast, lived in national, topic-based communities, with 29 of the 66 coded claims concentrated in a single conspiracy-focused community (r/conspiracy; Figure 5). A city-level map would therefore rest on too few points to be meaningful or honest. This infodemic was organized around community and topic rather than place, which is consistent with the amplification stations of SARF and points to community-level mapping, rather than geographic mapping, as the more natural lens for these data.

5.6. Future Work

Future studies could widen the lens to other platforms and to a longer window that includes the tournament itself, increase the number of coded claims, and test directly whether downplaying and inflating travel differently. A larger sample would also let us examine the theme and direction patterns seen here with proper statistical power, rather than reading them descriptively. Future work could also test how useful this two-sided monitoring is for the people who would put it to use. One step is to work with public-health risk communicators and agencies to see whether watching for denial as well as panic changes what they catch early. Another is to study other mass gatherings, such as concerts, festivals, and large sporting events, to learn whether the upside-down pattern seen here, where a low-risk disease drew the most attention, holds beyond the World Cup. A further step is to test whether tools that event and venue organizers and platform moderators already use could flag downplaying of a real threat, not just alarm. These are directions a larger study could pursue; the present sample is too small and too narrow to settle them.

6. Conclusions

This study used the run-up to the 2026 FIFA World Cup as a dated window and a real-risk benchmark, not as its subject, and asked how disease risk was distorted online within that window and whether that distortion ran in one direction or two. The conversation itself centered on the May 2026 hantavirus outbreak. Across 66 coded Reddit claims about Ebola, hantavirus, and measles, the answer was clear: risk was distorted in both directions, with real threats downplayed at least as often as minor ones were inflated, though the sample is too small to say which is larger. The disease that filled most of the conversation, hantavirus, was not the one that posed the most real danger, while measles, the genuinely rising threat, drew the least attention. Within this sample, the online picture ran counter to the epidemiological one.
By coding each claim against a fixed real-risk benchmark and by giving the downplaying of risk the same attention as its inflation, the study brings the often-overlooked attenuation side of the Social Amplification of Risk Framework into clear view. For risk communication around mass gatherings, the message is that denial of a real outbreak deserves as much attention as false alarm. Quiet, high-risk diseases also deserve attention, even when they are not trending. These findings come from a small, single-platform sample and are descriptive. Even so, they offer a clear and testable picture of an infodemic that does not simply mirror the epidemic.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org. Table S1: the fixed real-risk benchmark for Ebola, hantavirus, and measles, with the sources used to set each level. Table S2: the full coding rubric, including the Gate 1 to Gate 3 decision rules, the thematic tag categories, and the confidence rating. Figure S1: the daily volume of Reddit disease-risk claims across the March 1 to June 3, 2026 collection window, described in Section 4.5.

Author Contributions

Conceptualization, L.A. and M.Z.E.; methodology, L.A.; software, L.A.; validation, L.A. and N.R.; formal analysis, L.A.; investigation, L.A. and N.R.; data curation, L.A.; writing-original draft preparation, L.A.; writing-review and editing, L.A., N.R. and M.Z.E.; visualization, L.A.; supervision, M.Z.E.; project administration, M.Z.E. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethical review and approval were waived for this study because it analyzed only publicly available Reddit content and involved no interaction or intervention with human participants.

Data Availability Statement

The data analyzed in this study are public posts and comments collected through the Reddit API. To respect the platform's terms of service and protect user privacy, the authors do not redistribute the raw dataset. The complete search terms are provided in Appendix A, and the post identifiers and permalinks needed to reconstruct the dataset are available from the corresponding author on reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Search Terms

The collection script queried the following 46 search terms (Section 3.2). The conspiracy and health-misinformation vocabulary was informed by prior research and contemporary online conspiracy discourse [5,6,33,34,35]; the disease, venue, and World Cup terms specific to the 2026 tournament were developed by the authors. Terms are listed verbatim; the quotation marks are not part of the queries.
Ebola (10 terms): World Cup Ebola; FIFA Ebola; Ebola 2026 outbreak; Ebola bioweapon; Ebola hoax; Ebola conspiracy; Ebola depopulation; Marburg outbreak; Ebola migrants; Ebola lab leak.
Hantavirus (11 terms): World Cup hantavirus; FIFA hantavirus; hantavirus plandemic; hantavirus scripted pandemic; hantavirus hoax; hantavirus Covid 2.0; MV Hondius hantavirus; hantavirus ivermectin; hantavirus bioweapon; hantavirus depopulation; hantavirus cover up.
Measles (9 terms): measles World Cup; measles outbreak 2026 hoax; measles vaccine autism; vitamin A measles; MMR aborted fetal; measles vaccine dangerous; RFK measles; measles natural immunity; Texas measles hoax.
General mass gathering and cross-disease (16 terms): World Cup outbreak conspiracy; World Cup disease plandemic; FIFA 2026 virus; World Cup depopulation; World Cup superspreader; World Cup vaccine mandate; World Cup quarantine; World Cup biosecurity; MetLife stadium virus; World Cup pandemic 2026; released on purpose; great reset pandemic; population control virus; manufactured outbreak; next plandemic; fake outbreak 2026.

References

  1. Eysenbach, G. How to Fight an Infodemic: The Four Pillars of Infodemic Management. J. Med. Internet Res. 2020, 22, e21820. [CrossRef]
  2. Tangcharoensathien, V.; Calleja, N.; Nguyen, T.; Purnat, T.; D'Agostino, M.; Garcia-Saiso, S.; Landry, M.; Rashidian, A.; Hamilton, C.; AbdAllah, A.; Ghiga, I.; Hill, A.; Hougendobler, D.; van Andel, J.; Nunn, M.; Brooks, I.; Sacco, P.L.; De Domenico, M.; Mai, P.; Gruzd, A.; Alaphilippe, A.; Briand, S. Framework for Managing the COVID-19 Infodemic: Methods and Results of an Online, Crowdsourced WHO Technical Consultation. J. Med. Internet Res. 2020, 22, e19659. [CrossRef]
  3. Vosoughi, S.; Roy, D.; Aral, S. The spread of true and false news online. Science 2018, 359, 1146-1151. [CrossRef]
  4. Lazer, D.M.J.; Baum, M.A.; Benkler, Y.; Berinsky, A.J.; Greenhill, K.M.; Menczer, F.; Metzger, M.J.; Nyhan, B.; Pennycook, G.; Rothschild, D.; Schudson, M.; Sloman, S.A.; Sunstein, C.R.; Thorson, E.A.; Watts, D.J.; Zittrain, J.L. The science of fake news. Science 2018, 359, 1094-1096. [CrossRef]
  5. Suarez-Lledo, V.; Alvarez-Galvez, J. Prevalence of Health Misinformation on Social Media: Systematic Review. J. Med. Internet Res. 2021, 23, e17187. [CrossRef]
  6. Enders, A.M.; Uscinski, J.E.; Klofstad, C.; Stoler, J. The different forms of COVID-19 misinformation and their consequences. Harvard Kennedy School Misinformation Review 2020. [CrossRef]
  7. van der Linden, S. Misinformation: susceptibility, spread, and interventions to immunize the public. Nat. Med. 2022, 28, 460-467. [CrossRef]
  8. Kasperson, R.E.; Renn, O.; Slovic, P.; Brown, H.S.; Emel, J.; Goble, R.; Kasperson, J.X.; Ratick, S. The Social Amplification of Risk: A Conceptual Framework. Risk Anal. 1988, 8, 177-187. [CrossRef]
  9. Renn, O.; Burns, W.J.; Kasperson, J.X.; Kasperson, R.E.; Slovic, P. The Social Amplification of Risk: Theoretical Foundations and Empirical Applications. J. Soc. Issues 1992, 48, 137-160. [CrossRef]
  10. Pidgeon, N.; Kasperson, R.E.; Slovic, P. (Eds.) The Social Amplification of Risk; Cambridge University Press: Cambridge, UK, 2003. [CrossRef]
  11. Kasperson, R.E.; Webler, T.; Ram, B.; Sutton, J. The social amplification of risk framework: New perspectives. Risk Anal. 2022, 42, 1367-1380. [CrossRef]
  12. Moussaid, M.; Brighton, H.; Gaissmaier, W. The amplification of risk in experimental diffusion chains. Proc. Natl. Acad. Sci. USA 2015, 112, 5631-5636. [CrossRef]
  13. Gallotti, R.; Valle, F.; Castaldo, N.; Sacco, P.; De Domenico, M. Assessing the risks of 'infodemics' in response to COVID-19 epidemics. Nat. Hum. Behav. 2020, 4, 1285-1293. [CrossRef]
  14. Gassen, J.; Nowak, T.J.; Henderson, A.D.; Weaver, S.P.; Baker, E.J.; Muehlenbein, M.P. Unrealistic Optimism and Risk for COVID-19 Disease. Front. Psychol. 2021, 12, 647461. [CrossRef]
  15. Fjaeran, L.; Gould, K.P.; Goble, R. Before amplification: the role of experts in the dynamics of the social attenuation and amplification of risk. J. Risk Res. 2024, 27, 219-237. [CrossRef]
  16. Hopfer, S.; Fields, E.J.; Lu, Y.; Ramakrishnan, G.; Grover, T.; Bai, Q.; Huang, Y.; Li, C.; Mark, G. The social amplification and attenuation of COVID-19 risk perception shaping mask wearing behavior: A longitudinal twitter analysis. PLoS ONE 2021, 16, e0257428. [CrossRef]
  17. Oyeyemi, S.O.; Gabarron, E.; Wynn, R. Ebola, Twitter, and misinformation: a dangerous combination? BMJ 2014, 349, g6178. [CrossRef]
  18. Fung, I.C.-H.; Fu, K.-W.; Chan, C.-H.; Chan, B.S.B.; Cheung, C.-N.; Abraham, T.; Tse, Z.T.H. Social Media's Initial Reaction to Information and Misinformation on Ebola, August 2014: Facts and Rumors. Public Health Rep. 2016, 131, 461-473. [CrossRef]
  19. Sell, T.K.; Hosangadi, D.; Trotochaud, M. Misinformation and the US Ebola communication crisis: analyzing the veracity and content of social media messages related to a fear-inducing infectious disease outbreak. BMC Public Health 2020, 20, 550. [CrossRef]
  20. Vinck, P.; Pham, P.N.; Bindu, K.K.; Bedford, J.; Nilles, E.J. Institutional trust and misinformation in the response to the 2018-19 Ebola outbreak in North Kivu, DR Congo: a population-based survey. Lancet Infect. Dis. 2019, 19, 529-536. [CrossRef]
  21. Wakefield, A.; Murch, S.; Anthony, A.; Linnell, J.; Casson, D.; Malik, M.; Berelowitz, M.; Dhillon, A.; Thomson, M.; Harvey, P.; Valentine, A.; Davies, S.; Walker-Smith, J. RETRACTED: Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. Lancet 1998, 351, 637-641. [CrossRef]
  22. The Editors of The Lancet. Retraction: Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. Lancet 2010, 375, 445. [CrossRef]
  23. Phadke, V.K.; Bednarczyk, R.A.; Salmon, D.A.; Omer, S.B. Association Between Vaccine Refusal and Vaccine-Preventable Diseases in the United States. JAMA 2016, 315, 1149. [CrossRef]
  24. Majumder, M.S.; Cohn, E.L.; Mekaru, S.R.; Huston, J.E.; Brownstein, J.S. Substandard Vaccination Compliance and the 2015 Measles Outbreak. JAMA Pediatr. 2015, 169, 494. [CrossRef]
  25. Wilson, S.L.; Wiysonge, C. Social media and vaccine hesitancy. BMJ Glob. Health 2020, 5, e004206. [CrossRef]
  26. de Figueiredo, A.; Simas, C.; Karafillakis, E.; Paterson, P.; Larson, H.J. Mapping global trends in vaccine confidence and investigating barriers to vaccine uptake: a large-scale retrospective temporal modelling study. Lancet 2020, 396, 898-908. [CrossRef]
  27. Nilsson-Julien, E.; Paternoster, T. Staged claims and Israeli hoaxes: Debunking viral conspiracy theories about hantavirus. Euronews, 13 May 2026. Available online: https://www.euronews.com/my-europe/2026/05/13/staged-claims-and-israeli-hoaxes-debunking-viral-conspiracy-theories-about-hantavirus (accessed on 7 June 2026).
  28. Burke, S. From 'Covid 2.0' to 'crisis actors': How hantavirus myths spread. RTE, 20 May 2026. Available online: https://www.rte.ie/news/primetime/2026/0520/1574100-hantavirus-myths-spread/ (accessed on 7 June 2026).
  29. Harris, C.; Armien, B. Sociocultural determinants of adoption of preventive practices for hantavirus: A knowledge, attitudes, and practices survey in Tonosi, Panama. PLoS Negl. Trop. Dis. 2020, 14, e0008111. [CrossRef]
  30. Memish, Z.A.; Steffen, R.; White, P.; Dar, O.; Azhar, E.I.; Sharma, A.; Zumla, A. Mass gatherings medicine: public health issues arising from mass gathering religious and sporting events. Lancet 2019, 393, 2073-2084. [CrossRef]
  31. Al-Tawfiq, J.A.; Gautret, P.; Schlagenhauf, P. Infection risks associated with the 2022 FIFA World Cup in Qatar. New Microbes New Infect. 2022, 49-50, 101055. [CrossRef]
  32. Alhussaini, N.W.Z.; Elshaikh, U.A.M.; Hamad, N.A.; Nazzal, M.A.; Abuzayed, M.; Al-Jayyousi, G.F. A scoping review of the risk factors and strategies followed for the prevention of COVID-19 and other infectious diseases during sports mass gatherings: Recommendations for future FIFA World Cups. Front. Public Health 2023, 10, 1078834. [CrossRef]
  33. Brennen, J.S.; Simon, F.M.; Howard, P.N.; Nielsen, R.K. Types, Sources, and Claims of COVID-19 Misinformation; Reuters Institute for the Study of Journalism: Oxford, UK, 2020. Available online: https://reutersinstitute.politics.ox.ac.uk/types-sources-and-claims-covid-19-misinformation (accessed on 7 June 2026).
  34. Islam, M.S.; Sarkar, T.; Khan, S.H.; Mostofa Kamal, A.-H.; Hasan, S.M.M.; Kabir, A.; Yeasmin, D.; Islam, M.A.; Amin Chowdhury, K.I.; Anwar, K.S.; Chughtai, A.A.; Seale, H. COVID-19-Related Infodemic and Its Impact on Public Health: A Global Social Media Analysis. Am. J. Trop. Med. Hyg. 2020, 103, 1621-1629. [CrossRef]
  35. Wang, Y.; McKee, M.; Torbica, A.; Stuckler, D. Systematic Literature Review on the Spread of Health-related Misinformation on Social Media. Soc. Sci. Med. 2019, 240, 112552. [CrossRef]
  36. World Health Organization. Ebola Disease. Available online: https://www.who.int/news-room/fact-sheets/detail/ebola-virus-disease (accessed on 7 June 2026).
  37. World Health Organization. Hantavirus Cluster Linked to Cruise Ship Travel, Multi-country; Disease Outbreak News. Available online: https://www.who.int/emergencies/disease-outbreak-news/item/2026-DON600 (accessed on 7 June 2026).
  38. European Centre for Disease Prevention and Control. Andes Hantavirus Outbreak in Cruise Ship. Available online: https://www.ecdc.europa.eu/en/infectious-disease-topics/hantavirus-infection/surveillance-and-updates/andes-hantavirus-outbreak (accessed on 7 June 2026).
  39. Centers for Disease Control and Prevention. About Hantavirus. Available online: https://www.cdc.gov/hantavirus/about/index.html (accessed on 7 June 2026).
  40. Hviid, A.; Hansen, J.V.; Frisch, M.; Melbye, M. Measles, mumps, rubella vaccination and autism: A nationwide cohort study. Ann. Intern. Med. 2019, 170, 513-520. [CrossRef]
  41. World Health Organization. Measles. Available online: https://www.who.int/news-room/fact-sheets/detail/measles (accessed on 7 June 2026).
  42. Centers for Disease Control and Prevention. About Measles. Available online: https://www.cdc.gov/measles/about/index.html (accessed on 7 June 2026).
  43. Wardle, C.; Derakhshan, H. Information Disorder: Toward an Interdisciplinary Framework for Research and Policy Making; Council of Europe: Strasbourg, France, 2017. Available online: https://rm.coe.int/information-disorder-report-version-august-2018/16808c9c77 (accessed on 7 June 2026).
  44. Landis, J.R.; Koch, G.G. The Measurement of Observer Agreement for Categorical Data. Biometrics 1977, 33, 159-174. [CrossRef]
  45. McHugh, M.L. Interrater reliability: the kappa statistic. Biochem. Med. 2012, 22, 276-282. [CrossRef]
  46. Cohen, J. A coefficient of agreement for nominal scales. Educ. Psychol. Meas. 1960, 20, 37-46. [CrossRef]
  47. KFF. Amid Growing Measles Outbreak, More Americans Are Encountering False Claims About the Measles Vaccine, and Many Aren't Sure What to Believe. Available online: https://www.kff.org/health-information-trust/amid-growing-measles-outbreak-more-americans-are-encountering-false-claims-about-the-measles-vaccine-and-many-arent-sure-what-to-believe/ (accessed on 7 June 2026).
  48. Alarfaj, L.; Blackburn, J.; Amjad, M.; Patel, J.; Ertem, Z. Mapping the Infodemic: Geolocating Reddit Users and Unsupervised Topic Modeling of COVID-19-Related Misinformation. Information 2025, 16, 748. [CrossRef]
  49. Al Jazeera. US Says Congo Team Must Isolate Due to Ebola Before Arriving for World Cup. Available online: https://www.aljazeera.com/sports/2026/5/22/us-says-congo-team-must-isolate-due-to-ebola-before-arriving-for-world-cup (accessed on 3 June 2026).
Figure 1. Study-flow diagram for the Reddit content analysis, adapted from the PRISMA 2020 four-phase model. Records identified through the keyword search (n = 8,264) were screened to a topical candidate pool (n = 163), read in full for coding review (n = 111), and reduced by Gate 1 of the coding rubric to the final analytic set of coded claims (n = 66). The counts reconcile as 8,264 minus 8,101 equals 163, then 163 minus 52 equals 111, then 111 minus 45 equals 66. Exclusions at each stage, with reasons, are shown to the right. No record was re-coded at any stage.
Figure 1. Study-flow diagram for the Reddit content analysis, adapted from the PRISMA 2020 four-phase model. Records identified through the keyword search (n = 8,264) were screened to a topical candidate pool (n = 163), read in full for coding review (n = 111), and reduced by Gate 1 of the coding rubric to the final analytic set of coded claims (n = 66). The counts reconcile as 8,264 minus 8,101 equals 163, then 163 minus 52 equals 111, then 111 minus 45 equals 66. Exclusions at each stage, with reasons, are shown to the right. No record was re-coded at any stage.
Preprints 229610 g001
Figure 2. Risk distortion ran in both directions. Among the 41 claims judged inaccurate (of 66 coded), 24 downplayed the real risk (attenuation) and 17 inflated it (amplification), each measured against a fixed real-risk benchmark (Ebola near-zero, hantavirus very low, measles high). An exact binomial test cannot distinguish this split from an even one (p = 0.35). Within this corpus, attenuation was substantial rather than a minor share, and the two directions cannot be ranked. The remaining 25 claims judged true are not shown (24 accurate in direction; one overstated risk). Because this finding rests on per-claim direction coding rather than claim volume, it does not depend on how many claims each disease drew, though it describes this screened set rather than Reddit as a whole.
Figure 2. Risk distortion ran in both directions. Among the 41 claims judged inaccurate (of 66 coded), 24 downplayed the real risk (attenuation) and 17 inflated it (amplification), each measured against a fixed real-risk benchmark (Ebola near-zero, hantavirus very low, measles high). An exact binomial test cannot distinguish this split from an even one (p = 0.35). Within this corpus, attenuation was substantial rather than a minor share, and the two directions cannot be ranked. The remaining 25 claims judged true are not shown (24 accurate in direction; one overstated risk). Because this finding rests on per-claim direction coding rather than claim volume, it does not depend on how many claims each disease drew, though it describes this screened set rather than Reddit as a whole.
Preprints 229610 g002
Figure 3. Online attention ran counter to real risk in this sample. The number of coded claims per disease (online attention) set against each disease's fixed real-risk benchmark. Hantavirus, a low real risk to the general tournament public, drew the most claims (44); measles, the one high and rising real risk, drew the fewest (8); Ebola, near-zero real risk, drew 9 (5 claims tagged with two diseases are omitted from the figure). This figure describes the curated reading set, not Reddit as a whole: the keyword list was roughly balanced across the three diseases (Appendix A); the hantavirus lead is therefore not a search artifact. Even so, these counts come from a conspiracy-screened, keyword-driven sample and are illustrative rather than unbiased measures of discussion volume. No statistical comparison is made across diseases. Given the small and unbalanced per-disease counts, the figure sets the within-sample claim counts against the fixed real-risk rank by eye. Percentages in the figure are shares of the 61 single-disease claims. The real-risk side is an ordinal rank taken from the fixed Section 3.4 benchmark, not a measured quantity.
Figure 3. Online attention ran counter to real risk in this sample. The number of coded claims per disease (online attention) set against each disease's fixed real-risk benchmark. Hantavirus, a low real risk to the general tournament public, drew the most claims (44); measles, the one high and rising real risk, drew the fewest (8); Ebola, near-zero real risk, drew 9 (5 claims tagged with two diseases are omitted from the figure). This figure describes the curated reading set, not Reddit as a whole: the keyword list was roughly balanced across the three diseases (Appendix A); the hantavirus lead is therefore not a search artifact. Even so, these counts come from a conspiracy-screened, keyword-driven sample and are illustrative rather than unbiased measures of discussion volume. No statistical comparison is made across diseases. Given the small and unbalanced per-disease counts, the figure sets the within-sample claim counts against the fixed real-risk rank by eye. Percentages in the figure are shares of the 61 single-disease claims. The real-risk side is an ordinal rank taken from the fixed Section 3.4 benchmark, not a measured quantity.
Preprints 229610 g003
Figure 4. Attention to the three diseases tracked real events, not the tournament schedule. Relative attention over the collection window, with each Reddit series scaled to its own maximum and each Google Trends series shown on Google’s 0 to 100 index; solid lines are Reddit claim volume and dashed lines are Google search interest. The first shaded band marks the MV Hondius hantavirus outbreak, with World Health Organization notifications from May 2 to 15; the second marks the World Cup entry controversy over the Democratic Republic of the Congo Ebola outbreak, from about May 22. Hantavirus surged with the outbreak, with Reddit and search interest tracking closely at a Pearson correlation of 0.93; Ebola rose later with the tournament controversy; measles stayed low throughout. Both surges preceded the tournament opening on June 11. Because each Reddit series is scaled to its own maximum, the figure shows the timing of attention, not its size across diseases.
Figure 4. Attention to the three diseases tracked real events, not the tournament schedule. Relative attention over the collection window, with each Reddit series scaled to its own maximum and each Google Trends series shown on Google’s 0 to 100 index; solid lines are Reddit claim volume and dashed lines are Google search interest. The first shaded band marks the MV Hondius hantavirus outbreak, with World Health Organization notifications from May 2 to 15; the second marks the World Cup entry controversy over the Democratic Republic of the Congo Ebola outbreak, from about May 22. Hantavirus surged with the outbreak, with Reddit and search interest tracking closely at a Pearson correlation of 0.93; Ebola rose later with the tournament controversy; measles stayed low throughout. Both surges preceded the tournament opening on June 11. Because each Reddit series is scaled to its own maximum, the figure shows the timing of attention, not its size across diseases.
Preprints 229610 g004
Figure 5. Coded claims concentrated in a few topic-based communities. The number of coded claims by subreddit. A single conspiracy-focused community, r/conspiracy, held 29 of the 66 claims, while the rest were spread thinly across many communities. This concentration is consistent with the amplification stations of SARF and supports community-level rather than geographic mapping for these data (Section 5.5). Like the volume figure, it describes the curated, keyword-driven reading set.
Figure 5. Coded claims concentrated in a few topic-based communities. The number of coded claims by subreddit. A single conspiracy-focused community, r/conspiracy, held 29 of the 66 claims, while the rest were spread thinly across many communities. This concentration is consistent with the amplification stations of SARF and supports community-level rather than geographic mapping for these data (Section 5.5). Like the volume figure, it describes the curated, keyword-driven reading set.
Preprints 229610 g005
Table 1. Comparison of the closest prior studies on disease and health misinformation. The table sets each study against the five design features of the present work (multiple diseases, Reddit data, both directions of distortion, a real epidemiological-risk benchmark, and a mass-gathering context) and shows which features each study covers and where the gap lies.
Table 1. Comparison of the closest prior studies on disease and health misinformation. The table sets each study against the five design features of the present work (multiple diseases, Reddit data, both directions of distortion, a real epidemiological-risk benchmark, and a mass-gathering context) and shows which features each study covers and where the gap lies.
Features / Gaps Addressed Technique / Approach Study
False news spreads farther and faster than truth; focuses on spread (amplification) Large-scale analysis of Twitter rumor cascades Vosoughi et al. [3]
Prevalence and types of health misinformation across platforms Systematic review (69 studies) Suarez-Lledo & Alvarez-Galvez [5]
Correlates of COVID-19 conspiracy belief and links to behavior Survey of US adults Enders et al. [6]
Susceptibility, spread, and inoculation (prebunking); largely one-directional Narrative review van der Linden [7]
Tracks circulation of unreliable information during COVID-19 Infodemic Risk Index (Twitter) Gallotti et al. [13]
Defines bidirectional amplification and attenuation of risk Conceptual framework (SARF) Kasperson et al. [8]
Hazard events can heighten or attenuate risk perception Empirical SARF study Renn et al. [9]
Risk information amplified and distorted person to person Diffusion-chain experiment Moussaid et al. [12]
Unrealistic optimism leads to underestimating personal COVID-19 risk (attenuation) Survey Gassen et al. [14]
Revisits the attenuation side of risk Conceptual review (SARF) Fjaeran et al. [15]
Codes both amplification and attenuation; single disease, single platform, no real-risk benchmark Content coding of Twitter Hopfer et al. [16]
Early Ebola false treatment claims (for example, saltwater) Content analysis Fung et al. [18]
About 10% false; conspiracy, political, and discord-driven content in the US Ebola scare Content analysis (about 72,000 tweets) Sell et al. [19]
Outbreak disbelief and low trust reduce prevention (real-world attenuation) Population survey (961 adults, DRC) Vinck et al. [20]
Vaccine refusal associated with increased measles risk Review of case data Phadke et al. [23]
Organized anti-vaccine activity and foreign disinformation linked to lower coverage Cross-national analysis Wilson & Wiysonge [25]
Global trends in vaccine confidence Large-scale mapping (149 countries) de Figueiredo et al. [26]
Hantavirus knowledge gaps in an endemic area (real-world, not online) Knowledge-attitudes-practices survey (Panama) Harris & Armien [29]
Public-health risks at mass-gathering events Review Memish et al. [30]
Infection-prevention strategies for the FIFA World Cup Scoping review (34 studies) Alhussaini et al. [32]
Note. Studies are listed by author and year. Each row shows the design features one study covers; none of the prior studies covers all five features of the present design.
Table 2. Verdict by Direction for the 66 coded claims.
Table 2. Verdict by Direction for the 66 coded claims.
Total Accurate Inflates Downplays Verdict
28 0 10 18 False
13 0 7 6 Misleading
25 24 1 0 True
66 24 18 24 Total
Note. Cells show the number of claims. Direction is judged against the fixed real-risk benchmarks in the Coding Scheme. Internal-consistency check only: because a claim's accuracy verdict and its direction relative to real risk are linked by construction, a test on this table confirms that the two codes agree rather than revealing an independent association (chi-square = 63.86, df = 4, p < 0.001, Cramer's V = 0.70; minimum expected cell count 3.55; because one expected count fell below five, the same result was confirmed with the Fisher-Freeman-Halton exact test, p < 0.001). This is not a hypothesis test. The substantive test is the inflate-versus-downplay split within the 41 inaccurate claims (24 downplay, 17 inflate; exact binomial against an even split, p = 0.35).
Table 3. Rhetorical themes among the 41 misinformation claims.
Table 3. Rhetorical themes among the 41 misinformation claims.
Percent n Theme
37 15 Blame a country or group
20 8 Exaggerated severity
15 6 Made-up outbreak
12 5 Fake cure or prevention
10 4 Cover-up claim
7 3 Other
0 0 Travel panic
100 41 Total
Note. Themes are an author-developed coding scheme informed by prior categorizations of health and disease misinformation [33,34,35]; each claim received one theme. Counts are the primary figures; percentages are approximate shares of the 41 inaccurate claims and may not sum to 100 because of rounding. These counts describe the curated reading set and are not prevalence estimates for Reddit. Travel panic did not occur.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.