Preprint
Article

This version is not peer-reviewed.

Cross-Benchmark Social Bot Detection via Multimodal Graph Pretraining Under Distribution Shift

Submitted:

03 October 2026

Posted:

08 October 2026

You are already at the latest version

Abstract
Social bot detectors can perform well on a single benchmark but degrade when bot populations, feature availability, annotation practices, or graph structure change. This study evaluates multi-benchmark multimodal graph pretraining for cross-benchmark social bot detection under distribution shift. A shared encoder integrates profile, textual, behavioral, temporal, and relational signals and is assessed on TwiBot-20, Twi-Bot-22, Cresci-2015, and Multi-Relational Graph-Based Twitter Account Detection Benchmark (MGTAB) using within-benchmark classification, direct source-to-target transfer, leave-one-benchmark-out pretraining, and sequential adaptation. Mean within-benchmark Macro-F1 was 94.13%, whereas direct transfer produced generalization gaps of 2.78–16.13 percentage points. Pretraining on three benchmarks improved mean held-out F1 from 90.54% when training from scratch to 93.80% after fine-tuning. After sequential exposure to all four benchmarks, mean F1 remained 92.89%, with retrospective losses of 0.66–2.11 points on previously seen tasks. These results show that multi-benchmark pretraining improves target adaptation but does not eliminate distribution shift. The study separates in-domain accuracy, transferability, and retention, providing a focused evaluation framework for robust social bot detection.
Keywords: 
;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.